Intermediate level9 min readTo learn at another level, choose it before you start the course.
After this lesson you canExplain how diffusion language models fill in many tokens in parallel, and weigh their speed claims against outside tests and their quality trade-offs.
The short answerAlmost: a diffusion language model starts with an answer made of blanks and fills in many words at each step. It can be much faster than writing word by word, but so far its answers are usually weaker.
In simple words
- Normal models are autoregressive: they write one token after another, from left to right.
- Diffusion models start with hidden words and reveal many at once, over several steps.
- The idea comes from image generators, which turn noise into a picture step by step.
- They are fast, but makers’ own tests show lower quality than similar normal models.
The short answerMasked diffusion language models start from a fully masked answer and unmask the most confident tokens at each step, instead of writing one token at a time. They can be faster per user, but quality usually trails comparable autoregressive models.
In simple words
- Diffusion came from images: add noise, then learn to remove it step by step.
- For text, tokens are hidden behind a mask and filled back in; LLaDA 8B was trained this way from scratch.
- Fewer steps mean more speed but lower quality, because tokens fixed together can clash.
- Claims and tests differ: Inception claims 1,107 tokens a second for Mercury 2.5; Artificial Analysis measured 599.9 through its API.
The short answerDiscrete diffusion language models learn to reverse a corruption process, usually masking, and decode many tokens in parallel per step. They trade the token-by-token cost of autoregression for a fixed number of denoising steps, with open problems in variable length, caching, and dependencies between tokens decoded together.
Key points
- D3PM (2021) defined discrete corruption; SEDD (2023) beat GPT-2 on perplexity and matched its plain sampling with 32 times fewer passes.
- MDLM (2024) and LLaDA (2025) scaled simple masked diffusion; LLaDA 8B was trained from scratch.
- Block diffusion restores variable length and key-value caching; parallel decoding breaks dependencies.
- Measured speed lags the launch claims: 1,107 tokens a second claimed for Mercury 2.5, 599.9 measured.
1 The same sentence, two ways
After step 0, the diffusion model has filled in 0 of 12 words; a left-to-right model has written 0.
Words to know
- Autoregressive model
- A model that writes one token after another, from left to right.
- Masked diffusion
- Starting from a fully hidden answer and revealing tokens over several steps.
- Parallel decoding
- Fixing several tokens in one step instead of one at a time.
From pictures to words
Many image generators use diffusion: they start from random noise and clean it up, step by step, until a picture appears. Researchers made this work well for images around 2020.
Words are not like pixels, so text diffusion works a little differently. The model starts with an answer where every word is hidden behind a blank; this method is called masked diffusion. At each step, it fills in the blanks it is surest about, until none are left.
Faster, but weaker
Filling in many words per step is called parallel decoding, and it lets a diffusion model answer very quickly. The company Inception says its Mercury 2.5 model writes over 1,100 tokens a second.
An outside tester, Artificial Analysis, measured about 600 tokens a second through the company’s service, as of 8 October 2026. That is about half the claim. The two were measured differently, so this does not prove the claim false. It is still among the fastest speeds Artificial Analysis lists.
Quality is the weak spot. Google’s own tables show its diffusion models scoring below similar normal Gemini and Gemma models on most tests.
What they are good for
Speed matters most in tools such as code editors, where you wait for every suggestion. Inception’s first models were built for code, and Google aims its open DiffusionGemma model at editing text and filling gaps in code.
Because the model works on the whole answer, it can also fill a gap in the middle of a text.
Try it yourself
If you can, try a diffusion chat model, such as Inception’s Mercury, next to your usual chat app. Ask both for the same 300-word product FAQ, time them and compare the quality. Try twice, because server load changes speed.
Check yourself
How does a diffusion language model build its answer?
A company says its model writes 1,100 tokens a second. An outside tester measured about 600 through its service. What should you think?
What did makers’ own tests show about diffusion models’ answers?
Words to know
- Autoregressive model
- A model that writes one token after another, from left to right.
- Masked diffusion
- Starting from a fully hidden answer and revealing tokens over several steps.
- Parallel decoding
- Fixing several tokens in one step instead of one at a time.
From noise to masks
Diffusion was proposed in 2015: destroy data slowly with noise, then train a model to reverse the damage step by step. DDPM made it work well for images in 2020, and latent diffusion made it cheaper by working on a compressed image.
Text is made of separate tokens, not smooth pixel values. Austin and others (2021) proposed replacing tokens with a mask token instead of adding noise. That links diffusion to masked language models such as BERT, an early Google model that fills in hidden words.
How LLaDA writes
LLaDA 8B (2025) was trained from scratch on 2.3 trillion tokens. It starts from a fully masked answer and uses a transformer to predict the masked tokens, filling them in over several steps.
Its answer length and number of steps are set in advance; fewer steps are faster but lower in quality. On one A100 chip, it was 1.5 and 1.8 times faster than LLaMA3 8B on two maths tests at similar scores, but lagged on a coding test.
Block diffusion (2025) writes block by block and refines the tokens inside each block. Answers can then be any length, and the key-value cache of normal models works again.
Speed claims and outside tests
Inception’s Mercury models are diffusion models with a transformer inside. Inception claims 1,107 tokens a second for Mercury 2.5. As of 8 October 2026, Artificial Analysis measured 599.9 through Inception’s API, second fastest of 182 models, in a different setup.
Google gives 1,479 tokens a second for its experimental Gemini Diffusion, not counting 0.84 seconds of overhead. It says DiffusionGemma’s speed-up is strongest when one chip serves few requests at once, and may shrink in high-volume cloud serving.
The quality trade-off
Fixing many tokens in one step treats them as independent, so words that depend on each other can clash. An outside benchmark, ParallelBench, found “dramatic quality degradation” under parallel decoding on tasks that are easy for people.
Makers’ own tables show the gap. Gemini Diffusion scored 40.4% on GPQA Diamond, against 56.5% for Gemini 2.0 Flash-Lite. DiffusionGemma scored 69.1% on AIME 2026, against 88.3% for Gemma 4.
Try it yourself
If you can, try a diffusion chat model, such as Inception’s Mercury, next to your usual chat app. Ask both for the same 300-word product FAQ, time them and compare the quality. Try twice, because server load changes speed.
Check yourself
In LLaDA, what happens if you cut the number of steps to get answers faster?
Why can fixing many tokens in one step produce clashing words?
Google says DiffusionGemma’s speed-up is strongest when one chip serves few requests at once. What does that mean for a big cloud service?
Discrete diffusion
Sohl-Dickstein and others (2015) framed diffusion as learning to reverse a gradual noising process, and DDPM (2020) made it competitive for images. D3PM (Austin and others, 2021) defined corruption for discrete tokens, including a mask state that links diffusion to masked language modelling.
SEDD (Lou, Meng and Ermon, 2023) beat GPT-2 on perplexity, and matched the quality of GPT-2’s plain sampling with 32 times fewer network passes. MDLM (Sahoo and others, 2024) showed that a simple masked-diffusion recipe comes close to autoregressive models on standard benchmarks.
Scaling and decoding
LLaDA 8B (Nie and others, 2025) was trained from scratch on 2.3 trillion tokens, with a transformer predicting the masked tokens. Length and step count are fixed in advance, and its authors claim a win over GPT-4o at completing poems backwards.
Block diffusion (Arriola and others, 2025) decodes block by block and refines the tokens within each block, which restores variable length and key-value caching. DiffusionGemma refines 256-token blocks in up to 48 steps.
Fast-dLLM (2025) notes that decoding many tokens in one step treats them as independent, so dependent tokens can clash. ParallelBench (2025), an outside benchmark, found “dramatic quality degradation” under parallel decoding on tasks easy for people and for autoregressive models.
Claims and measurements
Inception’s 2025 report cites Artificial Analysis at 1,109 and 737 tokens a second for Mercury Coder Mini and Small on H100 chips. For Mercury 2, now marked deprecated, Inception claims 1,009 tokens a second on Blackwell; Artificial Analysis measured 448.8 through the API.
For Mercury 2.5 the figures are 1,107 claimed and 599.9 measured as of 8 October 2026. Its Artificial Analysis intelligence score is 12, below the median of 13 for comparable models. The setups differ, so a gap is not proof that a claim is false.
Quality and serving economics
Makers’ own tables show the trade-offs. Gemini Diffusion trails Gemini 2.0 Flash-Lite on GPQA Diamond, 40.4% against 56.5%, but leads slightly on LiveCodeBench, 30.9% against 28.5%. DiffusionGemma, with 25.2 billion total and 3.8 billion active parameters, scores below Gemma 4 on all but one listed test.
Google says the speed-up is strongest at small batch sizes on a single chip and may shrink in high-volume serving.
Several open models start from autoregressive weights instead of training from scratch, such as Dream 7B and LLaDA2.0-mini and -flash, at 16 and 100 billion parameters.
Try it yourself
If you can, try a diffusion chat model, such as Inception’s Mercury, next to your usual chat app. Ask both for the same 300-word product FAQ, time them and compare the quality. Try twice, because server load changes speed.
Check yourself
What did D3PM’s mask state connect text diffusion to?
What does block diffusion restore, compared with whole-answer masked diffusion?
Artificial Analysis measured Mercury 2.5 at 599.9 tokens a second; Inception claims 1,107. What is the sound conclusion?
Sources
- Deep Unsupervised Learning using Nonequilibrium Thermodynamics (Sohl-Dickstein and others, 2015, arXiv)
- Denoising Diffusion Probabilistic Models (Ho, Jain and Abbeel, 2020, arXiv)
- High-Resolution Image Synthesis with Latent Diffusion Models (Rombach and others, 2021, arXiv)
- Structured Denoising Diffusion Models in Discrete State-Spaces, D3PM (Austin and others, 2021, arXiv)
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution, SEDD (Lou, Meng and Ermon, 2023, arXiv)
- Simple and Effective Masked Diffusion Language Models, MDLM (Sahoo and others, 2024, arXiv)
- Large Language Diffusion Models, LLaDA (Nie and others, 2025, arXiv)
- Block Diffusion (Arriola and others, 2025, arXiv)
- Fast-dLLM (2025, arXiv)
- ParallelBench (2025, arXiv)
- Mercury: Ultra-Fast Language Models Based on Diffusion (Inception Labs, 2025, arXiv)
- Introducing Mercury 2 (Inception Labs)
- Mercury 2, measured speed and scores (Artificial Analysis)
- Introducing Mercury 2.5 (Inception Labs)
- Mercury 2.5, measured speed and scores (Artificial Analysis)
- Gemini Diffusion (Google DeepMind)
- DiffusionGemma model card (Google)
- DiffusionGemma: faster text generation, and where the speed-up shrinks (Google, June 2026)
- Dream 7B (2025, arXiv)
- LLaDA2.0 (2025, arXiv)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.