Intermediate level9 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how diffusion language models fill in many tokens in parallel, and weigh their speed claims against outside tests and their quality trade-offs.

The short answerMasked diffusion language models start from a fully masked answer and unmask the most confident tokens at each step, instead of writing one token at a time. They can be faster per user, but quality usually trails comparable autoregressive models.

In simple words

  • Diffusion came from images: add noise, then learn to remove it step by step.
  • For text, tokens are hidden behind a mask and filled back in; LLaDA 8B was trained this way from scratch.
  • Fewer steps mean more speed but lower quality, because tokens fixed together can clash.
  • Claims and tests differ: Inception claims 1,107 tokens a second for Mercury 2.5; Artificial Analysis measured 599.9 through its API.
Writing left to right, or filling in all at onceIllustration

1 The same sentence, two ways

Diffusion · step 0 of 4
Left to right · same 0 steps

After step 0, the diffusion model has filled in 0 of 12 words; a left-to-right model has written 0.

Each step fills in the words the model is surest of.
An illustration. Real diffusion language models choose how many words to fill in per step and use more steps: DiffusionGemma uses up to 48 for a block of 256 tokens. Each step reworks every word still being filled in, so it costs more than one left-to-right step.

Words to know

Autoregressive model
A model that writes one token after another, from left to right.
Masked diffusion
Starting from a fully hidden answer and revealing tokens over several steps.
Parallel decoding
Fixing several tokens in one step instead of one at a time.

From noise to masks

Diffusion was proposed in 2015: destroy data slowly with noise, then train a model to reverse the damage step by step. DDPM made it work well for images in 2020, and latent diffusion made it cheaper by working on a compressed image.

Text is made of separate tokens, not smooth pixel values. Austin and others (2021) proposed replacing tokens with a mask token instead of adding noise. That links diffusion to masked language models such as BERT, an early Google model that fills in hidden words.

How LLaDA writes

LLaDA 8B (2025) was trained from scratch on 2.3 trillion tokens. It starts from a fully masked answer and uses a transformer to predict the masked tokens, filling them in over several steps.

Its answer length and number of steps are set in advance; fewer steps are faster but lower in quality. On one A100 chip, it was 1.5 and 1.8 times faster than LLaMA3 8B on two maths tests at similar scores, but lagged on a coding test.

Block diffusion (2025) writes block by block and refines the tokens inside each block. Answers can then be any length, and the key-value cache of normal models works again.

Speed claims and outside tests

Inception’s Mercury models are diffusion models with a transformer inside. Inception claims 1,107 tokens a second for Mercury 2.5. As of 8 October 2026, Artificial Analysis measured 599.9 through Inception’s API, second fastest of 182 models, in a different setup.

Google gives 1,479 tokens a second for its experimental Gemini Diffusion, not counting 0.84 seconds of overhead. It says DiffusionGemma’s speed-up is strongest when one chip serves few requests at once, and may shrink in high-volume cloud serving.

The quality trade-off

Fixing many tokens in one step treats them as independent, so words that depend on each other can clash. An outside benchmark, ParallelBench, found “dramatic quality degradation” under parallel decoding on tasks that are easy for people.

Makers’ own tables show the gap. Gemini Diffusion scored 40.4% on GPQA Diamond, against 56.5% for Gemini 2.0 Flash-Lite. DiffusionGemma scored 69.1% on AIME 2026, against 88.3% for Gemma 4.

Try it yourself

If you can, try a diffusion chat model, such as Inception’s Mercury, next to your usual chat app. Ask both for the same 300-word product FAQ, time them and compare the quality. Try twice, because server load changes speed.

Check yourself

  1. In LLaDA, what happens if you cut the number of steps to get answers faster?

  2. Why can fixing many tokens in one step produce clashing words?

  3. Google says DiffusionGemma’s speed-up is strongest when one chip serves few requests at once. What does that mean for a big cloud service?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.