Beginner level8 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how diffusion language models fill in many tokens in parallel, and weigh their speed claims against outside tests and their quality trade-offs.

The short answerAlmost: a diffusion language model starts with an answer made of blanks and fills in many words at each step. It can be much faster than writing word by word, but so far its answers are usually weaker.

In simple words

  • Normal models are autoregressive: they write one token after another, from left to right.
  • Diffusion models start with hidden words and reveal many at once, over several steps.
  • The idea comes from image generators, which turn noise into a picture step by step.
  • They are fast, but makers’ own tests show lower quality than similar normal models.
Writing left to right, or filling in all at onceIllustration

1 The same sentence, two ways

Diffusion · step 0 of 4
Left to right · same 0 steps

After step 0, the diffusion model has filled in 0 of 12 words; a left-to-right model has written 0.

Each step fills in the words the model is surest of.
An illustration. Real diffusion language models choose how many words to fill in per step and use more steps: DiffusionGemma uses up to 48 for a block of 256 tokens. Each step reworks every word still being filled in, so it costs more than one left-to-right step.

Words to know

Autoregressive model
A model that writes one token after another, from left to right.
Masked diffusion
Starting from a fully hidden answer and revealing tokens over several steps.
Parallel decoding
Fixing several tokens in one step instead of one at a time.

From pictures to words

Many image generators use diffusion: they start from random noise and clean it up, step by step, until a picture appears. Researchers made this work well for images around 2020.

Words are not like pixels, so text diffusion works a little differently. The model starts with an answer where every word is hidden behind a blank; this method is called masked diffusion. At each step, it fills in the blanks it is surest about, until none are left.

Faster, but weaker

Filling in many words per step is called parallel decoding, and it lets a diffusion model answer very quickly. The company Inception says its Mercury 2.5 model writes over 1,100 tokens a second.

An outside tester, Artificial Analysis, measured about 600 tokens a second through the company’s service, as of 8 October 2026. That is about half the claim. The two were measured differently, so this does not prove the claim false. It is still among the fastest speeds Artificial Analysis lists.

Quality is the weak spot. Google’s own tables show its diffusion models scoring below similar normal Gemini and Gemma models on most tests.

What they are good for

Speed matters most in tools such as code editors, where you wait for every suggestion. Inception’s first models were built for code, and Google aims its open DiffusionGemma model at editing text and filling gaps in code.

Because the model works on the whole answer, it can also fill a gap in the middle of a text.

Try it yourself

If you can, try a diffusion chat model, such as Inception’s Mercury, next to your usual chat app. Ask both for the same 300-word product FAQ, time them and compare the quality. Try twice, because server load changes speed.

Check yourself

  1. How does a diffusion language model build its answer?

  2. A company says its model writes 1,100 tokens a second. An outside tester measured about 600 through its service. What should you think?

  3. What did makers’ own tests show about diffusion models’ answers?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.