Intermediate level9 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how a small student model learns from a big teacher, tell distillation apart from quantisation, and judge when a small model fits a task.

The short answerIn distillation, a small student model is trained to match a big teacher model’s outputs: either its probabilities for every next token, or answers the teacher wrote. Students are cheaper to run but usually trail their teachers, and distillation is not the same as quantisation.

In simple words

  • Soft labels, the teacher’s probabilities for every option, carry more information than the one right answer.
  • Meta trained Llama 3.2 1B and 3B on the outputs of Llama 3.1 8B and 70B.
  • DeepSeek fine-tuned six smaller models on about 800,000 answers made with R1.
  • Quantisation shrinks the same model by storing its numbers in fewer bits; it needs no separate teacher model.
A small model learning from a big oneIllustration + live data

1 What the student sees

She ordered a cup of

  1. coffee46%
  2. tea38%
  3. water8%
  4. milk5%
  5. all other words3%

The teacher also says “tea” was nearly as likely, and “milk” possible. The student learns how close the words are, from the same sentence.

2 Published teacher and student pairs

  1. Gemini 1.5 ProGemini 1.5 FlashGoogle: Flash was “trained by 1.5 Pro” · 2024
  2. Llama 3.1 8B and 70BLlama 3.2 1B and 3BCut down from the 8B, then trained on both teachers’ outputs · 2024
  3. DeepSeek-R1DeepSeek-R1-Distill, 1.5B to 70BFine-tuned on about 800,000 samples made with R1 · 2025
  4. Llama 4 BehemothLlama 4 MaverickMeta had not released the teacher · 2025
The percentages are made up to show the idea. The examples are the makers’ own descriptions of how they trained the small models.

Words to know

Knowledge distillation
Training a small “student” model to imitate a big “teacher” model’s answers.
Soft labels
The teacher’s percentages for every possible answer, not just its top pick.
Quantisation
Storing a model’s numbers with fewer bits so it needs less memory; no separate teacher model.

Soft labels

Hinton, Vinyals and Dean (2015) trained a small model on a big model’s probabilities for every answer, which they called soft targets. These show which wrong answers are near misses, which a plain label cannot.

In their tests, a small network for handwritten digits made 146 errors alone and 74 with soft targets.

Two ways to distil

Makers with the teacher’s probabilities can train on them directly. Google’s Gemma 2 2B and 9B learned this way instead of by plain next-token prediction. Meta cut Llama 3.2 1B and 3B down from the 8B model, then trained them on the 8B and 70B models’ outputs.

Others train on answers the teacher wrote. DeepSeek fine-tuned six Qwen2.5 and Llama models, from 1.5 to 70 billion parameters, on about 800,000 samples made with R1. R1’s MIT licence explicitly allows this.

How close do students get?

DistilBERT, in 2019, was 40% smaller and 60% faster than BERT, an early Google language model, and kept 97% of its language understanding in the authors’ tests. On the AIME 2024 maths contest, DeepSeek’s distilled 32-billion model scored 72.6%, against 79.8% for R1 itself.

Learning from the teacher can beat training the student alone: the same 32-billion base scored 47.0% when DeepSeek trained it with reinforcement learning directly.

Small models fit narrow, high-volume, fast or private tasks. Keep a large model for hard reasoning, and test the small one on your own cases.

Distillation, quantisation and the rules

Quantisation is a different trick: the same model stores its numbers with fewer bits. 8-bit weights halve the memory of 16-bit ones.

Using another company’s model as a teacher can breach its terms: Anthropic’s and Google’s forbid building competing models. In February 2026, Anthropic accused DeepSeek, Moonshot AI and MiniMax of over 16 million exchanges through about 24,000 fake accounts.

Anthropic still calls distillation itself a legitimate and widely used method. The problem it describes is using a rival’s model against its terms.

Try it yourself

Ask a chatbot: “A small white animal with long ears and a fluffy tail. Give a percentage for rabbit, hare, cat and dog.” The spread is what soft labels look like. The numbers it writes only show the idea; they are not its real inner probabilities.

Check yourself

  1. Why do soft labels teach a student more than plain labels?

  2. A team stores its model’s numbers in 4 bits instead of 16 to fit a laptop. What has it done?

  3. DeepSeek’s distilled 32-billion model scored 72.6% on AIME 2024, and R1 itself 79.8%. What does this show?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.