Beginner level8 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how a small student model learns from a big teacher, tell distillation apart from quantisation, and judge when a small model fits a task.

The short answerA small model can learn by copying a big model’s answers, like a student learning from a teacher; this is called distillation. It makes models small and cheap enough for phones and fast apps.

In simple words

  • The big model is the teacher; the small one is the student.
  • The student learns from the teacher’s answers, including how sure the teacher was about each option.
  • Many small models, such as the first Gemini Nano for Android phones, learned this way.
  • Students are usually a little weaker than their teachers.
A small model learning from a big oneIllustration + live data

1 What the student sees

She ordered a cup of

  1. coffee46%
  2. tea38%
  3. water8%
  4. milk5%
  5. all other words3%

The teacher also says “tea” was nearly as likely, and “milk” possible. The student learns how close the words are, from the same sentence.

2 Published teacher and student pairs

  1. Gemini 1.5 ProGemini 1.5 FlashGoogle: Flash was “trained by 1.5 Pro” · 2024
  2. Llama 3.1 8B and 70BLlama 3.2 1B and 3BCut down from the 8B, then trained on both teachers’ outputs · 2024
  3. DeepSeek-R1DeepSeek-R1-Distill, 1.5B to 70BFine-tuned on about 800,000 samples made with R1 · 2025
  4. Llama 4 BehemothLlama 4 MaverickMeta had not released the teacher · 2025
The percentages are made up to show the idea. The examples are the makers’ own descriptions of how they trained the small models.

Words to know

Knowledge distillation
Training a small “student” model to imitate a big “teacher” model’s answers.
Soft labels
The teacher’s percentages for every possible answer, not just its top pick.
Quantisation
Storing a model’s numbers with fewer bits so it needs less memory; no separate teacher model.

A teacher and a student

Big models are strong, but slow and expensive. Small models are fast and cheap, and some run on a phone. Distillation tries to get the best of both: a small “student” model is trained to imitate a big “teacher” model.

The idea is older than chatbots. Researchers did it in 2006, and in 2015 Geoffrey Hinton and two colleagues gave it the name “distillation”.

More than the right answer

A normal training text only shows which word came next. The teacher gives more: how likely it finds every possible word. These percentages are called soft labels. In the diagram, “coffee” is the word in the text, but the teacher says “tea” was nearly as likely.

Hinton’s team gave an example with pictures. A picture model may give a photo of a BMW a small chance of being a garbage truck, and almost none of being a carrot. Near misses like these teach the student that cars look more like trucks than vegetables.

Where you meet distilled models

Makers use distillation on their own models. Google says Gemini 1.5 Flash was “trained by 1.5 Pro”. The first Gemini Nano models, made for Android phones, were distilled from bigger Gemini models.

A phone model can work offline and keep your words on the device. But a student is usually somewhat weaker than its teacher. Small models fit narrow, frequent or private tasks; for hard questions, keep a big model, and test on your own cases.

Learning from another company’s model can break that company’s rules. Anthropic’s and Google’s terms, for example, forbid using their models to build competing ones.

Another way to shrink a model

A second way to make a model smaller is quantisation. The same model keeps all its numbers but stores each one less exactly, in fewer bits. It is like saving a photo at lower quality, and no second model is trained.

Makers often use both: the first Gemini Nano models were distilled and also stored at 4 bits for phones.

Try it yourself

Ask a chatbot: “A small white animal with long ears and a fluffy tail. Give a percentage for rabbit, hare, cat and dog.” The spread is what soft labels look like. The numbers it writes only show the idea; they are not its real inner probabilities.

Check yourself

  1. What does a student model learn from in distillation?

  2. Why do makers distil their big models into small ones?

  3. A phone app stores the same model’s numbers with fewer bits so that it fits. What is this called?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.