Intermediate level9 min readTo learn at another level, choose it before you start the course.
After this lesson you canExplain how a small student model learns from a big teacher, tell distillation apart from quantisation, and judge when a small model fits a task.
The short answerA small model can learn by copying a big model’s answers, like a student learning from a teacher; this is called distillation. It makes models small and cheap enough for phones and fast apps.
In simple words
- The big model is the teacher; the small one is the student.
- The student learns from the teacher’s answers, including how sure the teacher was about each option.
- Many small models, such as the first Gemini Nano for Android phones, learned this way.
- Students are usually a little weaker than their teachers.
The short answerIn distillation, a small student model is trained to match a big teacher model’s outputs: either its probabilities for every next token, or answers the teacher wrote. Students are cheaper to run but usually trail their teachers, and distillation is not the same as quantisation.
In simple words
- Soft labels, the teacher’s probabilities for every option, carry more information than the one right answer.
- Meta trained Llama 3.2 1B and 3B on the outputs of Llama 3.1 8B and 70B.
- DeepSeek fine-tuned six smaller models on about 800,000 answers made with R1.
- Quantisation shrinks the same model by storing its numbers in fewer bits; it needs no separate teacher model.
The short answerKnowledge distillation trains a student on a teacher’s outputs: temperature-softened probabilities, or sequences the teacher generated, including reasoning traces. Makers use it for their own small models, while distilling from another company’s model can break that company’s terms of service.
Key points
- Hinton and others (2015): a temperature softens the teacher’s probabilities so near misses show up.
- Sequence-level distillation: s1 fine-tuned Qwen2.5-32B on 1,000 Gemini traces in 26 minutes on 16 H100 chips.
- The first Gemini Nano models, Gemma 2 and 3, Llama 3.2 and Llama 4 Maverick were distilled from larger models.
- Anthropic’s and Google’s terms bar building competing models; copyright is a separate question.
1 What the student sees
She ordered a cup of
- coffee46%
- tea38%
- water8%
- milk5%
- all other words3%
The teacher also says “tea” was nearly as likely, and “milk” possible. The student learns how close the words are, from the same sentence.
2 Published teacher and student pairs
- Gemini 1.5 ProGemini 1.5 FlashGoogle: Flash was “trained by 1.5 Pro” · 2024
- Llama 3.1 8B and 70BLlama 3.2 1B and 3BCut down from the 8B, then trained on both teachers’ outputs · 2024
- DeepSeek-R1DeepSeek-R1-Distill, 1.5B to 70BFine-tuned on about 800,000 samples made with R1 · 2025
- Llama 4 BehemothLlama 4 MaverickMeta had not released the teacher · 2025
Words to know
- Knowledge distillation
- Training a small “student” model to imitate a big “teacher” model’s answers.
- Soft labels
- The teacher’s percentages for every possible answer, not just its top pick.
- Quantisation
- Storing a model’s numbers with fewer bits so it needs less memory; no separate teacher model.
A teacher and a student
Big models are strong, but slow and expensive. Small models are fast and cheap, and some run on a phone. Distillation tries to get the best of both: a small “student” model is trained to imitate a big “teacher” model.
The idea is older than chatbots. Researchers did it in 2006, and in 2015 Geoffrey Hinton and two colleagues gave it the name “distillation”.
More than the right answer
A normal training text only shows which word came next. The teacher gives more: how likely it finds every possible word. These percentages are called soft labels. In the diagram, “coffee” is the word in the text, but the teacher says “tea” was nearly as likely.
Hinton’s team gave an example with pictures. A picture model may give a photo of a BMW a small chance of being a garbage truck, and almost none of being a carrot. Near misses like these teach the student that cars look more like trucks than vegetables.
Where you meet distilled models
Makers use distillation on their own models. Google says Gemini 1.5 Flash was “trained by 1.5 Pro”. The first Gemini Nano models, made for Android phones, were distilled from bigger Gemini models.
A phone model can work offline and keep your words on the device. But a student is usually somewhat weaker than its teacher. Small models fit narrow, frequent or private tasks; for hard questions, keep a big model, and test on your own cases.
Learning from another company’s model can break that company’s rules. Anthropic’s and Google’s terms, for example, forbid using their models to build competing ones.
Another way to shrink a model
A second way to make a model smaller is quantisation. The same model keeps all its numbers but stores each one less exactly, in fewer bits. It is like saving a photo at lower quality, and no second model is trained.
Makers often use both: the first Gemini Nano models were distilled and also stored at 4 bits for phones.
Try it yourself
Ask a chatbot: “A small white animal with long ears and a fluffy tail. Give a percentage for rabbit, hare, cat and dog.” The spread is what soft labels look like. The numbers it writes only show the idea; they are not its real inner probabilities.
Check yourself
What does a student model learn from in distillation?
Why do makers distil their big models into small ones?
A phone app stores the same model’s numbers with fewer bits so that it fits. What is this called?
Words to know
- Knowledge distillation
- Training a small “student” model to imitate a big “teacher” model’s answers.
- Soft labels
- The teacher’s percentages for every possible answer, not just its top pick.
- Quantisation
- Storing a model’s numbers with fewer bits so it needs less memory; no separate teacher model.
Soft labels
Hinton, Vinyals and Dean (2015) trained a small model on a big model’s probabilities for every answer, which they called soft targets. These show which wrong answers are near misses, which a plain label cannot.
In their tests, a small network for handwritten digits made 146 errors alone and 74 with soft targets.
Two ways to distil
Makers with the teacher’s probabilities can train on them directly. Google’s Gemma 2 2B and 9B learned this way instead of by plain next-token prediction. Meta cut Llama 3.2 1B and 3B down from the 8B model, then trained them on the 8B and 70B models’ outputs.
Others train on answers the teacher wrote. DeepSeek fine-tuned six Qwen2.5 and Llama models, from 1.5 to 70 billion parameters, on about 800,000 samples made with R1. R1’s MIT licence explicitly allows this.
How close do students get?
DistilBERT, in 2019, was 40% smaller and 60% faster than BERT, an early Google language model, and kept 97% of its language understanding in the authors’ tests. On the AIME 2024 maths contest, DeepSeek’s distilled 32-billion model scored 72.6%, against 79.8% for R1 itself.
Learning from the teacher can beat training the student alone: the same 32-billion base scored 47.0% when DeepSeek trained it with reinforcement learning directly.
Small models fit narrow, high-volume, fast or private tasks. Keep a large model for hard reasoning, and test the small one on your own cases.
Distillation, quantisation and the rules
Quantisation is a different trick: the same model stores its numbers with fewer bits. 8-bit weights halve the memory of 16-bit ones.
Using another company’s model as a teacher can breach its terms: Anthropic’s and Google’s forbid building competing models. In February 2026, Anthropic accused DeepSeek, Moonshot AI and MiniMax of over 16 million exchanges through about 24,000 fake accounts.
Anthropic still calls distillation itself a legitimate and widely used method. The problem it describes is using a rival’s model against its terms.
Try it yourself
Ask a chatbot: “A small white animal with long ears and a fluffy tail. Give a percentage for rabbit, hare, cat and dog.” The spread is what soft labels look like. The numbers it writes only show the idea; they are not its real inner probabilities.
Check yourself
Why do soft labels teach a student more than plain labels?
A team stores its model’s numbers in 4 bits instead of 16 to fit a laptop. What has it done?
DeepSeek’s distilled 32-billion model scored 72.6% on AIME 2024, and R1 itself 79.8%. What does this show?
Probabilities and temperature
Buciluă, Caruana and Niculescu-Mizil (2006) compressed a large ensemble into one small network by training it on the ensemble’s predictions. Hinton, Vinyals and Dean (2015) generalised this as distillation, with a temperature that softens the teacher’s probabilities during training.
Soft targets carry a lot of information. With every “3” removed from its training set, their distilled digit model still recognised 98.6% of test 3s after one bias fix. A distilled speech model nearly matched a 10-model ensemble, at 60.8% against 61.1% frame accuracy.
Sequence-level and reasoning distillation
Without the teacher’s probabilities, a student can be fine-tuned on text the teacher wrote. Distilling step-by-step (2023) also used the teacher’s written reasons: a 770-million T5 beat few-shot 540-billion PaLM on one benchmark, using 80% of the data.
DeepSeek’s R1-Distill models, from 1.5 to 70 billion parameters, were fine-tuned on about 800,000 R1 samples without reinforcement learning. The s1 study fine-tuned Qwen2.5-32B-Instruct on 1,000 traces from Gemini 2.0 Flash Thinking in 26 minutes on 16 H100 chips.
Students usually trail their teachers, as R1’s distilled models do, so test a small model on your own cases before it replaces a large one.
Makers distilling their own models
The first Gemini Nano models (2023), at 1.8 and 3.25 billion parameters, were distilled from larger Gemini models and stored at 4 bits for devices. Gemma 2 and Gemma 3 use distillation in pre-training, and Qwen3’s small models draw on its flagship models.
Meta pruned Llama 3.2 1B and 3B from Llama 3.1 8B, then trained them on the 8B and 70B models’ outputs. Llama 4 Maverick was distilled from Llama 4 Behemoth, a teacher Meta had not released.
Whether GPT-6 Luna, Claude Haiku 5.5 or Gemini Flash-Lite are distilled is not published.
Not quantisation
Quantisation lowers numerical precision instead of training a new model. LLM.int8() (2022) halved memory with 8-bit weights, with no loss reported up to 175 billion parameters; GPTQ ran a 175-billion model on one GPU at 3 to 4 bits.
The two combine: the first Gemini Nano models were both distilled and stored at 4 bits. Apple trained its on-device model of about 3 billion parameters to work at 2 bits per weight.
Terms, copyright and accusations
Anthropic’s commercial terms bar using its services to train competing AI models, and Google’s Gemini API terms bar building competing models or extracting weights. These are contracts with the user.
Copyright is a separate question: the US Copyright Office says AI output is protected only where a human determined enough of its expression. Even unprotected output can still be covered by the terms under which it was obtained.
In February 2026, Anthropic accused DeepSeek, Moonshot AI and MiniMax of distilling through about 24,000 fake accounts. In September 2026, as our news reported, US agencies accused six China-based firms of “malicious distillation”, and China rejected the claims. These are accusations, not court findings.
Try it yourself
Ask a chatbot: “A small white animal with long ears and a fluffy tail. Give a percentage for rabbit, hare, cat and dog.” The spread is what soft labels look like. The numbers it writes only show the idea; they are not its real inner probabilities.
Check yourself
Why did Hinton and others raise the temperature when distilling?
A lab can reach a teacher only through an API that returns text, not probabilities. Which distillation can it do?
In a given case, AI output is not protected by copyright. Can a company then freely distil a provider’s model through its API?
Sources
- Model Compression (Buciluă, Caruana and Niculescu-Mizil, 2006, KDD)
- Distilling the Knowledge in a Neural Network (Hinton, Vinyals and Dean, 2015, arXiv)
- DistilBERT, a distilled version of BERT (Sanh and others, 2019, arXiv)
- Distilling Step-by-Step! (Hsieh and others, 2023, arXiv)
- DeepSeek-R1, including the R1-Distill models (DeepSeek, 2025, arXiv)
- DeepSeek-R1 model card and licence (DeepSeek, Hugging Face)
- s1: Simple test-time scaling (Muennighoff and others, 2025, arXiv)
- Gemini 1.5 Flash, “trained by 1.5 Pro” (Google, May 2024)
- Gemini: A Family of Highly Capable Multimodal Models (Google, 2023, arXiv)
- Gemma 2: Improving Open Language Models at a Practical Size (Google DeepMind, 2024, arXiv)
- Gemma 3 Technical Report (Google DeepMind, 2025, arXiv)
- Llama 3.2: on-device models made with pruning and distillation (Meta, 2024)
- The Llama 4 herd (Meta, 2025)
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (Dettmers and others, 2022, arXiv)
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (Frantar and others, 2022, arXiv)
- Apple Intelligence Foundation Language Models Tech Report 2025 (Apple)
- Gemini Nano on Android (Android Developers)
- Commercial Terms of Service (Anthropic)
- Gemini API Additional Terms of Service (Google)
- Copyright and Artificial Intelligence, Part 2: Copyrightability (US Copyright Office, January 2025)
- Detecting and preventing distillation attacks (Anthropic, February 2026)
- US agencies say Chinese AI firms copied top models at scale (Silicon AI News)
- Qwen3 Technical Report: small models that draw on the flagship models (Qwen, 2025, arXiv)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.