Learn · Fourth course
How AI models are built
The designs behind today’s models: transformers, experts, reasoning, small models, new architectures, Jev-style decision models and world models.
2 Start the course
The lessons
What is inside a large language model?
A large language model is one giant network of numbers. Your text goes in as tokens, passes through the same kind of block many times, and comes out as a guess for the next token.
Why do some huge models run cheaply?
Some huge models are split into many small parts called experts. For each piece of text, a router switches on only a few of them, so the model does far less work than its size suggests.
What do “reasoning” models do differently?
A reasoning model writes out hidden working steps before it gives its answer. That takes more time and costs more, but it helps most with problems whose answer can be checked, such as maths or code.
How do small models learn from big ones?
A small model can learn by copying a big model’s answers, like a student learning from a teacher; this is called distillation. It makes models small and cheap enough for phones and fast apps.
Are there alternatives to the transformer?
Yes: some newer designs read text the way a person reads a book, keeping a short running summary instead of looking back at every word. They are faster on very long texts but worse at finding one exact detail, so many new models mix both designs.
Can a model write a whole answer at once?
Almost: a diffusion language model starts with an answer made of blanks and fills in many words at each step. It can be much faster than writing word by word, but so far its answers are usually weaker.
What is a decision model like Jev?
A decision model does not write text: you give it a question and the possible answers, and it gives each answer a probability. That makes it fast and cheap for repeated choices, such as sorting messages.
What are world models and JEPA?
A world model is an AI’s inner simulator: it predicts what will happen next, for example after a robot moves its arm. JEPA is one way to build it, which predicts the meaning of what comes next instead of every pixel.
After this course you can
- Name the main parts of a language model, explain what attention and the context window do, and read the size figures that makers publish.
- Explain how a mixture-of-experts model uses only part of itself for each token, and judge what “total” and “active” parameters mean for cost and memory.
- Explain how reasoning models spend extra computing while they answer, choose when a higher effort setting is worth its cost, and say why the visible thinking is not proof.
- Explain how a small student model learns from a big teacher, tell distillation apart from quantisation, and judge when a small model fits a task.
- Explain why attention gets expensive on long inputs, describe how state space and recurrent designs keep a fixed-size memory, and say why most new designs are hybrids.
- Explain how diffusion language models fill in many tokens in parallel, and weigh their speed claims against outside tests and their quality trade-offs.
- Explain how a decision model differs from a model that writes text, read its probabilities with care, and decide when it may act without a person.
- Explain what a world model predicts, tell JEPA apart from models that generate video, and judge how far world models are from everyday robots.