- 1DataWeb pages, books, code, licensed archives
- 2Pre-trainingGuess the next token, trillions of times, for weeks on thousands of chips
- 3Fine-tuningExamples of good answers and people’s ratings
- 4ReleaseTested and shipped through an app or API
Training compute, in operations · compute ≈ 6 × parameters × training tokens
- EU, 1025 Above this, a general-purpose model is presumed to have systemic risk.In force
- California, Connecticut, New York, 1026 Above this, a model counts as frontier.CA in forceCT in forceNY from Jan 2027
The short answerA model learns by reading huge amounts of text and practising guessing the next word. Then people teach it to be helpful and safe.
In simple words
- First it practises on huge amounts of text, pictures or code.
- Then people show it good answers and rate its replies.
- Training takes thousands of chips and costs a lot.
- Some writers and publishers say their work was used without permission.
Practice, practice, practice
During training, the model guesses the next word a huge number of times and checks itself against the real text. Each wrong guess nudges it to do a little better.
This first stage takes weeks or months on thousands of powerful chips.
Learning good manners
After that, people show the model examples of good answers and rate its replies. This teaches it to follow instructions and to say no to harmful requests.
Whose data?
The data comes from the web, books, code and other sources. Some newspapers and artists have gone to court over this, and some AI companies now pay for data.
New laws also ask AI makers to say what data they used.
Check yourself
What does a model practise during training?
Why do people rate a model’s answers?
Why are some publishers in court with AI companies?
0 of 3 answered
The short answerDevelopers show a model huge amounts of data and let it adjust itself until its predictions improve. Then they fine-tune it with examples and human feedback, so it follows instructions.
In simple words
- Pre-training: the model learns from huge amounts of text, images or code.
- Fine-tuning: people teach it to follow instructions and refuse harmful requests.
- Training takes thousands of chips, so laws measure it in computing operations.
- Where the training data comes from is a growing legal fight.
Step 1: pre-training
In pre-training, the model reads a vast collection of text, code or images. It keeps guessing the next or missing piece, checks the real one, and nudges its parameters to do better next time.
This takes weeks or months on thousands of chips, and it is the most expensive step. The amount of computing is counted in operations, and Epoch AI tracks it for well-known models.
Step 2: fine-tuning
A pre-trained model knows a lot but is not yet a helpful assistant. Developers train it further on examples of good answers, and on people’s ratings of its answers.
This step teaches it to follow instructions, to admit uncertainty and to refuse harmful requests. It mostly shapes behaviour rather than adding knowledge.
Where the data comes from
Training data includes public web pages, books, code, licensed archives and data the developer creates. Publishers and artists have sued over the use of their work, and some now sell licences instead.
Training a model on another model’s answers is called distillation. AI labs accuse each other of it, and it is one of the stories we follow.
What the law asks
In the EU, makers of general-purpose models must publish a summary of their training content and respect publishers’ opt-outs. California asks developers to post a summary of their training data.
Laws also use training compute to spot the most powerful models: more than 10²⁵ operations in the EU, and more than 10²⁶ in California and New York.
Check yourself
What happens in pre-training?
What is distillation?
What must makers of general-purpose models publish in the EU?
0 of 3 answered
The short answerPre-training fits the parameters to minimise next-token prediction loss by gradient descent; post-training then aligns behaviour with demonstrations, preference learning and reinforcement learning. Compute, roughly six times parameters times tokens, is what laws measure.
In simple words
- Pre-training minimises cross-entropy loss on the next token, using gradient descent.
- Compute for a dense transformer is roughly 6 × parameters × training tokens.
- Post-training adds supervised examples, preference learning and reinforcement learning.
- Laws use compute thresholds: 10²⁵ in the EU, 10²⁶ in California, New York and Connecticut.
The objective
Pre-training minimises cross-entropy loss: the model is penalised by how unlikely it found the true next token. Backpropagation computes how each parameter should change, and an optimiser such as Adam applies small updates.
Repeated over trillions of tokens, this turns random weights into a model of language, code and the world described in its data.
Compute, data and scaling
A widely used rule of thumb puts training compute for a dense transformer at about 6 × parameters × training tokens operations. Scaling-law research (OpenAI, 2020) found loss falls smoothly as compute grows.
DeepMind’s Chinchilla study (2022) found that, for a fixed budget, parameters and data should grow together, at roughly 20 tokens per parameter.
Post-training
Supervised fine-tuning on demonstrations comes first. Preference learning follows: reinforcement learning from human feedback, as in OpenAI’s 2022 InstructGPT paper, or simpler methods such as direct preference optimisation.
Reasoning models are also trained with reinforcement learning on tasks whose answers can be checked automatically, such as maths and code.
Data, distillation and the law
Distillation trains a student model on a teacher model’s outputs. It is routine inside one company, but rivals’ terms often forbid it, which is behind OpenAI’s claims about Moonshot-linked users.
EU copyright law lets publishers opt out of text and data mining (Directive 2019/790, Article 4(3)), and the AI Act asks for a public training-content summary. California’s AB 2013 requires dataset documentation, and Connecticut privacy notices must say whether personal data trains language models.
Thresholds: the EU presumes systemic risk above 10²⁵ training operations (Article 51). California’s SB 53, New York’s RAISE Act and Connecticut use 10²⁶; California’s count includes later fine-tuning and reinforcement learning.
Check yourself
What does pre-training minimise?
Roughly how is training compute estimated for a dense transformer?
What does California count toward its 10²⁶ frontier threshold?
0 of 3 answered
Sources
- AI Act, Article 53 (EU AI Act Service Desk)
- Data on notable AI models and their training compute (Epoch AI)
- Civil Code section 3111 (California Legislature)
- Scaling Laws for Neural Language Models (Kaplan and others, 2020, arXiv)
- Training Compute-Optimal Large Language Models (Hoffmann and others, 2022, arXiv)
- Training language models to follow instructions with human feedback (Ouyang and others, 2022, arXiv)
- SB 53 (California Legislature)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.