From raw data to a released modelIllustration + live data
  1. 1DataWeb pages, books, code, licensed archives
  2. 2Pre-trainingGuess the next token, trillions of times, for weeks on thousands of chipsloss falls
  3. 3Fine-tuningExamples of good answers and people’s ratings
  4. 4ReleaseTested and shipped through an app or API

Training compute, in operations · compute ≈ 6 × parameters × training tokens

  • EU, 1025 Above this, a general-purpose model is presumed to have systemic risk.In force
  • California, Connecticut, New York, 1026 Above this, a model counts as frontier.CA in forceCT in forceNY from Jan 2027
The loss curve shows the idea, not a real run. The two lines on the scale come from the laws in our AI rules checker; each step on the scale is ten times more computing.

The short answerDevelopers show a model huge amounts of data and let it adjust itself until its predictions improve. Then they fine-tune it with examples and human feedback, so it follows instructions.

In simple words

  • Pre-training: the model learns from huge amounts of text, images or code.
  • Fine-tuning: people teach it to follow instructions and refuse harmful requests.
  • Training takes thousands of chips, so laws measure it in computing operations.
  • Where the training data comes from is a growing legal fight.

Step 1: pre-training

In pre-training, the model reads a vast collection of text, code or images. It keeps guessing the next or missing piece, checks the real one, and nudges its parameters to do better next time.

This takes weeks or months on thousands of chips, and it is the most expensive step. The amount of computing is counted in operations, and Epoch AI tracks it for well-known models.

Step 2: fine-tuning

A pre-trained model knows a lot but is not yet a helpful assistant. Developers train it further on examples of good answers, and on people’s ratings of its answers.

This step teaches it to follow instructions, to admit uncertainty and to refuse harmful requests. It mostly shapes behaviour rather than adding knowledge.

Where the data comes from

Training data includes public web pages, books, code, licensed archives and data the developer creates. Publishers and artists have sued over the use of their work, and some now sell licences instead.

Training a model on another model’s answers is called distillation. AI labs accuse each other of it, and it is one of the stories we follow.

What the law asks

In the EU, makers of general-purpose models must publish a summary of their training content and respect publishers’ opt-outs. California asks developers to post a summary of their training data.

Laws also use training compute to spot the most powerful models: more than 10²⁵ operations in the EU, and more than 10²⁶ in California and New York.

Check yourself

  1. What happens in pre-training?

  2. What is distillation?

  3. What must makers of general-purpose models publish in the EU?

0 of 3 answered

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.