Intermediate level9 min readTo learn at another level, choose it before you start the course.

After this lesson you canName the main parts of a language model, explain what attention and the context window do, and read the size figures that makers publish.

The short answerA language model is a stack of identical transformer blocks: in each, attention mixes in information from earlier tokens, and a feed-forward layer processes each token. At the top, the model scores every token in its vocabulary as the next one.

In simple words

  • Tokens are often pieces of words, and each one becomes a vector, a list of numbers called its embedding.
  • Each block has attention, where tokens share information, and a feed-forward layer, where each token is processed alone.
  • Llama 3 has 32 blocks at 8 billion parameters and 126 at 405 billion.
  • A chat assistant is the same kind of network after post-training, not a different design.
Inside a large language modelIllustration

1 What one attention head looks at

When the model is at “it”, this head looks most at “soup” (58%). Words to the right are hidden: a model that writes left to right cannot look ahead.

2 The same block, many times

  1. Tokens become numbersEach token turns into a long list of numbers, its embedding.
  2. Transformer blockAttention: each token gathers what it needs from earlier tokens.Feed-forward: each token is worked on alone, using what the model learned.repeated, block after block
  3. Next-token scoresA probability for every token the model knows.
Llama 3, by size
ModelBlocksNumbers per token
8B324,096
70B808,192
405B12616,384
The attention shares are made up to show the idea; a real model has many attention heads in every layer, and each looks for different links. Layer counts and widths are Meta’s published figures for Llama 3.

Words to know

Embedding
The list of numbers that stands for a token inside the model.
Attention
The step where each token takes in information from earlier tokens.
Context window
Everything the model can see at once: your text, files and its answer.

From text to vectors

Most models cut text into subword tokens, a method made popular by machine translation research in 2015. Common words stay whole and rare words are split into pieces, so the model can handle words it never saw.

Llama 3’s vocabulary has 128,000 tokens. Each token becomes an embedding, a vector of 4,096 numbers in the 8-billion-parameter Llama 3. The original transformer also added a signal for each position, because attention on its own has no sense of order.

The block that repeats

A transformer block has two parts. In attention, each token looks at itself and the earlier tokens and takes in what is relevant. In the feed-forward layer, each token is processed on its own by a larger network. In Llama 3, most of the parameters sit in these feed-forward layers.

The output of one block is the input of the next. Llama 3’s 8-billion-parameter model stacks 32 blocks; its 405-billion model stacks 126, with vectors of 16,384 numbers.

In models that write text, a mask stops each token from seeing later tokens. The model must predict the next token without seeing it.

Parameters, computing and size

The parameters are all the numbers in these blocks that training sets. Running a model costs about two operations per parameter for every token, plus extra for long contexts. So a dense model, which uses every parameter for every token, costs about twice as much to run when it is twice as big.

Bigger is not automatically better. In 2022, people preferred answers from a 1.3-billion-parameter InstructGPT model to those of the 175-billion-parameter GPT-3, because the small one had been trained to follow instructions.

The context window and its cost

The context window holds the prompt, any files and the answer. Attention compares tokens with each other, so its time and memory grow with the square of the context length. Double the text, and that part of the work grows four times.

Models also use long windows unevenly. A 2023 study found that they do best with key facts at the start or end of a long input, and worse with facts in the middle.

Try it yourself

Take a long public text, not a work file, and hide two made-up facts in its middle. Ask a chat app a question that needs both. Then move them to the start, open a new chat and ask again. Did the answer change?

Check yourself

  1. What happens inside each transformer block?

  2. Llama 3 405B has 126 blocks; Llama 3 8B has 32. What mostly makes the big one costlier to run?

  3. You paste a 200-page report and ask about a detail on page 100. What does research suggest?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.