Beginner level9 min readTo learn at another level, choose it before you start the course.

After this lesson you canName the main parts of a language model, explain what attention and the context window do, and read the size figures that makers publish.

The short answerA large language model is one giant network of numbers. Your text goes in as tokens, passes through the same kind of block many times, and comes out as a guess for the next token.

In simple words

  • Text is cut into tokens, small pieces of words, and each token becomes a list of numbers.
  • Inside, the same block repeats many times, like the floors of a tall building.
  • In each block, every word looks at the earlier words to work out what it means here.
  • The context window is the model’s working memory for one chat, not everything it knows.
Inside a large language modelIllustration

1 What one attention head looks at

When the model is at “it”, this head looks most at “soup” (58%). Words to the right are hidden: a model that writes left to right cannot look ahead.

2 The same block, many times

  1. Tokens become numbersEach token turns into a long list of numbers, its embedding.
  2. Transformer blockAttention: each token gathers what it needs from earlier tokens.Feed-forward: each token is worked on alone, using what the model learned.repeated, block after block
  3. Next-token scoresA probability for every token the model knows.
Llama 3, by size
ModelBlocksNumbers per token
8B324,096
70B808,192
405B12616,384
The attention shares are made up to show the idea; a real model has many attention heads in every layer, and each looks for different links. Layer counts and widths are Meta’s published figures for Llama 3.

Words to know

Embedding
The list of numbers that stands for a token inside the model.
Attention
The step where each token takes in information from earlier tokens.
Context window
Everything the model can see at once: your text, files and its answer.

Text in, numbers inside

A model cannot read letters. It first cuts your text into tokens: short words, or pieces of longer words. A token is usually shorter than a word: Anthropic says 100,000 tokens are about 75,000 English words.

Each token then becomes a long list of numbers, called its embedding. In Meta’s Llama 3 8B, each embedding has 4,096 numbers. The model does all its work on these numbers.

A tall building of the same floors

Almost every large language model today is a transformer, a design that Google researchers published in 2017. It is built from one kind of block, repeated many times. Llama 3 8B has 32 blocks, and the largest Llama 3 model has 126.

Each block does two things. First, every word (strictly, every token) looks back at the earlier words, to see which ones matter for its meaning. This is called attention. Then each word is worked on alone, using what the model learned in training.

In “The chef tasted the soup because it was too salty”, attention helps the model link “it” to “soup”, not to “chef”. You can try this in the diagram above.

Working memory, not a library

The context window is everything the model can see at once: your messages, any files, and its own answer. Claude’s newest models can take a million tokens, about 750,000 words.

The model does not look things up in a library of the texts it trained on. What it learned is stored in its numbers, called parameters, and it writes what is likely. Nothing inside the model checks the facts, so it can be wrong.

A long window is not read equally well, either. Researchers found that models use facts at the start or the end of a long text better than facts in the middle.

Model names often show the number of parameters: Llama 3 8B has 8 billion, and the largest Llama 3 has 405 billion. In a model like Llama, more parameters mean more computing for every token; lesson 2 shows an exception.

Try it yourself

Take a long public text, not a work file, and hide two made-up facts in its middle. Ask a chat app a question that needs both. Then move them to the start, open a new chat and ask again. Did the answer change?

Check yourself

  1. A news story says a model has a context window of one million tokens. What does that mean?

  2. In “The chef tasted the soup because it was too salty”, what does attention help the model do?

  3. A model is called “Llama 3 70B”. What does “70B” tell you?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.