Intermediate level9 min readTo learn at another level, choose it before you start the course.
After this lesson you canName the main parts of a language model, explain what attention and the context window do, and read the size figures that makers publish.
The short answerA large language model is one giant network of numbers. Your text goes in as tokens, passes through the same kind of block many times, and comes out as a guess for the next token.
In simple words
- Text is cut into tokens, small pieces of words, and each token becomes a list of numbers.
- Inside, the same block repeats many times, like the floors of a tall building.
- In each block, every word looks at the earlier words to work out what it means here.
- The context window is the model’s working memory for one chat, not everything it knows.
The short answerA language model is a stack of identical transformer blocks: in each, attention mixes in information from earlier tokens, and a feed-forward layer processes each token. At the top, the model scores every token in its vocabulary as the next one.
In simple words
- Tokens are often pieces of words, and each one becomes a vector, a list of numbers called its embedding.
- Each block has attention, where tokens share information, and a feed-forward layer, where each token is processed alone.
- Llama 3 has 32 blocks at 8 billion parameters and 126 at 405 billion.
- A chat assistant is the same kind of network after post-training, not a different design.
The short answerToday’s large language models are mostly decoder-only transformers: a token embedding, a stack of blocks with causal self-attention and a feed-forward network, and an output layer over the vocabulary. Their cost grows with the parameters used per token and, for attention over a whole text, with the square of its length.
Key points
- GPT (2018) used a 12-layer decoder; Meta calls Llama 3 a “relatively standard decoder-only transformer”.
- A forward pass costs about 2N operations per token for N parameters, plus a term for context.
- Attention’s time and memory grow with the square of the sequence length, so makers redesign it for long inputs.
- Closed labs often keep size and design secret; the GPT-4 report withheld both.
1 What one attention head looks at
When the model is at “it”, this head looks most at “soup” (58%). Words to the right are hidden: a model that writes left to right cannot look ahead.
2 The same block, many times
- Tokens become numbersEach token turns into a long list of numbers, its embedding.
- Transformer blockAttention: each token gathers what it needs from earlier tokens.Feed-forward: each token is worked on alone, using what the model learned.repeated, block after block
- Next-token scoresA probability for every token the model knows.
| Model | Blocks | Numbers per token |
|---|---|---|
| 8B | 32 | 4,096 |
| 70B | 80 | 8,192 |
| 405B | 126 | 16,384 |
Words to know
- Embedding
- The list of numbers that stands for a token inside the model.
- Attention
- The step where each token takes in information from earlier tokens.
- Context window
- Everything the model can see at once: your text, files and its answer.
Text in, numbers inside
A model cannot read letters. It first cuts your text into tokens: short words, or pieces of longer words. A token is usually shorter than a word: Anthropic says 100,000 tokens are about 75,000 English words.
Each token then becomes a long list of numbers, called its embedding. In Meta’s Llama 3 8B, each embedding has 4,096 numbers. The model does all its work on these numbers.
A tall building of the same floors
Almost every large language model today is a transformer, a design that Google researchers published in 2017. It is built from one kind of block, repeated many times. Llama 3 8B has 32 blocks, and the largest Llama 3 model has 126.
Each block does two things. First, every word (strictly, every token) looks back at the earlier words, to see which ones matter for its meaning. This is called attention. Then each word is worked on alone, using what the model learned in training.
In “The chef tasted the soup because it was too salty”, attention helps the model link “it” to “soup”, not to “chef”. You can try this in the diagram above.
Working memory, not a library
The context window is everything the model can see at once: your messages, any files, and its own answer. Claude’s newest models can take a million tokens, about 750,000 words.
The model does not look things up in a library of the texts it trained on. What it learned is stored in its numbers, called parameters, and it writes what is likely. Nothing inside the model checks the facts, so it can be wrong.
A long window is not read equally well, either. Researchers found that models use facts at the start or the end of a long text better than facts in the middle.
Model names often show the number of parameters: Llama 3 8B has 8 billion, and the largest Llama 3 has 405 billion. In a model like Llama, more parameters mean more computing for every token; lesson 2 shows an exception.
Try it yourself
Take a long public text, not a work file, and hide two made-up facts in its middle. Ask a chat app a question that needs both. Then move them to the start, open a new chat and ask again. Did the answer change?
Check yourself
A news story says a model has a context window of one million tokens. What does that mean?
In “The chef tasted the soup because it was too salty”, what does attention help the model do?
A model is called “Llama 3 70B”. What does “70B” tell you?
Words to know
- Embedding
- The list of numbers that stands for a token inside the model.
- Attention
- The step where each token takes in information from earlier tokens.
- Context window
- Everything the model can see at once: your text, files and its answer.
From text to vectors
Most models cut text into subword tokens, a method made popular by machine translation research in 2015. Common words stay whole and rare words are split into pieces, so the model can handle words it never saw.
Llama 3’s vocabulary has 128,000 tokens. Each token becomes an embedding, a vector of 4,096 numbers in the 8-billion-parameter Llama 3. The original transformer also added a signal for each position, because attention on its own has no sense of order.
The block that repeats
A transformer block has two parts. In attention, each token looks at itself and the earlier tokens and takes in what is relevant. In the feed-forward layer, each token is processed on its own by a larger network. In Llama 3, most of the parameters sit in these feed-forward layers.
The output of one block is the input of the next. Llama 3’s 8-billion-parameter model stacks 32 blocks; its 405-billion model stacks 126, with vectors of 16,384 numbers.
In models that write text, a mask stops each token from seeing later tokens. The model must predict the next token without seeing it.
Parameters, computing and size
The parameters are all the numbers in these blocks that training sets. Running a model costs about two operations per parameter for every token, plus extra for long contexts. So a dense model, which uses every parameter for every token, costs about twice as much to run when it is twice as big.
Bigger is not automatically better. In 2022, people preferred answers from a 1.3-billion-parameter InstructGPT model to those of the 175-billion-parameter GPT-3, because the small one had been trained to follow instructions.
The context window and its cost
The context window holds the prompt, any files and the answer. Attention compares tokens with each other, so its time and memory grow with the square of the context length. Double the text, and that part of the work grows four times.
Models also use long windows unevenly. A 2023 study found that they do best with key facts at the start or end of a long input, and worse with facts in the middle.
Try it yourself
Take a long public text, not a work file, and hide two made-up facts in its middle. Ask a chat app a question that needs both. Then move them to the start, open a new chat and ask again. Did the answer change?
Check yourself
What happens inside each transformer block?
Llama 3 405B has 126 blocks; Llama 3 8B has 32. What mostly makes the big one costlier to run?
You paste a 200-page report and ask about a detail on page 100. What does research suggest?
Encoder, decoder, decoder-only
The 2017 transformer had an encoder that read the source sentence and a decoder that wrote the translation. It dropped recurrence and convolutions and relied on attention, which let training run in parallel.
OpenAI’s GPT, in 2018, kept only a 12-layer decoder with masked self-attention, pre-trained on unlabelled text and then fine-tuned. GPT-3 followed in 2020 as an autoregressive model with 175 billion parameters. Meta calls Llama 3 a “relatively standard decoder-only transformer”.
Inside a block
Each block applies causal self-attention, with several heads in parallel, then a position-wise feed-forward network, with residual connections around both. Llama 3 8B has 32 layers, a model dimension of 4,096, a feed-forward dimension of 14,336, 32 query heads and 8 key-value heads.
From those figures, each feed-forward layer holds about 176 million parameters and each attention layer about 42 million. Most of a dense model’s weights therefore sit in the feed-forward layers, which mixture-of-experts designs replace (lesson 2).
Meta shares 8 key-value heads among the query heads, called grouped-query attention, to speed up answers and shrink the key-value cache while the model writes.
What scale costs
Kaplan and others (2020) estimate a forward pass at about 2N operations per token for a model with N parameters, plus a term that grows with context length. They also found that loss falls along smooth power laws as parameters, data and compute grow.
Parameter count alone does not decide quality. People preferred the 1.3-billion-parameter InstructGPT to the 175-billion-parameter GPT-3, after post-training on demonstrations and human rankings turned the same kind of network into an assistant.
Long contexts
Self-attention’s time and memory grow with the square of the sequence length, which makes long windows expensive. Makers change the design to cope: Gemma 3 uses more local attention layers, which look only at nearby tokens, to save memory on long inputs.
A window is not used evenly. Liu and others (2023) found results best when the relevant information sits at the start or end of the input, and lowest in the middle. Anthropic’s documentation calls the decline as context grows “context rot”.
What is not published
Open-weights makers publish their configurations, so figures like those above can be checked. Closed labs often do not: the GPT-4 report withheld the model size and architecture, citing competition and safety.
Google says Gemini 3 Pro is a sparse mixture-of-experts transformer, but gives no parameter count. Treat any size figure for a closed model as an outside estimate unless its maker published it.
Try it yourself
Take a long public text, not a work file, and hide two made-up facts in its middle. Ask a chat app a question that needs both. Then move them to the start, open a new chat and ask again. Did the answer change?
Check yourself
Llama 3 8B has a model dimension of 4,096 and a feed-forward dimension of 14,336. Where do most of each block’s parameters sit?
A dense model has 70 billion parameters. Roughly how many operations does one forward pass cost per token, ignoring the context?
A start-up says a closed rival “has 2 trillion parameters”. The rival never published a size. How should you treat the number?
Sources
- Attention Is All You Need (Vaswani and others, 2017, arXiv)
- Neural Machine Translation of Rare Words with Subword Units (Sennrich and others, 2015, arXiv)
- Improving Language Understanding by Generative Pre-Training (Radford and others, 2018, OpenAI)
- Language Models are Few-Shot Learners (Brown and others, 2020, arXiv)
- Scaling Laws for Neural Language Models (Kaplan and others, 2020, arXiv)
- The Llama 3 Herd of Models (Meta, 2024, arXiv)
- Introducing Meta Llama 3 (Meta)
- Training language models to follow instructions with human feedback (Ouyang and others, 2022, arXiv)
- FlashAttention: attention’s cost grows with the square of the length (Dao and others, 2022, arXiv)
- Lost in the Middle: How Language Models Use Long Contexts (Liu and others, 2023, arXiv)
- Gemma 3 Technical Report (Google DeepMind, 2025, arXiv)
- Context windows (Claude documentation)
- Introducing 100K context windows: 100,000 tokens is about 75,000 words (Anthropic, 2023)
- GPT-4 Technical Report (OpenAI, 2023, arXiv)
- Gemini 3 Pro model card (Google DeepMind, 2025)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.