Beginner level8 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain why attention gets expensive on long inputs, describe how state space and recurrent designs keep a fixed-size memory, and say why most new designs are hybrids.

The short answerYes: some newer designs read text the way a person reads a book, keeping a short running summary instead of looking back at every word. They are faster on very long texts but worse at finding one exact detail, so many new models mix both designs.

In simple words

  • A transformer looks back at every earlier word, so long texts get slow and use a lot of memory.
  • State space models, such as Mamba, keep one running summary of a fixed size.
  • That makes them fast on long texts, but worse at recalling one exact detail.
  • Many new models are hybrids: mostly the new layers, plus a few transformer layers.
Memory that grows, memory that stays the sameIllustration + live data

1 Notes kept for one conversation

  1. Transformer Llama 3.1 8B1.0 GB
  2. State space model fixed statesame

At 8,000 tokens, the transformer keeps 1.0 GB of notes for this one conversation, in 16-bit numbers. A state space model keeps one state of fixed size, however long the conversation gets.

2 How the designs developed

  1. LSTMA recurrent network: reads one token at a time and carries a running summary forward.
  2. TransformerAttention: every token can look back at every earlier token. Trains on a whole text at once.
  3. S4 and MambaState space models with a fixed-size memory; Mamba chooses what to keep as it reads.
  4. HybridsMostly state space or linear layers, plus a few attention layers for exact recall: Jamba, Granite 4.0-H, Qwen3.5, Nemotron 3.
The memory figure is worked out from Meta’s published settings for Llama 3.1 8B: 32 layers, 8 key-value heads of 128 numbers, 2 bytes each. It covers the notes for one conversation, not the model itself. Not every maker uses hybrids: MiniMax went back to full attention for its M2 model.

Words to know

State space model
A layer that keeps a fixed-size running summary of everything read so far.
Recurrent network
A network that reads one token at a time and carries a running summary forward.
Hybrid model
Mostly state space or linear layers, plus a few attention layers for exact recall.

Why long texts are hard for transformers

A transformer compares every word with every earlier word. Double the text, and that work grows four times.

It also keeps notes on every word it has read, and the notes grow with the text. In the diagram above, Meta’s Llama 3.1 8B needs about 17 GB of notes for one conversation of 128,000 tokens.

A running summary

Older designs, called recurrent networks, read one word at a time and carry a running summary that they update as they read. In 2017 the transformer took over, partly because it trains on a whole text at once, which is much faster.

State space models bring the running summary back in a new form. Mamba, from 2023, decides what to keep and what to forget as it reads. Its memory stays the same size however long the text gets, and its authors report about five times a transformer’s speed when writing.

The weak spot, and the mix

A summary of a fixed size cannot hold every detail. Researchers found that transformers are far better at copying text or finding one exact fact in a long input.

So many new models mix both: mostly state space or similar layers, plus a few transformer layers for exact recall. IBM’s Granite 4.0-H and NVIDIA’s Nemotron 3 models work this way.

Not every maker agrees. MiniMax went back to full attention for its M2 model.

Try it yourself

Ask a chat app for 20 made-up phone numbers. Read them once, cover them, and try to say the 12th. Your memory kept a short summary, like a state space model. Looking back at the list, as attention does, finds it at once.

Check yourself

  1. Why do very long texts slow a transformer down?

  2. What does a state space model such as Mamba keep while it reads?

  3. A lawyer needs a model to quote one exact clause from a 500-page contract. What is a weak point of a pure state space model here?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.