Intermediate level9 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain why attention gets expensive on long inputs, describe how state space and recurrent designs keep a fixed-size memory, and say why most new designs are hybrids.

The short answerAlternatives to attention, such as state space models, linear attention and new recurrent networks, process text with a fixed-size state, so their cost grows linearly with length. They are faster and leaner on long inputs but weaker at exact recall, which is why most new releases that use them are hybrids that keep some attention.

In simple words

  • Attention’s cost grows with the square of the length, and its key-value cache grows with every token.
  • Mamba (2023) made the state update depend on the input; RWKV and xLSTM are recurrent networks that train in parallel.
  • Fixed-size states struggle with copying and exact lookup.
  • Jamba, Granite 4.0-H, Qwen3.5 and Nemotron 3 are hybrids; MiniMax returned to full attention for M2.
Memory that grows, memory that stays the sameIllustration + live data

1 Notes kept for one conversation

  1. Transformer Llama 3.1 8B1.0 GB
  2. State space model fixed statesame

At 8,000 tokens, the transformer keeps 1.0 GB of notes for this one conversation, in 16-bit numbers. A state space model keeps one state of fixed size, however long the conversation gets.

2 How the designs developed

  1. LSTMA recurrent network: reads one token at a time and carries a running summary forward.
  2. TransformerAttention: every token can look back at every earlier token. Trains on a whole text at once.
  3. S4 and MambaState space models with a fixed-size memory; Mamba chooses what to keep as it reads.
  4. HybridsMostly state space or linear layers, plus a few attention layers for exact recall: Jamba, Granite 4.0-H, Qwen3.5, Nemotron 3.
The memory figure is worked out from Meta’s published settings for Llama 3.1 8B: 32 layers, 8 key-value heads of 128 numbers, 2 bytes each. It covers the notes for one conversation, not the model itself. Not every maker uses hybrids: MiniMax went back to full attention for its M2 model.

Words to know

State space model
A layer that keeps a fixed-size running summary of everything read so far.
Recurrent network
A network that reads one token at a time and carries a running summary forward.
Hybrid model
Mostly state space or linear layers, plus a few attention layers for exact recall.

Two costs of attention

Self-attention compares every token with every other one, so its computing cost grows with the square of the input length. A transformer also stores a key-value cache, notes on every token, which grows as the text gets longer.

AI21’s own table gives a bigger example. Llama 2 7B keeps four times as many notes per token as Llama 3.1 8B, because it shares no key-value heads. At 256K tokens, the table gives it a 128 GB cache, against 4 GB for AI21’s hybrid, Jamba.

Designs with a fixed-size state

State space models carry a fixed-size state forward. S4 (2021) solved a long-range test with sequences of 16,000 steps. Mamba (2023) made the state update depend on the input, so the model can choose what to keep; its authors report about 5 times a transformer’s speed when writing.

Other routes lead to the same place. Linear attention rewrites attention so its cost grows linearly, which turns a transformer into a kind of recurrent network. RWKV and xLSTM are recurrent networks redesigned to train in parallel, like a transformer.

Where they fall short

A fixed-size state limits exact recall. Jelassi and others (2024) found that transformers far outperform state space models at copying and retrieving from the input. Another study traced about 82% of one quality gap to recalling earlier context.

At 8 billion parameters, an NVIDIA-led study found pure Mamba behind transformers on some tasks, such as phonebook-style lookup. A hybrid with a few attention layers beat the transformer on all 12 standard tasks it tested.

Hybrids in today’s releases

Most new models of this kind are hybrids. AI21’s Jamba (2024) uses one attention layer for every seven Mamba layers, and IBM’s Granite 4.0-H uses nine Mamba-2 layers for every transformer layer. Alibaba’s Qwen3.5 and NVIDIA’s Nemotron 3 models are hybrids too.

Not everyone is convinced. MiniMax dropped its hybrid design for M2, reporting “clear deficits” in multi-step reasoning at larger scale. The layer designs of today’s top closed models are not published.

Our stories touch this without naming it: Salesforce’s Koa is built on Nemotron 3 Super, and Cloudflare’s Clef-flash on Qwen3.5-9B, both hybrids according to their makers.

Try it yourself

Ask a chat app for 20 made-up phone numbers. Read them once, cover them, and try to say the 12th. Your memory kept a short summary, like a state space model. Looking back at the list, as attention does, finds it at once.

Check yourself

  1. Mamba’s cost grows linearly with text length. What is its main weakness compared with a transformer?

  2. Why do most new releases with Mamba-style layers keep a few attention layers?

  3. What is the difference between a hybrid model and a mixture-of-experts (MoE) model?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.