Intermediate level9 min readTo learn at another level, choose it before you start the course.
After this lesson you canExplain why attention gets expensive on long inputs, describe how state space and recurrent designs keep a fixed-size memory, and say why most new designs are hybrids.
The short answerYes: some newer designs read text the way a person reads a book, keeping a short running summary instead of looking back at every word. They are faster on very long texts but worse at finding one exact detail, so many new models mix both designs.
In simple words
- A transformer looks back at every earlier word, so long texts get slow and use a lot of memory.
- State space models, such as Mamba, keep one running summary of a fixed size.
- That makes them fast on long texts, but worse at recalling one exact detail.
- Many new models are hybrids: mostly the new layers, plus a few transformer layers.
The short answerAlternatives to attention, such as state space models, linear attention and new recurrent networks, process text with a fixed-size state, so their cost grows linearly with length. They are faster and leaner on long inputs but weaker at exact recall, which is why most new releases that use them are hybrids that keep some attention.
In simple words
- Attention’s cost grows with the square of the length, and its key-value cache grows with every token.
- Mamba (2023) made the state update depend on the input; RWKV and xLSTM are recurrent networks that train in parallel.
- Fixed-size states struggle with copying and exact lookup.
- Jamba, Granite 4.0-H, Qwen3.5 and Nemotron 3 are hybrids; MiniMax returned to full attention for M2.
The short answerSub-quadratic sequence models replace softmax attention with a recurrent state: structured or selective state space models, linear attention, and recurrent networks that train in parallel. They cut memory and time on long contexts but lose exact recall, so production designs interleave them with a few attention layers.
Key points
- S4 (2021), Mamba (2023) and Mamba-2 (2024), which links state space models and attention mathematically.
- Linear attention, RWKV and xLSTM reach linear cost from the attention side and the recurrent side.
- Copying and associative recall are the known weak spots; state tracking is hard for transformers too.
- Hybrid ratios in releases: Jamba 1:7, Granite 4.0-H 1:9, Qwen3-Next and Qwen3.5 1:3.
1 Notes kept for one conversation
At 8,000 tokens, the transformer keeps 1.0 GB of notes for this one conversation, in 16-bit numbers. A state space model keeps one state of fixed size, however long the conversation gets.
2 How the designs developed
- LSTMA recurrent network: reads one token at a time and carries a running summary forward.
- TransformerAttention: every token can look back at every earlier token. Trains on a whole text at once.
- S4 and MambaState space models with a fixed-size memory; Mamba chooses what to keep as it reads.
- HybridsMostly state space or linear layers, plus a few attention layers for exact recall: Jamba, Granite 4.0-H, Qwen3.5, Nemotron 3.
Words to know
- State space model
- A layer that keeps a fixed-size running summary of everything read so far.
- Recurrent network
- A network that reads one token at a time and carries a running summary forward.
- Hybrid model
- Mostly state space or linear layers, plus a few attention layers for exact recall.
Why long texts are hard for transformers
A transformer compares every word with every earlier word. Double the text, and that work grows four times.
It also keeps notes on every word it has read, and the notes grow with the text. In the diagram above, Meta’s Llama 3.1 8B needs about 17 GB of notes for one conversation of 128,000 tokens.
A running summary
Older designs, called recurrent networks, read one word at a time and carry a running summary that they update as they read. In 2017 the transformer took over, partly because it trains on a whole text at once, which is much faster.
State space models bring the running summary back in a new form. Mamba, from 2023, decides what to keep and what to forget as it reads. Its memory stays the same size however long the text gets, and its authors report about five times a transformer’s speed when writing.
The weak spot, and the mix
A summary of a fixed size cannot hold every detail. Researchers found that transformers are far better at copying text or finding one exact fact in a long input.
So many new models mix both: mostly state space or similar layers, plus a few transformer layers for exact recall. IBM’s Granite 4.0-H and NVIDIA’s Nemotron 3 models work this way.
Not every maker agrees. MiniMax went back to full attention for its M2 model.
Try it yourself
Ask a chat app for 20 made-up phone numbers. Read them once, cover them, and try to say the 12th. Your memory kept a short summary, like a state space model. Looking back at the list, as attention does, finds it at once.
Check yourself
Why do very long texts slow a transformer down?
What does a state space model such as Mamba keep while it reads?
A lawyer needs a model to quote one exact clause from a 500-page contract. What is a weak point of a pure state space model here?
Words to know
- State space model
- A layer that keeps a fixed-size running summary of everything read so far.
- Recurrent network
- A network that reads one token at a time and carries a running summary forward.
- Hybrid model
- Mostly state space or linear layers, plus a few attention layers for exact recall.
Two costs of attention
Self-attention compares every token with every other one, so its computing cost grows with the square of the input length. A transformer also stores a key-value cache, notes on every token, which grows as the text gets longer.
AI21’s own table gives a bigger example. Llama 2 7B keeps four times as many notes per token as Llama 3.1 8B, because it shares no key-value heads. At 256K tokens, the table gives it a 128 GB cache, against 4 GB for AI21’s hybrid, Jamba.
Designs with a fixed-size state
State space models carry a fixed-size state forward. S4 (2021) solved a long-range test with sequences of 16,000 steps. Mamba (2023) made the state update depend on the input, so the model can choose what to keep; its authors report about 5 times a transformer’s speed when writing.
Other routes lead to the same place. Linear attention rewrites attention so its cost grows linearly, which turns a transformer into a kind of recurrent network. RWKV and xLSTM are recurrent networks redesigned to train in parallel, like a transformer.
Where they fall short
A fixed-size state limits exact recall. Jelassi and others (2024) found that transformers far outperform state space models at copying and retrieving from the input. Another study traced about 82% of one quality gap to recalling earlier context.
At 8 billion parameters, an NVIDIA-led study found pure Mamba behind transformers on some tasks, such as phonebook-style lookup. A hybrid with a few attention layers beat the transformer on all 12 standard tasks it tested.
Hybrids in today’s releases
Most new models of this kind are hybrids. AI21’s Jamba (2024) uses one attention layer for every seven Mamba layers, and IBM’s Granite 4.0-H uses nine Mamba-2 layers for every transformer layer. Alibaba’s Qwen3.5 and NVIDIA’s Nemotron 3 models are hybrids too.
Not everyone is convinced. MiniMax dropped its hybrid design for M2, reporting “clear deficits” in multi-step reasoning at larger scale. The layer designs of today’s top closed models are not published.
Our stories touch this without naming it: Salesforce’s Koa is built on Nemotron 3 Super, and Cloudflare’s Clef-flash on Qwen3.5-9B, both hybrids according to their makers.
Try it yourself
Ask a chat app for 20 made-up phone numbers. Read them once, cover them, and try to say the 12th. Your memory kept a short summary, like a state space model. Looking back at the list, as attention does, finds it at once.
Check yourself
Mamba’s cost grows linearly with text length. What is its main weakness compared with a transformer?
Why do most new releases with Mamba-style layers keep a few attention layers?
What is the difference between a hybrid model and a mixture-of-experts (MoE) model?
State space models
S4 (Gu, Goel and Ré, 2021) used structured state spaces and solved the Path-X task of Long Range Arena, at length 16,000. Mamba (Gu and Dao, 2023) made the state update input-dependent, or selective, with cost linear in length.
The Mamba authors report about 5 times the inference throughput of transformers. Mamba-2 (Dao and Gu, 2024) showed that state space models and attention are mathematically related, and its core layer runs 2 to 8 times faster than Mamba’s.
Mamba-3’s authors (2026) still note that many linear models trade quality for speed and fail at tasks such as state tracking. Merrill and others (2024) argue that state space models, like transformers, cannot express simple state tracking, so a recurrent state is no automatic advantage there.
Linear attention and new recurrent networks
Katharopoulos and others (2020) rewrote attention so that its cost grows linearly, showing that a transformer can run as a recurrent network. RWKV (2023) trains in parallel but runs with constant memory per step, and was scaled to 14 billion parameters.
xLSTM (2024) extends the 1997 LSTM with exponential gating and a matrix memory that trains in parallel. Newer releases use gated linear layers, such as the Gated DeltaNet in Qwen3-Next and Qwen3.5, and the KDA layers of Kimi Linear.
The recall problem
Jelassi and others (2024) showed that transformers far outperform state space models at copying and retrieving from context. The Zoology study traced about 82% of the gap between attention and gated-convolution models to recall of earlier context.
Waleffe and others (2024), an NVIDIA-led study with co-authors from state space model companies, found pure Mamba at 8 billion parameters behind on 5-shot MMLU, phonebook lookup and long-context reasoning.
Their hybrid of 43% Mamba-2, 7% attention and 50% feed-forward (MLP) layers beat the transformer on all 12 standard tasks, by 2.65 points on average. It is predicted to generate up to 8 times faster.
What the releases show
Ratios vary. Jamba has 1 attention layer per 7 Mamba layers, plus mixture of experts, and Granite 4.0-H has 1 transformer layer per 9 Mamba-2 layers. Qwen3-Next, Qwen3.5 and Kimi Linear have 1 attention layer per 3 linear ones. NVIDIA’s Nemotron 3 Nano has 23 Mamba-2, 23 mixture-of-experts (MoE) and 6 attention layers.
Efficiency figures are mostly makers’ own: IBM claims over 70% less memory, and Qwen claims 10 times the throughput above 32,000 tokens. In Google’s test, Gemma-7B ran out of memory at 8,192 output tokens, while RecurrentGemma-9B finished.
MiniMax dropped its lightning-attention hybrid for M2, citing “clear deficits” in multi-hop reasoning at scale and less mature serving software. Whether hybrids hold up at frontier scale is still open.
Try it yourself
Ask a chat app for 20 made-up phone numbers. Read them once, cover them, and try to say the 12th. Your memory kept a short summary, like a state space model. Looking back at the list, as attention does, finds it at once.
Check yourself
Why can a selective state space model not guarantee exact retrieval from a long prompt?
In Waleffe and others (2024), what did a hybrid with about 7% attention layers achieve at 8 billion parameters?
A vendor offers a pure state space model to find exact clauses in long contracts. What should you test first?
Sources
- Long Short-Term Memory (Hochreiter and Schmidhuber, 1997, Neural Computation)
- Transformers are RNNs: linear attention (Katharopoulos and others, 2020, arXiv)
- Efficiently Modeling Long Sequences with Structured State Spaces, S4 (Gu, Goel and Ré, 2021, arXiv)
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces (Gu and Dao, 2023, arXiv)
- Transformers are SSMs, Mamba-2 (Dao and Gu, 2024, arXiv)
- Mamba-3 (2026, arXiv)
- The Illusion of State in State-Space Models (Merrill, Petty and Sabharwal, 2024, arXiv)
- RWKV: Reinventing RNNs for the Transformer Era (Peng and others, 2023, arXiv)
- xLSTM: Extended Long Short-Term Memory (Beck and others, 2024, arXiv)
- Repeat After Me: Transformers are Better than State Space Models at Copying (Jelassi and others, 2024, arXiv)
- Zoology: Measuring and Improving Recall in Efficient Language Models (Arora and others, 2023, arXiv)
- An Empirical Study of Mamba-based Language Models (Waleffe and others, 2024, arXiv)
- Jamba: A Hybrid Transformer-Mamba Language Model (AI21, 2024, arXiv)
- IBM Granite 4.0: hybrid Mamba-2 and transformer models (IBM, October 2025)
- Qwen3-Next-80B-A3B model card (Qwen, Hugging Face)
- Qwen3.5-397B-A17B model card (Qwen, Hugging Face)
- Qwen3.5-9B model card: the base of Cloudflare’s Clef-flash (Qwen, Hugging Face)
- The Llama 3 Herd of Models: the settings behind the memory figure (Meta, 2024, arXiv)
- Llama 2: Open Foundation and Fine-Tuned Chat Models (Touvron and others, 2023, arXiv)
- Kimi Linear (Kimi team, 2025, arXiv)
- NVIDIA Nemotron 3 Nano model card (NVIDIA, Hugging Face)
- Why did M2 end up as a full attention model? (MiniMax, 2025)
- RecurrentGemma (Google DeepMind, 2024, arXiv)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.