Beginner level8 min readTo learn at another level, choose it before you start the course.
After this lesson you canExplain how a mixture-of-experts model uses only part of itself for each token, and judge what “total” and “active” parameters mean for cost and memory.
The short answerSome huge models are split into many small parts called experts. For each piece of text, a router switches on only a few of them, so the model does far less work than its size suggests.
In simple words
- A mixture-of-experts model has many experts, but uses only a few for each token.
- A small part called the router picks which experts to use.
- So it computes like a small model, but it still needs the memory of a big one.
- In a well-studied model, the experts were not split by subject; the router learns its own way to share the work.
The short answerA mixture-of-experts model replaces the feed-forward layer in each block with many expert networks and a router. The router picks a few experts for every token, so computing follows the active parameters, while memory still follows the total.
In simple words
- Mixtral 8x7B has 8 experts per layer and uses 2 per token: 47 billion parameters in total, 13 billion active.
- DeepSeek-V3 has 671 billion parameters but uses 37 billion per token.
- Memory follows the total, because every expert must stay loaded.
- Many of the largest open models use the design, and Google says Gemini does too.
The short answerSparse mixture-of-experts layers separate capacity from per-token computing: a learned gate sends each token to the top k of many feed-forward experts. The price is memory for every expert, routing and communication overhead, load balancing, and harder training and fine-tuning.
Key points
- Gating dates from Jacobs and others (1991); Shazeer and others (2017) scaled sparse gating to 137 billion parameters.
- The Switch Transformer sent each token to one expert and reported up to 7 times faster pre-training at equal compute.
- Experts on different chips need even load: capacity limits, dropped tokens and balancing losses.
- DeepSeekMoE adds many small experts plus shared experts that every token uses.
1 Each token gets two of eight experts
“The” goes to experts 3 and 6. The other six do no work for this token, so it costs about a quarter of the expert computing.
2 Stored versus used per token
- Active: used for each token
- Stored: all of it must fit in memory
Words to know
- Router
- The small part that scores the experts and picks a few for each token.
- Expert
- One of many small networks in a layer; only the chosen ones run.
- Active parameters
- The parameters used for one token. The others do no work for it, but must still sit in memory.
A big team, a few at a time
Imagine a firm with 128 workers, where each small task goes to just four of them. The firm is big, but each task is cheap. A mixture-of-experts model works in a similar way.
Inside each block, the model has many experts, small networks of their own. For every token, a router picks a few of them, and the rest stay idle. OpenAI’s open gpt-oss-120b model works like the firm: 128 experts in each block, 4 used per token.
So gpt-oss-120b has about 117 billion parameters, but uses only about 5 billion for each token. Those 5 billion are its active parameters.
Cheap to run, but not small
Using fewer parameters per token means less computing, so answers can come faster and cost less. Google says this is why its Gemini models use the design.
But every expert must stay loaded, because the next token may need a different one. So a huge model still needs a lot of memory, even if each token uses only a small part of it.
Not a team of subject experts
The name suggests a law expert, a history expert and so on. In Mistral’s Mixtral model, researchers found no such split: science papers, biology texts and philosophy texts used the experts in similar ways.
The router learns its own way to split the work during training. Its choices depend more on the kind of token, such as a word in computer code, than on the topic.
The idea is also not new: it comes from a research paper published in 1991.
Try it yourself
Open the model release tracker and find Mistral Large 4 and Reflection’s Beam. What share of each model is used per token? Would Beam fit on a laptop with 128 GB of memory? Tip: stored in 4 bits, each billion parameters needs about 0.5 GB.
Open the model release trackerCheck yourself
A model has 117 billion parameters but uses about 5 billion for each token. Why is it cheaper to run than a normal model of its size?
Why does a mixture-of-experts model still need a lot of memory?
Someone says a model has a “law expert” and a “history expert” inside. What did researchers find in Mixtral?
Words to know
- Router
- The small part that scores the experts and picks a few for each token.
- Expert
- One of many small networks in a layer; only the chosen ones run.
- Active parameters
- The parameters used for one token. The others do no work for it, but must still sit in memory.
Where the experts sit
In a normal, dense transformer, every token passes through every parameter. In a mixture-of-experts model, the feed-forward layer of each block is replaced by several expert networks. A small router scores the experts for each token and keeps only the top few.
In Mistral’s Mixtral 8x7B, each layer has 8 experts, and the router picks 2 for every token, in every layer. The pair can change from one token to the next.
Total and active parameters
Model cards now give two numbers. Total parameters are everything the model stores. Active parameters are those used for one token, and they set the computing cost. Mixtral has 47 billion in total and 13 billion active.
The gap has grown. DeepSeek-V3 has 671 billion parameters and 37 billion active. Moonshot’s Kimi K3 has 2.8 trillion and 104 billion active.
The counts are not always comparable. OpenAI counts gpt-oss’s output layer as active but not its input embeddings. Mistral gives 49 billion active for Large 4, or 52 billion with the embeddings and output layers.
Memory, not computing, is the limit
Every expert must stay in memory, because the next token may need any of them. The Mixtral paper notes that serving memory follows the total parameters, and that routing adds some overhead.
Makers work around this. OpenAI stores gpt-oss’s expert weights in about 4 bits each, a method called quantisation (lesson 4), so the 117-billion-parameter model fits on one 80 GB graphics chip. Meta says Llama 4 Scout fits on one H100 chip with its weights stored in 4 bits.
Why the design spread
The idea dates from 1991. In 2017, Google researchers used a trainable router to build a model with up to 137 billion parameters. For years, complexity, communication costs and unstable training held the design back.
Today many of the largest open models use it, including DeepSeek’s, Qwen’s, Kimi’s and OpenAI’s gpt-oss. Google says its Gemini 2.5 and Gemini 3 Pro models are sparse mixture-of-experts transformers, without giving their sizes.
Try it yourself
Open the model release tracker and find Mistral Large 4 and Reflection’s Beam. What share of each model is used per token? Would Beam fit on a laptop with 128 GB of memory? Tip: stored in 4 bits, each billion parameters needs about 0.5 GB.
Open the model release trackerCheck yourself
Mixtral 8x7B has 47 billion parameters and uses 13 billion per token. Roughly what does it cost to run, compared with dense models?
Why are two makers’ “active parameter” numbers not always directly comparable?
A team wants to run a 1-trillion-parameter mixture-of-experts model with 32 billion active parameters on its own servers. What decides whether it fits?
From 1991 to trillion-parameter models
Jacobs, Jordan, Nowlan and Hinton proposed adaptive mixtures of local experts in 1991. Shazeer and others (2017) added a trainable sparsely gated layer that picks a few of up to thousands of feed-forward sub-networks. It reached 137 billion parameters with minor efficiency losses.
GShard (2020) scaled a translation model past 600 billion parameters, and the Switch Transformer (2021) simplified routing to one expert per token. Switch reported up to 7 times faster pre-training for the same compute, and trained models with over a trillion parameters.
Routing, capacity and balance
A router scores all experts and keeps the top k: Mixtral uses 2 of 8, Qwen3-235B 8 of 128, and gpt-oss-120b 4 of 128. Experts often sit on different chips, called expert parallelism, so uneven routing leaves some chips idle and others overloaded.
Switch gives each expert a fixed capacity. Tokens beyond it are dropped from that layer and skip it, typically under 1%. An auxiliary loss pushes the router to spread tokens evenly; DeepSeek-V3 reports balancing load mostly without that extra loss.
Finer experts and shared experts
DeepSeekMoE splits experts into many smaller ones and adds shared experts that every token uses, so the routed experts can specialise. DeepSeek reports that its 16-billion model matched LLaMA2 7B with about 40% of the computation.
DeepSeek-V3 combines these ideas at 671 billion total and 37 billion active parameters, trained on 14.8 trillion tokens. DeepSeek reports 2.788 million H800 GPU hours for its final training, a company figure that leaves out earlier research and test runs.
The costs that remain
Memory follows the total. The Mixtral authors note that routing adds overhead and that the design suits batched serving: with many requests at once, each loaded expert gets enough tokens to keep the chips busy. gpt-oss stores its expert weights, over 90% of the parameters, at about 4 bits, so 116.8 billion parameters fit in a 60.8 GiB checkpoint.
Switch found that sparse models can overfit more when fine-tuned on small tasks. It also distilled one into a small dense model with over 90% less memory, keeping about 30% of its quality gain. ST-MoE (2022) still named instability and uncertain fine-tuning quality as limits.
Even “active” is a convention: gpt-oss counts the output layer but not the input embeddings, and Mistral quotes 49 or 52 billion for Large 4.
Try it yourself
Open the model release tracker and find Mistral Large 4 and Reflection’s Beam. What share of each model is used per token? Would Beam fit on a laptop with 128 GB of memory? Tip: stored in 4 bits, each billion parameters needs about 0.5 GB.
Open the model release trackerCheck yourself
In a Switch-style layer, an expert receives more tokens than its capacity. What happens to the extra tokens?
Why does serving a large mixture-of-experts model favour batched traffic?
What do DeepSeekMoE’s shared experts do?
Sources
- Adaptive Mixtures of Local Experts (Jacobs and others, 1991, Neural Computation)
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (Shazeer and others, 2017, arXiv)
- GShard (Lepikhin and others, 2020, arXiv)
- Switch Transformers (Fedus, Zoph and Shazeer, 2021, arXiv)
- ST-MoE: Designing Stable and Transferable Sparse Expert Models (Zoph and others, 2022, arXiv)
- Mixtral of Experts (Jiang and others, 2024, arXiv)
- DeepSeekMoE (Dai and others, 2024, arXiv)
- DeepSeek-V3 Technical Report (DeepSeek, 2024, arXiv)
- gpt-oss-120b and gpt-oss-20b model card (OpenAI, 2025, arXiv)
- Qwen3-235B-A22B model card (Qwen, Hugging Face)
- Kimi K3 model card (Moonshot AI, Hugging Face)
- Mistral Large 4 model card (Mistral AI, Hugging Face)
- Gemini 2.5 technical report (Google, 2025, arXiv)
- Gemini 3 Pro model card (Google DeepMind, 2025)
- The Llama 4 herd (Meta, 2025)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.