Beginner level8 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how a mixture-of-experts model uses only part of itself for each token, and judge what “total” and “active” parameters mean for cost and memory.

The short answerSome huge models are split into many small parts called experts. For each piece of text, a router switches on only a few of them, so the model does far less work than its size suggests.

In simple words

  • A mixture-of-experts model has many experts, but uses only a few for each token.
  • A small part called the router picks which experts to use.
  • So it computes like a small model, but it still needs the memory of a big one.
  • In a well-studied model, the experts were not split by subject; the router learns its own way to share the work.
A big model that uses a small part of itselfIllustration + live data

1 Each token gets two of eight experts

“The” goes to experts 3 and 6. The other six do no work for this token, so it costs about a quarter of the expert computing.

2 Stored versus used per token

  1. Mixtral 8x7B Mistral AI13B used of 47B · 28%
  2. gpt-oss-120b OpenAI5.1B used of 116.8B · 4.4%
  3. DeepSeek-V3 DeepSeek37B used of 671B · 5.5%
  4. Mistral Large 4 Mistral AI49B used of 1T · 4.9%
  5. Kimi K3 Moonshot AI104B used of 2.8T · 3.7%
  • Active: used for each token
  • Stored: all of it must fit in memory
The router’s choices are made up to show the idea. Real routers pick experts for every token in every layer; in Mixtral, researchers found the experts were not split by subject. The parameter counts come from the makers’ papers and model cards, and makers count active parameters in slightly different ways.

Words to know

Router
The small part that scores the experts and picks a few for each token.
Expert
One of many small networks in a layer; only the chosen ones run.
Active parameters
The parameters used for one token. The others do no work for it, but must still sit in memory.

A big team, a few at a time

Imagine a firm with 128 workers, where each small task goes to just four of them. The firm is big, but each task is cheap. A mixture-of-experts model works in a similar way.

Inside each block, the model has many experts, small networks of their own. For every token, a router picks a few of them, and the rest stay idle. OpenAI’s open gpt-oss-120b model works like the firm: 128 experts in each block, 4 used per token.

So gpt-oss-120b has about 117 billion parameters, but uses only about 5 billion for each token. Those 5 billion are its active parameters.

Cheap to run, but not small

Using fewer parameters per token means less computing, so answers can come faster and cost less. Google says this is why its Gemini models use the design.

But every expert must stay loaded, because the next token may need a different one. So a huge model still needs a lot of memory, even if each token uses only a small part of it.

Not a team of subject experts

The name suggests a law expert, a history expert and so on. In Mistral’s Mixtral model, researchers found no such split: science papers, biology texts and philosophy texts used the experts in similar ways.

The router learns its own way to split the work during training. Its choices depend more on the kind of token, such as a word in computer code, than on the topic.

The idea is also not new: it comes from a research paper published in 1991.

Try it yourself

Open the model release tracker and find Mistral Large 4 and Reflection’s Beam. What share of each model is used per token? Would Beam fit on a laptop with 128 GB of memory? Tip: stored in 4 bits, each billion parameters needs about 0.5 GB.

Open the model release tracker

Check yourself

  1. A model has 117 billion parameters but uses about 5 billion for each token. Why is it cheaper to run than a normal model of its size?

  2. Why does a mixture-of-experts model still need a lot of memory?

  3. Someone says a model has a “law expert” and a “history expert” inside. What did researchers find in Mixtral?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.