Intermediate level9 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how a mixture-of-experts model uses only part of itself for each token, and judge what “total” and “active” parameters mean for cost and memory.

The short answerA mixture-of-experts model replaces the feed-forward layer in each block with many expert networks and a router. The router picks a few experts for every token, so computing follows the active parameters, while memory still follows the total.

In simple words

  • Mixtral 8x7B has 8 experts per layer and uses 2 per token: 47 billion parameters in total, 13 billion active.
  • DeepSeek-V3 has 671 billion parameters but uses 37 billion per token.
  • Memory follows the total, because every expert must stay loaded.
  • Many of the largest open models use the design, and Google says Gemini does too.
A big model that uses a small part of itselfIllustration + live data

1 Each token gets two of eight experts

“The” goes to experts 3 and 6. The other six do no work for this token, so it costs about a quarter of the expert computing.

2 Stored versus used per token

  1. Mixtral 8x7B Mistral AI13B used of 47B · 28%
  2. gpt-oss-120b OpenAI5.1B used of 116.8B · 4.4%
  3. DeepSeek-V3 DeepSeek37B used of 671B · 5.5%
  4. Mistral Large 4 Mistral AI49B used of 1T · 4.9%
  5. Kimi K3 Moonshot AI104B used of 2.8T · 3.7%
  • Active: used for each token
  • Stored: all of it must fit in memory
The router’s choices are made up to show the idea. Real routers pick experts for every token in every layer; in Mixtral, researchers found the experts were not split by subject. The parameter counts come from the makers’ papers and model cards, and makers count active parameters in slightly different ways.

Words to know

Router
The small part that scores the experts and picks a few for each token.
Expert
One of many small networks in a layer; only the chosen ones run.
Active parameters
The parameters used for one token. The others do no work for it, but must still sit in memory.

Where the experts sit

In a normal, dense transformer, every token passes through every parameter. In a mixture-of-experts model, the feed-forward layer of each block is replaced by several expert networks. A small router scores the experts for each token and keeps only the top few.

In Mistral’s Mixtral 8x7B, each layer has 8 experts, and the router picks 2 for every token, in every layer. The pair can change from one token to the next.

Total and active parameters

Model cards now give two numbers. Total parameters are everything the model stores. Active parameters are those used for one token, and they set the computing cost. Mixtral has 47 billion in total and 13 billion active.

The gap has grown. DeepSeek-V3 has 671 billion parameters and 37 billion active. Moonshot’s Kimi K3 has 2.8 trillion and 104 billion active.

The counts are not always comparable. OpenAI counts gpt-oss’s output layer as active but not its input embeddings. Mistral gives 49 billion active for Large 4, or 52 billion with the embeddings and output layers.

Memory, not computing, is the limit

Every expert must stay in memory, because the next token may need any of them. The Mixtral paper notes that serving memory follows the total parameters, and that routing adds some overhead.

Makers work around this. OpenAI stores gpt-oss’s expert weights in about 4 bits each, a method called quantisation (lesson 4), so the 117-billion-parameter model fits on one 80 GB graphics chip. Meta says Llama 4 Scout fits on one H100 chip with its weights stored in 4 bits.

Why the design spread

The idea dates from 1991. In 2017, Google researchers used a trainable router to build a model with up to 137 billion parameters. For years, complexity, communication costs and unstable training held the design back.

Today many of the largest open models use it, including DeepSeek’s, Qwen’s, Kimi’s and OpenAI’s gpt-oss. Google says its Gemini 2.5 and Gemini 3 Pro models are sparse mixture-of-experts transformers, without giving their sizes.

Try it yourself

Open the model release tracker and find Mistral Large 4 and Reflection’s Beam. What share of each model is used per token? Would Beam fit on a laptop with 128 GB of memory? Tip: stored in 4 bits, each billion parameters needs about 0.5 GB.

Open the model release tracker

Check yourself

  1. Mixtral 8x7B has 47 billion parameters and uses 13 billion per token. Roughly what does it cost to run, compared with dense models?

  2. Why are two makers’ “active parameter” numbers not always directly comparable?

  3. A team wants to run a 1-trillion-parameter mixture-of-experts model with 32 billion active parameters on its own servers. What decides whether it fits?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.