// what changed
A conventional dense model runs every parameter for every token it produces. That is the baseline the whole idea was invented against. The Switch Transformer paper describes a mixture of experts as a model that instead "selects different parameters for each incoming example", producing a "sparsely-activated" network with "outrageous numbers of parameters — but a constant computational cost".
So the two numbers answer two different questions. The active count is how much work happens per token. The total is how much the model knows — and, as it turns out, how much room it needs.
What an expert actually is
It is narrower than the name suggests. The experts are not copies of the model and they are not specialists in topics you could name. They are the feed-forward blocks inside each layer, replicated, with a small router in front choosing which few to run.
The Mixtral paper is precise about it: "each layer is composed of 8 feedforward blocks (i.e. experts)", and for every token "a router network selects two experts to process the current state and combine their outputs". The result is that "each token has access to 47B parameters, but only uses 13B active parameters during inference".
And the choice is remade constantly. It is not one decision per token; it is one decision per layer, per token. In Qwen3-30B-A3B that is 8 experts out of 128, chosen again at each of 48 layers.
Why the labs all switched
Because it breaks a tradeoff that used to be fixed. Adding parameters to a dense model makes it slower and more expensive to serve, in direct proportion. Adding experts to a sparse one adds capacity while the per-token work stays where it was.
Mistral stated the consequence in product terms when they launched Mixtral: it "processes input and generates output at the same speed and for the same cost as a 12.9B model". Note carefully what that sentence claims — speed and cost. It says nothing about memory, and that gap is where most of the confusion lives.
The clearest demonstration
Meta shipped it in a single release. Llama 4 Scout and Llama 4 Maverick activate the same 17B parameters per token. Their totals differ by a factor of nearly four.
One variable moves, one does not. That is as close to a controlled experiment as a model release gets, and it settles the question: the active count tells you roughly how fast, and nothing at all about how big.
How far "everything is MoE now" actually goes
Far enough to be worth knowing, and not as far as the listicles claim. Counted from vendor cards only, across five labs:
The rule
The small number is the speed. The big number is the knowledge. A name like 30B-A3B is telling you both at once, and the one people plan around is the wrong one.
Which leads directly to the next question, and the one that catches people out: if all those experts have to be somewhere, where are they? Active parameters are not your RAM bill covers what the total actually costs you to run, and what to do when it won't fit.
Related: quantization is the bits-per-weight half of the same memory question, and whether to run one locally at all is the prior question.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
