// the misread
A card that says 30.5B total, 3.3B active is describing two different things. Only a small slice of the network runs for any given token, so the model generates roughly as fast as a small dense one. That is genuinely what you are buying.
What it is not telling you is how much memory to have. Same vendor, same family, same quantization:
The model that reads about as much per token as a 4B dense model needs about seven times the room to sit in.
Why every expert has to stay loaded
A mixture-of-experts layer holds many expert sub-networks and uses a few of them per token. The reason you cannot keep only the few is that the choice is remade constantly: the model picks a different handful at every layer, for every token. Over a full generation it reaches for essentially all of them, so runtimes simply keep them all resident.
The Mixtral paper puts the consequence plainly: the memory cost of serving the model is proportional to its 47B total, not its 13B active.
What the active count actually buys
Two things, both real:
Compute per token. Fewer parameters participate, so there is less arithmetic per token.
Bytes read per token. This is the one that matters most for a single stream on a consumer card. Generating one token at a time is dominated by the time spent pulling weights out of memory, so reading a small slice is what makes an MoE feel fast.
Expect "roughly" rather than exactly — routing and gathering the expert weights cost something, and the Mixtral paper flags that overhead explicitly.
When it doesn't fit
The instinct is to offload whole layers to system RAM. For an MoE that is the wrong cut. The experts are nearly all of the weight but only a small slice of the per-token work, while attention is small and touched by every single token. So move the experts and keep attention on the card.
One clean benchmark on an RTX 3090 went from 139.6 to 89.1 tokens/sec of generation while going from a 89.6K to the full 262K context — roughly two thirds of the speed for about three times the context.
Is it as good as a dense model of the same size?
No, and be wary of anyone who gives you a formula. The popular "geometric mean of total and active" rule has no traceable source and is badly wrong in both directions — it implies Mixtral 8x7B behaves like a 24B model, when Mistral reports it beating a 70B dense one.
The honest version: a 30B-A3B lands somewhere between its active and total size. Qwen never published a head-to-head against their own 32B dense model; the nearest first-party comparison is against their 14B, where the two trade blows.
So treat the total as your memory budget, the active count as your speed estimate, and the quality as something to measure on your own task rather than derive.
The rule
Size your hardware off the total. Expect the speed of the active part. If the name has an "A" number in it, that number is the read-per-token figure, not the download.
Related: quantization covers the bits-per-weight half of the memory question, and the KV-cache section on that same page covers the part that grows with your context window. Total parameters, bits per weight, and context are three separate lines on the same bill — and only the first is what an "A" number is talking about. Whether you need a local model at all is a prior question.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
