toolcall() ← all concepts

// concept · local models

Three billion active. Eighteen gigabytes.

Mixture-of-experts model names carry two numbers, and the smaller one is the one everybody plans around. It is the wrong one to plan around. The active count tells you how fast the model runs; the total tells you what you have to buy.

// the misread

A card that says 30.5B total, 3.3B active is describing two different things. Only a small slice of the network runs for any given token, so the model generates roughly as fast as a small dense one. That is genuinely what you are buying.

What it is not telling you is how much memory to have. Same vendor, same family, same quantization:

30.5B total · 3.3B active 18.6 GB 4B dense 2.5 GB # about the same bytes read per token # roughly seven times the storage

The model that reads about as much per token as a 4B dense model needs about seven times the room to sit in.

Why every expert has to stay loaded

A mixture-of-experts layer holds many expert sub-networks and uses a few of them per token. The reason you cannot keep only the few is that the choice is remade constantly: the model picks a different handful at every layer, for every token. Over a full generation it reaches for essentially all of them, so runtimes simply keep them all resident.

The Mixtral paper puts the consequence plainly: the memory cost of serving the model is proportional to its 47B total, not its 13B active.

One nuance worth knowing. Expert selection is not purely random — the Mixtral authors measured some positional locality, and expert prefetching is an active research area. So "you can't predict which experts you need" is too strong. The practical position is simpler: nothing shipping today tries to, so plan for full residency.

What the active count actually buys

Two things, both real:

Compute per token. Fewer parameters participate, so there is less arithmetic per token.

Bytes read per token. This is the one that matters most for a single stream on a consumer card. Generating one token at a time is dominated by the time spent pulling weights out of memory, so reading a small slice is what makes an MoE feel fast.

Expect "roughly" rather than exactly — routing and gathering the expert weights cost something, and the Mixtral paper flags that overhead explicitly.

When it doesn't fit

The instinct is to offload whole layers to system RAM. For an MoE that is the wrong cut. The experts are nearly all of the weight but only a small slice of the per-token work, while attention is small and touched by every single token. So move the experts and keep attention on the card.

# llama.cpp — keep expert tensors in system RAM llama-server -m model.gguf --n-cpu-moe 10 # --cpu-moe moves them all; --n-cpu-moe N does the first N layers

One clean benchmark on an RTX 3090 went from 139.6 to 89.1 tokens/sec of generation while going from a 89.6K to the full 262K context — roughly two thirds of the speed for about three times the context.

Two caveats. Prompt processing degrades much harder than generation (about −66% against −36% in that same benchmark), because prefill is compute-bound rather than bandwidth-bound — long prompts are where offloading stings. And Ollama does not expose this yet; the feature request is still open.

Is it as good as a dense model of the same size?

No, and be wary of anyone who gives you a formula. The popular "geometric mean of total and active" rule has no traceable source and is badly wrong in both directions — it implies Mixtral 8x7B behaves like a 24B model, when Mistral reports it beating a 70B dense one.

The honest version: a 30B-A3B lands somewhere between its active and total size. Qwen never published a head-to-head against their own 32B dense model; the nearest first-party comparison is against their 14B, where the two trade blows.

So treat the total as your memory budget, the active count as your speed estimate, and the quality as something to measure on your own task rather than derive.

The rule

Size your hardware off the total. Expect the speed of the active part. If the name has an "A" number in it, that number is the read-per-token figure, not the download.

Related: quantization covers the bits-per-weight half of the memory question, and the KV-cache section on that same page covers the part that grows with your context window. Total parameters, bits per weight, and context are three separate lines on the same bill — and only the first is what an "A" number is talking about. Whether you need a local model at all is a prior question.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click