// the wrong cut
The default move is to keep some layers on the GPU and put the rest in system RAM. For a dense model that is a reasonable trade. For a mixture of experts it is close to the worst available one, because a layer is not one kind of thing:
Offloading whole layers moves both. You pay to ship attention across the bus and get almost no memory back for it, because attention was never the thing filling your card.
The cut that actually works
Move only the expert tensors, and leave attention where it is. In llama.cpp there is a flag for exactly this, added in PR #15077 so people would stop hand-writing tensor regexes:
The asymmetry is the whole reason it works. Experts are the overwhelming majority of the weights and a small fraction of the per-token work; attention is small and touched by every single token. So the experts are the cheap thing to move and attention is the expensive thing to move, and the flag lets you separate them.
What it costs
One clean RTX 3090 benchmark, same model and quantization on both rows:
About two thirds of the generation speed for roughly three times the usable context. Whether that is a good trade depends entirely on what you were short of.
The catch that makes it a decision
Generation loses about a third. Prompt processing loses about two thirds. In the same benchmark, prefill degraded roughly 66% against generation's 36%, and the reason is structural rather than incidental.
Generating runs one token at a time, so it is dominated by the time spent pulling weights out of memory. Reading a prompt runs the whole thing at once, so it is dominated by arithmetic. Moving weights to the slower side of the bus hurts the arithmetic-bound half far more than the bandwidth-bound one.
This is the part most write-ups leave out, and it is the part that decides whether the flag helps you. Offload is not free; it is cheap in one direction and expensive in the other.
The rule
Offload experts, not layers — then check whether your prompts are short. If you mostly chat, you are trading a third of your generation speed for a model that fits and a context that doesn't. If you mostly paste in long documents, you are paying at the worst possible point.
This is the third of three: what an expert actually is and why every lab switched, then why the total is what you have to store, and now what to do when it won't fit. Related: bits per weight is the other lever on the same bill, and its KV-cache section covers the part that grows with your context window.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
