toolcall() ← all concepts

// concept · local models

Move the experts, not the layers

Your model doesn't fit on the card, so you push half of it into system RAM and it slows to a crawl. The offload wasn't the mistake. The half you picked was.

// the wrong cut

The default move is to keep some layers on the GPU and put the rest in system RAM. For a dense model that is a reasonable trade. For a mixture of experts it is close to the worst available one, because a layer is not one kind of thing:

# inside one layer attention small · runs for every token experts enormous · a handful run per token

Offloading whole layers moves both. You pay to ship attention across the bus and get almost no memory back for it, because attention was never the thing filling your card.

The cut that actually works

Move only the expert tensors, and leave attention where it is. In llama.cpp there is a flag for exactly this, added in PR #15077 so people would stop hand-writing tensor regexes:

# keep the MoE weights of the first N layers in the CPU llama-server -m model.gguf --n-cpu-moe 10 # --cpu-moe moves ALL expert weights instead

The asymmetry is the whole reason it works. Experts are the overwhelming majority of the weights and a small fraction of the per-token work; attention is small and touched by every single token. So the experts are the cheap thing to move and attention is the expensive thing to move, and the flag lets you separate them.

Ollama does not expose this. The feature request (issue #11772) is open with no maintainer response. If you are on Ollama, this particular lever isn't available to you yet — that is a llama.cpp capability, not a general one.

What it costs

One clean RTX 3090 benchmark, same model and quantization on both rows:

fully on GPU 139.6 tok/s · 89.6K context --n-cpu-moe 10 89.1 tok/s · 262K context (full)

About two thirds of the generation speed for roughly three times the usable context. Whether that is a good trade depends entirely on what you were short of.

Ignore the 4.87× figure. A widely-circulated speedup number for expert offload is confounded, and its own author says the 11 tok/s baseline "isn't GPU inference" — it was Windows driver paging. It measures escaping a broken configuration, not a like-for-like gain.

The catch that makes it a decision

Generation loses about a third. Prompt processing loses about two thirds. In the same benchmark, prefill degraded roughly 66% against generation's 36%, and the reason is structural rather than incidental.

Generating runs one token at a time, so it is dominated by the time spent pulling weights out of memory. Reading a prompt runs the whole thing at once, so it is dominated by arithmetic. Moving weights to the slower side of the bus hurts the arithmetic-bound half far more than the bandwidth-bound one.

# so, in practice chatting, short prompts good trade long documents, big pastes poor trade

This is the part most write-ups leave out, and it is the part that decides whether the flag helps you. Offload is not free; it is cheap in one direction and expensive in the other.

The rule

Offload experts, not layers — then check whether your prompts are short. If you mostly chat, you are trading a third of your generation speed for a model that fits and a context that doesn't. If you mostly paste in long documents, you are paying at the worst possible point.

This is the third of three: what an expert actually is and why every lab switched, then why the total is what you have to store, and now what to do when it won't fit. Related: bits per weight is the other lever on the same bill, and its KV-cache section covers the part that grows with your context window.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click