// the default split
llama.cpp's multi-GPU guide lists the ways it can spread a model across cards, and says which one you get when you pass nothing:
layer (default) | Pipeline parallelism. Each GPU holds a contiguous slice of layers. The KV cache for layer l lives on the GPU that owns layer l.Recipe 1 in the same guide is "Default - pipeline parallel across all visible GPUs". So on llama.cpp, a second card gets used even when the model would fit on the first. (Ollama behaves differently; more on that below.)
Two cards, taking turns
The guide explains what the default split does to a stream of tokens:
A reply is written one token after another, and each new token has to pass through the first card's layers and then the second card's. For a single chat there is only ever one token in flight, so while card one works card two waits, and then the other way round:
Each card now holds half the model, but the work per token is the same, so the reply comes out at roughly the pace one card would manage on its own. What you gained is room: memory for a bigger model or a longer context. What you did not gain is reply speed.
Ollama keeps it on one card, by default
Ollama's FAQ answers "How does Ollama load models on multiple GPUs?" directly:
Two things to keep straight. It is a default: the OLLAMA_SCHED_SPREAD setting ("Always schedule model across all GPUs") overrides it. And the second card isn't wasted: it stays free for another model. vLLM's docs give the same advice for serving: "if the model fits on a single GPU, distributed inference is probably unnecessary. Run inference on that GPU."
Don't carry Ollama's behaviour over to llama.cpp. llama.cpp's own default does split across every visible card, even when the model fits.
Where two cards do go faster
"Requires many tokens to scale well" cuts both ways. When lots of tokens go through together, both cards stay busy and the default split pays off. That happens when a long prompt is read in one go, or when many users are served at once. The pull request that added pipeline parallelism (llama.cpp #6017, a 7B model at F16 on NVIDIA A100 80GB PCIe cards) shows both sides:
So: long prompts and many users can go faster on two cards; the reply itself doesn't. This is a 2024 benchmark on datacenter cards, quoted for its shape, not as a prediction for your rig.
The mode that is aimed at reply speed is tensor parallelism, which splits every layer across the cards so they work on the same token at once. llama.cpp marks it EXPERIMENTAL: it "is much more bottlenecked by the GPU interconnect speed", "Performance should be good for multiple NVIDIA GPUs using the CUDA backend, no guarantees otherwise", and "--split-mode tensor is not implemented for all architectures" (a long list of MoE and state-space models fail at startup). The guide's troubleshooting table even has a row for "Performance is worse with multi-GPU than single-GPU", with the interconnect as the usual cause.
The verdict
The case for a second card is the first line of llama.cpp's own "When you need multi-GPU": "The model doesn't fit in a single GPU's VRAM ... Otherwise part of the model will need to be run off of the comparatively slower system RAM." If your model was spilling into system memory, moving that part onto a second card can be a real speed-up. If it already fits, part one's rule applies: reply speed follows the card's memory speed, so a faster card is what moves it.
The guide's second reason, "You want more throughput", is real too. It is the many-tokens case above: long prompts, many users, batch jobs. None of that is one person waiting on one reply.
Related: are local LLMs worth it covers whether to run models at home at all; quantization is the other way to make a model fit on the card you have; MoE offload covers what to move off the card when it doesn't fit.
Sources: llama.cpp docs/multi-gpu.md (split modes, pipeline vs tensor note, When you need multi-GPU, Troubleshooting); Ollama FAQ, "How does Ollama load models on multiple GPUs?"; ollama envconfig (OLLAMA_SCHED_SPREAD); vLLM docs, Parallelism and Scaling; llama.cpp PR #6017 benchmark tables. All verified 2026-09-26. Scope: one chat, the default split, a model that already fits.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
