// the default comes from your card
Ollama's context-length page states the rule directly:
The code behind it (server/routes.go) adds up total memory across every GPU, minus overhead, and uses slightly lower cut-offs "to account for small differences in the exact value":
Three details matter. It only applies when nothing else set num_ctx (the app, OLLAMA_CONTEXT_LENGTH, a Modelfile or the request). The window is capped at the model's own trained length. And if a model fails to load for lack of memory, Ollama steps the automatic window down ("reducing automatic context and retrying once": 256K to 32K, 32K to 4K), so a 24 GB+ card can still land on 4K. On a Mac, "VRAM" is Metal's recommended working set, a share of unified memory. The FAQ's older line, "By default, Ollama uses a context window size of 4096 tokens", matches the table only under 24 GiB.
The popular library models don't override it. On 2026-10-01 the params of qwen3-coder:30b, gpt-oss:20b, llama3.2, gemma3:12b and qwen3:8b set temperature, stop tokens and sampling, and no num_ctx. The server default is what you get.
What happens to a longer file
Paste a 9,000-token file into a 4,096-token window and the prompt doesn't fit. On the default path, Ollama's runner (llm/llama_server.go) shortens it and carries on:
How much survives? The limit comes from contextShiftPromptLimit, which "free[s] roughly half of the remaining context before generation needs the slot": numCtx - (numCtx - numKeep)/2. With the defaults (num_ctx 4096, num_keep 4, plus the BOS token) that is about 2,051 tokens: the first five tokens and the last ~2,046. So any single input over ~4,100 tokens loses more than half of itself, from the start. That is our arithmetic from the code, not a number Ollama publishes. Two things it does not mean: it doesn't keep "the last 4,000 tokens", and it doesn't promise your system prompt survives.
The function returns nil, so the HTTP response is a normal reply. No error reaches your app. The traces are a WARN line in the server log and, client-side, the token count: ollama run --verbose prints prompt eval count, and the API returns prompt_eval_count. A count that sits near 2,051 for a long input is the tell.
Where this breaks: it's "most library models", not all. deepseek2-family models don't context-shift and return "the prompt is longer than the context length currently available to the model". Models that run in llama.cpp's native chat mode get its "request (N tokens) exceeds the available context size" error. MLX models don't truncate. Chat history is trimmed differently (older messages are dropped first); this page is about one long input.
Same model, different window
Because the default follows the card, two people running the same model on the same prompt can be working with different windows. With neither of them setting a value:
So a teammate's copy can see parts of the file yours never received. That holds only for the defaults, and only if the bigger card's load didn't step down. The model, the weights and the prompt are identical. The difference is a number the server picked at startup and logged as vram-based default context.
How big should it be?
Ollama's own docs answer for the workloads that hit this hardest:
That's about 16× the small-card default (15.6×). The Claude Code integration page says the same: "For larger repositories, set the context length to 64k or higher."
The catch is in the same docs: "Setting a larger context length will increase the amount of memory required to run a model. Ensure you have enough VRAM available to increase the context length." On a card under 24 GB, a flat 64K may not fit next to the weights, and Ollama will offload to CPU or step down. So the advice is to trade memory on purpose: raise it as far as your card holds for the model you run. Quantization covers how the KV cache's memory grows with the window and how to shrink it.
Set it, then check it
Then confirm what actually loaded. ollama ps has a CONTEXT column (its header in cmd/cmd.go: NAME, ID, SIZE, PROCESSOR, CONTEXT, UNTIL). That column is the ground truth. It reflects the step-down, the model's cap and whatever setting won. Check PROCESSOR on the same line: if it no longer says 100% GPU, the bigger window pushed part of the model onto the CPU.
Related: context window limits covers what a window is in the first place; agent compaction covers what agents do when a long session outgrows it; are local LLMs worth it covers the bigger trade.
Sources: docs.ollama.com/context-length (tiers, the 64000 line, app slider, OLLAMA_CONTEXT_LENGTH, ollama ps); docs.ollama.com/faq ("By default, Ollama uses a context window size of 4096 tokens", num_ctx); docs.ollama.com/integrations/claude-code ("64k or higher"); docs.ollama.com/api/openai-compatibility (Modelfile num_ctx); Ollama v0.35.0 source: server/routes.go (VRAM tiers, 23/47 GiB), server/sched.go (OOM step-down, deepseek2 context shift), llm/llama_server.go (prompt truncation, half-window limit), cmd/cmd.go (ollama ps header), api/types.go (num_keep 4, prompt eval count), discover/gpu_info_darwin.m (recommendedMaxWorkingSetSize); llama.cpp b11081 tools/server/server-context.cpp (native-mode error); registry.ollama.ai params blobs. All verified 2026-10-01. Scope: local models on the default path, nothing set by the user.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
