toolcall() ← all concepts

// concept · workflows

Under 24GB VRAM? Ollama gives you 4K

Pull a model that supports a 128K or 256K context, run it on a 16 GB card, and Ollama will still give it a 4,096-token window. Nothing in the model decides that. The server picks the default from the video memory it detects, and anything longer than the window is cut from the front, quietly, on most library models. Checked against Ollama v0.35.0 (2026-09-28).

// the default comes from your card

Ollama's context-length page states the rule directly:

Ollama defaults to the following context lengths based on VRAM: < 24 GiB VRAM: 4k context; 24-48 GiB VRAM: 32k context; >= 48 GiB VRAM: 256k context

The code behind it (server/routes.go) adds up total memory across every GPU, minus overhead, and uses slightly lower cut-offs "to account for small differences in the exact value":

total VRAM (all GPUs, minus overhead) >= 47 GiB 262144 >= 23 GiB 32768 otherwise 4096 ← includes CPU-only machines

Three details matter. It only applies when nothing else set num_ctx (the app, OLLAMA_CONTEXT_LENGTH, a Modelfile or the request). The window is capped at the model's own trained length. And if a model fails to load for lack of memory, Ollama steps the automatic window down ("reducing automatic context and retrying once": 256K to 32K, 32K to 4K), so a 24 GB+ card can still land on 4K. On a Mac, "VRAM" is Metal's recommended working set, a share of unified memory. The FAQ's older line, "By default, Ollama uses a context window size of 4096 tokens", matches the table only under 24 GiB.

The popular library models don't override it. On 2026-10-01 the params of qwen3-coder:30b, gpt-oss:20b, llama3.2, gemma3:12b and qwen3:8b set temperature, stop tokens and sampling, and no num_ctx. The server default is what you get.

What happens to a longer file

Paste a 9,000-token file into a 4,096-token window and the prompt doesn't fit. On the default path, Ollama's runner (llm/llama_server.go) shortens it and carries on:

// keep num_keep tokens from the front, then the tail truncated = append(truncated, tokens[:nKeep]...) truncated = append(truncated, tokens[nKeep+discard:]...) slog.Warn("truncating input prompt", ...) return truncated, nil

How much survives? The limit comes from contextShiftPromptLimit, which "free[s] roughly half of the remaining context before generation needs the slot": numCtx - (numCtx - numKeep)/2. With the defaults (num_ctx 4096, num_keep 4, plus the BOS token) that is about 2,051 tokens: the first five tokens and the last ~2,046. So any single input over ~4,100 tokens loses more than half of itself, from the start. That is our arithmetic from the code, not a number Ollama publishes. Two things it does not mean: it doesn't keep "the last 4,000 tokens", and it doesn't promise your system prompt survives.

The function returns nil, so the HTTP response is a normal reply. No error reaches your app. The traces are a WARN line in the server log and, client-side, the token count: ollama run --verbose prints prompt eval count, and the API returns prompt_eval_count. A count that sits near 2,051 for a long input is the tell.

Where this breaks: it's "most library models", not all. deepseek2-family models don't context-shift and return "the prompt is longer than the context length currently available to the model". Models that run in llama.cpp's native chat mode get its "request (N tokens) exceeds the available context size" error. MLX models don't truncate. Chat history is trimmed differently (older messages are dropped first); this page is about one long input.

Same model, different window

Because the default follows the card, two people running the same model on the same prompt can be working with different windows. With neither of them setting a value:

16 GB card 4,096 tokens a 9K file: the start is gone 24 GB card 32,768 tokens 8× · the whole file fits 48 GB+ 262,144 tokens 64×

So a teammate's copy can see parts of the file yours never received. That holds only for the defaults, and only if the bigger card's load didn't step down. The model, the weights and the prompt are identical. The difference is a number the server picked at startup and logged as vram-based default context.

How big should it be?

Ollama's own docs answer for the workloads that hit this hardest:

Tasks which require large context like web search, agents, and coding tools should be set to at least 64000 tokens.

That's about 16× the small-card default (15.6×). The Claude Code integration page says the same: "For larger repositories, set the context length to 64k or higher."

The catch is in the same docs: "Setting a larger context length will increase the amount of memory required to run a model. Ensure you have enough VRAM available to increase the context length." On a card under 24 GB, a flat 64K may not fit next to the weights, and Ollama will offload to CPU or step down. So the advice is to trade memory on purpose: raise it as far as your card holds for the model you run. Quantization covers how the KV cache's memory grows with the window and how to shrink it.

Set it, then check it

# 1. the app Settings → context length slider # 2. the server, for every client OLLAMA_CONTEXT_LENGTH=32768 ollama serve # 3. one model, for OpenAI-style apps (the OpenAI API has no field for it) FROM qwen3:8b PARAMETER num_ctx 32768 $ ollama create qwen3-32k -f Modelfile # 4. per request, native API "options": { "num_ctx": 32768 }

Then confirm what actually loaded. ollama ps has a CONTEXT column (its header in cmd/cmd.go: NAME, ID, SIZE, PROCESSOR, CONTEXT, UNTIL). That column is the ground truth. It reflects the step-down, the model's cap and whatever setting won. Check PROCESSOR on the same line: if it no longer says 100% GPU, the bigger window pushed part of the model onto the CPU.

$ ollama ps NAME ID SIZE PROCESSOR CONTEXT UNTIL qwen3:8b … … 100% GPU 4096 … (illustration · run it on your own machine)

Related: context window limits covers what a window is in the first place; agent compaction covers what agents do when a long session outgrows it; are local LLMs worth it covers the bigger trade.

Sources: docs.ollama.com/context-length (tiers, the 64000 line, app slider, OLLAMA_CONTEXT_LENGTH, ollama ps); docs.ollama.com/faq ("By default, Ollama uses a context window size of 4096 tokens", num_ctx); docs.ollama.com/integrations/claude-code ("64k or higher"); docs.ollama.com/api/openai-compatibility (Modelfile num_ctx); Ollama v0.35.0 source: server/routes.go (VRAM tiers, 23/47 GiB), server/sched.go (OOM step-down, deepseek2 context shift), llm/llama_server.go (prompt truncation, half-window limit), cmd/cmd.go (ollama ps header), api/types.go (num_keep 4, prompt eval count), discover/gpu_info_darwin.m (recommendedMaxWorkingSetSize); llama.cpp b11081 tools/server/server-context.cpp (native-mode error); registry.ollama.ai params blobs. All verified 2026-10-01. Scope: local models on the default path, nothing set by the user.

One concept a week. Free.

The deeper, copy-paste version of each ToolCall short — in your inbox.

// total: 0.00 · spam: void · unsubscribe: one click