// what it actually is
The context window is the maximum number of tokens the model can attend to in one call — the prompt it reads plus the answer it writes. It's per-request working memory, not long-term recall. How many words fit in a token depends on the model's tokenizer: on Claude Opus 4.7 and later, 1M tokens is roughly 555k words (Anthropic's figure), where earlier Claude models fit about 750k. So the count is the model's, not yours.
Input and output share it
It's one budget, two claimants. A big input leaves less room for the reply — pour enough into the box and the model runs out of space to answer. Push past the edge on a hosted API and the call doesn't truncate quietly; it errors. (A local runtime can differ: on its default path Ollama cuts the prompt from the start with a server-log warning, "truncating input prompt", and no error reaches your app; see Ollama's 4K default.)
It resets every call
The model is stateless. Between turns it remembers nothing. The "it knows what we talked about" feeling is the app replaying the whole conversation on every request — re-sending past turns as input each time. Stop re-sending them and the memory is gone.
Sources: Anthropic — Context windows · OpenAI — Conversation state (stateless) · OpenAI — What are tokens · Liu et al. — Lost in the Middle
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
