// not a trim
The instinct is that something drops the oldest few messages and slides the window forward. That is not what happens. From the docs:
compaction block, continuing the conversation from the summary forward.All of them. What the model has after that point is one compaction block containing a summary, and whatever came after it.
So a small detail is either in the paragraph or nowhere
A summariser keeps what looks important. The trouble is that importance is judged without knowing what you are about to ask next.
This is the whole failure mode. The agent is not confused and has not malfunctioned — it is working from an accurate, lossy description of a conversation it can no longer see.
When it fires
Compaction is opt-in and configurable, which means the defaults are worth knowing before a long run rather than after one.
That is a broad model list, so this is not an exotic edge case — but it is one provider's parameter rather than an industry-wide behaviour, and it is beta.
Two levers, and a trap in the first
Tell it what to keep. Custom instructions steer the summariser — for example, "Focus on preserving code snippets, variable names, and technical decisions."
Pin the recent messages. pause_after_compaction: true hands control back to you, so you can rebuild the message list as the compaction block plus whichever recent turns you want preserved word for word.
Two things that will surprise you later
Your token count is not your bill. Compaction runs an extra sampling pass, and the top-level input_tokens and output_tokens do not include it. To get what you were actually charged you have to sum across usage.iterations.
Thinking blocks do not cross the boundary. On Fable 5.1 and Mythos 5.1, thinking blocks from before a compaction block are not carried forward; re-inserting earlier turns afterwards means stripping thinking and redacted_thinking blocks, or setting the prefix-mismatch behaviour to drop them.
Related: the durable answer to all of this is not to depend on the window at all — give the agent a scratchpad it writes to and reads back, so the important things live somewhere a summariser cannot edit. What a context window really limits is the fact underneath this, and why 20 tool calls cost you 200 is the other thing a long conversation does to you.
Source: Claude Platform docs — Compaction (beta). Verified 2026-09-07.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
