Context Compaction
Summarize, then keep going
A technique that lets an agent keep going when a conversation nears the context window limit, by replacing older history with a summary or clearing old tool results.
A long-running agent's history grows every turn; tool results in particular (file contents, command output) fill the window fast. When the window is full you have two options: stop, or make room. Compaction is making room.
There are two basic methods: - Summarization (compaction): the model summarizes older turns, the summary replaces that history, and work continues in a fresh, smaller context. It's lossy: details survive only as well as the summary keeps them. - Tool-result clearing: results of old tool calls are removed and replaced with a marker saying the content was cleared. Anthropic calls this one of the safest, lightest-touch forms of compaction: the model can simply call the tool again if it needs the data.
In Claude Code this is automatic: as it nears the limit it first
clears older tool outputs, then summarizes the conversation if needed
(auto-compact). You can trigger it by hand with a focus, like
/compact focus on the auth bug fix; set the automatic threshold with
/autocompact 500k; and see what's filling the window with
/context. Codex CLI has a /compact command too.
The Claude API offers two beta features: the compact_20260112
strategy writes a server-side summary once input tokens cross a
threshold (default 150,000, minimum 50,000 tokens), and
clear_tool_uses_20250919 clears old tool results (context editing).
Like the minutes of a long meeting. Nobody can reread every sentence of a three-hour discussion; a secretary writes a one-page summary of "decisions made, open issues, who owns what," and you start the second half from that page.
Good minutes save the meeting; bad minutes silently erase a critical decision made in the first hour. That's exactly the risk of compaction.
You're doing a big refactor in Claude Code. For two hours the agent has read dozens of files and run tests; the window is almost full.
1. The harness first clears old tool outputs — file contents and test
logs it has long since processed.
2. If that's not enough it summarizes the conversation: which files
changed, which tests are still failing, what's next.
3. The system prompt and CLAUDE.md reload; the agent continues from
the summary.
One risk: the "don't touch the legacy/ folder" rule you gave at the
very start might get lost in the summary. That's why you put rules like
that in CLAUDE.md, or give a focus before a key step:
/compact keep the legacy folder rule.
import anthropic
client = anthropic.Anthropic()
messages = []
def chat(user_message: str) -> str:
messages.append({"role": "user", "content": user_message})
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5-5",
max_tokens=16000,
messages=messages,
context_management={
"edits": [{
"type": "compact_20260112",
# Default 150_000; minimum 50_000
"trigger": {"type": "input_tokens", "value": 120_000},
# Optional: fully replaces the default summary prompt
"instructions": "Keep open bugs, modified files and user decisions.",
}]
},
)
# Append the WHOLE content, not just the text: the compaction
# block replaces the old history on the next request.
messages.append({"role": "assistant", "content": response.content})
return next(b.text for b in response.content if b.type == "text")response = client.beta.messages.create(
betas=["context-management-2025-06-27"],
model="claude-opus-5-5",
max_tokens=16000,
tools=tools,
messages=messages,
context_management={
"edits": [{
"type": "clear_tool_uses_20250919",
"trigger": {"type": "input_tokens", "value": 30_000},
"keep": {"type": "tool_uses", "value": 5}, # keep the last 5 calls
"clear_at_least": {"type": "input_tokens", "value": 5_000},
"exclude_tools": ["web_search"], # never clear these
}]
},
)- Long agent tasks that could overflow the window — hours of coding, research, data processing
- Workflows where tool results pile up fast — reading files, searching, inspecting logs
- Answer quality has started to drop as the conversation grows — a smaller context often performs better
- You need to stay in the same session but don't need every detail
- Short conversations — if the window will never fill, summarizing only loses information
- Work where every detail must be preserved verbatim (legal text, audit trails) — keep the full record separately
-
Switching to a new, unrelated task — start a clean session instead (
/clearin Claude Code)
Assuming the summary is lossless
A summary keeps what the model deemed important. Detailed instructions from early in the conversation may be lost. Put durable rules in the system prompt or CLAUDE.md rather than trusting the summary.
Dropping the compaction block
With the Claude API, if you append only the response text to history, the compaction block is lost and the compaction state silently breaks. Send the whole response.content back.
Forgetting the cache cost
Removing part of the history or replacing it with a summary invalidates the prompt cache from that point on. When clearing tool results, use clear_at_least so each clearing frees a meaningful amount.
Thrashing on one giant output
If a single file or tool output takes up most of the window, it refills right after every summary. Claude Code stops after a few attempts and shows an error; the fix is to trim the output or read it in pieces.