AI Atlas
EN TR
Advanced · ~2 min read #compaction #context #long-running-agents

Context Compaction

Summarize, then keep going

A technique that lets an agent keep going when a conversation nears the context window limit, by replacing older history with a summary or clearing old tool results.

CONTEXT COMPACTIONBEFORE · NEARLY FULLsystem + CLAUDE.mdusertool_result · 2.000 linestool_result · npm test logassistanttool_result · greptool_result · fileassistanttool_result · file/compactclear + summarizeAFTER · ROOM TO WORKsystem + CLAUDE.mduserassistantSUMMARY: decisions, changedfiles, open bugsfree spaceold tool outputs eat the windowdurable rules reloadsummaries are lossy — keep must-follow rules in a file
Definition

A long-running agent's history grows every turn; tool results in particular (file contents, command output) fill the window fast. When the window is full you have two options: stop, or make room. Compaction is making room.

There are two basic methods: - Summarization (compaction): the model summarizes older turns, the summary replaces that history, and work continues in a fresh, smaller context. It's lossy: details survive only as well as the summary keeps them. - Tool-result clearing: results of old tool calls are removed and replaced with a marker saying the content was cleared. Anthropic calls this one of the safest, lightest-touch forms of compaction: the model can simply call the tool again if it needs the data.

In Claude Code this is automatic: as it nears the limit it first clears older tool outputs, then summarizes the conversation if needed (auto-compact). You can trigger it by hand with a focus, like /compact focus on the auth bug fix; set the automatic threshold with /autocompact 500k; and see what's filling the window with /context. Codex CLI has a /compact command too.

The Claude API offers two beta features: the compact_20260112 strategy writes a server-side summary once input tokens cross a threshold (default 150,000, minimum 50,000 tokens), and clear_tool_uses_20250919 clears old tool results (context editing).

Analogy

Like the minutes of a long meeting. Nobody can reread every sentence of a three-hour discussion; a secretary writes a one-page summary of "decisions made, open issues, who owns what," and you start the second half from that page.

Good minutes save the meeting; bad minutes silently erase a critical decision made in the first hour. That's exactly the risk of compaction.

Real-world example

You're doing a big refactor in Claude Code. For two hours the agent has read dozens of files and run tests; the window is almost full.

1. The harness first clears old tool outputs — file contents and test logs it has long since processed. 2. If that's not enough it summarizes the conversation: which files changed, which tests are still failing, what's next. 3. The system prompt and CLAUDE.md reload; the agent continues from the summary.

One risk: the "don't touch the legacy/ folder" rule you gave at the very start might get lost in the summary. That's why you put rules like that in CLAUDE.md, or give a focus before a key step: /compact keep the legacy folder rule.

Code examples
Claude API · server-side summarization past a threshold (beta) python
import anthropic

client = anthropic.Anthropic()
messages = []

def chat(user_message: str) -> str:
    messages.append({"role": "user", "content": user_message})

    response = client.beta.messages.create(
        betas=["compact-2026-01-12"],
        model="claude-opus-5-5",
        max_tokens=16000,
        messages=messages,
        context_management={
            "edits": [{
                "type": "compact_20260112",
                # Default 150_000; minimum 50_000
                "trigger": {"type": "input_tokens", "value": 120_000},
                # Optional: fully replaces the default summary prompt
                "instructions": "Keep open bugs, modified files and user decisions.",
            }]
        },
    )

    # Append the WHOLE content, not just the text: the compaction
    # block replaces the old history on the next request.
    messages.append({"role": "assistant", "content": response.content})
    return next(b.text for b in response.content if b.type == "text")
Claude API · clearing old tool results (context editing, beta) python
response = client.beta.messages.create(
    betas=["context-management-2025-06-27"],
    model="claude-opus-5-5",
    max_tokens=16000,
    tools=tools,
    messages=messages,
    context_management={
        "edits": [{
            "type": "clear_tool_uses_20250919",
            "trigger": {"type": "input_tokens", "value": 30_000},
            "keep": {"type": "tool_uses", "value": 5},          # keep the last 5 calls
            "clear_at_least": {"type": "input_tokens", "value": 5_000},
            "exclude_tools": ["web_search"],                   # never clear these
        }]
    },
)
When to use
  • Long agent tasks that could overflow the window — hours of coding, research, data processing
  • Workflows where tool results pile up fast — reading files, searching, inspecting logs
  • Answer quality has started to drop as the conversation grows — a smaller context often performs better
  • You need to stay in the same session but don't need every detail
When not to use
  • Short conversations — if the window will never fill, summarizing only loses information
  • Work where every detail must be preserved verbatim (legal text, audit trails) — keep the full record separately
  • Switching to a new, unrelated task — start a clean session instead (/clear in Claude Code)
Common pitfalls

Assuming the summary is lossless

A summary keeps what the model deemed important. Detailed instructions from early in the conversation may be lost. Put durable rules in the system prompt or CLAUDE.md rather than trusting the summary.

Dropping the compaction block

With the Claude API, if you append only the response text to history, the compaction block is lost and the compaction state silently breaks. Send the whole response.content back.

Forgetting the cache cost

Removing part of the history or replacing it with a summary invalidates the prompt cache from that point on. When clearing tool results, use clear_at_least so each clearing frees a meaningful amount.

Thrashing on one giant output

If a single file or tool output takes up most of the window, it refills right after every summary. Claude Code stops after a few attempts and shows an error; the fix is to trim the output or read it in pieces.