AI Atlas
EN TR
Intermediate · ~2 min read #context #prompt #agent

Context Engineering

Designing what the model sees

The discipline of deliberately deciding what goes into the model's context window on each call, how much of it, and in what order.

WHAT GOES INTO THE CONTEXT WINDOW?DUMP EVERYTHINGsystem60 tools400-page manualhistory + raw outputswindow full · slow · costly · context rotWINDOW LIMITENGINEEREDsystem8 toolsrelevant partmemorylean historyheadroom = focused attentionthe smallest set of high-signal tokensjust-in-time retrieval · clear old outputs · compact · take notes · delegateit's not the size of the window — it's what you put in it
Definition

Prompt engineering asks "how do I write the instruction?" Context engineering asks a bigger question: what should all the tokens the model sees on this call be? In an agent, context is made of:

- System prompt — role, rules, working style - Tool definitions — names, descriptions, schemas - Retrieved documents — RAG results, files that were read - Memory — persistent notes like CLAUDE.md, summaries of past sessions - Message history — user messages, model replies, and the fastest growing item: tool results

Anthropic's September 2025 post "Effective context engineering for AI agents" states the guiding principle: find the smallest possible set of high-signal tokens that maximizes the likelihood of the desired outcome. The reason is context rot: as the window fills up, the model's ability to accurately recall what's in it declines. Because in a transformer every token attends to every other token (n² pairwise relationships for n tokens), attention is a finite budget, and every unnecessary token spends some of it.

Practical techniques: keep tools few and clear; instead of loading everything up front, fetch "just in time" through lightweight references like file paths or queries; clear old tool results; summarize when the window fills up (compaction); write durable facts to notes outside the window; hand sprawling research to subagents that have their own context.

Analogy

Like briefing a consultant. Drop a 300-page binder on their desk and say "it's all in there," and the three pages that matter get lost. Give them a crisp two-page summary, the location of the relevant contract and a note saying "ask the archive for X if needed," and they work faster and more accurately.

Context engineering is designing those two pages and the path to the archive. The size of the desk (the window) matters, but what you put on it matters more.

Real-world example

A typical scenario: a support agent embeds a 400-page product manual in its system prompt, ships with 60 tool definitions, and dumps every customer record into history as raw JSON. Result: a slow, expensive agent that often picks the wrong tool.

The redesign: - The system prompt shrinks to short, clear rules, and its static part is cached (prompt caching). - 60 tools become 8 well-described ones. - The manual leaves the prompt; a search_docs tool takes its place and the agent pulls only the relevant section. - The customer's past conversations live in a one-paragraph note. - Old tool results are cleared automatically past a threshold.

The model didn't change; what it sees on every turn did.

Code examples
Anthropic API · assemble context piece by piece, measure the budget python
import anthropic

client = anthropic.Anthropic()

SYSTEM = open("prompts/support.md").read()     # short, stable rules
tools = [search_docs_tool, get_ticket_tool]    # few, clear tools
notes = open("memory/customer_42.md").read()   # notes from past sessions

messages = [{
    "role": "user",
    "content": f"<notes>\n{notes}\n</notes>\n\n{question}",
}]

# Measure how many tokens the context takes before sending
count = client.messages.count_tokens(
    model="claude-opus-5-5",
    system=SYSTEM,
    tools=tools,
    messages=messages,
)
print("input tokens:", count.input_tokens)

response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    # The stable part comes first and is cached
    system=[{"type": "text", "text": SYSTEM,
             "cache_control": {"type": "ephemeral"}}],
    tools=tools,
    messages=messages,
)
CLAUDE.md · keep durable rules out of chat history markdown
# Project notes

- Package manager: pnpm. Tests: `pnpm test`.
- API layer in `src/api/`, business rules in `src/domain/`.
- Never edit the database schema directly; write a migration.

## Compact Instructions

When summarizing, always keep open bugs, the list of
modified files and decisions the user has made.
When to use
  • Any agent that runs over multiple turns — context grows every turn and degrades if not designed
  • The agent “knows” the right information yet acts wrong — the information is probably lost in noise
  • Cost and latency are high — shrinking context lowers both
  • When combining many tools, documents or memory sources
When not to use
  • A short, one-off call — a well-written prompt is enough
  • The problem is capability, not context — rearranging context won't rescue a task the model can't do
  • Optimizing without measuring — look at token counts and success rates first, then trim
Common pitfalls

Treating a big window as a dumping ground

A 1M-token window doesn't mean you should fill 1M tokens. Because of context rot, recall degrades as the window fills; every item has to earn its place.

Neglecting tool descriptions

Tool definitions are context too, and they're resent every turn. Overlapping tools with vague descriptions push the model toward the wrong choice.

Breaking your own cache

Putting something that changes per request, like a timestamp, at the top of the system prompt defeats prompt caching. Stable content first, variable content last.

Leaving durable rules in chat history

A rule stated early in the conversation can be lost during summarization. Put anything that must always hold in the system prompt or a file like CLAUDE.md.