Agent Memory
Short-term and long-term memory for agents
How an agent carries information beyond a single conversation: short-term memory in the context window, and long-term memory kept in files, memory tools or vector databases.
An LLM has no memory of its own; every request starts from scratch. The only reason it seems to "remember" is that the harness puts the relevant information back into the context window on every turn. Agent memory is the design of what gets put back, and how.
Short-term memory = the current context: conversation history, tool outputs, files read. Fast but finite; when the window fills up it gets compacted or truncated, and it disappears when the session ends.
Long-term memory survives across sessions. Three common forms:
- Instruction files: plain Markdown files like CLAUDE.md and
AGENTS.md, read automatically at the start of every session.
Written by humans, kept in git, shared by the team.
- Memory tools: files the agent reads and writes itself. The
memory tool in the Claude API lets the model create and edit files
under a /memories directory; Claude Code's auto memory keeps a
per-project directory indexed by a MEMORY.md file.
- Vector databases: for large volumes of past notes or
conversations; semantic search pulls only the relevant pieces into
context (the same mechanics as RAG).
The hard question isn't "how do we store it" but "what do we store". Good candidates: things not already in the code or git history that you keep having to re-explain — build commands, team conventions, decisions and their rationale. Bad candidates: anything readable from the code, transient state, and unverified guesses.
Like a shift handover in a hospital. What's in the doctor's head (short-term memory) leaves when the shift ends. The patient's chart (long-term memory) stays, and the next doctor reads it first. But if someone writes "fever down" and forgets to update it, the next doctor makes a bad call on stale information. Maintaining memory matters as much as having it.
A team is working with an agent on a Laravel project. They're tired of
telling it the same things every session: tests run with
php artisan test, there's no database, content lives as YAML files in
content/.
They write these into a CLAUDE.md at the project root (or AGENTS.md
if they use more than one agent). Every session now starts with that
knowledge. And when a developer corrects the agent ("no, we use Alpine
here"), auto memory adds the correction to its notes, so the same
mistake doesn't come back next session.
Three months later something breaks: the file still lists an old deploy script, and the agent confidently tries to run it. Lesson: memory files need review, just like code.
# Project notes
See @README.md for an overview.
## Commands
- Tests: `php artisan test`
- Front-end build: `npm run build`
## Conventions
- No database; content lives in YAML files under `content/`.
- Use Alpine.js for interactivity; don't add a new JS framework.
- Write commit messages in English.import anthropic
client = anthropic.Anthropic()
# The model requests view/create/edit operations under /memories;
# your application executes them against storage you control.
message = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=2048,
tools=[{"type": "memory_20250818", "name": "memory"}],
messages=[
{"role": "user", "content": "Let's pick up yesterday's support ticket."}
],
)- You keep re-explaining the same project facts every session (instruction file)
- Carrying progress across long tasks that span multiple sessions
- Persisting user preferences and past corrections
- Past notes and conversations are too large to fit in context (vector search)
- Information already readable from code, tests or git history — it creates two sources of truth
- Secrets (passwords, tokens) — memory files end up in context and often in the repo
- One-off tasks — short-term context is enough
Stale memory
Memory doesn't update itself. If a changed command or an abandoned decision stays in the file, the agent applies it as current truth. Review memory files regularly and delete what's outdated.
Turning memory into a junk drawer
Saving everything means spending tokens at the start of every session and drowning the important rule in noise. Even Claude Code's auto memory loads only the first 200 lines or 25 KB of MEMORY.md per session. Keep it tight.
Trusting memory blindly
A memory is a claim, not evidence. A note saying 'function X lives in file Y' is wrong once the file moves. Before acting on something important, the agent should check the memory against the current state.