Anatomy of an AI Harness: What Surrounds the Model
The layer around the model in tools like Claude Code, Codex CLI, Cursor and Gemini CLI: agent loop, tools, context management, subagents, skills, hooks, permissions, sandboxing, memory, and a sub-70-line harness you can build yourself.
TL;DR
When you open Claude Code, Codex CLI or Cursor, you are not talking to a "bare" language model. A thick software layer wraps the model and turns it into an agent: it calls the model in a loop, defines its tools, assembles and prunes its context, decides which commands may run, records the session and lets you rewind it. That layer is the agent harness.
A rough equation:
Agent = Model + Harness
Harness = loop + tools + context management + permissions/sandbox + extension points (skills, hooks, subagents, MCP) + memory + observabilityThis guide takes the harness apart piece by piece. Each section explains the concept first, then how popular tools implement it. At the end you'll find a comparison table of four popular harnesses, a working sub-70-line "build your own harness" example, and a checklist for choosing and configuring one.
Note: Every product-specific detail (command names, hook events, file names) was checked against official documentation as of October 2026. These tools ship weekly; details may change.
Model vs harness
Put the same model in two different harnesses and the results can differ dramatically. Not because the model got smarter or dumber, but because what it sees and what it can do changed:
- What does it see? System prompt, project instruction files, tool descriptions, previous tool results. All decided by the harness. Whatever lands in the context window is what the model reasons over.
- What can it do? Can it read files, edit them, run shell commands, reach the web? The toolset comes from the harness.
- When does it stop? How many turns it may take, how much it may spend, when it must ask the user: harness rules.
- What happens on failure? Does a failing test's output flow back to the model, or does the agent declare victory and stop? Feeding errors back is the harness's job.
The core idea: the model is the decision maker; the harness is the world in which those decisions are executed. Claude Code's documentation states the split explicitly: permission rules are enforced by Claude Code, not by the model. Instructions in CLAUDE.md shape what Claude tries to do, but they don't change what Claude Code allows.
So when an agent gives you poor results, check these before swapping the model: was the right information in context, were the right tools defined, were errors fed back?
The agent loop
Every harness has the same simple loop at its heart. The agent loop works like this:
1. User message + system prompt + tool definitions → send to model
2. Model responds:
a) Text + "done" → exit loop, show answer
b) One or more tool calls → go to 3
3. Harness checks each call (is it allowed?), executes it
4. Append tool results to the conversation → back to 1The model never executes anything itself; it only says "I want to call this tool with these arguments". In the Anthropic API that's a response with stop_reason: "tool_use" containing tool_use blocks. The harness runs the tool, sends the output back as a tool_result block, and the loop continues while stop_reason is "tool_use". The API-side mechanism is called function calling.
Stop conditions
How the loop ends matters as much as how it starts:
| Stop reason | What happens |
|---|---|
| Model returned text without tool calls | Normal finish, task considered done |
| Turn limit reached | Seatbelt against infinite loops |
| Budget limit reached | Cost ceiling hit |
| User interrupted (Esc / Ctrl+C) | Human intervention |
| A hook said "stop" | A deterministic rule kicked in |
SDKs expose these limits as parameters. The Claude Agent SDK (Python) offers max_turns (maximum agentic turns, i.e. tool-use round trips) and max_budget_usd (stops the query when the client-side cost estimate reaches that value). In the OpenAI Agents SDK, the loop ends when the model produces text output of the desired type with no tool calls; exceeding max_turns raises MaxTurnsExceeded.
Practical tip: Every agent that runs unattended (CI, cron) needs a turn and/or budget limit. Interactively, you are the limit; unattended, nothing else is.
Tools: the agent's hands
A harness's toolset almost always shares the same core:
| Tool family | Typical names | Purpose |
|---|---|---|
| Read | Read, Glob, Grep | Read files, find files, search content |
| Write | Edit, Write, apply_patch | Modify or create files |
| Shell | Bash, terminal | Tests, builds, git, package managers |
| Web | WebSearch, WebFetch | Docs, current information |
| Planning | todo/task lists | Track progress on long tasks |
| Delegation | Task/Agent | Spawn subagents |
MCP: extending the toolset from outside
When built-ins aren't enough, MCP (Model Context Protocol) steps in. An MCP server exposes a service (GitHub, a database, Sentry, a browser) as tools; the harness adds them to the model's tool list. All four major harnesses support MCP today: Claude Code, Codex CLI (a [mcp_servers.<name>] table in ~/.codex/config.toml, or codex mcp add), Cursor (mcp.json) and Gemini CLI (/mcp). See the MCP Server Catalog for ready-made servers and Build Your Own MCP Server to write one.
Tool descriptions are prompts too
The model knows a tool only through its name, description and parameter schema. A tool description is therefore an instruction the model reads. Two practical consequences:
- Bad description = wrong usage. "Reads a UTF-8 text file inside the project root; pass a path relative to the root" instead of "reads a file" noticeably cuts down on wrong-path attempts.
- Every tool costs context. Connect three MCP servers with 40 tools each and all 120 definitions ride along on every request. Disable servers you don't use; some harnesses (Claude Code, for example) also remove a tool from context entirely when a permission rule denies it by bare name.
Context engineering and management
Context engineering is the discipline of deciding what goes into the model's context window on every turn. Prompt engineering polishes a single message; context engineering manages the window's contents across a session of hundreds of turns.
The typical package a harness sends with each request:
[system prompt] ← the harness's own instructions
[tool definitions] ← built-in + MCP tools
[project instruction files] ← CLAUDE.md / AGENTS.md / GEMINI.md
[skill descriptions] ← name + description only
[conversation history] ← messages + tool calls + tool resultsSystem prompt
The harness's own system prompt tells the model who it is, how to use tools, when to ask. You rarely see it, but much of the agent's "personality" comes from it. If you're writing your own agent, the System Prompt Guide is a good start.
Project memory files
Every harness automatically adds a markdown file from your repo to the context:
| Harness | File | Behavior |
|---|---|---|
| Claude Code | CLAUDE.md (./, ./.claude/, ~/.claude/, CLAUDE.local.md) |
Files in and above the working directory load at launch; subdirectory files load as Claude reads files there. With no CLAUDE.md at all, AGENTS.md is read (v2.1.277+). |
| Codex CLI | AGENTS.md, AGENTS.override.md |
~/.codex first, then at most one file per directory from the project root down to the working directory; combined size capped at 32 KiB by default (project_doc_max_bytes). |
| Gemini CLI | GEMINI.md |
Global (~/.gemini/GEMINI.md), workspace, and "just-in-time": when a tool touches a directory, its GEMINI.md is scanned. /memory show prints the concatenated content. |
| Cursor | .cursor/rules/*.mdc, AGENTS.md |
Rules use description, globs, alwaysApply frontmatter; AGENTS.md is supported in the root and subdirectories. |
These files sit in context on every turn, so every line has a price. Claude Code's docs note that files over 200 lines consume more context and may reduce adherence. Rule of thumb: write what the model can't derive from the code (pitfalls, rationale, conventions that differ from defaults), not directory trees or dependency lists.
Compaction: summarizing history
A long session eventually fills the window. Context compaction frees space by replacing old conversation with a summary:
- Claude Code auto-compacts when approaching the context limit; trigger it manually with instructions, e.g.
/compact Focus on code samples and API usage. The project-rootCLAUDE.mdis re-read from disk and re-injected after compaction. - Codex CLI uses
/compact; Gemini CLI uses/compress. - You can hook into compaction:
PreCompact/PostCompactin Claude Code and Codex,PreCompressin Gemini CLI,preCompactin Cursor.
Compaction is lossy; the summary drops details. Having the agent write important decisions to a file (a plan file, notes) is more robust than trusting compaction.
Tool-result clearing
Tool results fill context faster than anything else: a 3,000-line log, a whole file, search hits. Once the model has processed them, it usually doesn't need them again. The Anthropic API offers context editing for this: the clear_tool_uses_20250919 strategy clears old tool results once context grows beyond a threshold you configure (beta, via the context-management-2025-06-27 header). In your own harness you can do something simpler: truncate tool output, return only the last N lines or first N characters.
Context budget
Treat context as a budget. Every item takes room: system prompt, tool definitions, instruction files, skill list, history. Codex makes this concrete: the initial skill list may use at most 2% of the context window (or 8,000 characters if the window size is unknown); beyond that, descriptions get shortened. For token-saving techniques see Token Reduction Techniques.
Subagents and parallelism
A subagent is a separate agent the main agent delegates a subtask to, running in its own context window. Its main benefit is context isolation more than parallelism: the subagent reads 50 files and scans 200 lines of search output, and only the summary comes back to the main conversation.
All four harnesses have them:
- Claude Code: defined as markdown files in
.claude/agents/(project) and~/.claude/agents/(user), each with its own system prompt, tool access and permissions. Ships with built-inExplore(read-only codebase exploration) andPlan(research during plan mode). - Codex CLI: parallel work is triggered by direct requests like "spawn three subagents: one for security, one for test gaps, one for maintainability".
- Gemini CLI:
.gemini/agents/*.md(project) and~/.gemini/agents/*.md(user); a subagent runs in a separate context loop. - Cursor: each subagent runs in its own context window and returns its result to the parent; three built-in subagents handle context-heavy operations.
Use them for: work with long output but a short result (codebase search, log analysis, independent research threads), and independent parallel checks (security / tests / style review).
Avoid them for: tasks needing frequent back-and-forth, phases sharing lots of context (plan → implement → test), small targeted changes. A subagent starts fresh and needs time to gather context, and each one spends its own tokens.
Skills: on-demand instructions
Agent skills are instruction packs made of a folder plus a SKILL.md file, optionally with scripts and reference files. The key idea is progressive disclosure:
At startup: only name + description in context (a few dozen tokens)
When needed: the full SKILL.md is loaded (hundreds to thousands of tokens)
If needed further: extra files in the skill folder are read / scripts runSo even with 50 skills installed, the context doesn't bloat; the model opens only what it needs. The format follows an open standard (agentskills.io) and is supported by Claude Code, Codex CLI, Gemini CLI (via the activate_skill tool) and Cursor.
Claude Code has two handy frontmatter fields:
---
name: deploy
description: Deploy the application to production
disable-model-invocation: true # Claude can't invoke it on its own; only you type /deploy
---disable-model-invocation: true matters for side-effecting workflows (deploy, commit, Slack messages): you don't want the agent deploying because the code "looks ready". Conversely, user-invocable: false turns a skill into background knowledge only the model uses.
For an introduction to writing skills, see the Claude Skills Guide.
Hooks: deterministic lifecycle automation
Writing "run the linter after every edit" in an instruction file is a request; the model sometimes forgets. Agent hooks are a guarantee: the harness runs your script at specific lifecycle points, every time.
| Harness | Before tool | After tool | Other notable events |
|---|---|---|---|
| Claude Code | PreToolUse |
PostToolUse |
SessionStart, UserPromptSubmit, PermissionRequest, Stop, SubagentStart/SubagentStop, PreCompact, SessionEnd |
| Codex CLI | PreToolUse |
PostToolUse |
SessionStart, UserPromptSubmit, PermissionRequest, Stop, SubagentStart/SubagentStop, PreCompact/PostCompact |
| Gemini CLI | BeforeTool |
AfterTool |
SessionStart, BeforeAgent/AfterAgent, BeforeModel/AfterModel, BeforeToolSelection, PreCompress |
| Cursor | beforeShellExecution, beforeMCPExecution |
afterShellExecution, afterFileEdit |
sessionStart, beforeSubmitPrompt, preCompact, stop |
Example: run a linter after every Edit/Write in Claude Code (.claude/settings.json):
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write",
"hooks": [
{ "type": "command", "command": "/path/to/lint-check.sh" }
]
}
]
}
}In Claude Code a hook script receives JSON on stdin; exit code 2 blocks the action and stderr goes back to the model as feedback. A PreToolUse hook can also make the permission decision by returning permissionDecision (allow/deny/ask/defer) in its JSON output. Handlers aren't limited to shell commands: command, http, mcp_tool, prompt (single-turn model evaluation) and the experimental agent.
One difference to watch: in Cursor, if a beforeShellExecution/beforeMCPExecution hook crashes, times out, or exits non-zero with anything other than 2, the action is allowed through by default (fail open). Set failClosed: true on security-critical hooks.
Typical uses: run formatters/linters, block commands like rm -rf or git push --force, block reads of .env, inject context at session start (git status, open issues), notify when the agent finishes, write every tool call to an audit log.
Permissions, sandboxing and approval modes
A coding agent runs shell commands under your user account. The harness's most critical job is to be the gate between the model's requests and the real world. The gate has three layers:
1. Permission rules and approval modes
These decide which tools run without asking and which need approval; you are the human-in-the-loop standing at the gate.
- Claude Code permission modes:
default(prompts on first use of each tool),acceptEdits(auto-accepts file edits and common filesystem commands likemkdir/mvin the working directory),plan(reads and explores but doesn't edit source files),auto(no routine prompts; a background classifier checks that actions align with your request before they run),dontAsk(auto-denies every call that would otherwise prompt) andbypassPermissions. Rules are evaluated deny → ask → allow; a broad deny rule beats a narrower allow rule. - Codex CLI uses two separate axes: sandbox mode (
read-only,workspace-write) and approval policy (on-request,never). The default "Auto" preset isworkspace-write+on-request: it reads, edits and runs commands in the workspace; writing outside it or accessing the network requires approval.--dangerously-bypass-approvals-and-sandbox(alias--yolo) turns off both sandbox and approvals. - Cursor uses "Run Modes": Auto-review (allowlisted calls run immediately, other shell commands run in the sandbox when possible, and calls that can't use the sandbox go to a classifier) and Allowlist (only listed actions run without approval). Cursor's docs explicitly state that Auto-review is not a security boundary.
- Gemini CLI has
/permissionsand/policiescommands and a plan mode.
2. OS-level sandbox
Permission rules answer "which tool?"; an agent sandbox answers "what can the tool touch once it runs?". A sandbox is a boundary the operating system enforces; no amount of model persuasion gets past it.
- Claude Code: the Bash tool is sandboxed with Seatbelt on macOS and bubblewrap on Linux and WSL2. Network traffic goes through a proxy that checks a domain allowlist. Toggle with
/sandbox. - Codex CLI: built-in sandbox for macOS, Linux and Windows; test how a command behaves inside it with
codex sandbox macos|linux|windows. - Gemini CLI:
sandbox-exec(Seatbelt) on macOS, or a Docker/Podman container; enable with-s/--sandbox, theGEMINI_SANDBOXenvironment variable, orsettings.json. - Cursor: terminal commands can run in a sandbox that restricts file and network access; a command that fails on a sandbox restriction can be rerun outside, and that rerun goes through the classifier.
3. Prompt injection: why this matters so much
Prompt injection means instructions embedded in content the model reads: a hidden line in a README, white-on-white text on a web page, an issue comment, data returned by an MCP tool. The model cannot reliably tell these apart from your instructions.
If an agent can simultaneously (1) access your private data, (2) read untrusted content and (3) send data out (network, git push, email), the door is open for an attacker. Simon Willison calls this combination the "lethal trifecta". The defense lives in the harness, not the model:
- Restrict the network with the sandbox and a domain allowlist.
- In sessions that read untrusted content, disable write/send tools or gate them behind approval.
- Read the code /
SKILL.mdof third-party MCP servers and skills before installing. - Put critical controls in hooks and deny rules, not in instruction files. More in Guardrails.
Memory across sessions
Every session starts with an empty context window. Agent memory carries what was learned in earlier sessions into the next one. There are two kinds:
- Hand-written instruction files (
CLAUDE.md,AGENTS.md,GEMINI.md): you write them, they load every session. Hard rules belong here. - Notes the agent keeps itself:
- Claude Code auto memory: Claude writes notes into a per-project directory,
~/.claude/projects/<project>/memory/. ItsMEMORY.mdis an index; the first 200 lines or 25 KB (whichever comes first) load at the start of every session. Topic files don't load at startup and are read on demand: progressive disclosure again. Browse and edit with/memory. - Codex Memories: local Codex clients use a separate local memory store;
/memoriescontrols whether the current chat may use existing memories or feed future ones. OpenAI's own advice: keep required rules inAGENTS.mdand treat memories as a helpful recall layer. - Gemini CLI: a memory tool for persistent notes, plus an experimental "Auto Memory" feature.
- Anthropic API memory tool (
memory_20250818): runs client-side. Claude requests operations on files under/memories; your application does the storage. If you're building your own agent, this is a ready-made memory layer.
- Claude Code auto memory: Claude writes notes into a per-project directory,
Memory is not the same as RAG: RAG searches an external knowledge base; agent memory comes from the agent's own experience (which test is flaky, which preference the user stated).
Observability: transcripts, cost, checkpoints
You can't trust an agent whose actions you can't see. A good harness gives you three things:
Transcripts
Claude Code writes a full record of each session to ~/.claude/projects/<project>/<session>.jsonl: every message, every tool call, every tool result. These records aren't encrypted at rest; if a tool reads .env or a command prints a secret, that value lands in the transcript. In Codex, /resume returns to a saved chat and /fork branches the current one. In Gemini CLI, /resume (alias /chat) opens the session browser. The OpenAI Agents SDK records runs as traces; trace_include_sensitive_data controls whether tool inputs/outputs are included.
Cost and token tracking
Token usage scales with context size: a bigger context means more tokens on every message. In Claude Code, /usage (alias /cost) shows session cost and plan usage; in Codex, /status shows token usage and remaining context capacity; in Gemini CLI, /stats shows session statistics.
Checkpoints and rewind
When the agent goes off track, you need a way back:
- Claude Code: a checkpoint is captured automatically before each prompt that starts a turn;
/rewind(aliases/checkpoint,/undo) rewinds code and/or conversation. Key limitation: files changed by Bash commands (e.g.rm,mv, script output) are not tracked; only changes made by the file-editing tools are. - Gemini CLI: when you approve a file-modifying tool, a snapshot is committed to a shadow git repo under
~/.gemini/history/<project_hash>;/restorebrings back files and conversation, and there's also a/rewindcommand. - Cursor: the agent creates checkpoints automatically before significant changes; restore them from the chat timeline.
Checkpoints are for quick session-level recovery; they don't replace git. Start the agent on a clean working tree and commit often.
Comparing popular harnesses
The table below only contains information verified against official documentation; cells I couldn't verify are marked "—". A dash does not mean the feature is missing, only that it wasn't verified.
| Component | Claude Code | Codex CLI | Gemini CLI | Cursor |
|---|---|---|---|---|
| Project instruction file | CLAUDE.md (else AGENTS.md) |
AGENTS.md, AGENTS.override.md |
GEMINI.md |
.cursor/rules/*.mdc, AGENTS.md |
| MCP | Yes | Yes (config.toml) |
Yes (/mcp) |
Yes (mcp.json) |
| Hooks | Yes (PreToolUse etc.) |
Yes (PreToolUse etc.) |
Yes (BeforeTool etc.) |
Yes (beforeShellExecution etc.) |
Skills (SKILL.md) |
Yes | Yes | Yes | Yes |
| Subagents | Yes (.claude/agents/) |
Yes | Yes (.gemini/agents/) |
Yes |
| Permission/approval model | 6 permission modes + allow/ask/deny rules | Sandbox mode + approval policy | /permissions, /policies, plan mode |
Run Modes (Auto-review, Allowlist) |
| OS sandbox | Seatbelt (macOS), bubblewrap (Linux/WSL2) | Built-in (macOS/Linux/Windows) | Seatbelt, Docker/Podman | Yes, for terminal commands |
| Manual compaction | /compact (+ automatic) |
/compact |
/compress |
— |
| Checkpoint / rewind | /rewind |
— | /restore, /rewind |
Checkpoints |
| Cross-session memory | Auto memory (MEMORY.md) |
Memories (/memories) |
Memory tool, Auto Memory (experimental) | — |
| Usage tracking | /usage (/cost) |
/status |
/stats |
— |
Gemini CLI note: Google is moving Gemini CLI to the new Antigravity CLI. For individual and free users, Gemini CLI stopped serving requests on June 18, 2026; enterprise customers keep access. According to Google, Antigravity CLI keeps Agent Skills, Hooks, Subagents and Extensions (now "Antigravity plugins"), though without 1:1 feature parity at launch. The Gemini CLI column reflects the Gemini CLI documentation.
Build your own harness in under 70 lines
The best way to understand a harness is to write the smallest possible one. The Python script below uses the Anthropic SDK for: the agent loop, two tools (read_file, run_command), a permission check (path-escape guard + user approval for every shell command), a turn limit and tool-output truncation.
pip install anthropic
export ANTHROPIC_API_KEY=...import pathlib
import subprocess
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment
MODEL = "claude-sonnet-5-5"
ROOT = pathlib.Path.cwd().resolve()
MAX_TURNS = 20 # seatbelt against infinite loops
MAX_OUT = 20_000 # context budget: max characters per tool result
TOOLS = [
{
"name": "read_file",
"description": "Read a UTF-8 text file inside the project root. Pass a path relative to the root.",
"input_schema": {"type": "object", "properties": {"path": {"type": "string"}}, "required": ["path"]},
},
{
"name": "run_command",
"description": "Run a shell command in the project root and return stdout+stderr. Every call needs user approval.",
"input_schema": {"type": "object", "properties": {"command": {"type": "string"}}, "required": ["command"]},
},
]
def allowed(name: str, args: dict) -> bool:
"""Permission check: the harness decides, not the model."""
if name == "read_file":
return (ROOT / args["path"]).resolve().is_relative_to(ROOT) # block ../../ escapes
if name == "run_command":
return input(f"\nRun `{args['command']}`? [y/N] ").strip().lower() == "y"
return False
def execute(name: str, args: dict) -> str:
if name == "read_file":
return (ROOT / args["path"]).read_text()[:MAX_OUT]
out = subprocess.run(args["command"], shell=True, cwd=ROOT, capture_output=True, text=True, timeout=120)
return (out.stdout + out.stderr)[-MAX_OUT:] or "(no output)"
def run(task: str) -> str:
messages = [{"role": "user", "content": task}]
for _ in range(MAX_TURNS):
resp = client.messages.create(
model=MODEL,
max_tokens=4096,
system="You are a careful coding agent. Read before you change anything; finish with a short summary.",
tools=TOOLS,
messages=messages,
)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use": # no tool calls → task finished
return "".join(b.text for b in resp.content if b.type == "text")
results = []
for block in resp.content:
if block.type != "tool_use":
continue
if not allowed(block.name, block.input):
content, is_error = "Permission denied (user/policy).", True
else:
try:
content, is_error = execute(block.name, block.input), False
except Exception as e: # feed the error back to the model, don't crash the loop
content, is_error = f"Error: {e}", True
results.append({"type": "tool_result", "tool_use_id": block.id, "content": content, "is_error": is_error})
messages.append({"role": "user", "content": results})
return "Stopped: turn limit reached."
if __name__ == "__main__":
print(run("Run this project's tests and summarize any failures."))Those few dozen lines are a harness skeleton. What you'd add to make it a real one maps one-to-one onto this guide's sections:
- Context management: as history grows, truncate or summarize old
tool_results (or use the API's context editing). - Project instructions: read
AGENTS.mdat startup and append it to the system prompt. - Better permissions: an allowlist instead of always asking (safe commands like
git status,pytestrun unprompted), a deny list for patterns likerm -rf. - Sandbox: run
run_commandinside a container or an OS sandbox. - Hook points: lists of functions called before and after
execute. - Observability: append
messagesto a JSONL file after each turn, count tokens viaresp.usage.
If you'd rather use a ready-made harness as a library, the Claude Agent SDK (Claude Code's foundation; tools, permission modes, hooks, subagents, max_turns, max_budget_usd and a can_use_tool callback come built in) or the OpenAI Agents SDK (loop, max_turns, guardrails, tracing) are good starting points.
Checklist for choosing and configuring a harness
When trying a new harness or auditing your current setup:
Loop and limits
- Do unattended runs have a turn and/or budget limit?
- Does failing test/command output flow back to the model?
Context
- Is there a project instruction file, is it short (roughly under 200 lines), and does it contain only what can't be derived from the code?
- If you use several tools, do they share one
AGENTS.md(or import it fromCLAUDE.md)? - Are unused MCP servers disabled?
- Do you know how compaction behaves in long sessions; are key decisions written to files?
Security
- What's the default permission/approval mode? Which commands run without asking?
- Is the OS sandbox on? Is network access limited by an allowlist?
- Are
.env, SSH keys and cloud credentials covered by deny rules? - In sessions that read untrusted content (web, issues, third-party MCP), are exfiltration paths closed?
- Did you read the contents of the skills and MCP servers you installed?
Extensibility
- Are "must happen every time" rules in hooks rather than instruction files?
- Is long research delegated to subagents?
- Are recurring workflows packaged as skills?
Observability
- Where are transcripts stored, and can secrets leak into them?
- Where do you track tokens/cost?
- What does rewind cover and what doesn't it (e.g. changes made via Bash)?
Continue reading
- Agent Harness — the short definition.
- Agent Loop and Function Calling — the API side of the loop.
- Context Engineering and Context Compaction — the theory of managing context.
- Subagent, Agent Skills, Agent Hooks — the harness's extension points.
- Agent Sandbox, Prompt Injection, Human-in-the-loop — the security layer.
- Agent Memory and AI Agent — the big picture.
- Claude Skills Guide, MCP Server Catalog, Build Your Own MCP Server, Token Reduction Techniques, System Prompt Guide.