AI Atlas
EN TR
Intermediate · ~2 min read #harness #agent #claude-code

Agent Harness

The software layer around the model

The surrounding software that turns a language model into a working agent: the tool loop, tools, context management, permissions and interface.

MODEL ≠ HARNESSHARNESS · THE SOFTWARE AROUND THE MODELMODELtext in, text outLOOPmodel → tool → resultTOOLSbash · files · MCPCONTEXTprompt · notes · memoryPERMISSIONSallow · ask · denyINTERFACEterminal · IDE · SDKsame model, different harness → a very different agent
Definition

On its own, a language model takes text in and produces text out. It can't open a file, run a command, or remember what happened in the last step. Everything that turns it into an agent that actually gets work done is called the harness (literally, the gear that turns a horse's strength into useful work).

A typical harness has these parts: - The loop — calls the model, executes the tool it asks for, feeds the result back, repeats until the job is done. - Tools — file read/edit, bash, web search, MCP servers; both the definitions and the code that actually runs them. - Context management — the system prompt, project instruction files (CLAUDE.md, AGENTS.md), summarizing history when the window fills up (compaction). - Permissions and safety — which commands run without asking, which need approval, which are forbidden; sandboxing. - Interface — terminal, IDE extension, desktop app, or an SDK.

Anthropic's Claude Code docs put it plainly: Claude Code is the layer around the model that provides the tools and manages the context the model sees, and that surrounding layer is what agentic harness refers to. Claude Code, OpenAI's Codex CLI and Cursor's agent mode are ready-made harnesses. The Claude Agent SDK ships Claude Code's harness as a library; the OpenAI Agents SDK gives you a loop to wire up with your own tools.

The key lesson is the split: model ≠ harness. The same model can perform very differently in a different harness. When an agent fails, the first question shouldn't be "is the model too weak?" but "did we give it the right tools, the right context and the right permissions?"

Analogy

Engine versus car. The model is a powerful engine; the harness is the chassis, steering, brakes, dashboard and seatbelt. A Formula 1 engine bolted to a test bench makes a lot of noise and goes nowhere.

Put the same engine in a truck or a race car and you get completely different vehicles. Same with agents: using Claude in a chat window and using it in a terminal agent is driving two different vehicles with the same engine.

Real-world example

You ask the same Claude model, from two different places, to "fix the failing tests in this repo."

In a chat interface: the model has no access to the code. It lists likely causes and asks you to paste the file. Running the tests, applying the fix and re-running is on you.

In Claude Code: the harness gives the model tools like Bash, Read, Edit and Grep. The model runs npm test, sees the error, reads the relevant file, fixes it and re-runs the tests. Meanwhile the harness checks the rules in .claude/settings.json: npm run test runs without asking, reading .env is denied.

Same model; the difference is entirely the harness.

Code examples
Claude Agent SDK · using a ready-made harness as a library python
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions


async def main():
    async for message in query(
        prompt="Run the tests and fix any failures",
        options=ClaudeAgentOptions(
            # Tools approved without prompting (not a restriction!)
            allowed_tools=["Read", "Grep", "Glob", "Edit", "Bash"],
            # These tools are removed from the model entirely
            disallowed_tools=["WebFetch"],
            permission_mode="acceptEdits",  # auto-accept file edits
            max_turns=30,                   # turn cap for the loop
        ),
    ):
        # The loop, tools and context management live inside
        # the SDK; you just read the result.
        if hasattr(message, "result"):
            print(message.result)


asyncio.run(main())
Claude Code · the harness's permission layer json
// ~/.claude/settings.json  (or .claude/settings.json in a project)
{
  "$schema": "https://json.schemastore.org/claude-code-settings.json",
  "permissions": {
    "allow": [
      "Bash(npm run lint)",
      "Bash(npm run test *)"
    ],
    "deny": [
      "Read(./.env)",
      "Read(./.env.*)"
    ]
  }
}
When to use
  • General-purpose agent work like coding, file edits, terminal tasks — a ready-made harness (Claude Code, Codex CLI) is usually enough
  • Embedding an agent in your own product — an Agent SDK gives you the loop, tools and context management without writing them from scratch
  • Diagnosing a failing agent — is the problem the model, or the tool, context or permission layer?
  • When you need auditing, logging and approval flows — all of those decisions live in the harness
When not to use
  • Single-shot classification, summarization, translation — a plain API call is enough; a harness is overhead
  • A deterministic workflow whose steps are known up front — a code-driven pipeline is more predictable and cheaper
  • The harness's built-in tools (shell, file writes) are too powerful for your environment and you can't restrict them
Common pitfalls

Blaming the model for every failure

Most of the time an agent gets stuck because of a missing tool, a vague tool description, a bloated context or a wrong permission rule. Look at what the harness shows the model before swapping models.

Rebuilding the harness from scratch

The core loop is 20 lines; but summarizing when context overflows, permission rules, error recovery, session persistence and interruption handling take weeks. Try an existing harness or SDK first.

Treating scores as harness-independent

A score on an agent benchmark also depends on the tool set and loop the model ran in. When comparing two models, keep the harness fixed.

Mistaking allowed_tools for a restriction

In the Claude Agent SDK, allowed_tools auto-approves tools; it does not forbid the ones it leaves out. Use disallowed_tools to remove a tool from the model's view.