Human-in-the-Loop
Approval checkpoints for agents
Designing an autonomous agent to pause at specific critical points and ask a human for approval, direction or a decision; the design that balances autonomy against safety.
A fully autonomous agent is fast, but when it makes a mistake it keeps going without anyone noticing. An agent that asks for approval at every step is safe but useless; the user would finish faster doing it themselves. Human-in-the-loop (HITL) means human checkpoints placed at the right spots in the agent's flow.
Common patterns: - Approval gate: the agent stops right before a risky action (deleting files, deploying, paying, sending email); a human approves, rejects, or approves a modified version. - Plan-then-execute: the agent first only reads and drafts a plan; execution starts once a human has reviewed and approved it. Misunderstandings get caught before any code is written. Claude Code's plan mode is exactly this. - Escalation: when uncertain, or facing something beyond its authority, the agent routes the question to a human instead of guessing ("two readings are possible — which one?"). - Post-hoc review: the agent works freely, but its output (a PR, a draft email) passes a human before it ships.
The core tension: every checkpoint adds safety but subtracts from the agent's value — saving a human's time. Good design reserves approval for actions with high irreversibility and blast radius, and lets the rest run automatically under sandboxing and permission rules.
Like a talented new intern. You don't hover while they draft a report or gather data. But before they email a client or sign a contract, you expect a "could you take a look?". As trust grows, checkpoints thin out; for the irreversible stuff, the signature always stays with someone senior.
A company builds an agent to handle customer refund requests. It can read order history, check the refund policy and draft a reply — all automatically.
But a human steps in at two points:
1. Refunds over $50: when the agent tries to call the
issue_refund tool, the harness pauses the call and the support team
gets an approval request with a summary and the reasoning.
2. Out-of-policy cases: when the agent can't find a rule, it hands
the request to a human with its reasoning instead of deciding.
The result: most requests close within minutes, and the human team only looks at the cases that actually need judgment.
import { query } from "@anthropic-ai/claude-agent-sdk";
// askHuman(): your own approval UI (CLI, Slack, web panel)
for await (const message of query({
prompt: "Clean up temp files and report back",
options: {
// Fires for any tool call not already approved by the permission flow
canUseTool: async (toolName, input) => {
console.log(`Tool: ${toolName}`, input);
const ok = await askHuman("Allow this action? (y/n)");
if (ok) {
return { behavior: "allow", updatedInput: input };
}
// The model sees this message and tries another approach
return { behavior: "deny", message: "User declined" };
},
},
})) {
if ("result" in message) console.log(message.result);
}# Start in plan mode: Claude reads, explores and writes a plan
# but doesn't edit source files. Approving the plan starts execution.
claude --permission-mode plan
# Inside a session, Shift+Tab cycles modes,
# or prefix a single prompt with /plan.- Irreversible actions: deleting, deploying, paying, sending messages outside
- Ambiguous tasks with several readings — a planning step catches misunderstandings early
- Regulated domains (finance, health, legal) — human sign-off is often mandatory
- When you give an agent a new kind of job, until trust is established
- Asking for approval on every tool call — it breeds approval fatigue and people start clicking without reading
- Reversible work that's safe to repeat inside a sandbox (running tests, reading files)
- Real-time flows — waiting on approval makes latency unacceptable
Approval fatigue
A user who sees 'allow?' a hundred times a day approves the hundredth without reading, and the checkpoint exists only on paper. Keep approvals rare but meaningful; put safe operations on an allow list.
Approval requests without context
'Run rm -rf build/?' is meaningless if the human doesn't know why the agent wants it. An approval request should summarise what will happen, why, and the likely impact.
Getting stuck after a rejection
If the agent doesn't know what to do after a 'no', it either asks for the same thing again or stalls. Feed the rejection and its reason back to the model so it can look for an alternative.