The harness is the product
TL;DR
- In April 2026, Anthropic traced a drop in Claude Code quality to an effort default, a caching bug and one line of system prompt. The API was not affected.
- The harness is everything around the model: prompts, tools, permissions, skills, hooks and defaults. It decides what users get.
- Keep it in git, pin what the vendor can change, run evals on every change and roll out slowly.
What happened
- Jun–Oct 2025: Claude Code adds hooks and subagents, then plugins and Agent Skills that load only when a task needs them.
- Sep 2025: Factory's Droid tops Terminal-Bench. Droid on Sonnet beats every other agent running on Opus.
- Feb 2026: Mitchell Hashimoto names harness engineering: when the agent makes a mistake, add an AGENTS.md line or a real tool so it never makes it again.
- Mar 2026: Anthropic writes that "every component in a harness encodes an assumption about what the model can't do on its own."
- Apr 2, 2026: Birgitta Böckeler sums it up as "Agent = Model + Harness", with guides that steer before the agent acts and sensors that check after.
- Apr 23, 2026: Anthropic's postmortem. On March 4 the default effort went from high to medium. On March 26 a change meant to clear old thinking from sessions idle for over an hour kept clearing it every turn. On April 16 a system-prompt line capped text between tool calls at 25 words. The fixes: "a broad suite of per-model evals for every system prompt change", soak periods and gradual rollouts.
- Sep–Oct 2026:
ant applymanages agents, skills and memory stores as files with a committed lockfile. Claude Code mods can rewrite prompts and gate tool calls, and are not sandboxed.
Take this: a harness release checklist
- One home. AGENTS.md, settings, permissions, hooks, skills and the MCP server list live in one repo and change through pull requests.
- Pin what the vendor can change. Set model and effort explicitly. The April regression started with a default.
- Eval every change. Run a fixed task set from your own codebase on
mainand on the branch; read some transcripts. Trigger it on every pull request touchingAGENTS.md,.claude/or.mcp.json. - Roll out slowly. Pilot group, soak period, one-commit revert.
- Record the version. Store the harness commit with every session log and eval result, so a regression can be bisected.
- Keep the agent out of its own harness. In CVE-2025-53773, a prompt injection made Copilot switch off its own approval prompts.
- Review extensions like dependencies. Plugins and mods run code on your machine.
- Delete scaffolding when a new model no longer needs it, then rerun the evals.
A checked-in .claude/settings.json that pins the model and effort and keeps the agent out of its config:
{
"model": "claude-opus-5-5",
"effortLevel": "high",
"permissions": {
"allow": ["Bash(npm run lint)", "Bash(npm run test *)"],
"deny": ["Read(./.env)", "Edit(/.claude/**)", "Edit(/AGENTS.md)", "Bash(git push *)"]
},
"hooks": {
"PostToolUse": [
{ "matcher": "Edit|Write", "hooks": [{ "type": "command", "command": "npm run lint" }] }
]
}
}
My take
Most of what I build is harness. The MCP servers that expose SAP ABAP development to agents have read-only profiles and per-action confirmation gates. The Eclipse plugin that brings a coding agent to SAP developers runs on internally hosted models, and its tool allow/deny lists decide what the agent can touch. Changes pass CI gates with automated tests and an LLM-as-judge eval. None of that is the model. The April postmortem is the case for giving a one-line config change the same gates as a code change.
Links
- An update on recent Claude Code quality reports (Anthropic)
- Harness engineering for coding agent users (martinfowler.com)
- My AI Adoption Journey (Mitchell Hashimoto)
- Harness design for long-running application development (Anthropic)
- Settings files and precedence (Claude Code docs)