Coding agents, October 2024 to October 2026: the 20 events that mattered
TL;DR
- In two years, the parts around the model changed more than the model: protocols, harness config and permissions.
- MCP, AGENTS.md and Agent Skills became formats that agents from several vendors read.
- The headline coding benchmark of 2024 lost its standing. Your own evals are what is left.
What happened
Protocols and shared formats
- Nov 2024: Anthropic open-sources the Model Context Protocol with SDKs and reference servers.
- Mar 2025: the 2025-03-26 MCP spec adds OAuth 2.1, Streamable HTTP and tool annotations that mark a tool read-only or destructive.
- Dec 2025: MCP and OpenAI's AGENTS.md move to the Agentic AI Foundation under the Linux Foundation.
- Dec 2025: Anthropic publishes Agent Skills as an open standard. The same day, GitHub Copilot starts reading
.claude/skills. - Jun 2026: ADT for VS Code ships a built-in ABAP MCP server, disabled by default (my notes).
- Jul 2026: the 2026-07-28 MCP spec drops the initialize handshake and the session ID, and adds headers that let gateways route requests.
The harness
- Feb 2025: Claude Code arrives as a research preview: a terminal agent that edits files, runs tests and commits.
- Sep 2025: Factory's Droid tops Terminal-Bench: "the right agent framework can lead to greater improvements than model selection."
- Feb 2026: Mitchell Hashimoto names harness engineering: when the agent makes a mistake, engineer a fix so it never makes it again.
- Apr 2026: Anthropic's postmortem traces a drop in Claude Code quality to three harness changes. The API was not affected.
- Sep 2026:
ant applymanages agents, skills and memory stores as files in the repository, with a committed lockfile.
Permissions and containment
- Jun 2025: Simon Willison names the lethal trifecta: private data, untrusted content and a way to communicate externally.
- Aug 2025: in CVE-2025-53773, a prompt injection makes Copilot turn on
chat.tools.autoApprovein its own settings, then run shell commands. - Oct 2025: Claude Code sandboxing isolates filesystem and network and cuts permission prompts by 84%.
- Mar 2026: Anthropic reports that users approve 93% of permission prompts and ships auto mode, where classifiers decide instead.
- Aug 2026: auto mode becomes the default for new Pro, Max and Team sessions. In a controlled test, 1,053 paid testers caught 13.6% of dangerous commands; auto mode blocked 89%.
Models and how we measure them
- Oct 2024: the upgraded Claude 3.5 Sonnet reaches 49.0% on SWE-bench Verified, up from 33.4%.
- Mar 2025: METR finds the length of tasks models finish at 50% success doubling about every 7 months. Claude 3.7 Sonnet is at about an hour.
- Jul 2025: in a METR randomised trial, experienced open-source developers took 19% longer with AI tools while believing they were 20% faster.
- Feb 2026: OpenAI stops reporting SWE-bench Verified: at least 59.4% of audited hard problems had flawed tests, and every frontier model tested had seen some solutions in training.
The pattern
Models got better: SWE-bench Verified went from a 49% headline to a benchmark OpenAI no longer reports. But most of what changed how I would build a coding agent sits around the model: a shared tool protocol, instruction and skill files that work across vendors, harness config you can version, and permissions moving from prompts to sandboxes and classifiers. The April 2026 postmortem shows it most clearly: the product got worse while the API did not. More in the harness is the product and long-running agents need a progress file.
Take this: four questions for your own setup
| Area | Ask | If not |
|---|---|---|
| Protocols | Do our MCP servers and clients state which spec version they speak? | Pin it, and plan the move to the stateless 2026-07-28 spec. |
| Harness | Are prompts, tools, permissions and skills in git, with an eval on every change? | Start with ten tasks from your own repo. |
| Permissions | Can the agent edit its own config, or combine private data, untrusted input and network access? | Deny writes to agent config. Split the trifecta. |
| Measurement | Do we pick models on our own tasks rather than a public leaderboard? | Keep a private task set. Read transcripts. |
My take
Most of my work on coding agents for SAP developers sits in these four rows: MCP servers that expose ABAP development with read-only profiles and per-action confirmation gates, an Eclipse plugin whose tool allow/deny lists decide what the agent may touch, and CI gates with automated tests and an LLM-as-judge eval. Models improve on the vendors' schedule. The parts around them are the ones you own.
Links
- An update on recent Claude Code quality reports (Anthropic)
- The lethal trifecta for AI agents (Simon Willison)
- Measuring AI ability to complete long software tasks (METR)