Long-running agents need a progress file, not a bigger context window
TL;DR
- Vendors report agent runs of 24 to 36 hours. METR's data says reliability drops long before that.
- The long runs that worked kept their state outside the context window: a feature list, a progress file and git.
- Start every session by reading that state, finish one feature, commit.
What happened
- Mar 2025: METR finds that the length of tasks models finish at 50% success has doubled about every 7 months. Claude 3.7 Sonnet is at about one hour.
- Sep 2025: Anthropic reports Claude Sonnet 4.5 keeping focus for more than 30 hours. Its context engineering post the same day names the tools for long tasks: compaction, structured note-taking, sub-agents.
- Nov 2025: OpenAI's GPT-5.1-Codex-Max is trained to work across context windows through compaction, and ran for more than 24 hours in OpenAI's tests.
- Nov 2025: Anthropic's long-running harness: an initializer agent writes
init.sh, a progress file and a JSON list of over 200 features. Each later session reads the git log and progress file, builds the top unfinished feature and commits. - Feb 2026: 16 parallel Claude agents build a C compiler that compiles Linux 6.9: about 2,000 sessions over nearly two weeks, 100,000 lines, just under $20,000. They kept progress files current and claimed tasks with lock files.
- Feb 2026: Cursor's long-running agents preview shows runs of 25 to 36 hours, with a plan approved up front and several agents checking each other's work.
- May 2026: METR's time-horizon data puts Claude Opus 4.6 at about 12 hours at 50% success but about 70 minutes at 80%, and warns that "measurements above 16 hrs are unreliable with our current task suite."
- Sep 2026: with GPT-6 Astra, Codex keeps notes across context windows, because "each compaction can leave out details about why a fix failed."
Take this: a progress file, a feature list and six rules
PROGRESS.md, which the agent appends to and people read:
# PROGRESS
Goal: <one line, link to the spec>
Done when: <exact command, e.g. make test-all>
## Now
Feature: F-014 Branch: agent/f-014 Last green commit: <sha>
State: tests written, 2 of 5 failing
## Log (newest first, append only)
- F-013 passes at <sha>. Next: F-014.
- F-013 blocked by a missing fixture; added it in <sha>.
## Traps
- Integration tests need `make db-up` first.
features.json, where the agent may only flip passes:
[
{"id": "F-014", "description": "Export totals per currency",
"verify": "pytest tests/test_export.py -k currency", "passes": false}
]
- Start every session the same way. Read
git log --oneline -20,PROGRESS.mdandfeatures.json, then run the smoke test. Fix a red baseline first. - One feature at a time: the first entry with
"passes": false. - Flip
passesonly whenverifypasses. Anthropic's harness put it bluntly: "It is unacceptable to remove or edit tests." - Commit after every green feature, with the progress entry in the same commit. If the context is lost, nothing else is.
- Write the file before compaction, not after. A summary can drop why a fix failed. The file keeps it.
- Keep the list in JSON. Anthropic found models less likely to overwrite JSON than Markdown.
My take
Plan unattended runs around METR's 80% number, not the 50% one, and let files carry the rest. My own delivery process works the same way: Codex specifies and reviews, Claude Code builds, and the acceptance tests exist before the build starts. A state file in the repository, not the chat, says whose turn it is. The reviewer reruns the tests on the exact candidate commit, and mutation testing checks the tests themselves. Details in Running Codex and Claude as a two-agent team.