Mert Sahin

Running Codex and Claude as a two-agent team: what broke, and the rules that fixed it

· 5 min read · Multi-agent · Process · Testing

I build software with two coding agents from two vendors. OpenAI Codex specifies and reviews. Claude Code builds and integrates. GitHub Spec Kit holds the specs, plans and tasks, and a WORK.md file in the repository records the active work item and whose turn it is.

One rule drives the setup: whoever writes the code does not accept it. I put another vendor's model in the reviewer seat so that "independent" means more than a second prompt to the same model. That is a design choice, not a measured result.

The eight rules, short version

  1. Nothing is done until it is on main and re-verified there.
  2. Re-derive every number in a report from the artifact or the logs, and run the exact command the spec names.
  3. Quote contract rules verbatim; never paraphrase them into assignments.
  4. Never narrow a measurement to make it pass.
  5. Treat the test harness as a product, and prove each control can fail.
  6. Name the defect's class and fix it at the choke point.
  7. Gates must be reachable; if findings stop converging, challenge the criterion.
  8. A file save is a signal, not a handoff. Act on commits.

The loop

SPEC         Codex: scope, allowed paths, acceptance tests + exact commands
 └ BUILD        Claude: product code + tests, one candidate commit
    └ REVIEW       Codex: ACCEPT | CHANGES | REPLAN | BLOCKED
       └ INTEGRATE    Claude: fast-forward main to the accepted commit
          └ DONE         Codex: re-run the tests on main, record VERIFIED

Every step appends an entry to WORK.md that ends in a header like this (values illustrative):

Work: W-012 export-totals
State: REVIEW
Next-Actor: codex
Base: 4f2c1e9
Candidate: a81d3b0
Review-Round: 1
  • Acceptance tests and their exact commands are fixed before BUILD. The builder may add tests, never narrow them.
  • The reviewer checks the exact candidate commit and runs the tests itself.
  • Two rounds of CHANGES at most; after that the work is replanned from a clean main.
  • Integration is fast-forward only, and the latest committed header decides who acts next. Chat is not authority.

Eight things that broke

1. Accepted work that never shipped

My first version had no integration step. An audit of one product found 30 work items opened in four days: 23 blocked, two with accepted executable code, none integrated into main.

Rule: done means on main and re-verified. Report authored, verified, accepted and integrated as separate stages.

2. Reports that described the plan, not the artifact

A build report said a test read nine files; it read eight. Another said eight named controls "all fail" when those IDs were not in the file. A spec named a production build script and the agent ran the plain build instead.

Rule: re-derive every number from the committed artifact or the logs with a script, and run the exact command the spec names.

3. Paraphrased rules became defects

An assignment shortened "a mixed-currency total becomes unknown" to "a currency mismatch is refused." Implementation and verification both followed the paraphrase, and the conflict surfaced two slices later.

Rule: when a contract already defines a rule, quote it verbatim with its location.

4. A filter that made a probe pass

A probe counting calls on a runtime built-in returned one. The agent filtered out the object it blamed, got green, and reported its theory as fact. The real caller was elsewhere, and a hostile input could slip through as an ordinary error.

Rule: never narrow a measurement to make it pass. Get a stack trace before forming a theory.

5. The verifier needed verifying

Seven of eleven mutation survivors were weak tests: .not.toBe('accepted') passes for any rejection, so deleting the rule under test still passed. The harness itself decoded Windows test output with the wrong codec and saw empty runs. The product had 752 tests; the harness that certified it had none.

Rule: treat the harness as a product. Prove every control fails on a known-bad input before trusting it.

6. Fixing the sentence instead of the class

Asked to remove a counted list from a document, the agent deleted the caption and left the list. Elsewhere, a fix covered one write path; a 15-line witness script found four more paths with the same bug.

Rule: name the defect's class, check every path that shares the mechanism, and fix it at the choke point.

7. A gate you can't pass by trying harder

Six adversarial review passes each found between 3 and 21 real defects, and the rate never fell. Rewriting the gate as a closed list of about forty named faults converged in one round.

Rule: if real findings stop converging, the criterion is the problem. Escalate it instead of grinding out more rounds.

8. Acting on a save instead of a commit

A file watcher woke Claude when WORK.md was saved, a moment before Codex committed. Claude nearly rejected a valid ACCEPT as unauthorized.

Rule: a save is a signal. Act only when the tree is clean and HEAD has moved.

If you adopt this

  • Run one work item through the whole loop by hand before automating anything.
  • Keep one active item and one writer. Git is the evidence store: no amend, rebase or stash in a delivery checkout.
  • Dry-run an assignment with a fresh read-only reviewer before dispatching it.
  • Escalate to people only for product, risk and release decisions. A failed review round is not one of them.

Where it's going

Next is a more autonomous version in which each model runs in its own harness and spawns its own subagents. I already use that pattern for review: before one freeze, forty read-only agents with a separate refuter for each claim confirmed three real defects that the tests had passed. More autonomy means more output nobody reads line by line, so these rules matter more, not less.

  • GitHub Spec Kit, the spec-driven toolkit used for specify, plan and tasks.