Pick coding agents on your own commits.
I take bug fixes your team already shipped, hold back the tests that proved them, and have every coding agent you're weighing try again. You get a scorecard of who fixed what, at what cost, and what I'd change.
| Model | Harness | Solved | Cost per run | |
|---|---|---|---|---|
| GPT-6 LunaCodex | Codex | 12 of 14 | $0.004 | |
| GPT-5.6 LunaCodex | Codex | 14 of 15 | $0.02 | |
| GLM-5.2pi | pi | 17 of 17 | $0.04 | |
| GPT-6 SolCodex | Codex | 13 of 14 | $0.15 | |
| GPT-5.6 SolCodex | Codex | 9 of 10 | $0.25 | |
| GPT-6 AstraCodex | Codex | 14 of 14 | $0.52 | |
| Opus 5.5Claude Code | Claude Code | 14 of 14 | $0.53 | |
| Sonnet 5Claude Code | Claude Code | 16 of 16 | $0.67 | |
| Opus 5Claude Code | Claude Code | 10 of 10 | $1.61 | |
| Fable 5.1Claude Code | Claude Code | 14 of 14 | $1.73 | |
How a bake works
- 1
You pick three commits
Recent bug fixes or features, each with the test that proved it.
- 2
I hold the tests back
Each commit is rolled back to before the fix, and its tests stay out of the agent's sight until grading.
- 3
Every agent tries it
Each model and harness you're weighing attempts the fix several times, with your build and test setup.
- 4
You get the scorecard
Pass rate, cost and time per fix, every session transcript, and what I'd change in your setup.
What the September bake found
400x between the cheapest and priciest passing fix
GPT-6 Luna in Codex averaged 0.4¢ a run and Fable 5.1 in Claude Code $1.73, on the same 7 fixes. At 1,000 automated fixes a day, the top of that range is about $50k a month.
A sandbox setting hid failures behind green scores
Codex's default sandbox stopped the test server from binding a port, so on 6 of the 7 fixes its models never saw an integration test pass. The scores looked healthy. The transcripts didn't.
The cheap model was usually enough
GLM-5.2 in pi passed all 17 of its runs at 4¢ each. Start cheap and escalate on a red test.
Who you'd work with
I'm Ari Wilson. I've spent 18 years building software and the engineering orgs that ship it, at Google, Celonis, and VideoAmp. Bakeoff, the eval behind these numbers, is about 7k lines of Go that agents wrote while I designed, reviewed, and argued with them (I typed maybe 10 lines of it).
More about me on ariwilson.comBring three commits
Bring three recent commits to a 30-minute call: bug fixes or features, each with the test that proved it. I'll walk through what a bake on each would measure, what it would take to run against your code, and what the results above suggest before you run anything.