Pick coding agents on your own commits.

I take bug fixes your team already shipped, hold back the tests that proved them, and have every coding agent you're weighing try again. You get a scorecard of who fixed what, at what cost, and what I'd change.

ClassBug fixes and features
Judged againstThe tests that shipped with each fix
Entries10 agent setups, 7 commits
ModelHarnessSolvedCost per run
GPT-6 LunaCodexCodex12 of 14$0.004
GPT-5.6 LunaCodexCodex14 of 15$0.02
GLM-5.2pipi17 of 17$0.04
GPT-6 SolCodexCodex13 of 14$0.15
GPT-5.6 SolCodexCodex9 of 10$0.25
GPT-6 AstraCodexCodex14 of 14$0.52
Opus 5.5Claude CodeClaude Code14 of 14$0.53
Sonnet 5Claude CodeClaude Code16 of 16$0.67
Opus 5Claude CodeClaude Code10 of 10$1.61
Fable 5.1Claude CodeClaude Code14 of 14$1.73
September 2026, on 7 commits from 3 of my own apps that no model has seen. Mean cost per run at list API prices. Read the write-up

How a bake works

  1. 1

    You pick three commits

    Recent bug fixes or features, each with the test that proved it.

  2. 2

    I hold the tests back

    Each commit is rolled back to before the fix, and its tests stay out of the agent's sight until grading.

  3. 3

    Every agent tries it

    Each model and harness you're weighing attempts the fix several times, with your build and test setup.

  4. 4

    You get the scorecard

    Pass rate, cost and time per fix, every session transcript, and what I'd change in your setup.

What the September bake found

400x between the cheapest and priciest passing fix

GPT-6 Luna in Codex averaged 0.4¢ a run and Fable 5.1 in Claude Code $1.73, on the same 7 fixes. At 1,000 automated fixes a day, the top of that range is about $50k a month.

A sandbox setting hid failures behind green scores

Codex's default sandbox stopped the test server from binding a port, so on 6 of the 7 fixes its models never saw an integration test pass. The scores looked healthy. The transcripts didn't.

The cheap model was usually enough

GLM-5.2 in pi passed all 17 of its runs at 4¢ each. Start cheap and escalate on a red test.

Read the full write-up on ariwilson.com

Who you'd work with

I'm Ari Wilson. I've spent 18 years building software and the engineering orgs that ship it, at Google, Celonis, and VideoAmp. Bakeoff, the eval behind these numbers, is about 7k lines of Go that agents wrote while I designed, reviewed, and argued with them (I typed maybe 10 lines of it).

More about me on ariwilson.com

Bring three commits

Bring three recent commits to a 30-minute call: bug fixes or features, each with the test that proved it. I'll walk through what a bake on each would measure, what it would take to run against your code, and what the results above suggest before you run anything.

Bring3 recent commits, each with its test
Length30 minutes
CostFree