Coding agent results, with the spread
Each row is one model on one suite from one sandbox. Solve rate is the headline, and the other columns are what a solve rate cannot tell you: whether the agent got there in a sensible number of steps, whether it recovered when a command failed, and whether it declared victory while the hidden tests were still red. Every number carries its standard deviation across seeds, and a cell with no runs behind it reads n/a rather than zero.
Dataset
- Last run
- 16 Sep 2026, 18:52 UTC
- Runs
- 180 (155 solved)
- Tasks
- 12
- Models
- 5
- Total spend
- $0.00
- Total wall clock
- 18m 9s
harness 0.1.1, schema 1, bundle built 16 Sep 2026, 19:07 UTC, rubric judge did not run
Results
5 rows. Click a header to sort.
| Model | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
stub:methodicalscripted policy | docker | 36 | 100.0% | 100.0%+/- 0.0% | 1.000+/- 0.000 | 100.0%+/- 0.0% | 100.0%+/- 0.0% | 6.2%+/- 0.0% | 0.0%+/- 0.0% | n/a | n/a | 5.99 s |
stub:sloppyscripted policy | docker | 36 | 100.0% | 100.0%+/- 0.0% | 0.766+/- 0.000 | 92.2%+/- 0.0% | 100.0%+/- 0.0% | 5.2%+/- 0.0% | 0.0%+/- 0.0% | n/a | n/a | 5.65 s |
stub:thrasherscripted policy | docker | 36 | 100.0% | 100.0%+/- 0.0% | 0.664+/- 0.000 | 100.0%+/- 0.0% | 84.7%+/- 13.7% | 49.3%+/- 0.0% | 0.0%+/- 0.0% | n/a | n/a | 8.75 s |
stub:recklessscripted policy72 destructive | docker | 36 | 97.2% | 98.1%+/- 2.7% | 0.766+/- 0.001 | 100.0%+/- 0.0% | 100.0%+/- 0.0% | 4.5%+/- 0.0% | 2.8%+/- 3.9% | n/a | n/a | 5.66 s |
stub:hastyscripted policy | docker | 36 | 33.3% | 67.4%+/- 16.7% | 1.000+/- 0.000 | 100.0%+/- 0.0% | 100.0%+/- 0.0% | 1.4%+/- 1.0% | 66.7%+/- 23.6% | n/a | n/a | 4.21 s |
What each column measures
- Metric 1: solve rate
- Every hidden test passed. Deliberately not derived from the run status: an agent can call finish on a run that solved nothing, and a run that hit its step ceiling can still have left the workspace in a passing state. Aggregated across seeds.
- Metric 2: partial credit
- Hidden tests passed divided by hidden tests discovered. Four failing tests taken down to one is progress, and a benchmark that reports that identically to no change at all has thrown away the only signal in the run.
- Metric 3: step efficiency
- Reference step count divided by steps taken, capped at 1. Measured over solved runs only, because on a run that failed a low step count means the agent gave up early rather than that it was efficient. A model that solved nothing has no efficiency at all, and the cell reads n/a.
- Metric 4: tool call validity
- Valid tool calls divided by total tool calls. A call counts as invalid when it named a tool that does not exist, sent arguments that failed validation, or emitted arguments that were not parseable JSON.
- Metric 6: recovery rate
- Of the commands that failed, the fraction the agent adapted after. A failure counts as recovered when at least one of the next two steps is something other than a re-issue of the same command. Measured over runs where something failed, so a row with nothing failing reads n/a.
- Metric 5: redundant action rate
- Byte identical repeat commands divided by total commands. Read this against the reference solution, not against zero: running a test suite again after changing the code is a repeat command and exactly the right thing to do. The signal is the gap between the agent's rate and the reference's.
- Metric 7: premature termination rate
- The share of runs where the agent called finish while the hidden tests still fail. This is the failure a pass rate is least able to see: the run looks like a clean completion from the outside, with no timeout, no error and no budget exhaustion.
- Metric 8: context drift
- Judge scored. 0 means the final third of the trajectory is still on task, 1 means it has drifted entirely. Reads n/a when the rubric judge did not run.
- Metric 9a: mean cost per solved task
- Summed provider cost divided by the number of solved runs. Null, and shown as n/a, when nothing was solved: a cost per solved task with no solved tasks behind it is not a number.
- Metric 9b: mean wall clock
- Mean wall clock seconds per run, including the verification phase.
- Backend
- Which sandbox produced the runs. Part of the row's identity: runs from the unisolated local backend form their own rows and are never averaged with container runs.
- Runs
- Runs behind the row, then the distinct tasks and seeds they cover. A mean over one seed has no spread to report, and every cell that aggregates across seeds says so.
Per task results and the step by step replay of every run are linked from the task list below. The failure taxonomy explains what the harness looks for when a run does not solve its task.
Tasks
Each task is a workspace with something wrong in it, a hidden test suite the agent never sees, and a reference solution that has to pass those tests in CI. The reference step count is what step efficiency is measured against.
- cli-implement-01Difficulty tier 3/5Implement a log summariser from its specification alonepythonref 10 stepsmax 4015 runs, 93% solved, 5 models
- git-surgery-01Difficulty tier 5/5Recover a commit dropped by a bad rebase and rebuild the branchanyref 11 stepsmax 4015 runs, 73% solved, 5 models
- go-race-01Difficulty tier 4/5Fix a data race in a metrics registry without serialising every readgoref 9 stepsmax 4015 runs, 93% solved, 5 models
- log-forensics-01Difficulty tier 2/5Find the root cause of an outage in 40,000 lines of logsanyref 11 stepsmax 4015 runs, 80% solved, 5 models
- py-dep-conflict-01Difficulty tier 2/5Resolve an unsatisfiable set of pinned dependenciespythonref 8 stepsmax 3015 runs, 93% solved, 5 models
- py-etl-schema-01Difficulty tier 3/5Load a vendor export whose schema changed, without losing rowspythonref 12 stepsmax 4015 runs, 93% solved, 5 models
- py-failing-suite-01Difficulty tier 1/5Fix the off by one in an inclusive date rangepythonref 6 stepsmax 2515 runs, 80% solved, 5 models
- py-flaky-test-01Difficulty tier 3/5Fix a test that fails one run in four without deleting itpythonref 11 stepsmax 4015 runs, 80% solved, 5 models
- py-perf-01Difficulty tier 3/5Make an accidentally quadratic rollup meet its two second budgetpythonref 11 stepsmax 4015 runs, 93% solved, 5 models
- sql-migration-01Difficulty tier 4/5Backfill a NOT NULL column on a large table without a long locksqlref 10 stepsmax 4515 runs, 93% solved, 5 models
- ts-api-contract-01Difficulty tier 3/5Regenerate a drifted OpenAPI client and fix every call sitetypescriptref 12 stepsmax 4515 runs, 80% solved, 5 models
- ts-build-break-01Difficulty tier 2/5Repair a TypeScript project reference build without weakening stricttypescriptref 10 stepsmax 3515 runs, 80% solved, 5 models