trajectory

Coding agent results, with the spread

Each row is one model on one suite from one sandbox. Solve rate is the headline, and the other columns are what a solve rate cannot tell you: whether the agent got there in a sensible number of steps, whether it recovered when a command failed, and whether it declared victory while the hidden tests were still red. Every number carries its standard deviation across seeds, and a cell with no runs behind it reads n/a rather than zero.

Dataset

Last run
16 Sep 2026, 18:52 UTC
Runs
180 (155 solved)
Tasks
12
Models
5
Total spend
$0.00
Total wall clock
18m 9s

harness 0.1.1, schema 1, bundle built 16 Sep 2026, 19:07 UTC, rubric judge did not run

Results

5 rows. Click a header to sort.

Aggregate results per model, suite and sandbox backend. Every metric column shows a mean with its standard deviation across seeds, and reads n/a when no runs sit behind it.
Model
stub:methodicalscripted policy
docker36100.0%100.0%+/- 0.0%1.000+/- 0.000100.0%+/- 0.0%100.0%+/- 0.0%6.2%+/- 0.0%0.0%+/- 0.0%n/an/a5.99 s
stub:sloppyscripted policy
docker36100.0%100.0%+/- 0.0%0.766+/- 0.00092.2%+/- 0.0%100.0%+/- 0.0%5.2%+/- 0.0%0.0%+/- 0.0%n/an/a5.65 s
stub:thrasherscripted policy
docker36100.0%100.0%+/- 0.0%0.664+/- 0.000100.0%+/- 0.0%84.7%+/- 13.7%49.3%+/- 0.0%0.0%+/- 0.0%n/an/a8.75 s
stub:recklessscripted policy72 destructive
docker3697.2%98.1%+/- 2.7%0.766+/- 0.001100.0%+/- 0.0%100.0%+/- 0.0%4.5%+/- 0.0%2.8%+/- 3.9%n/an/a5.66 s
stub:hastyscripted policy
docker3633.3%67.4%+/- 16.7%1.000+/- 0.000100.0%+/- 0.0%100.0%+/- 0.0%1.4%+/- 1.0%66.7%+/- 23.6%n/an/a4.21 s
What each column measures
Metric 1: solve rate
Every hidden test passed. Deliberately not derived from the run status: an agent can call finish on a run that solved nothing, and a run that hit its step ceiling can still have left the workspace in a passing state. Aggregated across seeds.
Metric 2: partial credit
Hidden tests passed divided by hidden tests discovered. Four failing tests taken down to one is progress, and a benchmark that reports that identically to no change at all has thrown away the only signal in the run.
Metric 3: step efficiency
Reference step count divided by steps taken, capped at 1. Measured over solved runs only, because on a run that failed a low step count means the agent gave up early rather than that it was efficient. A model that solved nothing has no efficiency at all, and the cell reads n/a.
Metric 4: tool call validity
Valid tool calls divided by total tool calls. A call counts as invalid when it named a tool that does not exist, sent arguments that failed validation, or emitted arguments that were not parseable JSON.
Metric 6: recovery rate
Of the commands that failed, the fraction the agent adapted after. A failure counts as recovered when at least one of the next two steps is something other than a re-issue of the same command. Measured over runs where something failed, so a row with nothing failing reads n/a.
Metric 5: redundant action rate
Byte identical repeat commands divided by total commands. Read this against the reference solution, not against zero: running a test suite again after changing the code is a repeat command and exactly the right thing to do. The signal is the gap between the agent's rate and the reference's.
Metric 7: premature termination rate
The share of runs where the agent called finish while the hidden tests still fail. This is the failure a pass rate is least able to see: the run looks like a clean completion from the outside, with no timeout, no error and no budget exhaustion.
Metric 8: context drift
Judge scored. 0 means the final third of the trajectory is still on task, 1 means it has drifted entirely. Reads n/a when the rubric judge did not run.
Metric 9a: mean cost per solved task
Summed provider cost divided by the number of solved runs. Null, and shown as n/a, when nothing was solved: a cost per solved task with no solved tasks behind it is not a number.
Metric 9b: mean wall clock
Mean wall clock seconds per run, including the verification phase.
Backend
Which sandbox produced the runs. Part of the row's identity: runs from the unisolated local backend form their own rows and are never averaged with container runs.
Runs
Runs behind the row, then the distinct tasks and seeds they cover. A mean over one seed has no spread to report, and every cell that aggregates across seeds says so.

Per task results and the step by step replay of every run are linked from the task list below. The failure taxonomy explains what the harness looks for when a run does not solve its task.

Tasks

Each task is a workspace with something wrong in it, a hidden test suite the agent never sees, and a reference solution that has to pass those tests in CI. The reference step count is what step efficiency is measured against.