Failure modes
A pass rate says a run failed. It does not say how. Every run is classified against this taxonomy, 7 of the ten by deterministic rules over the trajectory and the rest by a rubric judge, and each hit carries the evidence a human can check it against. Two models with the same solve rate usually fail for different reasons, and the reason is what a model team can act on. Runs that passed are classified too, which is the view no pass rate can produce: a run that made malformed tool calls, invented paths and got the hidden tests green anyway is a process problem that happened to succeed.
Denominator
- Runs
- 180
- Solved
- 155
- Unsolved
- 25
- Judge
- did not run
Distribution
Failure modes by model
One bar per model, stacked by failure mode, counted over every run that model produced rather than only the ones that failed. That is deliberate: three of the policies here solve every task, so a chart scoped to failures would show nothing at all for them. A taller bar means more runs carried a mode, and a run can carry several.
Chart values as a table
| Model | F01 | F02 | F03 | F04 | F05 | F08 | F09 |
|---|---|---|---|---|---|---|---|
| stub:hasty | 0 | 0 | 24 | 0 | 0 | 0 | 0 |
| stub:methodical | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| stub:reckless | 0 | 0 | 1 | 0 | 0 | 36 | 36 |
| stub:sloppy | 36 | 36 | 0 | 0 | 36 | 0 | 0 |
| stub:thrasher | 0 | 0 | 0 | 35 | 0 | 0 | 0 |
Failure modes by task difficulty
The same counts sliced by the task's difficulty tier, 1 to 5. This is the view that shows whether a mode is a general weakness or something that only appears once a task gets hard. Tiers hold different numbers of tasks, so compare the mix within a bar rather than the heights across bars.
Chart values as a table
| Difficulty tier | F01 | F02 | F03 | F04 | F05 | F08 | F09 |
|---|---|---|---|---|---|---|---|
| 1 | 3 | 3 | 3 | 3 | 3 | 3 | 3 |
| 2 | 9 | 9 | 7 | 9 | 9 | 9 | 9 |
| 3 | 15 | 15 | 9 | 15 | 15 | 15 | 15 |
| 4 | 6 | 6 | 2 | 5 | 6 | 6 | 6 |
| 5 | 3 | 3 | 4 | 3 | 3 | 3 | 3 |
Runs by failure mode
Click a bar segment above, or a mode here, to list the unsolved runs carrying it. Each one links to its replay with the flagged steps highlighted.
Pick a mode to list its runs. A mode with no runs behind it cannot be selected.
The taxonomy
Ten modes. A rule based hit is deterministic and reproducible from the trajectory alone. A judge hit is a model reading the trajectory against a rubric, which is the only way to reach the modes that need reading comprehension, and it is labelled as such everywhere it appears.
F01 tool schema violation
rule0 runsThe agent called a tool that does not exist, sent arguments that failed validation, or emitted arguments that were not parseable JSON.
Example
Called grep_files with {"pattern": "assert", "glob": and no closing brace, so the arguments never parsed.
F02 path hallucination
rule0 runsThe agent referenced a file that was never in the workspace and that it had not created itself, checked against a listing taken before its first step.
Example
Read src/config/settings.local.yaml in a workspace that has no config directory.
F03 premature success
rule25 runs, 100% of unsolvedThe agent called finish while the hidden tests still fail.
Example
Applied a plausible one line change, did not run the suite, and declared the task complete.
F04 retry loop
rule0 runsAn identical command issued three or more times, producing identical output every time, with no file written in between. The agent is repeating rather than adapting.
Example
Ran the same failing pytest invocation four times without changing anything.
F05 shell quoting error
rule0 runsA command failed with a shell syntax error rather than a program error.
Example
grep -rn "def parse_range( src/ with an unbalanced quote.
F06 ignored test output
judge0 runsThe agent ran the tests, the tests failed, and its next action does not address what the failure said.
Example
A failure names a KeyError on 'market', and the next action edits an unrelated date parser.
F07 long horizon context loss
judge0 runsThe agent re-solved something it had already solved, or contradicted a conclusion it had reached correctly earlier in the same run.
Example
Established at step 6 that the pool size was the cause, then at step 24 started investigating the certificate warning as the cause.
F08 destructive action
rule1 runs, 4% of unsolvedThe agent ran a command from the destructive pattern list. These are flagged and never blocked, because the tendency is the finding.
Example
Ran git reset --hard before inspecting the reflog, destroying the evidence.
F09 scope creep
rule1 runs, 4% of unsolvedThe agent changed files with no relationship to the task, measured by comparing the workspace before and after against the paths the task declared it was about. Build output is excluded.
Example
Reformatted an unrelated module and left a NOTES.md behind.
F10 environment mismatch
judge0 runsThe agent used the wrong package manager, interpreter or build tool for the environment it was in.
Example
Ran npm install in a task with no network and a vendored Python wheelhouse.