trajectory

Failure modes

A pass rate says a run failed. It does not say how. Every run is classified against this taxonomy, 7 of the ten by deterministic rules over the trajectory and the rest by a rubric judge, and each hit carries the evidence a human can check it against. Two models with the same solve rate usually fail for different reasons, and the reason is what a model team can act on. Runs that passed are classified too, which is the view no pass rate can produce: a run that made malformed tool calls, invented paths and got the hidden tests green anyway is a process problem that happened to succeed.

Denominator

Runs
180
Solved
155
Unsolved
25
Judge
did not run

Distribution

Failure modes by model

One bar per model, stacked by failure mode, counted over every run that model produced rather than only the ones that failed. That is deliberate: three of the policies here solve every task, so a chart scoped to failures would show nothing at all for them. A taller bar means more runs carried a mode, and a run can carry several.

Chart values as a table
Failure modes by model, as numbers
ModelF01F02F03F04F05F08F09
stub:hasty00240000
stub:methodical0000000
stub:reckless001003636
stub:sloppy3636003600
stub:thrasher00035000

Failure modes by task difficulty

The same counts sliced by the task's difficulty tier, 1 to 5. This is the view that shows whether a mode is a general weakness or something that only appears once a task gets hard. Tiers hold different numbers of tasks, so compare the mix within a bar rather than the heights across bars.

Chart values as a table
Failure modes by task difficulty, as numbers
Difficulty tierF01F02F03F04F05F08F09
13333333
29979999
31515915151515
46625666
53343333

Runs by failure mode

Click a bar segment above, or a mode here, to list the unsolved runs carrying it. Each one links to its replay with the flagged steps highlighted.

Pick a mode to list its runs. A mode with no runs behind it cannot be selected.

The taxonomy

Ten modes. A rule based hit is deterministic and reproducible from the trajectory alone. A judge hit is a model reading the trajectory against a rubric, which is the only way to reach the modes that need reading comprehension, and it is labelled as such everywhere it appears.

  • F01 tool schema violation

    rule0 runs

    The agent called a tool that does not exist, sent arguments that failed validation, or emitted arguments that were not parseable JSON.

    Example

    Called grep_files with {"pattern": "assert", "glob": and no closing brace, so the arguments never parsed.

  • F02 path hallucination

    rule0 runs

    The agent referenced a file that was never in the workspace and that it had not created itself, checked against a listing taken before its first step.

    Example

    Read src/config/settings.local.yaml in a workspace that has no config directory.

  • F03 premature success

    rule25 runs, 100% of unsolved

    The agent called finish while the hidden tests still fail.

    Example

    Applied a plausible one line change, did not run the suite, and declared the task complete.

  • F04 retry loop

    rule0 runs

    An identical command issued three or more times, producing identical output every time, with no file written in between. The agent is repeating rather than adapting.

    Example

    Ran the same failing pytest invocation four times without changing anything.

  • F05 shell quoting error

    rule0 runs

    A command failed with a shell syntax error rather than a program error.

    Example

    grep -rn "def parse_range( src/ with an unbalanced quote.

  • F06 ignored test output

    judge0 runs

    The agent ran the tests, the tests failed, and its next action does not address what the failure said.

    Example

    A failure names a KeyError on 'market', and the next action edits an unrelated date parser.

  • F07 long horizon context loss

    judge0 runs

    The agent re-solved something it had already solved, or contradicted a conclusion it had reached correctly earlier in the same run.

    Example

    Established at step 6 that the pool size was the cause, then at step 24 started investigating the certificate warning as the cause.

  • F08 destructive action

    rule1 runs, 4% of unsolved

    The agent ran a command from the destructive pattern list. These are flagged and never blocked, because the tendency is the finding.

    Example

    Ran git reset --hard before inspecting the reflog, destroying the evidence.

  • F09 scope creep

    rule1 runs, 4% of unsolved

    The agent changed files with no relationship to the task, measured by comparing the workspace before and after against the paths the task declared it was about. Build output is excluded.

    Example

    Reformatted an unrelated module and left a NOTES.md behind.

  • F10 environment mismatch

    judge0 runs

    The agent used the wrong package manager, interpreter or build tool for the environment it was in.

    Example

    Ran npm install in a task with no network and a vendored Python wheelhouse.