EVALUATION
Frontier Model Evaluation
Four years authoring benchmarks, grading agent trajectories and producing alignment training data for the labs building frontier models.
- RLHF
- DPO
- SFT
- Agentic Evaluation
- Terminal-Bench
- Rubric Design
- Red-Teaming
- Python
- TypeScript
- Go
Context
The labs training frontier models need problems the models cannot already solve, graded by someone who can tell a correct trajectory from a lucky one.
It is close to code review, with one difference that changes the job: a failing agent does not throw. It produces a plausible run that quietly assumed shell state survived a runner boundary, and the grade has to say which of the two it was.
Constraints
A benchmark task has to be hard for a frontier agent and still unambiguous and solvable, or it measures the prompt rather than the model.
Preference data is only worth what its consistency is worth, which is why blind audit rounds against senior reviewers are the real acceptance test.
Agent trajectories are long. Grading them means holding the whole run in view, not scoring the final answer.
What I built
Authored 200+ golden-solution software engineering tasks for Handshake AI's Project Helix, building multi-file repository problems with hidden test suites and graded difficulty tiers across Python, TypeScript and Go, and passed Helix Screening with a top-decile reviewer rating.
Evaluated long-horizon coding-agent trajectories on real software workflows, scoring tool-use correctness, sub-goal decomposition, state tracking and recovery from failure, and surfaced systematic failures in tool-call sequencing and shell command grounding that fed directly into agent fine-tuning.
Authored 50+ accepted Terminal-Bench tasks for Snorkel AI under the Terminus 2 agent scaffold: containerized environments with reference solutions and automated verification tests, calibrated to challenge frontier agents while staying unambiguous and solvable.
Designed 150+ model-stumping prompts validated against GPT, Claude and Gemini-class baselines, plus graduate-level STEM question-answer pairs for Humanity’s Last Exam-class reasoning benchmarks that passed multi-layer expert review.
Ran 3,000+ side-by-side preference comparisons producing RLHF and DPO training data across helpfulness, harmlessness, honesty and completeness.
Architecture
A Terminal-Bench task is a container, a reference solution and an automated verification test. The verification test is the specification, so it is written before the task is called finished.
Trajectory grading runs against a rubric with separate scores for tool-use correctness, sub-goal decomposition, state tracking and recovery from failure, which is what turns a failed run into a labelled failure mode.
Preference comparisons are scored on helpfulness, harmlessness, honesty and completeness as separate axes, then reconciled against senior reviewers in blind audit rounds.
Safety-tier review sits on top, flagging policy violations, jailbreak attempts and unsafe completions and recommending rewrites that hold policy without dropping helpfulness.
Results
200+ golden-solution engineering tasks accepted, with a top-decile reviewer rating at Helix Screening.
50+ Terminal-Bench environments accepted under the Terminus 2 scaffold.
3,000+ preference comparisons at 98 percent or better agreement with senior reviewers.
500+ AI-generated pull requests adjudicated, 150+ validated model-stumping prompts, 800+ expert coding prompts.
20+ enterprise workflow environments simulating accounting, procurement and sales operations, with submissions accepted at scoring tiers up to 40 percent above baseline.
What I would do differently
Write the verification test before the task prose, every time. The tasks that came back for revision were almost always the ones where the test was written last and quietly allowed two readings of the problem.