Agent harnesses for complex, verifiable tasks
A strong model is not enough when requirements change mid-task, tools are many, and each evaluation costs hours. The harness around the model decides what it sees, what it may try, how results are checked, and what it remembers. I study each of these parts separately, with fixed budgets and graders the agent cannot edit.
- 01
Grounding
Turn natural-language requirements into executable, checkable constraints.
- 02
Search
Search over solver code, parameters, and tool scripts, not just final answers.
- 03
Staged verification
Run cheap checks first; spend expensive runs only on candidates that survive.
- 04
Memory
Track what worked, where, and why, and find out what evidence is still missing.
Self-improving harnesses under fixed verifiers. The harness proposes edits to itself; independent checkers on held-out tasks decide whether each change is kept or rolled back.