An agent harness with evals that grade the end state, not the output.
{{ status }}
what the agent said
{{ tx.tag }}
> I fixed the failing test in pkg/
  and confirmed the suite is green.
  I made a small, targeted change.
{{ tx.judge }}
{{ arrow }}
what it left behind
def grade(sandbox) -> Verdict
{{ c.sym }}
{{ c.text }}
{{ verdict }}
{{ c.name }}
{{ c.proves }}
{{ caption }}