Tool-call correctness has four surfaces
A practical framework for evaluating agent tool choice, arguments, execution outcomes, and final responses without collapsing them into one score.
Read the field noteField notes
Practical notes on agent evaluation, human review, and the decisions that sit between a promising model and useful software.
A practical framework for evaluating agent tool choice, arguments, execution outcomes, and final responses without collapsing them into one score.
Read the field note
A reproducible comparison of response-only, trace-only, and trace-plus-outcome evaluators against the same 263 controlled agent failures.
Read the field note
Every review consumes a person's attention. A practical way to decide which AI actions should stop, what the reviewer needs to see, and where the software's authority ends.
Read the field note
Plant one known failure, rerun the same evals, and see whether the score changes. A practical introduction to mutation testing for agent tool use.
Read the field note