-
Fixing a Broken Evaluation Metric: Comparable Positions, Not Comparable Sequences (Ep:06.03)
The fix for a per-tool evaluation metric confounded by trace length isn’t more data — it’s evaluating only the one position that’s genuinely comparable across every tool. Module 6 of…
-
Catching an Undertrained Tool Before It Fails: Two Signals, and a Real Evaluation Trap (Ep:06.02)
Per-tool held-out loss and label-free generation confidence both aim to catch an undertrained agent tool automatically — but one of them walks straight into the exact evaluation-scope trap Module 5…