Eval Failures That Fooled Us
About this session
We trust evals to tell us the truth about our agents. Sometimes they confidently tell us the opposite. This talk is a set of real cases where an automated judge passed answers it should have failed and how long it took us to notice.
Two anchor stories. First, a judge throwing false positives because outputs were silently truncated at 512 characters, so it was grading a fragment, not the answer. Second, an entity-confusion failure where the judge evaluated the wrong case entirely and reported clean scores while the system was wrong in production.
The point isn't "evals are bad." It's that an eval is itself a system that can fail and most teams instrument the agent but never instrument the judge. I'll show the exact failure signatures, how we caught each (and how we should have caught them sooner) and the checks like input-length validation, entity grounding, spot audits that turn an eval from a comfort blanket into something you can actually trust.
Speaker
Key takeaways
- Checklist of silent eval failures including truncation, entity confusion and judges grading the wrong thing
- How to detect when your judge is lying before it reaches production
- Why the eval is itself a system that fails and why most teams instrument the agent but never the judge