Trust but Verify: Building Evaluation Systems for Production AI Agents
About this session
Most teams can build an agent demo in a weekend. Getting one to production, and keeping it there, is a different discipline. This session unpacks why evaluating AI agents is genuinely hard: non-deterministic outputs, multi-step trajectories where a correct final answer can hide a broken path, and the gap between “looks right” and “is right.” Drawing on production agents shipped for legal research and workflow automation, Rittika walks through a practical evaluation stack: trajectory-level evaluation that goes beyond final-answer accuracy, LLM-as-judge calibration and its common failure modes, observability with tools like Langfuse, and human-in-the-loop checkpoints that catch what automated evals miss. Attendees leave with a blueprint they can apply to their own agent systems immediately.
Speaker
Key takeaways
- How to evaluate agent trajectories, not just final answers, so you catch silent failures in multi-step reasoning
- Where LLM-as-judge breaks down, and how to calibrate it before you trust it in your pipeline
- A practical evaluation and observability stack you can assemble from tools available toda