From Model Quality to User Value: A Measurement Framework for AI

About this session

As AI systems become more capable, measuring whether they are actually improving becomes increasingly difficult. A model can score higher on benchmarks, achieve better LLM-judge scores, and produce more helpful responses, yet fail to improve the user's experience or the product's outcomes.

This talk presents a practical measurement framework for connecting model quality to user value. We will examine how to combine offline evaluations, automated regression testing, human feedback, online experimentation, and product metrics to build a more complete picture of AI performance.

The talk will explore questions such as: Which dimensions of AI quality should we measure? When should we trust offline evaluations versus online experiments? How do we avoid optimizing misleading proxy metrics? How should we evaluate tradeoffs between quality, safety, latency, and cost? And how do we connect improvements in model behavior to meaningful user and business outcomes?

The goal is not to define a single “AI quality score,” but to build an evidence-based measurement loop that helps teams make better decisions about what to ship, what to improve, and whether an AI system is genuinely creating value.

Speaker

Key takeaways

  • Build a measurement framework that spans across model quality, user experience, and product success
  • Design evals that combine deterministic metrics, LLM-as-a-judge, and human evaluation
  • Connect offline evals to online experimentation

Related sessions