The bias nobody budgets for. AI evaluation in 100+ languages

About this session

Most AI evaluation frameworks are built, tested, and validated in English, then assumed to generalize everywhere else. They don't. In this talk, I'll walk through a real case from our work building multilingual datasets and evaluation pipelines for enterprise AI: a model that scored well on standard benchmarks and failed badly the moment it hit a low-resource language. The focus will be on looking at what it actually took to catch and fix that, from annotation guidelines to evaluator selection to how the benchmark itself was built. There is a language blind spot at the center of most AI data pipelines. This session goes further, into the operational fixes that close the gap. Attendees will leave with a practical framework for stress-testing their own evaluation process for language, dialect, and demographic coverage.

Speaker

Key takeaways

  • Difference between standard eval and readiness for production
  • Upstream choices in annotation prevent failure in deployment
  • Multilingual and demographic considerations should not be an afterthought

Related sessions