Beyond RAG: Building Reliable AI Systems in Production

About this session

Building an LLM application that works in a demo is very different from building one that users can trust in production. This session explores practical lessons from deploying enterprise AI systems, with a focus on retrieval-augmented generation (RAG), LLM evaluation, and agentic workflows. We’ll examine how to design evaluation pipelines, measure answer quality, identify failure modes, and use continuous feedback to improve system reliability. Attendees will learn practical approaches for moving from “the model generated an answer” to measurable, trustworthy AI performance.

Speaker

Key takeaways

  • How to evaluate LLM and RAG systems beyond simple accuracy metrics
  • How to identify and diagnose failure modes in production AI systems
  • How to build continuous evaluation pipelines that improve reliability over time

Related sessions