When RAG Meets Reality: Scaling Retrieval for Production

About this session

Retrieval-Augmented Generation (RAG) moves beyond a simple retrieve-and-generate pipeline when applications need to support large scale enterprise documents, high query volumes, changing data, and production reliability. This talk presents practical retrieval architectures for building scalable RAG systems, covering hybrid search, multi-stage retrieval, reranking, query transformation, metadata filtering, caching, and evaluation. We examine the architectural trade-offs that shape retrieval quality, latency, cost, and operational complexity, with a focus on patterns that work in real-world production systems. The session connects these components into scalable architectures and highlights common retrieval bottlenecks, failure modes, and design decisions that determine whether a RAG application remains reliable as it grows.

Speaker

Key takeaways

  • Design retrieval as a multi-stage system, not just a vector search.
  • Balance quality, latency, cost, and scale through the right architecture.
  • Build evaluation and observability into RAG from day one.

Related sessions