Optimization techniques for training and inference of LLMs

About this session

As Large Language Models (LLMs) scale exponentially, the infrastructure required to serve them faces a critical efficiency wall. This session explores how Python serves as the primary control plane for massive AI workloads, demonstrating how a single parameter change in a Python line can optimize underlying serving frameworks like vLLM. We will illustrate how high-level Python code orchestrates complex techniques—such as PagedAttention, quantization, and chunked prefill—translating a simple flag into massive compute and memory savings across thousands of cluster machines, ultimately slashing latency and operational costs.

Prefix caching, Paged Attention and Chunked prefill reduces approximately 10x of TTFT, Quantization technique (AWD) increases throughput by 50% , Prefix caching and paged attention approximately increases prefix cache by 50% . Using optimization techniques, we can reduce the cost by 8%

Speaker

Key takeaways

  • TPU/GPU Architectural Efficiency
  • Financial and Operational Impact

Related sessions