Optimization techniques for training and inference of LLMs
About this session
As Large Language Models (LLMs) scale exponentially, the infrastructure required to serve them faces a critical efficiency wall. This session explores how Python serves as the primary control plane for massive AI workloads, demonstrating how a single parameter change in a Python line can optimize underlying serving frameworks like vLLM. We will illustrate how high-level Python code orchestrates complex techniques—such as PagedAttention, quantization, and chunked prefill—translating a simple flag into massive compute and memory savings across thousands of cluster machines, ultimately slashing latency and operational costs.
Prefix caching, Paged Attention and Chunked prefill reduces approximately 10x of TTFT, Quantization technique (AWD) increases throughput by 50% , Prefix caching and paged attention approximately increases prefix cache by 50% . Using optimization techniques, we can reduce the cost by 8%
Speaker
Key takeaways
- TPU/GPU Architectural Efficiency
- Financial and Operational Impact
Related sessions
- Context Validation: The Missing Layer for Coding Agents
- Your Inbox Is the New Interface: Building AI Agents on Chat applications
- Agents Don't Fail in the Sandbox: What Actually Breaks When You Ship AI Agents to 1000's of Users
- AI-GENERATED VOICE SYSTEMS: INTEGRATING COPYRIGHT LICENSING, BIOMETRIC DATA PROTECTION AND AUTOMATED ENFORCEMENT MECHANISMS