The Frugal Context Engineer

About this session

As AI agents move from prototypes to production, the defining bottleneck is no longer what you ask the model - it's what the model sees. This Technical Talk reframes cost optimization as a context engineering discipline, equipping builders with a practical toolkit to slash token spend without sacrificing output quality.

This technical talk introduces context engineering as a disciplined approach to token optimization, distinct from traditional prompt engineering. Attendees will learn a four-pillar framework (prompt caching, prompt compression, semantic caching, and optimized output formats) that can reduce token spend without sacrificing output quality. The talk will also cover five context-efficient agentic architecture strategies, including retrieval-based memory, message history management, tool metadata pruning, large result offloading, and dynamic context assembly. The talk will present actionable KPIs and budget alerting mechanisms to keep AI costs measurable and controllable at scale.

The Lightning Talk is designed for builders, solutions architects, and AI/ML practitioners operating LLM-powered agents in production.

Speaker

Key takeaways

  • Context engineering is distinct from prompt engineering - it focuses on curating the total information the model sees, not just phrasing instructions
  • Frameworks to reduce token spend
  • Context-efficient agentic architecture strategies

Related sessions