Ship the Model, Not the Headache: 10 Minutes to a Friction-Free AI Serving Pipeline

About this session

You trained a great model then export broke a layer, dynamic input shapes blew up at runtime, and a silent version mismatch quietly halved your throughput in prod. That's pipeline friction: the gap between "works in the notebook" and "serves at scale," and it taxes every AI team.

In 10 fast minutes I'll give you a mental model for the four biggest sources of serving friction export issues, unsupported ops, dynamic inputs, and version mismatches and the highest-leverage fixes for each: validating ONNX exports in CI, using TensorRT optimization profiles instead of rebuilding engines, pinning your full stack with NGC containers, and serving with Dynamo-Triton. You'll leave with a one-page deployment checklist you can drop into your own pipeline today.

Speaker

Key takeaways

  • A quick framework for diagnosing the 4 main sources of AI model-serving friction
  • Why version mismatches fail silently—and how NGC containers prevent them - A one-page deployment checklist to take home Talk Outline (10 min): 0:00–1:30 Pipeline friction: why "works in the notebook" breaks in prod 1:30–4:00 Export & dynamic inputs: CI export validation + TensorRT optimization profiles 4:00–6:30 The silent killer: version mismatches, dependency pinning, NGC containers 6:30–8:30 Serving fast with Dynamo-Triton: d
  • The two highest-leverage fixes: early export validation in CI and TensorRT optimization profiles

Related sessions