What Breaks When AI Workflows Go to Production
About this session
AI teams move fast, but production systems often fail in the spaces between models, tools, workflows, and deployment gates. This session shares lessons from building cloud-scale engineering platforms that orchestrate complex software delivery and AI-ready infrastructure workflows across validation, governance, automation, observability, and operational readiness. I will cover what breaks when workflows become long-running, distributed, and dependent on many downstream systems: unclear failure ownership, unreliable status, retry/cancellation edge cases, weak auditability, and manual recovery. Attendees will learn practical patterns for designing resilient workflow layers: stage abstraction, durable execution state, idempotent integrations, reconciliation loops, meaningful observability, and governance controls that help AI and platform teams ship faster without losing reliability.
Speaker
Key takeaways
- How to design reliable workflow orchestration layers for AI and cloud engineering teams.
- How to handle production failure modes such as retries, cancellations, stuck executions, inconsistent status, and unclear ownership.
- How observability, auditability, reconciliation, and governance make AI-ready platforms safer to scale.
Related sessions
- Your Inbox Is the New Interface: Building AI Agents on Chat applications
- Agents Don't Fail in the Sandbox: What Actually Breaks When You Ship AI Agents to 1000's of Users
- AI-GENERATED VOICE SYSTEMS: INTEGRATING COPYRIGHT LICENSING, BIOMETRIC DATA PROTECTION AND AUTOMATED ENFORCEMENT MECHANISMS
- Replace vibe-checks with real data for your AI products