How to Build AI That Does Not Fall Apart When the Data Gets Real
About this session
Every builder learns it eventually. The model works beautifully on clean sample data, then falls apart the moment it meets the real thing, messy, inconsistent, and never quite what you expected. I build the data systems behind AI in a national healthcare platform, where millions of records arrive every day and none of them are clean. In this talk I will share the discipline that separates AI that survives production from AI that quietly breaks. I will cover how to validate and clean data as it arrives, how to catch bad or drifting inputs before they reach your model, how to generate realistic synthetic data so you can build and test without risking real users, and how to design for the failures that only show up at scale. Healthcare is my proving ground, but every builder in the room will recognize the problem. The smartest model in the world cannot outrun the data you feed it.
Speaker
Key takeaways
- How to validate and clean data as it arrives, and catch bad or drifting inputs before they reach your model.
- How to generate realistic synthetic data so you can build and test without risking real user information.
- How to design for the data failures that only appear at scale, drawn from running AI data platforms at national scale.