Building AI Transcription for Places the Internet Doesn't Reach
About this session
Poket started as a field data collection tool. A few years in, we added AI-assisted transcription so field teams working in remote, low-literacy, low-connectivity settings could turn voice recordings into usable, searchable data. This talk is a builder's postmortem: how a 2-person engineering team, fully bootstrapped, shipped AI transcription across 100+ languages without a dedicated ML team, a big cloud budget, or reliable internet as a starting assumption.
I'll walk through the real tradeoffs — choosing hosted transcription APIs vs. smaller self-hosted models, what broke when models failed silently on rare dialects and noisy field audio, how we built human-in-the-loop verification to catch errors instead of trusting model output blindly, and what it actually cost — in dollars, engineering time, and accuracy tradeoffs — to go from prototype to a feature now used across 30+ countries.
Speaker
Key takeaways
- How to scope an AI feature when your team is 2 engineers and your users are offline by default — the architecture decisions that trade model size and accuracy for field reliability.
- What actually breaks in production: real failure modes of AI transcription on low-resource languages and noisy field audio, and the human-in-the-loop system built to catch them before they reach users.
- Real cost tradeoffs — what it takes in dollars and engineering time to ship and maintain an AI feature on a bootstrapped budget, and which corners are safe to cut.
Related sessions
- From AI Idea to Measurable Product: A Product Leader’s Playbook for Building AI That Actually Delivers
- The Last Mile Problem in Healthcare: How AI Can Improve Patient Access and Adherence
- Build with AI: from founder idea to working product
- From Problem to Production: What I Learned Building an AI Product That Delivered Real Commercial Value