Building AI Transcription for Places the Internet Doesn't Reach

About this session

Poket started as a field data collection tool. A few years in, we added AI-assisted transcription so field teams working in remote, low-literacy, low-connectivity settings could turn voice recordings into usable, searchable data. This talk is a builder's postmortem: how a 2-person engineering team, fully bootstrapped, shipped AI transcription across 100+ languages without a dedicated ML team, a big cloud budget, or reliable internet as a starting assumption.

I'll walk through the real tradeoffs — choosing hosted transcription APIs vs. smaller self-hosted models, what broke when models failed silently on rare dialects and noisy field audio, how we built human-in-the-loop verification to catch errors instead of trusting model output blindly, and what it actually cost — in dollars, engineering time, and accuracy tradeoffs — to go from prototype to a feature now used across 30+ countries.

Speaker

Key takeaways

  • How to scope an AI feature when your team is 2 engineers and your users are offline by default — the architecture decisions that trade model size and accuracy for field reliability.
  • What actually breaks in production: real failure modes of AI transcription on low-resource languages and noisy field audio, and the human-in-the-loop system built to catch them before they reach users.
  • Real cost tradeoffs — what it takes in dollars and engineering time to ship and maintain an AI feature on a bootstrapped budget, and which corners are safe to cut.

Related sessions