Designing for When AI Gets It Wrong: The UX Layer Every AI Product Needs

About this session

Most AI products are built to handle the case when the AI is right. The hard part is designing for when it's not.

I'm a UX designer at Bevaya, an AI-powered insurance automation platform. My job is to design the screen where trained reviewers catch, correct, and override AI extractions in real time, at volume, under time pressure, with real liability. That screen is the last line of defense between a hallucinated field value and a bad insurance decision.

Over the past year I've learned things about designing for human-AI collaboration that I couldn't have learned anywhere else. Show the AI output even when it's wrong, blank fields are harder to correct than bad ones. Confidence signals work, but the moment you show a percentage you lose the reviewer's trust. Hard stops before submission feel safe but kill throughput at scale. The document is evidence; the extracted data is the work.

These aren't UX opinions. They're decisions we made, broke, and remade with real reviewers on real workflows.

If you're building an AI product, at some point your users will need to catch what your model missed. This talk is about what happens in that moment, and what the UX of that experience determines about whether people trust your product or abandon it.

Speaker

Key takeaways

  • Most AI products are built to handle the case when the AI is right. The hard part is designing for when it's not. I'm a UX designer at Bevaya, an AI-powered insurance automation platform. My job is to design the screen where trained reviewers catch, correct, and override AI extractions in real time — at volume, under time pressure, with real liability. That screen is the last line of defense between a hallucinated field value and a bad insurance decision. Over the past year I've learned things about designing for human-AI collaboration that I couldn't have learned anywhere else. Show the AI output even when it's wrong — blank fields are harder to correct than bad ones. Confidence signals work, but the moment you show a percentage you lose the reviewer's trust. Hard stops before submission feel safe but kill throughput at scale. The document is evidence; the extracted data is the work. These aren't UX opinions. They're decisions we made, broke, and remade with real reviewers on real workflows. If you're building an AI product, at some point your users will need to catch what your model missed. This talk is about what happens in that moment — and what the UX of that experience determines about whether people trust your product or abandon it.
  • Most AI products are built to handle the case when the AI is right. The hard part is designing for when it's not. I'm a UX designer at Bevaya, an AI-powered insurance automation platform. My job is to design the screen where trained reviewers catch, correct, and override AI extractions in real time — at volume, under time pressure, with real liability. That screen is the last line of defense between a hallucinated field value and a bad insurance decision. Over the past year I've learned things about designing for human-AI collaboration that I couldn't have learned anywhere else. Show the AI output even when it's wrong — blank fields are harder to correct than bad ones. Confidence signals work, but the moment you show a percentage you lose the reviewer's trust. Hard stops before submission feel safe but kill throughput at scale. The document is evidence; the extracted data is the work. These aren't UX opinions. They're decisions we made, broke, and remade with real reviewers on real workflows. If you're building an AI product, at some point your users will need to catch what your model missed. This talk is about what happens in that moment — and what the UX of that experience determines about whether people trust your product or abandon it.
  • Most AI products are built to handle the case when the AI is right. The hard part is designing for when it's not. I'm a UX designer at Bevaya, an AI-powered insurance automation platform. My job is to design the screen where trained reviewers catch, correct, and override AI extractions in real time — at volume, under time pressure, with real liability. That screen is the last line of defense between a hallucinated field value and a bad insurance decision. Over the past year I've learned things about designing for human-AI collaboration that I couldn't have learned anywhere else. Show the AI output even when it's wrong — blank fields are harder to correct than bad ones. Confidence signals work, but the moment you show a percentage you lose the reviewer's trust. Hard stops before submission feel safe but kill throughput at scale. The document is evidence; the extracted data is the work. These aren't UX opinions. They're decisions we made, broke, and remade with real reviewers on real workflows. If you're building an AI product, at some point your users will need to catch what your model missed. This talk is about what happens in that moment — and what the UX of that experience determines about whether people trust your product or abandon it.

Related sessions