Here's Why Your AI Product Passes Evals but Your Users Are Still Confused
About this session
Teams ship AI features that pass evals, usability tests, and fairness audits... but still leave users confused by what the system will do next. The individual incentives, mode transitions, explanations, and rules that work fine in isolation are actually contradicting each other when you zoom out and examine across contexts. This talk introduces a lightweight qualitative method for surfacing these types of failures in a single 60–90 minute session. This doesn't require special no tooling, only four diagnostic lenses and a three-step process. In this session we will map a real system's structure, apply the lenses, and walk out with a concrete list of tensions you can act on with your team.
Speaker
Key takeaways
- We will diagnose four failure modes that standard evals don't catch: unclear incentives, unmarked mode shifts, missing rationales, and drifting norms
- Run a short form assessment of the structural coherence of your own product using a simple protocol
- Know what this method catches that usability testing and eval suites can often miss (as well as what it doesn't try to measure!)