Building AI for Work You Can’t Fully Specify
About this session
Many of the most valuable applications of enterprise AI involve tasks that cannot be evaluated like code. A report, recommendation, research summary, or analysis may be factually accurate and still fail because it emphasizes the wrong information, misses organizational context, misjudges its audience, or makes choices an expert would not have made. The challenge is not that these tasks have no right answers. It is that they often have many defensible answers, and quality depends on context, evidence, professional conventions, and human judgment.
Drawing on experience deploying AI for complex, high-stakes enterprise work, this talk examines what changes when correctness is necessary but insufficient. I’ll cover why generic evaluation and LLM-as-a-judge approaches break down, how to separate objective checks from contextual judgments, how to use expert feedback without reducing quality to a single score, and how to design systems that improve over time.
Attendees will leave with a practical framework for identifying what can be verified, what must be grounded, and where human judgment remains essential.
Speaker
Key takeaways
- Learn about AI deployment in the enterprise
- Understand how to approach not fully specifiable tasks
- How to work with tacit knowledge in the context of AI
Related sessions
- The Governance Layer Nobody's Building: How Client Context Becomes Language in AI-Driven Wealth Advice
- The Fees Nobody Puts on the Receipt: How ML Predicts the Cost of a Card Swipe
- AI × Energy: Where Will the Next Trillion-Dollar Opportunities Be?
- AI-Ready FDA Submissions: CDISC, Validation, and Machine-Readable Data