We Shipped by Making the LLM Do Less: Rebuilding a Clinical Data Review Workflow
About this session
Our first version was one prompt: here's the subject record, find the problems. It demoed beautifully and was dangerous in exactly the way that matters — plausible flags no reviewer could verify. So what happens if you keep narrowing what the model is allowed to do?
I'll walk through the iterations of a production clinical trial data review workflow, where each version moved work out of the LLM: broad review decomposed into narrow checks, ontology-driven context assembly, deterministic pre-filters reclaiming what a rules engine already did well, and an eval set built from years of adjudicated human review.
You'll leave with the precision/recall economics of review workflows — where one noisy flag costs you reviewer trust you only get to lose once — and a concrete pattern for putting LLMs inside a workflow that has to be right.
Production system in a regulated environment. The system and data are internal; the architecture, patterns, and failure modes are the talk, with synthetic examples throughout.
Speaker
Key takeaways
- Split the workflow at the right seam: LLM for interpretation and extraction, deterministic code for rule evaluation. The complex logic was never the hard part — unstructured input was
- Decomposition as an evaluation strategy: narrow checks are individually testable, tunable, and disableable
- Evaluating without labeled ground truth — expert review as the assessment mechanism, and why false positives cost more than misses in a human review workflow.