What Happens When Your AI Gets Women's Health Wrong
About this session
Most women's health apps collect rich longitudinal data, including cycle length, symptom severity, mood, bleeding patterns, and pain scores, and then display it back as a calendar. At Asele, we built an AI layer that reasons over this data to surface patterns, flag anomalies, generate structured appointment summaries, and deliver cycle-aligned nutrition and fitness recommendations personalised to what the user is actively reporting.
This talk covers the practical architecture behind that. I will share how we classify health tasks by consequence severity to determine model choice and output constraints, how we structure and embed months of symptom history for pattern detection without tipping into diagnosis territory, and where general purpose LLMs fell short for clinically adjacent tasks and what we did about it. I will also share the safety framework behind these decisions, including how we measure anomaly detection performance, track false positives, and define human review triggers. For example, higher-risk symptom combinations, uncertainty beyond a defined threshold, or requests that move from wellbeing support into medical decision-making require additional safeguards rather than a direct AI response.
You will hear real decisions around evaluation, safety thresholds, human review triggers, and the constant tension between what AI can do and what it should do when the person on the other end is trying to understand their own body.
Getting this wrong has consequences. This talk is about what it takes to get it right.
Speaker
Key takeaways
- A practical framework for tiering AI health tasks by consequence severity and how that decision shapes model selection, prompt design, and output constraints
- How to structure longitudinal symptom data for LLM reasoning including chunking, embedding, pattern detection, and clinician ready summarisation without overclaiming
- Where off the shelf models break for sensitive health contexts and how to close the gap with evaluation, targeted safeguards, and human review triggers