A Transcript Is Not a Person: Engineering Identity Into AI Clones
About this session
Delphi builds digital versions of people. The first answer is the product: if someone's Delphi cannot speak from their life on day one, trust never recovers. We scaled content training from about 5K to about 100K items per day at 99.8% workflow success using Temporal durable workflows so crashed workers resume instead of losing half-trained minds. Attribution is the silent failure: multi-hour podcasts mix speakers, so we use onboarding voice clips to train only on the right segments. Agents grow a temporal knowledge graph for retrieval, and we borrow stylometry to score whether a Delphi matches its person (about 70% on real conversations vs about 3% chance). A practitioner talk on pipelines, attribution, graphs, and identity evals.
Speaker
Key takeaways
- How to design durable content pipelines (Temporal) that survive IP bans, rate limits, and mid-job crashes while scaling training throughput.
- Why speaker attribution on multi-hour podcasts is the silent failure mode, and how onboarding voice clips keep training on the right person.
- How to use stylometry as an identity eval so release gates catch clones that drift toward the base model.