Democratizing Machine Learning at Netflix: Building the Model Lifecycle Graph

About this session

As Netflix's machine learning footprint grew across dozens of teams and business verticals, models, features, datasets, and pipelines became increasingly siloed — effectively black boxes that even ML practitioners inside the company struggled to discover, understand, or reuse. Answering a simple question like "which experiments are running this model?" or "which models share these features?" required manual archaeology across disconnected systems.This talk covers how Netflix addressed this problem by building the Model Lifecycle Graph, a metadata-driven architecture that models ML entities — models, features, pipelines, experiments, and datasets — as interconnected nodes rather than isolated pipeline stages. Powered by a real-time Metadata Service (MDS), the graph enables lineage tracing, impact analysis, and cross-domain discovery at scale, turning ML metadata into a first-class infrastructure concern rather than an afterthought. I'll walk through the architectural decisions behind the graph model, the real-time ingestion design that keeps metadata fresh as ML assets change, and the tradeoffs of choosing a graph-oriented approach over traditional pipeline-centric tooling.

Related blog post - https://netflixtechblog.com/democratizing-machine-learning-at-netflix-building-the-model-lifecycle-graph-5cc6d5828bb1

Speaker

Key takeaways

  • ML metadata is a graph problem, not a catalog problem — entities and relationships as first-class infra.
  • Real-time ingestion keeps lineage current as models, features, and experiments change.
  • Graph-based discovery cuts duplicated work and enables true self-service ML.

Related sessions