Model extraction is an attack that steals model capability or assets through interaction: systematically querying an API to clone behavior into a copycat model,
An attack that steals model capability or assets through interaction: systematically querying an API to clone behavior into a copycat model, or coaxing a system into revealing its proprietary prompts, training-data fragments, or embedded secrets.
By combining rate limiting and anomaly detection, output watermarking, contractual terms, and prompt-secrecy hygiene, never putting secrets in system prompts. For high-value models, extraction risk shapes how much capability is exposed per tier.