Study how model behavior can fail under adversarial conditions. Define a threat model, reproduce a bounded robustness experiment on a model you control, and doc
Last reviewed: 2026-10-03
Adversarial machine learning examines failures caused by deliberate manipulation of a model's inputs, training process, or surrounding system. NIST organizes the subject by lifecycle stage, attacker goals, capabilities, and knowledge. A useful assessment starts by stating those conditions explicitly.
Practice on a model and dataset you control. Compare clean and perturbed behavior within a bounded threat model, then evaluate a proposed mitigation under the same conditions. Passing one experiment does not establish security against other manipulations or deployment contexts.
No. Results depend on the tested threat model, data, and system. Document what was tested and what remains outside the evidence.