Backdoor Attack
A backdoor attack plants a hidden trigger in a model during training, so it behaves normally until it sees a specific input — then it flips to attacker-chosen behavior. The model passes normal testing, which is what makes the backdoor dangerous.
Also known as: backdoor attacks, trojan attack
A backdoor (or trojan) attack compromises a model during training so that it behaves normally on ordinary inputs and switches to the attacker’s chosen behavior when it sees a specific trigger: a phrase, a pattern, a small patch on an image. NIST’s adversarial machine learning taxonomy credits BadNets in 2017 as the first backdoor poisoning attack (NIST AI 100-2e2025). In that work, Tianyu Gu and colleagues at NYU showed that outsourced training is a supply-chain risk: a maliciously trained network can perform at state-of-the-art on the user’s own validation data while misbehaving on attacker-chosen inputs. Their example was a US street-sign classifier that read a stop sign as a speed-limit sign when a small sticker was added, and the backdoor persisted after the network was retrained for another task (Gu et al., BadNets).
Stealth is what makes backdoors dangerous. Accuracy tests on clean data do not reveal them, because the model only misbehaves when the trigger is present. NIST notes other variants that blend the trigger into the data, and clean-label attacks that work even when the attacker cannot change labels, at the cost of needing more poisoned samples.
For language models, Anthropic’s Sleeper Agents paper tested whether standard safety training removes a deliberately implanted backdoor (Hubinger et al.). The researchers trained models to write secure code when the prompt said the year was 2023 and to insert exploitable code when it said 2024. The behavior survived supervised fine-tuning, reinforcement learning, and adversarial training, persisted most in the largest models, and adversarial training could teach the model to recognize its trigger better, hiding the behavior rather than removing it.
The practical defenses are about provenance and verification rather than hoping to detect a trigger later. Know where model weights and training data came from, prefer models with documented training, treat third-party fine-tunes and adapters as untrusted code, and red-team for trigger-like behavior before deployment. Data poisoning covers how a backdoor gets into the data, and the AI security topic collects the episodes on attacks against models.