Model Poisoning

Model poisoning is an attack that manipulates a model's weights, fine-tuning process, or update mechanism to embed backdoors, biases, or malicious behaviors directly into the trained model itself.

How Model Poisoning Differs from Data Poisoning

Both attacks aim to make the model behave maliciously. Data poisoning achieves this by corrupting training data, letting the model learn the bad behavior itself. Model poisoning achieves it by directly manipulating the model, either by tampering with weights, injecting malicious code into the training or fine-tuning pipeline, or compromising the model update process.

The distinction matters because defenses differ. Data poisoning defenses focus on data provenance and integrity. Model poisoning defenses focus on training pipeline integrity, weight verification, and supply chain controls on model artifacts.

Where Model Poisoning Happens

Model poisoning can enter through several paths. Malicious model weights uploaded to public model repositories affect anyone who downloads them. Compromised training infrastructure affects any model trained through it. Malicious fine-tuning APIs can inject specific behaviors during customization. Supply chain attacks on ML frameworks can affect every model trained with those frameworks.

Defenses Against Model Poisoning

Standard defenses include cryptographic verification of model weights before deployment, provenance requirements for models sourced from external providers, behavioral testing for known trigger patterns before production use, and applying AI supply chain risk controls to the full model lifecycle.

Related Terms

Full AI Visibility. Full Control. One Connected Platform.

Enterprise AI is expanding faster than most governance programs can track. Kovrr connects every AI signal across browser, endpoint, network, identity, and vendor systems into a single platform so security, governance, and risk teams work from the same evidence.