Model Extraction

Model extraction is an attack in which an adversary reconstructs an AI model's parameters, decision boundaries, or capabilities by querying it strategically, enabling model theft or preparation of downstream attacks.

How Model Extraction Works

An attacker with query access to a target model, whether through an API, a web interface, or an embedded application, sends carefully chosen inputs and observes the outputs. Over many queries, the attacker builds up enough information to reconstruct a functional copy of the model, or at minimum to characterize its behavior in enough detail to plan attacks.

Extraction is generally not perfect. The attacker rarely recovers the exact model weights. What they recover is enough functional similarity to steal IP, replicate a commercial capability, or design more effective adversarial attacks against the original.

Why Model Extraction Matters

For enterprises deploying proprietary models, extraction is IP theft. For enterprises using foundation models, extraction against the provider can affect every downstream deployment. For any AI system, extraction is often a precursor to more effective attacks, since the attacker's copy can be probed exhaustively for weaknesses without any risk to the attacker.

See MITRE ATLAS for how extraction is catalogued in the AI threat landscape.

Defenses Against Extraction

Common defenses include query rate limiting, anomaly detection on query patterns, output perturbation (adding noise that degrades extraction quality without meaningfully affecting utility), and access controls that limit exposure to unauthenticated or high-volume querying.

Related Terms

Full AI Visibility. Full Control. One Connected Platform.

Enterprise AI is expanding faster than most governance programs can track. Kovrr connects every AI signal across browser, endpoint, network, identity, and vendor systems into a single platform so security, governance, and risk teams work from the same evidence.