Jailbreak

An AI jailbreak is an input crafted specifically to bypass a model's safety controls, guardrails, or usage policies, causing the model to produce content or take actions the developer had blocked.

How Jailbreaks Work

AI models, particularly LLMs, are typically trained and configured with safety filters that block specific categories of output (harmful content, sensitive information, restricted actions). A jailbreak is an input that circumvents those filters, often by exploiting how the model interprets context, roles, or instructions.

Common jailbreak patterns include role-play framing (asking the model to pretend to be an unrestricted system), obfuscation (encoding requests to evade filters), context flooding (overwhelming the model with content that shifts its behavior), and progressive escalation (small requests that gradually cross safety lines).

Jailbreaks and Enterprise Risk

Jailbreaks matter to enterprises because they expose failure modes in AI systems that are supposed to be safe. If a customer service AI can be jailbroken into producing offensive content, that becomes a customer-facing incident. If an internal AI can be jailbroken into revealing information from restricted sources, that becomes a data exposure.

See how organizations should prioritize AI security risks.

Jailbreak vs. Prompt Injection

Jailbreaks and prompt injection overlap but are not identical. Jailbreaks target safety controls, causing the model to violate its intended usage policies. Prompt injection targets the AI system's instruction-following, causing it to act on attacker-controlled instructions. Some attacks combine both.

Related Terms

Full AI Visibility. Full Control. One Connected Platform.

Enterprise AI is expanding faster than most governance programs can track. Kovrr connects every AI signal across browser, endpoint, network, identity, and vendor systems into a single platform so security, governance, and risk teams work from the same evidence.