What is an AI jailbreak attack?
A malicious exploit designed to bypass the built-in safeguards of an AI system for the purpose of causing the system to perform actions or generate content it would not under normal circumstances
A jailbreak attack is a malicious exploit designed to bypass the built-in safeguards of an AI system. The goal is to cause the system to generate inaccurate or harmful output, or to perform unauthorized actions. Jailbreaking is a form of prompt injection specifically designed to bypass a model’s safety controls, while prompt injection more broadly manipulates model behavior through malicious or untrusted instructions.
Common jailbreak techniques and vectors
Attackers use several techniques and vectors for jailbreaking. These are the most common approaches:
Roleplay and persona manipulation: Through an AI model prompt, attackersinstruct the model to adopt a persona that will disregard built-in safeguards. For example, attackers have instructed models to adopt the persona “DAN” (“do anything now”), which might deliver uncensored responses to prompts. Similarly, the persona “AIM” (“always intelligent and Machiavellian”) has been used to provide amoral answers that cross ethical boundaries.
Obfuscation and encoding: This jailbreak techniquetries to transform a malicious payload into a format that slips past any text-based safety filters or classifiers. Attackers have used Base64 code instead of natural language, ciphers (which use algorithms to shift or replace normal content), invisible characters, and characters from other alphabets. These modifications can be decoded and read by the models but not their safety filters.
Affirmative response forcing: Also called prefix injection or forced completion, this technique forces a model to start each response with an affirmative statement, like “Sure, I can help with that!” Once the model generates that response, its probabilistic approach makes it more likely to complete subsequent requests, even requests that would normally be refused.
Context flooding: For this technique, which is sometimes called context stuffing or attention saturation, attackers include a large amount of benign text or random data as part of their prompt. Large amounts of additional context can dilute or interfere with the influence of earlier instructions, potentially making it harder for a model to consistently follow its safety constraints.
The evolution of the threat
Initially, the primary targets of AI jailbreaking were large language models (LLMs), and the harmful results were largely limited to inappropriate text responses. But as organizations have increasingly adopted AI agents to handle an array of tasks autonomously, attackers have shifted their targets to those agents.
AI agent jailbreaking can have serious consequences for enterprises. Jailbroken agents with API access could, for example, exfiltrate sensitive data, delete database files, execute fraudulent financial transactions, send out phishing messages from corporate accounts, or crash critical systems.
Attackers are using AI themselves to find agent vulnerabilities that will improve their success with jailbreaking. For example, they are using large reasoning models (LRMs) to discover vulnerabilities. They can then deploy automated red teaming agents powered by those LRMs to continuously test agent vulnerabilities.
Why traditional security falls short
Traditional security solutions, such as web application firewalls (WAFs), are insufficient for preventing AI jailbreaks for two key reasons. First, WAFs examine structured data and code for known threat patterns. However, with jailbreaking, attackers use natural-language prompts to compel models and agents to execute actions. WAFs can’t decipher the meaning in that natural language. In addition, WAFs primarily inspect HTTP application traffic and do not have visibility into every source of content an AI system may ingest.
How to defend AI applications and agents
Defending AI applications and agents from jailbreaking requires a multi-layered strategy. Four essential practices can help build that strategy:
1. Implement dual (input/output) guardrails: Input guardrails can inspect prompts for obfuscation and encoding, and filter out attempts at roleplaying. Output guardrails, meanwhile, can evaluate generated output before it reaches the user and before the model executes any API calls. That evaluation might look for hate speech, data leakage, and unauthorized tool calls.
2. Deploy sentinel or moderation models: Instead of requiring a model to evaluate its own safety, organizations can deploy lightweight “sentinel” or moderation models that are designed for classification. Examining both inputs and outputs, these models can identify attempts at hijacking within prompts and determine whether generated ouput violates corporate policies (such as policies on hate speech or harassment).
3. Enforce system prompt fidelity: System prompt fidelity is the extent to which prompts adhere to system instructions without drifting or being overridden by other instructions. Organizations can enforce system prompt fidelity through multiple techniques, such as repeating core system instructions during user conversations (which combats context flooding) and fine-tuning instructional hierarchy, making sure the model will reject user commands to ignore system instructions.
4. Employing least-privilege API access for agents: Jailbroken AI agents can cause significant harm to enterprises, especially when those agents are connected to critical systems. Organizations can limit the potential damage of agent jailbreaking by employing least-privilege API access. Agents should be able to access only certain APIs required to complete their tasks.
Frequently asked questions
Why do AI models refuse benign questions (producing “false positives”)?
AI models sometimes refuse to answer benign user questions. These refusals are called “false positives” or “over-refusal.” The prompt might contain trigger words that signal danger in other contexts or contain framing that suggests an attempt at jailbreaking (like asking the model to adopt a persona). Or the model might calculate that its answer would likely generate a harmful request, which would be worse than refusing to answer.
How does character roleplaying trick safety guardrails?
Character roleplaying tricks safety guardrails by disguising malicious intent in seemingly benign context. If the guardrails view the prompt as only a creative exercise, they will not classify it as harmful. Meanwhile, the model will use its resources to comply with the command to take on a role and concentrate less on enforcing safety rules.
What is the difference between a “DAN” prompt and a universal jailbreak?
The infamous “DAN” prompt refers to the roleplay technique of telling a model to adopt a persona that will ignore safety rules (“do anything now”) and simply comply with user-provided instructions. A universal jailbreak attack uses a computer algorithm to create a string of characters that, when added to any prompt, overrides the model’s safety rules.
What are the legal risks and liabilities of jailbreaking AI systems?
Jailbreaking AI systems can create multiple risks and liabilities for the organizations running those AI systems. For example, jailbreaking could result in the exposure of sensitive customer data, dissemination of false information, compliance violations, or a variety of unlawful actions undertaken by autonomous agents. Organizations could face regulatory fines, lawsuits, and even criminal actions.
Can API safety settings completely stop jailbreak attempts?
No, API safety settings cannot completely stop jailbreak attempts. Attackers can bypass API filters using Base64 encoding or similar obfuscation techniques. In addition, API filters often fail to detect instructions disguised as ordinary text. Moreover, indirect jailbreak attempts hidden in third-party content can be particularly difficult to detect and require additional controls around external or retrieved content.
How do you secure AI systems and applications during inference and at runtime?
Securing AI systems during inference and at runtime requires controls that inspect prompts, responses, agent actions, and tool use as interactions occur. AI-specific runtime security can help block prompt injection, prevent sensitive data leakage, enforce policy, and govern how models and agents interact with users, data, APIs, and connected tools.
What F5 capabilities support enterprise AI security?
F5 AI Guardrails helps enforce runtime policies and protect against risks such as prompt injection, data leakage, harmful outputs, and unsafe agent actions. F5 AI Red Team uses adversarial testing to identify vulnerabilities and weaknesses before attackers find them. F5 AI Gateway provides a centralized control point for governing access to models, agents, and tools, while F5 Workforce AI Security provides visibility into sanctioned and unsanctioned AI use across the enterprise. Together, these capabilities form part of the F5 AI Security Platform. For broader protection across applications, APIs, and distributed environments, the F5 Application Delivery and Security Platform (ADSP) also brings together capabilities including web application and API protection, bot management, DDoS mitigation, and application delivery.

