LLM Security: Jailbreaks and Prompt Injection

Haohan Wang's group at the University of Illinois Urbana-Champaign studies the security of large language models: how aligned models can be jailbroken into harmful behavior, why current guardrails fail, and how to build defenses that hold up. The work spans text, vision, and audio models, and includes prompt-injection defenses for deployed LLM applications. The group's InfoFlood attack was covered by 404 Media, POLITICO, IT Brew, and the Illinois News Bureau.

Key questions

What is an LLM jailbreak?

A jailbreak is an input crafted to make an aligned language model bypass its safety training and produce content it would normally refuse. Haohan Wang's group has shown several new classes of jailbreaks, including information overload (InfoFlood), cipher characters that evade moderation guardrails, self-jailbreaking where a model guides its own compromise, and narrative-style audio attacks on audio-language models.

Why is defending against prompt injection hard?

The group's ICML 2026 paper identifies a security–fidelity tradeoff: defenses resist injected instructions largely by suppressing untrusted text, which breaks tasks that must preserve that text, such as translation and document editing. Across 1,168 examples and 48 configurations, no model or defense achieved both goals; the most secure defenses reached 99.3% security but only 71.0–73.9% fidelity.

Do existing safety guardrails stop these attacks?

Often not. InfoFlood achieved up to 3 times higher jailbreak success than baseline attacks on GPT-4o, GPT-3.5-turbo, Gemini 2.0, and LLaMA 3.1, and common post-processing defenses including OpenAI's Moderation API, Perspective API, and SmoothLLM failed to mitigate it.

Selected projects

Security–Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense

Hermon M, Gupta R, Ruan W, Sabir E, Wang H · International Conference on Machine Learning (ICML), 2026 · Spotlight

Introduces SecFid, a benchmark in which executing an injection, processing it as data, and ignoring it produce distinguishable outputs, making fidelity measurable alongside security.

  • No model or defense achieves both security and fidelity across 1,168 examples and 48 configurations
  • Highest-fidelity model: 96.5% fidelity at 47.8% security; most secure defenses: 99.3% security at 71.0–73.9% fidelity
  • A decision-theoretic analysis shows the right behavior depends on the deployment's relative cost of a hijack versus a dropped span

InfoFlood: Jailbreaking Large Language Models with Information Overload

Yadav A, Jin H, Luo M, Zhuang J, Wang H · arXiv preprint, 2025

Identifies a new vulnerability: excessive linguistic complexity can disrupt built-in safety mechanisms without any added prefixes or suffixes. InfoFlood automatically rewrites malicious queries into information-overloaded ones and refines them when they fail.

  • Up to 3× higher jailbreak success than baselines on GPT-4o, GPT-3.5-turbo, Gemini 2.0, and LLaMA 3.1
  • OpenAI Moderation API, Perspective API, and SmoothLLM fail to mitigate the attack

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting

Kulshreshtha D, Su H, Jin H, Hegde C, Wang H · Conference on Language Modeling (COLM), 2026

Introduces self-jailbreaking, a threat model in which an aligned LLM guides its own compromise with no external attacker model, and SLIP, a black-box tree-search attack that implements it. Also proposes the Semantic Drift Monitor defense.

  • 90–100% attack success rate (average 94.7%) across most of eleven tested models, including GPT-5.1, Claude-Sonnet-4.5, Gemini-2.5-Pro, and DeepSeek-V3
  • About 7.9 LLM calls on average, 3–6× fewer than prior methods
  • Semantic Drift Monitor detects 76% of attacks at a 5% false-positive rate

Now You Hear Me: Audio Narrative Attacks Against Large Audio–Language Models

Yu Y, Jin H, Yu Y, Zhuang J, Wang H · European Chapter of the ACL (EACL), 2026

Shows that embedding disallowed directives in a narrative-style synthetic speech stream circumvents safety mechanisms calibrated mainly for text.

  • 98.26% success rate against state-of-the-art models including Gemini 2.0 Flash, well above text-only baselines

Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks

Zhou A, Li B, Wang H · Advances in Neural Information Processing Systems (NeurIPS), 2024 · Spotlight

An optimization-based defense that incorporates the adversary directly into the defensive objective and learns a lightweight, transferable suffix that adapts to worst-case attacks.

  • Reduces attack success rate on JailbreakBench to 6% on GPT-4 and 0% on Llama-2

Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters

Jin H, Zhou A, Menke JD, Wang H · Advances in Neural Information Processing Systems (NeurIPS), 2024

Introduces JAMBench, 160 manually crafted instructions across four risk categories designed to trigger moderation guardrails, and JAM, an attack that uses cipher characters from a shadow guardrail model to bypass output filters.

  • About 19.88× higher jailbreak success and about 1/6 the filtered-out rate of baselines across four LLMs

JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models

Jin H, Hu L, Li X, Zhang P, Chen C, Zhuang J, Wang H · arXiv preprint, 2024

A survey of jailbreaking and defenses for LLMs and vision-language models.

  • Categorizes jailbreaks into seven distinct types and maps the corresponding defenses

Talks on this topic

  • Security is Task Dependent: The Security–Fidelity Tradeoff in Prompt Injection Defense — AICE, UIUC (Apr. 2026)
  • Practical Concerns of Prompt Injection in Deployed Language Models — Capital One (Dec. 2025)
  • Detecting and Neutralizing Prompt Injection via LLM Activation Monitoring — AICE Center, UIUC (Oct. 2025)
  • Guardian of Trust in Language Models: Automatic Jailbreak and Systematic Defense — UIC (July 2024); Virginia Tech (Apr. 2024); AI Talks (Apr. 2024); UIUC LLM Reading Group (Mar. 2024); VALSE (Jan. 2024)
  • Technical Advances of Jailbreaks — Protiviti (May 2024)
  • Tutorial on Trustworthy Machine Learning — ACM SIGKDD 2023

Research support

This work is supported by the Amazon-Illinois Center on AI for Interactive Conversational Experiences (AICE) and the NAIRR Pilot.

Further reading from the DREAM Lab

Invite a talk or collaborate

Haohan Wang gives talks on LLM jailbreaks, prompt-injection defense, and AI safety evaluation for academic, industry, and policy audiences, and welcomes collaborations on securing deployed LLM and agent systems. Contact Haohan Wang at haohanw at illinois.edu.