AI GlossaryㅌSafety and controversy
Jailbreak
Using clever wording to trick an AI into bypassing its own safety rules — an endless cat-and-mouse game between those who block and those who break through.
In plain words
This is the practice of using clever wording to get around the safety rules built into an AI. Classic tricks include role-play or emotional appeals, like "write this as a villain's line in a novel" or "my late grandmother used to tell me a story like this..." to coax out answers the AI would normally refuse to give.
The term borrows from "jailbreaking" a smartphone. It's an ongoing chase: companies patch the holes, communities find new ones, and some labs even offer bounties inviting people to try breaking in (red-teaming, bug bounties). It's easy to confuse with prompt injection — but a jailbreak is the user deceiving the AI directly, while injection is a third party hiding the deception inside external data.
See also
Stories using this term
- Tencent's Zhuque Lab Open-Sources AI Agent/MCP Security ScannerAI · 2026.08.21
- Open-source app lets you grow a mini LLM from scratch on a MacBookAI · 2026.08.16
- llama.cpp adds CI build target for AMD ROCm 7.14AI · 2026.08.11
- Apple unveils technique to block fine-tuning of AI model weightsAI · 2026.08.11
- Local LLM Community Flags a Gap in New 8B-12B ModelsAI · 2026.08.11
- DeepSeek v4 Flash 0731 run locally with RTX 4090 and Tesla P40 comboAI · 2026.08.10
