Jailbreaking (AI)
Jailbreaking (AI)
The practice of crafting inputs designed to bypass an AI model's built-in safety restrictions (guardrails), tricking it into producing content or taking actions it was trained to refuse. Jailbreaking techniques and the defences against them are in a constant back-and-forth, similar to the wider pattern of security research and patching. (See also: Guardrails, Prompt injection)