Jailbreaking (AI)

From ALT-TEXT
Revision as of 10:36, 7 September 2026 by imported>ALT-TEXT (Import: AI terminology and people glossary)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

Jailbreaking (AI)

The practice of crafting inputs designed to bypass an AI model's built-in safety restrictions (guardrails), tricking it into producing content or taking actions it was trained to refuse. Jailbreaking techniques and the defences against them are in a constant back-and-forth, similar to the wider pattern of security research and patching. (See also: Guardrails, Prompt injection)