Guardrails

From ALT-TEXT
Revision as of 10:36, 7 September 2026 by imported>ALT-TEXT (Import: AI terminology and people glossary)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

Guardrails

Rules, filters, and safety mechanisms built around an AI model to prevent it from producing harmful, illegal, or unwanted outputs, or from taking unwanted actions. Guardrails can be built into the model itself through training (see Reinforcement learning from human feedback) or added externally as a separate layer that checks inputs and outputs. (See also: Jailbreaking, Reinforcement learning from human feedback)