Guardrails
Guardrails
Rules, filters, and safety mechanisms built around an AI model to prevent it from producing harmful, illegal, or unwanted outputs, or from taking unwanted actions. Guardrails can be built into the model itself through training (see Reinforcement learning from human feedback) or added externally as a separate layer that checks inputs and outputs. (See also: Jailbreaking, Reinforcement learning from human feedback)