AI alignment
AI alignment
The effort to ensure that an AI system's behaviour and outputs actually reflect the goals, values, and safety expectations of the people deploying it, rather than pursuing its trained objective in ways that produce harmful or unintended side effects. Alignment is an active area of research and remains an unresolved problem as AI systems grow more capable and autonomous. (See also: Reinforcement learning from human feedback, Agentic AI)