Reinforcement learning from human feedback (RLHF)

From ALT-TEXT
Jump to navigation Jump to search

Reinforcement learning from human feedback (RLHF)

A training technique in which human reviewers rate or rank an AI model's outputs, and the model is adjusted to produce more of the highly rated responses and fewer of the poorly rated ones. RLHF is widely used to make large language models more helpful, less harmful, and more aligned with what users actually want, though it also embeds the values and judgement calls of whoever the reviewers are. (See also: AI alignment, Large language model)