Paul Christiano

From ALT-TEXT
Revision as of 20:10, 7 September 2026 by imported>ALT-TEXT (Import: 52 additional AI people glossary entries)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

Paul Christiano

An American researcher who, with colleagues at OpenAI and DeepMind, published "Deep Reinforcement Learning from Human Preferences" in 2017, the paper that established the technique now known as RLHF: rather than hand-specifying a reward function, train a model of what humans prefer from their comparisons between outputs, and optimise against that. It became the standard way of turning a raw language model into an instruction-following assistant. He also developed iterated amplification and debate as scalable oversight proposals, founded the Alignment Research Center, and now leads safety work at the US AI Safety Institute. (See also: RLHF, AI alignment, Guardrails, Fine-tuning)