Paul Christiano
Paul Christiano
An American researcher who, with colleagues at OpenAI and DeepMind, published "Deep Reinforcement Learning from Human Preferences" in 2017, the paper that established the technique now known as RLHF: rather than hand-specifying a reward function, train a model of what humans prefer from their comparisons between outputs, and optimise against that. It became the standard way of turning a raw language model into an instruction-following assistant. He also developed iterated amplification and debate as scalable oversight proposals, founded the Alignment Research Center, and now leads safety work at the US AI Safety Institute. (See also: RLHF, AI alignment, Guardrails, Fine-tuning)