
Paul Christiano
Founder, Alignment Research Center (ARC) | Pioneer of Reinforcement Learning from Human Feedback (RLHF)
Founder of the Alignment Research Center (ARC) and creator of ARC Evals (now METR). Former lead of the language model alignment team at OpenAI, where he pioneered Reinforcement Learning from Human Feedback (RLHF)—the core technology enabling InstructGPT, ChatGPT, Claude, and modern aligned LLMs. Appointed as head of AI safety for the US AI Safety Institute (NIST).
Areas of focus
Professional niches
Proof of Work
ARC Autonomous Replication & Threat Evaluation Framework (METR)
Created the standardized red-teaming benchmark testing whether frontier AI models possess autonomous cyber-offense capabilities, resource acquisition, self-replication, or model weight exfiltration skills.
Agent capabilities evolve rapidly with multi-agent orchestration, requiring continuous suite updates.
Context: Sandboxed agentic execution environments testing GPT-4, Claude 3, and open weights models.
View missionDeep Reinforcement Learning from Human Preferences (RLHF)
Authored the foundational paper establishing that training deep neural policies via pairwise human preference reward models allows agents to master complex tasks (Atari games and robotic backflips) without hard-coded programmatic reward functions.
Reward models can be gamified by agents via reward hacking and sycophantic responses.
Context: PPO actor-critic architectures and learned Bradley-Terry reward models.
View mission