Paul Christiano

Founder, Alignment Research Center (ARC) | Pioneer of Reinforcement Learning from Human Feedback (RLHF)

Sources checked
ABOUT

Founder of the Alignment Research Center (ARC) and creator of ARC Evals (now METR). Former lead of the language model alignment team at OpenAI, where he pioneered Reinforcement Learning from Human Feedback (RLHF)—the core technology enabling InstructGPT, ChatGPT, Claude, and modern aligned LLMs. Appointed as head of AI safety for the US AI Safety Institute (NIST).

Areas of focus

Professional niches

THE WORK BEHIND THE PROFILE

Proof of Work

implementationChecked Sep 23, 2026

ARC Autonomous Replication & Threat Evaluation Framework (METR)

Created the standardized red-teaming benchmark testing whether frontier AI models possess autonomous cyber-offense capabilities, resource acquisition, self-replication, or model weight exfiltration skills.

Scope & limitations

Agent capabilities evolve rapidly with multi-agent orchestration, requiring continuous suite updates.

Context: Sandboxed agentic execution environments testing GPT-4, Claude 3, and open weights models.

View mission
researchChecked Sep 23, 2026

Deep Reinforcement Learning from Human Preferences (RLHF)

Authored the foundational paper establishing that training deep neural policies via pairwise human preference reward models allows agents to master complex tasks (Atari games and robotic backflips) without hard-coded programmatic reward functions.

Scope & limitations

Reward models can be gamified by agents via reward hacking and sycophantic responses.

Context: PPO actor-critic architectures and learned Bradley-Terry reward models.

View mission

Guides to evaluating AI expertise