
Paul Christiano
Founder, Alignment Research Center (ARC) | Pioneer of Reinforcement Learning from Human Feedback (RLHF)
Verified Proof of Work Artifacts
2 items catalogedEach artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Paul Christiano. Every entry undergoes editorial source verification.
ARC Autonomous Replication & Threat Evaluation Framework (METR)
Created the standardized red-teaming benchmark testing whether frontier AI models possess autonomous cyber-offense capabilities, resource acquisition, self-replication, or model weight exfiltration skills.
Agent capabilities evolve rapidly with multi-agent orchestration, requiring continuous suite updates.
Deep Reinforcement Learning from Human Preferences (RLHF)
Authored the foundational paper establishing that training deep neural policies via pairwise human preference reward models allows agents to master complex tasks (Atari games and robotic backflips) without hard-coded programmatic reward functions.
Reward models can be gamified by agents via reward hacking and sycophantic responses.