Paul Christiano
VERIFIED TECHNICAL DOSSIERSources checked

Paul Christiano

Founder, Alignment Research Center (ARC) | Pioneer of Reinforcement Learning from Human Feedback (RLHF)

2 Verified ArtifactsSource Checked & Attributed

Verified Proof of Work Artifacts

2 items cataloged

Each artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Paul Christiano. Every entry undergoes editorial source verification.

#1
IMPLEMENTATION Checked Sep 23, 2026

ARC Autonomous Replication & Threat Evaluation Framework (METR)

Created the standardized red-teaming benchmark testing whether frontier AI models possess autonomous cyber-offense capabilities, resource acquisition, self-replication, or model weight exfiltration skills.

Model & Execution Context:Sandboxed agentic execution environments testing GPT-4, Claude 3, and open weights models.
Scope & Limitations

Agent capabilities evolve rapidly with multi-agent orchestration, requiring continuous suite updates.

#2
RESEARCH Checked Sep 23, 2026

Deep Reinforcement Learning from Human Preferences (RLHF)

Authored the foundational paper establishing that training deep neural policies via pairwise human preference reward models allows agents to master complex tasks (Atari games and robotic backflips) without hard-coded programmatic reward functions.

Model & Execution Context:PPO actor-critic architectures and learned Bradley-Terry reward models.
Scope & Limitations

Reward models can be gamified by agents via reward hacking and sycophantic responses.