Nathan Lambert, Ph.D.
Allen Institute for AI (Ai2)
Verified Proof of Work Artifacts
2 items catalogedEach artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Nathan Lambert, Ph.D.. Every entry undergoes editorial source verification.
RewardBench: Evaluating Reward Models for Language Model Alignment
Created RewardBench, the industry's primary benchmark for evaluating reward models and preference classifiers across chat capabilities, reasoning, and adversarial safety edge cases.
Static evaluation test sets face saturation as training data mixtures incorporate benchmark distributions.
Tülu 3: Pushing Frontiers in Open Language Model Post-Training
Co-led the Tülu 3 open post-training initiative, releasing complete training datasets, recipes, and checkpoint weights demonstrating open-source parity with proprietary frontier models through RLVR and DPO.
Verifiable rewards require deterministic verifiers (math/code compilers), limiting RLVR application in subjective writing.