Dr. Dan Hendrycks
Executive Director, Center for AI Safety (CAIS) | Creator of MMLU Benchmark
Verified Proof of Work Artifacts
2 items catalogedEach artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Dr. Dan Hendrycks. Every entry undergoes editorial source verification.
Natural Adversarial Examples and Out-of-Distribution Detection in Neural Networks
A foundational safety study by Dr. Dan Hendrycks introducing the ImageNet-A and ImageNet-O benchmarks, exposing fundamental vulnerabilities where vision models fail catastrophically on naturally occurring adversarial examples.
Vision classification benchmark; physical world adversarial perturbations present additional environmental dynamics.
Measuring Massive Multitask Language Understanding (MMLU)
The industry-standard academic benchmark measuring world knowledge and problem-solving across 57 subjects ranging from elementary mathematics to professional law and medicine, authored by Dr. Dan Hendrycks.
Multiple-choice evaluation format does not directly measure multi-step agent reasoning, tool usage, or long-form generation coherence.