Dr. Sayash Kapoor
Princeton AI Researcher
Sources checkedAI researcher at Princeton University's Center for Information Technology Policy. Co-author of 'AI Snake Oil' (with Arvind Narayanan). Named to TIME100 in AI. Specializes in auditing machine learning failures, quantifying data leakage, and evaluating whether predictive AI systems deliver reproducible value in critical social domains.
Areas of focus
Professional niches
Proof of Work
Leakage and the Reproducibility Crisis in Machine-Learning-Based Science
Published a comprehensive meta-analysis of machine learning scientific literature discovering data leakage across 329 papers spanning 17 fields, establishing formal taxonomy and verification criteria to prevent train-test contamination.
Meta-analyses depend on publicly available replication code; studies with proprietary or non-shared datasets could not be independently tested.
Context: Reproducibility audit across computer vision, clinical healthcare predictions, and natural language processing.
View missionAI Snake Oil: What Computers Can, Can't, and Shouldn't Do
Developed the analytical distinction between perception AI (generative synthesis, speech recognition) and predictive AI (predicting human recidivism, job success), providing rigorous frameworks for evaluating claims of AI accuracy in business and governance.
Practical applications in enterprise require navigating existing regulatory guidelines that may lag behind empirical machine learning consensus.
Context: Statistical risk assessment, cross-entropy evaluation, algorithmic bias metrics in high-stakes human systems.
View mission