Sebastien Bubeck

VP of GenAI Research, Microsoft | Author of Sparks of AGI & Phi Model Architect

Sources checked
ABOUT

Vice President of GenAI Research at Microsoft and former Professor at Princeton University. Renowned for leading the milestone 'Sparks of Artificial General Intelligence' empirical research team studying early GPT-4 checkpoints. Pioneer of the Phi series of Small Language Models (Phi-1, Phi-2, Phi-3, Phi-4), proving that high-quality synthetic educational data ('Textbooks are all you need') enables compact models to outcompete massive baselines.

Areas of focus

Professional niches

THE WORK BEHIND THE PROFILE

Proof of Work

researchChecked Sep 23, 2026

Textbooks Are All You Need (The Phi Model Series)

Demonstrated that a 1.3-billion parameter transformer trained on 6 billion tokens of curated synthetic 'textbook quality' data achieved 50.6% on HumanEval, dramatically outperforming models 10x its size trained on hundreds of billions of uncurated web tokens.

Scope & limitations

Synthetic data diversity must be aggressively managed to prevent model collapse and repetition loops on out-of-distribution prompts.

Context: Phi-1 and Phi-2 autoregressive models trained on synthetic Python exercises and educational textbooks curated by GPT-3.5.

View mission
researchChecked Sep 23, 2026

Sparks of Artificial General Intelligence: Early experiments with GPT-4

Authored the 154-page seminal investigation demonstrating that early non-multimodal GPT-4 exhibited emergent general intelligence across mathematics, coding, law, vision-via-code, psychology, and medicine without domain-specific fine-tuning.

Scope & limitations

Highlighted critical remaining flaws including lack of long-term planning, working memory degradation, and blind confidence in false premises.

Context: Unconstrained early GPT-4 base model checkpoint evaluated across novel out-of-distribution reasoning probes.

View mission

Guides to evaluating AI expertise