
Sebastien Bubeck
VP of GenAI Research, Microsoft | Author of Sparks of AGI & Phi Model Architect
Vice President of GenAI Research at Microsoft and former Professor at Princeton University. Renowned for leading the milestone 'Sparks of Artificial General Intelligence' empirical research team studying early GPT-4 checkpoints. Pioneer of the Phi series of Small Language Models (Phi-1, Phi-2, Phi-3, Phi-4), proving that high-quality synthetic educational data ('Textbooks are all you need') enables compact models to outcompete massive baselines.
Areas of focus
Professional niches
Proof of Work
Textbooks Are All You Need (The Phi Model Series)
Demonstrated that a 1.3-billion parameter transformer trained on 6 billion tokens of curated synthetic 'textbook quality' data achieved 50.6% on HumanEval, dramatically outperforming models 10x its size trained on hundreds of billions of uncurated web tokens.
Synthetic data diversity must be aggressively managed to prevent model collapse and repetition loops on out-of-distribution prompts.
Context: Phi-1 and Phi-2 autoregressive models trained on synthetic Python exercises and educational textbooks curated by GPT-3.5.
View missionSparks of Artificial General Intelligence: Early experiments with GPT-4
Authored the 154-page seminal investigation demonstrating that early non-multimodal GPT-4 exhibited emergent general intelligence across mathematics, coding, law, vision-via-code, psychology, and medicine without domain-specific fine-tuning.
Highlighted critical remaining flaws including lack of long-term planning, working memory degradation, and blind confidence in false premises.
Context: Unconstrained early GPT-4 base model checkpoint evaluated across novel out-of-distribution reasoning probes.
View mission