Stella Biderman
Executive Director of EleutherAI
Sources checkedExecutive Director at EleutherAI, the non-profit AI research lab that pioneered open-source large language model pre-training. Led the development of Pythia, GPT-NeoX-20B, and the Pile dataset. Former AI Researcher at Booz Allen Hamilton. Major contributor to foundation model interpretability, evaluation standards, and data governance.
Areas of focus
Professional niches
Proof of Work
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Directed the development and release of the Pythia model suite: 16 LLMs ranging from 70M to 12B parameters, all trained on identical data ordering with 154 intermediate checkpoints public for scientific analysis of memorization and bias.
Pre-training dataset contained early web corpora reflecting historical internet text distributions prior to modern RLHF curation.
Context: GPT-NeoX distributed training framework, FlashAttention, 300B tokens of the Pile dataset.
View missionThe Pile: An 800GB Dataset of Diverse Text for Language Modeling
Co-authored the curation and statistical characterization of The Pile, an 825GB open-source pre-training dataset combining 22 diverse text subsets (arXiv, PubMed, GitHub, StackExchange), establishing the foundation for modern open foundation models.
Includes academic text requiring high domain literacy; copyright and licensing complexities in mixed web-scraped sub-components.
Context: FastText deduplication, exact substring filtering, perplexity benchmark analysis.
View mission