Alec Radford
Research Scientist & Lead Architect
Research Scientist at OpenAI and primary architect of modern generative AI. Lead author of the original GPT (Improving Language Understanding), GPT-2, CLIP, and Whisper, establishing the scaling paradigms for unsupervised pre-training and contrastive vision-language representation.
Areas of focus
Professional niches
Proof of Work
Learning Transferable Visual Models From Natural Language Supervision (CLIP)
Authored the foundational ICML 2021 paper introducing CLIP, which trained dual vision and text encoders via symmetric cross-entropy contrastive loss on 400M image-text pairs, unlocking robust zero-shot image classification and powering Stable Diffusion.
Fine-grained spatial reasoning, counting, and typographic reading struggle without explicit spatial bounding box supervision.
Context: Vision Transformers (ViT-B/16, ViT-L/14) and ResNet baselines, 400M image-text dataset.
View missionRobust Speech Recognition via Large-Scale Weak Supervision (Whisper)
Architected Whisper, an open-weights sequence-to-sequence Transformer trained on 680,000 hours of multilingual audio, establishing zero-shot robustness across accents, background noise, and automated timestamp generation.
Long audio files can experience hallucination loops or timestamp drift during extended periods of ambient silence.
Context: Seq2Seq Transformer, 16kHz log-mel spectrogram audio input, 99 language recognition.
View mission