Tri Dao

Chief Scientist, Together AI & Assistant Professor, CMU | Inventor of FlashAttention

Sources checked
ABOUT

Tri Dao is the Chief Scientist at Together AI and an Assistant Professor in the Computer Science Department at Carnegie Mellon University. He is the principal inventor of FlashAttention and FlashAttention-2, breakthrough hardware-aware exact attention algorithms that transformed modern LLM training and inference efficiency. He is also the co-inventor of Mamba, pioneering linear-time state-space models.

Areas of focus

Professional niches

THE WORK BEHIND THE PROFILE

Proof of Work

Add work ↗
researchChecked Sep 20, 2026

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Seminal hardware-aware algorithm optimizing GPU SRAM and HBM memory reads/writes to compute exact attention in sub-quadratic memory, delivering 2-4x speedups across major foundation models.

Scope & limitations

Requires GPU compute capability >= 8.0 (NVIDIA Ampere, Ada, or Hopper) and compiled CUDA/Triton kernels.

Context: CUDA C++, Triton, PyTorch C++ extensions, and modern GPU architectures.

View mission

Guides to evaluating AI expertise