Tri Dao
Chief Scientist, Together AI & Assistant Professor, CMU | Inventor of FlashAttention
Sources checkedTri Dao is the Chief Scientist at Together AI and an Assistant Professor in the Computer Science Department at Carnegie Mellon University. He is the principal inventor of FlashAttention and FlashAttention-2, breakthrough hardware-aware exact attention algorithms that transformed modern LLM training and inference efficiency. He is also the co-inventor of Mamba, pioneering linear-time state-space models.
Areas of focus
Professional niches
Proof of Work
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Seminal hardware-aware algorithm optimizing GPU SRAM and HBM memory reads/writes to compute exact attention in sub-quadratic memory, delivering 2-4x speedups across major foundation models.
Requires GPU compute capability >= 8.0 (NVIDIA Ampere, Ada, or Hopper) and compiled CUDA/Triton kernels.
Context: CUDA C++, Triton, PyTorch C++ extensions, and modern GPU architectures.
View mission