Tri Dao

Chief Scientist, Together AI & Assistant Professor, CMU | Inventor of FlashAttention

Sources checked
ABOUT

Tri Dao is the Chief Scientist at Together AI and an Assistant Professor in the Computer Science Department at Carnegie Mellon University. He is the principal inventor of FlashAttention and FlashAttention-2, breakthrough hardware-aware exact attention algorithms that transformed modern LLM training and inference efficiency. He is also the co-inventor of Mamba, pioneering linear-time state-space models.

Areas of focus

Professional niches

THE WORK BEHIND THE PROFILE

Proof of Work

researchChecked Sep 20, 2026

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

A foundational architecture co-authored by Tri Dao and Albert Gu introducing selective state space models (SSMs). Mamba achieves linear scaling in sequence length with 5x higher inference throughput than standard Transformers while matching performance across million-token language contexts.

Scope & limitations

Requires specialized hardware-aware CUDA kernels for hardware utilization; memory benefits depend on hardware SRAM cache hierarchy.

Context: Custom CUDA kernels for selective scan operations; tested on NVIDIA A100 and H100 architectures.

View mission
researchChecked Sep 20, 2026

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Seminal hardware-aware algorithm optimizing GPU SRAM and HBM memory reads/writes to compute exact attention in sub-quadratic memory, delivering 2-4x speedups across major foundation models.

Scope & limitations

Requires GPU compute capability >= 8.0 (NVIDIA Ampere, Ada, or Hopper) and compiled CUDA/Triton kernels.

Context: CUDA C++, Triton, PyTorch C++ extensions, and modern GPU architectures.

View mission

Guides to evaluating AI expertise