Tri Dao
VERIFIED TECHNICAL DOSSIERSources checked

Tri Dao

Chief Scientist, Together AI & Assistant Professor, CMU | Inventor of FlashAttention

2 Verified ArtifactsSource Checked & Attributed

Verified Proof of Work Artifacts

2 items cataloged

Each artifact below represents an authenticated research publication, production code repository, or technical architectural framework directly authored or co-created by Tri Dao. Every entry undergoes editorial source verification.

#1
RESEARCH Checked Sep 20, 2026

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

A foundational architecture co-authored by Tri Dao and Albert Gu introducing selective state space models (SSMs). Mamba achieves linear scaling in sequence length with 5x higher inference throughput than standard Transformers while matching performance across million-token language contexts.

Model & Execution Context:Custom CUDA kernels for selective scan operations; tested on NVIDIA A100 and H100 architectures.
Scope & Limitations

Requires specialized hardware-aware CUDA kernels for hardware utilization; memory benefits depend on hardware SRAM cache hierarchy.

#2
RESEARCH Checked Sep 20, 2026

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Seminal hardware-aware algorithm optimizing GPU SRAM and HBM memory reads/writes to compute exact attention in sub-quadratic memory, delivering 2-4x speedups across major foundation models.

Model & Execution Context:CUDA C++, Triton, PyTorch C++ extensions, and modern GPU architectures.
Scope & Limitations

Requires GPU compute capability >= 8.0 (NVIDIA Ampere, Ada, or Hopper) and compiled CUDA/Triton kernels.