Tri Dao

From ALT-TEXT
Revision as of 20:10, 7 September 2026 by imported>ALT-TEXT (Import: 52 additional AI people glossary entries)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

Tri Dao

A Vietnamese-American researcher whose FlashAttention, published in 2022, rewrote the attention computation to account for how GPU memory actually works, avoiding writing the large intermediate attention matrix to slow high-bandwidth memory at all. The result is exact, not approximate, but dramatically faster and lighter, which is much of why long context windows became affordable. He has since produced FlashAttention-2 and 3, and co-created the Mamba state-space architecture with Albert Gu as an alternative to attention. He teaches at Princeton and is chief scientist at Together AI. (See also: Transformer, Context window, Inference, Compute)