Tri Dao
Tri Dao
A Vietnamese-American researcher whose FlashAttention, published in 2022, rewrote the attention computation to account for how GPU memory actually works, avoiding writing the large intermediate attention matrix to slow high-bandwidth memory at all. The result is exact, not approximate, but dramatically faster and lighter, which is much of why long context windows became affordable. He has since produced FlashAttention-2 and 3, and co-created the Mamba state-space architecture with Albert Gu as an alternative to attention. He teaches at Princeton and is chief scientist at Together AI. (See also: Transformer, Context window, Inference, Compute)