XFollowListPeople, ideas and original work on X
Tri Dao’s avatar

Chief Scientist at Together AI; ML systems researcher

Tri Dao

@tri_dao on X

Studies how model algorithms and GPU hardware fit together, with coauthored explanations of faster attention and alternatives to a growing attention cache.

Why read Tri Dao?

Read Dao when you want to understand why running a model costs memory and time. The coauthored FlashAttention-3 explanation starts with moving small blocks through GPU memory, then shows how overlapping calculations and transfers can make attention faster. Its speed comparisons are tied to specified hardware and operations.

Continue with the coauthored Mamba-3 overview to compare a different way of carrying information through a sequence: a fixed-size state rather than an expanding attention cache. The authors explain the tradeoff between efficient decoding and retrieving earlier details. The later sections assume familiarity with model architectures and GPU execution.

Start with the original

Selected work

  1. Research explainer ·

    FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

    Why can the same attention calculation run faster? This coauthored explanation connects memory traffic, overlapping GPU operations and lower-precision numbers. Start with the recap, then inspect the hardware and precision used in each comparison.

  2. Research explainer ·

    Mamba-3 Part 1

    What does a fixed-size memory change? This coauthored overview explains Mamba-3's design, decoding costs and retrieval comparisons. It also shows why the authors explore hybrids that combine state-space layers with attention.

Projects & roles

  • FlashAttention

    The implementation project for the attention algorithms developed by Dao and collaborators.

  • Mamba

    Code for state-space sequence models; Dao is a coauthor of the Mamba research and the selected Mamba-3 explanation.

Community rating

85

Vote once every 24 hours. No account needed.

Sources & further reading