XFollowListPeople, ideas and original work on X

AI reading list

LLM research & explanations

Large language models learn patterns in text and generate sequences of tokens — pieces of text. To understand what they can do, follow the architecture, training and tests behind a result. Start with an explanation or small implementation, then compare the original experiments. Each person here offers a specific route into that work.

Selected sources
12

People to read

Open the first material, then continue with the author on X. People appear once, grouped by their main reading use and ordered by handle within each group.

3 of 12 people

Training & model systems

See how a model is trained, adapted and run. Explore feedback, released checkpoints and the memory and hardware choices behind a training recipe.

  • Understand what happens after pretraining. Lambert's book distinguishes training on examples, preferred responses and verifiably correct answers; the team Tülu 3 paper shows how those stages fit into a released recipe with separate development and held-out tests.

    Start with this Web material · rlhfbook.com

    A map of language-model post-training

    Start with the recipe descriptions to separate examples, preference feedback and checked rewards. You can follow that overview before studying the equations.

    Material checked . Authorship source.

  • Rush connects coding models to their training and evaluation. His Composer 2 introduction leads to the coauthored report on continued pretraining, reinforcement learning and tests drawn from engineering work. His Annotated Transformer provides a deeper route into the architecture through Python code.

    Start with this Web material · cursor.com

    A technical report on Composer 2

    Rush introduces two training stages and explains why the team's evaluations include ambiguous requests and changes across several files. Continue to the linked report for the experiments and evaluation setup behind the claims.

    Published . Material checked .

  • Dao and collaborators explain how memory movement and GPU scheduling affect model speed. Start with the FlashAttention-3 recap, then explore Mamba-3 to compare a fixed-size state with an expanding attention cache. The examples connect algorithm choices to hardware and retrieval tradeoffs.

    Start with this Web material · tridao.me

    FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

    The coauthored recap explains processing attention in small memory blocks. Later sections show overlapping GPU operations and precision tradeoffs. Check the hardware and operation measured before applying a speed comparison to your own model.

    Published . Material checked .

How this list is selected

We selected people through their own explanations, implementations, interviews and coauthored research. Each card links an original starting material and explains the question it answers. Tutorials, research reviews and experimental papers serve different reading needs; the notes identify useful prerequisites and preserve team authorship. Roles come from current primary biographies. The same person keeps one profile across AI, crypto and coding topics.