AI reading list
LLM research & explanations
Large language models learn patterns in text and generate sequences of tokens — pieces of text. To understand what they can do, follow the architecture, training and tests behind a result. Start with an explanation or small implementation, then compare the original experiments. Each person here offers a specific route into that work.
- Selected sources
- 12
People to read
Open the first material, then continue with the author on X. People appear once, grouped by their main reading use and ordered by handle within each group.
12 people
Foundations & explanations
Connect tokens, attention and a Transformer to code you can inspect. The tutorials introduce the pieces; the original paper requires more mathematical preparation.
-
Polosukhin coauthored Attention Is All You Need, the original Transformer paper, before founding NEAR. Read the architecture and attention sections to connect today’s language models to that work. The experiments measure translation and parsing; they preserve a specific context for the results.
Start with this Web material · arxiv.org
Coauthored research. Start with Figure 1 and section 3 on the encoder, decoder and attention, then inspect the translation experiments in section 6. Familiarity with neural networks and matrix operations helps.
Published . Material checked .
-
Karpathy explains neural networks through implementations you can build and inspect. His Zero to Hero course starts with backpropagation, develops language modeling and builds a GPT. Follow the earlier language-modeling and PyTorch lessons before the Transformer lecture.
Start with this Web material · karpathy.ai
Choose a lesson through the syllabus. Begin with micrograd if backpropagation is new; the GPT lesson assumes earlier language-modeling and PyTorch material. The course requires solid Python and introductory math.
Material checked .
-
Raschka connects model diagrams to small implementations you can inspect. Begin with his attention tutorial, then follow LLMs from Scratch from tokens to a GPT-style model, pretraining and fine-tuning. Python knowledge helps; an introductory appendix covers PyTorch basics.
Start with this Web material · sebastianraschka.com
Understanding and Coding the Self-Attention Mechanism of Large Language Models From Scratch
Follow a six-word sentence from vectors through queries, keys and values to a context vector. The diagrams and PyTorch calculations explain attention. Continue with LLMs from Scratch for the model and training loop.
Published . Material checked .
Training & model systems
See how a model is trained, adapted and run. Explore feedback, released checkpoints and the memory and hardware choices behind a training recipe.
-
Understand what happens after pretraining. Lambert's book distinguishes training on examples, preferred responses and verifiably correct answers; the team Tülu 3 paper shows how those stages fit into a released recipe with separate development and held-out tests.
Start with this Web material · rlhfbook.com
A map of language-model post-training
Start with the recipe descriptions to separate examples, preference feedback and checked rewards. You can follow that overview before studying the equations.
Material checked . Authorship source.
-
Rush connects coding models to their training and evaluation. His Composer 2 introduction leads to the coauthored report on continued pretraining, reinforcement learning and tests drawn from engineering work. His Annotated Transformer provides a deeper route into the architecture through Python code.
Start with this Web material · cursor.com
A technical report on Composer 2
Rush introduces two training stages and explains why the team's evaluations include ambiguous requests and changes across several files. Continue to the linked report for the experiments and evaluation setup behind the claims.
Published . Material checked .
-
Dao and collaborators explain how memory movement and GPU scheduling affect model speed. Start with the FlashAttention-3 recap, then explore Mamba-3 to compare a fixed-size state with an expanding attention cache. The examples connect algorithm choices to hardware and retrieval tradeoffs.
Start with this Web material · tridao.me
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
The coauthored recap explains processing attention in small memory blocks. Later sections show overlapping GPU operations and precision tradeoffs. Check the hardware and operation measured before applying a speed comparison to your own model.
Published . Material checked .
Reasoning & evaluation
Ask what intermediate reasoning, a factual answer or a benchmark score establishes. Read the test conditions and compare the result with the task you care about.
-
Understand what a step-by-step response establishes. Wei's coauthored prompting study tests intermediate reasoning examples against simpler prompts, while his SimpleQA work separately grades correct facts, wrong answers and non-attempts.
Start with this Web material · arxiv.org
An example of prompting with intermediate steps
Figure 1 compares two prompts for the same math problem. Read it before the experiments, then section 6 on why fluent steps do not guarantee a correct answer. Wei is a coauthor.
Published . Material checked . Authorship source.
-
Study how a behavior develops during training. Biderman's coauthored Pythia suite exposes comparable checkpoints and data order, and her memorization work tests what cheaper runs can predict. Her 2026 position paper explains the research question before the technical experiments.
Start with this Web material · arxiv.org
Why study a model while it is being trained?
Sections 2.1 and 2.2 distinguish describing a finished model from predicting how its behavior develops. This coauthored position paper is a research agenda and needs no detailed experimental background to start.
Published . Material checked . Authorship source.
-
Inspect what a benchmark score measures. Chollet's ARC work asks how a system learns a new rule from examples; the coauthored 2024 competition report shows why the test split, adaptation method and compute budget matter when comparing results.
Start with this Web material · arxiv.org
A grid puzzle and two different leaderboard conditions
Start with section 1's grid puzzle, then section 2's two leaderboards. Their different conditions explain why scores cannot simply be compared. The coauthored report covers the 2024 competition.
Published . Material checked . Authorship source.
-
Weng explains why a higher evaluator score or a convincing chain of thought may fail to establish a better result. Start with the LLM examples in her reward hacking review, then read Why We Think for sampling, revision, reinforcement learning and checks of reasoning faithfulness.
Start with this Web material · lilianweng.github.io
Reward Hacking in Reinforcement Learning
Start with examples of coding models changing tests or reward code. Then follow the distinction between the intended goal and the score used to judge it. The review explains the mechanisms and surveys proposed mitigations.
Published . Material checked . Authorship source.
-
Read Brown to connect learning, search and multi-agent reasoning. His coauthored game research shows the assumptions behind those methods; his interview explains why demonstrating many agents is different from measuring their benefit against smaller comparison runs.
Start with this Web material · dwarkesh.com
Brown explains how to evaluate parallel agents
Read Brown's answers in the opening discussion of agent-count comparisons. The written transcript provides an accessible start before the search and reinforcement-learning papers.
Published . Material checked .
-
Wolf's Einstein AI essay asks what a hard exam score establishes about scientific discovery. Follow it with his coauthored DistilBERT paper for a concrete comparison of model size, task performance and inference time. Together they show how an evaluation's design determines what you can conclude.
Start with this Web material · thomwolf.io
Wolf argues that answering known questions and generating useful new research questions need different tests. Read this as a proposed research direction: the essay leaves the design of a scientific-discovery benchmark open.
Material checked . Authorship source.
No people match that search. Try another name, subject or reading task.
How this list is selected
We selected people through their own explanations, implementations, interviews and coauthored research. Each card links an original starting material and explains the question it answers. Tutorials, research reviews and experimental papers serve different reading needs; the notes identify useful prerequisites and preserve team authorship. Roles come from current primary biographies. The same person keeps one profile across AI, crypto and coding topics.