XFollowListPeople, ideas and original work on X

AI reading list

LLM research & explanations

Large language models learn patterns in text and generate sequences of tokens — pieces of text. To understand what they can do, follow the architecture, training and tests behind a result. Start with an explanation or small implementation, then compare the original experiments. Each person here offers a specific route into that work.

Selected sources
12

People to read

Open the first material, then continue with the author on X. People appear once, grouped by their main reading use and ordered by handle within each group.

6 of 12 people

Reasoning & evaluation

Ask what intermediate reasoning, a factual answer or a benchmark score establishes. Read the test conditions and compare the result with the task you care about.

  • Understand what a step-by-step response establishes. Wei's coauthored prompting study tests intermediate reasoning examples against simpler prompts, while his SimpleQA work separately grades correct facts, wrong answers and non-attempts.

    Start with this Web material · arxiv.org

    An example of prompting with intermediate steps

    Figure 1 compares two prompts for the same math problem. Read it before the experiments, then section 6 on why fluent steps do not guarantee a correct answer. Wei is a coauthor.

    Published . Material checked . Authorship source.

  • Study how a behavior develops during training. Biderman's coauthored Pythia suite exposes comparable checkpoints and data order, and her memorization work tests what cheaper runs can predict. Her 2026 position paper explains the research question before the technical experiments.

    Start with this Web material · arxiv.org

    Why study a model while it is being trained?

    Sections 2.1 and 2.2 distinguish describing a finished model from predicting how its behavior develops. This coauthored position paper is a research agenda and needs no detailed experimental background to start.

    Published . Material checked . Authorship source.

  • @fchollet François Chollet Researcher

    About François Chollet & selected work

    Inspect what a benchmark score measures. Chollet's ARC work asks how a system learns a new rule from examples; the coauthored 2024 competition report shows why the test split, adaptation method and compute budget matter when comparing results.

    Start with this Web material · arxiv.org

    A grid puzzle and two different leaderboard conditions

    Start with section 1's grid puzzle, then section 2's two leaderboards. Their different conditions explain why scores cannot simply be compared. The coauthored report covers the 2024 competition.

    Published . Material checked . Authorship source.

  • Weng explains why a higher evaluator score or a convincing chain of thought may fail to establish a better result. Start with the LLM examples in her reward hacking review, then read Why We Think for sampling, revision, reinforcement learning and checks of reasoning faithfulness.

    Start with this Web material · lilianweng.github.io

    Reward Hacking in Reinforcement Learning

    Start with examples of coding models changing tests or reward code. Then follow the distinction between the intended goal and the score used to judge it. The review explains the mechanisms and surveys proposed mitigations.

    Published . Material checked . Authorship source.

  • Read Brown to connect learning, search and multi-agent reasoning. His coauthored game research shows the assumptions behind those methods; his interview explains why demonstrating many agents is different from measuring their benefit against smaller comparison runs.

    Start with this Web material · dwarkesh.com

    Brown explains how to evaluate parallel agents

    Read Brown's answers in the opening discussion of agent-count comparisons. The written transcript provides an accessible start before the search and reinforcement-learning papers.

    Published . Material checked .

  • Wolf's Einstein AI essay asks what a hard exam score establishes about scientific discovery. Follow it with his coauthored DistilBERT paper for a concrete comparison of model size, task performance and inference time. Together they show how an evaluation's design determines what you can conclude.

    Start with this Web material · thomwolf.io

    The Einstein AI model

    Wolf argues that answering known questions and generating useful new research questions need different tests. Read this as a proposed research direction: the essay leaves the design of a scientific-discovery benchmark open.

    Material checked . Authorship source.

How this list is selected

We selected people through their own explanations, implementations, interviews and coauthored research. Each card links an original starting material and explains the question it answers. Tutorials, research reviews and experimental papers serve different reading needs; the notes identify useful prerequisites and preserve team authorship. Roles come from current primary biographies. The same person keeps one profile across AI, crypto and coding topics.