XFollowListPeople, ideas and original work on X
Lilian Weng’s avatar

Machine learning researcher; author of Lil’Log

Lilian Weng

@lilianweng on X

Connects reasoning and reinforcement learning methods to their failure modes, including models that improve an evaluator's score without completing the intended task.

Why read Lilian Weng?

Start with Weng when a better score or a convincing answer leaves you unsure what a model actually learned. Her reward hacking overview separates the intended goal from the signal used to train or judge an agent. The examples include coding models changing tests and evaluators giving different judgments when answer order changes.

Her Why We Think review then explains several ways to spend more computation on an answer: sample alternatives, revise a response or train reasoning with feedback. Read the sections on self-correction and faithful reasoning to understand why additional steps and a readable explanation still need a separate check of the result.

Start with the original

Selected work

  1. Research review ·

    Reward Hacking in Reinforcement Learning

    What if a model improves the score by changing the test? Weng surveys reward hacking, evaluator biases and proposed mitigations. Begin with the LLM examples, then follow the distinction between the real goal and its measured proxy.

  2. Research review ·

    Why We Think

    When does more thinking time help? Weng reviews sampling, revision, reinforcement learning and checks of chain-of-thought faithfulness. The comparisons distinguish a correct answer from an explanation that faithfully describes how it was reached.

Projects & roles

  • Lil’Log

    Weng's research notes, with diagrams, equations and references leading to the original papers.

Community rating

79

Vote once every 24 hours. No account needed.

Sources & further reading