Connects reasoning and reinforcement learning methods to their failure modes, including models that improve an evaluator's score without completing the intended task.
Why read Lilian Weng?
Start with Weng when a better score or a convincing answer leaves you unsure what a model actually learned. Her reward hacking overview separates the intended goal from the signal used to train or judge an agent. The examples include coding models changing tests and evaluators giving different judgments when answer order changes.
Her Why We Think review then explains several ways to spend more computation on an answer: sample alternatives, revise a response or train reasoning with feedback. Read the sections on self-correction and faithful reasoning to understand why additional steps and a readable explanation still need a separate check of the result.
Start with the original
Selected work
Research review ·
Reward Hacking in Reinforcement Learning
What if a model improves the score by changing the test? Weng surveys reward hacking, evaluator biases and proposed mitigations. Begin with the LLM examples, then follow the distinction between the real goal and its measured proxy.
Research review ·
Why We Think
When does more thinking time help? Weng reviews sampling, revision, reinforcement learning and checks of chain-of-thought faithfulness. The comparisons distinguish a correct answer from an explanation that faithfully describes how it was reached.
Projects & roles
Lil’Log
Weng's research notes, with diagrams, equations and references leading to the original papers.
Community rating
Sources & further reading
- Her own FAQ identifies @lilianweng as the account for updates to this blog.
- The dated byline and citation identify Weng as the review's author; she acknowledges John Schulman's feedback and direct edits.
- The original GPT-4 contribution record documents her historical applied research and evaluation work.