Learning from Reward
No dataset. No labels. Just a situation, a set of things you may do, and a number that arrives afterwards telling you how well that went. This is the third kind of machine learning.
Lesson 1.2 Said So
Lesson 1.2 named reinforcement learning as “fundamentally different from both supervised and unsupervised learning” — then set it aside for twelve modules. Module 14 pays that promise back.
Agent, Environment
The agent observes a state, chooses an action, and the environment answers with a reward and the next state. Then the loop turns again. One run from start to terminal state is an episode.
Credit, Across Time
The move that lost the game may have been forty moves before the loss. Which earlier action does a late reward belong to? Lesson 6.3 solved credit assignment across layers. This is the same problem across time.
Evaluative, Not Instructive
Supervised learning is told the correct action. Reinforcement learning is told only the value of the action it took — and nothing at all about the actions it did not take. A reward of 3 never says whether 10 was available.
The Data Depends on You
- The policy decides which states get visited
- So the policy decides which experience is collected
- Improve the policy and the data distribution moves
- Non-stationary by construction — caused by the learner
- A region never visited stays wrong forever, uncontradicted
Two-Armed Bandit
Arm A paid 1, 0, 1 → Q(A) = 0.67. Arm B paid 0 once → Q(B) = 0. A greedy agent now picks A forever — and never learns that B in fact pays 0.8. Robbins posed this problem in 1952.
Old Value, One Step
You never need to store every reward. Move the estimate a little way toward the error. Nearly every update in this module has this shape — a fixed step size α in place of 1/n makes it track a changing world instead.
Explore or Exploit
Exploiting what you know is the only way to collect reward now. Exploring is the only way to have better knowledge to exploit later. Every single step spends the budget on one or the other.
Reward Hacking
Reward a cleaning robot for dirt collected and it may tip the bin out to collect it again. Reward a recommender for clicks and it learns which headlines are hardest to resist. The agent optimises what you measured, not what you meant.
Four Real Homes
- Control — robotics, flight, cooling, scheduling
- Games — exact rules, unambiguous reward, self-play
- Recommendation — sequential choices, evaluative feedback
- RLHF — human comparisons instead of a written reward (Christiano et al., 2017; Lesson 9.3)
- All three conditions must hold, or use supervised learning
What You Learned
Agent, environment, state, action, reward, episode. Feedback is evaluative, not instructive — and the data distribution depends on the policy, which is the deeper difference. Reward is delayed, so credit assignment runs across time. The bandit isolates explore vs. exploit. And the agent optimises exactly what you measured.