ML 101
M14 · L01
Module 14

Learning from Reward

No dataset. No labels. Just a situation, a set of things you may do, and a number that arrives afterwards telling you how well that went. This is the third kind of machine learning.

01 / 13
ML 101
M14 · L01
An Old Promise

Lesson 1.2 Said So

Lesson 1.2 named reinforcement learning as “fundamentally different from both supervised and unsupervised learning” — then set it aside for twelve modules. Module 14 pays that promise back.

Reference
Sutton & Barto, 2018
New
A setting, not a model
02 / 13
ML 101
M14 · L01
The Loop

Agent, Environment

The agent observes a state, chooses an action, and the environment answers with a reward and the next state. Then the loop turns again. One run from start to terminal state is an episode.

One step
S_t \;\rightarrow\; A_t \;\rightarrow\; R_{t+1},\; S_{t+1}
03 / 13
ML 101
M14 · L01
Delayed Reward

Credit, Across Time

The move that lost the game may have been forty moves before the loss. Which earlier action does a late reward belong to? Lesson 6.3 solved credit assignment across layers. This is the same problem across time.

The hard part
There is no differentiable graph to trace back — only one sequence of what happened
04 / 13
ML 101
M14 · L01
Difference One

Evaluative, Not Instructive

Supervised learning is told the correct action. Reinforcement learning is told only the value of the action it took — and nothing at all about the actions it did not take. A reward of 3 never says whether 10 was available.

Consequence
The agent has to experiment. Fitting is not enough, because the alternatives were never scored.
05 / 13
ML 101
M14 · L01
Difference Two — The Deeper One

The Data Depends on You

  • The policy decides which states get visited
  • So the policy decides which experience is collected
  • Improve the policy and the data distribution moves
  • Non-stationary by construction — caused by the learner
  • A region never visited stays wrong forever, uncontradicted
06 / 13
ML 101
M14 · L01
The Smallest Real Example

Two-Armed Bandit

Arm A paid 1, 0, 1 → Q(A) = 0.67. Arm B paid 0 once → Q(B) = 0. A greedy agent now picks A forever — and never learns that B in fact pays 0.8. Robbins posed this problem in 1952.

Action value
Q_t(a) = \frac{1}{N_t(a)} \sum_{i=1}^{N_t(a)} R_i
07 / 13
ML 101
M14 · L01
Bookkeeping

Old Value, One Step

You never need to store every reward. Move the estimate a little way toward the error. Nearly every update in this module has this shape — a fixed step size α in place of 1/n makes it track a changing world instead.

Incremental mean
Q_{n+1} = Q_n + \frac{1}{n}\left[ R_n - Q_n \right]
08 / 13
ML 101
M14 · L01
The Dilemma

Explore or Exploit

Exploiting what you know is the only way to collect reward now. Exploring is the only way to have better knowledge to exploit later. Every single step spends the budget on one or the other.

ε-greedy
Best estimate almost always; with probability ε pick at random. Every action keeps probability ≥ ε/k, so every estimate keeps being corrected.
09 / 13
ML 101
M14 · L01
The Default Failure

Reward Hacking

Reward a cleaning robot for dirt collected and it may tip the bin out to collect it again. Reward a recommender for clicks and it learns which headlines are hardest to resist. The agent optimises what you measured, not what you meant.

Sutton & Barto
Reward says what you want achieved, never how. Method belongs in the initial policy or initial values.
10 / 13
ML 101
M14 · L01
Where It Pays

Four Real Homes

  • Control — robotics, flight, cooling, scheduling
  • Games — exact rules, unambiguous reward, self-play
  • Recommendation — sequential choices, evaluative feedback
  • RLHF — human comparisons instead of a written reward (Christiano et al., 2017; Lesson 9.3)
  • All three conditions must hold, or use supervised learning
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

Agent, environment, state, action, reward, episode. Feedback is evaluative, not instructive — and the data distribution depends on the policy, which is the deeper difference. Reward is delayed, so credit assignment runs across time. The bandit isolates explore vs. exploit. And the agent optimises exactly what you measured.

Lesson complete
Lesson 14.2: Markov Decision Processes →
13 / 13