ML 101
M14 · L03
Module 14

Q-Learning

Lesson 14.2 solved an MDP using transition probabilities. Nobody hands you those. So: can an agent learn to act well by simply acting, and never estimating a probability at all?

01 / 13
ML 101
M14 · L03
No Map Required

Model-Free

A robot does not know the probability a wheel slips. Value iteration needs that number; Q-learning does not. It uses what did happen instead of what might. Sutton & Barto (2018) draw exactly this line.

Needs a model
Value iteration
Needs experience
Q-learning
02 / 13
ML 101
M14 · L03
The Key Idea

Learn from One Step

Monte Carlo waits for the episode to end. Temporal-difference learning refuses to wait: after one step it compares its old opinion with the reward it just got plus the discounted value of where it landed.

TD error
\delta_t = R_{t+1} + \gamma \max_a Q(S_{t+1}, a) - Q(S_t, A_t)
03 / 13
ML 101
M14 · L03
Watkins, 1989

The Update Rule

Move the old estimate a fraction α toward the better one. Watkins’ 1989 thesis introduced it; Watkins & Dayan (1992) proved the tabular version converges to the optimal action values.

Q-learning
Q(S_t, A_t) \leftarrow Q(S_t, A_t) + \alpha \left[ R_{t+1} + \gamma \max_a Q(S_{t+1}, a) - Q(S_t, A_t) \right]
04 / 13
ML 101
M14 · L03
Worked by Hand

A Four-Cell Corridor

  • States A – B – C – G; G is the terminal goal
  • Actions left, right. Right from C enters G
  • Rewards 0 per step, +10 on entering G
  • α = 0.5, γ = 0.9, every Q entry starts at 0
  • G is terminal, so maxa Q(G, a) = 0
05 / 13
ML 101
M14 · L03
Update 1

The First Non-Zero

In C, move right, enter G. Reward +10, and G is terminal so the bootstrap term is 0.

Arithmetic
δ = 10 + 0.9×0 − 0 = 10
Q(C, right) ← 0 + 0.5×10 = 5.0

Note 5.0, not 10 — half the evidence, because α = 0.5.

06 / 13
ML 101
M14 · L03
Update 2

Learning from No Reward

In B, move right, arrive in C. Reward 0. But maxa Q(C, a) = max(0, 5.0) = 5.0, so there is something to learn.

Arithmetic
δ = 0 + 0.9×5.0 − 0 = 4.5
Q(B, right) ← 0 + 0.5×4.5 = 2.25

Zero reward, real learning. It all came from the agent’s own estimate of C.

07 / 13
ML 101
M14 · L03
Update 3

Value Walks Backwards

In C again, right into G. Same transition, different table.

Arithmetic
δ = 10 + 0 − 5.0 = 5.0
Q(C, right) ← 5.0 + 0.5×5.0 = 7.5

Each visit closes half the gap: 5.0, 7.5, 8.75, 9.375… toward Q* = 10, with Q*(B, right) = 9 and Q*(A, right) = 8.1 one discount step behind.

08 / 13
ML 101
M14 · L03
Exploration

Epsilon-Greedy

With probability 1 − ε take the best known action; with probability ε act at random. Crude, and it is what keeps every state-action pair reachable — the condition Watkins & Dayan’s proof requires.

Decaying it
ε = 1.0, ×0.995 per episode → ≈0.08 after 500 episodes, with a floor of 0.05. Too fast locks in a mediocre route; too slow never cashes in.
09 / 13
ML 101
M14 · L03
On-Policy vs Off-Policy

SARSA Beside Q-Learning

Rummery & Niranjan (1994). One term differs: the max over a, versus the action actually taken. Rerun update 2 with an exploratory left from C — Q-learning learns 2.25, SARSA learns 0.

SARSA
Q(S_t, A_t) \leftarrow Q(S_t, A_t) + \alpha \left[ R_{t+1} + \gamma Q(S_{t+1}, A_{t+1}) - Q(S_t, A_t) \right]
10 / 13
ML 101
M14 · L03
When the Table Runs Out

Deep Q-Networks

  • The corridor needs 8 table entries; an Atari screen needs astronomically more
  • Worse: a table has no notion of similarity between two nearly identical screens
  • DQN replaces the table with a network; the update rule is unchanged
  • Experience replay — random minibatches from a buffer break the correlation between consecutive frames, and reuse each transition
  • Target network — a frozen copy θ− computes the target, so it stops moving with every gradient step
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

What You Learned

Q-learning needs no model. It bootstraps after one step, moving the estimate a fraction α toward reward plus discounted best-next value. In the corridor: 5.0, then 2.25 from a zero-reward step, then 7.5. Epsilon-greedy explores and is decayed. SARSA is the on-policy twin. DQN replaces the table and needs replay plus a target network to survive — Mnih et al., Nature (2015).

Lesson Complete
Policy Methods →
13 / 13