Q-Learning
Lesson 14.2 solved an MDP using transition probabilities. Nobody hands you those. So: can an agent learn to act well by simply acting, and never estimating a probability at all?
Model-Free
A robot does not know the probability a wheel slips. Value iteration needs that number; Q-learning does not. It uses what did happen instead of what might. Sutton & Barto (2018) draw exactly this line.
Learn from One Step
Monte Carlo waits for the episode to end. Temporal-difference learning refuses to wait: after one step it compares its old opinion with the reward it just got plus the discounted value of where it landed.
The Update Rule
Move the old estimate a fraction α toward the better one. Watkins’ 1989 thesis introduced it; Watkins & Dayan (1992) proved the tabular version converges to the optimal action values.
A Four-Cell Corridor
- States A – B – C – G; G is the terminal goal
- Actions left, right. Right from C enters G
- Rewards 0 per step, +10 on entering G
- α = 0.5, γ = 0.9, every Q entry starts at 0
- G is terminal, so maxa Q(G, a) = 0
The First Non-Zero
In C, move right, enter G. Reward +10, and G is terminal so the bootstrap term is 0.
Q(C, right) ← 0 + 0.5×10 = 5.0
Note 5.0, not 10 — half the evidence, because α = 0.5.
Learning from No Reward
In B, move right, arrive in C. Reward 0. But maxa Q(C, a) = max(0, 5.0) = 5.0, so there is something to learn.
Q(B, right) ← 0 + 0.5×4.5 = 2.25
Zero reward, real learning. It all came from the agent’s own estimate of C.
Value Walks Backwards
In C again, right into G. Same transition, different table.
Q(C, right) ← 5.0 + 0.5×5.0 = 7.5
Each visit closes half the gap: 5.0, 7.5, 8.75, 9.375… toward Q* = 10, with Q*(B, right) = 9 and Q*(A, right) = 8.1 one discount step behind.
Epsilon-Greedy
With probability 1 − ε take the best known action; with probability ε act at random. Crude, and it is what keeps every state-action pair reachable — the condition Watkins & Dayan’s proof requires.
SARSA Beside Q-Learning
Rummery & Niranjan (1994). One term differs: the max over a, versus the action actually taken. Rerun update 2 with an exploratory left from C — Q-learning learns 2.25, SARSA learns 0.
Deep Q-Networks
- The corridor needs 8 table entries; an Atari screen needs astronomically more
- Worse: a table has no notion of similarity between two nearly identical screens
- DQN replaces the table with a network; the update rule is unchanged
- Experience replay — random minibatches from a buffer break the correlation between consecutive frames, and reuse each transition
- Target network — a frozen copy θ− computes the target, so it stops moving with every gradient step
What You Learned
Q-learning needs no model. It bootstraps after one step, moving the estimate a fraction α toward reward plus discounted best-next value. In the corridor: 5.0, then 2.25 from a zero-reward step, then 7.5. Epsilon-greedy explores and is decayed. SARSA is the on-policy twin. DQN replaces the table and needs replay plus a target network to survive — Mnih et al., Nature (2015).