ML 101
M14 · L04
Module 14

Policy Methods

Q-learning trained values and read a policy off them. The policy was a byproduct. This lesson trains the policy itself — and ends with where reinforcement learning is genuinely used, and where it is not.

01 / 13
ML 101
M14 · L04
When Values Lose

Learn the Policy Directly

  • Continuous actions — maxa Q over a steering angle is its own optimisation, inside every update
  • Stochastic optima — the best policy for rock-paper-scissors is random; an argmax cannot say that
  • Smoothness — a tiny value change flips an argmax; probabilities move gently
02 / 13
ML 101
M14 · L04
The Theorem

The Policy Gradient

Make good actions more likely, in proportion to how good they were. And notice what is absent: no term differentiates the environment, so this stays model-free. This is the undiscounted form — with discounting each step-t term carries a γt weight, routinely dropped in practice. Sutton & Barto (2018), policy-gradient chapter.

Policy gradient
\nabla_{\theta} J(\theta) = \mathbb{E}_{\pi_{\theta}}\left[ \nabla_{\theta} \log \pi_{\theta}(a \mid s) \, Q^{\pi}(s,a) \right]
03 / 13
ML 101
M14 · L04
Williams, 1992

REINFORCE

The simplest instance. Play an episode; use the return Gt that actually followed each step in place of Q. Gradient ascent — the sign is a plus, because reward is maximised rather than loss minimised.

REINFORCE
\theta \leftarrow \theta + \alpha \, G_t \, \nabla_{\theta} \log \pi_{\theta}(A_t \mid S_t)
04 / 13
ML 101
M14 · L04
The Catch

Unbiased, and Impossibly Noisy

Gt sums every reward after the action, so it absorbs every random event that followed. Two identical actions can draw wildly different returns — one strongly reinforced, one strongly discouraged.

And a second defect
If every reward is positive, every action ever taken gets reinforced. Learning only happens because good ones are reinforced harder.
05 / 13
ML 101
M14 · L04
The Fix

Baselines and Advantage

Subtract anything that depends on the state but not the action: the gradient stays unbiased, the variance drops. Use V(s) and you get the advantage — absolute scores become relative ones.

Advantage
A^{\pi}(s,a) = Q^{\pi}(s,a) - V^{\pi}(s)
06 / 13
ML 101
M14 · L04
Two Networks

Actor-Critic

The actor is the policy and follows the policy gradient. The critic estimates V(s) by ordinary TD regression — Lesson 14.3’s bootstrapping, reused. With a critic the advantage becomes the TD error, so no finished episode is needed.

REINFORCE
Unbiased, noisy
Actor-critic
Biased, workable
07 / 13
ML 101
M14 · L04
Schulman et al., 2017

PPO’s Clipped Objective

One oversized step wrecks the policy — and the policy generates the data, so it cannot recover. PPO clips the probability ratio rt(θ) into [1−ε, 1+ε]. Because it takes the minimum, the reward for pushing further goes flat.

Clipped objective
L^{\text{CLIP}}(\theta) = \mathbb{E}_t\!\left[ \min\!\left( r_t \hat{A}_t, \; \text{clip}(r_t, 1-\epsilon, 1+\epsilon)\hat{A}_t \right) \right]
08 / 13
ML 101
M14 · L04
Why It Became a Default

Boring, and It Works

  • First-order only — earlier trust-region methods needed curvature information; a clip and a minimum get a comparable effect
  • Few hyperparameters — a clip range, a learning rate, an epoch count
  • Data reuse — the clip is what makes several epochs over one batch safe
  • The properties Schulman et al. argued for — and where most people meet it is RLHF (next slide)
  • Note: this ε is a clip range, unrelated to exploration ε
09 / 13
ML 101
M14 · L04
Lesson 9.3, properly

RLHF

There is no reward function for “a helpful answer”, so it is learned. Christiano et al. (2017) established it; Ouyang et al. (2022) applied it to language models.

  • Preference data — humans rank pairs, not score them
  • Reward model — supervised fit to those rankings
  • PPO fine-tune — plus a penalty for drifting from the original model
10 / 13
ML 101
M14 · L04
Say It Plainly

Honest Limits

  • Sample inefficiency — millions of episodes for tasks a person learns in minutes
  • Reward specification — the agent optimises what you wrote, not what you meant
  • Sim-to-real — a policy trained in a simulator has learned the simulator, artefacts included
  • Reproducibility — different seeds, very different curves. One run is not evidence
11 / 13
ML 101
Knowledge Check

Check what stuck

Four questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

12 / 13
ML 101
Summary
Recap

Where This Goes

Policy methods train πθ directly, which values cannot do for continuous or genuinely random actions. REINFORCE is unbiased and noisy; a baseline gives the advantage; a critic makes single-step updates possible; PPO clips the step so the run survives. RLHF is that whole stack pointed at human preference. And if you already have labelled examples of correct behaviour, use supervised learning instead.

Module 14 Complete
You have reached the end of ML 101
13 / 13