Policy Methods
Q-learning trained values and read a policy off them. The policy was a byproduct. This lesson trains the policy itself — and ends with where reinforcement learning is genuinely used, and where it is not.
Learn the Policy Directly
- Continuous actions — maxa Q over a steering angle is its own optimisation, inside every update
- Stochastic optima — the best policy for rock-paper-scissors is random; an argmax cannot say that
- Smoothness — a tiny value change flips an argmax; probabilities move gently
The Policy Gradient
Make good actions more likely, in proportion to how good they were. And notice what is absent: no term differentiates the environment, so this stays model-free. This is the undiscounted form — with discounting each step-t term carries a γt weight, routinely dropped in practice. Sutton & Barto (2018), policy-gradient chapter.
REINFORCE
The simplest instance. Play an episode; use the return Gt that actually followed each step in place of Q. Gradient ascent — the sign is a plus, because reward is maximised rather than loss minimised.
Unbiased, and Impossibly Noisy
Gt sums every reward after the action, so it absorbs every random event that followed. Two identical actions can draw wildly different returns — one strongly reinforced, one strongly discouraged.
Baselines and Advantage
Subtract anything that depends on the state but not the action: the gradient stays unbiased, the variance drops. Use V(s) and you get the advantage — absolute scores become relative ones.
Actor-Critic
The actor is the policy and follows the policy gradient. The critic estimates V(s) by ordinary TD regression — Lesson 14.3’s bootstrapping, reused. With a critic the advantage becomes the TD error, so no finished episode is needed.
PPO’s Clipped Objective
One oversized step wrecks the policy — and the policy generates the data, so it cannot recover. PPO clips the probability ratio rt(θ) into [1−ε, 1+ε]. Because it takes the minimum, the reward for pushing further goes flat.
Boring, and It Works
- First-order only — earlier trust-region methods needed curvature information; a clip and a minimum get a comparable effect
- Few hyperparameters — a clip range, a learning rate, an epoch count
- Data reuse — the clip is what makes several epochs over one batch safe
- The properties Schulman et al. argued for — and where most people meet it is RLHF (next slide)
- Note: this ε is a clip range, unrelated to exploration ε
RLHF
There is no reward function for “a helpful answer”, so it is learned. Christiano et al. (2017) established it; Ouyang et al. (2022) applied it to language models.
- Preference data — humans rank pairs, not score them
- Reward model — supervised fit to those rankings
- PPO fine-tune — plus a penalty for drifting from the original model
Honest Limits
- Sample inefficiency — millions of episodes for tasks a person learns in minutes
- Reward specification — the agent optimises what you wrote, not what you meant
- Sim-to-real — a policy trained in a simulator has learned the simulator, artefacts included
- Reproducibility — different seeds, very different curves. One run is not evidence
Where This Goes
Policy methods train πθ directly, which values cannot do for continuous or genuinely random actions. REINFORCE is unbiased and noisy; a baseline gives the advantage; a critic makes single-step updates possible; PPO clips the step so the run survives. RLHF is that whole stack pointed at human preference. And if you already have labelled examples of correct behaviour, use supervised learning instead.