Skipping the Value Table
Lesson 14.3 learned action values and then read a policy off them: in each state, take whichever action has the highest Q. The policy was never trained — it was a byproduct. This lesson trains the policy itself.
A parameterised policy πθ(a | s) is a function with weights θ that takes a state and outputs a probability for each action, or the parameters of a distribution over actions. You then adjust θ to make good actions more probable. No Q-table, no argmax, no epsilon. Sutton and Barto’s Reinforcement Learning: An Introduction (2nd edition, 2018) devotes its policy-gradient chapter to exactly this shift, and it is the standard reference for the material below.
Three situations make this clearly better than learning values.
Continuous actions. A Q-learning agent must compute maxa Q(s, a). If the action is a steering angle or a joint torque — a real number, or a vector of forty of them — that maximisation is itself an optimisation problem, solved inside every single update. A policy network sidesteps it entirely: it outputs the action, or the mean and spread of a distribution over actions.
Stochastic optimal policies. Value-greedy policies are deterministic by construction. Sometimes the best policy genuinely is not. In a game like rock-paper-scissors, any deterministic policy is exploitable and the optimal policy is uniformly random. In a partially observed environment where two different states look identical, randomising is the best you can do. A policy that outputs probabilities can represent this; an argmax cannot.
Smoothness. A tiny change to a Q-value can flip an argmax and change the policy discontinuously. A policy gradient nudges action probabilities a little at a time, which is easier to make stable.
The Policy-Gradient Idea
The objective J(θ) is the expected return under the policy — how much reward the agent collects on average if it behaves according to πθ. We want to increase it, so we want its gradient with respect to θ. The obstacle is that θ influences the return only by changing which trajectories occur, and the environment’s contribution to that is unknown. The policy gradient theorem is what makes the gradient computable anyway.
The shape of that expression is worth pausing on, because it explains every algorithm in the rest of this lesson. ∇θ log πθ(a | s) is the direction in weight space that makes action a more likely in state s. Multiplying it by Qπ(s, a) scales that direction by how good the action was. Sum over experience and you get: make good actions more likely, in proportion to how good they were. Because it is an expectation, it can be estimated by sampling — just run the policy and average.
REINFORCE, and Its Variance Problem
Williams (1992) gave the simplest instance. Play a complete episode, and for each step use the actual return that followed it — call it Gt — in place of Qπ(s, a). The return really did occur, so it is an unbiased sample of the action value.
REINFORCE is unbiased, and in practice it is often unusably noisy. The problem is variance. Gt is the sum of every reward in the rest of the episode, so it absorbs every random event that happened after the action — environment noise, and every subsequent choice the stochastic policy made. Two identical actions in identical states can be followed by wildly different returns, so one gets strongly reinforced and the other strongly discouraged. The gradient estimate points in roughly the right direction on average, but any individual estimate can point almost anywhere, and averaging it down takes an enormous number of episodes.
There is a second, subtler defect. Suppose every reward in a task is positive — say every action earns between +1 and +10. Then Gt is always positive, so every action ever taken gets reinforced, including bad ones. Learning happens only because good actions are reinforced harder. That is a wasteful way to use a gradient signal, and it is what the next idea fixes.
Baselines and Advantage
Subtract a baseline b(s) from the return before using it: replace Gt with Gt − b(St). As long as the baseline depends only on the state and not on the action, this leaves the gradient unbiased — Sutton and Barto prove it, and the reason is that the extra term sums to zero when summed over actions. But it can cut the variance dramatically, because the quantity multiplying the log-probability is now a comparison rather than an absolute score.
The natural baseline is the state value Vπ(s) — how well the agent expected to do from that state under its current policy. Subtracting it produces the advantage.
This also disposes of the all-positive-reward problem. In a task where every action earns +1 to +10, a +2 action in a state worth +6 has advantage −4 and is made less likely. Absolute scores became relative ones.
Actor-Critic
The advantage needs Vπ(s), and nobody knows it. So learn it too. Train two networks side by side:
The actor is the policy πθ(a | s). It chooses actions, and it is updated in the direction of the policy gradient, scaled by the advantage the critic reports.
The critic estimates V(s). It is updated by ordinary temporal-difference regression — the same one-step bootstrapping idea as Lesson 14.3, minimising the squared TD error rather than a policy objective.
The critic buys more than variance reduction. Because it can estimate the value of the next state, the advantage can be approximated from a single transition: Rt+1 + γV(St+1) − V(St) — which is just the TD error again. So actor-critic methods, unlike REINFORCE, do not need a finished episode. That makes them applicable to continuing tasks with no episode boundary at all.
The cost is bias. The critic’s estimate of V is wrong, especially early in training, so the advantage it reports is wrong, so the gradient is biased. REINFORCE was unbiased and impossibly noisy; actor-critic is biased and workable. Every practical policy-gradient method sits somewhere on this trade-off, and the knob that positions it is how many real reward steps you use before handing over to the critic’s estimate.
PPO: Why It Became a Default
Policy-gradient methods have a specific fragility. The gradient tells you a direction but not how far to go, and one oversized update can wreck the policy. That is worse than it sounds: the policy also generates the data. A ruined policy collects garbage experience, from which it cannot recover, so a single bad step can end the run.
Schulman et al. (2017) proposed Proximal Policy Optimization, which addresses this by refusing to let the policy move far from the one that collected the data. It defines a probability ratio — how much more likely the new policy makes an action than the old one did — and clips it.
Why did it become a common default? Not because it is the most sophisticated option, but for the practical properties Schulman et al. set out to achieve — the ones that matter when you have to make something work:
Stability without heavy machinery. Earlier trust-region methods enforced a similar constraint but needed second-order information about the objective’s curvature. PPO gets a comparable effect from a clip and a minimum, using only first-order gradients — so it drops into any standard autodiff framework.
Few hyperparameters, forgiving defaults. A clipping range, a learning rate, an epoch count. Methods that need a dozen carefully tuned settings do not survive contact with new problems.
Data reuse. The clip is what makes it safe to take several gradient epochs over the same batch of collected experience, which matters when interacting with the environment is the expensive part.
Its status as a default is clearest in the RLHF pipeline below, where the InstructGPT work (Ouyang et al., 2022) uses PPO to fine-tune a language model from human feedback — the setting in which most practitioners now encounter it.
RLHF: Where Most People Actually Meet RL
Lesson 9.3 introduced reinforcement learning from human feedback as the step that turned a next-token predictor into an assistant. With policy gradients in hand, the mechanism can be stated properly. The obstacle RLHF solves is that there is no reward function for “a helpful answer.” You cannot write one, and RL needs one. So the reward function is learned from human comparisons.
Christiano et al. (2017) established the approach, training agents from human preferences between pairs of behaviour segments rather than from a hand-written reward. Ouyang et al. (2022), the InstructGPT paper, applied the same structure to language models. Three stages:
1. Preference data. Sample several responses to the same prompt and ask human annotators which they prefer. Note what is being collected: a ranking, not a score. People are far more consistent at saying which of two answers is better than at putting a number on either.
2. Reward model. Train a model to predict those preferences — it takes a prompt and a response and outputs a scalar, fitted so that preferred responses score higher. This is ordinary supervised learning, and it converts a pile of human judgements into the reward function RL requires.
3. PPO fine-tune. Now it is a policy-gradient problem: the policy is the language model, an action is emitting a token, and the reward comes from the reward model. PPO optimises it, with a penalty for drifting too far from the original model — which keeps the language model from collapsing into whatever degenerate text the reward model happens to overrate.
That third-stage penalty deserves attention, because it is the reward-hacking problem of Lesson 14.1 arriving in production. The reward model is an imperfect stand-in for human judgement, and a policy optimised hard enough against it will find its errors rather than satisfy the humans it was fitted to. The constraint is what keeps the optimisation honest. GenAI-101 covers the pipeline from the language-model side; the reinforcement learning inside it is this lesson’s material.
Honest Limits
Reinforcement learning has produced genuinely remarkable results, and it is also the least reliable tool in this course. Four limits are worth stating plainly, because the published successes tend not to advertise them.
Sample inefficiency. RL agents need staggering amounts of interaction — commonly millions of environment steps on tasks a person picks up in minutes. Where interaction is cheap and fast, as in a simulator or a game, this is an inconvenience. Where every sample costs a real robot movement, a real user session, or a real dollar, it is often disqualifying. This is the single biggest reason RL is rarer in production than its reputation suggests.
Reward specification. Supervised learning asks you for labels. RL asks you for a reward function, and that is a harder thing to be right about, because the agent optimises exactly what you wrote rather than what you meant. Lesson 14.1’s reward hacking is not an exotic failure mode; it is the default outcome of a slightly wrong reward function.
The sim-to-real gap. The standard response to sample inefficiency is to train in simulation, where samples are cheap and mistakes are free. But a policy trained in simulation has learned the simulator, including its inaccuracies. Contact friction, sensor noise, latency and unmodelled dynamics all differ in the real world, and a policy can exploit a simulator artefact that does not exist outside it. Domain randomisation — deliberately varying the simulator’s parameters so no single quirk can be relied on — is the usual mitigation, and it narrows the gap rather than closing it.
Reproducibility. RL results vary far more between runs than supervised results do, because the agent generates its own data: an early lucky exploration can lead to a good policy and an early unlucky one to failure, from identical code. The same algorithm on the same task with different random seeds can produce very different learning curves. Reporting a single run is therefore not evidence, and comparisons between algorithms require multiple seeds and reported spread.
Before reaching for RL, ask whether the problem really requires learning from consequences. If you have labelled examples of correct behaviour, supervised learning is cheaper, more stable and easier to debug. RL earns its cost when the right action is not known in advance but the outcome can be scored — games, control, and preference optimisation like RLHF. Where RL is genuinely used at scale, it is usually the last stage of a pipeline that is otherwise supervised, which is exactly the shape of RLHF.
- Policy methods train πθ(a | s) directly instead of deriving a policy from action values. They win on continuous actions, where maxa Q would be its own optimisation problem, and when the optimal policy is genuinely stochastic.
- The policy gradient is an expectation of ∇θ log πθ(a | s) weighted by how good the action was: make good actions more likely, in proportion to how good they were. No term differentiates the environment.
- REINFORCE (Williams, 1992) uses the actual episode return and is unbiased but very high variance, because the return absorbs every random event that followed the action. It also needs a finished episode.
- Subtracting a state-dependent baseline leaves the gradient unbiased and cuts variance. The natural choice gives the advantage A = Q − V: how much better the action was than the state was worth.
- Actor-critic learns the policy and the value function together. The critic makes single-step updates possible — the advantage becomes the TD error — at the cost of bias from an imperfect critic.
- PPO (Schulman et al., 2017) clips the probability ratio so the update cannot stray far from the policy that collected the data. It became a common default for engineering reasons: first-order only, few hyperparameters, and safe reuse of a batch for several epochs.
- RLHF is preference data, then a reward model fitted to those rankings, then a PPO fine-tune with a penalty for drifting from the original model — Christiano et al. (2017) and Ouyang et al. (2022).
- The honest limits: sample inefficiency, reward specification, the sim-to-real gap and poor reproducibility across random seeds. If you have labelled examples of correct behaviour, use supervised learning instead.