Reading
Stories Mode

Learning from Reward

~20 min read Lesson 1 of 4 in Module 14

The Promise Lesson 1.2 Made

Lesson 1.2 named three kinds of machine learning and said of the third that “reinforcement learning (RL) is fundamentally different from both supervised and unsupervised learning.” It then set that third kind aside. Every module since has been about a fixed dataset: rows that were collected before you arrived, and a model whose job is to map each row to an answer. This module returns to the promise.

The shift is not a new model family. It is a new problem setting. There is no dataset. There is a situation, a set of things you may do, and a number that arrives afterwards telling you how well that went. You must work out, from that number alone, what to do — and the only way to find out what an untried action pays is to try it.

The standard reference for everything in this module is Sutton and Barto, Reinforcement Learning: An Introduction (2nd edition, 2018). The notation used here is theirs, so that the book reads as a continuation of these four lessons rather than a translation exercise.

The Loop: Six Words That Define the Setting

Reinforcement learning has one diagram, and it is a loop with two boxes. The agent is the thing that decides. The environment is everything else — the world the agent acts in, including the parts of it the agent cannot see. The loop turns in discrete time steps.

At each step the agent observes a state St, a description of the situation right now. It chooses an action At from the actions available in that state. The environment responds with a reward Rt+1, a single scalar number, and a next state St+1. Then the loop turns again. An episode is one run of the loop from a start state to a terminal state — one game, one delivery, one dialogue. Some problems have no terminal state at all, and are called continuing.

The Trajectory
S_0,\, A_0,\, R_1,\, S_1,\, A_1,\, R_2,\, S_2,\, A_2,\, R_3,\, \ldots
Interaction produces a single interleaved sequence, not a table. Note the index convention, which is Sutton and Barto’s and is worth adopting now: the reward for acting at time t carries the subscript t+1, because it arrives with the next state. State, then action, then reward and the next state.

The rule the agent follows is its policy. A deterministic policy names one action per state; a stochastic one gives a probability to each. Learning, in this setting, means changing the policy so that the rewards it collects add up to more than they did before.

Where the Boundary Goes

The line between agent and environment is drawn at the limit of the agent’s control, not at the limit of its knowledge. A robot’s own motors and joints are part of the environment: the agent commands them and then finds out what happened. Anything the agent cannot change at will belongs on the environment side, however intimately it belongs to the robot.

Neither Supervised Nor Unsupervised — Twice Over

The usual explanation is that reinforcement learning has no labels. That is true and it is the shallower half. Supervised learning is told the correct action; reinforcement learning is told only the value of the action it took, and nothing about the actions it did not take. A reward of 3 does not say whether 10 was available. This is the difference between instructive feedback and evaluative feedback, and it is why the agent must experiment rather than simply fit.

The deeper half is usually left out: the data distribution depends on the policy. In supervised learning the training set is fixed, and every assumption in the course so far rests on that — the train/test split, cross-validation, even the meaning of overfitting. In reinforcement learning the agent’s current policy decides which states it visits, and therefore which experience it collects. Improve the policy and the distribution of the data moves. The learning problem is non-stationary by construction, and it is non-stationary because of something the learner itself did.

Two consequences follow immediately. First, a policy that never visits a region of the state space collects no information about it, so it can be arbitrarily wrong there forever without ever being contradicted. Second, the improvement loop can feed on itself: a slightly better policy visits slightly different states, which changes what is learned next. Most of the practical difficulty of reinforcement learning — and most of its reputation for instability — comes from this coupling rather than from the absence of labels.

Nor is this unsupervised learning. Clustering and dimensionality reduction in Module 5 look for structure in data that is already there. Here there is a goal, stated as a number, and the agent is judged against it. Sutton and Barto make the same point: reinforcement learning is a third paradigm, not a special case of either of the other two.

Delayed Reward and the Credit Assignment Problem

If every action paid immediately, the problem would be a lookup table. It does not. The move that lost the chess game may have been forty moves before the loss. The recommendation that made a user cancel their subscription was watched three weeks earlier. Reward is delayed, and the consequence is a version of the credit assignment problem Lesson 6.3 solved for network layers — here in time rather than in depth.

The two are worth holding side by side. Backpropagation assigns credit through a known, differentiable computation graph, and every path in that graph is available for inspection. Temporal credit assignment has no such graph: the agent has one sequence of states, actions and rewards, and must decide which of the earlier actions the later reward belongs to. It cannot re-run history with one action changed.

The whole of the rest of this module is machinery for this one problem. Lesson 14.2 defines the quantity that solves it in principle — the value of a state, meaning the reward you can expect from here on. Lesson 14.3 shows how to estimate that quantity from experience one step at a time. Lesson 14.4 goes after the policy directly.

Exploration vs. Exploitation: The Smallest Real Example

Strip the setting down until only one difficulty is left, and you get the multi-armed bandit: one state, k actions, and a reward drawn from an unknown distribution attached to whichever action you chose. There is nothing to plan, because nothing you do changes the situation you are in. The only question left is which arm to pull next, and that question is already hard. Robbins (1952) posed it in this form in the statistics literature, as the sequential design of experiments; Sutton and Barto open their book with it for the same reason it appears here.

Work a two-armed case by hand. You have pulled arm A three times and received rewards 1, 0 and 1; you have pulled arm B once and received 0. The obvious estimate of each arm’s worth is the average reward it has paid so far.

Sample-Average Action Value
Q_t(a) = \frac{1}{N_t(a)} \sum_{i=1}^{N_t(a)} R_i
The estimated value of action a at time t is the mean of the rewards received on the Nt(a) occasions it was chosen. For the two arms above: Q(A) = (1 + 0 + 1) / 3 = 0.67 and Q(B) = 0 / 1 = 0. The estimate is only as good as the count beneath it, and the count for B is 1.

A purely greedy agent now picks the arm with the highest estimate, which is A, receives another reward, updates Q(A), and picks A again. It will pick A for the rest of time. If arm B in fact pays 0.8 on average and that single 0 was bad luck, the agent will never discover it, because discovering it requires pulling an arm it has already decided is worse. That is the exploration–exploitation dilemma in its entirety: exploiting current knowledge is the only way to collect reward now, and exploring is the only way to have better knowledge to exploit later. Every step spends the budget on one or the other.

The simplest workable answer is ε-greedy: take the best-estimated action almost always, and with small probability ε pick uniformly at random instead.

ε-Greedy Action Selection
\pi(a) = \begin{cases} 1 - \varepsilon + \dfrac{\varepsilon}{|\mathcal{A}|}, & a = \arg\max_{a'} Q(a') \\[6pt] \dfrac{\varepsilon}{|\mathcal{A}|}, & \text{otherwise} \end{cases}
With k = |A| actions and ε = 0.1, the greedy arm is chosen with probability 0.1 / 2 + 0.9 = 0.95 in the two-armed case, and each arm is chosen with probability at least ε / k. That floor is the whole point: it guarantees every action keeps being sampled, so every estimate keeps being corrected. The cost is that a fraction ε of all steps are deliberately spent on an action believed to be worse.

Recomputing the mean from scratch after every pull means storing every reward. It is unnecessary: the average can be maintained one step at a time, and the form it takes is the form nearly every update rule in this module will take.

Incremental Update
Q_{n+1} = Q_n + \frac{1}{n}\left[ R_n - Q_n \right]
Read it as “old estimate, plus a step in the direction of the error.” The bracket is how surprising the latest reward was, and 1/n is the step size, which shrinks as evidence accumulates so that the estimate settles. Replace 1/n by a fixed α and the estimate instead tracks a changing world, forgetting old rewards geometrically — which is exactly the choice Lesson 14.3 makes for Q-learning.

Reward Hacking: You Get What You Measured

Everything above assumes the reward says what you meant. Specifying it is the hardest engineering task in applied reinforcement learning, and it is hard in a specific way: the agent is a relentless optimiser of the literal number, and it does not share your unstated assumptions about how that number is supposed to go up.

Reward a cleaning robot for the amount of dirt it collects and it may learn to tip the bin out and collect it again. Reward a recommender for clicks and it will learn which headlines are hardest to resist, which is not the same as which articles are worth reading. Reward a simulated runner for forward velocity and it may find a posture that exploits a bug in the physics engine rather than learning to run. In each case the agent did nothing wrong: it maximised the objective it was given.

Sutton and Barto’s Warning

The reward signal is how you tell the agent what you want achieved — not how you want it achieved. Sutton and Barto are explicit that putting subgoals into the reward is a mistake: reward the chess agent for taking pieces or controlling the centre and it may find ways to do exactly that while losing the game. Prior knowledge about method belongs in the initial policy or the initial value estimates, not in the reward.

Two practical habits follow. Write down how the agent could maximise your proposed reward while defeating your intent, before you train anything. And watch the behaviour, not only the reward curve — a reward curve going up is precisely what reward hacking looks like from the outside.

Where Reinforcement Learning Genuinely Pays

Reinforcement learning is the wrong tool far more often than it is the right one. It earns its cost when three things are true together: the decisions are sequential, so what you do now changes what you face later; the feedback is evaluative and delayed, so no one can hand you labelled correct actions; and experience is cheap enough to gather, whether from a real system or a good simulator. If any of the three fails, supervised learning is usually faster, cheaper and easier to debug.

Four settings where it genuinely fits:

Four Real Homes

Control. Robotics, flight, cooling and power systems, resource scheduling. The problem is natively sequential, a physical simulator can supply cheap experience, and the cost function is often genuinely known.

Games. The historical proving ground, because the rules are exact, the reward is unambiguous, and self-play can generate as much experience as you can afford. The clarity that makes games ideal is also what makes results there hard to transfer.

Recommendation and interaction. Which item to show next is a sequential decision under evaluative feedback, and the bandit formulation above is the honest description of a large part of it. Note the coupling from the section above is real here: the policy shapes the log, and the log is your only data.

RLHF. Reinforcement learning from human feedback, which Lesson 9.3 already covers as the step that turned a next-token predictor into a model that follows instructions. Christiano et al. (2017) established the pattern this module leads to: rather than hand-writing a reward function, collect human comparisons between pairs of short behaviour segments, fit a reward model to those preferences, and optimise the policy against the fitted reward. Lesson 14.4 returns to it once policy-gradient methods are on the table.

The honest limits belong here too, and Lesson 14.4 states them properly: reinforcement learning is sample-hungry, sensitive to reward specification, awkward to reproduce, and often defeated by the gap between simulator and reality. None of that is a reason to skip the module — it is a reason to recognise the setting when you are actually in it.

Key Takeaways
  • Reinforcement learning is the third setting Lesson 1.2 named: an agent takes actions in an environment, observes states, and receives a scalar reward — there is no dataset.
  • The reward for acting at time t is written Rt+1, because it arrives with the next state. An episode is one run of the loop to a terminal state.
  • Feedback is evaluative, not instructive: you learn the value of the action you took and nothing about the ones you did not.
  • The deeper difference from supervised learning is that the data distribution depends on the policy, so improving the policy changes the data — the problem is non-stationary by construction.
  • Reward is delayed, which makes temporal credit assignment the central problem; unlike backpropagation there is no differentiable graph to trace it through.
  • The multi-armed bandit (Robbins, 1952) isolates exploration vs. exploitation. ε-greedy keeps every action’s probability above ε/k so every estimate keeps being corrected.
  • Estimates are maintained incrementally as old value plus a step toward the error — the shape of nearly every update in this module.
  • Reward hacking is the default failure: the agent optimises what you measured. State what you want in the reward and put how in the initial policy, as Sutton and Barto advise.
  • It pays where decisions are sequential, feedback is evaluative and delayed, and experience is cheap: control, games, recommendation, and RLHF (Christiano et al., 2017; Lesson 9.3).
Previous Module 13, Lesson 4 Overview Next Markov Decision Processes