Markov Decision Processes
Lesson 14.1 described the loop. To compute anything you have to say precisely what the environment is. The MDP is that statement — and it buys you one recursive equation that the whole field is built on.
The Markov Property
The current state and action determine what comes next. History adds nothing. A chess position is Markov; one video frame of a moving ball is not — it has no velocity in it.
Five Pieces
- S — states, some of them terminal
- A — actions, possibly state-dependent
- p — transition probabilities of s′ and r
- r — the reward function, a property of the environment
- γ — discount factor, 0 ≤ γ ≤ 1
The Return
The agent is not maximising the next reward but everything after it. Acting at time t earns Rt+1, so the sum starts there — undiscounted, since γ0 = 1.
Three Reasons
- An endless sum of rewards can be infinite, and infinities cannot be compared
- Sooner is genuinely better — ten steps beats a thousand
- The far future is uncertain, so weight it less
- γ = 0 is myopic; near 1 is far-sighted and slower to learn; 0.9 to 0.99 is typical
- γ = 1 only when episodes are guaranteed to terminate
What a State Is Worth
A policy π(a | s) gives the probability of each action. The value of a state is the return you should expect from it — and only ever under some policy. The superscript is not decoration.
Value of an Action
Commit to action a first, then follow π. With V you must look ahead through the model to compare actions; with Q you compare numbers already indexed by action. That is why Lesson 14.3 learns Q.
Value, Recursively
The return splits into the next reward plus the discounted return from where you land. Average over the policy, then over the environment. One linear equation per state.
Average to Maximum
An optimal agent does not average over what it might do — it takes the best action. The expectation over the environment stays, because you cannot maximise over what the world does. The equation is no longer linear.
A Four-Cell Corridor
States 1, 2, 3 and goal G in a line. Right from 3 pays 1 and ends the episode; every other move pays 0. γ = 0.9, all values start at 0. Three sweeps of the optimality update and it is solved.
Value or Policy Iteration
- Value iteration — sweep the optimality update until nothing changes
- Policy iteration — evaluate the policy exactly, then make it greedy, repeat
- Value iteration is cheaper per sweep; policy iteration does more per round
- The policy usually becomes optimal before the values finish settling
- Both need p — and nobody hands you the dynamics of a warehouse
What You Learned
An MDP is S, A, p, r and γ, resting on the Markov property. The return discounts the future. V says what a state is worth under a policy; Q says what an action is worth. Bellman’s equations make both recursive, and dynamic programming solves them when p is known.