Move the temperature, then the decoding controls in the cross-course transformer demo, and predict — before you read it off the probabilities panel — how the final softmax turns a row of logits into a probability distribution, how raising T flattens it, why greedy decoding ignores T entirely, and what top-k does to the surviving probabilities.
The last thing a transformer does is turn a row of scores into the odds of each next token, and the knob that shapes those odds is temperature. This is arithmetic you can check by hand: softmax is a genuine probability distribution, so it must sum to 1; dividing every score by the same T is monotonic, so it flattens the odds without ever reordering them — which is exactly why greedy decoding, taking the single highest score, cannot be changed by T at all.
Open the demo (it lives in the GenAI 101 course, linked from this lab). It opens on exactly this state, so a Reset gets you here too. The output panel carries the live controls — temperature, the filter radios, and the greedy/sample switch.
Temperature T -> 0.80 Filter -> none
Decoding -> sampling Candidates listed -> 15
Watch the probabilities panel — every probability you predict is printed there. The 15 next-token candidates are the demo's own; you are checking what the softmax does to them, not their individual values.
On the opening state (T = 0.80, filter none), read the probabilities panel. There are 15 candidates. A softmax is defined to be a probability distribution, so its values must add up to something exact. Predict the sum of all 15 probabilities, then read them and add them.
They sum to exactly 1.00. That is what makes softmax a distribution rather than a list of scores: e^(z_i/T) divided by the sum of all e^(z_j/T) is normalised by construction, whatever the logits are.
Now drag the temperature through six stops — T = 0.3, 0.5, 0.8, 1.0, 1.5, 2.0 — and each time note the top token's probability. Dividing every logit by a larger T pulls the scaled values together, so the favourite loses ground at every step. Predict, across those six stops, how many of the five transitions show the top probability fall.
All 5 of them. The top probability falls at every single step as T climbs — this is flattening: at low T the softmax is peaked and near-greedy, and as T grows it spreads toward uniform. The ordering never changes, only the sharpness.
Switch decoding to greedy (argmax). Greedy always emits the single highest-probability token. Now sweep T across 0.1, 0.5, 1.0, 2.0. Because dividing by T cannot reorder the logits, the argmax cannot move. Predict how many different tokens greedy emits across those four temperatures.
Exactly 1. The emitted token is identical at T = 0.1 and T = 2.0 — temperature changes the odds, but greedy never rolls the dice, so it is invariant to T. Most explanations of temperature quietly imply otherwise; the arithmetic says one token, always.
Set decoding back to sampling, then choose the top-k filter with k = 3 (leave T = 0.80). Top-k keeps the 3 highest-probability candidates, zeroes the rest, and renormalises the survivors so they can be sampled from. Predict what those 3 surviving probabilities sum to after renormalisation.
Back to exactly 1.00. Truncation alone would leave the 3 survivors summing to less than 1, so the demo divides them by their own total — the same normalisation as the full softmax, now over just the kept set. top-k reshapes which tokens can be drawn without ever breaking the distribution.
Everything above is waiting in the demo. Drag the temperature and watch the whole distribution sharpen and flatten, flip between greedy and sampling to see the emitted token freeze or move, and toggle top-k and top-p to watch candidates drop out and the survivors renormalise in real time.
- Confirmed the softmax over all 15 candidates sums to 1.00 — a real distribution
- Watched raising T flatten the odds, the top probability falling at all 5 steps of the sweep
- Saw greedy emit 1 token, unchanged across T = 0.1 … 2.0 — temperature-invariant
- Filtered with top-k and watched the 3 survivors renormalise back to 1.00
- Every value you predicted is the demo's own softmax/top-k arithmetic — never a recovered or illustrative logit