AI-generatedAfter a published paperIllustrative numbers
This interactive diagram was generated by Claude (Anthropic) working from the paper
below, as part of the Wireless 101 course materials. It reproduces the visual design
of Figure 1 of:
Aeree Cho, Grace C. Kim, Alexander Karpekov, Seongmin Lee, Alec Helbling, Benjamin Hoover,
Zijie J. Wang, Minsuk Kahng, Duen Horng Chau.
Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and
Experimentation. CHI '26. DOI
10.1145/3772318.3791725.
CC BY 4.0.
It does not run a model. The original tool runs a live GPT‑2 in the browser; this
page is static HTML with no inference, so every number is one of three kinds and each is
labelled where it appears: recovered from the paper's own figures, a published
GPT‑2 Small architecture fact, or illustrative. The temperature, top‑k,
top‑p and softmax arithmetic is real; the attention values are not measurements.
Head behaviours reproduce shapes reported in probing studies, marked ● documented or
○ plausible-shape.
Transformer, end to end — Sankey prototype
After Fig. 1 of Transformer Explainer (Cho et al., CHI '26). Course demo —
linked from the Module 2 deck in both languages. Hover a token to follow its path;
click a ⚲ to expand a stage in place; page between blocks and heads.
Logits are recovered from the paper's Fig. 4B and the probabilities are computed live —
see the file header comment for which numbers are real and which are illustrative.
Transformer Block 1 of 12Head 1 of 12
Tokenizer — why it is not words
Byte-pair encoding merges the most frequent character pairs until it has a
fixed vocabulary. Common words survive whole; rarer ones get cut into familiar pieces.
The two things people get wrong.1.empowers is one word but two tokens —
em + powers. The model never
sees the word; it sees the pieces, and has to reassemble the meaning. That is what the
“Subword continuation” head is for.
2. The ␣ marks a leading space, which belongs to the
token after it. “ visualization” and
“visualization” are different tokens with different IDs.
Numbers: the split shown is real GPT‑2 BPE behaviour and the IDs
6681, 32784, 795
are the paper's (Fig. 5); the other three IDs are illustrative — this repo has
no tokenizer to look them up with.
Embedding — token ID, then two vectors added
Each token becomes an ID, the ID looks up a learned vector, a positional
vector is added, and the sum is what enters the first block.
token IDtoken embedding / + positional encoding (sinusoid — GPT‑2 learns
its own) / = indicative sum · vector(768)
Numbers: IDs 6681, 32784,
795 for Data, visualization, em are the paper's
(Fig. 5); the other three IDs and the token-embedding row are
illustrative. 768 is GPT‑2 Small's published width.
The positional row is computed from the real formula —
PE(pos,2i) = sin(pos / 100002i/d) (Vaswani et al. 2017
§3.5) — which is why its left edge oscillates fast and its right edge is nearly flat.
Caveat worth knowing: GPT‑2 actually learns its positional embeddings
rather than using this fixed sinusoid; the sinusoid is drawn because it is the version that
can be computed honestly here.
Grayscale is deliberate — §6.1 p.7, “to emphasize that they are the untransformed
initial representations.”
Self-attention, in three steps
The paper shows these side by side rather than in sequence, so the
intermediate results can be compared instead of remembered (§6.2 p.7).
Numbers: the raw dot products are illustrative, but their
shape is not arbitrary. Five reproduce head classes that are named and documented
— previous-token, broad/averaging, syntactic-dependency and rare-token
(Clark et al. 2019 What Does BERT Look At?; Voita et al. 2019
Analyzing Multi-Head Self-Attention; Olsson et al. 2022 for the decoder case)
and the attention sink (Xiao et al. 2023, not Clark). Three — self/identity,
recency decay and subword continuation — are plausible shapes rather than named
classes, and the panel marks which is which with ● and ○.
Cross-architecture caveat: Clark is a bidirectional encoder and Voita an NMT encoder,
so applying them to GPT‑2's decoder heads is an extrapolation.
And the ratio here is not the real ratio: eight named heads are drawn one per head so
you can see the catalogue, but in a real model the cleanly interpretable heads are the
minority — most look like heads 9–12 (Voita et al. found most heads
prunable; Michel et al. 2019 likewise).
Depth: clarity peaks early for positional heads and mid-stack for syntactic ones, and
decays by block 12 — page the block and watch the readout.
Ranges printed above each panel are computed from the matrix on screen, not copied from
the paper: its “−3.0…3.0” for panel 2 cannot follow from its
“−25.3…7.4” for panel 1, since dividing by √64 caps panel 2
near +0.9. Everything after panel 1 is real arithmetic: ÷√64, the causal mask,
softmax. Read the middle panel: the upper triangle goes to
−∞ because a token cannot attend to a token that has not
arrived yet — and it is drawn, not deleted. Panel 3 is softmax only; dropout is
training-time and would break the rows summing to 1, so it is not applied here.
Not shown, on purpose:induction heads — the best-documented class in
decoder-only models (Olsson et al. 2022: seeing “… A B … A”,
predict B). They need a repeated bigram in the context, and this six-token prompt
has none, so an induction head here would be indistinguishable from a broad one. Drawing it
would mean inventing a pattern the prompt cannot produce.
From one vector to a ranked guess
Left to right, in the order the operations actually run
(§6.2 p.8): logits → divide by temperature → filter → softmax.
One caveat about that order: top-k is rank-based, so it can act on the
logits directly — but top-p needs normalised probabilities to accumulate,
so it is computed on the softmax and the survivors are then renormalised. The columns
below are laid out in teaching order, not in strict execution order, for that one step.
Numbers: the five base logits are recovered from the paper's own
Fig. 4B (it prints −136.68 / 0.8 = −170.85 and
writes the softmax out in full), and at T = 0.8 they
reproduce its published percentages to within 0.07 points. Every other row's logit is
illustrative. The temperature, the filtering and the renormalisation above are
exact — arithmetic, not illustration.
What actually comes out
The two pagers move along different axes
Blocks — depth. Sequential. Each one is fed by the one before
it, so the order matters: swap two blocks and the model breaks.
Heads — width. Parallel, inside one block. All twelve
read the same input at the same time, so the order is meaningless: swap two heads and
nothing changes but the labels.
What each block does
Why the colours are what they are. §6.1 p.7: token embeddings are grayscale
“to emphasize that they are the untransformed initial representations”; then
“Q is blue and K is red, and the resulting attention scores appear in purple,
visually reflecting the combination of the two.” That purple is doing teaching work prose
cannot: it shows the score is the product of two things. The MLP stays a single hue
because it expands the representation rather than giving it a new role.
Our amber/ink palette has no blue, red or green, so the four hexes here were re-derived and
measured against cream and against the dark surface — all eight clear 4.5:1 as text and as
text-on-fill. The score grid's ramp inverts between themes, because pale purple is
8.77:1 on the dark surface and 1.74:1 on cream.
Ported or not. Nothing here is wired into a deck yet. Landing it in
m2-l1-slides.html means a ~590px-tall reveal.js box, so the end-to-end view will
have to become either a horizontally-scrolling band or three linked stages on the slides that
already exist (slide-pipeline, slide-attention-steps,
slide-temperature). That decision is deliberately deferred — see
ref/papers/transformer-explainer.md §9.3 R1/R5/R6.