AI-generated After a published paper Illustrative numbers

This interactive diagram was generated by Claude (Anthropic) working from the paper below, as part of the Wireless 101 course materials. It reproduces the visual design of Figure 1 of:

Aeree Cho, Grace C. Kim, Alexander Karpekov, Seongmin Lee, Alec Helbling, Benjamin Hoover, Zijie J. Wang, Minsuk Kahng, Duen Horng Chau. Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation. CHI '26. DOI 10.1145/3772318.3791725. CC BY 4.0.

It does not run a model. The original tool runs a live GPT‑2 in the browser; this page is static HTML with no inference, so every number is one of three kinds and each is labelled where it appears: recovered from the paper's own figures, a published GPT‑2 Small architecture fact, or illustrative. The temperature, top‑k, top‑p and softmax arithmetic is real; the attention values are not measurements. Head behaviours reproduce shapes reported in probing studies, marked ● documented or ○ plausible-shape.

Transformer, end to end — Sankey prototype

After Fig. 1 of Transformer Explainer (Cho et al., CHI '26). Course demo — linked from the Module 2 deck in both languages. Hover a token to follow its path; click a ⚲ to expand a stage in place; page between blocks and heads. Logits are recovered from the paper's Fig. 4B and the probabilities are computed live — see the file header comment for which numbers are real and which are illustrative.

Transformer Block 1 of 12 Head 1 of 12

What actually comes out

0.80

The two pagers move along different axes

Blocks — depth. Sequential. Each one is fed by the one before it, so the order matters: swap two blocks and the model breaks.
Heads — width. Parallel, inside one block. All twelve read the same input at the same time, so the order is meaningless: swap two heads and nothing changes but the labels.

What each block does

Why the colours are what they are. §6.1 p.7: token embeddings are grayscale “to emphasize that they are the untransformed initial representations”; then “Q is blue and K is red, and the resulting attention scores appear in purple, visually reflecting the combination of the two.” That purple is doing teaching work prose cannot: it shows the score is the product of two things. The MLP stays a single hue because it expands the representation rather than giving it a new role. Our amber/ink palette has no blue, red or green, so the four hexes here were re-derived and measured against cream and against the dark surface — all eight clear 4.5:1 as text and as text-on-fill. The score grid's ramp inverts between themes, because pale purple is 8.77:1 on the dark surface and 1.74:1 on cream.
Ported or not. Nothing here is wired into a deck yet. Landing it in m2-l1-slides.html means a ~590px-tall reveal.js box, so the end-to-end view will have to become either a horizontally-scrolling band or three linked stages on the slides that already exist (slide-pipeline, slide-attention-steps, slide-temperature). That decision is deliberately deferred — see ref/papers/transformer-explainer.md §9.3 R1/R5/R6.

The arithmetic, in full