This sandbox was generated by Claude (Anthropic) for the ML 101 course materials. It builds a real character-level n-gram language model on a small built-in text corpus and shows the four things the Module 8 lesson describes in words: the next-character distribution for a given context, how temperature sharpens or flattens sampling, the perplexity on held-out text, and text generated by sampling from the model. Nothing here is sketched — every bar, curve and generated character comes from the counts.
This is an n-gram count model, not a neural network. It has no embeddings, no attention and
no learned weights — it simply counts how often each character follows each context of the
previous n−1 characters and normalises, with add-k smoothing so an unseen
character is improbable rather than impossible. That is the transparent baseline the course builds
on. Nothing here is presented as a learned representation.
Perplexity is reported in bits (log base 2). A uniform guess over a vocabulary of
V characters then has perplexity exactly V; a perfectly predictable
sequence has perplexity 1. The equation block states the base explicitly, so the
numbers on screen mean what they say.
Computed: the n-gram counts on the training split, the add-k conditional distribution for any context, the vocabulary size, the number of distinct contexts, the training and held-out perplexity and average log-loss, the perplexity-vs-order curve, and every generated sample (from a seeded PRNG, so it is reproducible). Published figures: none — everything is a definition or a direct count. Chosen rather than computed: the built-in corpora, the train/held-out split fraction, the seed, and the defaults for order, smoothing, temperature and sample length. Raise the order on a small corpus and watch training perplexity fall while the counts grow sparse — the same overfitting trade-off the course warns about, computed on screen.
Course demo — linked from the Module 8 lesson deck; the page itself is English‑only for now. A language model predicts the next symbol from the ones before it. This page lets you set the order (how many previous characters count as context), the smoothing, the sampling temperature and the seed, and then computes the whole model and its perplexity from the text. One thing to take away: a higher order fits the training text better but needs far more data — which is exactly why real models are trained on enormous corpora.