LLM Fundamentals
How Large Language Models actually work — from raw text to intelligent responses.
How Large Language Models Work
At their core, LLMs do one thing: predict the next token. Given a sequence of words, what comes next?
Simple: It is the autocomplete on your phone keyboard, scaled up beyond recognition. It never decides what the sentence will say before it starts — it commits to one word, looks at everything so far including that word, and commits to the next.
Technical: An autoregressive decoder-only transformer. Each forward pass emits a probability distribution over the vocabulary conditioned on the whole preceding sequence; the chosen token is appended to the input and the pass repeats. Nothing in the architecture plans ahead, and no state is carried between passes other than the text itself.
Why Counting Words Wasn’t Enough
Long before neural networks, language models simply counted word sequences. An N-gram model predicts the next word from only the previous N−1 words. Drag N from 1 to 5 and watch the two curves fight each other: the writing keeps getting better while the table you would have to store keeps exploding.
Simple: Imagine memorizing every possible LEGO combination — impossible! Too many ways to put words together.
Technical: Combinatorial explosion: vocabulary^sequence_length makes exhaustive counting intractable for sparse data.
Next-Token Prediction
Watch how an LLM generates text — one token at a time, each choice based on probabilities. These are a neural LLM’s percentages for the prompt the trigram counter just handled; the numbers differ because the mechanism does — nothing here is counted from a table.
Simple: It is a raffle where the tickets are not shared out evenly. mat holds most of them, but roof holds a few — so pressing Generate twice on the same prompt can give you two different words, and neither is a mistake.
Technical: This widget samples from the distribution rather than taking the argmax, which is why the output is not deterministic. Greedy decoding would return the top candidate every time; sampling is what makes repeated calls to the same prompt diverge.
A Probability for Every Word
The predictor above hands you one sampled token. Step back and look at the whole ballot it sampled from: the model scores every word in its vocabulary — roughly 50,000 to 100,000 of them — and the scores have to add up to 100%.
Simple: Every word the model knows gets a score on every single turn — including the absurd ones. Nothing is ever struck off the list; rocks is simply given a very small number rather than no number at all.
Technical: Softmax over the vocabulary logits returns a strictly positive probability for every entry, so no token is ever exactly zero. That property is what makes the tail reachable: raise the temperature or widen the nucleus enough and an implausible token can be drawn. The percentages shown are illustrative, hand-authored to show the shape of such a distribution.
The Power of Scale
More parameters + more data = emergent abilities that smaller models simply don’t have.
Simple: Nobody sat down and added a reasoning feature. They built the same thing much bigger, fed it much more text, and abilities the smaller versions did not have started showing up.
Technical: Scaling laws describe loss falling predictably as parameters, data and compute grow together — predictable enough to plan a training run around. The two GPT-3 figures above are from Brown et al., 2020; frontier models since are trained on far more tokens, so read the trend rather than the numbers.
How a Model Gets Built: Three Stages
A finished assistant is not trained once. It goes through three separate passes, each with a different teacher and a very different price tag. Click a stage for the full picture.
Simple: The three stages answer three different questions. What does language look like? What does a good answer look like? And when two answers are both good, which one do people actually prefer?
Technical: Three different training objectives, not three helpings of the same one: a self-supervised next-token objective over raw text, then supervised imitation of demonstration pairs, then preference optimisation against a reward model learned from human comparisons. The figures are public order-of-magnitude estimates and date quickly.
What Are Tokens?
LLMs don’t read characters or words — they read tokens. A token is a chunk of text, often a word or part of a word.
Simple: The model reads in chunks about the size of syllables. And once the chunking is done, your letters are gone — each chunk is swapped for a number, and numbers are all the model ever sees.
Technical: Each token is an integer index into a fixed vocabulary, used to look up a row of the embedding matrix. The character sequence itself never reaches the network, which is why character-level tasks are awkward for a model that is otherwise fluent.
Guess Before You Look
Your intuition says “count the words”. The tokenizer disagrees. Guess first, then see the actual pieces — the gap between the two is the whole lesson.
Simple: Your instinct counts words because that is what you read. The tokenizer counts something else, and the gap is not random — a word it has seen a million times costs one piece, a word it has barely seen costs four.
Technical: Token count tracks corpus frequency, not word count or character count. Frequent strings survive as single vocabulary entries; rare ones are decomposed into subword units. This is why token count is the only honest unit for cost and context budgeting.
Tokenizer Playground
Type any text to see how it gets split into tokens. Each color is a different token.
Simple: Paste in prose and the pieces line up roughly with words. Paste in code and it shatters — brackets, indentation and underscores each cost their own piece, which is why the same amount of code costs more than the same amount of English.
Technical: The chars-per-token ratio is a workload property, not a constant: English prose runs around 4, code and heavily inflected or non-Latin scripts run far lower. Note the widget itself is a heuristic approximation splitting on common affixes, not a learned BPE vocabulary, so a production tokenizer will sometimes disagree with it.
Byte Pair Encoding (BPE)
The most common tokenization algorithm. It starts with characters and merges the most frequent pairs iteratively.
Simple: Nobody wrote the list of chunks by hand. It was found by reading a mountain of text and repeatedly gluing together whichever pair of neighbours turned up most often — so the chunks that exist are just the ones the text kept repeating.
Technical: BPE is a greedy bottom-up merge procedure. Start from a character-level vocabulary, count adjacent symbol pairs across the corpus, merge the most frequent, and repeat for a fixed number of merges. The ordered merge list is the tokenizer, and it is frozen before training — a model cannot change how it is chunked later.
Why Tokenization Matters
Simple: Tokenization is not a technicality you can leave to the plumbing. It sets your bill, how much of your document fits, and whether an Arabic sentence costs you two or three times what the same sentence costs in English.
Technical: The tokenizer is trained mostly on English, so scripts under-represented in that corpus fragment into more tokens per unit of meaning. That penalty compounds: billing is per token, the context window is denominated in tokens, and latency scales with sequence length — one asymmetry, charged three times.
Words as Numbers
Each token becomes a high-dimensional vector. Similar words cluster together. Drag the space to rotate it, or focus it and use the arrow keys.
Type a word or click an example to see where it lands in the semantic space…
Further reading and videos for this whole module are collected on the title slide.
Simple: Give every word a spot on a map, and put words that get used in similar situations near each other. Nobody drew the map — it fell out of the fact that king and queen keep turning up in the same kinds of sentences.
Technical: Each vocabulary entry maps to a dense vector, learned by gradient descent on the training objective rather than assigned. Proximity is measured by cosine similarity, and what it encodes is distributional similarity — similar contexts of use — which is not the same thing as similar meaning. The space above is a 3-D toy of a space with hundreds or thousands of dimensions.
Word Arithmetic
Embeddings encode relationships as directions. You can do math with meaning — pick an analogy and watch the vectors line up.
Simple: On this map, direction carries meaning too, not just position. The step that takes you from man to woman is roughly the same step, in the same direction, that takes you from king to queen.
Technical: Certain semantic and syntactic relations appear as approximately parallel offset vectors, so an analogy can be posed as vector addition and resolved by nearest-neighbour search. It is approximate, not exact: the arithmetic lands near the target, the query terms are normally excluded from the search, and the regularity holds far better for some relation types than others.
Inside an LLM
Text flows through a pipeline of transformations. Click each stage to learn more.
Simple: Text goes in the front and a next-word guess comes out the back, through four stations in a fixed order. The one marked ×N is the same station repeated — not a different one each time.
Technical: Tokenizer → embedding lookup → a stack of identical transformer blocks → an output projection to vocabulary logits. Each block is architecturally the same as every other — self-attention followed by a feed-forward network, with residual connections and normalisation — but each carries its own learned weights.
How Deep Is the Stack?
The “×N” in the diagram above is doing a lot of work. Drag the depth and send a token through.
Simple: Depth is not more kinds of thinking; it is the same operation applied again to its own output. Each pass refines a representation that is already partly built, the way successive edits sharpen a draft.
Technical: Every block reads the residual stream and writes back into it, so representations accumulate rather than being replaced. Depth costs compute and latency linearly, which is why depth is traded against width and context length rather than simply maximised. The block counts here are real orders of magnitude, not one product's specification.
Q, K, V — Three Roles
How does the model know that “it” refers to “cat”? Attention lets each word look at every other word.
Simple: Think of a room where everyone is holding up a card saying what they can answer, and everyone else is holding up a question. You read every card, decide who is worth listening to, and take a bit from each — more from the useful ones.
Technical: Q, K and V are three separate learned linear projections of the same input vector. Compatibility is the dot product of queries against keys, scaled by √dk to keep the pre-softmax scores in a range where gradients survive; softmax normalises them into weights, and the output is that weighted sum of value vectors.
Attention, Step by Step
The formula in one go is a lot. Here it is as the four things it actually does, one at a time.
Simple: Written as one formula it looks like a single leap. It is four small moves: ask, score everyone, turn the scores into shares that add to a whole, then blend.
Technical: The four steps and their order are exactly those of scaled dot-product attention (Vaswani et al., 2017, cited in full in the caption above), and only the third is non-linear — softmax is what turns arbitrary real-valued scores into a normalised, comparable set of weights. The weights in the widget are illustrative and hand-authored, not read out of a trained model.
Attention Heatmap
Each row is a word asking its question; each cell shows how much of the answer came from the word in that column. Different heads learn different jobs.
Simple: The model does not do this once. It does it several times side by side with different priorities — one pass chasing who a pronoun refers to, another watching for the word right next door — and then pools what they all found.
Technical: Multi-head attention runs several attention operations in parallel over lower-dimensional projections and concatenates the results. Head specialisation like coreference or syntactic dependency is an observed tendency in probing studies, not something specified in the architecture or trained for; the labels on the buttons above are an illustration of that finding.
Why Attention Wins
Transformers replaced RNNs and LSTMs because attention can process all tokens in parallel.
Simple: The older designs read a sentence the way you listen to a voicemail — strictly in order, one word at a time, no skipping. Attention reads it the way you read a page: everything is on the table at once, and word 3 can look straight at word 300.
Technical: Recurrent models impose a sequential dependency: step t cannot be computed before step t−1, so the sequence dimension cannot be parallelised across training, and the path between distant tokens grows with the distance between them. Self-attention makes that path length constant and the sequence dimension fully parallel. The cost is quadratic in sequence length — the trade that buys the speed-up.
Hallucination — When AI Makes Things Up
LLMs can generate text that sounds confident and fluent but is completely wrong. This is called hallucination.
Simple: There is no moment where the model checks whether what it is about to say is true. Fluent and correct feel identical from the inside — which is why a fabricated citation reads exactly as smoothly as a real one.
Technical: The training objective rewards likely continuations, and there is no separate truth signal to optimise against. Nothing in the forward pass consults a knowledge base or a verifier, so confidence in the output text is a property of the token distribution, not evidence about the world.
Why Long Chats Start Drifting
Past a point, more text in the window does the opposite of helping: the early details blur, and the model fills the blur with plausible inventions. Drag the slider.
Simple: The dangerous part is not that detail goes missing. It is that the gap gets filled. You get a confident, specific, plausible answer in the shape you asked for — and nothing marks the part that was invented.
Technical: Attention spreads over a longer sequence, and retrieval of any one earlier fact degrades as the window fills. The model still emits a high-probability completion, so the failure surfaces as substitution rather than as an abstention. This is hallucination, not forgetting — and the related “lost in the middle” position effect is covered in Module 5.
Fighting Hallucination
Several strategies can reduce (but not eliminate) hallucination.
Simple: You cannot make it stop inventing, so change what it has to invent from. Put the real document in front of it and the plausible answer and the true answer become the same answer.
Technical: These strategies attack the problem at different layers — input (grounding, retrieval), training (alignment to refuse under uncertainty), and output (citation plus external verification). None removes hallucination; each narrows the space in which a fabrication is the most probable continuation. Verification is the only one that catches a failure after it has happened.
Temperature — Controlling Creativity
Temperature controls how “creative” the model is. Drag the slider to see how probabilities change.
Simple: One dial, and it does not add or remove ideas — it decides how much of an underdog the model is willing to back. Turn it down and the favourite wins every time; turn it up and the long shots start getting picked.
Technical: Temperature divides the logits before the softmax. Below 1 it sharpens the distribution towards the mode; above 1 it flattens it, raising the mass on the tail this module showed you two slides in. It reshapes the existing distribution and never introduces a candidate that was not already scored. The outputs shown per band here are canned illustrations, not live generations.
Sampling Strategies
Temperature isn’t the only knob. Here are the main ways to control LLM output.
Simple: Two different questions, easily confused. Temperature asks how lopsided the odds should be. Top-k and top-p ask how many candidates are even allowed on the ballot before you draw.
Technical: Temperature rescales the distribution; top-k and top-p truncate it. Top-k keeps a fixed number of candidates, top-p keeps the smallest set whose cumulative probability reaches p — so top-p adapts its cut-off to how confident the model is at that step, and top-k does not. They compose, which is why most APIs expose several at once.
Knowledge Check
Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.