LLM Fundamentals

How Large Language Models actually work — from raw text to intelligent responses.

How Large Language Models Work

At their core, LLMs do one thing: predict the next token. Given a sequence of words, what comes next?

🔮
Next-Token Prediction
The fundamental operation: given context, predict the most likely continuation.
📚
Training Data
Trillions of words from books, websites, code, and conversations.
⚙️
Parameters
Billions of learned weights that encode knowledge and language patterns.
✨
Emergent Abilities
Reasoning, creativity, and understanding that arise from scale alone.
In Plain English

Simple: It is the autocomplete on your phone keyboard, scaled up beyond recognition. It never decides what the sentence will say before it starts — it commits to one word, looks at everything so far including that word, and commits to the next.

Technical: An autoregressive decoder-only transformer. Each forward pass emits a probability distribution over the vocabulary conditioned on the whole preceding sequence; the chosen token is appended to the input and the pass repeats. Nothing in the architecture plans ahead, and no state is carried between passes other than the text itself.

Why Counting Words Wasn’t Enough

Long before neural networks, language models simply counted word sequences. An N-gram model predicts the next word from only the previous N−1 words. Drag N from 1 to 5 and watch the two curves fight each other: the writing keeps getting better while the table you would have to store keeps exploding.

🔢
The N-gram Counting Problem
Predict from previous words by counting — then watch the combinations explode
1 Gibberish 2 Fragmented 3 Choppy 4 Readable 5 Natural
Possible sequences with a 50,000-word vocabulary = 50,000N
This is why we needed embeddings and transformers. Instead of memorizing counts for every sequence, models learn meaning (embeddings) and which earlier words matter (attention) — so they generalize to sentences they have never seen. Both arrive later in this module: embeddings in Lesson 3, attention in Lesson 4; everything in between is about getting text into the model in the first place. The percentages above are a trigram counter’s, straight from sequence counts — the next slide shows a neural LLM’s for the same prompt, which is why the two tables disagree.
In Plain English

Simple: Imagine memorizing every possible LEGO combination — impossible! Too many ways to put words together.

Technical: Combinatorial explosion: vocabulary^sequence_length makes exhaustive counting intractable for sparse data.

Next-Token Prediction

Watch how an LLM generates text — one token at a time, each choice based on probabilities. These are a neural LLM’s percentages for the prompt the trigram counter just handled; the numbers differ because the mechanism does — nothing here is counted from a table.

🎯
Token Predictor
Click “Generate” to sample the next token based on probabilities
The cat sat on the
The percentages are illustrative and hand-authored to show the shape of a next-token distribution — they are not read out of a trained model. The sampling itself is real: each click draws from exactly these probabilities, which is why repeating it gives different words.
In Plain English

Simple: It is a raffle where the tickets are not shared out evenly. mat holds most of them, but roof holds a few — so pressing Generate twice on the same prompt can give you two different words, and neither is a mistake.

Technical: This widget samples from the distribution rather than taking the argmax, which is why the output is not deterministic. Greedy decoding would return the top candidate every time; sampling is what makes repeated calls to the same prompt diverge.

A Probability for Every Word

The predictor above hands you one sampled token. Step back and look at the whole ballot it sampled from: the model scores every word in its vocabulary — roughly 50,000 to 100,000 of them — and the scores have to add up to 100%.

📊
The whole ballot, not the winner
Six of the candidates for one blank — including two nobody would pick
I love eating???
One prompt, one blank, one score for every possible filler.
Percentages are illustrative, hand-authored to show the shape of a next-word distribution — they are not read out of a trained model, the same convention this module uses for its trigram-versus-neural numbers. The point is the last two bars: rocks and clouds are marked implausible and drawn dashed and grey, yet they still carry a number. Nothing is ever zero. That is what “a probability distribution over the vocabulary” means, and it is why raising the temperature can eventually let one of them through.
In Plain English

Simple: Every word the model knows gets a score on every single turn — including the absurd ones. Nothing is ever struck off the list; rocks is simply given a very small number rather than no number at all.

Technical: Softmax over the vocabulary logits returns a strictly positive probability for every entry, so no token is ever exactly zero. That property is what makes the tail reachable: raise the temperature or widen the nucleus enough and an implausible token can be drawn. The percentages shown are illustrative, hand-authored to show the shape of such a distribution.

The Power of Scale

More parameters + more data = emergent abilities that smaller models simply don’t have.

175B
GPT-3 Parameters
~300B
GPT-3 Training Tokens
10K+
GPUs to Train — widely reported, never confirmed by OpenAI
Scaling laws: Performance improves predictably as you increase model size, data, and compute. This is why the race for bigger — and smarter — models continues. The two GPT-3 figures are from Brown et al., 2020; frontier models since then are trained on far more tokens — see the next slide.
In Plain English

Simple: Nobody sat down and added a reasoning feature. They built the same thing much bigger, fed it much more text, and abilities the smaller versions did not have started showing up.

Technical: Scaling laws describe loss falling predictably as parameters, data and compute grow together — predictable enough to plan a training run around. The two GPT-3 figures above are from Brown et al., 2020; frontier models since are trained on far more tokens, so read the trend rather than the numbers.

How a Model Gets Built: Three Stages

A finished assistant is not trained once. It goes through three separate passes, each with a different teacher and a very different price tag. Click a stage for the full picture.

🏗️
Read, then practise, then take feedback
Pre-training · supervised fine-tuning · alignment
Stage 1
Reading the library
pre-training · self-supervised
Guess the next word in real text, over and over, until grammar and facts stick. Nobody labels anything — the text is its own answer key.
Data10T+ tokens
Cost$100M+
Time3–6 months
Stage 2
Learning the job
fine-tuning · SFT
Show it worked examples of a good answer to a request, so it stops merely continuing text and starts replying to you.
Data100K+ examples
Cost$10K–1M
TimeDays
Stage 3
Taking feedback
alignment · RLHF
People compare two answers and pick the better one. The model learns to aim for the kind of answer people actually preferred.
Data500K+ comparisons
CostOngoing
TimeNever finished
Only Stage 1 needs a supercomputer. Stages 2 and 3 are comparatively cheap — which is why so many different assistants can be built on top of a handful of base models. Figures are public order-of-magnitude estimates for frontier models and date quickly.
In Plain English

Simple: The three stages answer three different questions. What does language look like? What does a good answer look like? And when two answers are both good, which one do people actually prefer?

Technical: Three different training objectives, not three helpings of the same one: a self-supervised next-token objective over raw text, then supervised imitation of demonstration pairs, then preference optimisation against a reward model learned from human comparisons. The figures are public order-of-magnitude estimates and date quickly.

What Are Tokens?

LLMs don’t read characters or words — they read tokens. A token is a chunk of text, often a word or part of a word.

“Tokenization”
→
“Token”
“ization”
Why not characters? Individual characters carry too little meaning. Why not words? Too many unique words, and morphology varies wildly. Tokens are the Goldilocks unit — meaningful chunks that keep vocabulary manageable.
In Plain English

Simple: The model reads in chunks about the size of syllables. And once the chunking is done, your letters are gone — each chunk is swapped for a number, and numbers are all the model ever sees.

Technical: Each token is an integer index into a fixed vocabulary, used to look up a row of the embedding matrix. The character sequence itself never reaches the network, which is why character-level tasks are awkward for a model that is otherwise fluent.

Guess Before You Look

Your intuition says “count the words”. The tokenizer disagrees. Guess first, then see the actual pieces — the gap between the two is the whole lesson.

🎲
Token Guessing Game
Four strings, canned answers from a real subword vocabulary
How many tokens is this string?
Score 0 / 0
Common words are single pieces; unusual words get chopped. That is why one long invented word can cost more than a whole short sentence. Algorithm: byte pair encoding — the next-but-one slide.
In Plain English

Simple: Your instinct counts words because that is what you read. The tokenizer counts something else, and the gap is not random — a word it has seen a million times costs one piece, a word it has barely seen costs four.

Technical: Token count tracks corpus frequency, not word count or character count. Frequent strings survive as single vocabulary entries; rare ones are decomposed into subword units. This is why token count is the only honest unit for cost and context budgeting.

Tokenizer Playground

Type any text to see how it gets split into tokens. Each color is a different token.

🧩
Tokenizer Playground
A heuristic approximation, not a real BPE vocabulary: it splits on common prefixes and suffixes instead of learned merges, so a production tokenizer will sometimes disagree with it. A leading · marks the space that belongs to the token, and the ID under each piece is an illustrative placeholder, not a real vocabulary ID.
Tokens: 0 Characters: 0 Ratio: 0 chars/token
Try:
In Plain English

Simple: Paste in prose and the pieces line up roughly with words. Paste in code and it shatters — brackets, indentation and underscores each cost their own piece, which is why the same amount of code costs more than the same amount of English.

Technical: The chars-per-token ratio is a workload property, not a constant: English prose runs around 4, code and heavily inflected or non-Latin scripts run far lower. Note the widget itself is a heuristic approximation splitting on common affixes, not a learned BPE vocabulary, so a production tokenizer will sometimes disagree with it.

Byte Pair Encoding (BPE)

The most common tokenization algorithm. It starts with characters and merges the most frequent pairs iteratively.

Step 1
Start with Characters
“lower” → [l, o, w, e, r] — every character is its own token
Step 2
Count Pairs
Find the most frequent adjacent pair across all training text (e.g., “e”+“r” appears 10,000 times)
Step 3
Merge & Repeat
Replace “e”+“r” with “er” everywhere. Repeat for 50K+ merge operations.
Result
A Vocabulary of Subwords
“lower” → [“low”, “er”] — common words stay whole, rare words get split.
In Plain English

Simple: Nobody wrote the list of chunks by hand. It was found by reading a mountain of text and repeatedly gluing together whichever pair of neighbours turned up most often — so the chunks that exist are just the ones the text kept repeating.

Technical: BPE is a greedy bottom-up merge procedure. Start from a character-level vocabulary, count adjacent symbol pairs across the corpus, merge the most frequent, and repeat for a fixed number of merges. The ordered merge list is the tokenizer, and it is frozen before training — a model cannot change how it is chunked later.

Why Tokenization Matters

💰
Cost
API pricing is per token. More tokens = higher cost for the same text.
🌍
Multilingual
Arabic and CJK can use 2–4× more tokens than English for the same meaning.
📏
Context Window
Models have a maximum token limit. Tokenization determines how much fits.
⚠️
Edge Cases
Numbers, code, and unicode can tokenize unexpectedly — affecting model behavior.
In Plain English

Simple: Tokenization is not a technicality you can leave to the plumbing. It sets your bill, how much of your document fits, and whether an Arabic sentence costs you two or three times what the same sentence costs in English.

Technical: The tokenizer is trained mostly on English, so scripts under-represented in that corpus fragment into more tokens per unit of meaning. That penalty compounds: billing is per token, the context window is denominated in tokens, and latency scales with sequence length — one asymmetry, charged three times.

Words as Numbers

Each token becomes a high-dimensional vector. Similar words cluster together. Drag the space to rotate it, or focus it and use the arrow keys.

Animals Royalty Emotions Tech Food
X
Y
Z
dog
cat
lion
tiger
wolf
puppy
king
queen
prince
princess
throne
crown
computer
laptop
phone
tablet
code
apple
banana
orange
pizza
grape
happy
sad
angry
joy
Try:

Type a word or click an example to see where it lands in the semantic space…

Positions here are hand-authored for teaching, not a projection of any trained model, and the word test is a small suffix-and-keyword heuristic over a 22-word list — not a real nearest-neighbour lookup.
Further reading and videos for this whole module are collected on the title slide.
📐
Cosine Similarity
Measure how close two word vectors are by the angle between them.
🎯
Semantic Clustering
Words with similar meanings naturally group together in vector space.
In Plain English

Simple: Give every word a spot on a map, and put words that get used in similar situations near each other. Nobody drew the map — it fell out of the fact that king and queen keep turning up in the same kinds of sentences.

Technical: Each vocabulary entry maps to a dense vector, learned by gradient descent on the training objective rather than assigned. Proximity is measured by cosine similarity, and what it encodes is distributional similarity — similar contexts of use — which is not the same thing as similar meaning. The space above is a 3-D toy of a space with hundreds or thousands of dimensions.

Word Arithmetic

Embeddings encode relationships as directions. You can do math with meaning — pick an analogy and watch the vectors line up.

🔄
Analogies
Paris:France :: Tokyo:? The vector math gives you “Japan”.
📊
High Dimensions
Real embeddings use 768–4096 dimensions to capture nuanced meaning.
In Plain English

Simple: On this map, direction carries meaning too, not just position. The step that takes you from man to woman is roughly the same step, in the same direction, that takes you from king to queen.

Technical: Certain semantic and syntactic relations appear as approximately parallel offset vectors, so an analogy can be posed as vector addition and resolved by nearest-neighbour search. It is approximate, not exact: the arithmetic lands near the target, the query terms are normally excluded from the search, and the regularity holds far better for some relation types than others.

Inside an LLM

Text flows through a pipeline of transformations. Click each stage to learn more.

📝 Tokenizer
→
📊 Embedding
→
⚡ Transformer ×N
→
🎯 Output
The Transformer Block is where the work happens — and it is a fixed sequence of operations, not magic. Each block has two parts: Self-Attention (understand context) and Feed-Forward Network (process information). Modern LLMs stack 32–120 of these blocks.
In Plain English

Simple: Text goes in the front and a next-word guess comes out the back, through four stations in a fixed order. The one marked ×N is the same station repeated — not a different one each time.

Technical: Tokenizer → embedding lookup → a stack of identical transformer blocks → an output projection to vocabulary logits. Each block is architecturally the same as every other — self-attention followed by a feed-forward network, with residual connections and normalisation — but each carries its own learned weights.

How Deep Is the Stack?

The “×N” in the diagram above is doing a lot of work. Drag the depth and send a token through.

🏗️
Text in, next token out
Watch one token fall through the stack
6
Block counts are real orders of magnitude — small open models run about 12–32 blocks, large ones over 100. Every block does the same two things; nothing new is introduced with depth.
In Plain English

Simple: Depth is not more kinds of thinking; it is the same operation applied again to its own output. Each pass refines a representation that is already partly built, the way successive edits sharpen a draft.

Technical: Every block reads the residual stream and writes back into it, so representations accumulate rather than being replaced. Depth costs compute and latency linearly, which is why depth is traded against width and context length rather than simply maximised. The block counts here are real orders of magnitude, not one product's specification.

Q, K, V — Three Roles

How does the model know that “it” refers to “cat”? Attention lets each word look at every other word.

Scaled Dot-Product Attention
Every word plays three roles at once, and the formula below is just those three roles combined:
Q — query: the question this word is asking. “I am a pronoun; who is my noun?”
K — key: what this word advertises it can answer. “I am a singular animal noun.”
V — value: the content this word hands over once it is chosen.
Read it left to right: multiply every query by every key (QKT) to score how well each word answers each question; divide by √dk to keep those scores in a workable range as vectors get longer; softmax softens them into percentages that add to 100%; then mix the values in exactly those proportions. That mixture is the word’s new, context-aware meaning.
In Plain English

Simple: Think of a room where everyone is holding up a card saying what they can answer, and everyone else is holding up a question. You read every card, decide who is worth listening to, and take a bit from each — more from the useful ones.

Technical: Q, K and V are three separate learned linear projections of the same input vector. Compatibility is the dot product of queries against keys, scaled by √dk to keep the pre-softmax scores in a range where gradients survive; softmax normalises them into weights, and the output is that weighted sum of value vectors.

Attention, Step by Step

The formula in one go is a lot. Here it is as the four things it actually does, one at a time.

🔍
What is this word looking at?
Choose a query, then step: keys → scores → softmax → weighted sum
output vector for “it”
Weights are illustrative, hand-authored to show the shape of the computation — they are not read out of a trained model. The four steps and their order are exactly those of scaled dot-product attention (Vaswani et al., Attention Is All You Need, 2017).
In Plain English

Simple: Written as one formula it looks like a single leap. It is four small moves: ask, score everyone, turn the scores into shares that add to a whole, then blend.

Technical: The four steps and their order are exactly those of scaled dot-product attention (Vaswani et al., 2017, cited in full in the caption above), and only the third is non-linear — softmax is what turns arbitrary real-valued scores into a normalised, comparable set of weights. The weights in the widget are illustrative and hand-authored, not read out of a trained model.

Attention Heatmap

Each row is a word asking its question; each cell shows how much of the answer came from the word in that column. Different heads learn different jobs.

🔍
Attention Heatmap
Hover cells to see which words attend to each other

Go deeperInteractive: the whole model end to end — tokenizer → Q/K/V → 12 heads → MLP → probabilities. AI-generated, after Fig. 1 of Cho et al., “Transformer Explainer” (CHI '26) ↗

In Plain English

Simple: The model does not do this once. It does it several times side by side with different priorities — one pass chasing who a pronoun refers to, another watching for the word right next door — and then pools what they all found.

Technical: Multi-head attention runs several attention operations in parallel over lower-dimensional projections and concatenates the results. Head specialisation like coreference or syntactic dependency is an observed tendency in probing studies, not something specified in the architecture or trained for; the labels on the buttons above are an illustration of that finding.

Why Attention Wins

Transformers replaced RNNs and LSTMs because attention can process all tokens in parallel.

Feature
RNN
LSTM
Transformer
Parallelism
❌ Sequential
❌ Sequential
✅ Fully parallel
An RNN hands a hidden state from each step to the next, so step t cannot begin before step t−1. An LSTM inherits that: its gates buy memory, not parallelism. The bottleneckA 1,000-token sequence costs an RNN or an LSTM 1,000 sequential steps. Self-attention lets every token attend to every other at once, so reading that sequence takes a Transformer 1 step. Writing is a different story: generation is still auto-regressive, one token at a time (Vaswani et al., 2017, §3), which is why a long answer streams instead of appearing at once.
Long-range
❌ Forgets
⚠️ Better
✅ Full context
An RNN’s hidden state has to squeeze all earlier context into one fixed-size vector, so it fades with distance. LSTM gates — forget, input, output — learn what to keep. Attention drops the compression step: every token stays directly reachable. Useful spanA vanilla RNN holds roughly 50 tokens — by token 500, token 1 is largely gone. An LSTM stretches that to about 200–500 tokens. Neither is a context window in the modern sense.
Training Speed
Slow
Slow
~7× less training cost
The speed-up is the parallelism, cashed in: with the sequence dimension free, a GPU works on all positions at once. Vaswani et al. put the base model at roughly a seventh of the training cost of the best previous system, reaching 27.3 BLEU in 12 hours on 8 GPUs — and they are explicit that the cost column is an estimate, not a stopwatch (2017, Table 2 and §6.1). Two more 2017 ideas came with it — multi-head attention (previous slide) and positional encodings, which inject word order without recurrence.
Era
2013–2017
2014–2017
2017–today
RNNs led NLP from 2013; LSTMs owned machine translation, speech recognition and text generation from 2014 — Google Translate ran on them until it moved to Transformers. Every large language model in wide use today (GPT, Claude, Gemini, Llama) is a Transformer.
“Attention Is All You Need” (2017) — the paper that started it all. Google researchers showed that attention alone, without recurrence, could achieve state-of-the-art results on translation tasks.
In Plain English

Simple: The older designs read a sentence the way you listen to a voicemail — strictly in order, one word at a time, no skipping. Attention reads it the way you read a page: everything is on the table at once, and word 3 can look straight at word 300.

Technical: Recurrent models impose a sequential dependency: step t cannot be computed before step t−1, so the sequence dimension cannot be parallelised across training, and the path between distant tokens grows with the distance between them. Self-attention makes that path length constant and the sequence dimension fully parallel. The cost is quadratic in sequence length — the trade that buys the speed-up.

Hallucination — When AI Makes Things Up

LLMs can generate text that sounds confident and fluent but is completely wrong. This is called hallucination.

🎭
Confident Fabrication
Invents facts, citations, or events that never happened — and the most common of the three. Where the training signal was weak, the model fills the gap with a plausible pattern instead of saying “I don’t know”.
Real exampleAsk about “the 1987 Treaty of Lisbon” and you can get a fluent summary of a landmark agreement. There was no such treaty: the Treaty of Lisbon was signed in 2007.
🔀
Source Confusion
Mixes up details from different sources or contexts, producing a chimeric answer that blends unrelated material. Most damaging in research and academic work, where attribution is the point.
Example“What did Einstein say about quantum computing?” — his real remarks on quantum mechanics get restaged as commentary on modern quantum computing, which he never discussed.
💥
Logical Errors
Chains of reasoning that look valid but contain hidden flaws — usually in arithmetic or multi-step deduction. The hardest kind to catch: every individual step reads as correct, and only the chain fails.
Example“All roses are flowers; some flowers fade quickly; therefore all roses fade quickly.” That is the undistributed middle fallacy, delivered in the same authoritative tone as a sound argument.
Why does this happen? LLMs predict plausible text, not true text. They have no internal fact-checker — they optimize for what sounds right based on patterns in training data.
In Plain English

Simple: There is no moment where the model checks whether what it is about to say is true. Fluent and correct feel identical from the inside — which is why a fabricated citation reads exactly as smoothly as a real one.

Technical: The training objective rewards likely continuations, and there is no separate truth signal to optimise against. Nothing in the forward pass consults a knowledge base or a verifier, so confidence in the output text is a property of the token distribution, not evidence about the world.

Why Long Chats Start Drifting

Past a point, more text in the window does the opposite of helping: the early details blur, and the model fills the blur with plausible inventions. Drag the slider.

🧠
How much of the chat is it still holding?
Illustrative — canned responses at each fill level, not a live model
In the window: 0K / 100K tokens
Perfect recall
empty 100K
clearfuzzyinvented
Earlier: My name is Sara and I am a civil engineer.
Earlier: I am working on the Cairo metro extension.
Earlier: My deadline is March.
Later, you ask: “remind me what I told you”
This is hallucination, not forgetting. A model rarely answers “I lost that”. It answers with something that fits the shape of the question — which is far harder to catch.
Keep the window tidy rather than full. Pasting everything you have is not free: it dilutes the part that mattered. Related failure mode: “lost in the middle” — Module 5.
In Plain English

Simple: The dangerous part is not that detail goes missing. It is that the gap gets filled. You get a confident, specific, plausible answer in the shape you asked for — and nothing marks the part that was invented.

Technical: Attention spreads over a longer sequence, and retrieval of any one earlier fact degrades as the window fills. The model still emits a high-probability completion, so the failure surfaces as substitution rather than as an abstention. This is hallucination, not forgetting — and the related “lost in the middle” position effect is covered in Module 5.

Fighting Hallucination

Several strategies can reduce (but not eliminate) hallucination.

Strategy 1
Grounding
Provide source documents in the prompt so the model answers from given context, not from its training data. Simple and effective, with no infrastructure — but you need the right documents up front, and the context window caps how much you can paste. How it works“Based on the following document, answer the question: [document] — Question: what was the revenue in Q3?” The model has nothing else to answer from.
Strategy 2
RAG (Retrieval-Augmented Generation)
Automatically search a knowledge base and inject relevant docs into the prompt — grounding, without you doing the fetching. It is the industry standard for reducing hallucination in production: a search engine’s retrieval bolted to a language model’s generation. The pipelineUser asks “what is our refund policy?” → the system searches company docs by embedding similarity → the top 3 matches are injected into the prompt → the model answers from those.
Strategy 3
Constitutional AI & RLHF
Train models to refuse when uncertain rather than making things up. Constitutional AI (Anthropic) trains against a written set of principles, including honesty about uncertainty; RLHF uses human feedback to reward accurate answers. Both act during training, not at inference — they shape behaviour at the edge of what the model knows. A principle, roughlyIf you are not sure about something, say so. Do not make information up. “I don’t know” beats a confident wrong answer.
Strategy 4
Citations & Verification
Ask the model to cite sources, then verify them. Use tool-use — web search, database queries — to fact-check. Strongest with agentic systems, which can run the check on their own output before you ever see it. The workflowAnswer with citations → use a search tool to confirm each one exists → cross-reference what it actually says against the claim → flag anything that does not check out.
In Plain English

Simple: You cannot make it stop inventing, so change what it has to invent from. Put the real document in front of it and the plausible answer and the true answer become the same answer.

Technical: These strategies attack the problem at different layers — input (grounding, retrieval), training (alignment to refuse under uncertainty), and output (citation plus external verification). None removes hallucination; each narrows the space in which a fabrication is the most probable continuation. Verification is the only one that catches a failure after it has happened.

Temperature — Controlling Creativity

Temperature controls how “creative” the model is. Drag the slider to see how probabilities change.

Balanced
0.7
nice
cold
hot
wild
banana
quantum
☀️
Sample output — “What is the capital of France?”
Softmax with Temperature
Rule of thumb: For code and factual tasks, use T=0. For creative writing, try T=0.8–1.0.
In Plain English

Simple: One dial, and it does not add or remove ideas — it decides how much of an underdog the model is willing to back. Turn it down and the favourite wins every time; turn it up and the long shots start getting picked.

Technical: Temperature divides the logits before the softmax. Below 1 it sharpens the distribution towards the mode; above 1 it flattens it, raising the mass on the tail this module showed you two slides in. It reshapes the existing distribution and never introduces a candidate that was not already scored. The outputs shown per band here are canned illustrations, not live generations.

Sampling Strategies

Temperature isn’t the only knob. Here are the main ways to control LLM output.

Strategy
How It Works
Use Case
Greedy
Always pick the top token
Factual Q&A, code
Top-k
Sample from top k tokens
Balanced creativity
Top-p (nucleus)
Sample from smallest set summing to p
Dynamic creativity
Temperature
Reshape the full distribution
Fine control
In practice: Most APIs combine these. Claude uses temperature + top-p together.
In Plain English

Simple: Two different questions, easily confused. Temperature asks how lopsided the odds should be. Top-k and top-p ask how many candidates are even allowed on the ballot before you draw.

Technical: Temperature rescales the distribution; top-k and top-p truncate it. Top-k keeps a fixed number of candidates, top-p keeps the smallest set whose cumulative probability reaches p — so top-p adapts its cut-off to how confident the model is at that step, and top-k does not. They compose, which is why most APIs expose several at once.

Knowledge Check

Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.

Question 1 of 0
Score 0/0

Key Takeaways

🔮
Next-Token Prediction
LLMs predict what comes next. Intelligence emerges from scale.
✂️
Tokens, Not Words
Text is split into subword tokens via BPE. This affects cost and capability.
📊
Embeddings Capture Meaning
Words become vectors where distance = semantic similarity.
⚡
Attention Is Everything
Transformers use attention to understand context in parallel.
🎭
Hallucination Is Real
LLMs sound confident even when wrong. Always verify critical claims.
🌡️
Temperature = Creativity Dial
Low = precise and deterministic. High = creative and varied.
TokenizationBPEEmbeddingsAttentionTransformerHallucinationTemperatureSampling