GENAI 102
M05 · L01
LLM Foundations

Tokenization & Generation

An LLM reads and writes in tokens — small chunks of text — and it produces one token at a time. Knowing just this explains a surprising amount of the model’s behavior and cost.

01 / 08
GENAI 102
M05 · L01
What Is a Token

Text Becomes Tokens

Before the model sees your text, it is split into tokens. GenAI‑101 covers the mechanics — here we care about what it costs.

  • Your text is split into tokens before the model reads it
  • A token is roughly a word-piece — often part of a word, not a whole one
  • Different text costs different token counts for the same character length
  • Non-English text often costs more tokens than the same idea in English
02 / 08
GENAI 102
M05 · L01
One Token at a Time

Generation Is Sequential

The model predicts the next token from everything written so far, appends it, and repeats. That single fact explains two things you will feel as an engineer.

Two Consequences
It can drift off course as small errors compound — and longer outputs cost more, because every token is another prediction step.
03 / 08
GENAI 102
M05 · L01
The Engineering Angle

Why This Matters

Tokens are the unit you are billed and timed in, so they turn into a design constraint.

  • Both cost and latency scale with the number of tokens
  • You pay for input tokens plus output tokens — the prompt is not free
  • Trimming prompts saves real money on anything that runs often
04 / 08
GENAI 102
M05 · L01
A Trust Boundary

Plausible, Not True

Because the model chooses the most plausible next token — not the true one — it can be confidently wrong. You do not fix that by trusting harder; you design verification around the output.

Foreshadowing
Module 8 is about exactly this — evaluating and checking what the model produces instead of taking it on faith.
05 / 08
GENAI 102
Build It
From Concept to Capstone

Build It

How to implement: paste a prompt into any tokenizer viewer and see how many tokens it is. Then shorten it — drop filler words and repetition — and watch the count drop.

  • Weekly AI Tasks tracker — the message-parse prompt runs on every message, so its token count is a recurring cost. Keep it lean.
  • Personal brand site — generation here is a one-time cost, so token count matters less than getting the content right.
06 / 08
GENAI 102
Knowledge Check

Check what stuck

Three questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

07 / 08
GENAI 102
Summary
Recap

What you learned

LLMs work in tokens and generate one at a time. That drives cost, latency, and drift — and because output is plausible rather than true, you build verification around it.

08 / 08