Tokenization & Generation
An LLM reads and writes in tokens — small chunks of text — and it produces one token at a time. Knowing just this explains a surprising amount of the model’s behavior and cost.
Text Becomes Tokens
Before the model sees your text, it is split into tokens. GenAI‑101 covers the mechanics — here we care about what it costs.
- Your text is split into tokens before the model reads it
- A token is roughly a word-piece — often part of a word, not a whole one
- Different text costs different token counts for the same character length
- Non-English text often costs more tokens than the same idea in English
Generation Is Sequential
The model predicts the next token from everything written so far, appends it, and repeats. That single fact explains two things you will feel as an engineer.
Why This Matters
Tokens are the unit you are billed and timed in, so they turn into a design constraint.
- Both cost and latency scale with the number of tokens
- You pay for input tokens plus output tokens — the prompt is not free
- Trimming prompts saves real money on anything that runs often
Plausible, Not True
Because the model chooses the most plausible next token — not the true one — it can be confidently wrong. You do not fix that by trusting harder; you design verification around the output.
Build It
How to implement: paste a prompt into any tokenizer viewer and see how many tokens it is. Then shorten it — drop filler words and repetition — and watch the count drop.
- Weekly AI Tasks tracker — the message-parse prompt runs on every message, so its token count is a recurring cost. Keep it lean.
- Personal brand site — generation here is a one-time cost, so token count matters less than getting the content right.
What you learned
LLMs work in tokens and generate one at a time. That drives cost, latency, and drift — and because output is plausible rather than true, you build verification around it.