Three Realities of Using Models
A model is not a free, all-knowing oracle. In practice it comes with three constraints you must design around: it can cache repeated work, it only knows the world up to a knowledge cutoff, and every call costs money, per token.
The Knowledge Cutoff
A model only “knows” what was in its training data, which stops at a fixed point in time — its cutoff. Ask about anything newer and it will guess, not recall.
- The model has no memory of events after its training cutoff.
- For anything recent or private, you must supply the information in the prompt.
- This is the core reason to ground the model with your own data — the subject of Module 6.
Reusing a Prefix
If many calls begin with the same text — a long system prompt, a fixed set of instructions — that shared start can be cached and reused, making repeated calls cheaper and faster.
You Pay Per Token
Most models bill by the token — you pay for what goes in and what comes out. One call is cheap; thousands of calls add up fast.
- Longer prompts and longer answers cost more, every single call.
- Repeated calls multiply the bill — usage, not the one-off, is what matters.
- Your levers: caching, shorter prompts, and smaller models — the focus of Module 9.
Budget From the Start
Cost is a design constraint, not an afterthought. A feature that quietly calls the model on every user action can look fine in a demo and produce a surprise bill at scale. Estimate the cost up front and treat it as a budget you design toward.
Build It
How to implement: estimate one feature's monthly model cost — calls per day times tokens per call times days — then note one way to cut it: cache the stable prefix, shorten the prompt, or use a smaller model.
- Weekly AI Tasks tracker — a stable system prompt is cacheable, so reuse it; and remember the model won't know what happened “this week” unless you tell it.
- Personal brand site — one-time generation cost is small, but the model won't know your latest project unless you supply it in the prompt.
What you learned
Three practical realities of using models: caching reuses an identical prefix to save time and money, the knowledge cutoff means you must supply anything new, and you pay per token — so budget for cost from the start.