GENAI 102
M05 · L04
Cache, Cutoff & Cost

Three Realities of Using Models

A model is not a free, all-knowing oracle. In practice it comes with three constraints you must design around: it can cache repeated work, it only knows the world up to a knowledge cutoff, and every call costs money, per token.

01 / 08
GENAI 102
M05 · L04
Reality One

The Knowledge Cutoff

A model only “knows” what was in its training data, which stops at a fixed point in time — its cutoff. Ask about anything newer and it will guess, not recall.

  • The model has no memory of events after its training cutoff.
  • For anything recent or private, you must supply the information in the prompt.
  • This is the core reason to ground the model with your own data — the subject of Module 6.
02 / 08
GENAI 102
M05 · L04
Reality Two

Reusing a Prefix

If many calls begin with the same text — a long system prompt, a fixed set of instructions — that shared start can be cached and reused, making repeated calls cheaper and faster.

Design rule
Put the stable, unchanging part of your prompt first, and the part that changes last — so the reusable prefix stays identical across calls.
03 / 08
GENAI 102
M05 · L04
Reality Three

You Pay Per Token

Most models bill by the token — you pay for what goes in and what comes out. One call is cheap; thousands of calls add up fast.

  • Longer prompts and longer answers cost more, every single call.
  • Repeated calls multiply the bill — usage, not the one-off, is what matters.
  • Your levers: caching, shorter prompts, and smaller models — the focus of Module 9.
04 / 08
GENAI 102
M05 · L04
Putting It Together

Budget From the Start

Cost is a design constraint, not an afterthought. A feature that quietly calls the model on every user action can look fine in a demo and produce a surprise bill at scale. Estimate the cost up front and treat it as a budget you design toward.

05 / 08
GENAI 102
Build It
From Concept to Capstone

Build It

How to implement: estimate one feature's monthly model cost — calls per day times tokens per call times days — then note one way to cut it: cache the stable prefix, shorten the prompt, or use a smaller model.

  • Weekly AI Tasks tracker — a stable system prompt is cacheable, so reuse it; and remember the model won't know what happened “this week” unless you tell it.
  • Personal brand site — one-time generation cost is small, but the model won't know your latest project unless you supply it in the prompt.
06 / 08
GENAI 102
Knowledge Check

Check what stuck

Three questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

07 / 08
GENAI 102
Summary
Recap

What you learned

Three practical realities of using models: caching reuses an identical prefix to save time and money, the knowledge cutoff means you must supply anything new, and you pay per token — so budget for cost from the start.

08 / 08