Cost & Latency Optimization
At scale, cost and latency stop being back-office details — they become the product. A great feature that is too slow to use or too expensive to run will not survive contact with real traffic.
The Levers
Most cost and latency wins come from a short list of moves — the same ones you met in Module 5.
- Model choice — use a smaller, cheaper model wherever it is good enough
- Caching — reuse results instead of paying for them twice
- Shorter prompts — fewer tokens in means less to pay for and less to process
- Fewer & parallel calls — batch what you can, run independent calls at once
- Workflow simplification — the cheapest call is the one you never make
Bootstrap, Then Move Down
A common pattern: use a strong model to get things working and to generate examples, then move the routine, high-volume work to a smaller, cheaper model that is good enough for the job.
Cutting Latency
Latency is what the user actually feels. Three moves shrink the wait, and one habit keeps you honest.
- Stream the response — show tokens as they arrive so waiting feels shorter
- Work in parallel — overlap independent steps instead of chaining them
- Cache — a hit returns instantly and costs nothing
- Measure it — use the metrics from Lesson 1 rather than guessing where the time goes
Measure First
Find the single biggest source of cost or latency, then cut that. Optimizing everything equally — or before you have measured — wastes effort on parts that were never the problem.
Build It
How to implement: find your tool’s single most expensive or slowest model call, then write one concrete way to cut it — a smaller model, a cache, or a shorter prompt.
- Weekly AI Tasks tracker — the per-message parse is the recurring cost; a smaller model plus a cached system prompt is your main lever.
- Personal brand site — generation is one-time and cheap, and latency does not matter, so do not optimize it.
Module 9, Wrapped
Module 9 covered operating AI in production — measuring behavior, watching for drift, handling failures, and now keeping it fast and affordable. Optimize what you measured, and only what matters.