Learning to See
Once your tool has real users, you can’t tell whether it’s working by guessing. Observability is how you actually see what’s happening in production — every request, every model call, every failure.
Four Signals
Good observability records enough to reconstruct what your tool did — and why.
- Logs — what happened, as it happened
- Traces — the steps of a single request, start to finish
- Metrics — latency, cost, and error rate over time
- The actual inputs & outputs — so you can do real error analysis
You Can’t Test Your Way Out
A model’s output is unpredictable. No amount of pre-launch testing can anticipate every prompt a real user will send — so you have to watch real usage to catch the quality problems that only show up in the wild.
Feeding the Evals
Observability isn’t just alarms — it’s the raw material for improvement.
- Production traces feed the error analysis from Module 8
- Real inputs & outputs become new eval cases
- Metrics tell you whether each change actually helped
- This is how you keep improving after launch, not just at launch
Logging Real People
Observability means logging real user data. That’s a responsibility: capture only what you need, protect it, and never log secrets — API keys, passwords, or tokens have no business in your logs.
Build It
How to implement: decide the three things you’d log for every model call in your tool — the input, the output, and the latency & cost — and where they go.
- Weekly AI Tasks tracker — log each parse: the message in, the task out, and whether it was corrected, so you can see failures and feed your evals.
- Personal brand site — log build-time generations so you can review what was produced before you publish it.
What you learned
Observability is how you see production: logs, traces, metrics, and real inputs & outputs. AI needs it more, because output is unpredictable — and it closes the loop back into your evals. Log responsibly, and never log secrets.