Evaluating Your Evals
You have built evals to catch your system's mistakes. But your evals can be wrong too. A green eval that doesn't reflect reality is worse than none — it gives you false confidence that the problem is solved.
How Evals Go Bad
A broken eval fails quietly, because a passing suite looks the same whether it is protecting you or not. Three ways it happens:
- They test the wrong thing — measuring something easy to check instead of what actually matters.
- They pass on broken output — the check is too loose to notice a bad answer.
- They drift — they no longer reflect what your users actually care about.
Check the Checker
To trust an eval, compare its verdicts against your own judgment on a sample of cases. Where they disagree, the eval is the thing to fix — not just the system.
Keep Evals Evolving
Every new failure mode you find — the way you learned to look for them in Lesson 2 — is a case your eval set didn't cover yet.
- Find a new failure in production or testing.
- Add it to your eval set before you fix the underlying bug.
- Now it can't come back — the eval catches any regression for good.
A Living Asset
Your eval suite grows with the product. A stale suite quietly stops protecting you: it still turns green, but it no longer describes the system you actually ship. As the saying goes — a green check is not evidence until it has failed once.
Build It
How to implement: take one eval and deliberately feed it a known-bad output. If it still passes, your eval is broken — fix it until the bad case fails.
- Weekly AI Tasks tracker — when a new kind of message breaks parsing, add it to the eval set before you fix the parser.
- Personal brand site — when a false claim slips through, add a check that would have caught it.
What you learned
Module 8 made evaluation the engine of improvement: you build evals, then you check that the evals themselves are honest, and you keep growing them as you find new ways to fail. Evals you trust are what let you change fast without breaking what worked.