Why Evals Are the Key Skill
Andrew Ng argues that the single trait that most separates great AI builders from the rest is a disciplined evals loop — actually measuring whether the system works, then using what you learn to improve it. This module is about building that habit.
Why It’s Hard, Why It Matters
AI output is unpredictable — the same prompt can give a good answer today and a bad one tomorrow. So “it looked fine when I tried it” is not evidence.
- Because behaviour varies (Module 1), one lucky run proves nothing.
- You can’t improve what you don’t measure — guesses just move the problem around.
- Evals turn vibes into measurement: a repeatable score you can trust.
The Eval Loop
The loop is what turns random tinkering into systematic progress — every pass tells you exactly what to fix next.
What Makes Evals Tricky
There is no single recipe you can copy. That is precisely why it counts as a skill rather than a checkbox.
- The right approach varies by project — a chatbot and a data extractor need different measures.
- It even varies by stage — a rough early check differs from a mature test suite.
- Choosing what to measure is judgement, and it improves with practice.
Test-First, Grown Up
You already met test-first thinking in Module 4. Evals are its grown-up version: for ordinary code a test is pass or fail, but AI output is a matter of degree, so you measure quality — how good, how often — not just whether it ran.
Build It
How to implement: pick the ONE thing your tool must get right, then write down how you’d measure it on 10 real examples. That’s your first eval.
- Weekly AI Tasks tracker — measure “did the message parse into the right task?” over a set of real messages.
- Personal brand site — measure “does every generated claim trace to a source?” over every section.
What you learned
A disciplined evals loop — measure, find the biggest failure, fix, measure again — is the habit that most separates great AI builders. There’s no one-size eval, so it’s a skill worth practising on your own capstones.