Regression & CI/CD
Every change you make can quietly break something that already worked. Regression testing plus CI/CD is how you keep shipping changes without holding your breath each time.
A Suite That Guards You
The mechanism is simple, and you already met it in Module 2.
- Keep a suite of tests and evals that says what “working” means.
- Run it automatically on every change, before it goes out.
- Block the change if the suite fails — that is CI/CD doing its job.
Why AI Regression Is Different
AI outputs vary — the same input can give slightly different results. So one pass/fail on one example is too blunt a signal.
Think In Rates
This is the same statistical mindset behind the evals in Module 8.
- One example passing — or failing — means little on its own.
- Measure the rate over a whole set of examples.
- Watch for a meaningful drop in that rate, not the ordinary noise between runs.
A Prompt Change Is A Code Change
Editing a prompt or swapping a model can change behaviour just as much as editing code — so it goes through the same eval gate.
Build It
How to implement: wire your Module 8 eval to run automatically before you deploy, and refuse to ship if the score drops.
- Weekly AI Tasks tracker — run the parse‑accuracy eval in CI so a prompt tweak can’t silently make parsing worse.
- Personal brand site — run the “every claim has a source” check on every build before publishing.
What you learned
Keep a suite that defines “working,” run it automatically, and block changes that fail. For AI, judge rates over many examples, not single cases — and treat a prompt or model change as a code change through the same gate.