Drift & Prompt-Injection Incidents
Production AI faces two failure classes traditional software does not: quiet quality drift, where behaviour degrades on its own, and active attacks like prompt injection. This lesson is about spotting both and having a plan.
Quality Drifts Quietly
Nothing crashed — the answers just got worse. Drift is slow degradation, so you only catch it if you are watching (Module 9, Lesson 1).
- Inputs shift — real user data moves away from what you tested on
- A change moves behaviour — a new model or an edited prompt shifts outputs
- Degradation is gradual — no single moment tells you it broke
Watch the Trend
Drift is invisible on any single request. The signal is in the trend — so keep scoring live traffic the way you scored your tests.
When It Is Attacked
The other class is active and adversarial — someone is trying to make your system misbehave (ties to Module 7, Lesson 6).
- Prompt injection — malicious input hijacks the model's instructions
- Data leaks — the system reveals information it should not
- Abuse — the tool is pushed to do things it was never meant to
- The rule — treat untrusted input as untrusted, always
Have a Response Plan
Incidents are not an emergency if you rehearsed the steps. The same discipline works for both failure classes.
- Detect — your monitoring surfaced it
- Contain — limit the blast radius fast
- Roll back — return to a known-good version (Module 2's version control)
- Add an eval — so it cannot recur silently (Module 8)
Build It
How to implement: name one drift signal and one attack your tool could face, and write down the one action you would take for each.
- Weekly AI Tasks tracker — a rising parse-failure rate is drift; a message crafted to trigger a bad action is injection. Validate input, and never execute message content as commands.
- Personal brand site — drift is stale or incorrect facts creeping in; the guardrail is re-verifying claims against their sources.
What you learned
Production AI has two failure classes: quiet drift, caught by watching eval scores trend on live traffic, and active attacks like prompt injection, met by treating untrusted input as untrusted. For both, have a plan: detect, contain, roll back, add an eval.