Safe for the World
A prototype that works for you is not the same as an agent that is safe for everyone. Shipping means putting up guardrails — protections against both misuse and honest mistakes.
Three Real Risks
These are generic risks that apply to any agent that reads outside input or takes actions.
- Prompt injection — malicious input hijacks the agent's instructions and makes it do something you never intended.
- Data exfiltration — the agent leaks private data it was trusted with.
- Harmful actions — it does something destructive it cannot take back.
Putting Up Guardrails
Validate what goes in and what comes out, and constrain what your tools are allowed to do.
Least Privilege & Review
Assume something will go wrong, and make sure it cannot go far.
- Least privilege — give the agent only the access it actually needs, nothing more.
- Human in the loop — keep a person's approval in front of any risky action.
- Logging — record what the agent does, so you can see what happened (this ties to observability in Module 9).
Writing Down Governance
Governance is deciding, out loud, who is accountable, what data is allowed, and what the agent must never do.
Build It
How to implement: for your own tool, name one adversarial input it might face and one guardrail that stops it — plus the one action it must always ask a human about first.
- Weekly AI Tasks tracker — a message is untrusted input; never let it trigger destructive actions; validate before storing, and keep secrets out of reach.
- Personal brand site — the guardrail is factual: never publish a claim that is not grounded in real source material.
What you learned
Prompt injection, data leaks and harmful actions are real risks. Guard against them by validating inputs and outputs, giving the agent least privilege, keeping a human in the loop, and writing your governance rules down before you ship.