Agentic AI & Claude Code
From chatbot to autonomous agent — tools, loops, and the future of AI-assisted work.
What Is Agentic AI?
An agent is an AI that doesn’t just answer questions — it takes actions to achieve goals.
Simple: A recipe tells you how to make dinner. A cook makes dinner. Both know the same things — only one of them comes back with a plate.
Technical: The model is unchanged. What changes is the harness around it: its output is routed to tools, and the tools’ results are fed back in, so it can observe the consequences of an action and act again.
Chatbot vs. Agent
Same AI model, completely different capabilities.
Chatbot — a line
Agent — a cycle
- □read the failing check and the code it covers
- □make an edit, run the check again — still failing
- □fix the real cause, run the check again — passing
Simple: Text a friend about your broken bike and you get instructions. Hand them the bike and you get back a working bike.
Technical: Same weights, different loop. A chatbot’s context holds only your turns; an agent’s also holds tool output, so its next prediction is conditioned on the real state of your files rather than on what it remembers about projects in general.
The Loop in Action
Watch how an agent handles a real task: “Fix the failing auth test.”
npm test to see the failure, then uses grep to find the relevant source file. It reads the test expectations and the actual implementation to understand the gap.
Simple: It is how you find a leak under a sink: look, tighten something, run the tap, look again. Nobody fixes it in one move, and the tap is what tells you whether you are done.
Technical: The loop closes on an external signal — the test’s exit code — not on the model’s own confidence. That is the whole difference between iterating and repeating: the environment, not the model, decides whether Verify passed.
See It Happen
A simulated agentic workflow — watch the agent gather, act, and verify.
Simple: Like watching over someone’s shoulder while they work. You see every step they take, not just the finished result they hand you.
Technical: Every line is a tool call and the output it returned, appended to the conversation. There is no hidden scratchpad here — this transcript is what the model is reasoning over on the next turn.
Built-in Tools
Every Claude Code session comes with powerful tools out of the box. Each is named the way the agent names it in its own output.
npm test, git diff, lsof -i :3000. The agent sees the full output.**/*.test.js); Grep searches contents by regex (TODO|FIXME). This is how it navigates an unfamiliar codebase.Simple: A kitchen already has a knife, a hob and a tap before you buy a single gadget. These six are that kitchen — nothing to install, nothing to configure.
Technical: Each is a function the model may call, with a declared schema for its arguments. Read, Edit, Write, Glob and Grep are deliberately narrower than Bash: a narrow tool fails loudly on the wrong input instead of doing something plausible and unintended.
One Plug, Every Tool
A standard socket, so a tool written once works with every AI that speaks it — the way any USB device works in any USB port. Each server below is one such tool. Click one to explore it.Model Context Protocol
query-docs("Next.js 15 app router") → Returns relevant sections with code examples.
Simple: A kettle plugs into any wall socket because everyone agreed on the shape of the socket. Nobody manufactures a kettle-shaped hole in the wall.
Technical: MCP standardises the interface — how a server lists its tools, how a tool is invoked, how a result comes back. A server author targets the protocol once instead of writing an integration per AI product (Anthropic, 2024).
How MCP Works
MCP is a standard protocol that connects AI models to external services.
Simple: Ordering in a restaurant. You talk to the waiter, the waiter talks to the kitchen, and you never walk in and shout at the chef.
Technical: JSON-RPC over stdio or HTTP with a client in the middle. The client owns discovery, the permission prompt and the log, which is why no server ever gets to talk to the model directly — and why a compromised server is contained rather than catastrophic.
Bash — The Universal Tool
The Bash tool can run any shell command. This is what makes agents truly powerful.
npm test, pytest -x, cargo test. This is the Verify step; without tests the agent is flying blind.git log --oneline -10 for recent context, git diff for what changed, git blame for why. Committing is fine; rewriting history is not, without permission.npm install zod, pip install requests, cargo add serde. It checks what is already installed first, so you do not end up with two libraries doing one job.node --version. “What is holding port 3000?” → lsof -i :3000. Most useful when a deploy or a config is misbehaving.rm -rf or git push --force. You stay in control.Simple: The other tools are specialised kitchen gadgets; the shell is the whole workshop. It is powerful for exactly the reason it needs supervision.
Technical: Bash makes the tool surface unbounded — any binary on your PATH becomes callable, and the agent cannot know in advance what each one does. That is why the permission gate sits on this tool rather than on a list of forbidden commands.
Read, Edit, Write
Dedicated file tools are safer and more precise than shell commands for code changes.
Simple: You can open a jar with a hammer. A jar opener is less exciting and breaks less glass.
Technical: Edit requires a unique match and errors when it finds none or two, so an ambiguous change fails instead of silently hitting the wrong line. A sed one-liner would have edited both and reported success.
Skills — Reusable AI Workflows
A skill is a pre-built prompt package that handles a specific task. Think of them as AI plugins.
Simple: A recipe card you wrote once because you make the same dish every Sunday. You are not inventing it again each week.
Technical: A skill is a stored prompt template invoked by name. The value is not that the model could not do it unprompted — it is that the instructions are identical every time, so the output is reproducible instead of re-improvised.
CLAUDE.md — The Project Brain
A CLAUDE.md file at your project root gives the agent persistent instructions across every session.
Simple: The note taped inside a shared kitchen cupboard: bins go out Tuesday, the oven runs hot. Nobody has to be told twice.
Technical: The file is read into context at the start of every session, so project conventions arrive as instructions rather than as something the model has to infer from the code — and inference is where a plausible-but-wrong convention gets invented.
Managing Context
Everything the agent has read this session sits in one space that does not grow. The commands that reclaim it are Lab 1’s material — what matters here is watching it fill, and choosing when to act.context window
Simple: A desk you cannot make bigger. Every new document you spread out means an older one has to be filed or binned.
Technical: The context window is a fixed token budget shared by instructions, files and tool output. These four commands are the only levers over what occupies it: add deliberately (@path), summarise (/compact), seed (/init), or persist outside it (Memory).
Which Model, When
Three separate dials that beginners turn as if they were one: which model, how hard it thinks, and the keys that stop it.
claude --list-models — that is the truth, not this table. Model names change; the roles do not. And the most common, most expensive beginner habit is reaching for the strongest model by default: on a question that was never hard, it is mostly just slower.Simple: Picking a courier. Same city, same parcel — a bicycle, a van and a lorry are not ranked best to worst, they fit different jobs. Sending everything by lorry is not care, it is waste.
Technical: Model and effort are orthogonal, and coupling them is the error — high effort does not rescue a poorly-scoped task, and the strongest model does not compensate for a context you filled with noise. Effort buys reasoning depth before acting; the model sets the ceiling on what that reasoning can reach.
Plan Mode & Thinking Mode
For complex tasks, slow down and think before acting.
Shift+Tab“think hard” in promptSimple: Two different ways to slow a decision down. One is asking a builder for drawings before they knock a wall through; the other is asking them to take a proper look first.
Technical: They act on different layers. Plan Mode is a permission constraint — the write tools are simply unavailable. Thinking Mode spends more output tokens on reasoning before the answer. You can run either alone, and they cost differently for that reason.
Hooks — Lifecycle Events
Hooks run your code before or after Claude Code takes an action.
rm -rf and DROP TABLE, validate paths, or rewrite the inputs. Exit non-zero and the call never happens.osascript / notify-send.Simple: The interlock on a microwave door. It is not asking the microwave to behave — it physically cannot run with the door open.
Technical: A hook is deterministic code on a lifecycle event, outside the model’s control. A PreToolUse hook that exits non-zero cancels the call, which is a guarantee; an instruction in CLAUDE.md saying “never do this” is only a strong preference.
Your First Hook
Real-world hooks you can add to .claude/settings.json.
matcher. It is the only field that decides when your script runs. "Bash" means “only before Bash calls”; the empty string means “every time this event fires”. Everything else here is boilerplate.Simple: A note on the fridge saying check this before you cook chicken is useless unless it says chicken. Otherwise you either check everything or check nothing.
Technical: matcher is the event filter. It is matched against the tool name, so "Bash" scopes the hook to shell calls and "" fires on every occurrence of that event — including the ones your script was not written to handle.
GitHub Integration
Automate your entire GitHub workflow with Claude Code.
/review-pr 123 checks four things at once: security (injection, vulnerabilities), logic (edge cases, races), style against your conventions, and performance (N+1 queries, leaks).claude "fix issue #42" → reads the issue and its comments, searches the codebase, writes the fix plus tests, opens a PR with a summary.-p means print mode — no interactive prompt. claude -p "fix lint errors", -p "update snapshots". Pair it with --allowedTools to bound what it may do.on: pull_request auto-reviews, on: issues auto-triages and labels, on: push runs quality checks. Official action: anthropics/claude-code-action.Simple: The difference between a colleague who reads your work and one who also files the paperwork, chases the reviewer and closes the ticket.
Technical: None of this is a new capability — it is the same Bash and MCP tools pointed at a version-control host. What makes it automatable is -p (print mode): no interactive prompt, so it runs where there is no terminal to answer one.
Pick a Scenario
Select a use case below and watch the agent work through it.
Simple: Watching six different repairs before you attempt one yourself. The tasks differ; the rhythm of the work does not.
Technical: These transcripts are hardcoded, not live. What is worth reading across all six is what each loop converges on: a bug fix ends when a test passes, a refactor ends when behaviour is unchanged, a review ends when there is nothing left to say.
Build Your Own Workflows
Combine everything you’ve learned to create custom agentic workflows.
Simple: A “leaving the house” routine instead of separately remembering keys, lights and the back door. The steps were never hard; the order and the not-forgetting were.
Technical: A workflow encodes sequencing and failure handling that a single prompt does not: deploy only if tests pass, notify only if deploy succeeded. That ordering constraint is the artefact, not the individual commands.
claude — The Full CLI
Claude Code is a powerful CLI with 50+ options. Here are the essential categories.
opus, sonnet, haiku, or full ID."Bash(git:*) Edit Read"manual, plan, auto, bypassPermissionstext (default), json, or stream-json for real-time streaming.text or stream-json for programmatic input.--max-budget-usd 5.00low, medium, high, xhigh, max--bg.stable, latest, or version number.Simple: Every appliance has more buttons than anyone presses. These five tabs are the ones worth finding.
Technical: The flags sort into three jobs: what the model is told (--system-prompt, --append-system-prompt), what it is permitted to do (--allowedTools, --permission-mode), and what shape comes back (--output-format, --json-schema). Learn the three jobs, not the fifty flags.
CLI Recipes
Copy-paste these real-world patterns into your terminal.
Simple: A bucket brigade: each person does one thing and hands the bucket on. None of them needs to know where the fire is.
Technical: -p turns claude into an ordinary Unix filter — reads stdin, writes stdout, exits with a status code. That is why it composes with cat, git diff and a CI runner without any of them knowing an AI is in the pipe.
claude mcp — Server Management
Manage MCP servers directly from the command line.
--transport. Every MCP server is either remote (an HTTP URL you point at) or local (a command after -- that Claude Code starts for you). Every other flag follows from that one choice — --header only applies to the first, -e only to the second.Simple: Adding a channel to a TV. Either you tune to a broadcast someone else runs, or you plug a box in beside your own sofa.
Technical: --transport decides which, and the difference is lifetime, not just address: an HTTP server is already running and you authenticate to it with --header; a stdio server is a child process Claude Code starts and stops for you, configured with -e.
Permission Modes
Control how much autonomy Claude has with --permission-mode.
manualplanacceptEditsautobypassPermissionsclaude --permission-mode auto --allowedTools "Bash(git:*) Read Glob Grep"Simple: A dial that runs from “ask me before every turn” to “just drive.” Where you set it should depend on how far the car can travel before you notice.
Technical: The gate sits at the tool-call boundary, evaluated after the model has decided on a call and before it executes — so a mode narrows what happens, never what the model may propose. bypassPermissions removes the gate, which is only defensible when the blast radius is already bounded by a sandbox.
Claude Code — Evolving Fast
Claude Code ships updates weekly. Here’s what’s changed recently and what to watch for.
ToolSearch — only discovered when needed. Reduces context overhead and enables infinite extensibility.--agent reviewer.claude -w creates git worktrees for isolated experiments. Combine with --tmux for parallel sessions.claude update regularly and read claude --help — that output, not this slide, is the authority on what flags exist today.Simple: A printed timetable at a bus stop. Accurate the week it went up, and worth checking against the live board.
Technical: Read these three as one direction rather than three features: each buys back a scarce resource. ToolSearch defers tool descriptions to save context; --agents narrows the tool set to reduce ambiguity; -w isolates the filesystem to contain mistakes. Time-sensitive: verify against claude --help.
The Agent Only Picks Up the Tools It Needs
Give the agent a hundred tools and it does not read a hundred manuals first. It looks up the one it needs, when it needs it — like a mechanic walking to the toolbox instead of carrying it.on-demand disclosureToolSearch
ToolSearch is a tool for finding tools. Servers register theirs as deferred; the agent sees only names, then loads the full descriptions it asks for. A library catalogue, not the whole shelf.ToolSearch("slack message"). Exact: ToolSearch("select:mcp__slack__send"). Scoped: ToolSearch("+slack send"). A deferred tool must be loaded before it can be called.Simple: A library does not post every book to you on the off chance. It sends you the catalogue, and you request the one you want.
Technical: Deferred registration: the server advertises tool names and withholds the full schemas until ToolSearch asks. The saving is in input tokens on every single turn, not once at startup — which is why it compounds over a long session.
Lab 0: Install Claude Code and Prove It Is Connected
node --version && git --versionTwo version lines. Node must be 18 or newer for the npm install in the next step. If node is not found, either install Node, or install Claude Code with a standalone installer (such as Homebrew’s claude-code) and skip the npm step.
npm install -g @anthropic-ai/claude-code && claude --versionnpm prints an added N packages line, then a version number. If npm fails with a permissions error, use a Node version manager rather than sudo — a global install run as root is the most common way to break a Node setup.
cd my-project && claudeA welcome banner, a one-time sign-in on first run, then a prompt waiting for input. You are now inside a session, not at your shell — the commands in Lab 1 only work here.
/statusA summary of the session: how you are signed in, which model is active, and which directory it is working in. This is the connection proof. If it shows you as not signed in, sign in before going further — an unauthenticated session fails on the first real request, and the error rarely says “you are not logged in”.
“What does this project do? Answer in three sentences.”A short answer, preceded by lines showing which files it opened. Read it as an exam you already know the answers to. If it is wrong about a project you know well, that is the single most useful thing you will learn today — and far better learned now than on work that matters.
/exit then claude --continue/exit returns you to your shell; claude --continue reopens the same conversation with its history intact. A session is a resumable thing on disk, not a window you lose by closing it.
claude --resume (and: claude --fork-session)A list of earlier conversations to pick from. --continue reopens the most recent one; --resume lets you choose; --fork-session opens a copy, so you can try something risky without spoiling a thread you depend on. Reach for the fork when you want a second attempt at the same problem while keeping the first — without it you overwrite whichever answer turns out to be the good one.
Lab 1: The Slash Commands That Matter
/helpA list of the commands available in this session, including any your project or its plugins added. The list is not fixed — it depends on where you launched from, which is why it is worth reading once per new project.
/initIt reads the repository and proposes a CLAUDE.md at the root for you to approve. Read it before accepting. It inferred your conventions from your code, and a plausible-but-wrong convention here is re-read by every future session.
/contextA breakdown of what is occupying the context window. The lesson is what sits there before you type anything: the system prompt, the tool definitions, any connected servers, and your CLAUDE.md. That is the real reason a session with many tools connected feels slower and more expensive than a bare one.
/clear versus /compact/clear discards the conversation and starts empty. /compact replaces it with a summary and carries on. Run /context after each to see the difference in the meter. These are not two strengths of one thing: one throws away a decision you may still need, the other keeps a lossy trace of it.
/rewindA list of earlier points in the session you can return to. This is the escape hatch for “I approved that too quickly” — and knowing it exists is what makes it safe to work quickly in the first place.
Lab 2: Permission Modes and Hooks — Two Guards, Opposite Directions
claude --permission-mode planThe session starts in plan mode. The real set is acceptEdits, auto, bypassPermissions, manual, dontAsk and plan — check with claude --help rather than trusting any list, including this one, since the set changes between versions.
“Add a one-line comment at the top of README.md”It asks before writing. Read every option before choosing: one of them approves this single action, and another approves every action of that kind from now on. Beginners pick the second by reflex and then wonder why they stopped being asked.
“Plan how you would add input validation here. Do not edit anything.”A written plan and no edits. Plan mode is read-only by construction, so it is the one mode where an expensive, wide-ranging question costs you nothing but time. A plan you did not read, though, is just latency.
A hook is written into your project settings: an event, a matcher that scopes it to a tool, and a command to run. The real events are PreToolUse, PostToolUse, UserPromptSubmit, SessionStart, Stop, SubagentStop, PreCompact and Notification. Because it is a file in the repo, a teammate gets the guard from a git pull.
“Run this: rm -rf /tmp/does-not-exist”The hook fires and the command is refused. An untested guard is not a guard. This failure is silent in one direction only: a matcher that does not scope to the tool you assumed never runs, and never tells you it did not.
Lab 3: Choosing a Model, and How Hard It Thinks
claude --list-modelsThe models your account can use. Treat this as the source of truth, not a table in a course — model names change, and a deck that hardcodes one is out of date within a quarter. /model switches between them inside a session.
Two answers. Compare latency and depth, not correctness — on a question that was never hard, the strong model is mostly just slower. Reaching for the strongest model by default is the most common and most expensive beginner habit.
claude --effort low (also: medium, high, xhigh, max)Effort controls how much the model reasons before acting, independently of which model you chose. Low suits mechanical edits; high earns its keep on “find the root cause”. Use /effort to change it inside a session.
/contextDifferent models give you different amounts of context. A larger window changes what you can fit, not how well it reasons — and you pay for the context you actually occupy, so a big window filled with noise is worse than a small one filled with the right two files.
Lab 4: Interrupting and Steering
It halts and returns you to the prompt. Work already written to disk stays; the step in flight does not finish. Interrupting is normal operation, not an error — and the sooner you are comfortable doing it, the less you pay for runs you already knew were wrong.
“Stop looking in tests/. The bug is in the request parser — read that first.”It resumes from where it was, with your correction in context. A correction is strictly cheaper than a restart, because starting again throws away the reading you already paid for.
/rewindPick an earlier point and the session returns to it. This is the counterweight to the persistent-approval option in Lab 2: the faster you let it work, the more you need a way back.
The boundary is the lesson: reversibility is a property of the tool’s own edits, not of the world. That boundary is also what decides which permission grants are safe to make permanent.
Lab 5: Getting Documents In
“Summarise @README.md in five bullets.”Typing @ opens a path completer, and the file is loaded directly. Compare this with asking it to “find the readme”: that pays for a search, and leaves the search output sitting in your context for the rest of the session.
“Compare @package.json and @README.md — does the README describe scripts that exist?”Both files load and one answer covers them. This is the shape of most real document work: a question that only makes sense across two sources, where the answer is the discrepancy.
The prompt shows an attachment and the answer describes what is in the image. Worth reaching for whenever screenshotting is faster than describing, which for a visual bug is almost always.
git log --oneline -30 | claude -p "Group these commits by theme."Standard input becomes the content, and -p prints one answer and exits. Any command that produces text is now a document source — logs, a CSV, the output of a build. This is also the door that composes with the rest of your shell, and therefore the door automation uses.
claude -p "List the three largest files as JSON" --output-format json ; echo $?A JSON object instead of prose, then a 0. --output-format takes text (the default), json for a single result, or stream-json for output as it arrives. Reach for this the moment another program has to read the answer — a CI step, a script, a dashboard. Prose is for you; JSON is for the next command. And check the exit code: a script that ignores it will treat a failed run as an empty result.
Lab 6: Make Something, Then Fix It When It Breaks
“Create NOTES.md summarising this project, one section per top-level directory.”The proposed file is shown for approval, then written. Open it and read it. You are the reviewer — and this is the first artefact in this track that you own and will keep.
“Add scripts/hello.sh that prints the current git branch, and make it executable.”Two operations — writing the file and changing its mode — each asking separately if you are in a mode that asks. Then run it: it should print your branch name. A file it claims to have made executable, but did not, only shows up when you run it.
git status && git diff --statExactly the files you asked for, and nothing else. If there is a third file you did not ask for, that is the finding — and /rewind from Lab 4 is how you take it back.
command not found → the global npm bin is not on your PATH (Lab 0). Nothing happens after a prompt → check the mode; you may be in plan mode (Lab 2). It ignores your conventions → there is no CLAUDE.md; run /init (Lab 1). Slow and forgetful → run /context, then /compact (Lab 1).Every one of these is an environment or mode fault, not a model fault — and each maps back to one earlier lab in this track, which is the real test of whether the track worked.
Quick Start Guide
npm install -g @anthropic-ai/claude-codecd my-project && claudeSimple: Assembling flat-pack furniture: three unglamorous steps, then the part you actually wanted.
Technical: Steps 1, 2 and 4 are per-session. Step 3 is the only one that persists — skip it and you re-explain the project on every future run, which is both slower and the main source of conventions being guessed at.
Knowledge Check
Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.