Power User

Advanced techniques for getting the most out of AI agents — skills, hooks, safety, and orchestration.

The Building Blocks

Eight pieces. Each solves one problem, and the whole module is about telling them apart.

Piece
What it is
Reach for it when
Tools
Built-in actions: Bash, Read, Edit, Grep
Always — these are the hands
MCP Servers
Plugins that add new tools
You need an API, database or browser
/ Commands
Shortcut that invokes a skill (/commit)
You run the same thing often
Skills
Instructions + tools bundled in a folder
A domain workflow or coding standard
Subagents
Fresh agents with their own context window
The task produces a lot of noise
CLAUDE.md
Project-wide context, always loaded
Stack, architecture, house rules
Hooks
Your code, on a lifecycle event
Validation, logging, notifications
Custom Agents
A named agent with a fixed role
A standing reviewer or tester
CLAUDE.md is who you are, skills are what you can do, hooks are what happens around you, and MCP is what you can reach.
In Plain English

Simple: A workshop has hand tools, a socket for plugging in new machines, a house rulebook, written recipes, and apprentices you can send on errands. Nobody confuses a recipe with a socket — and that is the whole slide.

Technical: The eight pieces differ along two axes: whether they add capability (tools, MCP) or instruction (CLAUDE.md, skills, hooks), and whether they execute inside your context window or in one of their own (subagents, custom agents).

How They Fit Together

A skill can use any combination of tools, MCP servers, and subagents. It’s the orchestration layer.

/ Command
→
Skill
→
Tools + MCP + Subagents
🔧
Tools (Built-in)
Always available. Bash, Read, Edit, Write, Glob, Grep, WebSearch. The foundation.
🔌
MCP (Plugins)
Add new tools from external services. Like npm packages for AI capabilities.
📋
Skills (Recipes)
Bundle instructions + tools into reusable workflows. Invoked via / commands or auto-triggers.
🤝
Subagents (Workers)
Fresh context windows for subtasks. Prevents bloat. Returns only the summary you need.
In Plain English

Simple: The recipe is in charge. It can pick up any tool, phone any supplier, and hand chopping to an assistant — but the recipe is what decides the order.

Technical: Skills sit at the orchestration layer. A slash command is the entry point; the skill body composes built-in tools, MCP-provided tools and delegated subagents into one sequence, so the composition is authored once rather than re-prompted each time.

See the Difference in Action

Same goal — different approaches. Watch how context changes with each method.

Goal: Check if tests pass
You: "Run the auth tests" ▹ Bash: npm test -- --run tests/auth.test.js ✓ PASS (3 tests, 0.8s)
Context impact: +50 tokens (command + output). Minimal. Direct action.
Goal: Get latest Next.js docs
You: "How does the new App Router work?" ▹ MCP context7: query-docs("Next.js 15 app router") ✓ 3 doc sections returned (1,200 tokens)
Context impact: +1,200 tokens. Medium. Docs injected into context. Worth it — avoids hallucination.
Goal: Commit all changes with a good message
You: /commit ▹ Skill loads: commit prompt template ▹ Bash: git diff (reads changes) ▹ Bash: git log --oneline -5 (reads style) ▹ Bash: git commit -m "fix: validate JWT..." ✓ Committed with auto-generated message
Context impact: +800 tokens. The skill orchestrates 3 tool calls. One command, multiple actions.
Goal: Understand a large codebase
You: "Explore the auth module architecture" ▹ Agent(Explore): spawned with fresh context · Reads 23 files in src/auth/ · Reads 8 test files · Analyzes imports and dependencies ✓ Returns: 15-line architecture summary
Context impact: +200 tokens (just the summary). The subagent used 15K tokens in its own window, but your context only gets the concise result. This is the key benefit.
In Plain English

Simple: Four ways to answer the same question, and they cost different amounts of desk space. Running one command leaves a sticky note; asking a helper to read the whole filing cabinet leaves you one paragraph.

Technical: The four tabs differ by an order of magnitude in tokens added to the caller’s window — roughly 50 for a direct tool call, 1,200 for injected MCP documents, 800 for a skill that chains three calls, and 200 for a subagent whose own 15K of reading never crosses back.

What Fills Your Window

In Plain English

Simple: Your desk is a fixed size, and it is not empty when you sit down — the reference books, the house style guide and the list of who to phone are already on it. Only the rest of the surface is yours.

Technical: The commands themselves are Lab 1’s material; what matters here is the bill. The system prompt, every connected server’s tool definitions, your skills and your CLAUDE.md are all charged before your first word, and on every turn thereafter, not once at startup. That is why a session with a dozen servers connected feels slower and costs more than a bare one doing identical work — and why turning one off is a real optimisation rather than tidying.

Every message, file read and tool result piles up in one space that does not grow. Five commands manage that space. Reading about them changes nothing — so type /compact, press Run, and watch the Conversation segment collapse.context window

Claude Code — context
Type a context command, or click a hint to fill it in…
Context: 45% used (90K / 200K tokens) 🤖 Sonnet tier
Context window90,000 / 200,000 tokens (45%)
Which one, when. /compact when the task continues but the history has gone stale. /clear when the next task has nothing to do with the last. /context before anything expensive. /init once, first, in every new project.
Time-sensitive: 200,000 tokens is the window this simulator models, and tier names change. The mechanic — a fixed budget you can inspect and reclaim — is the part that does not.

Skills — Reusable AI Workflows

Skills are prompt packages that encapsulate complex workflows into simple commands.

🧬
Skill Anatomy
A skill defines a prompt template, available tools, and trigger conditions.
📦
Built-in Skills
/commit, /review-pr, /simplify, /loop — ready to use out of the box.
🏗️
Creating Custom Skills
Define your own slash commands for team-specific workflows.
⚡
Auto-Triggers
Skills can activate automatically based on context — like importing a library.
In Plain English

Simple: Instead of explaining your house style from scratch every time, you write it down once and point at it. The note does not do the work — it tells the worker how you want the work done.

Technical: A skill is a versioned prompt package: instructions, tool permissions and trigger conditions in one folder. It is invoked either explicitly by a slash command or implicitly when its trigger matches the task.

Building an MCP Server

An MCP server is a program that exposes tools and resources via the Model Context Protocol.

// Minimal MCP server in TypeScript import { McpServer } from "@modelcontextprotocol/sdk"; const server = new McpServer({ name: "my-tools" }); server.tool("get_weather", { city: { type: "string" } }, async ({ city }) => { const data = await fetch(`api.weather.com/${city}`); return { text: await data.text() }; } ); server.run();
That’s it. This server exposes a get_weather tool that any MCP client (including Claude Code) can call. The protocol handles discovery, serialization, and error handling.
In Plain English

Simple: You are fitting a new plug on an appliance you already own. The appliance already works; the plug is what lets anything in the house switch it on without knowing how it was built.

Technical: An MCP server declares a named tool with a typed parameter schema and a handler. The protocol layer supplies discovery, JSON serialisation and error propagation, so the same server is usable by any conforming client with no client-side code.

The MCP Ecosystem

You mostly do not write MCP servers — you install them. Five are maintained by Anthropic and partners; the rest come from the community.

GitHub
Official · issues, PRs, code search
Slack
Official · read/send messages
PostgreSQL
Official · safe queries
Playwright
Official · browser automation
Filesystem
Official · file operations
Linear
Community · issue tracking
Notion
Community · documents
Jira
Community · issue tracking
Google Drive
Community · document access
AWS
Community · cloud infrastructure
Your deploy pipeline
Internal · you write it
Your internal API
Internal · you write it
Search npmjs.com for mcp-server packages. The highest-value ones are usually the last row: wrapping your own deploy pipeline, monitoring dashboard or ticket system gives the agent the same tools your team already uses every day.
In Plain English

Simple: Almost nobody builds their own kettle. You buy one that fits the socket. The interesting exception is the thing only your household owns — nobody sells that, so you wire it yourself.

Technical: The ecosystem splits three ways: vendor-maintained servers for common SaaS surfaces, community servers of varying quality, and internal servers wrapping systems with no public equivalent. The third category is where MCP earns its keep, because it exposes proprietary workflows through the same interface.

What an MCP Server Costs Before You Ask It Anything

In Plain English

Simple: Every appliance you plug in comes with a manual, and the manual sits on your desk whether or not you ever use the appliance. Plug in enough of them and the manuals leave no room to work.

Technical: Each connected server injects its full tool-schema catalogue into the system prompt on turn one — roughly roughly 150–200 tokens per tool, so a 42-tool server lands near 8,000 and a 2-tool server near 1,500. That is a standing cost, paid before any query, and it is retained through auto-compaction because it describes what the next turn can still call.

Connecting a server is not free. Each one hands the model a catalogue of everything it can do — roughly about 150–200 tokens for every tool a server exposes — and that is spent on turn one, before your first question. Then results land on top. Click read page repeatedly and watch what happens when the bar fills.tool schemas

Claude Code — MCP
Three servers are connected. Ask them for something…
3 MCP servers connected · 28% of context used 🤖 Sonnet tier
Context window56,000 / 200,000 tokens (28%)
The amber block is the standing charge. It is there before you type. Three servers with 42, 42 and 2 tools come to about 18,000 tokens of schema; this simulator models a heavier 40,000-token loadout so the effect is visible.
Try to break it. Around 180,000 tokens the agent compacts on its own: it summarises the conversation and keeps the tool schemas, because those describe what the servers will still offer next turn. Seven read page calls gets you there.

Subagents — Divide and Conquer

Spawn specialized agents for subtasks. Each gets its own context window.

Pattern
How It Works
Best For
Explore
Searches codebase, returns summary
Understanding large repos
Plan
Designs implementation strategy
Architecture decisions
Worktree
Works in isolated git copy
Parallel feature branches
Background
Runs while you continue working
Long-running tasks
Key insight: Subagents protect the main context window. The Explore agent might read 50 files, but only returns a concise summary. Your main context stays clean.
In Plain English

Simple: You send someone to the library instead of hauling the library to your desk. They read the shelf; you get the note.

Technical: Delegation is a context-isolation mechanism, not a speed one. The subagent runs in a separate window, so its intermediate reads are never charged against the caller’s budget — only the returned summary is.

How Sub-Agents Work

A sub-agent is an isolated assistant with its own context window that works independently and returns a summary.

Main Agent
→
Spawn Sub-Agent
→
Isolated Work
→
Summary Back
📦
Clean Context
Sub-agents have their own context window. Large search results stay there — only the summary returns to you.
⚡
Parallel Execution
Multiple sub-agents run simultaneously on independent tasks. 3 reviews at once instead of sequentially.
🎯
Focused Scope
Each sub-agent gets a clear, narrow task. Focused work = better results than one agent doing everything.
🔧
Tool Control
Restrict which tools sub-agents can use. Read-only agents for research, full access for implementation.
In Plain English

Simple: The helper goes away, does the job, comes back and tells you how it went. They cannot tap you on the shoulder halfway through to ask a question — so the brief has to be complete when you hand it over.

Technical: The lifecycle is spawn → isolated execution → single return. There is no mid-task turn back to the user, which is why the four properties on this slide — fresh context, parallelism, narrow scope and a restricted tool allow-list — all have to be fixed at spawn time.

Creating Custom Sub-Agents

Use the /agents command or --agents flag to build specialized sub-agents.

# Define custom agents in settings or CLI "agents": { "reviewer": { "description": "Reviews code for bugs and security", "prompt": "You are a senior code reviewer...", "allowedTools": ["Read", "Glob", "Grep"] }, "tester": { "description": "Writes and runs test suites", "prompt": "You are a test engineer...", "allowedTools": ["Read", "Write", "Bash"] }, "docs": { "description": "Generates documentation", "prompt": "You are a technical writer..." } }
In Plain English

Simple: A job description, written down. Who they are, what they are for, and which keys they are allowed to hold.

Technical: A custom agent is a named triple: a routing description, a system prompt that fixes the role, and an allowedTools allow-list. The reviewer here has no write tool at all — the restriction is enforced by configuration, not by asking the model nicely.

Designing Effective Sub-Agents

Patterns that make sub-agents reliable and predictable.

Pattern
What It Does
Example
Structured Output
Define exact return format
“Return JSON with fields: issues, suggestions, score”
Obstacle Reporting
Report what it couldn’t do
“If blocked, explain what you need”
Tool Restriction
Limit available tools
Read-only for researchers, no Bash for docs
Scope Boundaries
Define what NOT to do
“Only review auth files, skip tests”
Time Boxing
Set effort level
--effort low for quick scans
In Plain English

Simple: Tell them what “done” looks like, what to leave alone, and to speak up if they get stuck. Every one of these five rules is something you would say to a new colleague on their first morning.

Technical: All five patterns exist because a subagent cannot ask a clarifying question mid-run. A declared return schema makes the output parseable, an obstacle-reporting clause turns silent failure into a reported one, and scope and tool limits shrink the space in which it can go wrong unsupervised.

When to Use (and When Not To)

✅ Use sub-agents for
Large searches across big codebases
Parallel reviews of multiple files
Research tasks that produce verbose output
Independent work that doesn’t need your context
❌ Don’t use for
Simple queries — overhead isn’t worth it
Dependent tasks — that need main context
Quick edits — faster to do directly
Conversational work — sub-agents can’t ask you questions mid-task
Rule of thumb: If the task would consume >20% of your context window with intermediate results, delegate it to a sub-agent.
You do not have to sit and watch. --bg starts a session as a background agent and hands you your terminal back; claude agents is how you check on them afterwards. And when a long session reaches a fork in the road, --fork-session branches from this point instead of overwriting it, so you can try the risky version without losing the safe one — the same reason --from-pr exists for picking a review back up where it stopped.
In Plain English

Simple: Delegating has a cost too — the briefing, the wait, the handover. For a two-minute job you just do it yourself. But if it is a long job, you hand it over and go and do something else, rather than standing there watching.

Technical: The decision is a ratio, not a preference: delegate when the intermediate output a task generates is large relative to the summary it yields. Tasks that need the caller’s accumulated context, or that require a mid-run clarification, cannot be delegated at any size. Backgrounding is the orthogonal axis — it changes whether you have to wait, not whether the work is delegable, and branching a session is what makes an irreversible-looking experiment cheap to abandon.

Parallel Sub-Agents

The real power: run multiple sub-agents simultaneously on independent tasks.

Main Agent
↓
↓
↓
🔍 Reviewer
auth module
🧪 Tester
write tests
📖 Docs
update API docs
↓
↓
↓
Results Merged
3x faster: Three sub-agents working in parallel complete in the time of one. Each has its own context window, so there’s no interference.
In Plain English

Simple: Three people painting three different rooms finish in the time it takes to paint one room — as long as none of them needs the ladder the others are standing on.

Technical: Wall-clock time approaches that of the slowest branch rather than the sum, but only for genuinely independent subtasks. Shared state or an ordering dependency between branches serialises them again, and each branch still spends its own tokens even though the caller never sees them.

The Same Diagram, Happening

In Plain English

Simple: Watch the bar that measures your desk. Three people are working flat out and your desk barely fills up. That is the whole trick.

Technical: The three progress bars track work done inside three separate windows; the context bar underneath tracks only the caller’s. Roughly 17,400 tokens are spent across the branches and about 200 cross back, so the caller’s occupancy is decoupled from the total work performed.

The diagram on the previous slide is the map. This is the journey. Three helpers go off with their own notebooks, read far more than you ever see, and hand back a paragraph each. Watch the three bars fill — then read the last line, which is the whole reason to do this.sub-agent delegation

Claude Code — sub-agents
Describe something too big for one pass…
Task tool ready · your window: 17K / 200K 🤖 Sonnet tier
Explore agent
Waiting
Plan agent
Waiting
Bash agent
Waiting
Your context window17,200 / 200,000 tokens (8.6%)
Watch what does not happen. The three agents read dozens of files between them, and this bar barely moves. Their reading lived in their own windows and died with them; only the summaries crossed back.

Lab: Have Two Helpers Read Two Papers at Once

Intermediate ~20 min
Have one helper read one paper, then two read both at once, and see how little of the reading lands in your own conversation.
1
Download the two papers
Save both into one folder: Attention Is All You Need — arxiv.org/abs/1706.03762 — and BERT — arxiv.org/abs/1810.04805. Both are free.
Expected

Two papers in one folder. These are the two the course keeps contrasting — one writes forwards, the other reads a whole sentence at once — so you already know roughly what the answers should be. That is the point: you are testing the method, not learning the papers.

2
One helper, one paper, fixed answer shape
“Use a subagent to read the first paper and return only: three claims it makes, and one question it leaves unanswered.”
Expected

A short, structured answer. Now notice what did not happen: the paper’s text never came into your own conversation — only the summary did. Asking for a fixed answer shape is the part that matters, and the next step shows you why.

3
Now both at once, and compare them
“In parallel, one subagent per paper: same three-claims-and-one-question shape for each. Then give me one table comparing them.”
Expected

Two helpers start together and finish in a different order, then one comparison table. The whole thing takes about as long as the slower single paper, not both added up. Without the fixed shape from step 2 you would get two summaries that cannot be lined up side by side — and the comparison is the entire value.

4
Check how much room you used up
/context
Expected

Barely moved — against two full research papers. That gap is the entire reason to work this way. Had you pasted both papers in yourself, you would have spent most of your available room before asking your first real question.

Lab: Parallel Subagents Over Several Repositories

Advanced ~25 min
Ask one question of several repositories at once, and get back one comparable answer per repository.
1
Clone them side by side
Clone two or three repositories from GitHub into sibling folders. Pick ones you did not write — the point is to learn something you do not already know.
Expected

Three folders next to each other. This is the only step that needs the network; everything after it is local reading.

2
Put them all in reach of one session
claude --add-dir ../repo-b --add-dir ../repo-c
Expected

The session can now read outside the folder you started in. Confirm with /status. Without this, every subagent you spawn is confined to one repository and the comparison is impossible.

3
Write the question once, with fixed fields
“For each of the three folders, one subagent answers exactly: (a) what it is for, (b) how a newcomer would start using it, (c) how actively it is being worked on, (d) one risk. Nothing else.”
Expected

One helper per folder, running together, then a table. Four short fields is a deliberate choice — ask for a free-form report instead and you get three essays that cannot be lined up against each other.

4
Spot-check one answer before you trust any
“Prove your answer to (c) for the second folder by quoting the file and the line you got it from.”
Expected

A citation you can open and read for yourself. A batch of answers is exactly as trustworthy as its least-checked row, and checking one row per run is the cheapest honesty habit there is — far cheaper than finding the mistake after you have acted on it.

5
Find where it stops paying off
Re-run it with more directories than is comfortable, and watch where the answers get vaguer.
Expected

At some number the returns arrive slower and the merged table gets vaguer. Parallelism is bounded by the merge, not by the number of workers: whatever coordinates them still has to hold every return in one window.

The Hook Pipeline

Hooks intercept the agent at key moments. Click each stage to explore.

📨
PreTool
→
⚙️
Tool Runs
→
📬
PostTool
→
🔔
Stop
PreToolUse: Runs before a tool executes. Use it to validate, log, or block tool calls. Example: Block all Write calls to .env files — prevents the agent from accidentally overwriting secrets.
In Plain English

Simple: A door with a guard on each side. One checks what you are carrying in, the other checks what you are carrying out, and neither of them is the agent — they are your rules, and they cannot be talked round.

Technical: Hooks are your own programs fired on lifecycle events. Because a PreToolUse hook can veto the call before it executes, enforcement lives outside the model’s reasoning — a prompt instruction can be overridden by a persuasive context, a non-zero exit status cannot.

Hook Examples

✅
Auto-Lint
PostToolUse on Edit and Write:
npx eslint --fix $FILE && npx prettier --write $FILE
Failures go back to the agent, so it fixes them in the same loop pass.
📝
Audit Log
One line per tool call:
[14:23:01] Bash: npm test
[14:23:05] Read: src/auth/middleware.js
[14:23:06] Edit: src/auth/middleware.js:22
🛡️
Safety Guard
PreToolUse. Blocks rm -rf and git push --force, writes to .env or credentials, and DROP TABLE / DELETE FROM. The agent is told why it was blocked.
🔔
Slack Notify
Stop hook — it fires when the agent decides it is finished (M4 covers the full event set; Notification is the status-change event, not the end-of-task one). Posts to #dev-notifications: “Agent completed: fixed 3 failing tests in auth module. All 47 tests pass.”
In Plain English

Simple: Four small chores nobody wants to remember: tidy up after each edit, keep a diary of what was touched, refuse the genuinely dangerous requests, and tell the team when it is finished.

Technical: The four examples map to three distinct event points. Auto-lint is PostToolUse and feeds its failures back into the same loop pass, so the agent self-corrects; the safety guard is PreToolUse and must return a reason, because a silent block leaves the agent retrying blind.

Permissions & Safety

One question, asked once: how much may it do before checking with you? These are the real values --permission-mode accepts.

plan
Plan, do not act
Read-only by construction. The one mode where an expensive, wide-ranging question risks nothing but time.
manual
Ask me
You are the gate. Slowest, and the right place to start on a repository you care about.
acceptEdits
Edits without asking
File edits stop prompting. This is the mode most people settle into — and the one that makes a way back matter.
auto · dontAsk
Fewer interruptions still
Progressively less prompting. Check the current behaviour before relying on either — this is where the set has changed most between versions.
bypassPermissions
No checks at all
Every guard off. Defensible in a throwaway sandbox with nothing valuable in reach; almost never otherwise.
The mode is a dial, not a fence. The fence is --allowedTools and --disallowedTools, which name what may run at all, and a PreToolUse hook, which refuses a specific call even when nobody is watching. And when the configuration itself is the suspect, --safe-mode starts with every customisation disabled, while --setting-sources controls which settings files load at all.
📁
Where it may act
The mode says how much; this says where. A working directory is a boundary, and directories you add are additions to it — so an agent pointed at one service in a monorepo cannot quietly rewrite its neighbour. Deny rules beat allow rules, which is the property you want when the two disagree. path scoping
🌐
What it may reach
Anything that can run a command can usually open a socket, so "which tools" and "which hosts" are separate questions. Restricting outbound reach is what stops a fetched instruction becoming an exfiltration path — the risk that grows the moment an agent reads anything it did not write. network egress
In Plain English

Simple: How much shopping may someone do on your card before ringing you? The answer is different for a corner shop and a car dealership — and separately, some shops you simply never give them the card for.

Technical: Two different mechanisms, and confusing them is the common mistake. The mode is session state, and you are the gate — it changes when you are asked. An allowlist or a hook is configuration, and a rule is the gate — it holds in an unattended run, where there is nobody to answer a prompt at all. Loosen the mode for speed; rely on the allowlist and the hook for safety. Names and behaviour move between versions, so confirm with claude --help rather than trusting a slide.

What Are Agent Skills?

Skills are folders of instructions, scripts, and resources that agents discover and use on demand.

📚
Portable Knowledge
Package domain expertise, coding standards, and workflows into reusable, version-controlled packages.
🔍
On-Demand Loading
Skills are discovered dynamically — only loaded when relevant to the current task. No context waste.
🌐
Cross-Agent Standard
Works with Claude Code, Cursor, GitHub Copilot, VS Code, Gemini CLI, and 25+ other tools.
👥
Team & Enterprise
Capture organizational knowledge in portable packages. Share across teams, enforce standards at scale.
In Plain English

Simple: The know-how that normally lives in one experienced colleague’s head, written into a folder anyone can copy. When they go on holiday, the folder stays.

Technical: A skill is a filesystem-level artefact, so it inherits everything the filesystem already gives you: version control, diffs, review and distribution. Because the format is an open specification rather than one vendor’s config, the same folder is portable across conforming agents.

Anatomy of a Skill

Every skill lives in a folder with a SKILL.md file.

# SKILL.md — Deploy Reviewer --- name: deploy-reviewer description: Reviews deployment configs for best practices version: 1.0.0 trigger: - glob: "**/deploy*.yaml" - glob: "**/Dockerfile" --- ## Instructions When reviewing deployment configurations: 1. Check for hardcoded secrets 2. Verify resource limits are set 3. Ensure health checks are configured 4. Validate environment variable usage ## Resources - See checklist.md for the full review checklist - See examples/ for reference configurations
In Plain English

Simple: A recipe card with a label at the top. The label says when to reach for this card; the rest says what to do once you have.

Technical: The YAML frontmatter is machine-read for routing — name, version and trigger globs decide whether the skill is even offered. The Markdown body below is what the model actually reads, and it is only read once the frontmatter has matched.

Creating Your First Skill

Skills can be project-local, user-global, or shared across your team.

Step 1
Create the skill folder
mkdir -p .claude/skills/my-skill
Step 2
Write SKILL.md with triggers
Define when the skill activates: file globs, keywords, or manual invocation.
Step 3
Add supporting resources
Checklists, examples, templates, scripts — anything the agent needs.
Step 4
Test and iterate
Give the agent a task that should trigger the skill. Refine the instructions.
In Plain English

Simple: Make a folder, write the note, drop in anything the note refers to, then try it and fix what did not land. Step four is the one people skip and the only one that tells you the truth.

Technical: Two of the four steps are placement decisions, not authoring ones. .claude/skills/ in the project makes the skill travel with the repository and apply to collaborators; the same folder under your home directory keeps it personal and applies it to every project you open.

Configuration & Multi-File Skills

Complex skills use multiple files for different aspects of the workflow.

# Skill folder structure .claude/skills/code-review/ ├── SKILL.md # Main instructions + triggers ├── checklist.md # Detailed review checklist ├── security.md # Security-specific rules ├── performance.md # Performance review guide ├── examples/ # Good/bad code examples │ ├── good-auth.ts │ └── bad-auth.ts └── templates/ # Output templates └── review-report.md
Key insight: SKILL.md stays concise (under 500 lines). Details live in resource files that the agent reads only when needed — the folder is big and what gets read is small.on-demand disclosure
In Plain English

Simple: A cookbook with a contents page. You do not read all 300 pages to make one dish — you read the contents, then the one page you need.

Technical: Installed size and loaded size are different numbers. Only SKILL.md is guaranteed to enter the context window; the resource files are referenced by name and fetched on demand, so a large skill folder has a small resident cost.

Loaded 3 of 6 Files — Watch It Choose

In Plain English

Simple: You did not have to name the recipe. You said what you wanted for dinner and the right card came off the shelf — and only the pages that mattered got read aloud.

Technical: Two mechanisms are running at once here. Trigger matching selects the folder from a plain-language request with no slash command; progressive disclosure then loads a subset of that folder’s files, so the green segment grows by what was read rather than by what was installed.

A skill can install a hundred pages and read four of them. Four skills are installed below, collapsed. Describe a task in plain English: the folder that matches opens, every file is marked ✓ loaded or ⊘ skipped, and the green segment grows by exactly what was read — not by what was installed.progressive disclosure

Claude Code — skills
No slash command needed. Just say what you want…
4 skills installed · 17K / 200K used 🤖 Sonnet tier
.claude/skills/
commit/
📄 SKILL.md
code-reviewer/
📄 SKILL.md
📁 agents/
📄 pr-review.md
📄 security-check.md
📁 scripts/
📄 run-tests.sh
📄 lint-check.py
📄 checklist-long.md
diagram-builder/
📄 SKILL.md
📁 templates/
📄 sequence.md
📄 flowchart.md
📄 class.md
📄 state.md
report-builder/
📄 SKILL.md
📁 tools/
📄 fetch-issue.py
📄 update-status.py
📁 scripts/
📄 save-local.py
📄 analyse-local.py
📁 templates/
📄 summary.md
📄 triage.md
Run a prompt. The folder it matches opens, and each file is marked loaded or skipped.
Skill files loaded so far: 0 tokens
Context window17,000 / 200,000 tokens (8.5%)

The Skills Ecosystem

Agent Skills are an open standard adopted by 25+ AI development tools.

Claude Code
Anthropic
Cursor
Editor
GitHub Copilot
GitHub
VS Code
Microsoft
Gemini CLI
Google
OpenAI Codex
OpenAI
Roo Code
Open Source
Junie
JetBrains
OpenHands
Open Source
Kiro
AWS
Goose
Block
Amp
Sourcegraph
Write once, use everywhere. The same skill folder works across all compatible agents — no vendor lock-in. Browse community skills at github.com/anthropics/skills.
In Plain English

Simple: Like a plug shape that every country agreed on. Your appliance works when you move house, and you are not stuck with one manufacturer because your adapters only fit theirs.

Technical: Adoption across competing vendors is what turns a config format into a standard: the skill folder becomes a portable asset rather than switching-cost. Time-sensitive — the tool list on this slide is a snapshot; treat the count as illustrative, not current.

Sharing & Distributing Skills

Skills are just folders — share them like any other code.

📦
Git Repo
Commit to your repo. Team members get them on clone.
📤
npm Package
Publish as a package for cross-project sharing.
🌍
Global Skills
~/.claude/skills/ for personal skills across all projects.
📚
Community Registry
Browse skills at github.com/anthropics/skills.
In Plain English

Simple: It is a folder. You already know four ways to give someone a folder, and all four work here.

Technical: The four routes differ in scope, not mechanism. Committing to the repository binds the skill to one project and its collaborators; publishing as a package decouples the skill’s version from the consuming project’s; the home-directory location applies to every project but travels with you rather than with the code.

Troubleshooting Skills

❌
Not Loading
Four causes, in order of likelihood: wrong path (must be .claude/skills/), filename is not exactly SKILL.md, broken YAML frontmatter delimiters, or no trigger match. Debug with claude --debug "skills".
🎯
Wrong Trigger
**/*.yaml matches every YAML file — too broad. **/deploy*.yaml is deploy configs only. src/components/**/*.tsx is React components only. --verbose shows which skills fired.
⚠️
Conflicts
Two skills claiming the same files resolve by three rules: the more specific glob wins, a project skill overrides a global one, and a priority hint in the frontmatter overrides both.
⚡
Performance
Keep SKILL.md under 500 lines, keep trigger globs narrow, one skill per domain, and push the detail out into resource files that load only when read.
In Plain English

Simple: When a skill does nothing, it is almost never the writing — it is the filing. Wrong drawer, wrong label on the front, or a label so vague that it opens for everything.

Technical: Three of the four failure modes are discovery failures, not instruction failures: path, filename and frontmatter parsing all happen before the body is ever read. The fourth is over-broad globs, where the skill loads correctly and simply should not have.

Lab: Install a Skill Somebody Else Wrote

Intermediate ~15 min
Install a front-end design skill, prove it loaded, and watch it change the output of a prompt you did not change.
1
Get a baseline first — before the skill
“Build a single-file HTML pricing page for a small product. No frameworks.”
Expected

A working page that looks like default generated HTML. Save it. This is the only evidence you will have that installing a skill did anything at all — without it, “the skill improved things” is unfalsifiable.

2
Install it
A skill is just a folder with a SKILL.md in it. Put one at ~/.claude/skills/<name>/ for yourself, or .claude/skills/<name>/ to share it with the repository. A plugin marketplace is only a delivery mechanism for the same folder.
Expected

The folder exists with a SKILL.md in it. The frontmatter needs only two fields, a name and a description — and the description is not documentation, it is the trigger the model matches your request against.

3
Prove it loaded, in a fresh session
/exit then claude
Expected

The skill appears in your available skills with its one-line description. Skills are discovered when a session starts, so one installed mid-session may not be visible until you restart — the most common reason a correctly-written skill appears to do nothing.

4
Run the exact same prompt again
Re-run step 1 word for word, then diff the two files.
Expected

Visibly different output — a real type scale, deliberate spacing, a considered palette. You did not ask for any of that, and it applied anyway. That is what a skill is: instructions that fire on relevance rather than on being invoked.

Example: Presentation Skill Structure

# .claude/skills/presentation-builder/SKILL.md --- name: presentation-builder description: Generate HTML slide presentations from templates version: 1.0.0 trigger: - keyword: "presentation" - keyword: "slides" - keyword: "slide deck" --- ## Instructions When creating presentations: 1. Read all templates in templates/ folder 2. Read style-guide.md for design rules 3. Generate a self-contained HTML file 4. Include speaker notes in <aside> tags ## Resources - templates/ — HTML slide type templates - style-guide.md — Colors, fonts, animations - examples/ — Reference presentations
The one thing to notice: nothing here is HTML. The frontmatter says when to fire, the Instructions say what to read first, and the Resources say where the actual output format lives. The skill is a routing table, not a template — which is why you can change the look of every deck by editing style-guide.md alone.
In Plain English

Simple: The recipe card does not contain the cake. It tells you which tin, which oven setting and where the icing instructions are kept — so changing the icing changes every cake without rewriting the card.

Technical: Separating routing from content is what makes the skill maintainable. Output format lives in a resource file, so one edit to style-guide.md propagates to every artefact the skill produces, and SKILL.md stays under the size where it competes for context.

Lab: Author an HTML-Presentation Skill, and Let It Interview You

Advanced ~25 min
Have a skill-creation skill interview you, then write and test a skill that turns an outline into a self-contained HTML slide deck.
1
Install a skill-creation skill
Install a skill-creator skill the same way as the previous lab — a folder with a SKILL.md. Using a skill to write a skill is also the fastest way to learn what a good SKILL.md contains.
Expected

It shows up in your skill list. Nothing is generated yet — and that is correct: the next step is where you find out what it needs to know.

2
Seed it, and demand an interview
“Use skill-creator to build a skill that generates self-contained HTML presentations from an outline. Interview me first — ask about slide layouts, theme and colours, fonts, how to handle speaker notes, where to save output, and whether to deploy — then write it.”
Expected

Questions, not a file. That is the whole point of the seed: it asks about layouts, palette, fonts and notes before writing anything, so the SKILL.md ends up describing your deck rather than a generic one. A skill written without the interview encodes the model’s defaults, not your standard.

3
Read what it actually wrote
cat .claude/skills/html-presentation/SKILL.md
Expected

Frontmatter with a name and a description, a short instructions body, and pointers to any longer reference files. Check the description hardest — it is the trigger, so a vague one means a skill that never fires no matter how good the body is.

4
Test it on input it has never seen
“Using that skill, make a six-slide deck explaining how HTTP caching works.”
Expected

One HTML file that opens in a browser and matches the choices you made during the interview. If it does not match them, the gap is in SKILL.md, not in the deck — fix the instructions and regenerate.

5
Commit it, and it stops being yours alone
git add .claude/skills && git commit -m "Add html-presentation skill"
Expected

The skill is now a versioned artefact a teammate gets from a git pull — not session state that dies with your window. This is the difference between a trick you know and a capability your team has.

6
Make a subagent you can call again tomorrow
Ask for .claude/agents/reviewer.md whose YAML frontmatter sets name, a description saying when to use it, and tools: Read, Grep, Glob.
Expected

A short file with YAML frontmatter. Every subagent in this course so far was ad hoc — described in a sentence, gone when the session ended. This one is named, reusable and committed. The tools line is the whole point: a reviewer that cannot write cannot “helpfully” fix the thing it was asked to judge. Reach for a defined agent when you want the same role repeatedly, with the same limits, without re-describing it each time. Six mechanisms all add capability and are constantly confused: CLAUDE.md is what is always true about the project, a skill is a procedure with files, a slash command is a prompt you tired of retyping, a subagent is a second context doing a scoped job, a hook is a rule that fires whether or not anyone is watching, and an MCP server reaches a system outside your repository. Facts, procedures, scale, safety, reach — confirm the exact paths with /help in your own version rather than trusting any list, including this one.

Orchestrating Agents

A coordinator agent delegates subtasks to specialists. Click an agent to see its role.

🎯 Coordinator
Plans & delegates
Running
⚙️ Backend Agent
API & database
Running
🎨 Frontend Agent
UI components
Waiting
🧪 Test Agent
Integration tests
Waiting
👀 Review Agent
Code review
Waiting
Coordinator breaks the feature request into subtasks, assigns them to specialized agents, and synthesizes the results. It decides which agents run in parallel and which depend on each other. Example: Backend and Frontend run in parallel, then Test runs after both complete.
In Plain English

Simple: Somebody has to be the site manager. Not because they lay the bricks, but because someone must decide who starts now and who waits for the walls.

Technical: The coordinator holds the dependency graph. It decomposes the request, dispatches independent branches concurrently, blocks dependent ones until their inputs land, and synthesises the returns — which is why the test agent shows Waiting while backend and frontend show Running.

One Request, Fourteen Steps

In Plain English

Simple: Every piece in this module, in one picture, with your question walking through it. Follow the dot: it goes in one end as a sentence and comes out the other as an answer.

Technical: Two loops carry the whole architecture. The inner one is model ↔ runtime — the agent loop from Module 4. The outer one is runtime ↔ the world, reaching skills for instruction, MCP servers for data, and subagents for work that should not be charged to this window.

Everything in this module in one picture, with the request moving through it. You ask for something; the agent loops between the model and the code, pulls in a skill, calls out through MCP servers to real data, and hands two slices of the work to sub-agents. Run it once, then step through it.agent architecture

Current step
Press Run to watch the request travel through the system.
Step 0 of 14
Incident report
User request
Read the comments on issue 123, then check what the team said in chat
Skills
issue-triage
domain-notes
report-in-html
Agent
LLM
Claude Code
Tools / commands / hooks
MCPs
Docs MCP
Tracker MCP
Fetch MCP
Chat MCP
Data
Tickets
Design docs
Chat history
analyse-comments
analyse-logs

You Describe It, the Agent Types It

What changed is who writes the characters. You still decide what the program should do and you still have to check that it does it — but you say it in English instead of typing it in a language. Andrej Karpathy named both halves of that shift.vibe codingSoftware 3.0

🎵
Vibe Coding
“You fully give in to the vibes, embrace exponentials, and forget that the code even exists.”
3️⃣
Software 3.0
Natural language as the new programming language. Prompts are the source code.
🔄
The Human-Agent Loop
Humans set intent and verify. Agents research, code, test, and iterate autonomously.
🔬
Agentic Research
Agents that autonomously explore codebases, read papers, and synthesize findings.
Key shift: The developer’s job evolves from writing code to directing agents — specifying intent, reviewing output, and making judgment calls. The skills you’ve learned in this course are exactly what this future demands.
In Plain English

Simple: You are now the person who says what the building should do and walks round at the end to check it. Someone else is holding the trowel. Saying what you want and checking you got it are still your job — and they are the hard parts.

Technical: Karpathy’s framing (2025) is that natural language becomes the authoring layer while generated code remains the artefact that runs. The verification burden does not move with the typing: the reviewer still needs enough of the underlying model to tell working output from plausible output.

Resources & Further Learning

Deepen your skills with these official courses and references.

🎓
Anthropic Courses
Free official courses on skills, subagents, prompt engineering, and Claude Code at anthropic.skilljar.com.
📚
Agent Skills Spec
The open standard specification, examples, and reference library at agentskills.io.
📦
Example Skills Repo
Browse real-world skills at github.com/anthropics/skills — code review, testing, deployment, and more.
📖
Claude Code Docs
Official documentation at docs.anthropic.com — tools, MCP, hooks, permissions, and API reference.
Recommended learning path: Start with Anthropic’s free courses at anthropic.skilljar.com, then build your first skill using the lab above, and browse the community skills repo for inspiration.

Knowledge Check

Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.

Question 1 of 0
Score 0/0

Key Takeaways

⚡
Skills = Reusable Power
Package complex workflows into simple slash commands your whole team can use.
🔌
MCP = Infinite Reach
Connect your agent to any service via the standard Model Context Protocol.
🪝
Hooks = Guardrails
Automate safety checks, linting, logging, and notifications at every lifecycle stage.
🤝
Orchestration = Scale
Multiple specialized agents working together can tackle projects no single agent could.
SkillsMCPSubagentsHooksPermissionsOrchestrationSafetyWorktrees