Prompt Engineering

The art and science of communicating with AI — from simple questions to powerful agentic workflows.

The Prompt Is the Program

With LLMs, your prompt is your code. The quality of the output depends entirely on the quality of the input.

🧬
Prompt Anatomy
Every prompt has structure: system context, examples, instructions, and constraints.
🎯
Clarity Over Cleverness
Explicit instructions beat clever phrasing. Say exactly what you want.
🔄
Iterate, Don’t Guess
Great prompts are refined through testing, not written perfectly the first time.
🛠️
Beyond Chat
The real power comes from combining prompts with tools, files, and agentic workflows.
In Plain English

Simple: Ordering “the usual” works with your regular barista and nowhere else. Anywhere else you have to say the whole thing out loud: size, milk, hot, to take away.

Technical: The prompt is the only mutable input at inference time — the weights are frozen, so behaviour is conditioned entirely by the tokens you supply. That is what makes the prompt text the program, and everything downstream a function of it.

Interactive Prompt Builder

Toggle sections to build a structured prompt. Watch how each part shapes the output.

You are a helpful coding assistant. Always explain your reasoning step by step. Use code examples when helpful.
The user is working on a Python Flask API. The current error is: "TypeError: 'NoneType' object is not subscriptable" on line 42 of app.py.
Q: How do I handle a KeyError?
A: Use dict.get(key, default) or wrap in try/except. Example: value = data.get("name", "Unknown")
Fix the TypeError in my Flask app. Explain what caused it and how to prevent it.
Respond in this format:
1. Root cause (1 sentence)
2. Fix (code block)
3. Prevention tip (1 sentence)
In Plain English

Simple: A form, not a sentence. Five boxes: who you are talking to, what they need to know, one example of what good looks like, the actual ask, and the shape you want it back in.

Technical: The five toggles map to distinct roles in the assembled prompt — system message (persona and standing rules), pasted or retrieved context, few-shot exemplars, the user turn, and output-format constraints. They all become one token sequence in the end, but keeping them separate is what lets you change one without disturbing the others.

Six Boxes Worth Ticking

Before you send a prompt, run it past a short checklist — the same way you would brief a colleague who has never seen the task. Six boxes, one letter each.COSTAR framework

★
COSTAR Builder
Tick a box to add that ingredient. Watch the prompt fill out.
2 of 6 ticked
Not a magic word. COSTAR is a memory aid, not a required syntax — the model has no special handling for these labels. Its value is that it stops you forgetting the two people usually forget: audience and response format.
In Plain English

Simple: The pre-flight checklist a pilot reads out loud after ten thousand flights. Not because it is difficult — because the item you skip is always the one you assumed.

Technical: COSTAR (Context, Objective, Style, Tone, Audience, Response) is a mnemonic, not a parsed format: no tokeniser and no model gives these labels special handling. Its value is recall coverage — it forces the two habitually under-specified dimensions, audience and response schema, into the prompt.

The Same Job, Explained Three Ways

You already do this without thinking. You describe your job differently to a new colleague, to your manager, and to the person who has to plug into your system. Say which of those the model is writing for, and it picks the right depth, vocabulary and length by itself.audience specification

🧑‍💻
A new team member
Needs orientation, not depth. Wants to know where things live.
Explain how the login feature works. Audience: someone who joined this week. Focus on: the step-by-step flow, and which files to open first. Assume: no knowledge of our codebase.
💼
Senior leadership
Needs the decision, the risk and the date. Not the mechanism.
Summarise the state of the login work. Audience: senior leadership. Focus on: progress, blockers, and whether the date still holds. Format: 5 bullets max, no implementation detail.
🔧
Engineers integrating
Needs the contract: inputs, outputs, and what happens when it fails.
Document the login endpoint. Audience: engineers on another team calling it. Focus on: parameters, response shape, error codes, rate limits. Format: reference table.
Notice what changed and what did not. The underlying subject is identical in all three — one login feature. What moved is the depth, the vocabulary, and the shape of the answer. If you leave the audience unstated the model has to pick one, and it will usually pick the middle: too much detail for leadership, too little for whoever has to integrate.
In Plain English

Simple: You describe your job differently to a new colleague, to your manager, and to the team plugging into your system. Same job, three descriptions — and you have never once had to think about it.

Technical: Naming the audience conditions three things at once: lexical register, depth of exposition, and output length and shape. Leave it unstated and the model effectively averages over audiences and lands on the mode of its training distribution — the generic middle, which is wrong for both ends.

The Context Window

Every model has a finite memory. The context window is everything it can see at once — input + output combined.

📐
Context Window Calculator
See how different use cases fill the window
System
Context/Docs
User
Output
Free
4–8K
Typical, 2023
128–200K
Typical, 2024
1M+
Frontier, 2025
Read the trend, not the numbers. These are the orders of magnitude in general circulation at each point — roughly a few pages, then a whole codebase, then entire repositories. Individual products move faster than any slide can, so check the limit for whatever you are actually using. What does not change is the shape of the problem: the window is finite, input and output share it, and a bigger window is not the same as a better answer if what you put in it is noise.
In Plain English

Simple: A desk, not a filing cabinet. Everything you are working from has to be laid out on it at once — and whatever comes back has to fit on the same desk.

Technical: The context window is a hard token limit covering prompt and completion together, imposed by the position encoding and by attention’s quadratic cost in sequence length. The figures shown are orders of magnitude, as the slide says: read the trend, not the numbers.

Two Different Things You Are Telling It

Think about asking someone to cook dinner. Some of what you tell them is the situation: guests are coming, one of them is vegetarian, your kitchen has two pans and no oven. The rest is the rules: on the table by 7pm, under £50, nobody eats nuts. Both matter, and they do different jobs.context vs. constraints

A word used twice. On the previous slide “context” meant the token budget — how much text fits. Here it means the background you supply. Same word, two meanings; almost everyone writing about prompting uses both, so it is worth holding them apart.
Dimension
📦 Context — the situation
🚫 Constraints — the rules
The occasion
🍽 “I have guests coming for dinner.”Who it is for. Changes what counts as a good answer.
⏰ “It has to be ready by 7pm.”A limit the answer must not break.
The people
🥗 “One of them is vegetarian.”A fact about the world the answer has to survive.
🥜 “No nuts — allergy.”A hard exclusion. One violation makes the whole answer useless.
The resources
🏠 “I only have a basic kitchen.”What is actually available to work with.
💰 “Keep it under £50.”A budget you can check the answer against.
Why the split is worth making. Context makes the answer relevant; constraints make it usable. In a prompt, context is the paragraph before the ask; constraints are usually a short bulleted list after it. What happens with only one of them →
In Plain English

Simple: A taxi driver needs two different things from you. Where you are going and what the traffic is like — and separately, “no motorways, I get carsick”. Give only the second and they do not know the destination; give only the first and they take the road you cannot stand.

Technical: Context is conditioning information: it changes what counts as a relevant answer. Constraints are a filter on the output space, and because they are filters they are checkable after the fact — which is what makes them automatable. A violated constraint is a defect you can detect; missing context just yields a plausible answer about the wrong thing.

Iterating on Prompts

Nobody writes the perfect prompt on the first try. Iteration is the key skill.

Attempt 1
“Fix my code”
Too vague. Model doesn’t know what code, what error, or what “fix” means.
Attempt 2
“Fix the TypeError in app.py line 42”
Better. Specific error and location. But model still can’t see the actual code.
Attempt 3
“Here’s app.py [code]. Fix the TypeError and explain the root cause.”
Great. Includes context, specific request, and asks for explanation.
The Agentic Way
“Fix the failing test” + Claude reads the file, runs the test, reads the error, fixes it, runs again
Best. The AI uses tools to gather context itself, act, and verify. No manual copy-paste.
In Plain English

Simple: Nobody gives good directions on the first attempt either. “It’s near the shops” only becomes “third left after the roundabout, blue door” after the other person rings to say they are lost.

Technical: Each attempt removes ambiguity about the task: attempt 1 underspecifies the referent, attempt 2 fixes the location but not the content, attempt 3 supplies the artefact itself. The fourth removes you from the loop — the model acquires the same context by tool call, so context gathering stops being manual copy-paste.

Build It Up, One Turn at a Time

You do not have to describe the whole thing up front. Ask for the smallest useful version, check it, then ask for the next piece — the way you would brief a helper standing next to you.incremental prompting

🧱
Four Turns to a Log Analyser
Run the turns in order. Each one keeps everything the last one built.
Ready — turn 1 of 4
Why it beats one giant prompt: every turn is small enough to check, so a mistake surfaces immediately instead of being buried in a hundred lines you never read. The conversation itself becomes the specification.
In Plain English

Simple: You would not hand a builder the plans for the whole house and walk away. You watch the first wall go up, check it, then ask for the next one.

Technical: Incremental prompting keeps each turn’s output small enough to verify, and every previous turn stays in the context window, so the accumulated conversation is the running specification. Error-detection cost is paid per turn rather than deferred to the end, by which point a fault in step 1 has already propagated through everything after it.

The Same Request, Five Times Better

Watch one ordinary request grow. Each stage adds one kind of information, and the answer gets more usable every time.progressive refinement

🔀
Refinement Stepper
“Write me a recipe” → a prompt you could hand to anyone
Pick a stage
You rarely need all five. For a throwaway question stage 2 is plenty. Push to stage 5 only when the prompt is something you will reuse, hand to someone else, or run in a pipeline.
In Plain English

Simple: The difference between “get milk” and a written list with the brand, the size and which shop. Both work if you are the one going; only the second works if you are sending somebody else.

Technical: Each stage adds one class of information: conditioning (context), output filters (constraints), a structural skeleton (template), and a demonstration (exemplar). The marginal value falls off sharply — for a single throwaway query the first two dominate, and all five only repay the effort on a prompt that will be reused, handed over, or run in a pipeline.

Let the Model Fix Your Prompt

If you cannot tell what is missing from your request, ask. Hand the model your own half-formed prompt and ask it what it still needs to know.meta-prompting

Your prompt, as first written
How do I fix my sleep problem?
Too vague to answer well — and you may not know why. Whatever comes back will be generic, because the model has to guess at everything you left out.
Ask about the prompt instead
This prompt is too vague: "How do I fix my sleep problem?" Ask me 5 clarifying questions that would make it answerable. Do not answer yet.
It asks back: 1. Trouble falling asleep, or staying asleep? · 2. How long has this been going on? · 3. What is your current schedule — bedtime, wake time, naps? · 4. Any medication, caffeine or health conditions in play? · 5. What is the bedroom like — noise, light, temperature?
Answer those five and you have a specific, answerable prompt — built out of what the model told you it was missing. You spent one turn buying information instead of guessing at it.
Three phrasings that work: “Ask me 5 clarifying questions before answering.” · “What is missing from this prompt?” · “Rewrite this prompt so it is unambiguous, then wait.” The last word matters — without it you get the rewrite and a guessed answer.
In Plain English

Simple: Rather than guessing what the doctor needs to know, you let them ask. Five questions later they have the picture — and you never had to work out which five mattered.

Technical: You are inverting the direction of specification: instead of supplying context you have to guess at, you ask the model to enumerate the missing variables and then you fill them in. The load-bearing clause is “do not answer yet” — without it you get the questions and a guessed answer underneath them.

System Prompts — Setting the Stage

A system prompt defines the model’s persona, rules, and constraints before any user interaction.

🎭
Persona
“You are a senior Python developer who prioritizes clean, readable code.”
📋
Rules
“Always respond in JSON format. Never include comments in code.”
📎
Context
“You have access to bash, file read/write, and web search tools.”
🛡️
Safety
“If uncertain, ask for clarification rather than guessing.”
In Plain English

Simple: The briefing you give a new hire on their first morning, before anybody has asked them anything. It applies to every conversation they will ever have here, so you say it once.

Technical: The system message is a separately-roled turn placed ahead of the conversation, and providers train models to weight it above user turns when the two conflict. It persists across every turn in the session, which is why standing rules belong there rather than being restated in each message.

Few-Shot Prompting & XML Tags

Show the model what you want with examples. Use XML tags to structure complex prompts.

<examples> <example> <input>Convert 5 miles to km</input> <output>8.05 km</output> </example> <example> <input>Convert 100 kg to lbs</input> <output>220.46 lbs</output> </example> </examples> <!-- Now the model knows the exact format --> <input>Convert 72°F to Celsius</input>
XML tags give the model clear boundaries. Instead of saying “here’s some context,” wrap it in <context>...</context>. The model knows exactly where sections start and end.
In Plain English

Simple: Two filled-in forms teach the format faster than a page describing it. And putting each thing in its own labelled envelope means nobody has to guess where one thing ends and the next begins.

Technical: Few-shot exemplars condition the output distribution on a demonstrated input→output mapping with no weight update — this is in-context learning (Brown et al., 2020). Delimiters such as XML tags supply boundaries the model can see, which lets it tell instruction from data and reduces the chance that pasted content is read as a command.

Ask for a Shape, Not a Paragraph

Ask ten people to summarise the same article and you get ten differently-organised paragraphs. Hand them a form with labelled boxes instead and all ten come back the same. A schema is that form: you write out the exact fields you want, and the answer arrives already sorted into them.structured output / JSON schema

The ask — fields, spelled out
Summarise this article and return JSON in exactly this shape: { "title": <string>, "author": <string | null>, "year": <integer>, "summary": <string, max 40 words> } // Use null if a field is not stated. // Return the JSON and nothing else.
Why it helps
✓ The same shape every time. Field names and order stop drifting between runs, so a hundred articles come back comparable.
✓ A program can read it. No prose to parse — feed it straight into a spreadsheet, a database or the next step.
✓ Fewer invented fields. A named, closed list of keys leaves less room to helpfully add something you never asked for.
✓ You can check it automatically. Did every key arrive? Is year a number? That is a five-line validator, not a human reading output.
The rule that matters most. Say what to do when a field is missing (null, not a guess) — otherwise a blank turns into an invention. Three more ways to tighten it →
In Plain English

Simple: A parcel label, not a note to the postman. Fixed boxes for name, street and postcode, filled in the same order every time — so a machine can sort it without reading a word of English.

Technical: A declared schema constrains the output to a fixed key set with typed values, which makes the response machine-parseable and programmatically validatable. The critical clause is the null policy: absent an explicit instruction for missing fields, a blank is filled by generation rather than left empty.

Show, Don’t Tell

Try writing down every rule for tidying up a list of names. Then try showing four examples. The second one is shorter and handles cases you never thought of.few-shot vs. rule-based

📝
Name Formatting, Two Ways
Same task. Run both sides, then test them on an input neither covered.
Ready to compare
Approach A — list the rules
Approach B — show examples
One example replaces a paragraph of instructions; two establish a pattern; three or four pin down the awkward cases. Rules describe the boundary you thought of. Examples let the model infer the boundary you did not.
In Plain English

Simple: Teaching somebody to fold a shirt. You can write a page of instructions, or fold two in front of them. The second is shorter and it covers the shirt with the odd collar — the one you would never have thought to mention.

Technical: A rule list encodes only the boundary you explicitly enumerated. Exemplars let the model interpolate a boundary from the demonstrated mapping, which generalises to cases you did not enumerate. One exemplar sets the format, two establish a pattern, three or four pin the awkward cases; returns fall off quickly after that.

When the Examples Run Out

Sooner or later an input arrives that none of your examples covered. You get to choose in advance what happens next.corner-case handling

🧭
Guess, or Ask?
Input: “prince” — a single name. None of the examples had one.
Ready
Neither is the right answer. Guessing is faster and fine when a wrong result is cheap to spot and redo. Asking is slower and right when a wrong result is expensive — a database write, a customer email, a config change.
In Plain English

Simple: A locksmith arrives at a door your key will not open. Do you want them to force it, or to telephone you first? The right answer depends entirely on what is behind the door.

Technical: This is a policy decision about behaviour on out-of-distribution inputs, and it has to be stated in the prompt because the default is to generate something. Choose guessing when the error is cheap to detect and reverse; choose escalation when the action has side effects — a database write, an outbound email, a config change.

Give It Permission to Say “I Don’t Know”

Ask a helpful person a question they cannot answer and, if they think their job is to be helpful, they will have a go. Tell them up front that “I don’t know” is an acceptable answer and you get an honest one instead. The same sentence works on a model — and it is one clause you can paste into any prompt.abstention clause

Without the clause
What was the closing share price of a mid-cap listed company on 15 January 2025?
You might get: “$187.50” — stated flatly, with no hedge, no source and no date. It reads exactly like the answers that are right, which is the whole problem: nothing in the wording tells you which kind you got.
With the clause
What was the closing share price of a mid-cap listed company on 15 January 2025? If you do not know, say so. Do not estimate.
You get: “I don’t have real-time or historical market data, so I can’t give you that figure.” Less satisfying, far more useful — you now know to go and look it up.
When to reach for it
Anything factual and checkable: dates, prices, version numbers, specifications, quotations, who-said-what, and anything that happened recently.
Phrasings that work
“If you are not sure, say so.” · “Do not guess — mark anything uncertain.” · “Cite the source for each figure, or leave it blank.”
What it is not
Not a guarantee. It lowers the rate of confident invention; it does not remove it.
In Plain English

Simple: A confident wrong answer and a confident right answer look exactly the same on the page. That is the whole danger — you cannot tell which one you got by reading it.

Technical: The model has no calibrated abstention behaviour by default: a low-confidence continuation is emitted with exactly the fluency of a high-confidence one. An explicit abstention clause shifts the output distribution towards refusal on unanswerable inputs. As the slide says, it lowers the rate of confident invention; it does not eliminate it.

Where Prompt-Craft Stops Paying

Everything so far has been about the words you send. That takes you a long way, and then it stops — because three of the most common failures have nothing to do with how you phrased the request.

📅
It cannot know
Training data has a cutoff, and your files were never in it. No wording fixes a fact the model never saw.
🎭
It cannot check
With no source to consult, a gap gets filled with something plausible. An abstention clause lowers the rate; it cannot look anything up.
🔒
It cannot act
It can tell you what to do; running it is still your job. You end up as the wiring between the answer and the work.
So this lesson does the two things that are still yours to do. First, the techniques that squeeze the most out of a prompt on its own: making the reasoning explicit, giving the output a shape, and defining what “done” means precisely enough to hand the checking over. Second, it sets up the real answer to the three failures above — giving the model tools, so it can read, search, run and verify instead of remembering. That is Module 4: Agentic AI & Claude Code, and it is the next module, not this one.
In Plain English

Simple: However clearly you word the question, you cannot ask somebody about a letter they never received. Better wording does not create information that was never in the room.

Technical: Three failures sit outside the prompt’s reach: a training cut-off and private data the model never saw (no wording retrieves an absent fact), no retrieval path with which to verify a claim, and no execution capability. All three are answered by giving the model tools — which is Module 4, not better phrasing.

Make It Show Its Working

Remember being told to show your working in maths class? The point was never the marks — writing the middle steps down is what catches the mistake. Ask the model to do the same: add the sentence “think step by step.”chain-of-thought

Straight to the answer
A shop has 23 apples. It gives away 8, then buys 15 more. How many apples now?
You might get: “28 apples.” Wrong — and delivered with the same confidence as a right answer. The two operations got run together, and nothing on the page shows you where.
Working shown
A shop has 23 apples. It gives away 8, then buys 15 more. How many apples now? Think step by step.
You get: “Start with 23. Gives away 8: 23 − 8 = 15. Buys 15 more: 15 + 15 = 30.” Right — and if it had not been right, you could point at the exact line that broke.
When it earns its keep
Arithmetic, multi-step logic, debugging, anything where step 2 depends on step 1. Not needed for “summarise this” — there are no intermediate steps to get wrong.
The phrasing
“Think step by step” · “Work through this one step at a time” · “Show your reasoning, then give the answer” — the last one when you want the answer easy to find at the end.
Where it came from, and its limit
Wei et al., 2022 for the technique; the zero-shot form — just appending “let’s think step by step” — is Kojima et al., 2022. The limit: the steps are output, not a proof. A chain can be fluent and still wrong.
In Plain English

Simple: Doing a sum in your head against doing it on paper. On paper you get the same answer — but if it is wrong you can point at the line where it went wrong instead of starting over.

Technical: Chain-of-thought prompting elicits intermediate tokens, which give the later tokens something to condition on and improve accuracy on multi-step arithmetic and logic (Wei et al., 2022; the zero-shot form is Kojima et al., 2022). The chain is generated output, not a proof — as the slide warns, it can be fluent and still wrong.

Learn the Technique, Then Cook the Dish

Before asking about your problem, ask about the kind of problem. The general answer becomes the frame the specific answer is built on.step-back, then feed-forward

Turn 1 — step back
What are the top 3 strategies for debugging memory leaks in C++?
No mention of your code yet. You are asking for the principles, so the model commits to a framework in writing.
Turn 2 — feed forward
Now apply those three strategies to this specific heap dump from my service.
The analysis is now anchored to a stated method instead of whatever the model reaches for first.
Everyday version: read the technique for sautéing before attempting the dish. Engineering version: read the API contract before writing the call. Why it helps: the general answer is now in the context window, so the specific answer is reasoned from it and stays consistent across several follow-ups — a documented benefit of step-back prompting (Zheng et al., 2023).
In Plain English

Simple: Before asking a mechanic what is wrong with your car, ask what usually causes that noise. Now their answer about your car has to line up with what they just told you.

Technical: The first turn places a stated framework into the context window; the second turn’s answer is then conditioned on it and stays consistent across follow-ups instead of being reconstructed ad hoc each time. Documented as step-back prompting (Zheng et al., 2023).

Don’t Build the Castle All at Once

Split one big ask into ordered stages and check the output of each before the next one starts. A bad stage 1 quietly poisons stages 2, 3 and 4.task decomposition & validation gates

🧩
A Four-Stage Pipeline with Gates
Each stage has a check that must pass before the next begins.
01
Find the errors
gate: 3 or more found?
02
Sort by severity
gate: order defensible?
03
Find the root cause
gate: cause explains all?
04
Recommend fixes
gate: fix maps to cause?
Ready — 4 stages, 4 gates
The gate is the point, not the split. Four stages with no checking between them is just one long prompt with paragraph breaks. The check is what stops an error compounding — the same reason a build pipeline refuses to deploy a failing stage.
In Plain English

Simple: An assembly line only helps if somebody checks each station. Without the checks it is one long job with pauses in it, and a fault at the first station gets built into everything after.

Technical: Decomposition on its own does not improve reliability — the validation gate does. Each stage’s output is checked against a stated condition before it enters the next stage’s context, which bounds error propagation instead of letting a bad stage 1 silently condition stages 2 through 4.

Write the First Words Yourself

You can start the model’s answer for it. Because it only ever continues text, whatever you put at the start of its turn becomes a commitment it has to carry on from.prefill / assistant priming

No prefill
user: Extract the fields as JSON. assistant: Sure! Here’s the JSON you asked for: ```json { "name": "..." } ``` Let me know if you need any changes.
You now have to strip a greeting, a code fence and a sign-off before you can parse anything.
With prefill
user: Extract the fields as JSON. assistant: {"name":     ← you wrote this much "Ada Lovelace", "role": "mathematician"}
The reply is already mid-object, so there is nowhere to put a preamble. Output is parseable from the first character.
Three things prefill is good for: forcing a machine-readable shape ({, [, <answer>); locking a persona or language for the whole reply; and skipping the “Certainly! I’d be happy to…” opener. Two caveats: not every interface exposes an assistant turn you can write into (chat UIs generally do not; APIs generally do — check the one you are using), and a prefill that fights the request produces worse answers, not better ones.
In Plain English

Simple: Handing somebody a form with the first line already filled in. They carry on from where you stopped — they are not going to cross it out and start again with “Dear Sir”.

Technical: Prefill writes the opening tokens of the assistant turn, and because generation is autoregressive continuation those tokens become an unrevisable commitment. Opening with { or [ forces a parseable shape from the first character. As the slide notes, not every interface exposes a writable assistant turn.

Ask for the Plan Before the Work

You would not let a builder start knocking through walls before showing you a drawing. Same instinct here: for anything that touches more than one file, get the approach in writing first, read it, and only then say go. Nothing is changed while you are reading.plan mode

How to turn it on
/plan // or press Shift-Tab to toggle it
Not using Claude Code? The plain-language version works anywhere: “Before you change anything, tell me how you would do this and which files you would touch. Wait for my go-ahead.” The last sentence is the load-bearing one.
1️⃣
Review
Read the plan. You are checking whether it understood the problem — wrong plans are usually wrong about the goal, not the syntax.
2️⃣
Refine
Push back on the approach, not the code: “do not touch the migration”, “reuse the existing helper”, “split step 3 in two”.
3️⃣
Approve
Then let it run the whole thing. You have front-loaded the supervision instead of spending it one file at a time.
Worth it when a wrong turn is expensive: multi-file changes, refactors, anything architectural, anything touching data. Not worth it for a one-line fix — reading a plan for a typo costs more than the typo. Note how this relates to the two neighbouring techniques: meta-prompting (in the Iterating on Prompts lesson) gets the model to ask you questions; plan mode gets it to state its answer for you to check. Both buy the same thing — a cheap disagreement before an expensive one.
In Plain English

Simple: Reading a quote before agreeing to the work. It costs you five minutes, and it is the cheapest possible moment to discover that the two of you were talking about different jobs.

Technical: Plan mode separates proposal from execution: the model states its intended approach and the file set it would touch, with no writes performed, so the review happens before any state changes. The supervision cost is front-loaded once instead of paid per edit, and the disagreements it surfaces are about the goal rather than the syntax.

Say What “Done” Means

If you can write down how you would check the work, you can hand the checking over too — and stop being asked for approval every thirty seconds.success criteria & delegation

🎯
Measurable Done-Conditions
An antenna-array task with three numeric gates. Watch it iterate to green.
Calculate optimal beam angles for this antenna array. Success criteria: – Coverage ≥ 95% across the target area – SINR ≥ 10 dB at the 90th percentile – Interference zones ≤ 3% of cells Iterate until every criterion is met. Show me only the final result, with the validation numbers.
Ready
Three moves, in order: state criteria you could measure with a script, not adjectives (“good coverage” is not a gate); tell it to iterate rather than to answer once; then say how much of the working you want to see. Drop any one and you are back to supervising every step. This is the same pattern that makes an autonomous agent possible — Module 4 builds on it.
In Plain English

Simple: “Make it look nice” buys you a conversation every ten minutes. “Under two pages, no red, my name at the top” buys you the thing itself.

Technical: Delegation requires a checkable termination condition. Three moves, in order: criteria expressible as a measurement rather than an adjective, an instruction to iterate rather than to answer once, and a stated verbosity for the working. Drop any one and the loop comes back to you for adjudication — and this is the same structure an autonomous agent runs on.

Prompt Engineering Best Practices

🎯
Be Specific
“Summarize in 3 bullet points” beats “summarize this” every time.
📝
Show, Don’t Tell
Few-shot examples are worth a thousand words of instructions.
🔗
Chain of Thought
“Think step by step” dramatically improves reasoning on complex tasks.
🛠️
Use Tools First
Don’t ask the LLM to guess — give it access to read, search, and verify.
📏
Set Constraints
Define format, length, style. Constraints focus the output.
✅
Verify, Don’t Trust
Always validate outputs. Use tests, linters, and fact-checking tools.
In Plain English

Simple: Every one of these six is the same instinct wearing a different coat: say exactly what you want, show it rather than describe it, and check what comes back instead of trusting it.

Technical: The six reduce to two levers. One constrains the output distribution — specificity, exemplars, format constraints, elicited reasoning. The other closes the loop outside the model — tools for retrieval, tests and validators for verification. No amount of the first is a substitute for the second.

Knowledge Check

Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.

Question 1 of 0
Score 0/0

Key Takeaways

🧬
Structure Matters
System prompts, XML tags, and few-shot examples create reliable, repeatable outputs.
📐
Context Is Precious
The context window is finite. Be strategic about what you include.
🔗
Make the Working Visible
Asking for intermediate steps does not guarantee a right answer — it makes a wrong one findable (Wei et al., 2022).
🧩
Split, Then Gate
Ordered stages only help if something is checked between them. The gate is the point, not the split.
📝
Show, Don’t Tell
Four examples beat a page of rules — and cover cases you never wrote down.
🎯
Define “Done”
Measurable criteria plus “iterate until met” is what makes delegation possible.
System PromptsFew-ShotXML TagsCOSTARPrefillMeta-PromptingStep-BackTask DecompositionSuccess CriteriaContext WindowChain-of-ThoughtPlan ModeStructured OutputContext vs. ConstraintsAudience“I Don’t Know”