Prompt Engineering
The art and science of communicating with AI — from simple questions to powerful agentic workflows.
The Prompt Is the Program
With LLMs, your prompt is your code. The quality of the output depends entirely on the quality of the input.
Simple: Ordering “the usual” works with your regular barista and nowhere else. Anywhere else you have to say the whole thing out loud: size, milk, hot, to take away.
Technical: The prompt is the only mutable input at inference time — the weights are frozen, so behaviour is conditioned entirely by the tokens you supply. That is what makes the prompt text the program, and everything downstream a function of it.
Interactive Prompt Builder
Toggle sections to build a structured prompt. Watch how each part shapes the output.
A: Use dict.get(key, default) or wrap in try/except. Example: value = data.get("name", "Unknown")
1. Root cause (1 sentence)
2. Fix (code block)
3. Prevention tip (1 sentence)
Simple: A form, not a sentence. Five boxes: who you are talking to, what they need to know, one example of what good looks like, the actual ask, and the shape you want it back in.
Technical: The five toggles map to distinct roles in the assembled prompt — system message (persona and standing rules), pasted or retrieved context, few-shot exemplars, the user turn, and output-format constraints. They all become one token sequence in the end, but keeping them separate is what lets you change one without disturbing the others.
Six Boxes Worth Ticking
Before you send a prompt, run it past a short checklist — the same way you would brief a colleague who has never seen the task. Six boxes, one letter each.COSTAR framework
Simple: The pre-flight checklist a pilot reads out loud after ten thousand flights. Not because it is difficult — because the item you skip is always the one you assumed.
Technical: COSTAR (Context, Objective, Style, Tone, Audience, Response) is a mnemonic, not a parsed format: no tokeniser and no model gives these labels special handling. Its value is recall coverage — it forces the two habitually under-specified dimensions, audience and response schema, into the prompt.
The Same Job, Explained Three Ways
You already do this without thinking. You describe your job differently to a new colleague, to your manager, and to the person who has to plug into your system. Say which of those the model is writing for, and it picks the right depth, vocabulary and length by itself.audience specification
Simple: You describe your job differently to a new colleague, to your manager, and to the team plugging into your system. Same job, three descriptions — and you have never once had to think about it.
Technical: Naming the audience conditions three things at once: lexical register, depth of exposition, and output length and shape. Leave it unstated and the model effectively averages over audiences and lands on the mode of its training distribution — the generic middle, which is wrong for both ends.
The Context Window
Every model has a finite memory. The context window is everything it can see at once — input + output combined.
Simple: A desk, not a filing cabinet. Everything you are working from has to be laid out on it at once — and whatever comes back has to fit on the same desk.
Technical: The context window is a hard token limit covering prompt and completion together, imposed by the position encoding and by attention’s quadratic cost in sequence length. The figures shown are orders of magnitude, as the slide says: read the trend, not the numbers.
Two Different Things You Are Telling It
Think about asking someone to cook dinner. Some of what you tell them is the situation: guests are coming, one of them is vegetarian, your kitchen has two pans and no oven. The rest is the rules: on the table by 7pm, under £50, nobody eats nuts. Both matter, and they do different jobs.context vs. constraints
Simple: A taxi driver needs two different things from you. Where you are going and what the traffic is like — and separately, “no motorways, I get carsick”. Give only the second and they do not know the destination; give only the first and they take the road you cannot stand.
Technical: Context is conditioning information: it changes what counts as a relevant answer. Constraints are a filter on the output space, and because they are filters they are checkable after the fact — which is what makes them automatable. A violated constraint is a defect you can detect; missing context just yields a plausible answer about the wrong thing.
Iterating on Prompts
Nobody writes the perfect prompt on the first try. Iteration is the key skill.
Simple: Nobody gives good directions on the first attempt either. “It’s near the shops” only becomes “third left after the roundabout, blue door” after the other person rings to say they are lost.
Technical: Each attempt removes ambiguity about the task: attempt 1 underspecifies the referent, attempt 2 fixes the location but not the content, attempt 3 supplies the artefact itself. The fourth removes you from the loop — the model acquires the same context by tool call, so context gathering stops being manual copy-paste.
Build It Up, One Turn at a Time
You do not have to describe the whole thing up front. Ask for the smallest useful version, check it, then ask for the next piece — the way you would brief a helper standing next to you.incremental prompting
Simple: You would not hand a builder the plans for the whole house and walk away. You watch the first wall go up, check it, then ask for the next one.
Technical: Incremental prompting keeps each turn’s output small enough to verify, and every previous turn stays in the context window, so the accumulated conversation is the running specification. Error-detection cost is paid per turn rather than deferred to the end, by which point a fault in step 1 has already propagated through everything after it.
The Same Request, Five Times Better
Watch one ordinary request grow. Each stage adds one kind of information, and the answer gets more usable every time.progressive refinement
Simple: The difference between “get milk” and a written list with the brand, the size and which shop. Both work if you are the one going; only the second works if you are sending somebody else.
Technical: Each stage adds one class of information: conditioning (context), output filters (constraints), a structural skeleton (template), and a demonstration (exemplar). The marginal value falls off sharply — for a single throwaway query the first two dominate, and all five only repay the effort on a prompt that will be reused, handed over, or run in a pipeline.
Let the Model Fix Your Prompt
If you cannot tell what is missing from your request, ask. Hand the model your own half-formed prompt and ask it what it still needs to know.meta-prompting
Answer those five and you have a specific, answerable prompt — built out of what the model told you it was missing. You spent one turn buying information instead of guessing at it.
Simple: Rather than guessing what the doctor needs to know, you let them ask. Five questions later they have the picture — and you never had to work out which five mattered.
Technical: You are inverting the direction of specification: instead of supplying context you have to guess at, you ask the model to enumerate the missing variables and then you fill them in. The load-bearing clause is “do not answer yet” — without it you get the questions and a guessed answer underneath them.
System Prompts — Setting the Stage
A system prompt defines the model’s persona, rules, and constraints before any user interaction.
Simple: The briefing you give a new hire on their first morning, before anybody has asked them anything. It applies to every conversation they will ever have here, so you say it once.
Technical: The system message is a separately-roled turn placed ahead of the conversation, and providers train models to weight it above user turns when the two conflict. It persists across every turn in the session, which is why standing rules belong there rather than being restated in each message.
Few-Shot Prompting & XML Tags
Show the model what you want with examples. Use XML tags to structure complex prompts.
<context>...</context>. The model knows exactly where sections start and end.Simple: Two filled-in forms teach the format faster than a page describing it. And putting each thing in its own labelled envelope means nobody has to guess where one thing ends and the next begins.
Technical: Few-shot exemplars condition the output distribution on a demonstrated input→output mapping with no weight update — this is in-context learning (Brown et al., 2020). Delimiters such as XML tags supply boundaries the model can see, which lets it tell instruction from data and reduces the chance that pasted content is read as a command.
Ask for a Shape, Not a Paragraph
Ask ten people to summarise the same article and you get ten differently-organised paragraphs. Hand them a form with labelled boxes instead and all ten come back the same. A schema is that form: you write out the exact fields you want, and the answer arrives already sorted into them.structured output / JSON schema
year a number? That is a five-line validator, not a human reading output.null, not a guess) — otherwise a blank turns into an invention. Three more ways to tighten it →Simple: A parcel label, not a note to the postman. Fixed boxes for name, street and postcode, filled in the same order every time — so a machine can sort it without reading a word of English.
Technical: A declared schema constrains the output to a fixed key set with typed values, which makes the response machine-parseable and programmatically validatable. The critical clause is the null policy: absent an explicit instruction for missing fields, a blank is filled by generation rather than left empty.
Show, Don’t Tell
Try writing down every rule for tidying up a list of names. Then try showing four examples. The second one is shorter and handles cases you never thought of.few-shot vs. rule-based
Simple: Teaching somebody to fold a shirt. You can write a page of instructions, or fold two in front of them. The second is shorter and it covers the shirt with the odd collar — the one you would never have thought to mention.
Technical: A rule list encodes only the boundary you explicitly enumerated. Exemplars let the model interpolate a boundary from the demonstrated mapping, which generalises to cases you did not enumerate. One exemplar sets the format, two establish a pattern, three or four pin the awkward cases; returns fall off quickly after that.
When the Examples Run Out
Sooner or later an input arrives that none of your examples covered. You get to choose in advance what happens next.corner-case handling
“prince” — a single name. None of the examples had one.Simple: A locksmith arrives at a door your key will not open. Do you want them to force it, or to telephone you first? The right answer depends entirely on what is behind the door.
Technical: This is a policy decision about behaviour on out-of-distribution inputs, and it has to be stated in the prompt because the default is to generate something. Choose guessing when the error is cheap to detect and reverse; choose escalation when the action has side effects — a database write, an outbound email, a config change.
Give It Permission to Say “I Don’t Know”
Ask a helpful person a question they cannot answer and, if they think their job is to be helpful, they will have a go. Tell them up front that “I don’t know” is an acceptable answer and you get an honest one instead. The same sentence works on a model — and it is one clause you can paste into any prompt.abstention clause
Simple: A confident wrong answer and a confident right answer look exactly the same on the page. That is the whole danger — you cannot tell which one you got by reading it.
Technical: The model has no calibrated abstention behaviour by default: a low-confidence continuation is emitted with exactly the fluency of a high-confidence one. An explicit abstention clause shifts the output distribution towards refusal on unanswerable inputs. As the slide says, it lowers the rate of confident invention; it does not eliminate it.
Where Prompt-Craft Stops Paying
Everything so far has been about the words you send. That takes you a long way, and then it stops — because three of the most common failures have nothing to do with how you phrased the request.
Simple: However clearly you word the question, you cannot ask somebody about a letter they never received. Better wording does not create information that was never in the room.
Technical: Three failures sit outside the prompt’s reach: a training cut-off and private data the model never saw (no wording retrieves an absent fact), no retrieval path with which to verify a claim, and no execution capability. All three are answered by giving the model tools — which is Module 4, not better phrasing.
Make It Show Its Working
Remember being told to show your working in maths class? The point was never the marks — writing the middle steps down is what catches the mistake. Ask the model to do the same: add the sentence “think step by step.”chain-of-thought
23 − 8 = 15. Buys 15 more: 15 + 15 = 30.” Right — and if it had not been right, you could point at the exact line that broke.Simple: Doing a sum in your head against doing it on paper. On paper you get the same answer — but if it is wrong you can point at the line where it went wrong instead of starting over.
Technical: Chain-of-thought prompting elicits intermediate tokens, which give the later tokens something to condition on and improve accuracy on multi-step arithmetic and logic (Wei et al., 2022; the zero-shot form is Kojima et al., 2022). The chain is generated output, not a proof — as the slide warns, it can be fluent and still wrong.
Learn the Technique, Then Cook the Dish
Before asking about your problem, ask about the kind of problem. The general answer becomes the frame the specific answer is built on.step-back, then feed-forward
Simple: Before asking a mechanic what is wrong with your car, ask what usually causes that noise. Now their answer about your car has to line up with what they just told you.
Technical: The first turn places a stated framework into the context window; the second turn’s answer is then conditioned on it and stays consistent across follow-ups instead of being reconstructed ad hoc each time. Documented as step-back prompting (Zheng et al., 2023).
Don’t Build the Castle All at Once
Split one big ask into ordered stages and check the output of each before the next one starts. A bad stage 1 quietly poisons stages 2, 3 and 4.task decomposition & validation gates
Simple: An assembly line only helps if somebody checks each station. Without the checks it is one long job with pauses in it, and a fault at the first station gets built into everything after.
Technical: Decomposition on its own does not improve reliability — the validation gate does. Each stage’s output is checked against a stated condition before it enters the next stage’s context, which bounds error propagation instead of letting a bad stage 1 silently condition stages 2 through 4.
Write the First Words Yourself
You can start the model’s answer for it. Because it only ever continues text, whatever you put at the start of its turn becomes a commitment it has to carry on from.prefill / assistant priming
{, [, <answer>); locking a persona or language for the whole reply; and skipping the “Certainly! I’d be happy to…” opener. Two caveats: not every interface exposes an assistant turn you can write into (chat UIs generally do not; APIs generally do — check the one you are using), and a prefill that fights the request produces worse answers, not better ones.Simple: Handing somebody a form with the first line already filled in. They carry on from where you stopped — they are not going to cross it out and start again with “Dear Sir”.
Technical: Prefill writes the opening tokens of the assistant turn, and because generation is autoregressive continuation those tokens become an unrevisable commitment. Opening with { or [ forces a parseable shape from the first character. As the slide notes, not every interface exposes a writable assistant turn.
Ask for the Plan Before the Work
You would not let a builder start knocking through walls before showing you a drawing. Same instinct here: for anything that touches more than one file, get the approach in writing first, read it, and only then say go. Nothing is changed while you are reading.plan mode
Simple: Reading a quote before agreeing to the work. It costs you five minutes, and it is the cheapest possible moment to discover that the two of you were talking about different jobs.
Technical: Plan mode separates proposal from execution: the model states its intended approach and the file set it would touch, with no writes performed, so the review happens before any state changes. The supervision cost is front-loaded once instead of paid per edit, and the disagreements it surfaces are about the goal rather than the syntax.
Say What “Done” Means
If you can write down how you would check the work, you can hand the checking over too — and stop being asked for approval every thirty seconds.success criteria & delegation
Simple: “Make it look nice” buys you a conversation every ten minutes. “Under two pages, no red, my name at the top” buys you the thing itself.
Technical: Delegation requires a checkable termination condition. Three moves, in order: criteria expressible as a measurement rather than an adjective, an instruction to iterate rather than to answer once, and a stated verbosity for the working. Drop any one and the loop comes back to you for adjudication — and this is the same structure an autonomous agent runs on.
Prompt Engineering Best Practices
Simple: Every one of these six is the same instinct wearing a different coat: say exactly what you want, show it rather than describe it, and check what comes back instead of trusting it.
Technical: The six reduce to two levers. One constrains the output distribution — specificity, exemplars, format constraints, elicited reasoning. The other closes the loop outside the model — tools for retrieval, tests and validators for verification. No amount of the first is a substitute for the second.
Knowledge Check
Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.