Context & Retrieval

Managing the most precious resource in AI — what the model can see, and how to make every token count.

Context — The Scarce Resource

The context window is everything the model can see at once. Input + output combined. It’s finite, expensive, and the #1 bottleneck.

📐
Context Window Visualizer
See how different tasks fill the window
System
Context/Docs
User
Output
Free
Simple chat uses only ~15% of the window. Most of the context is free for the conversation to grow.
In Plain English

Simple: It is a desk, not a filing cabinet. Only what is laid out on the desk right now can be worked on — and the finished page has to fit on the same desk, next to the source material it was written from.

Technical: The context window is a hard token limit spanning prompt and completion, re-supplied in full on every request. The model is stateless between calls, so nothing outside that window exists to it; everything you think of as memory is text someone put back in.

What Actually Fills It

One bag, one fixed size — when it is full, the only way to add something is to take something out. And packing for a trip, the socks are not the problem, the boots are. Same here: instructions and tool definitions are small, and files and code are what fill the bag. Toggle each item to pack or unpack it.

🎒
Pack the Context Backpack
Four things go in every request — one of them is huge
Backpack empty
Nothing packed yet.
Percentages are illustrative, not measured — the point is the arithmetic. The window is one fixed budget shared by the system prompt, tool definitions, retrieved documents, conversation history and the answer itself. Pack all four here and you are at 80% before the model has written a word. This is why retrieve the relevant chunk beats paste the whole file.
In Plain English

Simple: Packing is not about what you want to bring, it is about what you are willing to leave behind. The moment the bag is full, every new thing you add is a decision to remove something else — and the bag fills long before you expected it to.

Technical: The window is one shared budget across system prompt, tool schemas, retrieved passages, conversation history and the completion. Tool definitions in particular are a fixed per-request tax paid whether or not a tool is called, and the reserve left for output is the constraint people forget: overrun it and the answer is truncated, not the input.

What the Model Forgets

Think of skimming a long email: you remember the opening line and the final ask, and the paragraphs in between blur. Long context behaves the same way — a fact buried mid-document is the one most likely to be missed.primacy & recency bias

🫥
Lost in the Middle
Watch the model read a long passage, then check what it kept
Ready — full passage in context
—
Start recall
—
Middle recall
—
End recall
This is the “lost in the middle” effect, named and measured by Liu et al., Lost in the Middle: How Language Models Use Long Contexts (2023). Their finding: retrieval accuracy is highest when the needed fact sits at the very start or the very end of the context, and dips when it sits in the middle. The percentages above are illustrative of that U-shape, not figures from the paper. Practical rule: put the instruction and the critical fact at the edges.
In Plain English

Simple: The dangerous part is not that the middle is forgotten — it is that nothing tells you it was. You get a confident, complete-sounding answer built on the two-thirds of the document the model actually used.

Technical: Retrieval accuracy over long contexts traces a U-shape against the position of the needed fact, strongest at the extremes and weakest in the middle (Liu et al., 2023). It is a positional effect, not a capacity one: the fact is inside the window and still under-attended, so a larger context window does not fix it. Ordering the input does.

Context Management Strategies

Every token counts. Here’s how to use them wisely.

🗜️
Compress, Don’t Dump
Summarize long docs instead of pasting everything. Focus on what’s relevant.
🎯
Retrieve, Don’t Include
Use RAG to pull only the relevant chunks instead of the entire knowledge base.
🔧
Tools Over Context
Let the agent read files on demand instead of pre-loading everything into context.
✂️
Prune Conversation
Long conversations accumulate noise. Summarize and restart for complex tasks.
In Plain English

Simple: All four strategies are one instinct: bring the page, not the library. A researcher who carries every book they might need cannot lift any of them — the skill is knowing which page the question turns on.

Technical: The four differ in when the reduction happens. Compression reduces at write time and is lossy but cheap. Retrieval defers selection to query time and is only as good as the retriever. Tool access defers it to the model's own judgement mid-loop. Pruning reduces after the fact and discards conversation state, so anything not summarised out is gone.

Choose Your Model

Same family, three sizes — and the difference that matters is not a benchmark, it is what job you would hand each one. Click a card.

🧠
Opus
The specialist
⚡
Sonnet
The all-rounder
🚀
Haiku
The sprinter
Opus — the specialist you book for the hard one.most capable tier
Slower and dearer, and worth it exactly when the thinking is the work rather than the typing.

Try it: read three conflicting reports and tell me where they actually disagree.
Tier names and relative speed are stable; specific model versions, prices and speeds change often, so treat any figure you read elsewhere as dated. The shape of the choice — cheap-and-fast, balanced, slow-and-deep — is what to remember, and it applies to every vendor’s line-up, not just this one.
In Plain English

Simple: You do not book the consultant surgeon to take out a splinter, and you do not send the sprinter to run the marathon. Picking the biggest model for everything is not caution, it is paying specialist rates for splinter work — and waiting longer for it.

Technical: Tiers trade capability against latency and cost per token, and the three axes move together. The right selector is the difficulty of a single unit of work, not the total volume: a thousand trivial classifications belong on the fastest tier, one genuinely ambiguous judgement belongs on the most capable. Version names, prices and speeds are the volatile part; the shape of the trade-off is not.

When to Use What

Sort by the difficulty of one unit of work, not by how much of it there is.

Task
Best Model
Why
Tag 5,000 reviews happy / unhappy
Haiku
One tiny judgement, repeated thousands of times — speed and cost win
Summarise a long meeting
Haiku
The material is already there; nothing has to be worked out
Draft a reply, fix a spreadsheet formula
Sonnet
Real thought needed, but not deep thought — the everyday default
Fix a bug in code you did not write
Sonnet
Good reasoning, fast turnaround
Reconcile three reports that disagree
Opus
The answer is not in any one source; it has to be reasoned out
Design a system, or restructure a large codebase
Opus
Nuanced, interdependent decisions with expensive mistakes
Rule of thumb: Start with Sonnet for most tasks. Upgrade to Opus for complex reasoning. Use Haiku for high-volume, simple operations.

The RAG Pipeline

Click each step to see how RAG transforms a user query into a grounded answer.

01
💬
Query
→
02
🔢
Embed
→
03
🔎
Search
→
04
📎
Augment
→
05
✨
Generate
Query: The user asks a question: "How do I set up authentication in Next.js 15?" This natural language query is the starting point of the pipeline.
In Plain English

Simple: Ask a librarian a question and they do not recite an answer from memory. They work out what you are really after, walk to the right shelf, bring back three pages, and then answer — with the pages open in front of them.

Technical: Five stages, and the one that decides quality is Search, not Generate. The query is embedded into the same vector space as the indexed chunks, nearest neighbours are retrieved, and the top-k are concatenated into the prompt. The generator can only be as right as what Search handed it — retrieval failures surface as fluent, well-formed wrong answers.

How Close Are Two Meanings?

Search step 3 asks a question you can put a number on: how close are these two meanings? Every word becomes a point in space, and closeness is measured as the angle between them — not the letters they share.cosine similarity

📐
Compare Two Words
Click a pair to score it
No comparisons yet
Pick a pair above to see its score.
Scores are illustrative values chosen to show the shape of the result, not measurements from a specific embedding model. Note king / queen: opposites in one respect, yet close overall, because they share almost everything else. And car / happiness near zero — unrelated meanings sit at right angles.
In Plain English

Simple: Two people can point in almost the same direction from opposite ends of a field. What is being compared is the direction they are facing, not where they are standing — which is how a short note and a long report can score as being about the same thing.

Technical: Cosine similarity is the dot product of two vectors divided by their magnitudes, so it measures angle and discards length. Length in an embedding tends to track frequency and document size rather than meaning, so normalising it away is the point, not a simplification. Identical direction scores 1, unrelated is near 0, and opposed is −1 — though genuine −1 is rare in practice.

Matching Words vs. Matching Meaning

Old-style search looks for the words you typed. Semantic search looks for what you meant. Run both against the same three-document library and watch where the first one falls over. One honest caveat: the left column below is naive exact matching — real keyword search stems words, so it would match “bake” to “baking” and this exact failure would be softer. The shape still holds when the query shares no vocabulary with the document at all, and that is why the next slide keeps keyword search rather than replacing it.

🔎
Two Searches, One Library
Library: “Best Chocolate Cake Recipe” · “Cookie Baking Guide” · “Homemade Brownie Tips”
Pick a query
Naive exact matching no stemming
Waiting for a query.
Semantic search meaning
Waiting for a query.
Match percentages are illustrative, and the left column is deliberately the naive version: it compares surface forms only. A production keyword scorer such as BM25 stems its terms, so bake/baking would match and this particular miss would not happen. What survives stemming is the harder case — a query whose vocabulary simply does not overlap the document’s (“something sweet” vs. “brownie”), which no amount of word matching can bridge. That is the gap embeddings close, and the reason the answer is to run both.
In Plain English

Simple: Ask a librarian for “that book with the blue cover about the war” and they find it. Ask a computer that searches for the literal word “blue” and it hands you a book on paint. But ask either of them for a part number and the literal matcher wins outright.

Technical: Lexical retrieval scores on term overlap and cannot bridge a vocabulary mismatch; dense retrieval scores on embedding proximity and bridges it easily, but blurs rare exact strings — identifiers, error codes, part numbers — because they carry little semantic signal. The two fail on disjoint cases, which is precisely why hybrid search runs both and fuses the rankings rather than choosing.

Why RAG Beats Fine-Tuning

Aspect
Fine-Tuning
RAG
Data freshness
Frozen at training
Always current
Cost
Expensive to retrain
Just update the index
Traceability
Black box
Cite exact sources
Setup time
Hours to days
Minutes with embeddings
Hallucination
Still common
Grounded in sources
In practice: RAG is the default choice for most production AI applications. Fine-tuning is reserved for changing model behavior (tone, format), not adding knowledge.
In Plain English

Simple: Teaching someone a fact and teaching them a manner are different lessons. If your problem is “it does not know our returns policy”, hand it the policy. If your problem is “it does not sound like us”, no amount of handing it documents will fix that.

Technical: Fine-tuning updates weights and is the right tool for shifting style, format adherence and task framing. RAG leaves weights untouched and injects knowledge at inference, so the corpus is updatable, auditable and citable. Facts baked into weights cannot be revised without retraining, cannot be attributed, and cannot be revoked — three properties that decide the choice more often than accuracy does.

Inside “Search the Vector Database”

Start with a problem you can feel. A user types ERR_4032 into your support search. Ask for the meaning of that string and you get back documents about errors. Count the letters instead and you get the one document that actually names ERR_4032. Neither is enough alone, and neither can afford to compare your query against ten million documents one at a time. Three jobs fix that.
🎯
Close enough, instantly ANN
You do not want the provably nearest café — you want a good one in ten seconds. Same trade here: give up the single closest document and a very close one arrives thousands of times faster.
🕸️
A map with motorways and side streets HNSW
You do not drive to another city street by street: motorway across the country, then the A-roads, then the local streets to the door. This index is built the same way, and most vector databases use it.
🔤
Ctrl-F that knows “the” doesn’t count BM25
Counts how often your words appear, discounts the words that appear in every document, and stops rewarding length. This is the one that finds ERR_4032. Run it alongside the vector search and you have hybrid search.
Each card opens the paper behind it. ANN is approximate nearest neighbour — a class of method, not one algorithm.

Agentic Search Patterns

Agents don’t just search once — they iteratively refine their search strategy based on results.

Level 1
Naive Search
Single query, return top-K results. Simple but often misses nuance.
Level 2
Query Rewriting
AI rewrites the query for better embedding match before searching.
Level 3
Multi-Query
Generate multiple perspectives of the same question, search each, merge results.
Level 4
Agentic Search
Agent searches, evaluates results, refines query, searches again — looping until satisfied.
In Plain English

Simple: The four levels are four answers to one question: who is allowed to decide the search was no good? At level 1, nobody — you get what you get. By level 4, the agent itself reads the results, judges them poor, and goes again.

Technical: Levels 1–3 are all single-pass: the query may be rewritten or fanned out, but the number of retrieval rounds is fixed before execution. Level 4 puts retrieval inside a control loop, so round count becomes data-dependent. That is what buys recall on under-specified queries and what makes latency and cost unpredictable — which is why an agentic searcher needs a step cap in a way the other three do not.

Following the Clues

A detective does not solve the case with one question. They ask, look at what came back, and let that shape the next question. An agent searches the same way — think, act, look at the result, think again.ReAct

🕵️
Watch an Agent Refine Its Search
One vague question, three searches, one usable answer
Pick a question
Pick a question to watch the loop run.
The reason→act→observe pattern was named ReAct by Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022). Note what the first search does not produce: an answer. It produces a better question. That is the whole difference from single-shot retrieval.
In Plain English

Simple: The detective’s first question is rarely the useful one — it is the one that tells them what to ask next. Skip the looking-at-what-came-back step and you are not investigating, you are guessing out loud in a sequence.

Technical: ReAct interleaves reasoning traces with tool calls so that each observation is in context before the next action is chosen (Yao et al., 2022). The observation step is what distinguishes it from chain-of-thought: reasoning alone can only elaborate on what the model already believes, whereas an observation can contradict it. That is also the failure mode — if observations are never contradicting, the loop is only confirming itself.

Chunking Strategies

How you split documents determines retrieval quality. Adjust the chunk size below.

✂️
Chunk Size Explorer
See how different sizes affect document splitting
Medium
3 chunks — balanced between context and precision.
In Plain English

Simple: Cut a recipe into single lines and “bake for 40 minutes” arrives without saying what you are baking. Keep it as one page and every search for anything in that recipe drags the whole page along. The cut has to fall where the meaning already changes.

Technical: Chunk size sets a precision/context trade-off at both ends of the pipeline: the chunk is the unit that gets embedded, so a large one averages several topics into a single vector and retrieves vaguely, while a small one retrieves sharply but may land without the surrounding sentences needed to use it. Overlap between adjacent chunks buys back some of the lost boundary context at the cost of a larger index, and splitting on structure — headings, paragraphs — usually beats splitting on a fixed count.

Long-Horizon Task Management

Complex tasks can span hundreds of tool calls. How agents stay on track.

📋
Plan First
Break the task into steps before executing. Agents use internal plans to track progress.
🗜️
Context Compression
As context fills up, older messages are summarized to make room for new information.
🧠
Persistent Memory
Save key facts to files so they persist across context window resets.
🤝
Subagents
Delegate subtasks to fresh agents with their own context windows.
In Plain English

Simple: On a job that runs for days you do not rely on remembering — you keep a notebook. The plan is on paper, the findings are on paper, and when your head is full you hand the next piece to someone with a clear one.

Technical: All four are responses to the same constraint: the window is fixed but the task is not. A written plan externalises state that would otherwise have to stay resident. Compression trades fidelity for room and is irreversible, so what it drops is gone. Files move state out of the window entirely and are the only one of the four that survives a reset. Subagents partition the budget instead of stretching it — each gets a clean window, and only the summary comes back.

Two Shapes of Task

A librarian and a detective are both good at their jobs. You would not send the detective to fetch a book, or the librarian to solve the case. RAG is a librarian; an agent is a detective.

 
RAG
Agentic
Shape
Retrieve, then answer — once
Loop: plan, act, check, repeat
Fits when
The answer already exists in documents you hold
The task needs action, iteration, or reasoning past lookup
Typical job
Q&A over a known corpus; policy and doc lookup; citation
Debug and fix; multi-source research; anything with steps
Steps per request
One
Unbounded — the model decides when to stop
Cost & latency
Low and predictable
Higher and variable
Failure mode
Retrieves the wrong chunk and answers confidently
Loops, wanders, or acts on a wrong conclusion
Not either/or at the system level. An agent typically has RAG as one of its tools — it retrieves when retrieval is what the current step needs, then goes on to do something with the result. So the real question is never “RAG or agents?” but “does this task finish in one retrieval, or does it need a loop?” If one retrieval answers it, do not pay for a loop.
In Plain English

Simple: Both can be wrong, and they are wrong in different ways. The librarian brings you the wrong book and tells you about it with total confidence. The detective follows a bad lead for an hour and bills you for the hour.

Technical: The two shapes have different cost curves and different failure signatures. Single-pass retrieval has bounded latency and a single point of failure at the retriever, so a bad retrieval yields one confident wrong answer. An agentic loop has unbounded step count, so its errors compound across steps and its cost is variable by construction. That asymmetry, not capability, is why the loop needs a step cap and an explicit stop condition.

Which Should I Use?

One question decides it: can this be answered by finding the right passage, or does something have to happen? Classify each task, then check yourself.

🧭
Classify the Task
Pick RAG or Agentic for each — feedback is immediate
0 of 5 classified
Further watching: RAG vs. agentic AI — external explainer video ↗ — an outside take on the same distinction; not required for this module.

Knowledge Check

Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.

Question 1 of 0
Score 0/0

Key Takeaways

📐
Context Is Finite
Every token costs money and attention. Files and code dominate the budget — be strategic about what goes in.
🫥
Mind the Middle
Models recall the start and end of long context best. Put the instruction and the critical fact at the edges.
🧠
Right Model, Right Task
Opus for deep reasoning, Sonnet for daily work, Haiku for speed.
🔎
RAG Is Essential
The standard for production AI. Under the hood: ANN indexes (usually HNSW) plus BM25 hybrid search, then re-ranking.
✂️
Chunk Wisely
Chunk size determines retrieval quality. Too small loses context, too large adds noise.
🧭
Match the Task Shape
One retrieval answers it → RAG. Needs action, iteration or self-correction → agentic. Agents use RAG as a tool, so it is not either/or.
Context WindowLost in the MiddleOpusSonnetHaikuRAGEmbeddingsCosine SimilarityANNHNSWBM25Hybrid SearchChunkingReActSubagents