Context & Retrieval
Managing the most precious resource in AI — what the model can see, and how to make every token count.
Context — The Scarce Resource
The context window is everything the model can see at once. Input + output combined. It’s finite, expensive, and the #1 bottleneck.
Simple: It is a desk, not a filing cabinet. Only what is laid out on the desk right now can be worked on — and the finished page has to fit on the same desk, next to the source material it was written from.
Technical: The context window is a hard token limit spanning prompt and completion, re-supplied in full on every request. The model is stateless between calls, so nothing outside that window exists to it; everything you think of as memory is text someone put back in.
What Actually Fills It
One bag, one fixed size — when it is full, the only way to add something is to take something out. And packing for a trip, the socks are not the problem, the boots are. Same here: instructions and tool definitions are small, and files and code are what fill the bag. Toggle each item to pack or unpack it.
Simple: Packing is not about what you want to bring, it is about what you are willing to leave behind. The moment the bag is full, every new thing you add is a decision to remove something else — and the bag fills long before you expected it to.
Technical: The window is one shared budget across system prompt, tool schemas, retrieved passages, conversation history and the completion. Tool definitions in particular are a fixed per-request tax paid whether or not a tool is called, and the reserve left for output is the constraint people forget: overrun it and the answer is truncated, not the input.
What the Model Forgets
Think of skimming a long email: you remember the opening line and the final ask, and the paragraphs in between blur. Long context behaves the same way — a fact buried mid-document is the one most likely to be missed.primacy & recency bias
Simple: The dangerous part is not that the middle is forgotten — it is that nothing tells you it was. You get a confident, complete-sounding answer built on the two-thirds of the document the model actually used.
Technical: Retrieval accuracy over long contexts traces a U-shape against the position of the needed fact, strongest at the extremes and weakest in the middle (Liu et al., 2023). It is a positional effect, not a capacity one: the fact is inside the window and still under-attended, so a larger context window does not fix it. Ordering the input does.
Context Management Strategies
Every token counts. Here’s how to use them wisely.
Simple: All four strategies are one instinct: bring the page, not the library. A researcher who carries every book they might need cannot lift any of them — the skill is knowing which page the question turns on.
Technical: The four differ in when the reduction happens. Compression reduces at write time and is lossy but cheap. Retrieval defers selection to query time and is only as good as the retriever. Tool access defers it to the model's own judgement mid-loop. Pruning reduces after the fact and discards conversation state, so anything not summarised out is gone.
Choose Your Model
Same family, three sizes — and the difference that matters is not a benchmark, it is what job you would hand each one. Click a card.
Slower and dearer, and worth it exactly when the thinking is the work rather than the typing.
Try it: read three conflicting reports and tell me where they actually disagree.
Simple: You do not book the consultant surgeon to take out a splinter, and you do not send the sprinter to run the marathon. Picking the biggest model for everything is not caution, it is paying specialist rates for splinter work — and waiting longer for it.
Technical: Tiers trade capability against latency and cost per token, and the three axes move together. The right selector is the difficulty of a single unit of work, not the total volume: a thousand trivial classifications belong on the fastest tier, one genuinely ambiguous judgement belongs on the most capable. Version names, prices and speeds are the volatile part; the shape of the trade-off is not.
When to Use What
Sort by the difficulty of one unit of work, not by how much of it there is.
The RAG Pipeline
Click each step to see how RAG transforms a user query into a grounded answer.
"How do I set up authentication in Next.js 15?" This natural language query is the starting point of the pipeline.
Simple: Ask a librarian a question and they do not recite an answer from memory. They work out what you are really after, walk to the right shelf, bring back three pages, and then answer — with the pages open in front of them.
Technical: Five stages, and the one that decides quality is Search, not Generate. The query is embedded into the same vector space as the indexed chunks, nearest neighbours are retrieved, and the top-k are concatenated into the prompt. The generator can only be as right as what Search handed it — retrieval failures surface as fluent, well-formed wrong answers.
How Close Are Two Meanings?
Search step 3 asks a question you can put a number on: how close are these two meanings? Every word becomes a point in space, and closeness is measured as the angle between them — not the letters they share.cosine similarity
Simple: Two people can point in almost the same direction from opposite ends of a field. What is being compared is the direction they are facing, not where they are standing — which is how a short note and a long report can score as being about the same thing.
Technical: Cosine similarity is the dot product of two vectors divided by their magnitudes, so it measures angle and discards length. Length in an embedding tends to track frequency and document size rather than meaning, so normalising it away is the point, not a simplification. Identical direction scores 1, unrelated is near 0, and opposed is −1 — though genuine −1 is rare in practice.
Matching Words vs. Matching Meaning
Old-style search looks for the words you typed. Semantic search looks for what you meant. Run both against the same three-document library and watch where the first one falls over. One honest caveat: the left column below is naive exact matching — real keyword search stems words, so it would match “bake” to “baking” and this exact failure would be softer. The shape still holds when the query shares no vocabulary with the document at all, and that is why the next slide keeps keyword search rather than replacing it.
Simple: Ask a librarian for “that book with the blue cover about the war” and they find it. Ask a computer that searches for the literal word “blue” and it hands you a book on paint. But ask either of them for a part number and the literal matcher wins outright.
Technical: Lexical retrieval scores on term overlap and cannot bridge a vocabulary mismatch; dense retrieval scores on embedding proximity and bridges it easily, but blurs rare exact strings — identifiers, error codes, part numbers — because they carry little semantic signal. The two fail on disjoint cases, which is precisely why hybrid search runs both and fuses the rankings rather than choosing.
Why RAG Beats Fine-Tuning
Simple: Teaching someone a fact and teaching them a manner are different lessons. If your problem is “it does not know our returns policy”, hand it the policy. If your problem is “it does not sound like us”, no amount of handing it documents will fix that.
Technical: Fine-tuning updates weights and is the right tool for shifting style, format adherence and task framing. RAG leaves weights untouched and injects knowledge at inference, so the corpus is updatable, auditable and citable. Facts baked into weights cannot be revised without retraining, cannot be attributed, and cannot be revoked — three properties that decide the choice more often than accuracy does.
Inside “Search the Vector Database”
ERR_4032 into your support search. Ask for the meaning of that string and you get back documents about errors. Count the letters instead and you get the one document that actually names ERR_4032. Neither is enough alone, and neither can afford to compare your query against ten million documents one at a time. Three jobs fix that.ERR_4032. Run it alongside the vector search and you have hybrid search.Agentic Search Patterns
Agents don’t just search once — they iteratively refine their search strategy based on results.
Simple: The four levels are four answers to one question: who is allowed to decide the search was no good? At level 1, nobody — you get what you get. By level 4, the agent itself reads the results, judges them poor, and goes again.
Technical: Levels 1–3 are all single-pass: the query may be rewritten or fanned out, but the number of retrieval rounds is fixed before execution. Level 4 puts retrieval inside a control loop, so round count becomes data-dependent. That is what buys recall on under-specified queries and what makes latency and cost unpredictable — which is why an agentic searcher needs a step cap in a way the other three do not.
Following the Clues
A detective does not solve the case with one question. They ask, look at what came back, and let that shape the next question. An agent searches the same way — think, act, look at the result, think again.ReAct
Simple: The detective’s first question is rarely the useful one — it is the one that tells them what to ask next. Skip the looking-at-what-came-back step and you are not investigating, you are guessing out loud in a sequence.
Technical: ReAct interleaves reasoning traces with tool calls so that each observation is in context before the next action is chosen (Yao et al., 2022). The observation step is what distinguishes it from chain-of-thought: reasoning alone can only elaborate on what the model already believes, whereas an observation can contradict it. That is also the failure mode — if observations are never contradicting, the loop is only confirming itself.
Chunking Strategies
How you split documents determines retrieval quality. Adjust the chunk size below.
Simple: Cut a recipe into single lines and “bake for 40 minutes” arrives without saying what you are baking. Keep it as one page and every search for anything in that recipe drags the whole page along. The cut has to fall where the meaning already changes.
Technical: Chunk size sets a precision/context trade-off at both ends of the pipeline: the chunk is the unit that gets embedded, so a large one averages several topics into a single vector and retrieves vaguely, while a small one retrieves sharply but may land without the surrounding sentences needed to use it. Overlap between adjacent chunks buys back some of the lost boundary context at the cost of a larger index, and splitting on structure — headings, paragraphs — usually beats splitting on a fixed count.
Long-Horizon Task Management
Complex tasks can span hundreds of tool calls. How agents stay on track.
Simple: On a job that runs for days you do not rely on remembering — you keep a notebook. The plan is on paper, the findings are on paper, and when your head is full you hand the next piece to someone with a clear one.
Technical: All four are responses to the same constraint: the window is fixed but the task is not. A written plan externalises state that would otherwise have to stay resident. Compression trades fidelity for room and is irreversible, so what it drops is gone. Files move state out of the window entirely and are the only one of the four that survives a reset. Subagents partition the budget instead of stretching it — each gets a clean window, and only the summary comes back.
Two Shapes of Task
A librarian and a detective are both good at their jobs. You would not send the detective to fetch a book, or the librarian to solve the case. RAG is a librarian; an agent is a detective.
Simple: Both can be wrong, and they are wrong in different ways. The librarian brings you the wrong book and tells you about it with total confidence. The detective follows a bad lead for an hour and bills you for the hour.
Technical: The two shapes have different cost curves and different failure signatures. Single-pass retrieval has bounded latency and a single point of failure at the retriever, so a bad retrieval yields one confident wrong answer. An agentic loop has unbounded step count, so its errors compound across steps and its cost is variable by construction. That asymmetry, not capability, is why the loop needs a step cap and an explicit stop condition.
Which Should I Use?
One question decides it: can this be answered by finding the right passage, or does something have to happen? Classify each task, then check yourself.
Knowledge Check
Eight questions on this module. Answer to see why — the explanation appears whether you were right or wrong.