RAG & Vector Search
RAG — Retrieval‑Augmented Generation — is the classic grounding recipe: find the relevant text, put it in the prompt, then answer. Instead of hoping the model already knows, you hand it the facts it needs.
The RAG Loop
Four steps — done once up front, then once per question.
- Embed — turn your documents into vectors
- Store — keep those vectors so you can search them
- Retrieve — at question time, pull the closest chunks
- Generate — answer using those chunks in the prompt
Search By Meaning
An embedding turns text into a vector, so “similar meaning” becomes “close together.” You search by nearness, not by matching exact words.
Why Approximate Is Fine
You want a good match fast, not the provably best one after a long wait.
- Approximate nearest‑neighbour search finds close matches without scanning everything
- It trades a little recall for a lot of speed
- For grounding an answer, a near‑best chunk is almost always as useful as the best
Retrieval Sets the Limit
Retrieval quality caps answer quality. If you retrieve the wrong chunks, the model answers from the wrong facts — confidently. Garbage in, garbage out.
Build It
How to implement: take a handful of your own documents, split them into chunks, and think hard about what a good “retrieve the right chunk” query looks like — that query is where most of the quality lives.
- Weekly AI Tasks tracker — RAG over your notes answers “what did I do last week?” by retrieving the relevant entries.
- Personal brand site — retrieve the right résumé or project snippet so generated copy stays grounded and specific.
What you learned
RAG grounds a model: embed, store, retrieve, generate. Vector search finds chunks by meaning, approximate search keeps it fast, and retrieval quality caps the answer. For the deeper RAG treatment, see GenAI‑101.