Prompt vs Retrieve-on-Demand
One of the first decisions in any AI feature: what do you put directly in the prompt, and what do you let the model fetch on demand with a tool? Get this right and your feature stays simple, fresh, and affordable.
Put It In the Prompt
Include information directly when it is small, stable, and needed on almost every request.
- Small — it fits comfortably without crowding out the question.
- Stable — it rarely changes, so it will not go stale between requests.
- Always needed — instructions, an output schema, a few key facts.
Retrieve on Demand
Fetch information at question time when the data is large, changes often, or is only sometimes needed.
The Tradeoff
Neither choice is free — you are trading simplicity against freshness and scale.
- In-prompt is simple, but it fills the context window and can go stale.
- Retrieval is fresh and scales to large data, but adds a moving part that can fail.
Agentic Retrieval
Instead of pre-fetching everything, let the agent decide what to look up as it works — the idea we build on in Module 7. It is powerful, but harder to predict, since you no longer control exactly what gets fetched.
Build It
How to implement: for one feature, sort the information it needs into two columns — “always in the prompt” vs “fetch when needed” — and note who does the fetching.
- Weekly AI Tasks tracker — the task schema goes in the prompt; the actual tasks are retrieved on demand, since they change constantly.
- Personal brand site — the tone and rules go in the prompt; specific project facts are retrieved per section.
What you learned
Put information in the prompt when it is small, stable, and always needed; retrieve on demand when it is large, changing, or occasional. In-prompt is simpler; retrieval is fresher and scales — and letting the agent decide is more powerful but harder to predict.