When One Window Isn’t Enough
An agent working through a long task can’t hold everything in its context window at once. To stay coherent across many steps, it needs a memory architecture — a plan for what to keep, where, and how to get it back.
Short‑Term, Working, Long‑Term
Most agent designs use three complementary layers.
- Short-term — the current context window: what the model can see right now.
- Working notes — a scratchpad the agent re-reads: a plan, a running summary, decisions so far.
- Long-term store — facts saved and retrieved later, often via retrieval (Module 6).
The Window Is Not Memory
The context window is temporary. As the session grows, older content scrolls out — and once it’s gone, the model no longer knows it existed.
Managing a Long Session
The same three moves keep a long-running session from drowning in its own history — the discipline from managing long agent sessions (Module 3, Lesson 8) and context windows (Module 5).
- Summarize & compact — fold finished work into a short summary so the window stays lean.
- Persist decisions — save choices and results outside the window before they scroll away.
- Retrieve on demand — pull back only what the current step actually needs.
Match Memory to the Task
Memory is not free — every stored fact is something to write, index, and search. So you size it to the job.
Build It
How to implement: decide what your agent must remember beyond one turn, and where it lives — a file, a database, or a vector store.
- Weekly AI Tasks tracker — long-term memory is your task database; the agent retrieves the relevant week rather than holding all of history in context.
- Personal brand site — little memory needed; the “store” is the source files it re-reads per section.
What you learned
Memory is three layers — short-term window, working notes, long-term store. The window is not memory, so keep what matters by writing it down and retrieving it, and size the whole design to the task.