Memory: Why Agents Forget
People are often surprised when an agent forgets something it clearly knew ten minutes ago, or starts a new session with no idea what you worked on yesterday. This is not a bug in one product. It comes from how the models underneath work.
Agents have two very different kinds of memory: the short-term working space the model can see right now, and longer-term memory stored outside the model. Knowing the difference explains most "why did it forget?" moments, and how to work around them.
What You'll Learn
- What the context window is and why it acts as an agent's short-term memory
- Why long agent tasks run out of room, and what agents do about it
- How long-term memory works, and why it is really "notes plus search"
- Practical habits that help agents remember what matters
Short-term memory: the context window
A language model does not remember anything between calls. Each time the agent loop asks "what next?", the software sends the model everything it needs to know again: the instructions, the goal, the tool menu, and the full history of steps and results so far.
All of that has to fit in the model's context window, a fixed limit on how much text it can take in at once. Think of it as the model's desk. Only what is on the desk can be used. Anything not on the desk, the model simply cannot see.
In a normal chat, the desk rarely fills up. Agents are different. Every step adds more to the pile: search results, whole web pages, file contents, error messages. A long task can fill the desk surprisingly fast.
How LLMs Actually Work explains tokens and context windows in more detail if you want the mechanics.
What happens when the desk is full
Agent software has a few ways to make room. Each has a cost.
Every way of freeing space trades away some information.
| Criteria | What the agent does | What can go wrong |
|---|---|---|
| Drop the oldest steps | Removes early history | Forgets your original instructions or constraints |
| Summarize the history | Squeezes old steps into a short summary | Small but important details get lost |
| Keep only key results | Saves findings, drops raw pages | May drop the one detail needed later |
| Start a sub-task fresh | Gives a helper agent a clean desk | The helper lacks the bigger picture |
What the agent does
- Drop the oldest steps
- Removes early history
- Summarize the history
- Squeezes old steps into a short summary
- Keep only key results
- Saves findings, drops raw pages
- Start a sub-task fresh
- Gives a helper agent a clean desk
What can go wrong
- Drop the oldest steps
- Forgets your original instructions or constraints
- Summarize the history
- Small but important details get lost
- Keep only key results
- May drop the one detail needed later
- Start a sub-task fresh
- The helper lacks the bigger picture
This is why a long agent task can drift. Early on, you said "only free courses" or "never email the client directly." Fifty steps later, that instruction has been summarized away or pushed off the desk, and the agent quietly breaks it.
There is a second, quieter effect. Even when everything fits, models tend to pay less attention to details buried in the middle of a very long context. A crowded desk is harder to use well, even if nothing has fallen off.
Long-term memory: notes plus search
Some assistants and agents remember you across sessions: your name, your job, your preferences, a project you worked on last week. This is not the model learning. The model itself does not change when you use it.
Instead, the software keeps notes outside the model, and brings the relevant ones back onto the desk when needed.
- SaveUseful facts are written to a memory store
- SearchAt the next task, the store is searched for relevant notes
- LoadMatching notes are placed in the context window
- UseThe model can now see and use them
This explains long-term memory's quirks:
- It remembers what was saved, not everything. If a detail was never written to memory, it is gone.
- Search can miss. If the search does not find the right note, the agent acts as if it never knew.
- Old notes can be wrong. A preference you changed months ago may still be stored and still be used.
- You can often see and edit it. Many assistants have a memory settings page. Checking it now and then is worth it.
The same idea, searching a store of documents and loading the relevant parts, is how agents answer questions about your files. It is often called RAG, short for retrieval-augmented generation.
Habits that help agents remember
You cannot change the model's limits, but you can work with them:
- Put key rules in the right place. Many products have custom instructions, project instructions, or a memory feature. Rules there are loaded every time, instead of being lost in the history.
- Restate the important constraints when you give a long task: "Remember: only free options, and ask before sending anything."
- Break big jobs into smaller tasks. Several focused runs usually beat one giant run that fills the desk.
- Ask for a written summary at the end of a work session, then paste it in at the start of the next one.
- Start a fresh session when a long conversation starts going off track. A clean desk often fixes strange behavior.
Key Takeaways
- The context window is the agent's short-term memory. The model can only use what fits in it right now.
- Agent tasks fill the window fast because every step adds results. When it is full, old details get dropped or summarized.
- This is why long tasks drift and break instructions given early on.
- Long-term memory is notes saved outside the model and searched when needed. The model itself does not learn from your chats.
- Keep key rules in instructions or memory settings, restate constraints, split big tasks, and start fresh when things drift.

