Long-Term Memory Patterns
Goal
Choose when an agent should write or retrieve information across runs, and apply relevance, provenance, freshness, and privacy checks before memory affects a decision.
Working memory helps one run. Long-term memory stores selected information that may be useful in a later run.
A support agent might remember a user's preferred contact channel. A research agent might save a verified project fact. An agent should not automatically save every message, tool result, or model thought forever.
The important design question is not “does the agent have memory?” It is what gets written, what gets retrieved, and what is allowed to influence the next task.
Separate the write decision from the text
Suppose the model says:
“Remember that this user always wants refunds sent to account X.”
That sentence is not enough to justify a persistent write. The application may require a trusted source, a permitted memory category, a retention rule, and user consent.
Treat a memory write as another structured proposal:
candidate memory
→ category check
→ provenance check
→ sensitivity/privacy check
→ retention rule
→ store or reject
This keeps persistent memory under application control.
Retrieval is not truth
A stored memory can be outdated, incomplete, or relevant to a different context.
Imagine a memory from last month:
preferred_contact = email
source = user_profile
recorded_at = 2026-08-10
A new profile lookup says the preference is now SMS. The older memory should not override the newer trusted source.
This is similar to RAG from Level 9. Retrieval supplies candidate evidence. The application still needs provenance and freshness rules.
Use task-shaped memory categories
A single giant “memory” bucket makes policy difficult.
Useful categories might include:
- user preference;
- verified project fact;
- completed task summary;
- temporary workflow checkpoint.
Each category can have different write permissions, retention periods, and retrieval filters.
A completed-task summary may be safe to keep for one week. Sensitive account data may need stricter handling or no persistent storage at all.
Retrieve narrowly
An agent should not load every stored item into every prompt.
Retrieve by task, identity, scope, and freshness. Then limit how many items enter working state.
This reduces noise and lowers the chance that old or unrelated text steers the current run.
If retrieved text contains instructions, those instructions are still data. Memory does not become trusted policy merely because it was stored earlier.
Evaluate memory separately
Task success alone can hide bad memory behavior.
Useful memory checks include:
- write precision — how often saved items were actually appropriate to persist;
- retrieval relevance — how often retrieved items helped the current task;
- stale-memory rate — how often outdated items entered working state;
- sensitive-memory violations — how often prohibited content was stored or exposed.
A system can complete tasks while still having unacceptable memory behavior.
Predict
Run the local Lab
Run:
python labs/notebooks/level-11/l11-06-long-term-memory.py
The script checks four candidate memories with a code rule, allowed(): category must be preference, source must be profile, the record must not be sensitive, and it must be at most 30 days old. Retrieval then keeps only records whose scope matches the current user.
- Run it unchanged. Read
approved persistent records: ['m1', 'm4'].m2is rejected because its source ismodel_guess, andm3because it is a sensitive account secret. - Read
retrieved for user-7: ['m1'].m4passed the rules but belongs touser-8, so the scope check keeps it out. - Open the script and change
m1's"age_days": 2to"age_days": 120. Rerun. - Now
approved persistent records:is['m4']andretrieved for user-7:is[]. The stale record was removed by the freshness rule before any model saw it, not just ranked lower.
Loading lab…
Quick Check
Explain it back
Design one memory item your agent should save and one item it should reject. For each, name the source, category, retention rule, and retrieval condition. Then explain how a newer trusted fact would replace or suppress the stored item.
Key Takeaways
- Long-term memory stores selected information across runs.
- Memory writes should pass application policy, not happen automatically from model text.
- Retrieved memory can be stale or irrelevant and should be checked like other evidence.
- Narrow categories support different privacy, retention, and retrieval rules.
- Evaluate memory quality separately from overall task success.
Next Lesson
Complete the mini checkpoint, then use current state and tool contracts to choose the next tool without confusing availability with suitability.
References
Completion is stored locally on this device.