Agent systems · Memory · Operations
Your agent's memory is a landfill, not a library
Storage and retrieval help agents find facts, but changing knowledge also needs curation — deciding what to keep, revise, and trust.
Your agent said something wrong last week. You went to fix it, and you could not find where the belief lived. It was buried in a store you had to go looking through record by record, extracted by a model that never told you why, sitting next to a hundred other "facts" you never approved.
Or you went the other way. Six months ago you gave your agent a tidy folder of markdown notes. You have added to it faithfully ever since. Yet the agent still repeats a mistake you thought those notes had settled.
These are curation problems, not proof that you picked the wrong tool. Storage and retrieval decide where facts go and how people or agents find them. Those are useful jobs. A searchable knowledge base or a well-kept folder of notes can be valuable on its own. Not every team needs an agent, an automated workflow, or the four-tier approach below.
The harder question arises when an agent relies on knowledge that changes. Who decides which correction replaces an old belief? Who checks that the source still supports the conclusion? Someone needs to tend that collection. We call that responsibility the librarian. It is a defined review process, not a claim that software can take over a person's judgment.
Two approaches, and the curation question
Two common approaches make different parts of this work easier.
One approach extracts memory automatically. Your agent talks, a model pulls out what it judged to matter, and writes it somewhere for later retrieval. Calling these tools black boxes misses the inspection features available across this category: memory browsers, per-record histories, editable blocks, and traversable graphs. Those features make review easier.
Inspection still needs a process. Reviewing every extracted record can become impractical as volume grows. Decide which changes need approval, which can be sampled, and how someone can correct an error. More memory may improve recall. It does not, by itself, establish that the retained conclusions are sound.
Another approach stores editable notes. Human-readable files. Plain markdown. You can search them, version them in git, and revise them directly. Some tools also provide an editing surface: rewrite a section in place, replace a block, or run a consolidation pass on demand. For a focused collection with a clear maintainer, this may be all you need.
The risk is leaving that maintenance undefined. If nobody reconciles the notes, the same lesson can appear in several conflicting forms. Scheduled consolidation can help. A timer alone does not settle which version to trust or who reviews a disputed merge. Those decisions need rules that fit the collection and the consequences of an error.
Some systems combine editable, git-backed markdown with a background pass that reviews recent sessions and merges duplicates. That can be a useful curation tool, not just a storage feature. The evaluation question is practical: can you set the review rules, inspect changes, and correct a bad merge? For lower-risk work, reviewing the diff afterward may be enough. Higher-consequence decisions may need approval first.
Growth is not intelligence
More stored context does not guarantee better answers. Organization, filtering, and the ability to update old knowledge matter too. LongMemEval, a benchmark for long-term interactive memory, examines challenges including knowledge updates and reasoning about when something was true. These are useful tests for a collection that changes over time.
More notes can mean more useful evidence. They can also mean more contradiction, duplication, and noise to search past. An append-only source history is valuable. The problem is treating every historical statement as equally current when answering today's question.
A collection's value depends on the job it serves, not just its size. A library needs a catalog that helps readers find a book. A maintained reference collection also needs decisions about what gets kept, merged, or promoted to the reference shelf. Search and curation support different parts of that work.
The idea is not ours. Curation already appears in work on context curation and reflective memory management. Tiered memory architectures promote records between levels; graph-based approaches can supersede facts without deleting their history. Our focus is the operating discipline around those mechanisms: a promotion ladder, a revision log, preserved sources, and a person accountable for the review rules. That is the approach we use.
The ladder, worked
In this approach, memory gets promoted up a ladder. Raw at the bottom, distilled at the top. Nothing is auto-extracted. A curator makes each promotion on purpose, and every level points back down to where it came from. This is one way to handle knowledge that needs a traceable review history, not a requirement for every search tool.
- L0, transcript. The raw session. Immutable, append-only, never edited after capture. Evidence of what was said, not proof that every statement was correct.
- L1, atom. One extracted idea, one claim, with a pointer that resolves back to the exact line in L0 it came from.
- L2, principle. The generalized rule the atoms taught, named by concept, revised in place as it learns.
- L3, doctrine. The rule that governs the principles. Same machinery, one level up.
Abstract, that is just a diagram. So watch one real lesson climb it.
L0, the raw line. In a debugging session, the agent logs:
[00:22] Confirmed: verifier clock was 45s BEHIND the issuer. Freshly issued tokens carried an nbf the verifier had not reached yet, so auth failed for ~45s after issuance.This lives inL0-transcripts/2026-01-05-auth-debug.md, at line anchor#L2. It is never touched again.L1, the atom. The librarian extracts one idea and files it as
L1-0001-jwt-clock-skew.md: "JWT verification failed intermittently because the verifying service's clock ran behind the issuer's, so a token'snbfsat in the verifier's future at the moment it arrived. A bounded validation leeway absorbs real-world clock skew." Its frontmatter carriessource_ref: L0-2026-01-05-auth-debug#L2, a drill-down pointer. You can always fall from the atom back to the verbatim line. No "extracted, source unknown."L2, the principle, v1. The librarian promotes the atom into a principle file named by concept, not by date:
validate-time-sensitive-tokens-with-leeway.md. Version 1 reads: "JWTs need a bounded clock-skew leeway on validation." Itssources:list points down toL1-0001.
For JWT validation, RFC 7519 explicitly permits implementers "some small leeway, usually no more than a few minutes," to account for clock skew.
Now the part that separates a librarian from a note-taker. Two weeks later a different failure, a signed URL rejected for the same clock-skew reason, produces a new atom, L1-0007. The note-taker writes a second principle: signed URLs need leeway too. Now two rules say almost the same thing, and next quarter there will be five.
In this process, the curator finds the principle that already covers the issue and revises that one file in place:
L2, the principle, v2.
validate-time-sensitive-tokens-with-leeway.mdis rewritten toward generality: "Any time-sensitive credential, whether JWT, signed URL, OTP or nonce, must be validated with a bounded leeway that absorbs real clock skew in either direction between the issuing and verifying clocks. The skew must also be monitored, because leeway hides drift until it exceeds the window."revision_countticks to 2. A revision log records what changed and why.L1-0007folds into thesources:list.
One file became more general. The file count did not grow. The rule for the maintained reference layer is revise in place, do not append a duplicate. revision_count: 2 and a git log show that the same file was edited. They do not prove the revised rule is correct: the broader credential claim still needs review against each protocol's requirements.
And the atom that got folded up is not deleted. It is marked retired, its pointer now aims at the principle that absorbed it, and a tombstone is written so any old reference still resolves to where the content went. That is the second rule: never auto-delete history. The curator's only deletions are a status flip and a tombstone.
Neither rule is our invention, and pretending otherwise would be the kind of claim this ladder exists to catch. Revise-in-place exists as an update operation in the extraction tools. Non-destructive supersession exists too: at least one graph-based memory service already closes a superseded fact's validity window with an invalidation timestamp instead of deleting the fact, so point-in-time queries still work. The mechanisms are prior art.
The operating question remains: who decides which fact was superseded, by what, and why? A tool may make that decision easier or automate a low-risk part of it. The review responsibility still needs to be clear.
Where this leaves you
Start with the job. If people need to find a current policy or an equipment record, a search tool or knowledge base may be enough. If an agent needs to apply lessons from changing records, add review rules for what becomes a trusted principle. Neither job has to automate an entire operation to be worthwhile.
If your agent keeps repeating a corrected mistake, open its memory and ask: can we find the source, identify the current rule, and see who approved the change? The answer will tell you whether the gap is retrieval, data quality, curation, or some combination. Fix that gap rather than adding a process you do not need.
We use this custodial approach for agent memory. If it fits your work, we can discuss the curation rules, human review points, and support covered by an agreement. A starting plan can help identify questions to explore. It is not a validated build specification or a final quote.