Short answer
Because it was deleted, not mislaid. A benchmark published 30 September 2026 found every failure under bounded recency memory was an eviction, none a ranking error. Better search cannot surface a fact the memory already dropped, so what a team keeps has to live somewhere nothing ages out of.
Every team that has wired an assistant into its own knowledge has had the same bad afternoon. You ask it something the team definitely decided three months ago, it answers confidently and wrongly, and you go looking for the reason. Almost everyone looks in the same place first: the search. Better embeddings, a reranker, a bigger k, chunking the documents differently. A paper published on 30 September 2026 suggests that for a large class of these failures you are tuning the half that was never broken. The thing that lost the fact was not the ranking. It was the deletion that happened weeks earlier, quietly, under a policy nobody chose on purpose. Baalda is built on the opposite assumption: the store is plain markdown files on your own disk, and nothing in it ages out.
What did the retention and retrieval paper actually measure?
What Should an Agent Remember? Disentangling Retention from Retrieval in Bounded-Memory Evaluation is a single-author preprint by Juli Huang, submitted to arXiv on 30 September 2026 as 2610.00366. The code is public at TheClassicTechno/MemoryLLMAgentProject.
The argument is methodological, and it is the kind of argument that turns out to be practical. A persistent agent makes two different decisions. It decides what to retain as information arrives, which it has to do online, before it knows what anyone will ask. Then it decides what to surface once a query shows up, which is query-aware and happens later. Most memory evaluations compare systems that differ in both at once, so the number at the end does not tell you which decision earned it.
So the paper builds a streaming-recall benchmark that crosses the two. Five retention rules (recency, random, frequency, type-oracle, unbounded) against five selection rules (random, recency, query-similarity, dense, oracle), every condition run over the same 300 seeded episodes. The bounded setting keeps K equals 5 items out of a stream of 10 facts, and exposes k equals 3 of them to selection.
Two results from the abstract, both the author's own:
- Holding access fixed, query-aware selection is worth 15.5 percentage points of required-fact recall (95% confidence interval 12.8 to 18.2). That is the honest value of better retrieval.
- A mixed comparison that also changes history access reports a 68.7-point advantage, of which 53.2 points are attributable to access. Three quarters of the headline was the system being allowed to see more, not being better at looking.
And then the finding that matters most once you stop reading it as a benchmark result:
- Under bounded recency retention, all 319 observed failures are evictions. Displacement and ranking errors account for zero of them.
- Recall runs 100% to 58.5% to 0.0% across delays of 0 to 2, 3 to 5, and 6 to 9 facts. The older the fact, the more certainly it is simply gone.
Why can better retrieval not fix an eviction?
Because retention sets a ceiling that selection cannot exceed. This sounds obvious written down, and it is routinely ignored in practice, because the two failures look identical from the outside. The assistant says it does not know. You cannot see which thing happened.
The paper makes the ceiling visible by measuring both halves separately, and the shape it finds is stark: under bounded retention, query-aware, dense and oracle selection all converge on the same number. Oracle selection means a selector that cannot be beaten, a cheat. When the cheat and the ordinary ranker score the same, there is nothing left to win in ranking. Everything that was lost was lost before the query arrived.
This is the same gap we wrote about when OpenAI shipped agent memory nobody can open, approached from a different side. There the problem was that you could not inspect or correct an individual memory. Here the problem is one step earlier: you cannot even establish whether there was a memory to correct.
One caveat, said plainly, because the numbers above are the sort that get quoted loose. This is a seeded synthetic benchmark, with a 10-fact stream and a budget of 5, by one author, as a preprint with no stated venue and no peer review as of 7 October 2026. It is evidence about how memory evaluations should be designed. It is not a measurement of any shipped product, and "recall falls to 0%" describes a recency-eviction rule under a tight budget, not your assistant.
How would you tell the two apart in your own team's notes?
Here is the test, and it is worth running against whatever you currently use. Pick a decision your team made two months ago that the assistant should know. Ask it. When it gets the answer wrong, answer this question: is the fact in the store at all?
If you can answer that in under a minute, your knowledge system reports retention and selection separately, which is exactly what the paper's closing recommendation asks for: hold access fixed across compared methods, and report the retention rate beside every recall number. If you cannot answer it, you are about to spend a quarter tuning a reranker over a hole.
Most AI memory products cannot answer it, for a structural reason rather than a careless one. The memory is derived. It is summaries of conversations, embeddings in an index, extracted entities, consolidated after the fact by a model that made its own choices about what was worth keeping. There is no canonical list of what is in there, because the thing in there is not a list of anything, and a vector index returns the nearest rows to your query whether or not the right row exists.
What does a store with no eviction policy look like?
In Baalda the two halves the paper says to measure separately are two different tools, and they behave differently on purpose. An agent reaches them over the Model Context Protocol, against the same markdown files a person has open; how MCP works for notes covers the protocol itself.
list_notes is the retention half. It takes a vault id and an optional folder id. It takes no query and no k. It returns every note you can read, ordered by path:
[
{ "docId": "note_7f2a", "title": "Auth provider",
"relPath": "decisions/auth-provider.md",
"permission": "edit", "updatedAt": "2026-08-12T09:41:08Z" },
{ "docId": "note_91c4", "title": "Deploy runbook",
"relPath": "runbooks/deploy.md",
"permission": "view", "updatedAt": "2026-10-02T14:22:51Z" }
]That is a complete enumeration, not a ranked sample. The only things filtered out are notes someone deleted and notes you do not have permission to see, and the second one is reported rather than hidden, because permission comes back on every row. Behind it is a Postgres query with no limit clause and no recency term, sorted by rel_path. There is no retention budget, no TTL, no consolidation pass, no summarise-and-discard. A note leaves the vault when a person or an agent calls delete_note, which is an action with an author, and not otherwise.
search_notes is the selection half, and it is bounded, which is the honest part. It takes a query, runs semantic and keyword scoring together over notes and the extracted text of files, and returns at most k hits, default 10 and capped at 50. Each hit carries docId, title, relPath, score and kind.
So when the assistant gets your two-month-old decision wrong, you have a one-minute diagnostic that most memory architectures structurally cannot offer. Call list_notes on the folder, or open the folder in Finder, because it is a folder of .md files. The note is there or it is not. If it is there, you have a selection problem, and a reranker or a better-worded query is the right place to spend. If it is not there, no amount of retrieval work will ever reach it, and the fix is that somebody has to write it down. The paper's point, as a Tuesday afternoon.
Where does an unbounded vault not help?
In several places, and they are worth being direct about.
It does not decide what is worth keeping. Removing the eviction policy does not remove the retention decision, it moves it onto a person. If nobody writes the note, list_notes returns a complete and perfectly accurate listing of nothing useful. The paper's unbounded condition wins because retention was free in a benchmark; in a company, retention costs somebody twenty minutes of writing. That cost is the actual problem and no tool removes it.
Unbounded retention creates its own failure mode. Keep everything and selection has more to sort through, which is the trade the paper's bounded conditions exist to study. Search is genuinely capped at 50 results. A vault where every meeting got an auto-filed summary is a vault where search returns fifty plausible documents and the right one is forty-first. Somebody has to prune, and nothing here does that for you.
None of this touches the context window. A vault is storage. It does not compact a transcript, decide what to evict from a live run, or cut your token bill. The Context Language Models work is about what a model does with its working memory inside a single run, and that is a different problem with different answers.
And if you are one person. Obsidian is a local markdown vault with no eviction policy either, it is excellent, and it costs nothing. The real-time multi-user editing and the per-file permissions are the part you are paying for here, and if nobody is handing anything over to anybody, you are paying for machinery you will not use.
What should a team check this week?
One thing, and it takes an hour.
Take the five questions your team actually asks an assistant about its own work. The architecture decision and why the other option lost. What the on-call runbook says for the thing that broke in August. What the client agreed to in the kickoff. Ask all five. For every wrong answer, go and establish whether the fact exists in writing anywhere the assistant can reach.
Our guess, from watching teams do this, is that most of the failures are retention and almost none are ranking. Which means the work is not a better retrieval stack. It is the unglamorous business of a team writing down what it decided, somewhere durable, in a format a person and a model can both open. A team second brain is the name for that place. The research is just the part that tells you to stop tuning the search first.
