Short answer
In a file the team keeps. A UW and Meta paper published 29 September 2026 found a model manages its context better as a file it rewrites than as an append-only transcript. That file is scratch and dies with the run, so the durable version has to be a markdown note on your own disk.
There are two ways to give a model something to work from. One is a transcript: everything that happened, in order, appended forever, and when it gets too long you summarise it and throw the original away. The other is a file: a working document the model rewrites in place, where a stale search result can be deleted and a plan can be corrected rather than contradicted eight turns later. A paper published on 29 September 2026 put a number on which one works better, and the file won by a wide margin. The uncomfortable part is that the winning file is scratch. It belongs to one run, on one machine, and the paper never claims it survives the run ending. Baalda exists because the durable version of that file is the thing a team actually runs on, and it should be a markdown note at a path somebody can open.
What did the Context Language Models paper actually do?
Context Language Models comes from Rulin Shao and colleagues at the University of Washington and Meta Superintelligence Labs, with co-authors including Mike Lewis, Wen-tau Yih, Luke Zettlemoyer and Pang Wei Koh. It was submitted to arXiv on 29 September 2026 as 2609.37725.
The move is small to describe. Instead of a harness deciding what stays in the context window, the context is handed to the model as a file. In the authors' words, the model "can edit this file using general Bash commands, just as it would edit other files in storage". Edits to the file are synchronised back into the live context. So the model deletes a dead branch of the search, rewrites its own plan, collapses twenty tool calls into two lines, keeps a scoreboard at the top and updates it in place.
The reported results, all from the paper's abstract:
- 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, against state-of-the-art context management strategies.
- 5% higher scores with 59% fewer FLOPs on a 12-hour EdgeBench run.
- 65% greater improvement at the same compute on a 24-hour multi-repository agent-swarm task.
- Steering with natural-language instructions improved held-out accuracy by up to 35.9 points on a context-management task.
- An online reinforcement learning method improved Qwen3.5-9B on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs.
Those are the authors' own figures, published alongside the paper, and as of 6 October 2026 nobody outside the group has reproduced them. Treat them as claims. The direction is still the interesting bit, and the direction is not in dispute by anyone who has run an agent for more than an hour.
Why does a rewritable file beat an append-only transcript?
Because an append-only log cannot be wrong in a recoverable way. Once a bad assumption is in the transcript it stays in the transcript, and every later turn reads it. Summarisation is the usual patch, and summarisation is lossy in the one direction that hurts: it compresses detail and keeps narrative, when what a long task needs is the opposite.
A file does not have that problem. The wrong line gets replaced by the right line. The state of the work is wherever the file says it is, not distributed across four hundred turns that have to be re-read to be understood. This is the same reason teams keep a decision log rather than a chat archive, and the paper is, in effect, a measurement of that intuition at the scale of one model and one run.
There is a cost, and the paper is straight about it. Editing the middle of a context breaks prefix caching, which is why the authors co-design something they call Suffix Cache Reuse, reporting a further 35% reduction in server-side compute relative to standard SGLang at matched performance. Rewriting in place is not free. It is just cheaper than carrying everything.
What happens to the context file when the run ends?
The paper does not say. It describes edits being synchronised with the live context during execution, and it does not discuss what becomes of the file afterwards, because that is not the question it set out to answer. For the benchmark, it does not matter. For a team, it is the entire question.
Think about what is in that file by the end of a 24-hour agent-swarm run. The plan as it ended up, not as it started. Which approaches were tried and dropped, and why. A scoreboard of what each sub-agent was doing. That is a better artefact than most teams produce by hand, and under the arrangement the paper describes it is a scratch file on a machine nobody logs into, holding the only honest account of a day's work.
This is the same gap we wrote about when OpenAI shipped agent memory you cannot open, arriving from the opposite direction. There, the memory was a product feature with no edit button. Here, the memory is a plain file with no owner. Either way the correction a human wants to make has nowhere to land.
Can an agent edit a shared note the same way?
Yes, and this is the mechanism rather than a slogan. Baalda is a team second brain built on plain markdown files on your own disk, which several people edit at once in real time. It serves the Model Context Protocol, so an agent reaches the same files a person has open. How MCP works for notes covers the protocol itself; what matters here is the write surface.
edit_note is the same four operations the paper's models reach for, exposed as an MCP tool against a named note:
{
"docId": "note_7f2a",
"expectedRevision": "rev_41",
"edits": [
{ "type": "replace", "find": "Provider: Auth0 (decided 12 Aug)",
"replace": "Provider: Clerk (decided 2 Oct, Auth0 pricing)" },
{ "type": "delete", "find": "- [ ] benchmark Auth0 SSO latency\n" },
{ "type": "insert_after", "anchor": "## Open questions",
"text": "\n- Does Clerk cover our SAML requirement?" }
]
}Replace, delete, insert before, insert after. The same in-place rewrite, pointed at decisions/auth-provider.md instead of at a scratch buffer. Three things follow from the file being real rather than temporary:
- An anchor that does not match exactly once refuses the whole call and writes nothing. No partial edit, no guessed position. If the agent's picture of the note is stale, it finds out before it changes anything.
- `read_note` returns a `revision`, and passing it back as `expectedRevision` refuses the write if the note moved in between. A teammate editing the same paragraph is not silently overwritten.
- The note has a path, a history and permissions. Someone who was not in the run can open it, see what changed, and fix the line that is wrong. Access is set per folder and per note, so the agent reaches exactly what the person whose token it carries can reach.
Baalda is open source under Apache-2.0 and self-hostable, and it is free to run locally, with a managed Team plan for hosted sync. The files stay markdown on disk whichever way you run it, which is the part that makes the vault a durable answer rather than another place the knowledge can get stuck.
Does this replace context engineering?
No, and it is worth being plain about that, because the two things are easy to blur.
A vault does not manage your context window. It will not compact a transcript, decide what to evict, pick what to re-read, or cut your token bill. Every gain in the Context Language Models paper happens inside a single run, in the model's working memory, and a notes application has nothing to say about it. If your problem is that a long agent run is expensive and drifts, the answer is in that paper, or in the compaction your harness already ships, not in a second brain.
Two more honest limits. A note is only as good as the agent's judgement about what was worth writing down, and an agent that files a summary after every task produces a vault nobody reads, which is worse than no vault. Someone has to prune. And if your team is one person who never hands anything over, Obsidian is a better fit and costs nothing; the multiplayer and permissions machinery is overhead you are not using.
What should a team of ten change on Monday?
Not much, which is the point. Pick the three or four documents that an agent should be allowed to correct rather than append to: the architecture decisions, the runbook, the open-questions list for whatever is in flight. Put them in a folder the whole team can open. Give the agent a token scoped to that folder and let it use targeted edits, not appends, so the document stays the current state of the work rather than a log of every time something was considered.
Then read it on Friday. The paper's finding is that a model does better work when its context is a document it maintains instead of a history it carries. The same has been true of teams for a long time. The only new part is that the document can now have more than one kind of author, and a team second brain is what you call the place where that is allowed.
