AI + MCP

Give an AI agent memory that a person can open and correct

By Baalda Team · · 10 min read

Short answer

Store it outside the model and feed it back each run. The usual answer is a memory framework that extracts facts into its own database. For a team, make the memory a markdown note instead: the agent amends it with anchored edits over MCP, and a person can open the same file and fix a wrong line.

A model keeps nothing between calls, so giving an AI agent memory means writing what it learns somewhere outside the model and feeding the relevant part back in on the next run. Every guide that ranks for this question picks the store for you: a vector database, a temporal knowledge graph, a managed memory service. What none of them puts in front of you is the decision that matters the moment a second person is involved, which is whether the thing the agent writes is something a human can open, read and correct. That is the whole subject of this post, and it is the problem Baalda is built for.

What does giving an agent memory actually involve?

Three moving parts, and they are the same three whatever product you use.

  • A store. Somewhere outside the model that survives the end of the run.
  • A write path. Something that decides what goes in, either the agent calling a tool deliberately or a background process extracting facts from the transcript.
  • A read path. Something that selects the relevant slice and puts it in the prompt before the model runs.

The read path is where most of the writing on this topic lives, because it is the part that looks like an engineering problem: embeddings, ranking, recency weighting, graph traversal. The write path gets much less attention and it is the one that decides whether the memory is usable by anybody other than the agent that wrote it.

Here is the split that organises the rest of this post. There are two honest answers to "how do I give an agent memory", and they are not competing implementations of one idea, they are different products:

  • Memory as inferred state. The system watches the conversation and writes down what it concluded. Good at personalisation, good at "remember that I prefer aisle seats", and the record it keeps is a derived artefact in a schema the vendor chose.
  • Memory as a document. The agent writes into a file, on purpose, in prose. Good at decisions, runbooks and the things a team has to agree on, and the record it keeps is something a person can read without an API.

Most teams end up needing both. Almost everything written about this question only describes the first.

What do the memory frameworks actually store?

Not the conversation. Extracted records, in a store shaped for retrieval.

Mem0's own documentation is direct about it: "By default, Mem0 stores extracted memories, not a verbatim transcript." The example it gives is the input "I prefer aisle seats" becoming the stored memory "User prefers aisle seats", written by an LLM that "extracts preferences, decisions, plans, and other details your agent can reuse", across three backends, a SQL database holding "Facts and metadata" as "The source of truth for each memory", a vector database holding embeddings, and an entity store.

Amazon's Bedrock AgentCore Memory works the same way at a larger scale. Its documentation states that "Long-term memory generation is an asynchronous process that runs in the background and automatically extracts insights after raw conversation/context is stored in short-term memory", and names the two operations it performs: Extraction, then Consolidation, which "Consolidates newly extracted information with existing information". What comes out is a memory record, reached through GetMemoryRecord, ListMemoryRecords or a semantic RetrieveMemoryRecords.

Zep, which ranks first for this query, proposes a different structure and the same shape of answer. Its page argues for a bi-temporal knowledge graph over a vector store, where "Each fact records where it came from (provenance) and when it was valid", so the agent can ask what is true now and what was true then. It is a genuinely better data model for a lot of problems, and worth reading. It also does not mention a second person anywhere, because the unit it is built around is one user's graph.

LangChain's store is the least opinionated of the four and the most revealing. Its long-term memory documentation describes store.put(namespace, key, data) with a namespace "similar to a folder" and a key "like a file name", holding JSON documents, and notes that "Namespaces often include user or org IDs or other labels that makes it easier to organize information". It gives you a folder and a filename and then stores JSON in a Postgres table, which is a fair summary of the whole category.

What it writesWho writes itCan a person read it without an API?
Mem0Extracted facts plus entitiesAn LLM, automaticallyNo
AgentCore MemoryMemory records from background extractionThe service, asynchronouslyNo
ZepFacts in a bi-temporal graphThe memory layer, on ingestNo
LangChain storeJSON under a namespace and keyYour codeOnly by querying the store
A markdown vaultProse in a noteThe agent, by calling a toolYes, it is a file

None of those four is doing anything wrong. They are answering the question for one agent serving one user, which is the question most people asking it actually have.

Why does that break the moment a second person needs it?

Because the memory becomes a second copy of what the team knows, in a format nobody on the team can open.

Watch what happens over three months. Your agent has been working with the team on a deployment process. It has accumulated memory: that the staging database is restored from a snapshot and not seeded, that one service has to be deployed before another, that the person who knows the billing migration left in July. All of it is real, all of it is correct, and all of it is in a vector store behind an API.

Now a teammate asks "how do deploys actually work here?" and the answer is in a place they cannot look. They ask the agent instead, which is fine until the agent is wrong, and then they are debugging a memory record by having a conversation with the thing that wrote it. The failure the team hits is not retrieval quality. Agent memory fails by throwing things away rather than by searching badly covers that half. This one is simpler: the knowledge is real, it is retrievable, and it is unreadable to everyone except one process.

And correction is worse than reading. Mem0 is honest about the shape of the problem, noting that "The automatic extraction path is additive" and that you should "Use explicit update or delete operations when your application needs to correct or remove a memory". That is an engineering task, not something the person who noticed the error can do. OpenAI's dots have the sharper version of it, covered in whose disk shared agent memory sits on: you cannot modify an individual memory at all, only reset the agent.

So the question underneath "how do I give my agent memory" turns out to be a question about ownership. If the memory is worth keeping for three months, it is team knowledge. Team knowledge that only one program can read is not knowledge the team has.

How do you keep memory in a file without the file rotting?

By letting the agent change one line instead of rewriting the document.

This is the real objection to file-based memory and it deserves a straight answer, because "just have the agent write markdown" fails in a predictable way. The agent reads a 400-line note, decides one sentence is now wrong, and writes the whole file back. Every time it does that it re-renders the other 399 lines from its own understanding, so formatting drifts, a teammate's edit from an hour ago disappears, and the diff is useless. Do that fifty times and the file is agent output, not a document.

Baalda is a team second brain: plain markdown files on your own disk, several people editing the same note at once in real time, and an AI reading and writing those same files over MCP. The tool that makes it work as memory is edit_note, which takes targeted edits rather than a body. Each edit names an anchor, a piece of exact text, and the server refuses the whole call if that anchor is not unique:

json
{
  "docId": "note_deploy_runbook",
  "expectedRevision": "9f2c…",
  "edits": [
    {
      "type": "replace",
      "find": "Staging is seeded from fixtures.",
      "replace": "Staging is restored from last night's snapshot, not seeded."
    },
    {
      "type": "insert_after",
      "anchor": "## Deploy order",
      "text": "\n- Billing must go out before the gateway."
    }
  ]
}

The server's own comment on that code says it is "Strict on purpose: an anchor must occur EXACTLY once unless the edit says all." If the anchor is missing, or matches twice, the call fails with anchor_not_found or anchor_ambiguous and nothing is written. Not the first edit, not a partial document, nothing. An agent working from a stale reading of the file cannot half-apply a change to it.

The second half is the precondition. read_note returns a revision, which is the sha256 of the note's text at the moment it was read, and passing it back as expectedRevision refuses the write if the note changed in between. The comment in the writer explains why it is a hash rather than a counter: the note is a CRDT, so there "has no single version number; the content hash is the honest equivalent". The effect for memory is the one you want. If a teammate corrected that line while the agent was thinking, the agent's amendment fails loudly rather than reverting them.

There is a third piece that is less obvious and does more work than it looks. Anchors come back to the server as text an LLM copied out of JSON, and that copy is often not byte-exact, so the matcher has a fallback fold for exactly that. The source comment lists what it absorbs: a model "rewrites a non-breaking or ideographic space as a plain space, drops zero-width characters it cannot see, straightens curly quotes, turns an en/em dash into a hyphen, collapses doubled spaces, loses the trailing spaces before a newline, or hands back NFC where the note holds NFD", and those cases "were the bulk of edit_note's failures". Exact matching still runs first and wins, the fold only engages when the exact anchor is absent, the exactly-once rule applies to the folded text too, and the fold never touches the note itself. That is the difference between a file-based memory that works in a demo and one that works on the fiftieth amendment.

What you get at the end of it is a .md file in a folder. A teammate opens it in the editor, sees the agent's correction in context, fixes a word, and the agent reads the fixed version on its next call because both of them are editing the same document. The memory is not a copy of the team's knowledge. It is the team's knowledge, with an agent as one more contributor to it. Connecting an AI to your notes over MCP covers the endpoint itself, and what breaks when a second person joins a Claude Code second brain covers the concurrency in more detail.

Where a markdown vault is the wrong answer

It does not do the first kind of memory at all, and that is a design decision rather than a gap waiting to be filled.

There is no extraction. Nothing watches a transcript and writes down what it concluded. If the agent does not call create_note, append_note, update_note or edit_note, nothing is remembered, which means memory depends on the agent being instructed to keep it. There is no consolidation pass merging today's note with last month's, no decay, no recency weighting, no TTL, and no per-user personalisation layer. If what you want is an assistant that quietly learns that this user prefers short answers, AgentCore and Mem0 do that and a vault does not.

Retrieval is bounded too. search_notes runs semantic and keyword scoring together across notes and the text extracted from the files beside them, and returns at most k results, default 10 and capped at 50. A vault with ten thousand notes in it can still bury the right one below the cut, and no amount of file ownership fixes a ranking problem.

And prose is the wrong container for some things. Structured state an agent updates on every turn, a task queue, a counter, a session's working context: those belong in a database, and writing them into markdown would be an abuse of both. The vault is for the things that are worth a paragraph and worth keeping after the agent that wrote them is switched off.

So which one do you actually need?

Split it by how long the memory has to live and who else needs it.

Personalisation and session state, the things that are true about one person's use of one agent, belong in whatever memory layer your framework already ships. Nobody else needs to read them, nobody is going to correct them by hand, and the extraction that writes them is doing real work you would otherwise do yourself.

Anything a new teammate would need in their first week belongs in a file. Decisions and the reasons behind them, runbooks, why that service is named after a fish, the thing you only find out by breaking it. If an agent learns one of those, the useful place to put it is the same document a person would have written, in a vault the team already reads, amended one line at a time so that the file survives being written by a machine.

The test is simple enough to apply in the moment. When the agent remembers something, ask who else on your team would want to know it. If the answer is anyone, it should not be in a store only the agent can open.

FAQ

Frequently asked questions

Do I still need a vector database if the agent writes to notes?

Usually yes, for a different job. A memory framework such as Mem0 or Bedrock AgentCore Memory extracts personalisation and session state automatically, which a vault does not do at all. The vault holds the things a teammate would also want to read: decisions, runbooks, the reason a service is built the way it is. Run both and give each the kind of memory it is shaped for.

What stops an agent from overwriting a teammate's edit?

Two things, and they compose. `read_note` returns a `revision`, the sha256 of the note's text, and passing it back as `expectedRevision` refuses the write if the note changed in between. On top of that, `edit_note` requires every anchor to match exactly once, so a missing or ambiguous anchor fails the whole call with nothing written rather than applying part of it. An agent working from a stale read fails loudly instead of quietly reverting someone.

Does the agent decide for itself what to remember?

No, and that is the honest limit. Nothing in a Baalda vault watches a transcript and extracts facts from it. A note exists because something called `create_note`, `append_note`, `update_note` or `edit_note`, so what gets remembered depends on the agent being told to keep it. The upside is that nothing is written that nobody asked for. The cost is that you have to ask.

Can an agent remember something a teammate is not allowed to read?

It can hold it in a note that teammate cannot open. An MCP token is minted per person and scoped to one vault, and access is set per folder and per note, so a note the person cannot read is a note the agent acting for them cannot reach either. What it does not do is partition the agent's own working context: anything you paste into a session is in that session.

How many results does the agent get back when it searches its memory?

At most `k`, which defaults to 10 and is capped at 50. `search_notes` runs semantic and keyword scoring together across notes and the text extracted from the files beside them, each hit tagged note or file. Retention in a vault is unbounded, because nothing ages out, but selection is not, so a large vault can still bury the right note below the cut.

Start your team’s brain

Free and open source. No account needed.