Short answer
Gemini CLI keeps what it knows in GEMINI.md, which it concatenates into every prompt and stores on one machine. Connecting it to a Baalda vault over MCP moves the team's notes behind search_notes and read_note, so they are fetched when needed rather than carried on every request, and written back through anchored edits.
Gemini CLI has exactly one place to keep what it knows about your work, and it is a markdown file called GEMINI.md. That is a better default than a hidden vector store, because you can open it, read it and fix it. It has two properties that stop working the moment a second person is involved: it is loaded in full into every request, and it sits on one machine.
Baalda, the app this site is for, is a team second brain: plain markdown files on your own disk, edited together in real time, with an MCP endpoint over the same files. The useful thing it does for Gemini CLI is not "more memory". It is moving the team's written context from something the model carries to something the model asks for.
What does Gemini CLI actually remember between sessions?
The context files, and nothing else you did not write down.
Google's documentation is direct about what they are: "Context files, which use the default name GEMINI.md, are a powerful feature for providing instructional context to the Gemini model." The CLI looks for them in three places, in order:
- Global, at
~/.gemini/GEMINI.md, "in your user home directory". - Workspace, where "The CLI searches for
GEMINI.mdfiles in your configured workspace directories and their parent directories". - Just-in-time, where "When a tool accesses a file or directory, the CLI automatically scans for
GEMINI.mdfiles in that directory and its ancestors".
Then the important sentence, which is the one most write-ups skip: the CLI "loads various context files from several locations, concatenates the contents of all found files, and sends them to the model with every prompt."
Every prompt. Not when relevant, not when retrieved. The whole concatenation rides along on each request, whether you asked about the deployment runbook or asked it to rename a variable. /memory show prints the result, which is worth running once before you believe any estimate of how big yours is.
Why is that a problem for a team rather than for you?
Because the two obvious ways to scale it both fail, in different directions.
The first is to write more down. This works, right up until the file is large, and then you are paying for all of it on every request. Gemini CLI's free tier is "60 requests/min and 1,000 requests/day with personal Google account" against "Gemini 3 models with 1M token context window", so you will not hit a wall quickly. You will just carry a document about the billing service into every conversation about CSS. Context you always send is context you cannot target.
The second is to share it. A GEMINI.md committed to a repo is genuinely shared, and for repo-shaped rules that is the right answer. But the things that make an experienced teammate fast are mostly not repo-shaped: why the vendor was chosen, what broke last quarter, which customer cannot be migrated yet, what was decided in the meeting nobody wrote up. That material does not belong in a code checkout, and the file where it usually ends up is ~/.gemini/GEMINI.md, which is in a home directory and is in no repository at all.
So the knowledge is real, it is written down, it is in markdown, and the second person still does not have it. This is the same shape covered in Claude Code's second brain is a folder, not an integration and in how to write an AGENTS.md file: an unconditionally loaded instruction file is a budget every teammate pays on every run, and it is the wrong container for a growing body of shared knowledge.
How do you connect Gemini CLI to a Baalda vault over MCP?
Gemini CLI is an MCP client, so the team's notes can sit behind a tool call instead of inside the prompt.
Mint a token first: in the Baalda desktop app, Vault settings → MCP, and create one. Tokens are per person and scoped to one vault.
Then register the server. The CLI has a command for it:
gemini mcp add --transport http baalda https://api.baalda.com/api/mcp \
--header "Authorization: Bearer mcp_…"Or write it into settings.json yourself, at ~/.gemini/settings.json for every project or .gemini/settings.json for one:
{
"mcpServers": {
"baalda": {
"httpUrl": "https://api.baalda.com/api/mcp",
"headers": {
"Authorization": "Bearer mcp_…"
}
}
}
}Use `httpUrl`, not `url`. This is the one detail that will cost you an afternoon. Gemini CLI's own configuration reference defines them as two different transports: url is the "SSE endpoint URL" and httpUrl is the "HTTP streaming endpoint URL". Baalda speaks the second one and deliberately does not speak the first. Its route file says so and the server behaves accordingly: POST /api/mcp is the endpoint, and GET/DELETE on the same path answer 405 Method Not Allowed, because there is no server-to-client SSE stream to open. Put the Baalda endpoint in url and the client will try to open a stream against something that answers 405.
Self-hosting changes one string. The endpoint is your server's URL plus /api/mcp, and because Gemini CLI runs on your machine it can reach a server on localhost perfectly well. That is a real difference from the hosted assistants: an agent running in someone else's datacentre cannot see http://localhost:3010/api/mcp, so a free local install is invisible to it. A terminal agent on your own machine does not have that problem.
Run /mcp inside the CLI to check. It "displays server list, connection status, server details, available tools, and discovery state", which is how you tell a bad token from a bad transport.
What changes once the notes are behind a tool call?
The cost model inverts, and the write path stops being destructive.
Reading becomes conditional. GEMINI.md is paid on every request. A tool call is paid when the model decides it needs one. search_notes runs semantic and keyword scoring together across the notes and the text extracted from the files beside them, and read_note pulls one back in full. A hundred decision records cost nothing on the request where they are irrelevant.
Writing becomes additive rather than replacing. This is the part that matters once two people share the vault. edit_note makes targeted edits at exact anchors, and its contract is strict on purpose: each anchor "must match exactly once", and "a missing or ambiguous anchor refuses the whole call with nothing written". On top of that, read_note returns a revision, and passing it back as expectedRevision refuses the write if the note changed in between. An agent working from a stale read fails loudly instead of quietly reverting a teammate, which is the failure mode behind why merging after the fact is too late.
Permissions come with it. The token is minted per person and scoped to one vault, and access is set per folder and per note. A note a person cannot open is a note the agent acting for them cannot reach.
And because the file on disk is the durable copy, the same note is a file your teammate's editor has open, with a live CRDT keeping the open copies in sync. If someone has the note open while Gemini CLI edits it, they watch it change.
There is also an OAuth path on both sides. Gemini CLI "supports OAuth 2.0 authentication for remote MCP servers using SSE or HTTP transports" with authProviderType defaulting to dynamic_discovery, and the Baalda server publishes both /.well-known/oauth-authorization-server and /.well-known/oauth-protected-resource. The token flow above is the one Baalda's README documents, so that is the one to start with.
Where does this not help?
Four places, and the first one is the one people get wrong.
It does not replace `GEMINI.md`, and it should not. Instructions that must apply unconditionally belong in a file that is loaded unconditionally. How to run the tests, which package manager, the house style, do not touch the migrations folder. That is exactly what a context file is for. What moves out is the growing body of reference material that is only needed sometimes. Keep the rules in GEMINI.md and point it at the vault in one line, so the model knows the notes exist.
Retrieval is capped, so a big vault can bury the right note. search_notes returns at most k results, which defaults to 10 and tops out at 50. Nothing ages out of a vault, so retention is unbounded while selection is not. That trade is better than the alternative, but it is a trade.
Nothing extracts memories on its own. A note exists because something called create_note, append_note, update_note or edit_note. Gemini CLI will not notice that a session contained a decision worth keeping and file it for you. You have to ask, or tell it to in GEMINI.md. The upside is that nothing is written that nobody asked for.
The token sits in a config file. Gemini CLI's documented environment-variable expansion covers the env block, which is a stdio concern, so a bearer header for a remote server is a value in settings.json. Treat the file accordingly. And think twice about trust: true on this server: it "bypasses all tool call confirmations for this server", which is a reasonable setting for a read-only tool and a less reasonable one for something that can write to your team's notes. includeTools is the narrower instrument if you want the agent to read the vault and not edit it.
Is this better than pointing Gemini CLI at a folder?
For one person on one machine, usually not, and it is worth saying plainly.
If your notes are a folder on the same disk the CLI is running on, the CLI already has file tools. It can read and write them with no server, no token and no MCP at all, and the just-in-time context loading means a GEMINI.md inside that folder gets picked up when a tool touches it. That is the cheapest correct answer and a lot of people should stop there.
MCP earns its keep at the point where the notes are not simply yours: when a second person edits the same files, when some folders are not for everyone, when you want the agent's reach to be the same as the person's reach rather than whatever the filesystem happens to allow. That is the line, and it is the same line described in what a team second brain is. Below it, a folder wins. Above it, a folder has no way to say no.
