# Gemini CLI loads GEMINI.md into every request. Your team's notes belong behind a tool call

> Gemini CLI concatenates every GEMINI.md into every prompt, so written context is a budget paid per request. How to put a team's notes behind MCP instead.

Source: https://baalda.com/blog/gemini-cli-second-brain
Site: Baalda, https://baalda.com
Published: 2026-10-11

**Short answer.** Gemini CLI keeps what it knows in GEMINI.md, which it concatenates into every prompt and stores on one machine. Connecting it to a Baalda vault over MCP moves the team's notes behind search_notes and read_note, so they are fetched when needed rather than carried on every request, and written back through anchored edits.

Gemini CLI has exactly one place to keep what it knows about your work, and it is a markdown file called `GEMINI.md`. That is a better default than a hidden vector store, because you can open it, read it and fix it. It has two properties that stop working the moment a second person is involved: it is loaded in full into every request, and it sits on one machine.

Baalda, the app this site is for, is a team second brain: plain markdown files on your own disk, edited together in real time, with an MCP endpoint over the same files. The useful thing it does for Gemini CLI is not "more memory". It is moving the team's written context from something the model carries to something the model asks for.

## What does Gemini CLI actually remember between sessions?

The context files, and nothing else you did not write down.

Google's documentation is direct about what they are: "Context files, which use the default name `GEMINI.md`, are a powerful feature for providing instructional context to the Gemini model." The CLI looks for them in three places, in order:

- **Global**, at `~/.gemini/GEMINI.md`, "in your user home directory".
- **Workspace**, where "The CLI searches for `GEMINI.md` files in your configured workspace directories and their parent directories".
- **Just-in-time**, where "When a tool accesses a file or directory, the CLI automatically scans for `GEMINI.md` files in that directory and its ancestors".

Then the important sentence, which is the one most write-ups skip: the CLI "loads various context files from several locations, concatenates the contents of all found files, and sends them to the model with every prompt."

Every prompt. Not when relevant, not when retrieved. The whole concatenation rides along on each request, whether you asked about the deployment runbook or asked it to rename a variable. `/memory show` prints the result, which is worth running once before you believe any estimate of how big yours is.

## Why is that a problem for a team rather than for you?

Because the two obvious ways to scale it both fail, in different directions.

The first is to write more down. This works, right up until the file is large, and then you are paying for all of it on every request. Gemini CLI's free tier is "60 requests/min and 1,000 requests/day with personal Google account" against "Gemini 3 models with 1M token context window", so you will not hit a wall quickly. You will just carry a document about the billing service into every conversation about CSS. Context you always send is context you cannot target.

The second is to share it. A `GEMINI.md` committed to a repo is genuinely shared, and for repo-shaped rules that is the right answer. But the things that make an experienced teammate fast are mostly not repo-shaped: why the vendor was chosen, what broke last quarter, which customer cannot be migrated yet, what was decided in the meeting nobody wrote up. That material does not belong in a code checkout, and the file where it usually ends up is `~/.gemini/GEMINI.md`, which is in a home directory and is in no repository at all.

So the knowledge is real, it is written down, it is in markdown, and the second person still does not have it. This is the same shape covered in [Claude Code's second brain is a folder, not an integration](/blog/claude-code-second-brain) and in [how to write an AGENTS.md file](/blog/how-to-write-agents-md): an unconditionally loaded instruction file is a budget every teammate pays on every run, and it is the wrong container for a growing body of shared knowledge.

## How do you connect Gemini CLI to a Baalda vault over MCP?

Gemini CLI is an MCP client, so the team's notes can sit behind a tool call instead of inside the prompt.

Mint a token first: in the Baalda desktop app, **Vault settings → MCP**, and create one. Tokens are per person and scoped to one vault.

Then register the server. The CLI has a command for it:

```bash
gemini mcp add --transport http baalda https://api.baalda.com/api/mcp \
  --header "Authorization: Bearer mcp_…"
```

Or write it into `settings.json` yourself, at `~/.gemini/settings.json` for every project or `.gemini/settings.json` for one:

```json
{
  "mcpServers": {
    "baalda": {
      "httpUrl": "https://api.baalda.com/api/mcp",
      "headers": {
        "Authorization": "Bearer mcp_…"
      }
    }
  }
}
```

**Use `httpUrl`, not `url`.** This is the one detail that will cost you an afternoon. Gemini CLI's own configuration reference defines them as two different transports: `url` is the "SSE endpoint URL" and `httpUrl` is the "HTTP streaming endpoint URL". Baalda speaks the second one and deliberately does not speak the first. Its route file says so and the server behaves accordingly: `POST /api/mcp` is the endpoint, and `GET`/`DELETE` on the same path answer `405 Method Not Allowed`, because there is no server-to-client SSE stream to open. Put the Baalda endpoint in `url` and the client will try to open a stream against something that answers 405.

Self-hosting changes one string. The endpoint is your server's URL plus `/api/mcp`, and because Gemini CLI runs on your machine it can reach a server on `localhost` perfectly well. That is a real difference from the hosted assistants: an agent running in someone else's datacentre cannot see `http://localhost:3010/api/mcp`, so a free local install is invisible to it. A terminal agent on your own machine does not have that problem.

Run `/mcp` inside the CLI to check. It "displays server list, connection status, server details, available tools, and discovery state", which is how you tell a bad token from a bad transport.

## What changes once the notes are behind a tool call?

The cost model inverts, and the write path stops being destructive.

**Reading becomes conditional.** `GEMINI.md` is paid on every request. A tool call is paid when the model decides it needs one. `search_notes` runs semantic and keyword scoring together across the notes and the text extracted from the files beside them, and `read_note` pulls one back in full. A hundred decision records cost nothing on the request where they are irrelevant.

**Writing becomes additive rather than replacing.** This is the part that matters once two people share the vault. `edit_note` makes targeted edits at exact anchors, and its contract is strict on purpose: each anchor "must match exactly once", and "a missing or ambiguous anchor refuses the whole call with nothing written". On top of that, `read_note` returns a `revision`, and passing it back as `expectedRevision` refuses the write if the note changed in between. An agent working from a stale read fails loudly instead of quietly reverting a teammate, which is the failure mode behind [why merging after the fact is too late](/blog/obsidian-sync-conflicts).

**Permissions come with it.** The token is minted per person and scoped to one vault, and access is set per folder and per note. A note a person cannot open is a note the agent acting for them cannot reach.

And because the file on disk is the durable copy, the same note is a file your teammate's editor has open, with a live CRDT keeping the open copies in sync. If someone has the note open while Gemini CLI edits it, they watch it change.

There is also an OAuth path on both sides. Gemini CLI "supports OAuth 2.0 authentication for remote MCP servers using SSE or HTTP transports" with `authProviderType` defaulting to `dynamic_discovery`, and the Baalda server publishes both `/.well-known/oauth-authorization-server` and `/.well-known/oauth-protected-resource`. The token flow above is the one Baalda's README documents, so that is the one to start with.

## Where does this not help?

Four places, and the first one is the one people get wrong.

**It does not replace `GEMINI.md`, and it should not.** Instructions that must apply unconditionally belong in a file that is loaded unconditionally. How to run the tests, which package manager, the house style, do not touch the migrations folder. That is exactly what a context file is for. What moves out is the growing body of reference material that is only needed sometimes. Keep the rules in `GEMINI.md` and point it at the vault in one line, so the model knows the notes exist.

**Retrieval is capped, so a big vault can bury the right note.** `search_notes` returns at most `k` results, which defaults to 10 and tops out at 50. Nothing ages out of a vault, so retention is unbounded while selection is not. That trade is better than the alternative, but it is a trade.

**Nothing extracts memories on its own.** A note exists because something called `create_note`, `append_note`, `update_note` or `edit_note`. Gemini CLI will not notice that a session contained a decision worth keeping and file it for you. You have to ask, or tell it to in `GEMINI.md`. The upside is that nothing is written that nobody asked for.

**The token sits in a config file.** Gemini CLI's documented environment-variable expansion covers the `env` block, which is a stdio concern, so a bearer header for a remote server is a value in `settings.json`. Treat the file accordingly. And think twice about `trust: true` on this server: it "bypasses all tool call confirmations for this server", which is a reasonable setting for a read-only tool and a less reasonable one for something that can write to your team's notes. `includeTools` is the narrower instrument if you want the agent to read the vault and not edit it.

## Is this better than pointing Gemini CLI at a folder?

For one person on one machine, usually not, and it is worth saying plainly.

If your notes are a folder on the same disk the CLI is running on, the CLI already has file tools. It can read and write them with no server, no token and no MCP at all, and the just-in-time context loading means a `GEMINI.md` inside that folder gets picked up when a tool touches it. That is the cheapest correct answer and a lot of people should stop there.

MCP earns its keep at the point where the notes are not simply yours: when a second person edits the same files, when some folders are not for everyone, when you want the agent's reach to be the same as the person's reach rather than whatever the filesystem happens to allow. That is the line, and it is the same line described in [what a team second brain is](/blog/team-second-brain). Below it, a folder wins. Above it, a folder has no way to say no.

## FAQ

### Should I use url or httpUrl for an MCP server in Gemini CLI?

For Baalda, `httpUrl`. Gemini CLI's configuration reference treats them as two different transports: `url` is the "SSE endpoint URL" and `httpUrl` is the "HTTP streaming endpoint URL". Baalda speaks the second one only. `POST /api/mcp` is the endpoint, and `GET` or `DELETE` on the same path answer `405 Method Not Allowed`, because the server offers no server-to-client SSE stream. Putting the endpoint in `url` makes the client attempt a stream against a path that refuses one, which shows up as a connection that never finishes discovering. Run `/mcp` inside the CLI to see the status and the tool list.

### Does connecting a vault mean I can delete GEMINI.md?

No, and you should not want to. Instructions that have to apply unconditionally belong in a file that is loaded unconditionally: how to run the tests, which package manager, the house style, which directories not to touch. That is what a context file is for. What moves out to the vault is the reference material that is only relevant sometimes, which is precisely the part that is expensive to carry on every request. The usual arrangement is a short GEMINI.md that includes one line telling the model the vault exists.

### Can Gemini CLI reach a Baalda server running on localhost?

Yes. Gemini CLI runs on your own machine, so a self-hosted server at `http://localhost:3010/api/mcp` is reachable in the same way any local service is. This is a genuine advantage of a terminal agent over a hosted assistant, which runs in someone else's datacentre and cannot see anything on your loopback interface. For a self-hosted deployment the endpoint is just your server's URL plus `/api/mcp`; on the managed service it is `https://api.baalda.com/api/mcp`.

### What stops the agent from overwriting a note a teammate just changed?

Two separate checks. `edit_note` makes targeted edits at exact anchors, and every anchor must match exactly once, so a missing or ambiguous anchor refuses the whole call with nothing written rather than applying part of it. Separately, `read_note` returns a `revision`, and passing it back as `expectedRevision` refuses the write if the note changed in between. An agent working from a stale read fails loudly instead of quietly reverting someone.

### Will Gemini CLI decide for itself what to save into the vault?

No. A note exists because something called `create_note`, `append_note`, `update_note` or `edit_note`. Nothing watches a session and extracts facts from it, so if a decision worth keeping comes up and nobody asks for it to be written, it is not written. The cost is that you have to ask, or put the instruction in GEMINI.md. The benefit is that nothing lands in the team's notes that nobody requested.
