Team second brain

Internal documentation: write the decisions, index the files you already have

By Baalda Team · · 8 min read

Short answer

Write the things only a person knows: the decision, the reason, the convention, the exception. Do not retype the spreadsheet, the contract or the exported report. Put those where one search reaches their words as well as your notes, under the same permissions, and keep the index tied to the bytes it describes.

Most advice about internal documentation is advice about writing. Pick an audience, use plain language, break it into sections, build templates, name an owner, set a review cadence. All of that is fine, and none of it touches the thing that decides whether anybody finds an answer, which is that a large part of what your team already knows is written down and sitting in a format nobody is ever going to retype.

The dividing line is not well-written docs against badly written ones. It is the documentation you author against the documentation you inherit. The first is notes somebody sat down and wrote. The second is the pricing spreadsheet, the signed contract, the architecture document a consultant exported to Word, the CSV of error codes, the schema file. Every guide that tells you to put your documentation in a wiki is quietly telling you to retype the second pile into the first. Teams start that migration and stop at about forty percent, and the result is two documentation systems, one searchable and one a folder.

Baalda treats those as one corpus. It is a team second brain built on plain markdown files on your own disk, and the search over a vault ranks the notes people wrote and the text pulled out of the files they dropped in, in one result set. That is the part of "how to write internal documentation" that the writing advice cannot reach, so it is most of what follows.

What should you actually write, and what should you leave alone

Write the things that exist only in somebody's head. The decision and why it went that way. The convention and the exception to it. The thing that looks like a bug and is deliberate. The order the steps have to happen in. None of that is recoverable from an artifact, which is why it is worth a person's hour.

Do not rewrite an artifact into prose. The spreadsheet is already the authority on its numbers, and a wiki page summarising it is a second authority that will disagree with the first inside a month. What the spreadsheet needs is not a rewrite, it is to be findable by its contents and to sit under the same permissions as everything else.

That split gives you a shorter writing job than any of the guides imply. A team of ten does not need two hundred pages. It needs maybe thirty notes that only a person could have written, plus every file it already has, reachable.

Why does the "put it all in the wiki" plan never finish

Because the cost is per artifact and the benefit is only felt in aggregate. Nobody gets credit for the forty-third page. The migration stalls, and then the wiki is both incomplete and authoritative-looking, which is worse than an honest folder.

It also assumes the only reader is a person who can click around. That stopped being true. A lot of internal documentation is now read by an assistant pulling context before it answers, and an assistant cannot click around a shared drive. It gets what the search gives it. If half your knowledge is in files the search never opened, the assistant will confidently answer from the half it can see, which is the specific failure that makes teams distrust the whole arrangement.

What does one search over notes and files actually return

In Baalda the MCP tool search_notes runs over a vault and scores two stores in the same query: note_index, derived from the note bodies, and blob_text, the words a desktop client extracted out of a .docx, .xlsx, .pptx, .csv, a JSON or YAML file, or a source file. Results come back interleaved and ranked together. Each hit carries a kind of note or file, and a file hit also carries blobId and ext, so the caller can tell a page someone wrote from a spreadsheet somebody dropped in. Default k is 10, maximum 50, and includeFiles defaults to true.

A file hit is then read with read_attachment_text, which returns the extracted words and not the bytes. An assistant asking for a 20 MB workbook over JSON-RPC wants its text, so that is what it gets: 20,000 characters by default, 200,000 at most, with a truncated flag saying whether there was more. list_attachments enumerates the files instead, each one carrying hasText, which tells you whether there is anything to read before you ask.

The extraction itself happens on the desktop, in Rust, not in the browser window and not on the server. The reason is in the module's own comment: the index has to be correct with no window open at all, because a rebuild runs on a background thread when a vault opens and the file watcher indexes headless. The server receives a copy as ranking fuel rather than re-parsing the container itself.

Who is allowed to see what the search found

This is where a documentation search differs from a desktop search tool, and it is the half that tools like DocFetcher, which genuinely does index inside fifty-odd formats and keeps that index current, do not attempt. They run as you, on your disk, over everything you can open.

A team store cannot do that. In Baalda the visibility rule sits inside the SQL WHERE clause of the file search, not in a filter applied afterwards, and the code says why: if unreadable files were scored and then removed, their contents would still influence the ranking of the readable ones, and "your query scored higher when I added a word that only appears in a file you cannot see" is a slow read of that file. Filtering after the fact leaks through the ranking. Filtering in the query does not.

Two kinds of file resolve differently. A registered file in the tree follows its folder's permissions through its doc_id, exactly as the notes beside it do. A hash-named drop in attachments/ with no document of its own is only a candidate for someone who can read the whole vault. The same per-folder rules govern a person in the app and an assistant holding a token, which is the point covered at more length in per-file permissions on a shared vault.

The file half of a search is capped at 2,000 rows per query, deliberately, because a blob_text row can be a megabyte where a note body is a few kilobytes. Notes are not capped the same way, and in neither case do the bodies leave the database.

How does documentation drift, and what catches it

Documentation does not usually go wrong by being written badly. It goes wrong by quietly ceasing to describe the thing. A review cadence is a calendar reminder, not a signal: it tells you to look, not that anything changed.

There is one place where this can be enforced rather than scheduled. When a client uploads the text it extracted from a file, the request carries the sha256 of the bytes it extracted from, and the server checks that hash against the bytes the blob actually holds. If they disagree, it refuses with a 409 and a sha_mismatch code. The reasoning in the route is the whole thesis of internal documentation in one line: storing it would make search describe content the vault does not hold.

So the index cannot drift away from the file underneath it. Somebody edits the spreadsheet, the hash changes, and the old extracted text is not allowed to keep standing in for it.

The permission check on that write runs the same way. Uploading extracted text requires edit rights on the file's own document, not merely write access to the vault, because that text is what the vault's search will say the file contains. Changing what search reports about a document is a write to that document.

Both of these are the kind of guarantee a style guide cannot give you. They are also, honestly, the only two parts of documentation decay that software can fix. The rest is still people.

What does this not solve

A fair amount, and the limits are specific rather than vague.

What you haveWhat happens to it
.docx, .xlsx, .pptx, .csv, JSON, YAML, source filesText extracted, ranked beside your notes
A PDFNot extracted by default. Recorded as unsupported, filled in only if a client that can read PDFs supplies the text
A scanned PDF or a photo of a whiteboardNothing. There is no OCR
A plain text file over 2 MB, or an Office file over 20 MBSkipped for size, still findable by its name
An imageAn empty result, which is a finished answer and not a failure

The PDF gap is the one most likely to bite, and it is a deliberate choice rather than an oversight: there is no dependency-free Rust extractor the project is willing to run on malformed input, so a PDF is marked unsupported instead of being parsed unsafely, and a viewer-side extractor can fill those rows in later. If your documentation is mostly scanned PDFs, a tool built for that, with OCR, is the honest answer and Baalda is not.

Stored text is capped too: 500,000 characters per file locally, and one megabyte on the server, past which the upload is refused outright.

Beyond extraction, Baalda is not a document management system. There is no approval workflow, no retention policy, no records management, and no importer from Confluence or Google Docs, so moving an existing wiki in is an export and a conversion you run yourself. If a one-click migration decides it, pick a tool with an importer. The storage question behind that choice is worked through in Confluence alternatives that use markdown.

And none of it writes the thirty notes for you. Search can stop you retyping a spreadsheet. It cannot work out why your team made the decision it made in March, and that remains the only documentation that was ever worth the hour.

Where to start, concretely

If you want a first week rather than a project:

  • Put the files you already have somewhere a single search covers them, under the permissions you actually want, and stop there. Do not reorganise them yet.
  • Write down the five decisions a new hire asks about in their first fortnight. That is your highest-value documentation and it fits in an afternoon.
  • For each artifact you were about to summarise into a page, write one short note saying what it is for and who owns it, and link to the file. The file stays the authority.
  • Connect your assistant to the same store over MCP so it answers from the corpus rather than from whatever it was last told.

The rest of the writing advice still applies once you are there. It just was never the part that was stopping you. More on what a shared store looks like for a whole team is in the team second brain.

FAQ

Frequently asked questions

What should internal documentation include?

The things that are not recoverable from any artifact you already own: why a decision went the way it did, the convention and its exceptions, the order steps have to happen in, and the behaviour that looks like a bug and is deliberate. Numbers, contracts and reports are already authoritative where they live, so link to them rather than summarising them into a page that will disagree with the original.

Do we have to rewrite our PDFs and spreadsheets as wiki pages?

No, and teams that try tend to stop part way, which leaves two documentation systems where one looks complete and is not. The useful move is to make the files findable by their contents and governed by the same permissions as your notes, then write only the context that is not in them.

Does Baalda search inside PDFs?

Not by default. Text is extracted on the desktop for Office files, CSVs, JSON, YAML and source files, but PDF is recorded as unsupported rather than parsed, because there is no dependency-free Rust extractor the project is willing to run on malformed input. A client that can read PDFs can supply the text afterwards. There is no OCR at all, so scanned documents and photographed whiteboards stay unsearchable.

Who should own internal documentation?

Ownership works better per artifact than per team. A note that records a decision belongs to whoever made it, and a file belongs to whoever maintains the thing it describes. The failure mode of a single documentation owner is that they become a bottleneck for content they did not write and cannot verify.

How do you stop internal documentation going stale?

A review cadence tells you to look, not that anything changed, so it catches drift late. What can be enforced is the link between an index and the thing it indexes: when extracted text is uploaded it carries the hash of the bytes it came from, and the server refuses it if that hash no longer matches the file, so search cannot keep describing a version the vault does not hold.

Start your team’s brain

Free and open source. No account needed.