AI news

Google open-sourced the embedder. Your notes app kept the index

By Baalda Team · · 7 min read

Short answer

An embedding index is derived data, so the team that owns the source files can rebuild it with any model. Google released EmbeddingGemma 2 under Apache-2.0 on 6 October 2026. It only helps your team's knowledge if the notes are files you hold and the indexer is a component you can replace.

There is a line in every notes product that nobody shows you, and it runs between the files and the index built over them. On one side are the things you wrote. On the other is a pile of numbers that decides what comes back when you search, built by a model somebody else picked, living somewhere you cannot read. For most teams that line is the ownership boundary, whatever the marketing page says about export. Baalda draws it in a place you can check: the notes are plain markdown on your own disk, and the index over them is a cache with a swappable model behind it, because a search index you cannot rebuild is a search index you do not own.

Google made that distinction matter this week by giving away a good embedder.

What did Google actually ship on 6 October 2026?

EmbeddingGemma 2, published on 6 October 2026 under the Apache 2.0 licence, with weights on Hugging Face.

The figures below are from Google's own model card and the developer guide published the same day:

  • It is modular. Text and code only is 270M parameters, text plus vision is 440M, text plus audio is 570M, and the full multimodal model is 740M.
  • Everything lands in one 768-dimensional space: text, code, images, video and audio.
  • It is trained with Matryoshka Representation Learning, so that 768-dimensional vector can be truncated to 512, 256 or 128 dimensions and still work.
  • The context window is 8,192 tokens, shared across modalities. It covers 100+ languages, with training data spanning over 140.
  • Google's reported scores at 768 dimensions, full precision: 61.36 on MTEB multilingual, 78.68 on MTEB code, 64.64 on MIEB lite for images, 50.67 on MMEB v2 for video, 69.54 on MSEB for audio retrieval.
  • The developer guide lists it running under sentence-transformers, Hugging Face Transformers, vLLM, SGLang, MLX, Ollama, LM Studio and LiteRT.

Read that list as an infrastructure change rather than a model release. A team can now run a competent, multilingual, multimodal embedder on a laptop, with no API key, no per-token bill and no data leaving the building. The thing that used to be a line item from OpenAI or Cohere is a file you download.

Which raises the question the announcement does not answer: point it at what?

Why can a local embedder not reach most teams' notes?

Because in most products the index is not yours and the files are not files.

An open-weight embedder is a function. It turns text into a vector. To be useful on your team's knowledge it needs two things that have nothing to do with the model: a corpus it can read, and somewhere to put the result that your search actually consults. Cloud notes tools give you neither.

Notion, Confluence and the rest hold your pages as rows in their database and expose no way to supply your own embeddings or read theirs. Their search is a finished feature, not a seam. You can export your workspace to markdown, embed that with EmbeddingGemma 2, and build a perfectly good index on the side, and you will then have two systems that disagree the moment somebody edits a page. The copy you indexed is already stale, and the product's own search box is still running whatever it was running.

This is the shape of a lot of 2026 lock-in and it has nothing to do with export buttons. Export is not ownership: you can get the text out, and you still cannot change how the thing you use every day finds it. An open model does not fix a closed index.

The teams that can use EmbeddingGemma 2 on their own knowledge today are the ones whose knowledge is already a directory of files. That is the whole eligibility test, and it is worth saying that Obsidian passes it cleanly. A vault is markdown in a folder, and anyone can point a local indexer at a folder. What Obsidian does not give you is the same thing for a team: several people editing one note at once, with permissions, over a shared index. That is the gap Baalda is built in.

What does a replaceable index look like in code?

It looks like an interface with a default behind it. Baalda is open source under Apache-2.0, so this is readable rather than a claim.

The embedder in the server is a two-method contract:

ts
export const EMBED_DIM = 256;

export interface Embedder {
  readonly dim: number;
  embed(text: string): number[];
}

export let embedder: Embedder = localEmbedder;

export function setEmbedder(next: Embedder): void {
  embedder = next;
}

The shipped default, localEmbedder, is deliberately unambitious: a 256-dimension hashed bag-of-words. Tokens are lowercased word runs, each hashed into a bucket with FNV-1a, counts accumulate and the vector is L2-normalised so cosine similarity is a dot product. It is deterministic and it makes no network call, which is the point. A self-hosted or air-gapped deployment has working search on the day it comes up, with no API key and no vendor.

The index it writes into is explicitly disposable. The file header says note_index and note_links are "a rebuildable cache derived from the canonical Yjs state". Indexing runs whenever a note's document is stored, debounced two seconds so a burst of keystrokes collapses into one write, and it embeds title plus body together so a query matching the title still ranks the note.

Three details make the swap to something like EmbeddingGemma 2 a normal piece of work rather than a migration:

  • `note_index.vector` is a `JSONB` column. It is not a fixed-width vector type, so nothing in the schema asserts a dimension. A 768-dimensional vector fits where a 256-dimensional one was.
  • 256 is one of EmbeddingGemma 2's own truncation sizes. If you want the existing storage footprint, Matryoshka gives you a 256-dimensional cut of the same model, and the column shape does not change at all.
  • The source of truth is not the index. The notes are markdown, synced as CRDT documents and written to disk. Delete every row in note_index and you have lost a cache, not knowledge. That is what makes re-indexing against a new model a chore rather than a risk.

The same files are what an AI reads and writes over MCP, through one endpoint on a server you can run yourself (how that works). So the embedder, the index, the notes and the assistant all sit inside one boundary you drew, and each of the four can be replaced without touching the other three.

Where this does not help

Baalda ships no EmbeddingGemma 2 adapter, and the default embedder is worse than it sounds.

A hashed bag-of-words matches words, not meaning. It does not know that "deploy" and "release" are related, it has no notion of a synonym, and 256 buckets over a whole vocabulary means unrelated words collide in the same slot by design. Calling that semantic search is generous. It is a cheap, honest, offline default that gets you ranked results with zero setup, and it is nowhere near what a real embedding model does.

So the seam described above is a seam, not a feature. setEmbedder is an export with no production caller in the server. Using EmbeddingGemma 2 means you write the adapter, run the model somewhere the server can reach, re-index the vault, and own the result. Nobody has done that work for you, and this post is not pretending otherwise.

Two more limits worth stating:

  • If you want the best search quality today with no work, this is the wrong argument. A cloud product with a tuned retrieval stack will beat Baalda's default on day one and probably on day one hundred. The case here is about who gets to change it later, which is a different thing from who is better now.
  • A local model is not automatically cheaper or faster. Embedding a large vault with a 740M-parameter model on your own hardware costs real compute and real time, and somebody has to operate it. "Free weights" is a licence fact, not an infrastructure one.

Three questions, and they apply to whatever you are using now, not just to us.

  1. Can you read the index? Not the documents, the derived data. If the answer is no, your search behaviour is a black box you rent.
  2. Can you rebuild it against a different model? This is the one EmbeddingGemma 2 just made concrete. There is a free, open, competent embedder on the table as of 6 October 2026. A system that cannot accept it is a system that will not accept the next one either.
  3. What happens if the index is lost? If the answer is anything other than "it gets rebuilt from the files", then the index is not a cache, it is a second copy of your knowledge, and you now have two things to keep honest.

A team second brain is partly an argument about storage and partly an argument about control. Open weights moved one piece of that into reach for everybody this week. Whether your team can pick it up depends on a decision made much earlier, about whether the things you know are files you hold or rows in somebody's product.

FAQ

Frequently asked questions

What is EmbeddingGemma 2 and what licence is it under?

An open multimodal embedding model from Google DeepMind, published 6 October 2026 under Apache 2.0 with weights on Hugging Face. Per Google's model card it is modular: 270M parameters for text and code, 440M with vision, 570M with audio, 740M for the full model. Text, code, images, video and audio all map into one 768-dimensional space, with an 8,192-token context window and 100+ languages.

Can I use a local embedding model for my team's notes?

Only if your notes are files you can read and your search index accepts vectors you supply. Cloud tools like Notion and Confluence hold pages in their own database and expose no way to provide your own embeddings, so the most you can do is export, index the export on the side, and watch that copy go stale. A vault of markdown files passes the test; so does Obsidian, for one person.

What embedder does Baalda ship by default?

A 256-dimension hashed bag-of-words that runs entirely offline, in `apps/server/src/index/embedder.ts`. Tokens are lowercased word runs hashed into buckets with FNV-1a, counts accumulate, and the vector is L2-normalised so cosine similarity is a dot product. It makes no network call, so a self-hosted or air-gapped deployment has working search from the first boot with no API key.

Would swapping in EmbeddingGemma 2 need a database migration?

No. `note_index.vector` is a `JSONB` column rather than a fixed-width vector type, so nothing in the schema asserts a dimension and a 768-dimensional vector fits where a 256-dimensional one was. EmbeddingGemma 2 is also trained with Matryoshka Representation Learning and truncates to 512, 256 or 128 dimensions, so the 256-dimension cut keeps the existing storage footprint exactly.

Does Baalda come with an EmbeddingGemma 2 integration?

No, and that is the honest limit. `setEmbedder` is an exported seam with no production caller in the server. Using an open embedder means writing the adapter, running the model somewhere the server can reach, and re-indexing the vault yourself. A cloud product with a tuned retrieval stack will give you better search today; the argument here is about who gets to change it later.

Start your team’s brain

Free and open source. No account needed.