Short answer
Gemini 4 Argon can emit up to 1 million tokens in one response, roughly fifteen times its predecessor's 64,000. That moves the bottleneck from generating work to checking it. Output that lands in a chat transcript cannot be reviewed by a team, so it has to land in shared files several people can edit at once.
The ceiling that mattered this year was never how much a model could read. It was how much it could write in one go, and for most of 2026 that number was small enough that nobody had to think about it. Google just moved it by a factor of fifteen, and the part of a team's work it lands on is the part nobody has built for: what happens to machine-written text after it exists.
What did Google actually ship on 30 September?
Gemini 4 Argon, announced on 30 September 2026 by Koray Kavukcuoglu, SVP of Google DeepMind. The headline number is an output limit of 1 million tokens in a single response, up from 64,000.
Read that twice, because most of the coverage blurred it. This is not a context window figure. Plenty of models already take a million tokens of input. Argon is allowed to emit a million tokens in one turn, which is a different claim and a much rarer one.
The rest of the announcement, as of 5 October 2026:
- Pricing. An introductory $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period. Cached input tokens are priced at 95% off the input rate.
- Access. Restricted to a limited group of trusted cyber defenders through Google's Fairwind Program. Wider rollout is said to begin with paid API customers and Google AI Ultra subscribers, with no public release date given.
- Benchmarks. Google reports leads on the Vals Index, AutomationBench and DeepSWE v1.1. Those are vendor-reported numbers published alongside the announcement, and because access is still gated there has been very little opportunity for anyone outside Google to reproduce them.
So the most expensive single response the model can physically produce costs about ten dollars at the introductory rate. That is the whole economic barrier to generating something the length of a long novel, on demand, in one request.
Why is a bigger output ceiling a review problem?
Because the constraint on how much AI-written work a team can actually use was never the model's ceiling. It was the number of people willing to read the result and put their name under it.
A 64,000 token response was already more than anyone reads carefully in a sitting, but it stayed human-shaped. A long memo. One big refactor. A report with a conclusion at the end. A million tokens is a different category of object: a migration across an entire codebase, a complete set of runbooks, every contract in a folder re-drafted at once. One person presses enter and the output arrives faster than the team can form an opinion about it.
Generation scaled. Reading did not. That gap is the actual story of this release, and it is a knowledge problem before it is a model problem, which is where Baalda comes in. Baalda is a team second brain built on plain markdown files on your own disk, where several people edit the same notes in real time and an AI reads and writes those same files over MCP. The reason that shape matters here is not storage. It is that review has to be able to run in parallel, and almost nothing a model currently writes into lets it.
Where does a million tokens of output land today?
In a chat transcript. Which means, concretely:
- One account can see it. Nobody else on the team has a link.
- Nobody can edit it. You can reply to it, which is not the same thing.
- It has no addresses. There is no way to say "section 14 is wrong" in a way a tool or a teammate can act on.
- It does not survive. The next run starts from nothing and writes its own near-duplicate.
The usual rescue is to copy and paste it into a document, and that step is where the work quietly dies. The document and the thing the model wrote immediately diverge, the corrections live in the copy, and the model never sees any of them. You get the volume without ever getting the corrections back. This is the problem a team second brain exists to solve, and a fifteenfold jump in output just made it fifteen times more expensive to ignore.
| Where output lands | Several people can edit it | Edits survive the session | A later model run can read it back |
|---|---|---|---|
| Chat transcript | No | No | No |
| Exported doc | One at a time, or in a cloud suite | Yes | Only if you wire it up |
| Shared markdown vault | Yes, simultaneously | Yes, as files on disk | Yes, over MCP |
How does a shared markdown vault change the review?
By making the destination something a team can be inside while it is still being written.
Baalda serves the Model Context Protocol at POST /api/mcp, and the write tools are the plain ones: create_note, append_note, edit_note, update_note. The part that matters for this post is what happens underneath them. A write from an MCP client does not go to a separate AI copy of the note. If the note is open in somebody's editor, the write mutates the live shared document and broadcasts to every connected client, which is the identical path a human keystroke takes. Baalda merges keystroke by keystroke through a CRDT, so there is no conflict dialog and no second version of the file.
What that buys you on a 700,000 word draft is simple and quite hard to get any other way. The model can still be writing section 40 while three people are already fixing sections 1 through 12, each in a different note, all of them watching the text arrive. Review stops being a thing that happens after generation finishes and becomes a thing that happens alongside it.
It is also file-shaped, which is the other half. A vast generated draft lands as a folder of .md files on your own disk. Each piece has a path, so somebody can own it. A correction is an edit to a line, not a new message in a thread. And because the files are the real artifact rather than an export of one, the next run reads the corrected version back through search_notes and read_note instead of starting again from whatever the model believed last week. The setup for connecting an AI over MCP is one token and one endpoint.
What does this not fix?
Four things, and they are worth saying plainly.
It does not make the output correct. A vault full of confident, fluent, wrong text is still wrong, and it now arrives faster and sits somewhere the team is more inclined to trust. Baalda moves where the review happens. It does not do the review.
It does not reduce how much there is to read. Parallel review divides the work, it does not shrink it. Five people against 700,000 words is still 140,000 words each, which is a bad afternoon. The only real answers to volume are asking for less of it or deciding, out loud, which parts nobody is going to check.
It cannot get you Argon. Access is Google's to grant, through the Fairwind Program now and paid API and AI Ultra tiers later. Nothing about where your notes live changes your place in that queue.
And if you are one person, you do not have this problem yet. A solo vault and a transcript work fine at one reader. Obsidian is still the better single-user editor and nothing in this release changes that, it simply has no answer for several people inside one document at the same time, which is the gap that opens up the moment review has to be shared. If your knowledge already lives in Notion, Notion's hosted MCP server is genuinely easier to stand up: it is OAuth, and there is no server of yours to run.
Should you do anything about this today?
Probably nothing about Argon specifically. As of 5 October 2026 you most likely cannot use it, and the benchmark claims have not been checked by anyone outside Google.
But the thing it makes visible is already in the building with the models you do have. Ask where machine-written output lands by default on your team right now. If the honest answer is "a transcript in somebody's account", that is the piece worth changing, and it is far cheaper to change while the volume is still small. The teams that will get something out of a million token response are the ones that already had somewhere for a long one to go, and who had an AI writing into the same files they edit rather than into a window beside them.
