AI news

Mistral Large 4 scored 59.9% on business workflows someone else had already written down

By Baalda Team · · 6 min read

Short answer

Only where the work is already written down. Mistral Large 4, in public preview since 6 October 2026, scores 59.9% on AutomationBench across 657 business workflows that the benchmark specifies for it. Inside a company nobody specified them, so the fix is written context the agent can read and add to.

Read the Mistral Large 4 announcement as a list of things the model was given rather than a list of things it can do, and it changes shape. Every headline score comes from an evaluation that handed the model a task somebody had already specified. That is not a criticism of the benchmarks, which have to work that way. It is the gap between a benchmark and a company, and it is the whole reason Baalda exists: the specification a benchmark supplies for free is the thing your team never wrote down.

What did Mistral actually ship on 6 October 2026?

Mistral released Mistral Large 4, internally "Le Chonk", as a public preview on 6 October 2026. In Mistral's own words it is "a 1 trillion-parameter natively multimodal model with 49 billion active parameters" and "our largest and most capable model to date". It was "trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own datacenters in Europe", covers more than 160 languages, and lists at $1.36 per million input tokens and $4.18 per million output tokens on Mistral Studio. Mistral says "we will release the weights by the end of the month", and that for organisations needing "sovereign, auditable AI for security operations, it will be able to run on private cloud or on-premise".

The announcement does not state a context window. Third-party listings disagree with each other on that number as of 7 October 2026, so this post does not quote one.

The part worth reading twice is the benchmark table, because it is weighted towards work rather than puzzles:

BenchmarkMistral's reported result
AutomationBench59.9% across "657 business workflows across apps like Gmail, Google Sheets, Slack, and Salesforce"
AA-Briefcase (long-horizon knowledge work)1,393 Elo, "ahead of DeepSeek V4 Pro"
HarveyAI Legal Agent benchmark"ML4 outperforms all open-source models"
DeepSWE v1.161.7%
Terminal-Bench 428.3%
Cybench93% of 40 security challenges

Mistral's framing is that "ML4 runs general-purpose agents that gather information, use tools, and produce finished deliverables across complex workflows". All of those numbers are Mistral's own, published with the preview, and none of them are independently reproduced yet.

Why does a 59.9% on business workflows not transfer to your company?

Because AutomationBench writes the workflow down and your company does not.

Look at what an agentic eval has to contain to be an eval at all. It needs a task statement precise enough to grade. It needs the tools declared. It needs a success condition someone agreed on in advance. 657 business workflows means 657 pieces of specification written by people whose job was to write specifications. The model is being measured on execution, with the hard, slow, human half already handed over.

Now take the same model into a team of five to fifty. Ask it to run the deploy. The deploy procedure is four steps in one engineer's muscle memory, a Slack thread from March about the migration order, and a caveat about Fridays that nobody has ever typed. Ask it why the auth provider was swapped. The reason was a pricing change discussed in a call that was not minuted. Ask it to draft the client update. The house style is "whatever the last one looked like", and the last one is in someone's sent folder.

None of that is a model capability problem. A 1T-parameter model and a 7B model are equally ignorant of why your team chose Postgres. Raising the score on the executor does not produce the specification, and the specification is what is missing. This is also why "we tried an agent and it was disappointing" is such a common result: the eval arrives fully briefed and the company arrives with nothing.

So the honest question Mistral Large 4 raises for a team is not which model to call. It is: where would the briefing live, who can change it, and can the model add to it.

Can an agent write the specification instead of just consuming it?

It can, if the place it writes to is a file rather than a transcript.

Baalda is a team second brain: plain markdown files on your own disk, edited by several people at once in real time, with an AI reading and writing those same files over the Model Context Protocol. How MCP works for notes covers the protocol end. The part that matters here is that the write tools are shaped for an agent that has to add to a document a human also owns.

Creating the note takes a vault-relative path, so the thing the agent produces has an address from the first write:

json
{
  "name": "create_note",
  "arguments": {
    "vaultId": "vault_3b1",
    "relPath": "Runbooks/deploy.md",
    "content": "# Deploy\n\n## Order of operations\n..."
  }
}

Adding to it later is append_note, which takes an idempotencyKey:

json
{
  "name": "append_note",
  "arguments": {
    "docId": "note_7f2a",
    "text": "\n## 2026-10-07 · migration ran before the restart, not after\n",
    "expectedRevision": "r41",
    "idempotencyKey": "deploy-note-2026-10-07-a"
  }
}

Two details in that call do the work. The idempotencyKey means "a repeat with the same key returns the first result instead of appending again", so an agent loop that retries after a timeout cannot leave the same paragraph in the runbook twice. The expectedRevision, read back from read_note, refuses the write outright if someone edited the note in between, rather than landing on top of them. For a partial change there is edit_note, where each anchor has to match exactly once and "a missing or ambiguous anchor refuses the whole call with nothing written".

The result is that the artifact the benchmark handed the model for free is the artifact the team slowly accumulates. The first time an agent works out the deploy order, that goes in Runbooks/deploy.md at a path anyone can open, not into a chat log that closes. The second time, it reads it before it starts. A human who disagrees fixes one line in one file, and every subsequent run reads the fix.

Where does a shared vault not help?

It does not make the specification true, and it does not make anyone write one.

An agent with append access can append something confidently wrong, and a wrong runbook is read exactly as obediently as a right one. What a vault changes is that the wrong thing now has an address and a revision, so it is correctable once instead of being re-litigated in five threads. That is a smaller claim than "the agent maintains your documentation", and it is the honest one.

It also does nothing about the reason most teams have no runbook, which is usually that nobody has twenty minutes. Installing a notes application does not buy the twenty minutes. The teams this helps are the ones where the writing happens and then rots, not the ones where it never happens.

And on Mistral's own pitch: a vault does not make you sovereign either. If you call the hosted preview API, your prompt leaves your boundary whatever your notes are stored on. The weights landing at the end of October is what makes the inference half movable, and self-hosting is what makes the knowledge half movable. Those are two separate decisions and a team that makes only one of them has moved half the problem.

What should a team do with this release?

Treat the model as the easy variable, because it now is. There are several models at this level, one more arrives most weeks, and swapping which one reads your notes is a client-side change.

Treat the written context as the hard one. Before running an agent against real work, pick the three questions it will be asked most and write the answers as notes at stable paths, however roughly. Then give the agent write access to those same notes and let it add what it learns, with a person reading the diffs. If you want the cautious version of that last step, keep the agent's writes to a folder of its own, with its own permissions, until you trust them.

The benchmark table will keep going up. The 657 workflows will keep being supplied. Your own are not going to write themselves, and the model is not what is stopping them.

FAQ

Frequently asked questions

What is Mistral Large 4?

Mistral's largest model, put into public preview on 6 October 2026. Mistral describes it as "a 1 trillion-parameter natively multimodal model with 49 billion active parameters" and "our largest and most capable model to date", trained "from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own datacenters in Europe", covering more than 160 languages. The announcement lists $1.36 per million input tokens and $4.18 per million output tokens on Mistral Studio, and says "we will release the weights by the end of the month". It does not state a context window, and third-party listings disagreed on that number as of 7 October 2026.

What does AutomationBench actually measure?

Execution against workflows that have already been specified. Mistral's announcement describes it as "657 business workflows across apps like Gmail, Google Sheets, Slack, and Salesforce" and reports 59.9%. For the benchmark to be gradable, each of those workflows needs a task statement, declared tools and an agreed success condition, all written in advance by people whose job was to write them. The score is Mistral's own, published with the preview, and is not yet independently reproduced.

Are Mistral Large 4's weights open?

Not yet. As of 7 October 2026 the model is a public preview on the Mistral Studio API, and the announcement says "we will release the weights by the end of the month" and that for organisations needing "sovereign, auditable AI for security operations, it will be able to run on private cloud or on-premise". The announcement does not name the licence the weights will carry, so there is nothing to check yet on that front.

How does an AI agent add to a shared note without overwriting a teammate?

Through the write tools, not by replacing the file. `read_note` returns a `revision`; passing it back as `expectedRevision` to `append_note`, `update_note` or `edit_note` makes the write fail rather than land if the note changed in between. `append_note` also takes an `idempotencyKey`, so a retried call "returns the first result instead of appending again". `edit_note` anchors have to match exactly once, and "a missing or ambiguous anchor refuses the whole call with nothing written".

Will a stronger model fix an agent that keeps getting our internal processes wrong?

No, and this is the limit worth stating. A 1T-parameter model and a 7B model are equally ignorant of why your team chose its database. If the agent is wrong about your deploy order, the missing input is a written deploy order, not more parameters. A vault gives that input an address and a revision history; it does not write it, and it does not make it correct once written.

Start your team’s brain

Free and open source. No account needed.