
Analysis
By Marta Reinders
Published on August 6, 2026
Sovereign AI Registry — batch 6, 2026-08-06. Thirteen meeting and transcription products graded; the registry now holds 83 rows.
Batch 5 graded the category that reads your repository and found that it would not say who reads it. This batch graded the category that records the room, and found the opposite. These vendors disclose well. Otter publishes an entity-by-entity sub-processor table with countries and an effective date. Fireflies, Fathom, AssemblyAI and Happy Scribe all run live trust centres with named entities. Speechmatics lists roughly thirty-five processors inline in its privacy policy. Only two of thirteen keep the list back.
So the finding is not that they hide it. It is what the lists say once you read them, and what the training clauses say when you read both halves.
| Rows | Median grade |
|---|---|---|
EU-headquartered, hosted | 5 | 63 |
US-headquartered, hosted | 6 | 35 |
Twenty-eight points, and the two populations barely touch: the weakest European row scores 51, the strongest American one 55. Nothing else in this registry splits that cleanly on where the company is incorporated. Coding assistants did not — JetBrains is Czech and scores C 46 because its AI sub-processors are American.
Here the correlation holds because the European vendors actually kept the audio in Europe. Amberscript stores in Western Europe. jamie runs inference on Vertex AI in EEA data centres, contracted through Google Cloud EMEA Limited in Dublin. Happy Scribe uses Cloudflare EU, Heroku Ireland and DeepL Germany, with Slack in the US for internal communications under standard contractual clauses.
One of six American vendors offers any EU residency at all — AssemblyAI, which lets a customer switch processing to Dublin self-serve from the dashboard, and which is also the best-disclosed US row in the batch.
The six US-headquartered hosted rows name 54 sub-processors between them. The five European ones name 23. That is not a disclosure gap in the Europeans' favour. It is a supply-chain gap: there are simply more companies in the American data path.
Fireflies is the clearest case. Its security page leads with a zero-day retention policy and "we don't train on it by default", and both claims are true. Its trust centre then names seventeen sub-processors, every one American, including two independent transcription vendors (AssemblyAI and Soniox) and five separate LLM vendors (OpenAI, Anthropic, Groq, Perplexity, Exa) plus ElevenLabs. A recorded meeting touches all of them. No EU residency is offered and no Chapter V transfer mechanism is named.
Fathom publishes the same information more starkly: four sub-processors, each tagged with whether it handles meeting content. Google Cloud, Anthropic, OpenAI and ElevenLabs — all USA, all handling the recording or the transcript.
This is a service to the buyer. It is also the reason both rows score D.
Read the clause that follows it.
Vendor | The promise | The next sentence |
|---|---|---|
"We do not authorize third parties to use your personal information or User Content to train their artificial intelligence models" | "We may use and create de-identified data generated from your User Content to train, customize or improve our in-house artificial intelligence models" | |
"We do not allow third parties (like OpenAI or Anthropic) to use your data to train their AI models" | "Granola trains on your anonymized data so we can keep making Granola better. You can opt out of this in your Settings" | |
Workspace API data is not used for "generalized/non-personalized" models | uses information "to train and improve our models within our Services" | |
Pro and Enterprise data is never used for training | "Only users in the Free Plan are subject to data used for model training" |
Granola's privacy policy adds the part that makes the opt-out weaker than it looks: once de-identified content is in the weights it stays there, because removal "may not be technically feasible without complete model retraining". Fathom publishes no opt-out from its in-house training at all.
Of the thirteen rows, exactly one makes training opt-in: Happy Scribe, which trains "only where you have opted in to this use at signup and only for as long as you remain opted in", and says consent can be withdrawn.
This pattern is why the rubric changed. Methodology v1.3 adds training_default: trains-deidentified at 3 points, between opt-out-clear (5) and opt-out-buried (2): the de-identification is the vendor's own unaudited claim, and on their own account the contribution cannot be withdrawn once trained.
It also corrects a bug we had been carrying. opt-in had scored 0 since v1, alongside "trains by default", and was never defined in the methodology text. Training that only happens on explicit consent means no training happens by default. It now scores 8. No published grade moved on either change, because no existing row had used either value — the bug stayed invisible until a vendor with a genuinely opt-in policy was graded.
The top of the batch is not a compromise.
Grade | Row | Why |
|---|---|---|
A 95 | MIT weights, MIT reference implementation, runs on the operator's own hardware. No vendor in the path at all. | |
A 86 | Kubernetes containers inside the customer's own cluster. ISO 27001:2022, SOC 2 Type II, roughly 35 processors published inline. |
Every hosted row in this batch is selling the product built on top of that capability: the bot that joins the call, the calendar integration, the summaries, the search, the support contract. That is a real product and worth paying for. But a buyer who has been told that meeting transcription inherently requires an American vendor has been told something untrue, and the two rows above are the receipt.
The pattern matches batch 5 exactly, and it is now this registry's most repeated result: the A band is always a deployment the customer controls. Mistral reached A 94 among coding assistants by shipping weights on-premises, not by being French. Speechmatics reaches A 86 here the same way, from Cambridge.
Method: every value on every row traces to a first-party document fetched and snapshotted at grading time, stored under evidence/<source_id>/. Scores are computed by score.py from recorded field values with no judgment at scoring time; all 83 live rows recompute exactly. Rubric: methodology v1.3.
This article was researched and written during the registry's batch 6 harvest, with AI assistance and human review before publication. The hero image is AI-generated. Every figure is computed from the registry's live records. Corrections: see the methodology page.