Every field is source-linked and dated.See the rubric behind the grades.

How we grade
Sovereign AI Registry
ExploreBlogGov accessCertsCountries

Footer

Sovereign AI Registry

The compliance registry for AI vendors. Data residency, training defaults, retention, subprocessors and EU AI Act posture — one row per vendor, product and deployment, every claim linked to its source.

Registry

  • Explore vendors
  • Deployment models
  • Categories
  • Countries

Compliance

  • Gov access exposure
  • EU AI Act roles
  • Certifications

Resources

  • FAQ
  • Methodology
Built with ShipMore·Build yours →

© 2026 Sovereign AI Registry. All rights reserved.

Schematic line drawing of a meeting room whose recording fans out into many lines crossing a dashed boundary

Analysis

The meeting recorders tell you exactly who hears you

By Marta Reinders

Published on August 6, 2026

Sovereign AI Registry — batch 6, 2026-08-06. Thirteen meeting and transcription products graded; the registry now holds 83 rows.

Batch 5 graded the category that reads your repository and found that it would not say who reads it. This batch graded the category that records the room, and found the opposite. These vendors disclose well. Otter publishes an entity-by-entity sub-processor table with countries and an effective date. Fireflies, Fathom, AssemblyAI and Happy Scribe all run live trust centres with named entities. Speechmatics lists roughly thirty-five processors inline in its privacy policy. Only two of thirteen keep the list back.

So the finding is not that they hide it. It is what the lists say once you read them, and what the training clauses say when you read both halves.

The widest home-country gap in the registry

Rows

Median grade

EU-headquartered, hosted

5

63

US-headquartered, hosted

6

35

Twenty-eight points, and the two populations barely touch: the weakest European row scores 51, the strongest American one 55. Nothing else in this registry splits that cleanly on where the company is incorporated. Coding assistants did not — JetBrains is Czech and scores C 46 because its AI sub-processors are American.

Here the correlation holds because the European vendors actually kept the audio in Europe. Amberscript stores in Western Europe. jamie runs inference on Vertex AI in EEA data centres, contracted through Google Cloud EMEA Limited in Dublin. Happy Scribe uses Cloudflare EU, Heroku Ireland and DeepL Germany, with Slack in the US for internal communications under standard contractual clauses.

One of six American vendors offers any EU residency at all — AssemblyAI, which lets a customer switch processing to Dublin self-serve from the dashboard, and which is also the best-disclosed US row in the batch.

Good disclosure, read carefully, describes an American supply chain

The six US-headquartered hosted rows name 54 sub-processors between them. The five European ones name 23. That is not a disclosure gap in the Europeans' favour. It is a supply-chain gap: there are simply more companies in the American data path.

Fireflies is the clearest case. Its security page leads with a zero-day retention policy and "we don't train on it by default", and both claims are true. Its trust centre then names seventeen sub-processors, every one American, including two independent transcription vendors (AssemblyAI and Soniox) and five separate LLM vendors (OpenAI, Anthropic, Groq, Perplexity, Exa) plus ElevenLabs. A recorded meeting touches all of them. No EU residency is offered and no Chapter V transfer mechanism is named.

Fathom publishes the same information more starkly: four sub-processors, each tagged with whether it handles meeting content. Google Cloud, Anthropic, OpenAI and ElevenLabs — all USA, all handling the recording or the transcript.

This is a service to the buyer. It is also the reason both rows score D.

"We don't let third parties train on your data" is a narrower sentence than it sounds

Read the clause that follows it.

Vendor

The promise

The next sentence

Fathom

"We do not authorize third parties to use your personal information or User Content to train their artificial intelligence models"

"We may use and create de-identified data generated from your User Content to train, customize or improve our in-house artificial intelligence models"

Granola

"We do not allow third parties (like OpenAI or Anthropic) to use your data to train their AI models"

"Granola trains on your anonymized data so we can keep making Granola better. You can opt out of this in your Settings"

Read AI

Workspace API data is not used for "generalized/non-personalized" models

uses information "to train and improve our models within our Services"

Gladia

Pro and Enterprise data is never used for training

"Only users in the Free Plan are subject to data used for model training"

Granola's privacy policy adds the part that makes the opt-out weaker than it looks: once de-identified content is in the weights it stays there, because removal "may not be technically feasible without complete model retraining". Fathom publishes no opt-out from its in-house training at all.

Of the thirteen rows, exactly one makes training opt-in: Happy Scribe, which trains "only where you have opted in to this use at signup and only for as long as you remain opted in", and says consent can be withdrawn.

This pattern is why the rubric changed. Methodology v1.3 adds training_default: trains-deidentified at 3 points, between opt-out-clear (5) and opt-out-buried (2): the de-identification is the vendor's own unaudited claim, and on their own account the contribution cannot be withdrawn once trained.

It also corrects a bug we had been carrying. opt-in had scored 0 since v1, alongside "trains by default", and was never defined in the methodology text. Training that only happens on explicit consent means no training happens by default. It now scores 8. No published grade moved on either change, because no existing row had used either value — the bug stayed invisible until a vendor with a genuinely opt-in policy was graded.

Transcription does not require sending the recording anywhere

The top of the batch is not a compromise.

Grade

Row

Why

A 95

Whisper, self-hosted

MIT weights, MIT reference implementation, runs on the operator's own hardware. No vendor in the path at all.

A 86

Speechmatics on-premises containers

Kubernetes containers inside the customer's own cluster. ISO 27001:2022, SOC 2 Type II, roughly 35 processors published inline.

Every hosted row in this batch is selling the product built on top of that capability: the bot that joins the call, the calendar integration, the summaries, the search, the support contract. That is a real product and worth paying for. But a buyer who has been told that meeting transcription inherently requires an American vendor has been told something untrue, and the two rows above are the receipt.

The pattern matches batch 5 exactly, and it is now this registry's most repeated result: the A band is always a deployment the customer controls. Mistral reached A 94 among coding assistants by shipping weights on-premises, not by being French. Speechmatics reaches A 86 here the same way, from Cambridge.

What to check in your own stack

  1. Open your note-taker's sub-processor list and count the entries. Fireflies names seventeen. Every one is a company with a technical path to your recorded meetings.
  2. Read the training clause twice, and look for the word "third". A promise about third parties is not a promise about the vendor.
  3. Ask whether the AI processing region is a setting or a property. tl;dv stores in Germany and lets the account choose Europe or the United States for AI processing. Storage location does not settle processing location.
  4. Ask who consented. Every other category in this registry grades data the customer chose to send. A meeting recorder captures the voice of everyone in the room, including people from the other company who never saw the vendor's terms. Only Fireflies publishes a retention schedule that treats voice as a biometric identifier.

Method: every value on every row traces to a first-party document fetched and snapshotted at grading time, stored under evidence/<source_id>/. Scores are computed by score.py from recorded field values with no judgment at scoring time; all 83 live rows recompute exactly. Rubric: methodology v1.3.

This article was researched and written during the registry's batch 6 harvest, with AI assistance and human review before publication. The hero image is AI-generated. Every figure is computed from the registry's live records. Corrections: see the methodology page.