Every field is source-linked and dated.See the rubric behind the grades.

How we grade
Sovereign AI Registry
ExploreBlogGov accessCertsCountries

Footer

Sovereign AI Registry

The compliance registry for AI vendors. Data residency, training defaults, retention, subprocessors and EU AI Act posture — one row per vendor, product and deployment, every claim linked to its source.

Registry

  • Explore vendors
  • Deployment models
  • Categories
  • Countries

Compliance

  • Gov access exposure
  • EU AI Act roles
  • Certifications

Resources

  • FAQ
  • Methodology
Built with ShipMore·Build yours →

© 2026 Sovereign AI Registry. All rights reserved.

Line diagram: a grid of vendor record cards under a magnifying glass

"We don't train on your data" is table stakes

By Marta Reinders

Published on August 4, 2026

Sovereign AI Registry — analysis, 4 August 2026. Every figure below recomputes from /api/records and score.py.

Ask an AI vendor whether they train on your data and you will almost always get the same answer. We checked all 58 products in the registry:

53 of 58 — 91% — commit in writing that they never train on customer data by default.

That number is the whole problem. A claim that 91% of the market makes is not a differentiator, it is a floor. It has been absorbed into the standard enterprise contract the way "we use TLS" was a decade ago. If it is the question your procurement checklist leads with, your checklist is measuring something every plausible vendor already passes.

The claim predicts nothing

Here is what happens when you take those 53 vendors and look at how they actually score across the full rubric:

Score

Lowest-scoring vendor that promises never to train

27 / 100

Median

59 / 100

Highest

96 / 100

A 69-point spread. Among the 53 vendors making the same promise there are 9 A grades — and also 5 Ds and an F. Replicate commits never to train on your data and scores F 27. Nextcloud Assistant makes the same commitment and scores A 96.

Both statements are true. Only one of those two vendors is a defensible procurement decision, and the training question cannot tell you which.

What actually separates them

Split the registry into its top and bottom bands and the picture resolves immediately.

A-band (≥80) · 15 rows

D/F band (<50) · 16 rows

No non-EU government can compel access

15 of 15

0 of 16

EU-only or customer-controlled residency

13 of 15

0 of 16

Publishes no sub-processor list

0 of 15

6 of 16

Offers a self-host path

7 of 15

3 of 16

The separation is total on two axes. Every single A-band row is out of reach of non-EU government compulsion. Not one D/F row is. And no A-band vendor hides its sub-processors, while more than a third of the bottom band publishes nothing at all.

That is the real question. Not will you train on my data — everyone says no. It is whose jurisdiction is my prompt sitting in, and who else touches it on the way.

The jurisdiction number

Across all 58 rows:

  • 33 (57%) are exposed to the US CLOUD Act.
  • 19 (33%) are reachable by no non-EU government.
  • 3 carry PRC-linked exposure — two under the National Intelligence Law, one under the Hong Kong National Security Law.
  • 30 of 58 are headquartered in the United States.

Median score for a US-headquartered vendor: 53.5. Median for an EU-headquartered one: 84.0.

Before anyone reads that as a flag-waving result — it is not. The gap is not caused by where the letterhead is. It is caused by what a company controls. We published a separate finding showing German and French vendors landing in the C band precisely because they broker American models: SAP's Generative AI Hub scores C 52, thirty-eight points below STACKIT, its German competitor. European incorporation buys you nothing if the prompt still lands in us-east-1.

The exit-cost axis nobody scores

One more split, and it is the one no sovereignty framework currently in force measures:

  • 36 of 58 (62%) serve open-weight models.
  • 14 (24%) are proprietary-only.
  • 17 (29%) offer some self-host path.

This matters more than certification stacks. If a vendor serves open weights, your cost of leaving is changing a base URL. If it is proprietary-only, leaving is a rebuild — and every commitment they have made to you is revocable on their schedule, not yours. The EU Cloud Sovereignty Framework scores data centres, staff vetting and legal control. It does not score whether the model you depend on can be run anywhere else.

What to ask instead

Three questions that actually discriminate, in order:

  1. Which governments can lawfully compel access to this data? Not where the company is registered — whose jurisdiction the inference runs in and which corporate entity signs your contract.
  2. Who is your full sub-processor list, with locations? Ten percent of this registry publishes nothing. Silence is a finding, not a blank.
  3. Can I run this model somewhere else? Open weights make every other promise enforceable, because you can leave.

Ask "do you train on my data" if you like. Just know that 91% of the market will pass, and the answer will not have told you anything.

Method and limits

This is an analysis of stated positions. The registry records what a vendor commits to in writing, with an evidence URL, a claim basis (stated where the vendor says it, inferred where we concluded it) and a confidence level per field. It does not audit whether vendors honour those commitments — no public dataset can.

So "91% say they never train" is a finding about the market's contractual floor, not a certificate of good behaviour. That is precisely why the other axes matter: jurisdiction and portability are structural, and structure does not rely on trust.

All 58 rows recompute deterministically from methodology v1.2 and the published field values. Browse the full dataset in the registry, or by government access exposure.