Every field is source-linked and dated.See the rubric behind the grades.

How we grade
Sovereign AI Registry
ExploreBlogGov accessCertsCountries

Footer

Sovereign AI Registry

The compliance registry for AI vendors. Data residency, training defaults, retention, subprocessors and EU AI Act posture — one row per vendor, product and deployment, every claim linked to its source.

Registry

  • Explore vendors
  • Deployment models
  • Categories
  • Countries

Compliance

  • Gov access exposure
  • EU AI Act roles
  • Certifications

Resources

  • FAQ
  • Methodology
Built with ShipMore·Build yours →

© 2026 Sovereign AI Registry. All rights reserved.

Schematic line drawing of a code repository fanning out to eight destinations, five of which are empty outlines

Analysis

The coding assistants that will not tell you who reads your code

By Marta Reinders

Published on August 6, 2026

Sovereign AI Registry — batch 5, 2026-08-06. Fifteen coding assistants graded; twelve new rows added to the registry, which now holds 70.

A coding assistant is not a chatbot with a different prompt. It is the only category in this registry where the vendor receives the repository — imports, type declarations, related files, the whole retrieval context — rather than a question a human chose to type. So the disclosure standard should be higher here than for an inference API, and it is measurably lower.

Median sovereignty grade across the fifteen coding assistants is 49 (C). The other 55 rows in the registry sit at 58. The category that ingests the most sensitive input scores nine points below everything else.

A third of them publish no sub-processor list

Five of the fifteen do not name a single sub-processor. Registry-wide the figure is 11 of 70 — so this one category, a fifth of the registry, holds nearly half of every undisclosed row in it.

Vendor

What the sub-processor page actually says

Cognition (Windsurf)

windsurf.com/subprocessors returns HTTP 200 with the body "This page does not exist"

Replit

The page exists, is linked from the footer, and names categories — "cloud infrastructure, payments processors, analytics providers" — not entities

Augment Code

"For more details about our subprocessors or Data Processing Agreement (DPAs), please contact us"

Poolside

The DPA defines an "Authorized Subprocessor" approval mechanism; no roster accompanies it

Qodo

trust.qodo.ai returns HTTP 403 to every automated path attempted

Four of those five returned HTTP 200. These are not blocked crawls or JS-rendered tables — the failure mode that produced two false undisclosed values in this registry on 2026-08-03 and forced a correction pass. They are pages that exist and have no names on them.

The comparison is the point. GitLab, Sourcegraph, JetBrains and AWS all publish complete tables with per-entity processing locations, and AWS commits to 30 days' notice before adding one. The split does not track company size, funding, or price. Sourcegraph and Augment are the same kind of company. One published the list; the other did not.

Being established in the Union says nothing about where the code goes

JetBrains is the cleanest test case the registry has found for the difference between establishment and data path. JetBrains s.r.o. is a Czech company. It says so on its own privacy pages, notes that it is therefore "directly regulated by the EU General Data Protection Regulation", and publishes its AI sub-processors better than almost anyone in this batch.

Read that published table and the AI features call:

  • OpenAI — US
  • Anthropic — US
  • xAI — US
  • Baseten — US, UK
  • Tavily — US
  • Google — EU, US and Asia

Five of six are US-only, and no EU-residency option for AI Assistant inference is published anywhere. The contractual protections are real and unusually strong — a written undertaking not to train on inputs, extended by contract to every AI subcontractor, plus SOC 2 Type II and a published DPA — but they are protections against use, not against jurisdiction. JetBrains scores C 46.

The counter-example is Mistral, at A 94 — and it does not get there by being French. It gets there by shipping Codestral, Codestral Embed, Devstral and Mistral Medium onto the customer's own GPUs, so the question of whose jurisdiction the inference sits in never arises. Sovereignty here is an architecture, not an address.

The open-source substrate is unmaintained

The three rows at the top of this batch are Continue (A 96), GitLab Duo Self-Hosted (A 96) and Mistral Vibe for Code (A 94). All three score where they do for the same reason: the model runs where the customer put it, so there is no vendor in the inference path to disclose.

Continue is the reference open-source coding agent — Apache-2.0 client, any model backend including a fully local one, with documented guides for running without internet access and for self-hosting a model. As of 2026-08-06 the continuedev/continue repository is read-only and carries the note "no longer actively maintained". Version 2.0.0 is described as a final release, and one of its changes was removing anonymous telemetry on the way out.

Mistral Code — now sold as Vibe for Code — is, in Mistral's own launch words, "built on the proven open-source project Continue". So the most sovereign commercial coding assistant available to a European buyer sits on an upstream that has stopped moving.

The registry does not deduct for this, and the reason matters: the rubric grades the data path, and an archived Apache-2.0 client has exactly the data path it had six months ago. It is recorded instead as a procurement condition on the row. A buyer standing Continue up in 2026 is adopting a fork, not a product, and needs an owner for it.

What the fifteen look like

Grade

Vendor — product

A 96

GitLab — Duo Self-Hosted

A 96

Continue Dev — Continue

A 94

Mistral AI — Vibe for Code

C+ 56

Tabnine Enterprise (private installation)

C+ 56

GitHub Copilot (Business / Enterprise)

C 54

Poolside Assistant (self-managed)

C 52

AWS — Q Developer Pro

C 49

Sourcegraph — Amp

C 46

JetBrains — AI Assistant / Junie

C 46

Anysphere — Cursor with Privacy Mode

D 38

Cognition — Windsurf

D 37

Anysphere — Cursor, default settings

D 34

Augment Computing — Augment Code

D 32

Qodo — Qodo Gen and Qodo Merge

F 20

Replit — Replit Agent

Every row in the A band is a deployment the customer controls. Every row below C sends the repository somewhere the customer cannot see, and four of the five will not say where.

Two things this batch deliberately did not do

Poolside carries opt-out-clear, not never. Nothing Poolside publishes says it will not train on customer source code. The only training statement in its entire legal estate is a privacy-policy clause saying collected personal data may be used for "training our machine learning models (unless you opt-out)". In a self-managed deployment the practical exposure is small — inference runs on the customer's own metal — but for a vendor whose whole pitch is deployment control, the missing commitment is the finding, and missing evidence scores zero here.

Qodo and Tabnine's SOC 2 claims score nothing. Both say "SOC 2" without naming a type, and neither trust centre could be read: trust.qodo.ai 403s automated requests, and no reachable Tabnine trust-centre path was found. Both are recorded as SOC 2 (type unverified) and flagged for a human with a browser, rather than converted into a claim in either direction. That is the same handling Cerebras got in batch 3, and it is the rule that keeps this registry from manufacturing findings out of a scraper's field of view.

What to check in your own stack

  1. Open your assistant's sub-processor page and count the company names on it. Not categories, not "industry-standard providers" — names, with a processing location beside each. Five of fifteen fail this.
  2. Ask where inference runs, separately from where the vendor is incorporated. A European supplier that brokers US model APIs gives you a European contract and a US data path. JetBrains at C 46 is the shape to look for.
  3. If the answer is "we don't train on your code", read the actual clause. Poolside's legal estate says the opposite of its positioning.
  4. If you are buying sovereignty, buy an architecture. Every row above 90 in this category runs the model on infrastructure the customer controls. None of them gets there by nationality.

Method: every value on every row traces to a first-party document fetched and snapshotted at grading time, stored under evidence/<source_id>/. Scores are computed by score.py from recorded field values with no judgment at scoring time; all 70 live rows recompute exactly. Rubric: methodology v1.2.

This article was researched and written during the registry's batch 5 harvest, with AI assistance and human review before publication. The hero image is AI-generated. Every figure is computed from the registry's live records. Corrections: see the methodology page.