EU AI Act · Article 10

Prove your AI training data's provenance.

The EU AI Act asks a hard question of every high-risk model: what data trained it, where did that data come from, and can you prove it hasn't changed? AT-1 seals a dataset so you can answer all three byte-exactly — which data, whose lineage, how much is synthetic, and provably unaltered — years after the model shipped.

What the Act asks of your data

Article 10 — data governance

High-risk AI systems must use training, validation and test data that is governed and documented — relevant, representative, and traceable. You have to be able to show what data went in.

Which exact data trained this model?

Not 'roughly this dataset' — the specific bytes. AT-1 seals a dataset to a hash, so you can prove byte-exactly which data produced a given model version.

Lineage & composition

Where did each part come from, and how much of it is synthetic? AT-1 records provenance and the real-vs-synthetic split, sealed alongside the data.

Unaltered, and provable later

An auditor two years from now can verify the training set is byte-for-byte the one you documented — or see, precisely, that it isn't.

How AT-1 makes it provable

  • Seal the dataset. Each training/validation/test set becomes a verifiable archive bound to a content hash — a byte-exact fingerprint of exactly what you trained on.
  • Record lineage & composition. Provenance of each part, and the real-vs-synthetic split, sealed with the data — not kept in a side spreadsheet that can drift.
  • Verify years later. Anyone with the archive can confirm it is byte-for-byte the documented dataset — or pinpoint that it isn't. Independent of us.
  • Bind data to model. Store the dataset hash with the model version for a tamper-evident link between what went in and what shipped.

Honest scope: AT-1 is a technical evidence layer, not a conformity assessment. It does not make a high-risk AI system compliant on its own — it makes the training-data governance, provenance and integrity obligations demonstrable and auditable. We are not your legal advisor.

EU AI Act data questions, answered

How does AT-1 help with EU AI Act data governance (Article 10)?
AT-1 seals your training, validation and test datasets into verifiable archives. Each is bound to a cryptographic hash, so you can prove exactly which data trained a model, demonstrate its lineage and real-vs-synthetic composition, and let an auditor verify years later that the dataset is byte-for-byte the one you documented. It is the technical evidence layer under an Article 10 data-governance programme.
Can I prove which exact dataset trained a specific model version?
Yes. A sealed dataset has a stable content hash; record it with the model version and you have a byte-exact, tamper-evident link between data and model. That turns 'we think it was this data' into documentary proof — the kind of traceability high-risk AI conformity assessments ask for.
How do you show how much training data is synthetic?
AT-1's corpus governance records the composition — which records are real, which are synthetic, and where each came from — sealed alongside the data so the split is provable, not asserted. Relevant both to Article 10 representativeness and to disclosure obligations around synthetic data.
Does this make my AI system EU AI Act compliant?
No — and avoid any tool that claims it does. Conformity for high-risk AI is a full programme (risk management, technical documentation, human oversight, and more). AT-1 makes one hard, auditable part demonstrable: that your training data is governed, documented, provenance-tracked and unaltered.

Building a high-risk model? Let's make its training data provable.

Get in touch