Prove your AI training data's provenance.
The EU AI Act asks a hard question of every high-risk model: what data trained it, where did that data come from, and can you prove it hasn't changed? AT-1 seals a dataset so you can answer all three byte-exactly — which data, whose lineage, how much is synthetic, and provably unaltered — years after the model shipped.
What the Act asks of your data
Article 10 — data governance
High-risk AI systems must use training, validation and test data that is governed and documented — relevant, representative, and traceable. You have to be able to show what data went in.
Which exact data trained this model?
Not 'roughly this dataset' — the specific bytes. AT-1 seals a dataset to a hash, so you can prove byte-exactly which data produced a given model version.
Lineage & composition
Where did each part come from, and how much of it is synthetic? AT-1 records provenance and the real-vs-synthetic split, sealed alongside the data.
Unaltered, and provable later
An auditor two years from now can verify the training set is byte-for-byte the one you documented — or see, precisely, that it isn't.
How AT-1 makes it provable
- Seal the dataset. Each training/validation/test set becomes a verifiable archive bound to a content hash — a byte-exact fingerprint of exactly what you trained on.
- Record lineage & composition. Provenance of each part, and the real-vs-synthetic split, sealed with the data — not kept in a side spreadsheet that can drift.
- Verify years later. Anyone with the archive can confirm it is byte-for-byte the documented dataset — or pinpoint that it isn't. Independent of us.
- Bind data to model. Store the dataset hash with the model version for a tamper-evident link between what went in and what shipped.
Honest scope: AT-1 is a technical evidence layer, not a conformity assessment. It does not make a high-risk AI system compliant on its own — it makes the training-data governance, provenance and integrity obligations demonstrable and auditable. We are not your legal advisor.
EU AI Act data questions, answered
- How does AT-1 help with EU AI Act data governance (Article 10)?
- AT-1 seals your training, validation and test datasets into verifiable archives. Each is bound to a cryptographic hash, so you can prove exactly which data trained a model, demonstrate its lineage and real-vs-synthetic composition, and let an auditor verify years later that the dataset is byte-for-byte the one you documented. It is the technical evidence layer under an Article 10 data-governance programme.
- Can I prove which exact dataset trained a specific model version?
- Yes. A sealed dataset has a stable content hash; record it with the model version and you have a byte-exact, tamper-evident link between data and model. That turns 'we think it was this data' into documentary proof — the kind of traceability high-risk AI conformity assessments ask for.
- How do you show how much training data is synthetic?
- AT-1's corpus governance records the composition — which records are real, which are synthetic, and where each came from — sealed alongside the data so the split is provable, not asserted. Relevant both to Article 10 representativeness and to disclosure obligations around synthetic data.
- Does this make my AI system EU AI Act compliant?
- No — and avoid any tool that claims it does. Conformity for high-risk AI is a full programme (risk management, technical documentation, human oversight, and more). AT-1 makes one hard, auditable part demonstrable: that your training data is governed, documented, provenance-tracked and unaltered.
Building a high-risk model? Let's make its training data provable.
Get in touch