If you build a high-risk AI system, two obligations sit at its core: your data must be governed, and your system must be documented. In plain terms, you have to be able to show that the data you trained and tested on was suitable and checked for bias, and you have to keep detailed records that let a regulator understand your system without taking it apart. "We have lots of data" is not compliance. Neither is "it works well." You need written evidence of both. This falls on the builder of the system, so if you only use a high-risk tool someone else made, most of this is their job, not yours.
In the last episode we explained that high-risk AI systems carry the heaviest obligations, and that the heavy ones sit with the provider, the company that builds the system, rather than the deployer who merely uses it. This week we open up two of those provider obligations in detail, because they are where good engineering teams most often assume they are compliant when they are not. They are the rules on data, and on documentation.
Rule one: govern your data
The first obligation is about the data your high-risk system learns from and is tested on. The rule requires that this data, the training, validation, and test sets, be relevant, representative, as free of errors as possible, and complete, and suited to the context where the system will actually be used. Behind that plain sentence sits real work.
In practice it means you must be able to show several things: where your data came from and how it was collected, what you did to prepare it, cleaning, labeling, annotation, and that you examined it for bias and acted on what you found. It also means articulating what your data actually represents, and just as importantly, what it does not. A dataset that looks fine technically can still be unsuitable if it does not match the real-world context the system operates in. The common trap is assuming that having a large amount of data, or having data lineage, is the same as governing it. It is not. Governance means documented, deliberate choices you can defend, not data that simply happens to exist.
Rule two: document your system
The second obligation is technical documentation. A high-risk system must come with detailed records, laid out in the Act's technical annex, that allow a regulator to understand and assess the system without having to reverse-engineer it. This documentation is the backbone of the formal check a high-risk system must pass before going to market, so it is not an afterthought; it is what proves everything else.
The documentation covers what the system is and does, how it was built, the data governance just described, the risk-management and human-oversight measures from earlier episodes, and how the system performs. And crucially, it is not a one-time document. It has to be kept current across the system's whole life. If you retrain the model, you update the records. If you discover a new bias or problem, you document it and what you did about it. The test to keep in mind is simple: could someone outside your team understand and evaluate your system from your documents alone? If not, the documentation is not yet doing its job.
First, "it performs well" is not compliance. A model can score highly and still fail these rules, because the rules are about governed data and evidence, not just accuracy. Second, buying a compliant component does not automatically make your system compliant, if you build a high-risk system on top of a third-party model, you are the provider and you carry these data and documentation duties. Your vendor's contract does not transfer them to them. And the one that catches almost everyone: a GDPR data agreement does not satisfy these AI Act data rules. GDPR governs personal data; the AI Act adds separate, AI-specific data governance. The two apply at the same time, and meeting one does not mean you have met the other.
If you only use the system
As in the last episode, most businesses are users, not builders, and this is lighter for you. If you deploy a high-risk system someone else built, the heavy data-governance and technical-documentation work is the provider's responsibility. Your related duties are to use the system according to its instructions, keep the records and logs it generates, and hold on to the information the provider gives you. You are not expected to reconstruct the provider's datasets or write their technical file. But you should confirm the provider has done their part, because if they have not, a high-risk system without proper data governance and documentation is a system you cannot stand behind.
What is at stake
These obligations are serious both legally and practically. Failing them sits in the penalty band that can reach millions of euros or a percentage of worldwide turnover, and inadequate documentation is one of the most common ways a high-risk system fails its assessment. But the deeper point is that data governance and documentation are what let you defend your system if it ever makes a harmful or contested decision. Without them, you cannot show why the system behaved as it did, or that you built it responsibly. With them, you can. The paperwork is not bureaucracy for its own sake; it is the evidence that protects your business.
Your action this week
For any high-risk system you build, ask two blunt questions. First, on data: can we show, in writing, where our training and test data came from, how we prepared it, and that we checked it for bias? Second, on documentation: could an outsider understand and assess our system from our records alone, and are those records kept up to date when the system changes? If the answer to either is no, that is your priority, and for anything near the line, get professional help, because this is detailed work. If you only use a high-risk system built by someone else, your action is simpler: confirm the provider has done this, and keep the documentation and logs they give you. Next week, we step back from the high-risk tier to something every business can do now: building an inventory of the AI systems you use.
Frequently asked questions
What does data governance mean under the EU AI Act?
For high-risk AI, it means your training, validation, and test data must be relevant, representative, as error-free as possible, complete, and suited to the system's real context, and you must document all of this. That includes where the data came from, how it was prepared, and evidence that you examined and addressed bias. Simply having a large dataset, or data lineage, does not count; governance means documented, defensible choices.
What technical documentation does a high-risk AI system need?
Detailed records, defined in the Act's technical annex, that let a regulator understand and assess the system without reverse-engineering it. They cover what the system is and does, how it was built, its data governance, its risk-management and human-oversight measures, and its performance. The documentation must be kept current: if you retrain or discover a problem, you update it. It is the backbone of the conformity check.
I use a third-party AI model. Am I responsible for its data governance?
If you build a high-risk system on top of that model, yes, you become the provider of the high-risk system and carry the data-governance and documentation duties. Your contract with the model vendor does not transfer these to them. If you are only a deployer using a finished high-risk tool, the heavy work is the provider's; your job is to use it correctly and keep the records and information they supply.
Does GDPR compliance cover the AI Act data rules?
No. This catches many teams. GDPR governs the processing of personal data; the EU AI Act adds separate data governance obligations specific to AI systems, covering training-data quality, bias, and documentation. The two frameworks apply at the same time, and a GDPR-compliant data agreement does not satisfy your AI Act Article 10 obligations, or the reverse. You have to meet both.
Isn't a model that performs well already compliant?
No. Strong performance does not equal compliance. These rules are about governed data and documented evidence, not accuracy alone. A model can score highly and still fail, if you cannot show where the data came from, that you checked it for bias, and that your system is properly documented. Compliance is proven on paper, not just in results.
Follow the series, get compliant one rule at a time
This is Part 6 of our weekly guide to AI regulation, breaking down one rule at a time so compliance feels manageable instead of overwhelming. Explore more clear, honest guides on AISetApp and follow along each week.
Explore more on AISetApp- EU AI Act (Regulation (EU) 2024/1689), Article 10 (data and data governance) and Article 11 with Annex IV (technical documentation)
- 2026 practitioner guides on Article 10 data governance and Article 11 documentation from NeuralTrust, AI Governance Desk, Vigilia, and Teleport
- Analyses on the overlap between GDPR and AI Act data obligations, and on provider versus deployer responsibility for third-party models, 2026
Reviewed September 2026. This is an explainer, not legal advice. The law is evolving; verify specifics with a qualified professional before acting.
Researched and drafted with AI assistance, reviewed and edited by Yasser El Hardouz, who takes editorial responsibility for this article.