EU AI Act Art. 10 — Data and data governance
Article 10 is about the fuel: training, validation, and testing data. High-risk AI must be trained on data that is relevant, representative, complete, and as error-free as possible — with documented governance to prove it.
At a glance
What this article requires
- Training, validation, and testing data must be relevant, sufficiently representative, and free of errors 'to the best extent possible'.
- Governance practices must cover data design, collection, processing, and bias detection and correction.
- Data must be examined for biases that could affect health, safety, or fundamental rights.
- Special-category personal data (Art. 9 GDPR) can be processed to monitor and correct bias — but only under strict conditions and with safeguards.
- The data quality bar is measured against the 'accepted state of the art'.
Scope
Who this applies to
Providers of high-risk AI who train models on data. Deployers are indirectly affected: they should demand data documentation when procuring, and their own deployment data choices affect outcomes.
Obligations
What you must actually do
Build a data sheet
For each dataset: provenance, collection method, cleaning steps, known gaps, and how representativeness was assessed. This is the evidence market surveillance authorities will ask for first.
Run bias detection and correction
Examine datasets for bias that could harm people — protected attributes, demographic skew, proxy variables. Document what you found and what you did about it.
Handle special-category data carefully
Processing biometrics, health, ethnicity, or other Art. 9 GDPR data for bias monitoring is permitted only where strictly necessary, with safeguards and appropriate retention limits.
Action plan
Practical first steps
- 1
Inventory every dataset feeding the system and score it on relevance, representativeness, and error levels.
- 2
Write the bias-analysis methodology before you run it, so the results are credible.
- 3
Record decisions to use or drop datasets, with rationale — empty spots in the record read as gaps.
- 4
Put a data-owner in place with a refresh cadence, because stale data is a drift risk.
Penalty exposure
Article 10 failures sit in the general tier: up to €15 million or 3% of global annual turnover.
FAQ
Questions about Art. 10
Can I use personal data to fix bias in my hiring model?
Article 10(5) allows processing of special-category personal data strictly for detecting and correcting bias in high-risk AI, under Art. 9 GDPR conditions — proportionate, with safeguards, and only for that purpose.
What does 'sufficiently representative' mean in practice?
Your training data should reflect the population your system will actually see — for hiring AI, that means the applicant pool, not just your best-performing employees. Document the demographic coverage analysis.
Sources
Citations & further reading
Related
More article explainers
Wondering which articles apply to your AI?
Describe your system in the free Risk Scanner and get a preliminary risk read with the obligations that likely apply — in seconds.
Check my use casePreliminary EU AI Act clarity summary. Not legal advice.