Skip to content
Call, 0511 – 47 55 58 11

Data quality and discrimination risks in AI

Article 10 AI Act obliges providers of high-risk systems to run data governance that detects and mitigates bias. Under Article 26(4) deployers owe input data that matches the intended purpose and is sufficiently representative.

An AI system rarely discriminates because someone wanted it to. It discriminates because the data it learned from carries an existing inequality forward. Recital 67 of the AI Act names precisely that connection, and Article 10 draws the consequence. Providers of high-risk systems need data governance that looks for bias, prevents it and mitigates it.

Four criteria for training, validation and testing data

Where bias arises

In development

  • Skewed sampling, because certain groups are missing from the dataset
  • Biased or erroneous annotation of the data
  • A choice of features that describes the target variable crookedly
  • Technical constraints in the development process itself

In operation

  • Bias that only appears in the specific deployment context
  • Misreading of the outputs by the people using them
  • Drift caused by continued learning during operation
  • Input data that has shifted away from the provider assumptions

Bias becomes visible through statistical analysis. The AI Act leaves the choice of method open and requires only that the parameters used to measure potentially discriminatory effects be documented. Fairness metrics show whether particular groups are treated systematically differently. They do not replace the legal and ethical assessment, not least because different fairness goals are in part mathematically incompatible. Correction also regularly costs accuracy, and both have to be documented under Annex IV AI Act.

German equal treatment law applies to AI deployment unchanged, in employment as well as in mass-market civil transactions such as credit decisions, insurance pricing or the allocation of housing. What matters in practice is indirect discrimination. The protected characteristic is not part of the model, but a characteristic correlating with it is. Familiar examples are the postcode and gaps in a career history.

Section 22 AGG eases the burden of proof. Anyone presenting indicia of discrimination shifts that burden to the other side. Conversely it can be argued that a documented examination of the data for bias rebuts the presumption, while a breach of Article 10 AI Act can serve as an indicium. Those affected also have the right to an explanation of a high-risk system’s decision under Article 86 AI Act.

Selection, scoring or pricing with AI in the business?

We review the data for bias and the documentation for whether it would hold up.

Get in touch

Frequently asked questions

Does Article 10 AI Act apply to us if we only deploy the system?

The requirements of Article 10 AI Act address providers of high-risk systems. Deployers carry the separate duty under Article 26(4) AI Act to ensure that input data matches the intended purpose and is sufficiently representative, so far as they control it. Independently of the AI Act, Article 5(1)(d) GDPR requires personal data to be accurate.

Does training data have to be free of errors?

Only as far as possible. Absolute accuracy was recognised during the legislative process as unattainable and was qualified accordingly. That moves the process into the foreground, with annotation, labelling, cleaning and enrichment to be documented under Article 10(2)(c) AI Act. The same applies to completeness.

What does bias mean here?

The AI Act does not define the term. Recital 67 addresses bias inherent in the dataset, for example where historical data is used, and bias generated during implementation. In development the familiar categories include skewed sampling, biased labelling and the choice of describing features. In operation, bias can appear in the deployment context, in the interpretation of outputs, and through continued learning.

How does German equal treatment law fit in?

The AGG applies to AI deployment without qualification, above all in employment and in mass-market civil transactions. What matters in practice is indirect discrimination. The protected characteristic is not part of the model, but a characteristic correlating with it is, such as a postcode or gaps in a career history. Under section 22 AGG indicia suffice, after which the other side bears the burden of proof.

May we process data on ethnicity or gender to detect bias?

Article 10(5) AI Act contains a standalone legal basis for this, subject to strict cumulative conditions. It permits the processing of special categories of personal data so far as strictly necessary to detect and correct bias. Without those conditions Article 9 GDPR governs.

Related