Source-linked AI summary

Why Fairness Cannot Be Automated: Bridging the Gap Between EU Non-Discrimination Law and AI

Sandra Wachter, Brent Mittelstadt, Chris Russell

arXiv:2005.05906v1cs.AI

TL;DR

The paper examines the gap between context-sensitive EU non-discrimination law and statistical fairness measures designed for automated systems. It analyzes legal evidential requirements and fairness tests, then proposes conditional demographic disparity (CDD) as baseline statistical evidence aligned with the ECJ’s standard. CDD is intended to regularize assessment while preserving case-specific judicial interpretation.

  • Problem

    EU non-discrimination law uses contextual, intuitive, and case-specific evidential requirements, while automated discrimination is difficult to detect through the human signals and experiences traditionally supporting claims.

  • Method

    The paper reviews EU jurisprudence and existing machine-learning fairness measures, comparing them with procedures for assessing prima facie discrimination and developing CDD as a baseline measure.

  • Results

    CDD aligns with the ECJ’s ‘gold standard’ and reports the magnitude of disparity in an affected population rather than providing a binary pass/fail decision.

  • Takeaways & Limitations

    CDD can support consistent statistical assessment of potential automated discrimination while leaving contextual interpretation and final legal judgments to courts.

Abstract

from arXiv · show

This article identifies a critical incompatibility between European notions of discrimination and existing statistical measures of fairness. First, we review the evidential requirements to bring a claim under EU non-discrimination law. Due to the disparate nature of algorithmic and human discrimination, the EU's current requirements are too contextual, reliant on intuition, and open to judicial interpretation to be automated. Second, we show how the legal protection offered by non-discrimination law is challenged when AI, not humans, discriminate. Humans discriminate due to negative attitudes (e.g. stereotypes, prejudice) and unintentional biases (e.g. organisational practices or internalised stereotypes) which can act as a signal to victims that discrimination has occurred. Finally, we examine how existing work on fairness in machine learning lines up with procedures for assessing cases under EU non-discrimination law. We propose "conditional demographic disparity" (CDD) as a standard baseline statistical measurement that aligns with the European Court of Justice's "gold standard." Establishing a standard set of statistical evidence for automated discrimination cases can help ensure consistent procedures for assessment, but not judicial interpretation, of cases involving AI and automated systems. Through this proposal for procedural regularity in the identification and assessment of automated discrimination, we clarify how to build considerations of fairness into automated systems as far as possible while still respecting and enabling the contextual approach to judicial interpretation practiced under EU non-discrimination law. N.B. Abridged abstract

I. INTRODUCTION

The paper argues that EU non-discrimination law is too contextual and interpretive to automate directly, while AI makes discrimination harder to detect through intuition. It proposes standardized statistical evidence, centered on CDD, to regularize assessment without replacing case-specific judicial interpretation.

  • I. INTRODUCTION: AI discrimination is harder to identify intuitively because algorithms operate at scales and levels of complexity that can produce unfamiliar and unintuitive disparities.This weakens intuition as the primary mechanism for identifying potentially discriminatory actions.
  • I. INTRODUCTION: The paper proposes conditional demographic disparity (CDD) as a baseline statistical measure aligned with the European Court of Justice’s assessment standard.CDD is presented as a way to support consistent evidence procedures while preserving contextual judicial interpretation.
  • I. INTRODUCTION: European discrimination law relies on contextual, case-specific normative choices about groups, harm, and evidence that cannot be converted into a static automated test.The paper identifies these choices as matters for judicial interpretation rather than system design alone.
  • I. INTRODUCTION: Standardized statistical evidence could provide procedural regularity for judges, regulators, industry, and claimants without freezing fairness thresholds into code.The proposal seeks consistency in identifying and assessing potential discrimination, not uniformity in legal interpretation.

II. THE UNIQUE CHALLENGE OF AUTOMATED DISCRIMINATION

Automated discrimination weakens the intuitive experiences and comparative information that traditionally help people bring EU discrimination claims. AI can also generate novel proxies and disparities affecting groups that do not map neatly onto legally protected characteristics.

  • II. THE UNIQUE CHALLENGE OF AUTOMATED DISCRIMINATION: Automated discrimination is abstract, subtle, and intangible, making it difficult for victims to detect or prove disadvantage.Individuals may not realize they were disadvantaged, while system information may be limited by opacity or intellectual-property concerns.
  • II. THE UNIQUE CHALLENGE OF AUTOMATED DISCRIMINATION: The loss of intuitive or comparative experiences can prevent individuals from recognizing disadvantage and initiating a non-discrimination claim.Consumers may not know whether they received the best price or whether certain advertisements were withheld from them.
  • II. THE UNIQUE CHALLENGE OF AUTOMATED DISCRIMINATION: AI systems may use new, counterintuitive proxies for protected characteristics that are not readily detected.Their ability to process data at scale and find unfamiliar connections means discrimination need not resemble familiar human patterns.
  • II. THE UNIQUE CHALLENGE OF AUTOMATED DISCRIMINATION: Disparities may affect groups that do not correspond to legally protected characteristics, challenging the scope of existing non-discrimination law.The paper notes that such groups may experience disparity levels that would be considered discriminatory if applied to protected groups.
  • II. THE UNIQUE CHALLENGE OF AUTOMATED DISCRIMINATION: Complex AI systems are difficult to decompose into isolated rules, while claimants may lack information about their optimization conditions, decision rules, and outputs.This limits the ability to define the contested rule, comparator group, and relevant evidence.

III. CONTEXTUAL EQUALITY IN EU NON-DISCRIMINATION LAW

EU non-discrimination law deliberately preserves contextual equality through flexible, case-specific interpretation across varied legal frameworks and Member States. Claims require evidence of harm, protected-group impact, and disproportionate disadvantage, but the relevant assessments remain interconnected and context dependent.

  • III. CONTEXTUAL EQUALITY IN EU NON-DISCRIMINATION LAW: National implementation produces fragmented protection because the four EU non-discrimination directives establish minimum standards that vary across Member States.The directives also differ in the groups and sectors they protect.
  • III. CONTEXTUAL EQUALITY IN EU NON-DISCRIMINATION LAW: EU non-discrimination law distinguishes direct discrimination from indirect discrimination based on apparently neutral rules that disproportionately disadvantage protected groups.Indirect discrimination can reveal systematic and structural unfairness rather than only isolated individual treatment.
  • III. CONTEXTUAL EQUALITY IN EU NON-DISCRIMINATION LAW: A prima facie claim requires evidence of particular harm, significant manifestation within a protected group, and disproportionate impact compared with similarly situated people.The alleged offender may then justify the contested rule or refute the claim.

A. COMPOSITION OF THE DISADVANTAGED GROUP

EU law defines disadvantaged groups in relation to the contested rule and the facts of each case, making composition difficult to automate. This difficulty increases when a rule’s reach is unclear, as with online platforms and multinational companies.

  • Indirect discrimination requires showing that an apparently neutral provision significantly disadvantages a legally protected group.
  • The disadvantaged group is defined case by case, using broad or narrow traits according to the contested rule and its potential comparators.
  • The ECJ provides some consistency by linking group composition to the contested provision implemented by the alleged offender.
  • The contested rule’s reach determines who may be affected and therefore shapes the disadvantaged and comparator groups.
  • An ad hoc approach can work for intuitive discrimination but is difficult to apply when online platforms or multinational companies have uncertain reach.

1. Multi-dimensional discrimination

EU law has struggled to address discrimination based on multiple protected characteristics. The Parris judgment required age and sexual orientation to be assessed separately, a position criticized for its effects on intersectional and additive discrimination.

  • Multi-dimensional discrimination concerns unequal treatment based on more than one protected characteristic, including additive and intersectional forms.
  • In Parris, the Court held that combined age and sexual-orientation discrimination could not be established when neither ground caused discrimination in isolation.
  • The judgment was criticized for harming people affected by intersectional and additive discrimination, while relevant case law remains scarce.

B. COMPOSITION OF THE COMPARATOR GROUP

Comparator groups are necessary under EU non-discrimination law but are defined contextually rather than by clear universal rules. Socially constructed characteristics, fragmented protections, and uncertainty over hypothetical comparators complicate automation.

  • EU equality law requires comparing people in comparable situations, yet directives do not clearly specify how advantaged comparator groups should be composed.
  • Choosing legitimate comparators involves deciding who is similarly situated and whether a concrete or hypothetical comparator is sufficient.
  • Gender, ethnicity, and disability are socially constructed characteristics whose meanings and legitimate comparators depend on social context.
  • Protection for socially constructed characteristics varies across Member States because EU directives require national interpretation and transposition.
  • Comparator selection becomes especially difficult in multi-dimensional discrimination and cases where no binary comparator is apparent, such as age discrimination.
  • EU law permits hypothetical comparators, but their legitimacy varies across ECJ and Member State jurisprudence, including disputes involving national origin.

C. REQUIREMENTS FOR A PARTICULAR DISADVANTAGE

Establishing a particular disadvantage requires assessing the harm’s nature, severity, and significance, but EU law leaves their thresholds flexible and case-specific. The Court also allows different assessment times and has not consistently resolved whether hypothetical harm is sufficient.

  • A particular disadvantage concerns the harm’s nature, severity, and significance, including who is affected and how many protected-group members are disadvantaged.
  • EU non-discrimination law provides no consistent explicit thresholds for these elements, and significance may outweigh individual severity when many people are affected.
  • Although hypothetical harm can suffice in some jurisprudence, whether abstract harm is sufficient remains open and particular disadvantage appears practice-specific.
  • The Court’s reasoning varies on whether disadvantage requires a concrete comparison tied to a specific practice or rule.
  • EU law has not established clear-cut severity or significance thresholds, instead using flexible phrases such as “considerably more” or “far greater number.”
  • Discrimination may be assessed at different points in time, including when a law was enacted or when the discriminatory effect occurred.
  • The ECJ’s inconsistent requirements for evaluating evidence have produced divergent national standards, complicating automated fairness assessment.

E. A COMPARATIVE ‘GOLD STANDARD’ TO ASSESS DISPARITY

EU non-discrimination law’s comparative “gold standard” calls for assessing both disadvantaged and advantaged groups, but courts have applied this inconsistently. Full comparison is necessary to reveal the nature and magnitude of disparity, especially in intersectional cases.

  • The ECJ’s “best approach” compares the respective proportions meeting and failing a contested requirement within both disadvantaged and comparator groups.This approach assesses group composition and effects rather than relying only on affected-person counts.
  • The ECJ and national courts have not applied the full-comparison standard consistently, often examining only the disadvantaged group despite available comparator statistics.Examples include Gerster and other cases in which the courts did not consistently compare both groups.
  • Only examining the advantaged group can likewise fail to establish the full disparity relevant to a discrimination claim.The Court has sometimes followed this approach, including in cases where statistics concerning the disadvantaged group were available.
  • Assessing only the disadvantaged group can conceal intersectional or additive disparity and create the impression that all groups were treated fairly.Comparison with advantaged groups is a minimal requirement for revealing disparity affecting people belonging to more than one group.
  • A flexible and pragmatic legal test must accommodate complete comparisons between disadvantaged and advantaged groups when assessing prima facie discrimination.The paper presents this comparative procedure as the basis an automated assessment would need to facilitate without eliminating judicial flexibility.

F. REFUTING PRIMA FACIE DISCRIMINATION

After prima facie discrimination is established, the alleged offender may refute causation or justify indirect discrimination through a legitimate, necessary, and proportionate interest. These assessments remain context-dependent and open to competing evidence and judicial interpretation.

  • Once prima facie discrimination is established, the burden of proof shifts to the alleged offender to refute the claim.The shift may follow a convincing case, probable or likely causation, or refusal to provide relevant information.
  • For indirect discrimination, alleged offenders can refute a causal link or acknowledge differential results while providing a legitimate, necessary, and proportionate justification.Both parties may submit statistical evidence relevant to these arguments.
  • Examining the advantaged group can help refute causation, including by showing that protected-group members receive favorable outcomes in greater proportions than their workforce representation.The paper illustrates this with Bilka’s pension statistics.
  • The Court has also encouraged advantaged-group evidence as a possible defence in direct-discrimination cases such as Feryn and Accept.In those cases, employer statements concerning immigrants or gay players were assessed alongside potential evidence about outcomes.
  • The content of legitimate interests and acceptable measures depends on case context, national-court interpretation, and relevant legislation.National courts receive a high margin of appreciation, subject to legal limits.

IV. CONSISTENT ASSESSMENT PROCEDURES FOR AUTOMATED DISCRIMINATION

EU courts and national courts use statistical evidence to prove prima facie discrimination rarely and inconsistently, without standardized disparity thresholds. This creates a mismatch with automated systems, where subtle indirect discrimination is likely to make statistics increasingly important.

  • Statistical evidence is used rarely and inconsistently in EU discrimination jurisprudence, which lacks well-defined thresholds for illegal disparity across cases.Fairness has historically been specified contextually through judicial intuition rather than a coherent statistical procedure.
  • The absence of consistent legal requirements leaves developers, controllers, regulators, and users without clear standards for detecting, remedying, and preventing automated discrimination.The paper connects this gap to the inconsistent assessment of prima facie discrimination.
  • Automated systems are more likely to produce widespread and subtle indirect discrimination than direct discrimination, increasing the importance of statistical evidence.AI can also increase the capacity to identify new proxies for protected attributes.

A. TOWARDS CONSISTENT ASSESSMENT PROCEDURES

Automated systems should support, rather than replace, contextual judicial interpretation by producing consistent statistical evidence. The paper therefore advocates cross-disciplinary procedures that preserve normative authority while improving detection and assessment of automated discrimination.

  • Statistical evidence can support automated-discrimination assessment, but EU law’s contextual character prevents the necessary evidence from being specified universally.National laws often lack uniform tests, thresholds, or metrics for illegal disparity.
  • Normative questions about rule reach, group composition, harm, admissible evidence, and discrimination thresholds remain matters for judicial, legislative, or regulatory interpretation.The paper argues that system developers and controllers should not set these thresholds alone.
  • The technical community can contribute statistical evidence and fairness tools while preserving the judiciary’s democratic legitimacy to determine contextual equality.The proposed collaboration combines technical support with legal interpretation rather than replacing one with the other.
  • An “early warning system” should consistently generate statistical evidence for judicial assessment and for controllers’ systematic detection of potential discrimination.The paper frames this as a practical alternative to automatically detecting, evaluating, and correcting discrimination independently of local guidance.
  • Automated fairness procedures must accommodate the flexible and pragmatic assessment approach developed by EU courts.Judicial intuition suited obvious human discrimination less well than subtle, complex, and widespread algorithmic patterns.

A. STATISTICAL FAIRNESS TESTS IN EU JURISPRUDENCE

EU jurisprudence uses demographic disparity and negative dominance as related statistical approaches, but fixed evidential thresholds remain unsettled. Their requirements diverge especially when protected groups are minorities.

  • Scope: These tests are intended for broad scenarios, including settings without ground-truth data where accuracy-based group comparisons cannot be applied.
  • Demographic parity: Demographic parity requires the protected group's proportion to match across advantaged, disadvantaged, and overall populations.For example, if 35% of white applicants are admitted, parity requires 35% of black applicants to be admitted.
  • Negative dominance: Negative dominance requires that most of the disadvantaged group not belong to a protected class and that only a minority of the protected class be advantaged.The legal threshold may exceed a simple majority, such as 80% or 90%.
  • Comparing the tests: Negative dominance is harder to satisfy than demographic disparity, so every finding of negative dominance also establishes demographic disparity.
  • Comparing the tests: When the protected group comprises 50% of the population, demographic disparity and negative dominance are functionally equivalent.
  • Minority-group cases: Negative dominance becomes inappropriate for minority groups because its requirements can differ sharply from demographic disparity.In the Company B example, demographic disparity requires disadvantaging all white employees to match the disadvantaged black employees, whereas negative dominance requires only a matching number of white employees.

B. SHORTCOMINGS OF NEGATIVE DOMINANCE

Negative dominance can make legal protection depend on protected-group size and can permit disparity to be defended through subgroup fragmentation. Demographic disparity avoids these specific problems and is presented as the preferable statistical standard.

  • Judicial interpretation: Statistical testing cannot fully replace the intuition courts use when interpreting non-discrimination law.
  • Group size: Negative dominance can reduce protection for minorities because its requirements become more problematic as the protected group becomes smaller.The paper identifies this as particularly ill-suited to ethnicity and sexual orientation.
  • Divide and conquer: The divide-and-conquer strategy could justify disparity by splitting a protected group into subgroups too small to satisfy negative dominance.The paper gives intersectional subgroups of women as an example, even when aggregate testing would show negative dominance.
  • Demographic disparity: Demographic disparity does not create the same intersectional loophole because subgroup testing increases the alleged offender's burden.
  • Preferred standard: The authors recommend demographic disparity because it enables realistic comparisons aligned with the ECJ's statistical gold standard.

VI. CONDITIONAL DEMOGRAPHIC DISPARITY: A STATISTICAL ‘GOLD STANDARD’ FOR AUTOMATED FAIRNESS

The paper proposes conditional demographic disparity as a minimal statistical standard for automated discrimination cases. CDD supplies baseline comparisons for contextual judicial assessment rather than replacing that assessment.

  • Proposal: Conditional demographic disparity is proposed as a minimal statistical evidence standard for automated-system discrimination cases.
  • Role of automation: The paper redefines automating fairness as generating baseline evidence for external adjudication while preserving contextual, case-specific judicial interpretation.
  • Function: CDD can provide measurements that identify comparable protected groups and quantify their outcome disparities across the affected population.These measurements support examination of potentially illegal disparity and the composition of disadvantaged and comparator groups.
  • Definition: CDD compares outcome disparities between protected groups after conditioning on one or more additional attributes.For instance, equal admission rates among applicants with the same GPA would be required across protected groups.
  • Definition: Unlike a binary pass/fail test, the proposed CDD measure reports the magnitude of disparity to support contextual analysis.
  • Caveat: Searching every attribute combination can produce false positives, making an apparently unbiased system appear biased by chance.

A. CONDITIONAL DEMOGRAPHIC PARITY IN PRACTICE

The Berkeley admissions example shows how conditioning on department changes the apparent direction and interpretation of gender disparity. It also illustrates why statistical significance and legal significance require contextual judgment.

  • Berkeley’s aggregate admissions data showed strong gender disparity, but conditioning on department made the apparent bias disappear.The example uses departmental admissions statistics to demonstrate Simpson’s paradox.
  • Across Berkeley, women comprised 32% of the advantaged group and 46% of the disadvantaged group, satisfying demographic disparity but not the 50% negative-dominance threshold.
  • After conditioning on department, women sometimes constituted a greater proportion of rejected applicants than admitted applicants in individual departments.
  • Aggregated conditional statistics revealed a small bias in favour of women rather than against them.Bickel et al. and Freedman et al. reached consistent conclusions using different aggregation approaches.
  • Whether disparity is sufficiently serious to be illegal remains a contextual and political question, while very large samples can make practically trivial effects statistically significant.
  • CDD testing may be restricted to a single protected group when prior evidence indicates likely bias, but the choice remains contextual.

B. SUMMARY STATISTICS TO SUPPORT JUDICIAL INTERPRETATION

The paper recommends conditional demographic disparity as baseline descriptive evidence for automated discrimination cases. CDD is intended to support contextual legal assessment without replacing judicial decisions.

  • The authors recommend CDD as a baseline evidential standard for summary statistics assessing potential automated discrimination.
  • Descriptive summary statistics are preferred to statistical tests so semi-technical audiences can understand the magnitude and context of effects.
  • Prior tests based on negative dominance and non-conditional demographic disparity can incompletely capture the nature, severity, and significance of disparity.
  • CDD does not determine whether disparity is illegal or justified; it provides a roadmap for further contextual investigation.
  • A common CDD evidence baseline could align machine-learning fairness work with EU law and reduce the burden of producing relevant statistics.
  • CDD summary statistics can help identify disadvantaged groups, comparator groups, and particular disadvantages across affected populations.

VII. AUTOMATED FAIRNESS: AN IMPOSSIBLE TASK?

EU non-discrimination law treats fairness as contextual, while algorithmic systems produce unintuitive and difficult-to-detect disparities. The paper therefore supports systematic statistical evidence that assists, but does not automate, judicial interpretation.

  • European fairness law lacks a static, homogeneous framework suited to automated discrimination testing because its evidential requirements depend on context and judicial interpretation.
  • AI challenges existing protections because algorithmic discrimination lacks the human signalling mechanisms that can alert victims to discriminatory treatment.
  • The paper concludes that procedural regularity should improve detection while preserving contextual, case-specific judicial interpretation.
  • Current tests may enable a “divide and conquer” strategy that fragments intersectional disadvantaged groups below legal significance thresholds.
  • CDD offers a statistical baseline for identifying potential disparities while leaving comparator selection, thresholds, justification, and other normative judgments to courts.

APPENDIX 1 – MATHEMATICAL PROOF FOR CONDITIONAL DEMOGRAPHIC PARITY

The proof shows that conditional demographic parity yields equal proportions of advantaged people across groups after conditioning on selected factors. It establishes this by defining group counts, applying the parity relation, and algebraically rearranging it.

  • The proof considers a group defined by attributes R and represents advantaged and disadvantaged people with group-specific counts.The notation distinguishes advantaged black people (B_A), advantaged white people (W_A), disadvantaged black people (B_D), and disadvantaged white people (W_D).
  • Starting from conditional demographic parity, the proof applies the relation for any choice of R.
  • Assuming B_A and B_D are nonzero, the proof inverts both sides of the relation before continuing the algebraic rearrangement.
  • The resulting equality states that advantaged black and white people occur in the same proportion, denoted k_R%.
Loading 2005.05906v1…