Source-linked AI summary

Fairness Perceptions of Algorithmic Decision-Making: A Systematic Review of the Empirical Literature

Christopher Starke, Janine Baleis, Birte Keller, Frank Marcinkowski

arXiv:2103.12016v1cs.HCcs.AIcs.CY

TL;DR

Unfair algorithmic decision-making can harm individuals and social groups, motivating closer attention to perceived fairness. This systematic review synthesizes empirical research on algorithmic fairness and finds that the evidence is drawn almost exclusively from Western democracies, supporting calls for broader research and society-in-the-loop approaches.

  • Problem

    Biased input data or faulty algorithms can produce unfair ADM systems that denigrate certain members of society.

  • Method

    The paper conducts a systematic literature review that sheds light on theoretical concepts of fairness and the empirical literature on perceived algorithmic fairness.

  • Results

    Over 25,000 unique observations of citizens’ fairness perceptions of ADM are synthesized, with fairness evidence collected almost exclusively in Western democracies.

  • Takeaways & Limitations

    The findings support more research from non-Western contexts and consideration of human intervention throughout the ADM decision cycle.

  • Takeaways & Limitations

    The reviewed insights were drawn almost exclusively from Western democracies, predominantly the US.

Abstract

from arXiv · show

Algorithmic decision-making (ADM) increasingly shapes people's daily lives. Given that such autonomous systems can cause severe harm to individuals and social groups, fairness concerns have arisen. A human-centric approach demanded by scholars and policymakers requires taking people's fairness perceptions into account when designing and implementing ADM. We provide a comprehensive, systematic literature review synthesizing the existing empirical insights on perceptions of algorithmic fairness from 39 empirical studies spanning multiple domains and scientific disciplines. Through thorough coding, we systemize the current empirical literature along four dimensions: (a) algorithmic predictors, (b) human predictors, (c) comparative effects (human decision-making vs. algorithmic decision-making), and (d) consequences of ADM. While we identify much heterogeneity around the theoretical concepts and empirical measurements of algorithmic fairness, the insights come almost exclusively from Western-democratic contexts. By advocating for more interdisciplinary research adopting a society-in-the-loop framework, we hope our work will contribute to fairer and more responsible ADM.

1 Heinrich Heine University Düsseldorf

The paper concerns algorithmic decision-making, fairness perceptions, justice, discrimination, and literature review.

  • The review focuses on fairness perceptions and justice in algorithmic decision-making, including discrimination.

1 Introduction

Algorithmic decision-making increasingly affects important decisions and can produce discriminatory harms, motivating empirical attention to citizens’ fairness perceptions. This review synthesizes 39 empirical studies, using systematic searches and coding across four dimensions.

  • Algorithmic decision-making increasingly shapes important decisions across domains including public administration, lending, law, and medicine.
  • Biased data or faulty algorithms can reinforce racial and gender stereotypes, marginalize minorities, and denigrate members of society.
  • These societal implications require empirical understanding of when and why citizens perceive algorithmic decision-making as fair or unfair.
  • The review synthesizes 39 empirical studies and over 25,000 unique observations of citizens’ fairness perceptions of algorithmic decision-making.
  • The authors combine systematic online and citation searches, screen more than 2,500 entries, and use two coding steps to identify relevant literature.
  • The literature is systemized by algorithmic predictors, human predictors, human-versus-algorithmic comparisons, and consequences of algorithmic decision-making.

2 Bias in Algorithmic Decision-Making

Algorithmic systems may reduce some human biases but can also reproduce or amplify societal bias and create harms through data, model design, or implementation. The cited examples span beneficial applications and damaging exclusions or misclassifications.

  • Algorithmic systems can potentially reduce human biases because they do not grow tired or inattentive and lack human agency.
  • Machine-learning applications have improved emergency response and refugee integration outcomes in cited examples.
  • Other AI-based systems have decreased fairness by excluding citizens from support programs or mistakenly reducing benefits.
  • ADM bias can arise during data collection, processing, model selection, design, or specification, with historical data reproducing existing social biases.
  • Models may perform fairly on some tasks but unfairly on others, so bias reduction is not merely a technical challenge.
  • Implementation can create harms through privacy violations or using algorithms for sensitive decisions that should not be automated.

3 Concepts of Algorithmic Fairness

Algorithmic fairness encompasses competing formal and social-science concepts, whose incompatibility makes social context central to evaluation. The literature distinguishes multiple fairness dimensions and emphasizes how inequalities are produced, not only distributed.

  • Formal fairness definitions include statistical, similarity-based, causal, and preference-based approaches, each reflecting different fairness notions.
  • Many fairness conceptions are incompatible with one another, creating trade-offs that require attention to social context.
  • Algorithmic unfairness should consider how inequality is produced, including historical structural injustices against minorities, rather than only unequal distribution.
  • Organizational-justice approaches distinguish distributive, procedural, informational, and interpersonal fairness in algorithmic decision-making.
  • Interpersonal fairness includes refraining from protected data use and respecting privacy rights.
  • Social-science concepts such as equality of resources and capability of functioning have not been adequately addressed in machine-learning literature.
  • Because ADM systems operate within societies, fair predictions require calibration to specific social contexts and a human-centric perspective.
  • This paper reviews the growing empirical literature on human perceptions of algorithmic fairness.

4 Method

The review used a predefined, transparent systematic-review procedure to identify and assess empirical studies of perceived algorithmic fairness. Searches across academic databases and gray literature were followed by staged screening, citation-based expansion, and inter-rater reliability testing.

  • The review followed seven systematic-review steps, from defining the question and inclusion criteria through searching, screening, evaluation, synthesis, and dissemination.
  • 4.1 Establishing the research question: The PICOC framework organized the question into population, intervention, comparison, outcome, and context components.
  • 4.1 Establishing the research question: The research question asked how individuals perceive the fairness of algorithmic decision-making, without narrowing the country or domain context.
  • 4.2 Inclusion Criteria: Searches combined algorithmic-decision and fairness-related terms across Web of Science, PsycINFO, IEEE Xplore, and Scopus, supplemented by pearl-growing and gray-literature searches.
  • 4.3 Comprehensive Literature Search: 4,045 contributions were identified, 2,467 remained after duplicate filtering, and 99 potentially relevant articles came from manual gray-literature searches.

5 Results

The reviewed evidence spans diverse concepts, methods, domains, and human and algorithmic predictors of perceived fairness. However, the literature is concentrated in Western democracies and shows inconsistent effects across contexts and fairness criteria.

  • The review synthesizes empirical results on individuals’ perceptions of algorithmic fairness across multiple domains and disciplines.
  • 5.1 Descriptive Results: 23 studies were conducted in the United States, two each in the United Kingdom and Netherlands, and single studies in Germany, South Korea, and China.
  • 5.1 Descriptive Results: Seven studies used qualitative methods, 22 used quantitative methods, and ten used mixed-method designs.
  • 5.1 Descriptive Results: 11 studies addressed criminal justice, especially pretrial risk assessment, while seven examined work-related decisions, particularly hiring.
  • 5.2 Concepts and measurements: The literature primarily measures perceived distributive fairness of outcomes, while concepts and measurements of algorithmic fairness remain highly heterogeneous.
  • 5.4 Algorithmic predictors: Fairness preferences varied across equality, equity, and efficiency, and respondents sometimes favored demographic parity or false-positive-rate equalization.
  • 5.5 Human predictors: Perceptions depended on features, domains, and context: sensitive inputs were often viewed as unfair, but support for their use increased when better minority outcomes were explained.
  • 5.4 Algorithmic predictors: Transparency effects were ambiguous: one study found increased fairness perceptions, whereas another found no significant effect from different transparency levels.

6 Discussion

The review finds that perceived algorithmic fairness is highly context-dependent and that the empirical literature lacks coherent theoretical and measurement foundations. It identifies major scope gaps and calls for diversified, harmonized, interdisciplinary research involving broader societal stakeholders.

  • Context dependence: Perceived algorithmic fairness depends on algorithmic design, application area, and task characteristics, including whether decisions are high- or low-stakes.
  • Theoretical groundwork: Existing fairness theories are used inconsistently, and people may evaluate algorithmic and human decisions using different factors.
  • Research diversification: The evidence base is dominated by studies from Western democracies, predominantly the United States, limiting the generalizability of its insights.
  • Research diversification: Fairness perceptions vary considerably across algorithmic designs and application areas, making systematic cross-domain and cross-task comparisons necessary.
  • Measurement: Measurement harmonization could make findings more comparable and support more nuanced interpretations of perceived algorithmic fairness.
  • Society-in-the-loop: Society-in-the-loop research should involve multiple stakeholders in system design, ethical standards, and institutional and legal frameworks.

7 Conclusion

The review consolidates empirical research on perceived algorithmic fairness across four dimensions and highlights the need for broader, more coordinated inquiry. Its conclusion emphasizes non-Western research, harmonized concepts and measurements, and interdisciplinary society-in-the-loop approaches.

  • Motivation: ADM’s increasing penetration across society has intensified concerns about the fairness of algorithmic systems.The authors frame fairness perceptions as relevant to a human-centric approach to designing and implementing ADM.
  • Review synthesis: The review crystallizes insights from 39 empirical studies along four dimensions of perceived algorithmic fairness.These dimensions cover algorithmic predictors, human predictors, comparative effects between human and algorithmic decision-making, and consequences of ADM.
  • Research gaps: The authors call for more research from non-Western contexts because existing insights are geographically limited.The conclusion specifically identifies non-Western contexts as a priority for future research.
  • Research gaps: Further theoretical and methodological groundwork is needed to harmonize concepts and measurements of algorithmic fairness perceptions.The review identifies conceptual and measurement alignment as a remaining research need.
  • Future direction: The authors advocate interdisciplinary research adopting a society-in-the-loop framework to support fairer and more responsible ADM.This recommendation links the framework to the paper’s broader goal of improving algorithmic decision-making.

Perceived fairness of algorithmic

The reviewed studies operationalize perceived algorithmic fairness through diverse fairness concepts and measurement approaches. Measures include distributive, procedural, interactional, informational, substantive, and group fairness, using single-item, multi-item, content-analysis, and other designs.

  • Measurement approaches included single-item measures, short multi-item Likert scales, four-item scales, and multidimensional instruments.
  • Studies measured multiple fairness dimensions, including distributive, procedural, interactional, informational, substantive, and group fairness.
  • Fairness concepts varied from equality, equity, and meritocratic distribution to demographic parity, individual fairness, and equality of opportunity.
  • The overview includes studies examining algorithmic decision-making in contexts such as government decision-making, loan decisions, and criminal-justice risk assessment.
  • Some studies did not measure fairness perceptions directly, instead examining decision-process preferences, perceived advertising problems, comprehension, model approval, or perceived predictors.
Loading 2103.12016v1…