Source-linked AI summary

Human Perceptions of Fairness in Algorithmic Decision Making: A Case Study of Criminal Risk Prediction

Nina Grgić-Hlača, Elissa M. Redmiles, Krishna P. Gummadi, Adrian Weller

arXiv:1802.09548v1stat.MLcs.CYcs.LG

TL;DR

Algorithmic fairness research has largely prescribed how fair decisions should be made, leaving limited descriptive evidence about how people perceive and reason about feature fairness. This paper surveys 576 people and proposes an eight-property framework for fairness judgments. In the COMPAS bail-decision scenario, the framework captures concerns beyond discrimination, reveals disagreement in property assessments, and supports a common fairness predictor with high accuracy.

  • Problem

    Existing algorithmic fairness studies are largely normative, while this paper investigates what people perceive as fair in context and the moral reasoning behind those perceptions.

  • Method

    The authors conduct scenario-based surveys with 576 people and model feature-use fairness through assessments of eight latent properties.

  • Results

    The study finds multidimensional fairness concerns beyond discrimination, substantial disagreement in property assessments, and strong consensus in reasoning from those assessments to fairness judgments.

  • Takeaways & Limitations

    Fairness research should consider concerns beyond discrimination and may benefit from more objective causal data combined with a common human fairness heuristic.

  • Takeaways & Limitations

    Self-report responses may be affected by self-report and generalizability biases, and fairness perceptions may differ across contexts.

Abstract

from arXiv · show

As algorithms are increasingly used to make important decisions that affect human lives, ranging from social benefit assignment to predicting risk of criminal recidivism, concerns have been raised about the fairness of algorithmic decision making. Most prior works on algorithmic fairness normatively prescribe how fair decisions ought to be made. In contrast, here, we descriptively survey users for how they perceive and reason about fairness in algorithmic decision making. A key contribution of this work is the framework we propose to understand why people perceive certain features as fair or unfair to be used in algorithms. Our framework identifies eight properties of features, such as relevance, volitionality and reliability, as latent considerations that inform people's moral judgments about the fairness of feature use in decision-making algorithms. We validate our framework through a series of scenario-based surveys with 576 people. We find that, based on a person's assessment of the eight latent properties of a feature in our exemplar scenario, we can accurately (> 85%) predict if the person will judge the use of the feature as fair. Our findings have important implications. At a high-level, we show that people's unfairness concerns are multi-dimensional and argue that future studies need to address unfairness concerns beyond discrimination. At a low-level, we find considerable disagreements in people's fairness judgments. We identify root causes of the disagreements, and note possible pathways to resolve them.

1 INTRODUCTION

The paper descriptively studies how people judge fairness in algorithmic feature use, addressing a gap in predominantly normative fairness research. Surveys of COMPAS bail-decision features reveal multidimensional concerns, disagreement in judgments, and a common reasoning pattern linking latent property assessments to fairness judgments.

  • Existing algorithmic fairness research largely prescribes fair decisions, whereas this paper empirically studies what people perceive as fair and the moral reasoning behind those perceptions.
  • The study asks whether using a feature F is fair in a decision-making scenario S, recognizing that fairness perceptions are multidimensional and context-dependent.
  • The authors survey 576 people about the fairness of using features from COMPAS, a criminal risk tool used to assist US bail decisions.
  • The proposed framework models fairness reasoning through eight latent feature properties, including concerns beyond discrimination such as privacy sensitivity and volitionality.
  • Most respondents judged half of the COMPAS features unfair to use, while the properties considered were mostly unrelated to discrimination.
  • Disagreement largely reflects differing assessments of latent properties, especially causal relationships, while a single simple classifier accurately predicts judgments from those assessments.
  • The findings suggest that objective data about latent properties might be combined with consistently elicited human moral reasoning to inform fair algorithmic decision making.
  • The study does not examine COMPAS uses in criminal sentencing or parole because those settings involve additional factors, including the societal role of long-term incarceration.

2 JUDGING FEATURE USAGE FAIRNESS

The framework treats fairness judgments as a two-part heuristic: people assess latent properties of a feature and then morally reason from those assessments. It includes eight proposed properties, several of which address causal and non-discrimination concerns.

  • People may use implicit or explicit assessments of underlying feature properties as a heuristic for judging whether feature use is fair.
  • The framework separates assessing eight latent properties from weighting those properties during individual moral judgments.
  • The properties were drawn from social, economic, political, moral, philosophical, and legal scholarship.
  • Reliability concerns whether a feature can be assessed reliably, while Causes Outcome concerns whether it may increase or mitigate risky behavior.
  • Causes Vicious Cycle concerns whether feature use may trap people in increasingly risky behaviors, illustrated through friends’ criminal history.
  • The eight properties captured users’ considerations, and six were statistically significant for predicting fairness judgments.

3 METHODOLOGY

The authors conducted online surveys in September and October 2017 to collect judgments about algorithmic fairness and the framework’s latent feature properties. The methodology received institutional ethics review approval.

  • The study used a series of online surveys conducted in September and October 2017.
  • The survey methodology was approved by the authors’ institutional ethics review board.

3.1 Survey Design

The study uses scenario-based surveys about COMPAS bail decisions, asking participants to rate feature-use fairness and assess eight latent properties. Randomization, pre-testing, attention checks, and separate pilot designs support questionnaire validity and analysis of reasoning.

  • Scenario and features: Participants evaluated a real-world scenario in which COMPAS information helps judges decide whether defendants can be released on bail.
  • Scenario and features: The scenario covered ten COMPAS questionnaire features, with feature presentation order randomized between respondents.
  • Pilot Survey 1: Pilot survey 1 asked participants to rate fairness on a 7-point Likert scale and select reasons using the eight latent properties or an Other option.
  • Pilot Survey 1: Each proposed property was used by at least 15% of respondents, while less than 3% selected Other; most Other responses mapped to one of the eight properties.
  • Pilot Survey 2: Pilot survey 2 omitted fairness questions and asked participants to rate the eight properties independently on the same 7-point Likert scale.
  • Latent properties: The eight assessments covered reliability, relevance, privacy, volitionality, causal effects on outcomes, vicious cycles, disparities, and sensitive-group membership.
  • Main Survey: The Main survey combined fairness and latent-property questions for each feature and randomized both their order and the order of features and properties.
  • Quality controls: Researchers pre-tested questionnaires through cognitive interviews with five demographically diverse participants, refined items iteratively, randomized question order, and added an attention check.

3.2 Survey Samples and their Demographics

The study used two U.S. respondent samples to balance response quality and demographic representativeness. AMT respondents differed from census demographics, while SSI respondents were broadly census-representative across several measures.

  • Survey samples: The main survey included 196 U.S. AMT master workers and 380 U.S. respondents recruited through SSI.The two samples were selected to address response quality and representativeness concerns.
  • AMT demographics: AMT respondents included 43% females, 76% Caucasians, 51% with at least a college degree, and 57% liberals.These figures differed considerably from the U.S. population benchmarks.
  • SSI demographics: SSI respondents were within 5% of census benchmarks for gender, education, and political leaning.The comparison benchmarks were 55% female, 32% with a BS or above, and 37% liberal.
  • SSI demographics: SSI respondents were 71% Caucasian, exceeding the population comparison, possibly because the race question did not allow multiple selections.The paper offers this as a possible explanation rather than a confirmed cause.

3.3 Analysis Methods

The analysis quantified agreement in fairness and latent-property ratings, then tested whether the eight-property framework could predict binary fairness judgments.

  • Consensus measurement: Consensus was measured as 1 − normalized Shannon entropy, where 1 indicates complete consensus and 0 indicates complete disagreement.The normalized entropy measure ranges from 0 to 1 before conversion into the reported consensus score.
  • Predictive analysis: A logistic regression classifier with L2 regularization predicted whether a feature was judged fair or unfair from respondents’ latent-property evaluations.Fair ratings included completely, mostly, slightly, and neutral; unfair ratings included completely, mostly, and slightly unfair.

3.4 Discussion of Limitations

The study acknowledges self-report and generalizability limitations while using controls and dual samples to assess potential bias. It also tests whether latent-property assessments were affected by asking fairness questions.

  • Self-report limitations: Self-report bias may affect the data, although the study attempted mitigation through pre-testing and question randomization.This limitation applies to the survey-based measurement strategy.
  • Measurement-bias check: KL-divergence was below 0.1 for 90% of questions and below 0.14 for the remaining 10% when comparing control and main-survey latent-property ratings.The paper interprets these values as evidence that fairness questions minimally affected latent-property assessments.
  • Generalizability: The study recruited both a census-representative population and an AMT sample because survey participants may not represent the general population.The two samples were intended to balance generalizability and data quality.
  • External validity: Future work should test whether models based on self-reported inputs align with real-world fairness perceptions in ecologically valid decision-making situations.The stated boundary concerns transferring survey-based perceptions to more realistic algorithmic settings.

4 ANALYZING FAIRNESS JUDGMENTS

Respondents judged half of the ten COMPAS features unfair despite none directly encoding race or gender, but consensus varied substantially across features and samples.

  • More than half of respondents judged five of ten COMPAS features unfair to use for criminal-risk prediction.
  • 95% and 94% agreed that Current Charges and Criminal History, respectively, were fair to use.
  • Personality and Criminal Attitudes produced little consensus, with neither fairness judgment receiving more than a slender 51% majority.
  • Respondents showed low consensus even for features rated least fair, including Education & School Behavior.
  • Feature rankings by mean fairness were similar across AMT and SSI samples, but SSI respondents generally reached less consensus.The paper suggests this may reflect SSI’s more-random and diverse population.

5 ANALYZING FAIRNESS REASONING

The authors explain fairness disagreement through assessments of eight latent feature properties and model judgments from those assessments. Causal properties were more controversial, while a single classifier predicted most respondents’ judgments accurately.

  • Respondents disagreed on latent-property assessments for at least one feature, especially properties concerning causal relationships and discrimination.
  • Consensus was above 0.5 for relevance, reliability, and privacy on at least some features, and aligned with consensus in fairness judgments.
  • 88% accuracy for AMT and 87% for SSI were achieved when predicting fairness judgments from respondents’ latent-property ratings.
  • The model made fewer mistakes for very unfair or very fair ratings but performed close to random for neutral ratings.
  • More than 85% of AMT respondents received predictions with at least 80% accuracy from the single classifier.
  • Volitionality, relevance, reliability, and causality increased fairness judgments, whereas privacy sensitivity and vicious cycles decreased them.
  • Relevance had the strongest effect: each one-point increase on its 7-point rating scale multiplied fair-rating odds by 2.47.
  • The framework’s odds ratios illustrate fairness reasoning in this scenario but are not expected to hold necessarily in other scenarios.

6 CONCLUDING DISCUSSION

The paper contrasts its descriptive study of fairness perceptions with prior normative work focused largely on discrimination. It finds that people consider additional feature properties and that causal disagreement challenges causal approaches to fairness.

  • Prior algorithmic-fairness research largely prescribes nondiscriminatory decisions, whereas this paper descriptively studies what people perceive as fair.
  • People’s unfairness concerns extend beyond discrimination to properties including feature relevance and reliability.
  • Disagreement about causal relationships between features and outcomes raises challenges for fairness approaches requiring a known causal structure.
Loading 1802.09548v1…