Source-linked AI summary

Factors Influencing Perceived Fairness in Algorithmic Decision-Making: Algorithm Outcomes, Development Procedures, and Individual Differences

Ruotong Wang, F. Maxwell Harper, Haiyi Zhu

arXiv:2001.09604v1cs.HC

TL;DR

The paper addresses limited evidence on what shapes people’s perceptions of algorithmic fairness, a question relevant to acceptance of systems affecting people’s lives. It uses a randomized between-subjects MTurk experiment manipulating outcomes, group error rates, development procedures, and participant characteristics. Favorable personal outcomes strongly increase perceived fairness, exceeding the effect of described demographic-group bias, and the paper argues that user-feedback evaluations should account for outcome favorability bias.

  • Problem

    Research has focused on building formally fair algorithms, but perceived fairness in practical algorithmic decision-making has not been systematically studied.

  • Method

    A randomized between-subjects MTurk experiment manipulated personal outcomes, demographic-group error rates, development procedures, and participant characteristics.

  • Results

    Favorable personal outcomes increased fairness ratings more than describing strong biases against particular demographic groups.

  • Takeaways & Limitations

    Systems using user feedback to evaluate algorithmic fairness should account for outcome favorability bias.

  • Takeaways & Limitations

    The study used fixed bias-rate values and a simplistic error-rate representation, so other rates or bias measures may change the relative importance of bias.

Abstract

from arXiv · show

Algorithmic decision-making systems are increasingly used throughout the public and private sectors to make important decisions or assist humans in making these decisions with real social consequences. While there has been substantial research in recent years to build fair decision-making algorithms, there has been less research seeking to understand the factors that affect people's perceptions of fairness in these systems, which we argue is also important for their broader acceptance. In this research, we conduct an online experiment to better understand perceptions of fairness, focusing on three sets of factors: algorithm outcomes, algorithm development and deployment procedures, and individual differences. We find that people rate the algorithm as more fair when the algorithm predicts in their favor, even surpassing the negative effects of describing algorithms that are very biased against particular demographic groups. We find that this effect is moderated by several variables, including participants' education level, gender, and several aspects of the development procedure. Our findings suggest that systems that evaluate algorithmic fairness through users' feedback must consider the possibility of outcome favorability bias.

INTRODUCTION

The paper examines how algorithm outcomes, development procedures, and individual differences shape perceived fairness, addressing a gap between formal fairness research and behavioral judgments. It uses a randomized MTurk experiment manipulating algorithm descriptions, group error rates, and personally favorable or unfavorable outcomes.

  • Research questions: The study investigates individual and group outcomes, development procedures, and evaluator characteristics as factors influencing perceived fairness.Group bias is operationalized through differing error rates, whereas unbiased treatment uses similar error rates across demographic groups.
  • Method: The researchers conducted a randomized online MTurk experiment in which workers evaluated an algorithm determining eligibility for a Masters Qualification.Participants judged an algorithmic decision relevant to their own workplace context.
  • Method: Participants received algorithm descriptions varying development procedures and demographic-group error rates, then were randomly shown either a pass or fail output before rating fairness.The study included manipulation checks and a debriefing explaining that the decision was hypothetical and randomly generated.
  • Findings: Favorable personal outcomes and unbiased group error rates increased perceived fairness, with favorable outcomes having the larger effect.The outcome effect exceeded the negative effect of describing strong demographic-group biases.
  • Research motivation: Algorithmic fairness research has emphasized formal constraints, while perceived fairness in practical decision-making has not been systematically studied.The paper connects theoretical fairness-aware machine learning with behavioral research on how affected stakeholders judge algorithms.

• Algorithm Outcomes:

The study operationalizes algorithm outcomes and group treatment through binary indicators for unfavorable versus favorable individual outcomes and biased versus unbiased error rates.

  • Algorithm Outcomes: Unfavorable outcome is coded 1, while favorable outcome is coded 0.
  • Algorithm Outcomes: Biased treatment is coded 1, while unbiased treatment is coded 0.
  • Algorithm Development: Algorithm development varies transparency, design team, model type, and decision process.The design conditions include outsourced, computer-scientist, and mixed teams; model conditions contrast machine learning with rules; decision conditions vary algorithm-only versus mixed decisions.

• Individual differences:

Individual-difference measures include education, computer literacy, age, gender, and race, using grouped or coded participant characteristics for analysis.

  • Individual differences: Education is grouped into above Bachelor’s degree, Bachelor’s degree, and below Bachelor’s degree, with above Bachelor’s degree as the baseline.
  • Individual differences: Computer literacy is measured using eight questions assessing computer skills, familiarity, and knowledge.Six questions use a 7-point scale, two use a 4-point scale normalized to 7 points, and responses are composited into a final score.
  • Individual differences: Age is grouped into below 25, between 25–45, and above 45, using between 25–45 as the baseline.
  • Individual differences: Gender analysis includes participants identifying as male or female.Participants could provide their preferred gender through an optional text input box.
  • Individual differences: Race is represented with five dummy variables based on participants’ selected racial categories.Each racial-group variable is coded 1 for belonging to that group and 0 otherwise.

Control Variables:

The analysis controls for participants’ expectations about whether they would pass the Master qualification before receiving the algorithmic outcome.

  • Control Variables: Self-expectation is coded 1 for participants expecting to pass and 0 for participants expecting to fail.

Participants

The study analyzed 579 responses after excluding 11 participants who failed an attention check, using regression models of perceived fairness.

  • 590 MTurk participants were recruited in the United States, and 579 responses remained after excluding 11 failed attention checks.Participants were recruited in January and February 2019; each received $1.50 after completing the experiment.
  • The sample included 153 participants expecting to fail and 426 expecting to pass the qualification test.
  • 292 participants received passes and 287 received fails, while 287 outcomes matched self-expectations and 292 did not.
  • Linear regression models predicted perceived fairness from algorithm outcomes, development procedures, and individual differences while controlling for self-expectation.The analyses reported coefficients, p-values, standard errors, R2, and adjusted R2 values.

(Un)favorable outcome vs. (Un)biased treatment

Both personal outcomes and group-level algorithm bias influenced perceived fairness, but favorable outcomes had the stronger effect. Unfavorable decisions reduced fairness ratings regardless of participants’ self-expectations.

  • 1.040 points: participants told they failed rated algorithm fairness lower on the 7-point scale than those told they passed (p<0.001).The 95% CI was [0.819, 1.260], and the unfavorable-outcome effect was significant for participants expecting either failure or passage.
  • The unfavorable-outcome effect was significant regardless of whether participants expected themselves to pass or fail.
  • 0.410 points: participants rated a biased algorithm lower on the 7-point fairness scale than an unbiased algorithm (p<0.01).The 95% CI was [0.175, 0.644].
  • 1.034 versus 0.396: in the combined model, the unfavorable-outcome effect was twice the size of the biased-versus-unbiased algorithm effect.The coefficients were significant at p<0.001; their 95% CIs were [-1.253, -0.816] and [-0.615, -0.177], respectively.
  • Model 1 explained more perceived-fairness variance than Model 2, with adjusted R2 values of 0.161 and 0.055, respectively.

Algorithm Creation and Deployment

Development procedures had no significant direct effect on perceived fairness, but outsourcing and transparency changed how strongly biased outcomes affected ratings.

  • The development-procedure manipulations had no significant main effects on perceived fairness.The five procedure variables collectively explained less than 1% of the variance.
  • Outsourcing strengthened the negative effect of algorithm bias compared with internal development by computer science experts or mixed internal staff.Interaction coefficients were 0.971 and 1.070 for the two internal-development comparisons, both with p<0.001.
  • Higher transparency also strengthened the negative effect of a biased algorithm on perceived fairness.The interaction coefficient was -0.486 (p<0.05; 95% CI: [-0.925, -0.046]).

Individual difference

Computer literacy was positively associated with perceived fairness, while gender and education moderated responses to unfavorable outcomes. Lower-literacy participants generally rated the algorithm less fairly.

  • Low-computer-literacy participants rated algorithm fairness 0.285 points lower than high-literacy participants (p<0.05).The 95% CI was [-0.515, -0.055].
  • Gender, education level, age, and race showed no main effects on perceived fairness, while computer literacy did.
  • Participants with education beyond a Bachelor’s degree barely changed their fairness ratings in response to the outcome.
  • Female participants reacted more strongly to unfavorable outcomes than male participants (Coef.=0.558, p<0.05).The 95% CI was [0.082, 1.034].
  • Participants with lower education levels reacted more strongly to unfavorable outcomes than participants with higher education levels (Coef.=-1.144, p<0.01).The 95% CI was [-1.916, -0.372].

DISCUSSION

Personal outcomes strongly shape perceived algorithmic fairness: unfavorable decisions reduce ratings more than strongly biased group error rates, while expectations and education moderate these judgments. The findings imply that fairness assessments based on user feedback must account for outcome favorability.

  • Algorithm Outcomes: A “fail” outcome lowered perceived fairness by one point on average on a seven-point scale.
  • Algorithm Outcomes: Biased prediction error rates across demographic groups reduced fairness evaluations by about 0.4 points, less than the unfavorable personal outcome effect.
  • Algorithm Outcomes: The lowest aggregate fairness ratings came from participants with a “fail” expectation who also received an unfavorable outcome.
  • Algorithm Outcomes: Participants expecting “fail” rated the process 0.7 points less fair than participants expecting “pass,” regardless of the eventual outcome.
  • Implications: User-feedback systems evaluating algorithmic fairness should measure or incorporate outcome favorability bias.
  • Implications: Interactive tools showing how algorithmic decisions affect different people may provide perspective on fairness and induce empathy across users.

Development Procedures

Development procedures had no main effect on fairness evaluations, although outsourcing and high transparency intensified reactions to algorithmic bias. Individual differences mattered: lower computer literacy and education were associated with more outcome-sensitive judgments.

  • Development Procedures: The four development-process manipulations produced no main effect on fairness evaluations.The authors suggest the manipulation may have been weak or that nonexperts may not understand how development choices affect fairness.
  • Development Procedures: Outsourcing and high transparency exacerbated the negative effect of algorithmic bias on perceived fairness.
  • Individual Differences: People with low computer literacy gave lower fairness evaluations than people with high computer literacy.
  • Individual Differences: Lower education levels amplified the fairness-rating difference between favorable and unfavorable outcomes.Participants with the highest education level, more than a Bachelor’s degree, barely changed ratings based on outcome.
  • Individual Differences: Education may help people consider algorithmic fairness beyond their own self-interest when decisions affect others.

Limitations & Future Work

The study’s limitations concern participant attention, simplified and fixed bias manipulations, and uncertainty about whether findings generalize beyond the MTurk scenario.

  • Participant attention: Some workers may not have read the scenario carefully or may have forgotten details before evaluating fairness.The study used an attention check and five quiz questions to reinforce understanding.
  • Bias manipulation: The study tested only two fixed levels of demographic-group error-rate disparity, so other bias magnitudes may alter the relative importance of bias and other factors.The authors also describe error rate as a simplistic representation because false-positive and false-negative rates have different implications.
  • Generalizability: The MTurk Master-qualification scenario may not generalize to recidivism prediction, hiring, or school-admission settings.The authors note that scenarios may strongly influence fairness judgments and call for further work on generalizability.

CONCLUSION

The study uses a between-subjects randomized MTurk survey to examine how algorithm outcomes, development procedures, and individual differences shape perceived fairness. Evaluations are especially sensitive to whether participants personally receive a positive outcome, an effect that can exceed the negative effect of strong demographic-group bias.

  • A between-subjects randomized MTurk survey examines algorithm outcomes, development procedures, and individual differences as influences on perceived fairness.
  • Personal positive outcomes strongly affect fairness evaluations, even surpassing the negative effect of describing strong bias against particular demographic groups.
Loading 2001.09604v1…