Source-linked AI summary
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Alexandra Chouldechova
TL;DR
Recidivism prediction instruments are increasingly used despite concerns about discriminatory bias and disparate impact. The paper connects psychometric test fairness to classification error rates and shows that predictive fairness can produce disparate impact when recidivism prevalence differs across groups. Its conclusions depend on the specific use context, especially whether high-risk classifications lead to stricter penalties.
Problem
Growing use of recidivism prediction instruments raises concerns about discriminatory bias and whether fairness standards adequately capture disparate impact.
Method
The paper connects psychometric test fairness to classification error rates and analyzes risk-assessment use cases through prevalence and penalty policies.
Results
When recidivism prevalence differs across groups, a test-fair instrument generally yields higher false positive and lower false negative rates for the higher-prevalence group, producing greater penalties under stricter-penalty policies.
Takeaways & Limitations
Fairness criteria and desired error-rate balance should be selected for the specific application context and at an appropriate level of granularity.
Takeaways & Limitations
The impact analysis specifically addresses use cases where high-risk individuals receive stricter penalties, whereas some systems use high-risk assessments to support risk-reduction benefits.
Abstract
from arXiv · showhide
Recidivism prediction instruments provide decision makers with an assessment of the likelihood that a criminal defendant will reoffend at a future point in time. While such instruments are gaining increasing popularity across the country, their use is attracting tremendous controversy. Much of the controversy concerns potential discriminatory bias in the risk assessments that are produced. This paper discusses a fairness criterion originating in the field of educational and psychological testing that has recently been applied to assess the fairness of recidivism prediction instruments. We demonstrate how adherence to the criterion may lead to considerable disparate impact when recidivism prevalence differs across groups.
1 Introduction
Recidivism prediction instruments are increasingly used in criminal justice, raising concerns about discriminatory bias and disparate effects. This paper connects psychometric fairness standards to recidivism prediction and examines how predictive fairness can coexist with disparate impact.
- Risk assessment instruments are used or considered for pretrial, parole, and sentencing decisions, where high-risk misclassification can adversely affect defendants.
- Psychometric fairness standards have been applied to COMPAS and PCRA, with initial findings indicating gender but not racial predictive bias.
- A ProPublica analysis reported that non-recidivating Black defendants were nearly twice as likely as White defendants to receive high-risk assessments.
- The paper shows that racial differences in false positive and false negative rates can follow from predictive fairness when recidivism prevalence differs across groups.
- Fairness and disparate impact are treated as social and ethical concepts rather than purely statistical properties.
- The analysis links statistically quantifiable RPI features to disparate impact in hypothetical use cases.
2 Assessing fairness
The paper defines test fairness as equal recidivism likelihood at a given score across groups and examines its implications for classification error rates. It argues that differing group prevalence constrains false positive and false negative rates even when COMPAS appears well calibrated.
- Test fairness: Test fairness requires equal conditional recidivism probabilities across groups at every score value.
- Test fairness: COMPAS appears to adhere well to the test-fairness condition based on observed recidivism rates across its score values.
- Coarsened classification: The coarsened score thresholds the continuous risk score to classify defendants as high or low risk, enabling confusion-matrix analysis.
- Implied constraints: Test fairness implies that the coarsened score’s positive predictive value does not depend on group membership.
- Implied constraints: Recidivism prevalence is a group-specific constraint that the instrument does not directly control.
- Implied constraints: When group prevalence differs, a test-fair score cannot generally have equal false positive and false negative rates across groups.
3 Assessing impact
The paper models disparate impact when risk assessments assign stricter penalties to high-risk defendants, showing how group differences in recidivism prevalence and error rates shape expected penalties. It also examines finer-grained error-rate differences and connects disparate impact to distributional overlap measures.
- The analysis focuses on settings where a high-risk assessment leads to a stricter penalty, including bail, parole, or sentencing decisions.
- For the Broward data, Black defendants had a 45% FPR and 28% FNR, compared with 23% FPR and 48% FNR for White defendants.
- The MinMax policy assigns the lower penalty tL to low-risk defendants and the higher penalty tH to high-risk defendants within guideline bounds.
- Disparate impact is measured as the expected penalty difference between defendants from groups b and w with specified observed outcomes.
- When recidivism prevalence differs across groups, a test-fair risk instrument generally produces higher FPR and lower FNR for the higher-prevalence group, resulting in greater penalties among both recidivators and non-recidivators.
- 3.1 Connections to measures of effect size: False-positive-rate differences persisted across prior-record subgroups among misdemeanor defendants, rather than appearing only among defendants with more serious charges or extensive records.
- 3.1 Connections to measures of effect size: The paper relates disparate impact to distributional overlap and notes that COMPAS decile scores are not normally distributed, motivating total-variation-based analysis.
4 Discussion
The paper shows that disparate impact can arise even when a binary recidivism instrument is free from predictive bias. It argues that error-rate balance may be desirable in some contexts, but must be evaluated at an appropriate level of granularity.
- The paper’s primary contribution is showing that a predictive-bias-free recidivism instrument can nevertheless produce disparate impact.
- The analysis focuses on a binary risk assessment informing a binary penalty policy, while the formulas have natural analogs for nonbinary scores and penalties.
- Balancing error rates across groups may be desirable in some use cases, even though it generally produces instruments that are not free from predictive bias.
- Equal overall false-positive rates do not guarantee equal error rates within prior-record score categories, so the desired granularity of balance must be chosen explicitly.
- The authors argue against abandoning data-driven assessment because of negative headlines and instead call for context-specific evaluation of quantifiable biases that could lead to disparate impact.