Source-linked AI summary

The age of secrecy and unfairness in recidivism prediction

Cynthia Rudin, Caroline Wang, Beau Coker

arXiv:1811.00731v2stat.APcs.CY

TL;DR

The paper asks whether transparency, rather than competing definitions of fairness, is the more attainable requirement for secret decision algorithms. By partially reverse engineering COMPAS, it examines the algorithm’s age and race relationships and identifies cases where scores conflict with criminal histories. It concludes that transparency would let defendants and the public scrutinize risk-score methodology and calculations.

  • Problem

    Secret algorithms can be described incompletely or misleadingly, while competing fairness definitions lack a clear, compatible standard.

  • Method

    The authors partially reverse engineered COMPAS using processed public criminal-justice data and examined its age, race, and criminal-history relationships.

  • Results

    COMPAS does not seem to depend linearly on age, and many defendants with long criminal histories receive low risk scores.

  • Takeaways & Limitations

    Transparency provides defendants and the public an opportunity to scrutinize the methodology and calculations behind recidivism risk scores.

  • Takeaways & Limitations

    COMPAS scores are not single risk numbers because current charges are separate, leaving substantial freedom in their interpretation.

Abstract

from arXiv · show

In our current society, secret algorithms make important decisions about individuals. There has been substantial discussion about whether these algorithms are unfair to groups of individuals. While noble, this pursuit is complex and ultimately stagnating because there is no clear definition of fairness and competing definitions are largely incompatible. We argue that the focus on the question of fairness is misplaced, as these algorithms fail to meet a more important and yet readily obtainable goal: transparency. As a result, creators of secret algorithms can provide incomplete or misleading descriptions about how their models work, and various other kinds of errors can easily go unnoticed. By partially reverse engineering the COMPAS algorithm -- a recidivism-risk scoring algorithm used throughout the criminal justice system -- we show that it does not seem to depend linearly on the defendant's age, despite statements to the contrary by the algorithm's creator. Furthermore, by subtracting from COMPAS its (hypothesized) nonlinear age component, we show that COMPAS does not necessarily depend on race, contradicting ProPublica's analysis, which assumed linearity in age. In other words, faulty assumptions about a proprietary algorithm lead to faulty conclusions that go unchecked without careful reverse engineering. Were the algorithm transparent in the first place, this would likely not have occurred. The most important result in this work is that we find that there are many defendants with low risk score but long criminal histories, suggesting that data inconsistencies occur frequently in criminal justice databases. We argue that transparency satisfies a different notion of procedural fairness by providing both the defendants and the public with the opportunity to scrutinize the methodology and calculations behind risk scores for recidivism.

1 Introduction

The paper argues that transparency should take priority in evaluating recidivism algorithms because proprietary, complicated systems are difficult to verify, allowing errors and misleading methodological claims to persist. A partial reconstruction of COMPAS challenges its documented age dependence, questions conclusions about racial dependence, and identifies low-risk scores among people with long criminal histories.

  • Motivation: Competing fairness definitions are incompatible, so the paper redirects attention toward transparency as a more readily obtainable procedural safeguard.Transparent models make fairness easier to debate and give defendants and the public information needed to scrutinize tools used for safety and justice.
  • Motivation: COMPAS is difficult to verify because it is proprietary, uses 137 questionnaire variables, and does not reveal how collected data contribute to assessments.Its calculations cannot be double-checked for individual cases, while manually entered data can contain typographical, integration, missing-data, and other errors.
  • Motivation: Errors in complicated scoring systems can remain hidden, making it unclear how often calculation or data-entry problems have affected criminal-justice decisions.The paper identifies the frequency of calculation errors as a central question and notes that errors in complicated models are harder to find.
  • Contributions: Partial reverse engineering suggests COMPAS depends nonlinearly on age, contradicting its stated methodology and invalidating ProPublica’s linear-age regression conclusion about race.Adding a nonlinear age term could address that specific regression issue, but would not resolve the broader problem of unverifiable model descriptions.
  • Contributions: The analysis finds that COMPAS depends heavily on age but not as strongly on criminal history or race proxies, raising concern that it relies on unwanted variables.The paper also pinpoints individuals whose scores seem unusually low given their criminal histories, while the proprietary system prevents determining whether errors caused those scores.
  • Implications: Interpretable models can predict recidivism as accurately as black-box models, weakening the claimed need for proprietary systems in criminal justice.The cited interpretable models use age and prior-crime counts, and the authors argue that high-performing models can be developed without cost to the criminal justice system.
  • Implications: Transparency provides defendants and the public an opportunity to scrutinize risk-score methodology and calculations, constituting a distinct form of procedural fairness.The authors argue that life-changing decisions should not rely on error-prone systems without entitlement to clear information.

2 Reconstructing COMPAS

The paper partially reconstructs COMPAS from its documentation and available data, then tests whether its documented additive structure matches observed scores. It finds nonlinear age dependence, weak apparent dependence on criminal history and race after age adjustment, and important data limitations.

  • COMPAS as described by its creator: COMPAS combines 137 questionnaire variables into subscales, then linearly combines those subscales with age and age-at-first-arrest to form raw risk scores.The final integer score is normalized from a raw score, with higher scores indicating higher risk.
  • Data and reconstruction: The analysis supplements the ProPublica dataset with Broward probation data, but subjective questionnaire items and other features remain unavailable.The authors therefore cannot compute all COMPAS inputs or fully verify subscale construction.
  • Age dependence: Observed lower bounds indicate that COMPAS depends on age through approximately linear splines rather than the linear age terms stated in its documentation.The authors fit age functions after identifying likely age outliers and use the lower-bound structure to test the documented model.
  • Age dependence: After subtracting the fitted age component, adding age produces nearly identical prediction accuracy, suggesting that age does not participate in the remainder term.This pattern appears for both the general and violence recidivism scores.
  • Criminal history: The general-score remainder shows no clear relationship with criminal history, and criminal-history variables predict it surprisingly poorly.The authors therefore describe COMPAS dependence on criminal-history and related subscale components as weak.
  • Race and unmeasured features: Including race does not improve prediction accuracy, so the authors hypothesize at most weak dependence on race after conditioning on age and criminal history.They caution that analyses using only observed features can misestimate COMPAS dependence when unmeasured survey features are correlated with them.

3 COMPAS sometimes labels individuals with long or serious criminal histories as low-risk

The paper identifies multiple cases in which COMPAS assigns low scores to people with long or serious criminal histories, indicating possible input or scoring problems. Public records and comparisons with a machine-learning model further expose discrepancies that secrecy makes difficult to investigate.

  • COMPAS’s 137 questionnaire variables and common data-entry errors create opportunities for incorrect scores, while proprietary scoring prevents determining how errors affect outcomes.
  • A defendant with trafficking and aggravated-battery convictions received COMPAS violence score 1, the lowest possible risk.
  • Several individuals have low COMPAS scores despite long criminal histories, and their scores appear inconsistent with the available inputs.
  • Long criminal histories with low COMPAS violence scores should be impossible under the analyzed age dependence unless inputs were entered incorrectly or omitted.
  • A defendant charged with kidnapping but lacking prior crimes received the lowest-risk COMPAS score of 1 because current charges are excluded from the general and violence scores.
  • A boosted decision tree often disagreed with COMPAS, including cases where the tree indicated high recidivism risk but COMPAS indicated lower risk.

4 Is age unfair? Fairness through the lens of transparent models

The paper argues that age-based risk prediction can display the same group disparities used to criticize COMPAS, while transparency makes the underlying trade-offs easier to examine. It also contends that ProPublica’s regression analysis relied on an invalid linear-age assumption.

  • ProPublica’s regression analysis was invalid because it assumed a linear dependence on age, while the first analysis did not condition on age and criminal history.
  • Adults’ recidivism risk decreases with age, but African-American people in the dataset were assessed at younger median ages than Caucasians.
  • The age-only model produced approximately 10% higher FPR for African-Americans and approximately 10% higher FNR for Caucasians.
  • A simple model using age and number of priors was as “unfair” as COMPAS under ProPublica’s true/false positive-rate analysis.
  • Removing age and criminal history because they produce disparities could leave less useful predictors and return decisions to non-transparent human judgment.
  • Transparent age models make competing fairness definitions easier to debate and can reveal disagreement arising from different age distributions.

5 Discussion

The discussion argues that COMPAS’s opacity creates procedural risks beyond disputed group fairness, including faulty interpretations, miscalculations, privacy concerns, and unnecessary resource use. Because transparent models can match COMPAS’s predictive accuracy, the authors find no good reason to retain complicated proprietary models for recidivism assessment.

  • Reverse engineering suggests COMPAS’s dependence on age is nonlinear, undermining ProPublica’s linear-age methodological assumptions and resulting conclusions.The authors report that subtracting a hypothesized nonlinear age component changes the interpretation of COMPAS’s possible race dependence.
  • Because COMPAS is a black box, practitioners may struggle to combine its scores with current charges or other outside information, leaving score interpretation highly flexible.The discussion notes that the current charge is separate from the COMPAS score and that the appropriate weighting is unclear.
  • COMPAS can assign low risk scores to individuals with long criminal histories, suggesting frequent data inconsistencies in criminal justice databases.The authors link such cases to possible miscalculation and potentially dangerous decision-making.
  • Transparent models incur no predictive-accuracy loss in criminal justice applications, while proprietary complexity can produce errors and procedural injustice.The authors argue that simpler interpretable models are no less useful for recidivism prediction than COMPAS.
  • COMPAS’s collection of extensive private information may be unnecessary for estimating risk, creating potential privacy unfairness for people compelled to disclose it.The discussion specifically raises concerns about socioeconomic proxies, family criminal history, and poverty information.
  • Opacity can affect many parties by obscuring bias, allowing errors from typos, misallocating public resources, and limiting scrutiny of methodology and calculations.The authors frame transparency as a form of procedural fairness for defendants and the public, while noting that removing COMPAS without a transparent alternative would still leave a black box.

Supplementary materials

The supplementary analyses examine age-related recidivism patterns, demographic age distributions, and whether race changes recidivism-prediction performance.

  • The probability of a new charge within 2 years decreases as age increases.
  • African-Americans in the Broward County dataset tend to be younger when their COMPAS scores are calculated.
  • Predictions of general and violent recidivism are very similar with and without race as a feature.
  • Tables 7 and 8 report misclassification errors for general and violent recidivism, respectively, comparing models with and without race.

Fitting fage and fviol age

The authors fit age-based lower bounds for COMPAS scores using individuals assumed to have the lowest possible scores at each age. Their analyses support nonlinear current-age functions, while age-at-first-arrest does not define a comparable smooth lower bound.

  • The analysis assumes that some individuals have the lowest possible COMPAS score for each age despite unobserved inputs.
  • The COMPAS lower bounds are defined using current age because its relationship with the score is clearest in the data.
  • Age-at-first-arrest does not produce a smooth lower bound, unlike current age, and is therefore not used to define an initial lower bound.
  • Many individuals share identical raw COMPAS scores at particular ages, and most have current age equal to age-at-first-arrest.
  • Alternative explanations involving nonzero unobserved inputs or manual overrides are considered but judged unlikely under the stated assumptions.
  • The authors argue that these individuals lie on the true age functions and have the lowest possible COMPAS scores for their ages.

An alternative explanation of the nonlinearity in age

The authors test whether uneven age sampling could explain the observed nonlinear lower bound. Subsampling age groups does not remove the nonlinearity, mitigating this alternative explanation.

  • The analysis considers whether more extreme values at densely sampled ages induce the observed nonlinear lower bound.
  • Exploratory analysis mitigates the concern that uneven age distributions explain the nonlinearity.
  • Capping each age group at min(150, number of individuals per age) does not alter the lower-bound nonlinearity.
  • Age splines fitted on the sampled data are very close to those fitted on the full data.

Logistic Regression

The authors replicate ProPublica’s logistic-regression analysis of COMPAS risk categories using features including race. They note that the analysis uses recidivism information unavailable when COMPAS is calculated.

  • The analysis models medium-or-high versus low COMPAS risk categories using features including race.
  • The authors exclude observations lacking at least two years of data beyond the screening date because recidivism is used as a covariate.
  • Using 2-year recidivism in ProPublica’s model relies on information unavailable when the COMPAS score is calculated.

Data processing

The analysis reconstructs COMPAS-related features from public criminal-justice data, defining screening dates, offenses, recidivism, and feature inputs through explicit processing rules. It also documents assumptions and data limitations that constrain these reconstructions.

  • Data sources and feature construction: The researchers processed raw COMPAS-related data into person-level features, partly to ensure feature quality and partly to create features matching COMPAS subscale components.Their data included COMPAS scores and public criminal-justice records; ProPublica published analysis code and raw data but not its feature-processing code.
  • Screening dates: Each feature set corresponds to an individual on a particular COMPAS screening date, with earlier events included when computing features for later screening dates.The screening date is when the COMPAS score was calculated.
  • Data assumptions and limitations: Missing charges and probation events, duplicate same-day scores, and inferred offense classifications required explicit assumptions and processing rules.The study also excluded degree “(0)” charges, inferred offense types from statute numbers, and used thresholds of 365 days after probation onset and 30 days before termination.
  • Offense data: The analysis used charge data rather than arrest data because statute numbers needed for the Violence Subscale were absent from arrest records.The authors assumed charge data would provide similar results even though COMPAS subscales appeared to rely on arrest data.
  • Current-offense identification: The current offense was identified from the most recent charge on or before screening, using charges within 30 days when the exact triggering offense was unclear.Observations with missing current offenses were excluded from analyses requiring criminal history.
  • Outcome and demographic definitions: Recidivism was defined as any charge within two years of screening, and only observations with two years of subsequent data were used when recidivism was required.Age was measured in completed years, and juvenile charges were offenses before age 18.

Machine learning implementation

The paper predicts COMPAS raw-score remainders and recidivism using several machine-learning methods after removing hypothesized nonlinear age components. It compares regression and classification settings using specified feature sets and cross-validation procedures.

  • Raw COMPAS score remainders are predicted with linear regression, random forests, Extreme Gradient Boosting, and SVM.
  • The models predict raw rather than decile scores after subtracting age polynomials for general and violence scores.
  • XGBoost and SVM hyperparameters are selected by 5-fold cross-validation, random forests use defaults, and the code is written in R.
  • General-score remainder models use Criminal Involvement features, while violence-score models use History of Violence and History of Noncompliance features.
  • Race and screening-age variables may be included depending on the analysis, and age-at-first-arrest is used for both raw-score types.
  • Recidivism prediction uses the same methods and features with classification adaptations, includes the current offense, and applies logistic regression instead of linear regression.

Subscale tables

The paper maps COMPAS general and violent recidivism scores to specific subscale inputs and documents which components can be computed from the available data. Several subscales are only partially or not computable.

  • General recidivism uses Criminal History, Substance Abuse, and Vocation/Education subscales.
  • Violent recidivism uses History of Violence, History of Noncompliance, and Vocation/Education subscales.
  • The History of Violence table reports computed components in bold, unavailable components, and a family-violent-arrests feature that is always zero.
  • The History of Noncompliance table likewise distinguishes computed components from unavailable ones.
  • Criminal Involvement components are all computed, with a charge interpreted as an arrest for the number-of-arrests feature.
  • No Vocation/Education or Substance Abuse components can be computed because the necessary data are unavailable.
Loading 1811.00731v2…