Source-linked AI summary

Interventions over Predictions: Reframing the Ethical Debate for Actuarial Risk Assessment

Chelsea Barabas, Karthik Dinakar, Joichi Ito, Madars Virza, Jonathan Zittrain

arXiv:1712.08238v2cs.LGcs.CYstat.AP

TL;DR

Actuarial risk assessments are debated as predictive tools whose fairness and accuracy may be compromised by historical and structural bias. The paper reframes their purpose toward diagnosis, using machine learning to surface covariates and causal inference to evaluate interventions addressing social, economic, psychological, and systemic drivers of crime.

  • Problem

    Current fairness debates emphasize predictive accuracy and parity, while risk assessments may reproduce narrow subgroup assumptions and overlook structural drivers of crime.

  • Method

    The paper distinguishes predictive regression and machine learning from causal inference, proposing causal models and complementary qualitative analysis for intervention-oriented assessment.

  • Results

    The paper concludes that risk assessments should be diagnostic tools for understanding crime drivers and evaluating interventions, supported by causal inference and empirical studies of downstream consequences.

  • Takeaways & Limitations

    Risk assessment should move from predicting risk scores toward empirically grounded risk mitigation that addresses criminogenic effects at individual and systemic levels.

  • Takeaways & Limitations

    Regression models may fail to contextualize subgroup-specific risk and intervenable factors, especially because validation often centers on white male populations.

Abstract

from arXiv · show

Actuarial risk assessments might be unduly perceived as a neutral way to counteract implicit bias and increase the fairness of decisions made at almost every juncture of the criminal justice system, from pretrial release to sentencing, parole and probation. In recent times these assessments have come under increased scrutiny, as critics claim that the statistical techniques underlying them might reproduce existing patterns of discrimination and historical biases that are reflected in the data. Much of this debate is centered around competing notions of fairness and predictive accuracy, resting on the contested use of variables that act as "proxies" for characteristics legally protected against discrimination, such as race and gender. We argue that a core ethical debate surrounding the use of regression in risk assessments is not simply one of bias or accuracy. Rather, it's one of purpose. If machine learning is operationalized merely in the service of predicting individual future crime, then it becomes difficult to break cycles of criminalization that are driven by the iatrogenic effects of the criminal justice system itself. We posit that machine learning should not be used for prediction, but rather to surface covariates that are fed into a causal model for understanding the social, structural and psychological drivers of crime. We propose an alternative application of machine learning and causal inference away from predicting risk scores to risk mitigation.

1 INTRODUCTION

The 2016 ProPublica article alleged that COMPAS, a proprietary risk assessment tool used in the U.S. criminal justice system, was racially biased and sparked national debate.

  • In 2016, ProPublica alleged that COMPAS was racially biased.The article concerned a proprietary tool used in the U.S. criminal justice system and triggered a national debate.

2 CURRENT DEBATES ON THE FAIRNESS OF RISK ASSESSMENTS

Fairness debates about risk assessments focus on transparency, human discretion, predictive bias, and trade-offs among competing fairness criteria. The paper argues that this predictive framing overlooks assessments’ diagnostic role in guiding interventions that can reduce dynamic risk.

  • Defendants typically receive risk scores without access to the calculations or input data used to produce them.Critics argue that proprietary interests should not override defendants’ due-process rights.
  • Human discretion enters risk assessment through score misapplication, misinterpretation, and deliberate manipulation.Researchers have documented administrators shading scores and using needs scores for decisions they were not designed to inform.
  • Critics debate predictive parity and whether risk models distribute inaccuracies unevenly across groups, while proponents emphasize calibration and equal accuracy.These positions reflect competing technical definitions of bias rather than a settled fairness standard.
  • Fairness criteria can conflict with one another and with optimal predictive accuracy, especially when group base rates differ.The literature identifies multiple fairness concepts whose simultaneous achievement is generally impossible except under restrictive conditions.
  • Risk assessments should be treated as diagnostic tools because risk is dynamic and can be lowered through semi-personalized interventions.The paper argues that their value extends beyond predicting future crime to informing punishment and treatment decisions.

3 RISK ASSESSMENTS: PREDICTIVE OR DIAGNOSTIC TOOLS?

Actuarial risk assessment evolved from subjective clinical tools toward regression-based prediction and then toward models incorporating intervenable needs. The paper argues that current prediction-oriented methods should instead support diagnosis and intervention across the criminal justice process.

  • First-generation assessments used semi-structured clinical evaluations to identify rehabilitative treatment options.They standardized clinical items but were later criticized for subjectivity and low predictive accuracy.
  • Second-generation assessments used regression to optimize predictive accuracy from largely static historical factors.Regression identifies predictive variables without requiring an explanation of why those variables matter.
  • Prediction-oriented policies relying heavily on criminal history were linked by scholarship to mass incarceration and growing racial disparities.Critics also describe a ratchet effect in which incarceration’s social consequences can increase future crime risk.
  • Third-generation tools added criminogenic needs such as employment and substance-abuse history, reframing risk as dynamic and intervenable.These factors could inform treatment interventions beyond incapacitation, even when they were less predictive than static attributes.
  • Pretrial decisions should be treated as interventions because bail can worsen long-term recidivism rather than simply serving as a forecasting problem.The paper cites evidence that high bail amounts, often resulting in detention, drive higher long-term recidivism rates.
  • Regression and machine learning are ill-suited to diagnosis when intervention requires distinguishing correlational variables from causal drivers.The paper therefore calls for statistical techniques that identify causal relationships between risk factors and future crime.

4 REGRESSION VS. MACHINE LEARNING VS. CAUSAL INFERENCE

The paper distinguishes predictive regression and machine learning from causal inference, arguing that current models inadequately represent subgroup-specific and structural drivers of crime. It proposes causal methods for selecting and designing interventions rather than merely improving prediction.

  • Regression-based assessments can obscure risk and intervenable factors that differ across subgroups.The paper notes that models are often validated on populations skewed toward white men and may omit gendered or racialized pathways to crime.
  • RNR assessments focus on individualized treatment variables and give limited attention to broader social or structural drivers of crime.This narrow scope can conflate causes of individual differences in crime with causes of crime itself.
  • Regression selects covariates for predictive significance, which can exclude needs such as mental health when statistical significance is not established.Its limited ability to test competing theories also leaves no principled basis for choosing among divergent subgroup models.
  • Machine learning expands regression-like prediction across more features but remains focused on accurate prediction rather than causal explanation.Interpretability methods can explain model logic without identifying which features causally drive criminal behavior or how much an intervention should change them.

5 TOWARDS A CAUSAL FRAMEWORK FOR RISK ASSESSMENT

The paper proposes using causal inference to evaluate interventions that reduce specified criminal-justice risks, while using machine learning to surface predictive covariates and causal questions rather than produce risk predictions alone. This framework addresses the limits of historical-data models and supports analysis of intervention timing, duration, and effects.

  • Regression and machine learning rely on historical data, which may omit variables needed to design effective interventions.
  • Causal inference is better suited than regression and predictive supervised machine learning to answer which interventions reduce risks such as recidivism and failure to appear.
  • Randomized control trials are the gold standard for measuring causal effects, and a Philadelphia trial found low-intensity community supervision more effective than high-intensity law-enforcement supervision at reducing new criminal activity.
  • RCTs may be impractical in criminal justice because researchers face data constraints, legal resistance, resource limitations, and difficulty obtaining treatment and control outcomes for the same unit.
  • Random assignment balances covariates and helps isolate whether an intervention causes an outcome, while causal analysis can separate unaffected covariates from treatment-induced intermediate outcomes.
  • Machine learning can identify features predictive of recidivism that generate hypotheses about interventions and their timing for later causal testing.

6 IMPLICATIONS AND CONCLUSION

The paper argues that risk assessments should move from predictive technologies toward diagnostic tools grounded in causal, qualitative, and quantitative analysis. It concludes that observational causal inference is an important alternative when randomized experiments are infeasible or ethically risky, although implementation remains sparse and institutionally difficult.

  • Predictive accuracy can leave systems without grounded guidance for intervention and may reinforce mass incarceration and inequality when treated as the primary objective.
  • Risk assessments should shift from prediction toward diagnostic methods that examine criminal justice system effects and evaluate interventions intended to interrupt cycles of crime.
  • The proposed diagnostic framework combines causal inference with qualitative and quantitative analysis to understand social, economic, and psychological drivers of crime.
  • Only 16% of surveyed criminal-justice interventions were evaluated using randomized experiments, reflecting the sparse practical use of the paper’s preferred causal method.
  • Observational causal inference should be considered when randomized experiments are infeasible or ethically fraught, and studies have examined consequences such as pretrial detention’s effects on plea bargaining and labor-market participation.
  • Structural barriers to causal research include limited mentorship, reduced funding for randomized experiments, and the convenience of estimating regression and supervised-prediction models.
  • Data-driven tools can support fair punishment and future crime prevention when they are used to understand and address individual and systemic drivers of crime.
Loading 1712.08238v2…