Source-linked AI summary
A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores
Maria De-Arteaga, Riccardo Fogliato, Alexandra Chouldechova
TL;DR
This paper asks whether experts can recognize and override erroneous algorithmic risk recommendations in a sensitive child welfare decision process. Using retrospective data from a deployed screening tool and a technical glitch that produced incorrect scores, it finds that workers changed their behavior after deployment and were less likely to follow incorrect recommendations. The findings support decision pipelines that preserve human autonomy rather than fully automating these decisions.
Problem
The paper examines whether humans can identify cases where algorithmic recommendations are wrong and appropriately override them in sensitive decision-making domains.
Method
The authors retrospectively analyze child welfare call workers’ screening decisions before and after deployment of a risk assessment tool, using a technical glitch that produced misestimated scores.
Results
Call workers changed their behavior after deployment and were less likely to adhere to recommendations when a technical glitch caused an incorrect score to be displayed.
Takeaways & Limitations
Decision pipelines should preserve human autonomy so workers can identify and correct erroneous algorithmic recommendations.
Takeaways & Limitations
The study is retrospective, although its real-world deployment setting provides field validity and access to a phenomenon unsuitable for a randomized field trial.
Abstract
from arXiv · showhide
The increased use of algorithmic predictions in sensitive domains has been accompanied by both enthusiasm and concern. To understand the opportunities and risks of these technologies, it is key to study how experts alter their decisions when using such tools. In this paper, we study the adoption of an algorithmic tool used to assist child maltreatment hotline screening decisions. We focus on the question: Are humans capable of identifying cases in which the machine is wrong, and of overriding those recommendations? We first show that humans do alter their behavior when the tool is deployed. Then, we show that humans are less likely to adhere to the machine's recommendation when the score displayed is an incorrect estimate of risk, even when overriding the recommendation requires supervisory approval. These results highlight the risks of full automation and the importance of designing decision pipelines that provide humans with autonomy.
INTRODUCTION
Algorithmic risk tools are entering sensitive expert decision pipelines, but their use raises questions about human reliance and the ability to correct erroneous recommendations. This study examines those questions in child welfare screening, using misestimated scores caused by a technical glitch.
- Risk assessment tools condense case information into scores intended to reflect the likelihood of adverse outcomes across sensitive domains.
- Research suggests machine-assisted decisions may not improve outcomes and can increase racial and socioeconomic disparities.
- The study asks whether experts can identify cases where a machine recommendation is wrong and appropriately override it.
- Allegheny County child welfare call workers decide whether allegations of neglect or maltreatment should be screened in for investigation.
- A technical glitch incorrectly calculated some model inputs, producing misestimated risk scores and enabling analysis of decisions under erroneous algorithmic advice.
- The authors emphasize that transparency about technical issues is uncommon despite such issues not being uncommon.
BACKGROUND AND RELATED WORK
Related work frames human use of algorithmic recommendations as a problem of complementarity, shaped by aversion, over-reliance, task type, and decision context. This paper extends that literature by studying correction of misestimated scores in real-world risk assessment.
- Prior research finds that human decisions aided by decision support systems are often no better than decisions made by humans alone.
- Algorithm aversion describes under-compliance after users observe erroneous recommendations, whereas automation bias describes following recommendations despite contradictory information.
- Automation bias includes omission errors from missed flags and commission errors from acting on erroneous recommendations.
- Complex tasks, time pressure, experience, confidence, and social accountability have been associated with users’ reliance on decision support.
- Algorithm aversion and automation bias are opposing tendencies whose prevalence varies with task type and level of automation.
- Diagnostic tasks have knowable ground truth, whereas prognostic tasks concern uncertain future outcomes and therefore inevitably involve mistakes.
- In child welfare, adherence to a score is not equivalent to trust because workers must consider other relevant factors and information.
- This study uses retrospective observations before and after deployment to examine whether humans correct misestimated risk scores in a real-world setting.
DEPLOYMENT SETUP AND DATA
Child welfare hotline workers decide whether alleged maltreatment or neglect warrants investigation using referral-call information and extensive administrative data about associated children and adults.
- Call workers assess whether referrals alleging potential child maltreatment or neglect should be screened in for investigation.
- Workers use referral-call information alongside hundreds of administrative data elements covering demographics, welfare involvement, criminal history, and related information.
Predictive model
The deployed tool predicts two child welfare outcomes and combines their results across all children in a referral into a single displayed risk score.
- Two predictive models estimate each child’s probabilities of out-of-home placement and future referral from demographic, welfare, justice, and behavioral-health features.
- The predicted probabilities are converted into integer scores from 1 to 20 corresponding to ventiles.
- The displayed score S ∈{1,...,20} is the maximum score across both models and all children in the referral.
Deployment
The deployed pipeline calculates two risk scores for each child associated with a referral and shows the maximum to the call worker. Screening out high-placement-risk calls requires supervisory approval.
- Two scores estimate re-referral risk and out-of-home placement risk for every child associated with a referral.
- Screening out a referral with an out-of-home placement score of at least 18 requires supervisor approval.
- The maximum score across associated children is shown to the call worker.Workers also have access to historical records and information conveyed during the call.
- Call workers combine the displayed score with other available information when deciding whether to screen in a call for investigation.
Score misestimation
A deployment glitch caused some realtime model inputs to be calculated incorrectly, making displayed scores differ from the scores that should have been shown. The paper uses this discrepancy to study responses to inaccurate recommendations.
- A realtime database issue incorrectly returned zero counts and indicators for some model inputs.
- Other mismatches arose because adults’ roles and associated information changed as cases evolved.
- Figure 2 maps the fraction of referrals with each shown score ˜S conditional on their assessed score S.
- Most shown scores were equal or close to the assessed scores, but some displayed scores were inaccurate because of the glitch.
Data
The study compares worker behavior before and after deployment using hotline referral data, while restricting analysis to referrals with displayed scores and excluding cases governed by mandatory investigation rules.
- The analysis uses January 2015–July 2016 data before deployment and August 2016–December 2017 data after adoption.
- Figure 3 examines how frequently each feature’s realtime value differed from its retrospectively recalculated value.
- The study is restricted to the 92.5% of referrals with an associated score shown.
- The analysis excludes the 19% of referrals subject to regulations requiring investigation, because workers had no screening discretion in those cases.
Change in call workers’ behavior
Workers changed which referrals they screened in after deployment, even though the overall screen-in rate remained around 45%. Post-deployment decisions were more aligned with assessed risk and showed evidence of correcting inaccurate displayed scores.
- Overall screen-in rates stayed around 45% before and after deployment, while the types of cases investigated changed.
- Post-deployment screen-in decisions were better aligned with assessed score S, especially for very high and very low risk cases.
- Screen-in rates increased for the highest-risk cases and decreased for low- and moderate-risk cases after deployment.
- Screen-in rates for assessed mandatory cases rose from 58% before deployment to 71% afterward.
- Post-deployment decisions aligned more closely with assessed score S than with shown score ˜S, indicating correction of some score miscalculations.
Overrides of erroneous scores
Workers changed screening behavior after deployment and did not blindly follow inaccurate scores. They used other information to override recommendations, including when supervisory approval was required.
- Workers’ decisions were better calibrated to the assessed score than to the shown score.
- Among cases with shown scores between 11 and 15, underestimated cases were screened in at almost 60%, versus around 30% for other cases.
- When shown scores were high but assessed scores were lower, workers were less likely to screen in cases in the highest-risk bucket.
- For assessed mandatory cases, screen-in rates remained approximately constant across shown-score buckets, even when shown scores were more than 12 points lower.
- Post-deployment, accepted-for-service rates increased from 18% to 21%, while referrals connected to existing cases increased from 19% to 23%.The share not accepted for service fell from 63% to 56%.
- Cases identified as underestimated were more likely to be accepted for service after investigation, including more than half versus 15% for other cases with shown scores between 11 and 15.
Disparities in decision-making
The analysis examines whether tool deployment changed screening disparities across race and socioeconomic status. It reports no apparent poverty correlation among assessed mandatory cases and no difference in racial adherence that would compound prior injustices.
- Overall screen-in rates slightly decreased for Black children and slightly increased for White children after deployment.
- For mandatory-screen-in cases, screen-in rates increased for both races, with a slightly sharper increase for White children.
- The results did not indicate a racial difference in willingness to adhere to assessed mandatory-screen-in recommendations that would compound prior racial injustices.
- Among assessed mandatory cases, screen-in rates did not appear to correlate with neighbourhood poverty levels.
DISCUSSION
In the deployed child welfare setting, workers changed their decisions and showed partial rather than blind adherence to the tool. Their autonomy and access to additional case information helped them respond to erroneous scores, although the retrospective design limits causal conclusions and further work is needed to improve error detection.
- Workers were less likely to adhere to recommendations when a technical glitch caused an incorrect score to be displayed.
- Call workers changed their screening decisions after deployment but did not follow the tool’s recommendations in every instance.
- The study’s retrospective design provides field validity but prevents the controlled variation available in randomized experiments.The real-world setting offers an opportunity to observe trained experts making high-stakes decisions that may differ from laboratory behavior.
- The retrospective study could not perturb decision-pipeline elements to assess their individual effects on decision outcomes.The authors call for controlled research on how trust and other pipeline factors influence errors and human-machine decision making.
- Access to referral calls and administrative data gave workers case information beyond what entered the risk score calculation.Workers could still view correct child welfare history when corresponding inputs were miscalculated in real time.
- Providing humans autonomy to contradict the machine mitigated the effects of miscalculated scores in child maltreatment call screening.The authors recommend designing systems to strengthen humans’ ability to identify and correct model mistakes.