Source-linked AI summary
Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision Support
Anna Kawakami, Venkatesh Sivaraman, Hao-Fei Cheng, Logan Stapleton, Yanghuidi Cheng, Diana Qing, Adam Perer, Zhiwei Steven Wu, Haiyi Zhu, Kenneth Holstein
TL;DR
Public agencies need evidence about how workers use AI decision support in high-stakes child welfare settings. Through contextual inquiries and interviews with AFST users, this paper finds that reliance is shaped by contextual knowledge, beliefs, organizational pressures, and mismatched objectives, highlighting tensions in the tool’s current design.
Problem
Workers’ practical experiences integrating AI decision support into child welfare decisions remain poorly understood, despite increasing public-sector adoption.
Method
The study uses contextual inquiries and semi-structured interviews with child welfare hotline workers and supervisors using the AFST.
Results
Workers’ reliance on the AFST is guided by contextual knowledge, beliefs about its capabilities, organizational pressures, and misalignments between algorithmic predictions and decision objectives.
Takeaways & Limitations
The findings support reconsidering how ADS are designed, evaluated, and integrated to better complement workers’ abilities in public-sector decision-making.
Takeaways & Limitations
Comparisons of AFST and worker accuracy are constrained when their predictive targets and decision objectives differ.
Abstract
from arXiv · showhide
AI-based decision support tools (ADS) are increasingly used to augment human decision-making in high-stakes, social contexts. As public sector agencies begin to adopt ADS, it is critical that we understand workers' experiences with these systems in practice. In this paper, we present findings from a series of interviews and contextual inquiries at a child welfare agency, to understand how they currently make AI-assisted child maltreatment screening decisions. Overall, we observe how workers' reliance upon the ADS is guided by (1) their knowledge of rich, contextual information beyond what the AI model captures, (2) their beliefs about the ADS's capabilities and limitations relative to their own, (3) organizational pressures and incentives around the use of the ADS, and (4) awareness of misalignments between algorithmic predictions and their own decision-making objectives. Drawing upon these findings, we discuss design implications towards supporting more effective human-AI decision-making.
1 INTRODUCTION
This study examines how child welfare workers use AI-assisted screening in practice, addressing limited evidence about workers’ experiences and effective human–AI partnerships in this contentious setting.
- The study uses interviews and contextual inquiries to investigate workers’ day-to-day practices and challenges with the Allegheny Family Screening Tool.
- Workers’ reliance on the ADS is guided by contextual knowledge, beliefs about its capabilities, organizational pressures, and awareness of mismatched decision objectives.
- Although the AFST had been used for half a decade, many workers viewed its design as a source of tension and a missed opportunity to complement their abilities.
- The findings motivate reconsideration of how public-sector ADS should be designed, evaluated, and integrated.
- The paper provides the first in-depth qualitative investigation of workers’ current AFST practices and challenges, complicating prior narratives about its workplace role.
2 BACKGROUND AND RELATED WORK
Prior work describes growing use of decision support in child welfare but provides limited insight into how workers integrate ADS into practice. This paper extends that literature through an on-the-ground study of the AFST.
- Child welfare decision support ranges from checklist tools to machine-learning systems such as AFST and Rapid Safety Feedback.
- Several ADS deployments have failed amid concerns from affected communities or workers, motivating greater participation during design.
- Research on child-welfare ADS has often examined affected communities or retrospective worker decisions, while workers’ day-to-day integration of these systems remains understudied.
- This study extends prior quantitative evidence that workers can use discretion to compensate for erroneous algorithmic risk assessments by examining those practices in context.
- Child welfare decisions balance child safety against family preservation under high referral volumes and limited ability to systematically use administrative data and case histories.
3 METHODS
The study combines contextual inquiry, interviews, and thematic analysis to examine how call screeners and supervisors use the AFST. Researchers observed workplace decisions, elicited participants’ perspectives, and developed four themes.
- Researchers conducted contextual inquiries and post-interviews with nine call screeners and four supervisors across two visits to Allegheny County CYF.
- Observations covered information gathering, AFST execution, screening recommendations, supervisory review, and final screening decisions.
- All 13 workers participated in post-interviews, with 12 consenting to audio recording and one consenting to note-taking.
- Thematic analysis combined approximately 9.5 hours of transcripts with 92 pages of notes and coding by two to three researchers per participant.
- Affinity diagramming refined 1,529 unique codes into four top-level themes covering workers’ practices and perceived design opportunities.
- Researchers protected participants through voluntary anonymous participation, aggregate reporting, and acknowledgment of positionality and follow-through responsibilities.
4 RESULTS
The results are organized around four themes explaining how workers calibrate reliance on the AFST: contextual information, beliefs about the tool, organizational pressures, and alignment between predictions and decision objectives.
- 4.4 Misaligned objectives: Workers’ reliance is also guided by perceived misalignments between algorithmic predictions and their own decision-making objectives.
- 4.1 Contextual information: Workers calibrate reliance on the AFST using contextual information that the AI model cannot capture.
- 4.2 Beliefs about the AFST: Workers’ beliefs about the AFST are shaped by and shape their day-to-day interactions with the tool.
- 4.3 Organizational pressures and incentives: The findings examine organizational pressures and incentives surrounding workers’ use of the ADS.
4.1 How workers calibrate their reliance on the AFST based on knowledge beyond the model
Workers calibrated reliance on the AFST by integrating qualitative case knowledge and social context that the model captured only coarsely or not at all.
- Workers incorporated rich causal narratives about motives and cultural misunderstandings that were absent from AFST data and manual inputs.These narratives were often communicated to supervisors to support interpretation and decision-making.
- Workers based overrides and screening decisions primarily on qualitative allegations and the social context surrounding each case.They considered relationships among callers, children, families, and alleged perpetrators, along with cultural background and motivations.
- A low AFST score did not prevent screening-in when allegations revealed a potential safety risk the model could not observe.In one case, the model lacked information that a visiting relative might pose a risk to the child.
- Workers viewed allegation categories as too coarse to represent important distinctions, especially when cultural context changed the meaning of parent-child conflict.They regarded imminent risks and caregiver substance abuse as stronger safety indicators, while parent-child conflict required additional context.
- Adding allegation-type variables did not eliminate the underlying communication gap because workers could select only from a fixed set of categories.The available mechanism constrained how workers could convey relevant contextual knowledge to the model.
4.2 How workers’ beliefs about the AFST shape and are shaped by their day-to-day use of the algorithm, in the absence of detailed training
With little formal training or model transparency, workers learned about the AFST through day-to-day use, informal collaboration, and score experiments, forming useful but imperfect beliefs about its behavior.
- Most workers understood little about the AFST’s features, weighting, or data sources because training was procedural and the interface exposed only a score or decision message.Workers reported receiving surface-level instruction on running and viewing the score rather than authoritative explanations of model mechanics.
- Minimal model information intended to prevent gaming led workers to improvise their own ways of learning the AFST’s capabilities and limitations.Workers developed individual and collective knowledge through daily use and discussion.
- Workers informally predicted scores and discussed surprising cases together, using collaborative sensemaking to infer how the AFST behaved.These practices included guessing expected scores from case histories and comparing the model’s output with workers’ own impressions.
- Workers compared scores from slightly modified referral records to infer how particular factors affected the AFST output.For example, they omitted a family member, reran the score, and compared results while preserving the main record’s thoroughness.
- These learned intuitions were often accurate but remained incomplete, with unexpected scores sometimes making the AFST appear hyper-sensitive or unreliable.One supervisor observed a change from 14 to 7 without apparent input changes, contributing to distrust.
- Workers wanted clearer explanations and feedback about how the AFST generated scores, including which data sources and factors it used.They sought information that could help them avoid faulty assumptions and interpret scores more effectively.
- Workers’ beliefs about re-referrals, service involvement, family size, and age shaped how they interpreted AFST scores, sometimes treating accumulated system history as evidence of risk.Their interpretations included beliefs that more re-referrals, service involvement, or people in a report increased the score.
4.3 Beyond “trust” in AI: Influences of organizational pressures and incentives
Workers’ reliance on the AFST was shaped not only by trust, but also by organizational pressures, override procedures, workload concerns, and whether feedback was treated as consequential.
- Workers perceived pressure to avoid overriding AFST scores too often because analytics meetings appeared to monitor override rates.They did not know whether an official acceptable rate existed, but some believed they had crossed an unspoken threshold.
- High scores could create pressure to assign cases despite workers’ judgments that the allegation posed little immediate safety risk.Supervisors sometimes helped screeners reason through whether a mandatory high-risk protocol had to be followed.
- Workers sometimes accepted high-risk protocols to avoid the additional work required to document an override.The required open-text rationale was intended to create friction and encourage reflection, but could instead influence compliance through workload concerns.
- Workers questioned the value of feedback narratives when they received no acknowledgment or useful explanation of how their input was used.Some felt that writing feedback unnecessarily increased their workload without producing a meaningful response.
- Dismissive responses to reported limitations made workers feel their expertise was undermined and discouraged further feedback.Workers described the process as explaining why the score was right rather than considering what the algorithm might have missed.
4.4 Navigating value misalignments between algorithmic predictions and human decisions
Workers navigated a mismatch between the AFST’s longer-term risk predictions and their own focus on immediate child safety, using the score selectively when it complemented rather than replaced judgment.
- Workers’ decisions targeted immediate child safety, whereas the AFST predicted particular adverse outcomes over a two-year period.Several workers therefore viewed the current AFST as only partly relevant to their day-to-day decisions.
- Some workers rejected longer-term risk prediction as fundamentally misaligned with their roles, while others saw it as potentially complementary to immediate safety assessment.The disagreement reflected different views of the decision problem, not simply different levels of trust in the technology.
- Most workers assigned the AFST a minor, non-driving role, describing it as a nudge rather than a determinant of recommendations.Some workers said they would not change their recommendation because of the score, while others saw it as modest directional input.
- Workers believed the AFST could help counteract personal cultural or moral biases in cases where their own judgment might underplay risk.This perceived benefit coexisted with awareness of the algorithm’s biases and value misalignments.
- Some workers still valued the AFST when uncertain or pressed for time, using a high or low score to make an initially ambiguous decision more confidently.All supervisors and some screeners reported using the score in at least these circumstances.
- Middle-range scores between 10 and 14 were especially difficult to interpret because workers lacked enough transparency to see how the score should complement their judgment.Workers described these yellow reports as a recurring source of frustration and uncertainty.
5 DISCUSSION
The discussion shows that workers calibrate AFST reliance through contextual expertise, imperfect model understanding, organizational pressures, and differing decision objectives. These findings complicate existing accounts of algorithmic reliance and motivate broader reconsideration of ADS design and evaluation.
- Contextual expertise: Workers compensate for AFST’s gaps by using qualitative, case-specific knowledge absent from the administrative data underlying its predictions.They form causal narratives about referred cases to calibrate reliance on the score.
- Understanding the ADS: Minimal training led workers to develop sophisticated but imperfect intuitions about AFST behavior through everyday algorithm auditing.Workers learned from day-to-day interactions and from comparing predictions with contextual case details.
- Organizational pressures: Workers’ decisions to follow or contradict AFST recommendations were shaped by organizational pressures and downstream caseworker workload, not trust alone.Some complied to avoid administrative criticism, while others considered caseworker capacity when screening ambiguous cases.
- Decision objectives: Workers recognized that AFST’s longer-term predictive targets differed from their focus on immediate safety and short-term risk.The two-year prediction horizon did not clearly show workers how to complement their own judgment in practice.
- Evaluation implications: Because workers and AFST may pursue different objectives, comparisons of their predictive accuracy can evaluate workers on a task they are not actually performing.The discussion therefore calls for evaluation methods that account for differences between human decision objectives and AI predictive targets.
6 DESIGN IMPLICATIONS
The proposed design implications focus on making ADS more responsive to worker expertise, more understandable through training and interfaces, and more aligned with meaningful measures of decision quality.
- Worker expertise: ADS should let workers provide contextual feedback and learn from patterns in workers’ override decisions over time.These mechanisms would incorporate worker expertise into algorithmic recommendations and system improvement.
- Training and interfaces: Training tools and interfaces should help workers understand the boundaries of an ADS’s capabilities and limitations.
- Decision quality: Agencies and researchers should co-design decision-quality measures with workers so those measures are meaningful and expose counterproductive shortcomings.Worker involvement may improve motivation to use the measures and help identify problematic evaluation criteria.
7 CONCLUSION
The conclusion calls for rethinking the interfaces, models, and organizational processes shaping ADS use in child welfare. It emphasizes involving the broader stakeholder ecosystem throughout the ADS lifecycle to reconcile values and serve workers and families.
- Future research: Future research should redesign the interfaces, models, and organizational processes that shape ADS use in child welfare and other public-sector contexts.
- Stakeholder ecosystem: Meaningful stakeholder involvement across design, development, deployment, use, and maintenance can help resolve value misalignments and better serve families and child welfare workers.Relevant stakeholders include administrators, leadership, caseworkers, families, affected communities, and ADS developers.