Source-linked AI summary
Ignore, Trust, or Negotiate: Understanding Clinician Acceptance of AI-Based Treatment Recommendations in Health Care
Venkatesh Sivaraman, Leigh A. Bukowski, Joel Levin, Jeremy M. Kahn, Adam Perer
TL;DR
Clinician acceptance is a barrier to deploying AI treatment recommendations when treatment quality is uncertain and decisions unfold over time. This paper presents an interpretable sepsis-treatment CDS interface and studies 24 intensive care clinicians making AI-assisted decisions on real patient cases. Explanations increased confidence, but clinicians’ responses extended beyond binary acceptance or rejection, often involving negotiation over which recommendation aspects to follow, reject, or delay.
Problem
Treatment recommendations lack a clear objective ground truth, and relatively few studies examine how clinicians accept or use AI-generated treatment advice.
Method
The study developed an interpretable sepsis-treatment CDS interface and analyzed AI-assisted decisions by 24 intensive care clinicians using mixed methods.
Results
Explanations increased clinicians’ perceived AI usefulness and confidence, while interaction patterns extended beyond binary concordance to include four forms of reliance, especially negotiation.
Takeaways & Limitations
Treatment-focused AI adoption should account for clinicians’ partial and selective reliance on recommendations rather than treating use as simple acceptance or rejection.
Takeaways & Limitations
The patient cases were limited to structured MIMIC-IV data, and the SHAP explanations covered only the model’s state-clustering component.
Abstract
from arXiv · showhide
Artificial intelligence (AI) in healthcare has the potential to improve patient outcomes, but clinician acceptance remains a critical barrier. We developed a novel decision support interface that provides interpretable treatment recommendations for sepsis, a life-threatening condition in which decisional uncertainty is common, treatment practices vary widely, and poor outcomes can occur even with optimal decisions. This system formed the basis of a mixed-methods study in which 24 intensive care clinicians made AI-assisted decisions on real patient cases. We found that explanations generally increased confidence in the AI, but concordance with specific recommendations varied beyond the binary acceptance or rejection described in prior work. Although clinicians sometimes ignored or trusted the AI, they also often prioritized aspects of the recommendations to follow, reject, or delay in a process we term "negotiation." These results reveal novel barriers to adoption of treatment-focused AI tools and suggest ways to better support differing clinician perspectives.
1 INTRODUCTION
Health care AI must earn clinician acceptance and support transparent, calibrated collaboration, especially when treatment decisions lack a single correct answer. In a mixed-methods study of AI-assisted sepsis decisions, explanations increased confidence, while clinicians’ interactions extended beyond simple acceptance or rejection.
- Clinician-AI collaboration: Clinicians may distrust AI recommendations despite their quality, while novices may over-rely on incorrect advice, making trust calibration a central adoption challenge.Explanations are one proposed strategy, but they can increase trust even when that trust is unwarranted.
- Treatment-focused AI: Treatment-focused decision support differs from diagnostic support because the best treatment may be uncertain, disputed, and distributed across sequential decisions.Treatment outcomes can be difficult to evaluate because clinical evidence is limited and expert disagreement is common.
- Study focus: The study used an interactive, interpretable CDS interface to examine how ICU clinicians interacted with AI recommendations for sepsis treatment.Sepsis was selected because it is life-threatening, has limited evidence-based protocols, and shows substantial variation in treatment practices.
- Main findings: Explanations improved clinicians’ perceptions of AI usefulness and confidence in their own decisions, but did not appear to change overall binary concordance with recommendations.The mixed-methods analysis combined think-aloud transcripts with structured decision responses.
- Behavior patterns: Clinicians displayed four interaction patterns: Ignore, Negotiate, Consider, and Rely.Negotiation involved weighing and prioritizing individual recommendation aspects, while Consider involved dichotomously deferring to or overriding the recommendation.
- Implications: Partial reliance, particularly negotiation, may affect the efficacy of selected treatments in undetermined ways and exposes obstacles for improving treatment-focused AI tools.The findings also identified ways the model’s formulation hindered effective clinician use.
2 BACKGROUND AND RELATED WORK
Sepsis treatment AI addresses a consequential clinical problem, but its deployment is constrained by uncertain decision quality, explainability risks, and clinician acceptance. The background motivates evaluating human-AI interaction directly rather than assuming that accurate predictions or retrospective policy value will translate into bedside use.
- Clinical context: Sepsis affects over 1.7 million U.S. adults annually and is the leading cause of hospital death, making timely and appropriate management consequential.Treatment includes infection control, IV fluids, and vasopressors.
- Sepsis AI: Existing sepsis AI has focused mainly on early identification, but implemented early-warning systems have not generally changed treatment decisions or patient outcomes.Such systems may fall short when they provide information that is not novel or actionable.
- Treatment recommendations: Sepsis treatment models use historical patient trajectories to standardize care, yet their potential benefit depends on clinicians acting on recommendations at the bedside.The AI Clinician was associated with a potential mortality reduction from around 13% to around 5% according to cited prior work, but prior evaluations were retrospective.
- Explainability: Interpretability methods can increase trust in AI even when explanations are misleading or the underlying prediction is incorrect.Explanations may also interact with confirmation and availability biases.
- Evaluation: Decision quality is difficult to evaluate because real-world correctness can be unknowable or contentious without an objective ground truth.Non-prescriptive systems avoid direct recommendations but may also be easier for clinicians to ignore.
- Study rationale: The present study addresses these challenges by combining behavioral and attitudinal measures to assess interaction with an imperfect AI system when correct decisions are unclear.This design evaluates users within a high-stakes setting rather than along a single quality axis.
- Research gap: Prior research has examined relatively few AI-generated treatment recommendations, leaving their acceptability and influence on clinician decisions less established than for diagnostic tools.The cited studies include antidepressant selection, device implantation, and intravenous fluid administration.
3 DESIGN OF AN INTERACTIVE AI-DRIVEN CDS SYSTEM
The AI Clinician Explorer combines a reinforcement-learning treatment model with interactive visualizations of patient trajectories, recommendations, historical actions, and explanations. The system was built as a research and educational interface and as a foundation for clinician-facing decision support.
- 3.1 Reinforcement Learning for Sepsis Treatment: The AI Clinician uses historical patient trajectories discretized into 750 clustered states and 25 treatment actions at 4-hour intervals.Its Q-values estimate future rewards for actions, and the policy selects the action with the largest value estimate in each state.
- 3.1 Reinforcement Learning for Sepsis Treatment: The study used Komorowski et al.’s AI Clinician because it was a well-known treatment model despite criticism of off-policy evaluation and uncertain performance benchmarks.The authors note that newer deep-learning methods may improve accuracy, but their reliance on off-policy evaluation leaves reliable comparisons elusive.
- 3.1 Reinforcement Learning for Sepsis Treatment: The replicated model was trained on 18,143 MIMIC-IV sepsis patients and produced a policy value of 83.8.The value was estimated using weighted importance sampling, whose possible range is −100 to 100; the original MIMIC-III values ranged from 80 to 90.
- 3.2 AI Clinician Explorer: The AI Clinician Explorer lets users search MIMIC-IV cases, inspect disease trajectories, and compare model predictions with bedside treatment decisions.Its components include patient filtering, trajectory charts, recommendation and historical-action heatmaps, state interpretation, and treatment-value comparisons.
- 3.2 AI Clinician Explorer: The interface explains recommendations through state-feature interpretations and alternative-treatment comparisons while exposing uncertainty in the model’s treatment values.The design was motivated partly by the observation that the AI Clinician often assigns similar values to multiple actions rather than strongly preferring one.
4 STUDY METHODS
The study examined how 24 ICU clinicians used sepsis treatment recommendations under progressively richer visualization conditions. Participants made decisions on four MIMIC-IV cases while thinking aloud, and responses were analyzed using quantitative ratings, concordance standards, and qualitative transcripts.
- Study Design: The mixed-methods study investigated clinicians’ perceptions of AI-supported decision-making, explanation effects on acceptance, and challenges incorporating treatment recommendations.The study recruited 24 ICU clinicians, including attending physicians, advanced practice providers, and critical care fellows.
- Study Procedure: Participants assessed four MIMIC-IV patient cases in a simplified AI Clinician Explorer interface while thinking aloud and reviewing longitudinal clinical data.Patients were presented in randomized order, while visualization conditions followed a fixed progression to reduce cognitive burden and familiarize participants with the interface.
- Visualization Conditions: The visualization conditions progressed from No AI and Text Only to Feature Explanation and Alternative Treatments.Feature Explanation used SHAP attributions, whereas Alternative Treatments displayed five ranked actions with AI quality scores and historical clinician frequencies.
- Case Selection: The authors deliberately selected cases where AI recommendations differed substantially from historical clinician actions rather than selecting cases by a target accuracy level.This design reflected the difficulty of determining treatment-recommendation accuracy without a ground-truth correct decision.
- Analysis: Usefulness, confidence, and related perceptions were analyzed with Likert-scale outcomes, OLS regression, and respondent-clustered standard errors.The study also used think-aloud transcripts and a closing semi-structured interview to characterize how clinicians interpreted and used the visualizations.
- Analysis: Treatment choices were compared with the AI recommendation, the actual MIMIC-IV clinician action, and the majority attending-physician action in the No AI condition.These three standards represented model concordance, historical bedside practice, and an approximation of clinical consensus.
5 RESULTS
The results section combines quantitative comparisons of attitudes across visualization conditions with qualitative analysis of clinicians’ decision-making processes and reflections on the AI system.
- Results: The study first reports participants’ attitudes toward each visualization condition, then analyzes think-aloud decision patterns and clinicians’ interrogation of the AI’s assumptions.The final analysis also considers how clinicians believed the system could better assist them.
5.1 Perceptions of Decision-Making with AI and Explanations
Explanatory visualizations improved clinicians’ perceptions of the AI’s usefulness and its effect on their confidence, but they also made cases feel more difficult. Explanations did not significantly change confidence in the chosen treatment or binary concordance with recommendations.
- Usefulness of the AI: AI usefulness differed by visualization condition, with Feature Explanation rated higher than Text Only by Δ = 0.83, 95% CI [0.24, 1.43], p = 0.018.The overall condition effect was F(2, 69) = 4.251, p = 0.03; Feature Explanation was directionally higher than Alternative Treatments by Δ = 0.75.
- Effect of AI on Confidence: The AI’s effect on clinician confidence differed by condition, with Feature Explanation exceeding Text Only by Δ = 1.08, 95% CI [0.51, 1.66], p < 0.001.The overall effect was F(2, 69) = 7.946, p = 0.002; Feature Explanation was directionally higher than Alternative Treatments by Δ = 0.67.
- Qualitative Responses: Clinicians valued explanatory evidence, particularly Alternative Treatments’ comparisons of outcomes associated with multiple possible decisions.Participants described these comparisons as convincing for changing clinical decision-making.
- Confidence in Treatment Choice: Confidence in the chosen treatment did not differ significantly across visualization conditions, although ratings directionally increased when explanations were provided.The condition effect was F(3, 92) = 2.220, p = 0.11, and adjusted pairwise comparisons were not statistically meaningful.
- Perception of Case Difficulty: Cases were perceived as more difficult with AI explanations: No AI was rated less challenging than Alternative Treatments by Δ = 1.08, 95% CI [0.46, 1.71], p = 0.003.The authors interpret this pattern as evidence that explanations prompted clinicians to consider more factors, especially when the explanation did not align with their expectations.
5.2 Patterns of Interaction with the AI
Clinicians interacted with treatment recommendations in more nuanced ways than simply accepting or rejecting them. Many selectively adopted treatment components, dosage levels, or timing while balancing AI input against clinical judgment.
- Concordance: 42% of decisions matched the AI’s full treatment recommendation across visualization conditions, compared with a 33% concordance base rate without the recommendation.Including agreement with either fluids or vasopressors produced similarly stable AI concordance across visualization conditions.
- Concordance: Explanations changed perceived usefulness and confidence but did not meaningfully change clinicians’ actual decisions across visualization conditions.Participants nevertheless reported that explanatory visualizations improved the AI’s usefulness and increased confidence in their own decisions.
- Ignore: Seven participants made 21 AI-assisted decisions without meaningful AI influence, relying primarily on their initial clinical assessments.These clinicians often rejected recommendations despite explanatory visualizations because they were already confident in their decisions.
- Negotiate: Negotiation balanced prioritized recommendation components with clinicians’ intuition, and the Negotiate group rated the AI’s usefulness 4.6 versus 2.8 on a 7-point scale.Clinicians used explanatory visualizations and aggregate clinician behavior as evidence, but could still reject recommendations they could not justify.
- Other reliance patterns: Some clinicians fully relied on the AI in uncertain cases, while others considered its recommendations in every decision, particularly when they viewed the underlying data as objective.These patterns show that reliance varied across participants and decisions rather than following a single acceptance model.
5.3 Perspectives on AI for Treatment Decision-Making
Clinicians viewed AI treatment recommendations alongside bedside information, personal practice, guidelines, and prior evidence about model quality. Their perspectives were constrained by missing clinical context and by dosage and timing representations that did not always fit ICU practice.
- Bedside information: Eleven participants identified additional bedside data collection as a helpful next step when the AI conflicted with their judgment.Suggested assessments included a straight leg raise and bedside ultrasound for evaluating fluid responsiveness.
- Bedside information: Participants considered bedside information more reliable than the data available to the AI and used this distinction to preserve their role as expert decision-makers.They emphasized that the algorithm lacked access to dynamic assessments and general patient appearance.
- Model representation: Quantile-based discretization into 25 dosage bins produced fluid recommendations that clinicians found unusually low compared with familiar practice.One example was 75 mL over four hours, which a participant described as an implausibly small fluid dose.
- Model representation: The AI’s 4-hour recommendation intervals fit trajectory review and relatively stable cases but made clinicians uncomfortable with higher-risk decisions requiring shorter reassessment cycles.Clinicians described one- to two-hour follow-up after fluid boluses as a buffer against uncertainty about treatment response.
- Clinical expectations: Clinicians expected recommendations to align with guidelines or their personal practices, and deviations could make the AI seem less useful or nonsensical.Eight clinicians expected the AI to recapitulate guideline-based practice, while five experienced clinicians used their own practice as the comparison standard.
- Evaluation and trust: Trust depended partly on upfront evidence about the model’s methodology, developers, validation, and effect on patient care, but such evidence would not replace case-specific clinical judgment.Participants proposed randomized trials, outcome associations, or replication of clinician decisions as possible evaluation approaches.
6 DISCUSSION
Clinicians used treatment recommendations in varied ways, often negotiating among recommendation components rather than simply accepting or rejecting them. Explanations increased confidence but did not consistently change reliance, while study design and setting constrained evaluation and generalizability.
- Reliance patterns: Most clinicians negotiated partial reliance, considering multiple plausible next steps and overriding recommendation components when contextual reasons supported doing so.This approach may better fit treatment decisions without an evident right answer, but clinicians had to determine which components to trust.
- Design implications: Supporting negotiation could help clinicians compare concurrent and sequential treatment strategies instead of following a rigid multi-hour treatment plan.The proposed interface would prioritize credible recommendation aspects and support comparisons clinicians may already make.
- Explainability: Explanations increased perceived usefulness and confidence, but explanations alone did not significantly change reliance on the AI.Feature explanations helped participants decide how much weight to place on recommendations, consistent with prior explainable-AI findings.
- Explainability: Explainable recommendations imposed mixed cognitive effects: they could waste time for clinicians who ignored the AI but prompt consideration of previously neglected options for negotiators.The authors suggest adjusting recommendation visibility and complexity according to user confidence or disagreement with the AI.
- Clinician characteristics: Observed reliance patterns did not appear correlated with clinician seniority or experience level, contrary to some prior findings about expertise and AI familiarity.The authors therefore do not support focusing adoption efforts only on novice clinicians.
- Evaluation challenges: Realistic integration complicates outcome validation because clinicians’ adoption and the AI’s effect on decisions become difficult to separate.Binary acceptance or rejection is also inadequate because participants often credited the AI without following its recommendation completely.
7 CONCLUSION
This paper examines clinician interaction with a real AI system that predicts treatment effects under uncertainty. It finds that clinicians who viewed the system as additional evidence found it most useful, motivating more nuanced reliance and evaluation approaches.
- Contribution: The study rigorously assesses clinicians’ interactions with a real AI system that predicts treatment-strategy effects under uncertainty.The system is intended to complement human decision-makers by revealing patterns in historical outcomes.
- Conclusion: Clinicians found the AI Clinician most useful when treating it as additional evidence alongside their own assessment.The authors propose individually validated recommendations as one way to clarify intended use and facilitate evaluation.
- Implications: The authors frame the work as a step toward appropriate reliance supported by human-centered algorithm design and more nuanced decision metrics.