Source-linked AI summary
Unremarkable AI: Fitting Intelligent Decision Support into Critical, Clinical Decision-Making Processes
Qian Yang, Aaron Steinfeld, John Zimmerman
TL;DR
Most clinical decision support tools fail to fit clinical workflows and collaborative practice. This paper designs and evaluates a meeting-integrated DST that embeds prognostics into automatically generated slides, finding that clinicians may more readily encounter and embrace unobtrusive support while highlighting unresolved validation and design-balance challenges.
Problem
Most deployed clinical decision support tools have failed in practice, with poor consideration of clinicians’ workflows and the collaborative nature of clinical work identified as a likely reason.
Method
The paper designs a DST that automatically generates slides for multidisciplinary VAD decision meetings, embedding prognostic support in a corner so it can be encountered naturally and ignored unless disagreement matters.
Results
The field evaluation suggests clinicians can encounter and embrace a DST integrated into their existing decision meetings and presented in an unobtrusive form.
Takeaways & Limitations
DST effectiveness should be evaluated as an integrated experience situated in social and physical contexts, not only by prediction accuracy.
Takeaways & Limitations
The study could not fully assess the design because clinicians considered an unvalidated DST unethical and required prospective validation, creating a chicken-and-egg problem.
Abstract
from arXiv · showhide
Clinical decision support tools (DST) promise improved healthcare outcomes by offering data-driven insights. While effective in lab settings, almost all DSTs have failed in practice. Empirical research diagnosed poor contextual fit as the cause. This paper describes the design and field evaluation of a radically new form of DST. It automatically generates slides for clinicians' decision meetings with subtly embedded machine prognostics. This design took inspiration from the notion of "Unremarkable Computing", that by augmenting the users' routines technology/AI can have significant importance for the users yet remain unobtrusive. Our field evaluation suggests clinicians are more likely to encounter and embrace such a DST. Drawing on their responses, we discuss the importance and intricacies of finding the right level of unremarkableness in DST design, and share lessons learned in prototyping critical AI systems as a situated experience.
1 INTRODUCTION
Clinical decision support tools promise diagnostic, treatment, and prognostic insights, yet most have failed in practice because they fit poorly with clinical workflows and social contexts. This paper proposes a workflow-situated DST that embeds prognostic support into required decision-meeting slides.
- 1 INTRODUCTION: Poorly fitting DSTs assume individual clinicians will recognize when help is needed, leave their workflow, and trust a separate system.These assumptions overlook clinicians’ workflow and the collaborative nature of clinical work.
- 1 INTRODUCTION: VAD implantation decisions are consequential because many recipients die shortly after receiving the device, despite VADs offering some patients their only chance to extend life.A prognostic DST could help identify patients most likely to benefit from implantation.
- 1 INTRODUCTION: Clinicians rarely encountered or actively engaged with computational support during VAD decisions, which usually did not occur at a computer.Most cases were not perceived as challenging, while hierarchical roles separated decision-making physicians from computer-using mid-level clinicians.
- 1 INTRODUCTION: The proposed DST automatically generates decision-meeting slides with prognostic support embedded unobtrusively in a corner.The design aims to present computational advice at a relevant time and place while slowing decisions only when it adds value.
- 1 INTRODUCTION: A field evaluation at three VAD hospitals suggests clinicians are more likely to encounter and embrace decision support bound to their existing work routine.The paper also examines whether this design might generalize beyond VAD decisions and discusses the right level of unremarkableness.
- 1 INTRODUCTION: Most clinician-facing DSTs failed when moving from research labs into clinical practice.Healthcare researchers attributed these failures primarily to insufficient HCI consideration rather than poor technical performance.
VAD Decision-Making and Its Context
VAD decisions occur through routine, collaborative interactions rather than isolated computer use, within a hierarchical culture that separates decision makers from computer users. The design therefore embeds support into existing workflow and makes it easy to ignore unless disagreement warrants attention.
- VAD Decision-Making and Its Context: Clinicians perceived little need for computational support because they considered most patient cases textbook cases following standard therapy escalation.This reduced motivation to seek decision support during routine cases.
- VAD Decision-Making and Its Context: VAD implant decisions occur during daily rounding, hallway conversations, and multidisciplinary meetings, but rarely in front of a computer.The multidisciplinary meeting is a required collective decision context involving the VAD team.
- VAD Decision-Making and Its Context: The workplace culture was strongly hierarchical yet highly collaborative, separating senior physician decision makers from mid-level computer users.These groups rarely overlap during the decision-making process.
- VAD Decision-Making and Its Context: The design sought to embed DST output into clinicians’ current workflow because they were unlikely to recognize when help was needed and walk to a computer.This reverses the conventional model of a separate support system waiting for clinicians to initiate use.
- VAD Decision-Making and Its Context: DST outputs should be easily ignored for textbook cases but present enough to slow decisions when clinicians’ judgments and the DST disagree.The goal is selective interruption rather than constant intervention.
- VAD Decision-Making and Its Context: Unremarkable Computing frames technology as significant yet natural and subservient to everyday routines, an orientation the authors apply to VAD decision making.The paper treats implant decisions as life-and-death decisions that are nevertheless part of clinicians’ work routines.
Design Process
The design places prognostic support where multidisciplinary VAD decisions already occur: in automatically generated meeting slides. Small, corner-positioned visualizations aim to build trust when predictions agree and prompt attention when they conflict, but synthetic cases were required for design completion.
- Design Process: The multidisciplinary patient evaluation meeting was selected because it brings decision participants together while using a computer at a common hospital touch point.Regularly scheduled multidisciplinary meetings are common across hospital sites.
- Design Process: The DST was integrated into a meeting-slide generator that automatically extracts patient information from electronic medical records.This reduced clinicians’ data-entry effort and augmented paperwork already used in meetings.
- Design Process: The final prognostic display used a small line chart showing predicted survival and likely causes of death.The design was iterated with an attending cardiologist and a nurse practitioner.
- Design Process: The chart was placed in the slide’s top-right corner so agreement could support trust without slowing routine decisions, while conflict could prompt reconsideration.Subtlety was deliberately chosen to achieve the intended level of unremarkableness.
- Design Process: Because policies and legal regulations barred use of real patient data, collaborators populated the slides with synthetic cases assembled from former cases.Creating prototypical synthetic patients was challenging because each case required dozens of vital signs and test results.
- Design Process: The slide combined DST outputs with a summarized patient history, categorized test results, demographics, and social and financial evaluation links.The detailed contents were developed with clinicians and referenced existing meeting printouts and workup checklists.
4 DESIGN ASSESSMENT
The assessment examined workflow encounter, public acceptance, and selective interruption across three US VAD hospitals. Access restrictions prevented full evaluation in actual implant decisions, so the researchers combined interviews, limited meeting presentations, observation, and thematic analysis.
- 4 DESIGN ASSESSMENT: The assessment asked whether clinicians would encounter the DST naturally, accept it in public meetings, and treat its corner placement as appropriately unremarkable.It also examined whether aligned predictions would be ignored and conflicting predictions would slow decisions.
- 4 DESIGN ASSESSMENT: The study involved three geographically and organizationally varied US hospitals performing approximately 40 to over 100 VAD implants annually.The sites included two hospitals from the formative field study and one new hospital.
- 4 DESIGN ASSESSMENT: None of the hospitals allowed slides containing information about patients currently undergoing implantation to be presented in an actual decision meeting.Clinicians and site policies cited potential effects on life-and-death decisions, workload, and meeting observation restrictions.
- 4 DESIGN ASSESSMENT: The researchers adapted the assessment to site constraints, omitting meeting presentation at hospital B and physician-and-surgeon interviews at hospital A.All procedures were conducted at hospital C.
- 4 DESIGN ASSESSMENT: Researchers interviewed mid-level clinicians and attending physicians, presented the DST in two hospitals’ meetings, observed responses, and conducted follow-up interviews.The study included nine attending cardiologists or surgeons and eight mid-level clinicians, with interviews lasting at least one hour.
- 4 DESIGN ASSESSMENT: Interview data were audio-recorded, transcribed, and analyzed using affinity diagrams and thematic analysis.Field notes were also recorded during the assessment.
Assessing Generalizability of the DST Design
The study probed whether embedding the DST in interdisciplinary decision meetings might generalize beyond VAD implantation. Interviews with six physicians across other medical domains informed this assessment alongside observations at three VAD hospitals.
- Rationale: Decision meetings were selected partly because they are established practices in other critical medical domains.
- Cross-domain interviews: The researchers interviewed six physicians whose practices included decision meetings in six medical domains.Domains included pediatric surgery, pediatric critical care, adult cardio-thoracic surgery, emergency internal medicine, orthopedic surgery, and obstetrics/gynecology.
- Three-site evaluation: The field evaluation compared three hospitals’ cultures, facilities, practices, encounter likelihood, acceptance, unremarkableness, and generalizability.
- Site variation: Hospital A’s limited technology infrastructure included a recent transition from paper records and blocked common web services.
- Site variation: Hospital B combined strong risk-modeling expertise with a minimalist approach to using models in its own implant decisions.
- Site variation: Hospital C had accessible meeting technology and an experienced nurse practitioner who operated the presentation system.
- Prior computational support: Hospital C discontinued a manually entered DST practice after concerns about model miscalibration, while other EMR models went unused because they required manual data entry.
Likelihood of Encountering DST in Workflow
Clinicians were likely to encounter DST output in recurring decision meetings, where it could add factual context without replacing clinical judgment. Slides also offered mid-level clinicians a formal vehicle for communicating concerns and supporting their influence.
- Encounter opportunities: Weekly implant decision meetings occurred at all three hospitals and were among the few events shared across sites.
- Encounter opportunities: Meetings placed senior clinicians near shared computers, unlike other decision points that were often informal discussions without EMR access or records.
- Acceptance: No interview participants resisted including DST output in decision meetings, although one hospital had abandoned manual inclusion after losing confidence in model quality.
- Clinical judgment: A prognostic DST could provide a more factual view when emotionally difficult decisions made objectivity challenging.
- Clinical judgment: Senior physicians valued a DST’s statistical perspective across many cases as additional context for decisions based on fewer recent experiences.
- Roles and hierarchy: Mid-level clinicians described their role as informing and supporting discussions rather than making implant decisions.
- Automation: A slide generator could automate non-billable preparation work while removing irrelevant EMR data from physician presentations.
- Roles and hierarchy: Formal meeting slides could amplify mid-level clinicians’ voices by presenting their concerns as visible facts rather than unsupported opinions.
Intricacies of Making DST Unremarkable
Clinicians appreciated support that intervened only when necessary, but the evaluation exposed limits in judging responses to conflicting predictions and in interpreting synthetic, model-centered cases.
- Right level of unremarkableness: Clinicians appreciated DSTs that could slow decisions only when necessary, but the evaluation could not establish whether this design achieved that goal.
- Evaluation boundary: Clinicians said paper-based patient data could not reproduce their experience of making critical clinical decisions, limiting assessment of reactions to conflicting predictions.
- Clinical context: Patient history alone was insufficient for confident implant decisions because clinicians needed to see, talk to, and care for the patient as a whole.
- Interpretive variation: The same prognostics for two synthetic cases produced sharply different interpretations, ranging from futility to immediate implantation.
- Synthetic-case effects: Synthetic cases made real patient discussion more difficult and redirected attention toward the DST’s provenance, mechanism, and quality.
- Credibility: Clinicians viewed an unvalidated model as ethically problematic and sought evidence of clinical validation, publication, and local relevance.
Is the Model Validated by Clinical Trials?
Clinicians questioned whether prognostic outputs reflected validated clinical evidence, causal factors, or merely historical associations. They wanted models that supported modifiable interventions while preserving human management of uncertainty.
- Validation: Physicians wanted models validated with local hospital data, published in reputable journals, and tested nationally across implant centers.
- Human judgment: Clinicians stated that predictive models could not replace human decision-making because VAD implantation inherently involves uncertainty management.
- Actionability: Clinicians contrasted the DST’s static patient view with their interest in future actions that might improve modifiable risk factors.
- Actionability: They considered understanding which features drive predictions important for identifying what can be changed for an individual patient.
- Causality: Clinicians did not consistently distinguish predictive features from causal factors, although they regarded that distinction as important.
- Causality: Clinicians viewed correlation-versus-causality distinctions as central to deciding whether prognostic outputs should inform care.
- Interpretation: They wanted prognostics treated as facts grounded in historical data rather than predictions carrying agency or subjectivity.
Are Data-Driven Prognostics Facts OR Predictions?
Clinicians treated prognostic DST outputs as population-level averages rather than individualized facts, and interpreted them alongside patient-specific and institutional factors. Ambiguity about prediction timing further complicated their use in implant decisions.
- Clinicians commonly viewed DST outputs as averages, making personalized predictions difficult to grasp.
- Some clinicians considered applying population statistics to individual patients unethical.
- A displayed 21-day life-expectancy prediction was unclear because “now” did not match the likely implantation date.Clinicians questioned whether the estimate began on the meeting date or after implantation.
- Clinicians treated DST output as only one factor among patient-specific and institutional “X factors.”
- For some cardiac surgeries, officially defined risk models had already made decision meetings center on surgeon and care-team ratings, unlike VAD implants.
Generalizability Beyond VAD
The authors examined whether a DST situated in multidisciplinary decision meetings could extend beyond VAD care. Interviews indicated that such meetings are widespread across high-consequence clinical domains, while workflow integration and routine support may improve adoption.
- Multidisciplinary decision meetings occur across many clinical domains for aggressive, last-option interventions.
- Examples span cancer, chronic disease, pediatric syndromes, surgery, medication management, and emergency-room care.
- Clinical DSTs have mostly failed outside laboratories when their designs lacked contextual integration with healthcare practice.
- The paper argues that DSTs should be designed as integrated experiences whose effectiveness includes social and physical context, not only prediction accuracy.
- The proposed design places prognostic output in meeting slides and keeps it unobtrusive unless predictions conflict with a seasoned physician’s course.
- Decision meetings provide an existing routine, a socially aggregated decision point, and time for clinicians to deliberate over prognostics.
- Because multidisciplinary meetings are promoted in VAD care and occur across clinical domains, the design could potentially transfer across hospitals and practices.
- The authors identify other socially aggregated, deliberative routines as future opportunities rather than claiming decision meetings are the only integration path.
Interaction Form.
The paper frames Unremarkable AI as support embedded in existing clinical routines: visible enough to matter when needed, but unobtrusive during ordinary decisions. Field responses suggest this reduced resistance while also exposing a trade-off between augmentation and transformation.
- Designing a Right Level of Unremarkableness: Unremarkable AI embeds DST interactions in an existing decision-making routine instead of requiring clinicians to seek help separately.
- Designing a Right Level of Unremarkableness: The DST is designed to be noticed mainly when its information may add value to the decision.
- Designing a Right Level of Unremarkableness: Field assessment found positive indicators that an unremarkable DST reduced clinicians’ resistance to clinical decision support.
- Designing a Right Level of Unremarkableness: Clinicians did not appear threatened or concerned that the technology would replace them, and valued its role in informing discussion without dominating it.
- Designing a Right Level of Unremarkableness: Despite its visually unremarkable form, the DST introduced predictive reasoning into a culture rooted in facts and statistical significance.
- Designing a Right Level of Unremarkableness: When risk models officially measured patient risk and clinician skills, decision making became centered on those models and performance pressure became more substantial.
- Designing a Right Level of Unremarkableness: The authors identify the preferred DST role and the right level of unremarkableness as unresolved design questions.
- Designing a Right Level of Unremarkableness: Future research must examine the trade-off between naturally augmenting decisions and transforming clinical decision making.
Experience Prototyping DST In-Situ
The field assessment treated situated DST prototyping as a constrained design problem requiring generalizable systems, experience-focused evaluation, and flexible prototypes. Realistic in-situ assessment remained limited because clinicians needed real patient data and validated functioning models.
- Restricted access to clinical environments creates fundamental challenges for iterative UX design and evaluation.
- The study designed for a class of structurally similar decisions to improve generalizability across hospitals and clinical settings.
- The evaluation methods aimed to describe and unpack complex, subtle, multifaceted experience rather than explicitly measure it.
- Researchers used prototypes instead of functioning DST models so they could probe outputs and incorporate participant feedback easily.
- Assessing whether the DST achieved the right level of remarkableness was impossible without real patient data and fully functioning ML systems.
- Clinicians needed to see their own patients’ data to connect DST information with actual decision making and assess its effect on care.
- Clinicians regarded an unvalidated DST as unethical and misleading and requested validation through randomized trials on retrospective and prospective populations.
- The authors expect similar assessment challenges for other critical, high-consequence decisions and call for new design-assessment methods.