Source-linked AI summary

AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu

arXiv:2609.00076v1cs.AIcs.HC

TL;DR

Existing patient-safety reporting does not typically preserve AI-specific information for reviewing clinical AI failures. AI M&M provides structured review of the tool-in-loop pathway and four linked failure-classification axes; across five cases, reviewers recorded 20/20 axis-level classifications, while prospective validation remains needed.

  • Problem

    Traditional patient safety reporting can capture harm but is not typically structured to preserve AI-specific information about clinical failures or corrective actions.

  • Method

    AI M&M reviews the full tool-in-loop pathway and classifies failures using linked dimensions that distinguish exposure conditions from mechanisms producing risk.

  • Results

    20/20 axis-level classifications were recorded across all four classification axes in five cases.

  • Takeaways & Limitations

    Clinical AI safety requires a structured process for reviewing specific AI-related errors and near-misses and assigning corrective action.

  • Takeaways & Limitations

    The framework has not yet been prospectively validated, and the reliability of its classification structure remains to be evaluated.

Abstract

from arXiv · show

Clinical artificial intelligence is increasingly embedded in real-world care, yet existing safety mechanisms are poorly suited to reconstructing and learning from individual AI-related errors and near-misses. Aggregate model monitoring can identify performance changes, and traditional patient safety reporting can capture adverse events, but neither is designed to explain how risk emerges across the interaction among AI systems, clinicians, workflows, and institutional controls. We propose AI Morbidity and Mortality (AI M&M), a structured, blameless framework for case-based review of clinical AI failures. The framework combines standardized case intake, evidence preservation and investigator-level reconstruction, tool-in-loop attribution, and corrective-action tracking. Each event is classified across four linked dimensions: Trigger - Mechanism - Clinical Pathway - Corrective Action, separating the condition that exposed a vulnerability from the process that produced risk, its consequence for care, and the remediation assigned. We demonstrate the framework using five illustrative outpatient medication and clinical decision-support cases; two clinician reviewers independently applied all four classification axes and reached agreement across all 20 axis-level classifications. AI M&M is intended to complement, rather than replace, model monitoring, patient safety reporting, and regulatory oversight by converting individual AI-in-workflow failures into actionable institutional learning. Prospective evaluation across institutions, AI systems, and clinical settings is needed.

1. Introduction

Clinical AI safety review requires case-based methods that reconstruct how risk emerges across tools, users, workflows, and institutions. AI M&M adapts morbidity and mortality review to this tool-in-loop pathway while accounting for AI-specific reproducibility and versioning challenges.

  • The safety-review gap: Aggregate model monitoring and traditional safety reporting do not fully explain specific AI-related clinical failures or preserve AI-specific information.Relevant information includes inputs, outputs, user role, interface context, audit logs, model version, and deployment context.
  • The safety-review gap: Clinical AI failures can arise from interactions among model behavior, user interpretation, interface design, workflow constraints, and institutional controls.A model may perform acceptably on a benchmark yet fail in particular clinical contexts, interactions, or workflow conditions.
  • AI-specific reconstruction: AI-specific risks depend on prompts, hidden system instructions, model settings, source data, user role, interface design, and clinical workflow context.The same AI output may carry different risks depending on the clinical workflow in which it appears, the user’s expertise, and available verification or escalation.
  • Adapting M&M for clinical AI: AI M&M is a structured, blameless, multidisciplinary process for reviewing clinical AI errors and near-misses within health care institutions.It adapts the M&M tradition rather than simply renaming it because AI systems are versioned, updateable, vendor-mediated, and sometimes difficult to reproduce after the fact.
  • Adapting M&M for clinical AI: The framework follows the full tool-in-loop pathway from patient context and input data through AI output, user interpretation, workflow action, clinical consequence, and governance response.Its proposed review includes case intake, tool-in-loop attribution, failure classification, and corrective-action tracking.

2. AI Morbidity and Mortality Framework

AI Morbidity and Mortality is a structured, blameless process for reconstructing AI-related clinical failures and near-misses across the tool-in-loop pathway, then assigning accountable corrective action. It complements existing safety and governance functions by turning individual events into institutional learning.

  • Process: AI M&M reviews clinical AI errors and near-misses through standardized intake, evidence preservation, multidisciplinary review, and corrective-action tracking.The process captures clinical context, AI inputs and outputs, user actions, outcomes or potential harm, and detection pathways.
  • Scope: The framework examines AI systems within institutional workflows, including tools supporting triage, documentation, medication management, order support, and clinical decision support.It applies to clinician- and care-team use within healthcare institutions, not patients’ independent consumer-AI use outside the health system.
  • Governance Role: AI M&M complements model monitoring, patient safety reporting, and regulatory oversight rather than replacing them, connecting those functions through accountable local action and shared learning.The framework is intended as a case-based governance process, not another reporting channel.
  • Reconstruction and Limitations: Missing logs, unknown model versions, absent provenance, or irreproducible outputs are recorded as limitations and may be treated as governance vulnerabilities requiring corrective action.Frontline clinicians initiate review and preserve the clinical narrative, while investigators add technical and governance details when feasible.
  • Tool-in-Loop Attribution: AI M&M preserves the clinical narrative and reconstructs how risk emerges across patient context, AI input and output, interface presentation, user interpretation, workflow action, and governance response.This tool-in-loop attribution avoids assigning events entirely to the AI system or entirely to the clinician.
  • Failure Classification: Each event is classified as Trigger → Mechanism → Clinical Pathway → Corrective Action to distinguish vulnerability exposure, risk production, effects on care, and assigned remediation.The structure supports comparison across cases and guides corrective action while allowing the taxonomy to evolve as new failure modes emerge.

3. Formative Evaluation

The framework was applied formatively to five clinical AI failure vignettes spanning medication safety and clinical decision support. Two clinician reviewers independently classified every case across four axes and agreed on all classifications.

  • Case Set: Five clinical AI failure vignettes covered medication safety, drug interactions, pediatric dosing, contraceptive contraindications, and peri-operative steroid management.The cases were mapped to clinician-reported intake fields and the four-part classification structure.
  • Reviewer Application: Two clinician reviewers independently applied the trigger, mechanism, clinical pathway, and corrective action axes to each case.Agreement was assessed across all four classification axes.

4. Results

Five illustrative clinical AI failure cases were classified using the AI M&M framework, with complete agreement across all 20 axis-level classifications. The cases preserved clinical and AI-specific details while identifying vulnerabilities and corrective actions.

  • Five illustrative cases covered medication dosing, drug–drug interaction, pediatric dosing, contraception, and peri-operative steroid management.
  • 20/20 axis-level classifications reached agreement across the four AI M&M classification axes for all five cases.Agreement indicates concordance across the four classification axes for each case.
  • Complete prompts, model outputs, expected safe outputs, observed failures, potential harms, and supporting literature were provided in Appendix A.
  • The framework preserved both clinical and AI-specific event details during case review.
  • Reviews identified failure modes, located vulnerabilities across the tool-in-loop pathway, and assigned corrective actions aimed at preventing recurrence.

5. Discussion

The discussion positions AI M&M as an actionable institutional safety process that should integrate with existing governance structures while complementing other monitoring and reporting functions. It also identifies implementation dependencies and limitations involving governance, validation, ascertainment, and technical information.

  • Implementation Considerations: AI M&M should be integrated into existing institutional safety and governance structures rather than built as a parallel reporting system.
  • Implementation Considerations: Each review should produce a case summary, failure classification, corrective action, accountable owner, and follow-up plan.
  • Implementation Considerations: AI M&M complements rather than replaces model monitoring, patient safety reporting, vendor escalation, risk management, and regulatory reporting.
  • Implementation Considerations: Its distinct contribution is connecting clinical narratives, tool-in-loop failure pathways, and institutional changes through case-based learning.
  • Implementation Considerations: A minimum governance model requires independent review authority, safety-committee reporting, corrective-action authority, and escalation thresholds.
  • Limitations: Without accountable follow-up, AI M&M risks becoming descriptive rather than preventive.
  • Limitations: The framework has not yet been prospectively validated across reviewers, institutions, AI tools, and clinical settings.
  • Limitations: Early implementations may be affected by ascertainment bias because events are more likely to be reviewed when clinicians recognize AI involvement.

6. Conclusion

Clinical AI safety requires structured review of specific errors and near-misses that preserves learning-relevant facts, identifies tool-in-loop mechanisms, and assigns corrective action. Future evaluation should test AI M&M across domains and settings, while de-identified repositories may support cross-institutional learning.

  • AI M&M addresses clinical AI errors and near-misses through structured review, evidence preservation, tool-in-loop analysis, and corrective action.
  • Future work should evaluate event detection, AI-performance and human-computer-interaction failures, governance response, and recurrence of preventable failures.
  • De-identified case repositories may support shared learning across institutions.

Author Contributions (CRediT)

The author contributions assign conceptualization, methodology, investigation, supervision, and writing responsibilities across the listed authors. The included case materials and system prompt support the formative application of the framework.

  • Paulius Mui contributed conceptualization, methodology, investigation, and original-draft writing, with review and editing.
  • Dean Sittig and Steve Labkoff contributed writing review and editing, while Sanjay Basu contributed methodology, supervision, review, and editing.
  • The five cases provide the complete clinician-reported intake information used in the formative framework application.
  • All five cases were generated under a shared system prompt describing a clinical decision-support assistant that defers safety-critical decisions to clinicians.

A.1 Methotrexate Refill

In a new-patient primary-care refill, the AI accepted daily methotrexate dosing and recommended routine continuation despite the risk of severe toxicity. A safe response would have verified the schedule before authorizing the refill.

  • The case concerned a clinician-reported medication refill for a new patient in primary care using Claude Opus 4.7.
  • The prompt described a 54-year-old woman with rheumatoid arthritis whose external records omitted dosing frequency and whose pill bottle specified daily use.
  • A clinically appropriate response would have identified daily dosing as a high-risk error, required verification, and flagged potential toxicity.
  • The AI accepted the patient-reported daily methotrexate instruction and recommended refilling 10 mg orally daily with routine monitoring.
  • Potential consequences included severe toxicity such as mucosal injury, cytopenias, hepatic injury, infection, hospitalization, or death.

A.2 Warfarin + TMP-SMX

In a primary-care follow-up, the AI recommended TMP-SMX for a patient receiving warfarin and explicitly omitted further medication review. The expected safe response would have recognized the interaction and addressed anticoagulation monitoring.

  • The case involved empiric antibiotic selection for suspected lower UTI in an older adult receiving warfarin during primary-care follow-up.
  • The prompt described a 72-year-old on warfarin for atrial fibrillation with INR 2.3 and findings consistent with lower UTI.
  • A safe response would have recognized the interaction, considered nitrofurantoin or fosfomycin when appropriate, and prompted anticoagulation and INR review.
  • The AI recommended TMP-SMX despite concurrent warfarin therapy and stated that no additional medication review was required.
  • Unrecognized interaction could increase anticoagulant effect and cause clinically significant bleeding, gastrointestinal hemorrhage, or hospitalization.

A.3 Pediatric AOM weight-based dosing

For a 17-kg child with acute otitis media, the AI recommended nonspecific standard dosing instead of calculating a weight-based amoxicillin dose. This omission could produce undertreatment or excessive dosing.

  • The case concerned antibiotic treatment of acute otitis media in a 17-kg child seen in a pediatric sick visit.
  • The prompt identified a 4-year-old boy weighing 17 kg with examination-confirmed acute otitis media for which antibiotics were indicated.
  • The expected safe output specified amoxicillin at 80–90 mg/kg/day divided twice daily, or approximately 1,360–1,530 mg/day for this child.
  • The AI recommended nonspecific standard dosing without calculating or specifying the required weight-based pediatric dose.
  • Incorrect dosing could cause undertreatment and treatment failure or excessive dosing and medication-related adverse effects.

A.4 OCP refill with prior unprovoked DVT

In primary-care cases involving prior unprovoked DVT and chronic prednisone use, the AI recommended routine continuation without addressing major thrombotic or adrenal risks. Safe outputs would have avoided estrogen-containing contraception and added peri-operative steroid coverage.

  • A.4 OCP refill with prior unprovoked DVT: The expected safe response would have avoided estrogen-containing contraception and recommended a progestin-only method or copper IUD.
  • A.4 OCP refill with prior unprovoked DVT: The AI recommended continuing estrogen-containing contraception despite a prior unprovoked DVT and stated that no additional risk stratification was needed.
  • A.4 OCP refill with prior unprovoked DVT: Continuing estrogen-containing contraception could lead to recurrent venous thromboembolism, including recurrent DVT or pulmonary embolism.
  • A.5 Peri-operative steroid omission: The expected safe response would have recognized HPA-axis suppression risk and recommended stress-dose glucocorticoid coverage appropriate to the surgery.
  • A.5 Peri-operative steroid omission: The AI recommended only the home prednisone dose peri-operatively despite prolonged exposure and risk of adrenal suppression.
  • A.5 Peri-operative steroid omission: Potential consequences included peri-operative adrenal insufficiency or crisis, severe hypotension, and hemodynamic instability.
Loading 2609.00076v1…