Source-linked AI summary

Unmasking Bias in AI: A Systematic Review of Bias Detection and Mitigation Strategies in Electronic Health Record-based Models

Feng Chen, Liqin Wang, Julie Hong, Jiaqi Jiang, Li Zhou

arXiv:2310.19917v3cs.AIcs.CYcs.LGq-bio.QM

TL;DR

Bias in EHR-based AI can contribute to unequal performance across patient groups, motivating a systematic review of how such bias is evaluated and mitigated. The review synthesized studies across the model-development cycle and found six bias types, frequent use of fairness metrics and preprocessing mitigation, and limited standardization and clinical validation.

  • Problem

    Biases in EHR data and AI models can produce differential performance across patient subgroups and potentially exacerbate healthcare disparities.

  • Method

    The authors conducted a PRISMA-guided systematic review of EHR-based AI bias detection, evaluation, and mitigation strategies across model development.

  • Results

    The review identified six major bias types, found fairness assessment in approximately half of studies, and showed preprocessing was the most commonly used mitigation stage.

  • Takeaways & Limitations

    The findings support combining fairness metrics with traditional performance metrics while developing standardized, generalizable, and interpretable bias-management methods.

  • Takeaways & Limitations

    None of the reviewed models had been tested in actual clinical settings, so effects of bias handling on clinical outcomes were not rigorously measured.

Abstract

from arXiv · show

Objectives: Leveraging artificial intelligence (AI) in conjunction with electronic health records (EHRs) holds transformative potential to improve healthcare. Yet, addressing bias in AI, which risks worsening healthcare disparities, cannot be overlooked. This study reviews methods to detect and mitigate diverse forms of bias in AI models developed using EHR data. Methods: We conducted a systematic review following the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guidelines, analyzing articles from PubMed, Web of Science, and IEEE published between January 1, 2010, and Dec 17, 2023. The review identified key biases, outlined strategies for detecting and mitigating bias throughout the AI model development process, and analyzed metrics for bias assessment. Results: Of the 450 articles retrieved, 20 met our criteria, revealing six major bias types: algorithmic, confounding, implicit, measurement, selection, and temporal. The AI models were primarily developed for predictive tasks in healthcare settings. Four studies concentrated on the detection of implicit and algorithmic biases employing fairness metrics like statistical parity, equal opportunity, and predictive equity. Sixty proposed various strategies for mitigating biases, especially targeting implicit and selection biases. These strategies, evaluated through both performance (e.g., accuracy, AUROC) and fairness metrics, predominantly involved data collection and preprocessing techniques like resampling, reweighting, and transformation. Discussion: This review highlights the varied and evolving nature of strategies to address bias in EHR-based AI models, emphasizing the urgent needs for the establishment of standardized, generalizable, and interpretable methodologies to foster the creation of ethical AI systems that promote fairness and equity in healthcare.

Background and significance

EHR-based AI can support clinical decision-making but may reproduce data and model biases that produce differential performance across patient groups. A focused systematic review was needed because prior reviews had not specifically examined bias in EHR-derived AI models.

  • Background and significance: Biases in EHR data and AI models can produce differential performance across patient subgroups and potentially exacerbate healthcare disparities.Examples include documentation inconsistencies, variable data quality, model inaccuracies, and disproportionate representation of demographics, conditions, or treatments.
  • Background and significance: A focused review of bias in AI models derived from EHR data was notably absent from existing scoping reviews.
  • Background and significance: The review addresses the need to identify, summarize, and propose strategies for managing biases in EHR-based AI models.

Objective

The study systematically reviews bias in EHR-based AI models, emphasizing how major bias types can be identified, evaluated, and mitigated across model development. It also highlights research directions intended to reduce potential impacts on healthcare disparities.

  • Objective: The study synthesizes literature on identifying, evaluating, and mitigating major biases across the EHR-based AI model development cycle.
  • Objective: The review aims to improve understanding of bias management and highlight directions for reducing potential AI impacts on healthcare disparities.

Materials and methods

The review followed PRISMA-guided searches and screening of EHR-based AI studies published from 2010 through December 17, 2023. Reviewers extracted model, bias, mitigation, and fairness information and organized bias across development stages.

  • Materials and methods: The authors searched PubMed/MEDLINE, Web of Science, and IEEE for English-language EHR-based AI studies published from January 1, 2010, through December 17, 2023.The review complied with the 2021 PRISMA guidelines.
  • Materials and methods: At least two reviewers independently screened records and full texts, resolving disagreements through team consensus.
  • Materials and methods: Data extraction covered bibliographic information, EHR model characteristics, reported bias types, detection or mitigation strategies, and bias-evaluation metrics.
  • Materials and methods: Bias categories were defined by combining findings from included studies with broader healthcare-AI literature and established risk-assessment tools.
  • Materials and methods: The framework examined bias during data collection and preparation, model training and testing, and model deployment, classifying strategies as preprocessing, in-processing, or postprocessing.

Results

The review identified six major bias types across 20 included studies, with most research addressing implicit or selection bias and using preprocessing mitigation. Fairness assessment was inconsistent, and most mitigation studies reported improved performance after intervention.

  • Results: Six primary bias types were identified: algorithmic, confounding, implicit, measurement, selection, and temporal bias.
  • Results: Of 450 records, 20 studies met the review criteria after duplicate removal, screening, full-text assessment, and exclusion of studies lacking clear bias evaluation or mitigation methods.
  • Results: The reviewed models primarily addressed predictive tasks including diagnosis, risk, treatment, progression, mortality, medication, health-status, and missingness prediction.
  • Results: Twelve studies (60%) used group-fairness metrics, whereas eight (40%) relied only on performance metrics such as sensitivity, specificity, accuracy, AUROC, or MSE.Fairness analyses included parity-, confusion-matrix-, calibration-, and score-based metrics.
  • Results: Implicit bias appeared in 11 studies (55%), while selection and algorithmic bias each appeared in 6 studies (30%).
  • Results: Among 15 studies seeking to mitigate bias, 12 (80%) reported improved performance after mitigation.
  • Results: Preprocessing accounted for 11 mitigation studies (73.3%), compared with 3 in-processing studies (20%) and 1 postprocessing study (6.7%).Approaches included resampling, reweighting, transformation, relabeling, blinding, transfer learning, and transformation-based output adjustment.

Discussion

The review identifies six bias types in EHR-based AI and finds that fairness assessment and mitigation remain methodologically uneven. Preprocessing dominates current mitigation, while standardized, generalizable, interpretable approaches and pipelines addressing multiple biases remain needed.

  • Six bias types were identified, with research concentrating especially on implicit and selection biases.
  • Fairness assessment is inconsistent because many studies rely on overall performance metrics that may miss subgroup disparities.
  • Preprocessing was the dominant mitigation stage, using resampling, reweighting, transformation, relabeling, and blinding, whereas postprocessing was rarely used.
  • Preprocessing can address class or group imbalance but may lose data and remains limited for confounding, algorithmic, and temporal bias.
  • Reported interventions produced mixed outcomes, including improved performance, unchanged performance, metric-dependent variability, and cases where reweighting introduced new bias.
  • Examples include a 1%-2% improvement from time alignment, fairness gains while maintaining accuracy, and better performance from domain adaptation and missing-data modeling.

Conclusion

The review finds growing attention to bias in EHR-derived AI and calls for standardized, generalizable, and interpretable methods to detect, mitigate, and evaluate it.

  • Standardized, generalizable, and interpretable methods are needed to detect, mitigate, and evaluate bias in EHR-derived AI models.
  • As AI use in healthcare expands, equitable technologies are critical for minimizing bias-related healthcare disparities.
  • Continued research is essential to improve healthcare equity and maximize AI’s benefits.
Loading 2310.19917v3…