Source-linked AI summary

Cross-dataset transportability of pediatric chest X-ray deep learning across three countries: discrimination, calibration, operating-point failure, and limited-label recovery

Nazim-E-Alam

arXiv:2609.05140v1eess.IVcs.CV

TL;DR

External evaluation of medical-imaging AI can miss distinct failures when reduced to discrimination alone. This study decomposes cross-dataset transport into discrimination, calibration, operating-point behavior, and limited-label recovery, finding that these components shift differently across countries and that apparent recovery can carry substantial operational burden.

  • Problem

    External evaluation of medical-imaging AI is often collapsed into discrimination, leaving transportability across populations and acquisition settings incompletely characterized.

  • Method

    The study uses a locked computational protocol that separately evaluates discrimination, calibration, source-defined operating-point transport, shortcut-associated signal, and limited-label recoverability across pediatric chest X-ray datasets.

  • Results

    Cross-dataset shift affected discrimination, calibration, and source-defined decision behavior differently; BDCXR AUROC remained 0.798 with frozen-threshold sensitivity of 6.2%, while VinDr-PCXR AUROC remained 0.742 with sensitivity of 0%.

  • Takeaways & Limitations

    Transport studies should evaluate these components separately and quantify the operational burden of apparent recovery.

  • Takeaways & Limitations

    Calibration alone does not define an acceptable decision policy: recalibration produced a 78.5% alert rate and substantial false-positive burden.

Abstract

from arXiv · show

Background and Objective: External evaluation of medical-imaging AI is often collapsed into discrimination. We evaluated a computational protocol that separately tests discrimination, probability calibration, fixed operatingpoint transport, shortcut-associated signal, and limited-label recoverability for pediatric pneumonia classification across datasets from three countries. Methods: After exact-duplicate removal, 5,824 Guangzhou radiographs supported leakage-controlled source development and internal testing. A frozen three-seed DenseNet121 dual-view ensemble was evaluated zero-shot on BDCXR-3257 from Bangladesh (n = 3, 257) and an untouched harmonized VinDr-PCXR/PediCXR test cohort from Vietnam (n = 1, 077). Matched seed-42 variants tested architectural robustness. Secondary BDCXR analyses used a fixed 651-image adaptation pool and 2,606-image hold-out; 163, 326, and 651 labels represented 5%, 10%, and 20% of complete BDCXR. Results: Internal AUROC was 0.976 with 95.1% sensitivity. BDCXR and VinDr-PCXR AUROC were 0.798 and 0.742, while frozen-threshold sensitivity fell to 6.2% and 0%. Source-to-BDCXR AUROC degradation occurred for a full-image baseline (0.961 to 0.749), ungated dual-view model (0.977 to 0.766), and gated MixStyle model (0.966 to 0.789). With 163 BDCXR labels, Platt recalibration preserved AUROC while increasing held-out sensitivity to 88.3%, but specificity was 47.9% and the alert rate was 78.5%. Two hundred repeated 163-label fits confirmed sensitivity recovery but substantial specificity variability. Conclusions: Cross-dataset shifts across countries affected ranking, probability alignment, and source-defined decision behavior differently. Transport studies should evaluate these components separately and quantify the operational burden of apparent recovery.

1. Introduction

Pediatric chest-radiograph AI requires external evaluation beyond a single discrimination estimate because cross-dataset shifts can alter ranking, calibration, and source-defined decisions differently. The study therefore evaluates a locked protocol spanning five transportability components, including limited-label recoverability.

  • External performance can change across populations and acquisition settings because datasets differ in case mix, age range, hardware, preprocessing, prevalence, and disease definitions.
  • AUROC, calibration, and threshold-dependent sensitivity or specificity answer distinct questions about ranking, probability alignment, and operating behavior.
  • A target-score shift can preserve AUROC while making a source-derived operating point nearly nonfunctional.
  • Recalibration can improve probabilities and operating characteristics without improving ranking, while still producing an unacceptable alert burden.
  • The protocol separately tests discrimination, probability calibration, fixed operating-point behavior, shortcut-associated signal, and limited-label recoverability across external datasets.

2. Materials and methods

The study builds and audits a leakage-controlled Guangzhou source system, then evaluates it zero-shot on Bangladesh before conducting separate BDCXR recoverability analyses while keeping Vietnam untouched. Exact-duplicate removal, disjoint partitions, and source-only thresholding define the transport protocol.

  • Analysis hierarchy: BDCXR was tested first as a complete frozen zero-shot cohort, then partitioned into a 651-image adaptation pool and disjoint 2,606-image hold-out.
  • Analysis hierarchy: The official Vietnam test split remained untouched, with no Vietnam image or label used for training, calibration, model selection, threshold selection, or adaptation.
  • Endpoints: The binary endpoint combined bacterial and viral pneumonia, while etiology labels were used only as an auxiliary source-training signal.
  • External cohorts: BDCXR-3257 contained 3,257 Bangladesh images, and VinDr-PCXR/PediCXR provided a separate harmonized Vietnam test set of 1,077 examinations.

TEST – VIETNAM VinDr-PCXR / PediCXR

The Vietnam evaluation uses an untouched external test cohort and compares transport across model and adaptation analyses. Its design distinguishes frozen source operating behavior from recalibration, fine-tuning, shortcut stress tests, and uncertainty assessment.

  • Model and evaluation: A frozen three-seed DenseNet121 ensemble combined full-radiograph and lung-conditioned views with learned gating and source-only MixStyle training.
  • Model and evaluation: Source temperature scaling and threshold selection produced T = 0.8313 and a primary operating threshold of 0.9999728.
  • Metrics: Primary external evaluation reported discrimination, calibration, and threshold-dependent operating metrics, including AUROC, AUPRC, Brier score, NLL, ECE, sensitivity, and specificity.
  • Adaptation: Secondary BDCXR recovery compared threshold adjustment, scaling, Platt recalibration, isotonic regression, linear probing, and progressively broader fine-tuning across target-label budgets.
  • Stability: Repeated 163- and 326-label recalibration fits, plus 651-label bootstrap resamples, assessed variability in calibrator performance and uncertainty.
  • Robustness: Shortcut stress tests compared full, lung-conditioned, background-only, and border-only views, while matched seed-42 models tested architectural dependence.

3. Results

External transport separated ranking from decision behavior: AUROC remained above chance, but the frozen source threshold failed across both target datasets. Limited-label Platt recalibration recovered sensitivity and calibration on BDCXR while imposing substantial alert and specificity burdens.

  • Internal performance and architecture robustness: 0.9759 internal-test AUROC accompanied 95.1% sensitivity and 90.9% specificity at the frozen source threshold.The three-seed ensemble achieved 0.9820 AUPRC and 93.0% balanced accuracy.
  • Internal performance and architecture robustness: Source-to-BDCXR AUROC fell from 0.9606 to 0.7492 for full-image DenseNet121, 0.9772 to 0.7664 for ungated dual-view, and 0.9655 to 0.7888 for gated MixStyle.The deterioration was therefore not specific to the gated model.
  • Zero-shot external transport: 6.2% BDCXR sensitivity at the frozen threshold contrasted with 99.9% specificity, a 4.5% alert rate, and 68.5 false negatives per 100 radiographs.BDCXR AUROC remained 0.7984, while calibration deteriorated with ECE 0.258.
  • Zero-shot external transport: 0% VinDr-PCXR sensitivity at the frozen threshold coexisted with AUROC 0.7421 and AUPRC 0.3960.None of the 170 pneumonia-family examinations crossed the frozen threshold.
  • Limited-label recovery: With 163 labels, Platt recalibration raised sensitivity to 88.3% and reduced ECE to 0.044, but specificity was 47.9% and 78.5% of radiographs generated positive alerts.At 326 labels, sensitivity was 91.9% with specificity 38.6% and ECE 0.029; at 651 labels, sensitivity was 90.7%, specificity 42.5%, and ECE 0.030.
  • Limited-label recovery: Across 200 repeated 163-label fits, mean sensitivity was 91.5% while mean specificity was 37.1%, with specificity ranging from 21.0% to 49.3%.Mean ECE was 0.028, and repeated sampling showed reproducible calibration recovery but sample-dependent operating points.
  • Operating-point robustness: A liberal source-only threshold reached 100% Guangzhou sensitivity but only 69.0% BDCXR and 15.9% VinDr-PCXR sensitivity.The unusual primary threshold contributed to the Bangladesh failure but did not explain the larger cross-dataset score shift.

4. Discussion

The study decomposes pediatric chest-X-ray transportability into discrimination, calibration, operating-point behavior, shortcut-associated signal, and limited-label recovery. Across Bangladesh and Vietnam, ranking often remained useful while source-defined operating behavior failed, and recalibration improved alignment without ensuring acceptable workload.

  • Cross-dataset shifts affected discrimination, calibration, and source-defined operating behavior differently.The protocol separates these components rather than treating AUROC or a single thresholded metric as complete deployment performance.
  • The same qualitative source-to-BDCXR degradation appeared for conventional full-image DenseNet121 and both dual-view variants.This architecture-robustness analysis makes the transport result less dependent on one implementation.
  • The untouched Vietnam cohort replicated the central pattern without target fitting.This preserved an external validation setting distinct from BDCXR analyses used for calibration or fine-tuning.
  • With 163 labels, Platt recalibration preserved AUROC and improved probability alignment, but repeated fits showed substantial specificity variability.The resulting 78.5% alert rate and false-positive burden show that apparent recovery does not equal clinical readiness.
  • Background-only and border-only images remained predictive on BDCXR, while public datasets lacked controlled metadata or interventions needed to identify a specific causal shortcut.These findings provide context for shortcut-associated signal rather than causal proof.

5. Conclusions

Across pediatric chest X-ray datasets from China, Bangladesh, and Vietnam, cross-dataset shift affected discrimination, calibration, and source-defined operating behavior differently. Limited target labels improved probability alignment but did not ensure a clinically acceptable operating policy, supporting separate evaluation of transport components and operational burden.

  • Cross-dataset shift affected discrimination, calibration, and source-defined operating behavior differently across China, Bangladesh, and Vietnam.
  • Limited target labels consistently improved probability alignment but did not ensure a clinically acceptable operating policy.
  • External medical-imaging AI studies should evaluate discrimination, calibration, operating-point transport, shortcut-associated signal, and recoverability separately.
  • Untouched external testing should remain distinct from later target adaptation when assessing transportability.

Ethics statement

This study analyzed de-identified research datasets without recruiting participants, performing interventions, or accessing identifiable information. Ethics, consent, and data-governance conditions were handled by the original dataset providers and publications.

  • The analysis used de-identified research datasets and recruited no new participants.
  • No intervention was performed and no identifiable information was accessed.
  • Original dataset providers and publications described ethics approvals, consent procedures, and data-governance conditions.
  • No additional participant consent was obtained for this secondary analysis.

CRediT authorship contribution statement

Nazim-E-Alam contributed across the study’s conceptualization, methodology, implementation, validation, analysis, data work, visualization, writing, and administration.

  • Nazim-E-Alam contributed to conceptualization, methodology, software, and validation.
  • The contribution included formal analysis, investigation, data curation, and visualization.
  • The contribution also included original drafting, review and editing, and project administration.

Funding

The research received no specific grant funding from public, commercial, or not-for-profit sectors.

  • The research did not receive any specific grant from funding agencies.
  • No specific grant came from public, commercial, or not-for-profit sectors.

Declaration of competing interest

The author declares no known competing financial interests or personal relationships that could have influenced the work reported in this paper.

  • No known competing financial interests are declared.
  • No known personal relationships are declared as potential influences on the work.
  • The declaration covers the work reported in this paper.

Data and code availability

Source datasets are available through their original repositories under applicable licenses and access conditions, while reproducibility workflows are supplied with the submission. Raw radiographs and restricted VinDr-PCXR DICOM data are not redistributed, and the author describes ChatGPT use for literature synthesis, debugging, structuring, and language refinement with human review and verification.

  • Source datasets are available from their original repositories subject to licenses and access conditions.
  • The reproducibility-code archive contains executable audit, adaptation, external-test, threshold-robustness, and repeated-recalibration workflows with integrity hashes and derived summaries.
  • Raw radiographs and restricted VinDr-PCXR DICOM data are not redistributed.
  • ChatGPT was used for literature synthesis, code debugging, manuscript structuring, and language refinement.
  • The author reviewed and edited the content, verified numerical claims against preserved analysis outputs, and retained responsibility for the article.
  • No generative-AI system was used to create or alter scientific image data.
Loading 2609.05140v1…