Source-linked AI summary

Auditable CT Phenotyping Through Report-derived Radiological Observations

Riga Wu, Walter Witschey, Yicheng Li, Felix Barajas Ordonez, Keno K. Bressem, Lisa C. Adams, Gary E. Weissman, Li Shen, Christos Davatzikos, Eduardo Barbosa, Daniel Truhn, Tianyu Han

arXiv:2608.25948v1cs.CV

TL;DR

CT foundation models can predict EHR phenotypes, but it remains unclear whether their probes use disease-specific radiological evidence or correlated shortcuts. This paper introduces ACT, which builds predictions from report-derived observations and audits and restricts the evidence used. ACT outperformed CT-CLIP across 221 phenotypes, while audits found nonspecific observations leading many probes and restriction redirected probes without reducing accuracy.

  • Problem

    Existing CT phenotype probes have not been tested for reliance on phenotype-related, named observations, leaving their evidentiary basis unclear.

  • Method

    ACT uses a report-derived observation bank as a common language for CT annotation, phenotype prediction, probe auditing and evidence restriction.

  • Results

    ACT exceeded CT-CLIP across 221 unseen-CTPA phenotypes, while only 97 observations ranked first across probes and one calcification phrase led 20 phenotypes.

  • Takeaways & Limitations

    Restricting probes to clinician-specified evidence preserved performance while redirecting them toward phenotype-related observations, separating predictor semantics from accuracy.

  • Takeaways & Limitations

    The audit is observational and global, so it does not show that a confounder was used or removed in any individual patient.

Abstract

from arXiv · show

Medical image foundation models can predict clinical phenotypes from computed tomography (CT), but strong performance leaves open whether they read disease-specific findings or shortcuts that correlate with the diagnosis. We tested this in 221 electronic-health-record (EHR) phenotypes using Auditable CT phenotyping (ACT), built on report-derived radiological observations. We trained ACT on 38,317 patients, mined 376,194 observations and evaluated it in 25,183 held-out patients. ACT exceeded five vision-language baselines on zero-shot annotation, and CT-CLIP across 221 phenotypes from unseen CT pulmonary angiography, both under zero-shot scoring (0.651 versus 0.572) and under linear probing (0.709 versus 0.662). Reading each probe exposes what accuracy conceals: only 97 observations occupy the 221 rank-1 positions, and one phrase describing aortic and coronary calcification ranks first for 20 phenotypes, including osteoporosis, urinary tract infection and major depressive disorder. Restricting the bank to clinician-specified evidence redirects those probes onto phenotype-related observations in 86 phenotypes at no accuracy cost (0.751 versus 0.741). Accurate CT-based EHR phenotyping can therefore rest on observations that are not valid evidence for the coded phenotype and that ACT can identify and intervene on.

Introduction

ACT uses report-derived radiological observations as a shared language linking CT annotation, EHR phenotyping, model auditing, and clinical restriction. The study tests this framework across annotation, organization, transfer to 221 CTPA-derived EHR phenotypes, and audit with clinical restriction.

  • Framework: ACT combines a native volume-report model with a bank of 376,194 atomic observations extracted from reports.The model was trained on 38,317 patients from CT-RATE and Merlin, spanning non-contrast chest and contrast-enhanced abdominal CT.
  • Motivation: Existing volumetric image-text models leave untested whether EHR phenotype probes depend on phenotype-related, named observations.ACT instead takes a single observation bank through the chain from CT to phenotyping and audit.
  • Study design: The study evaluates ACT through zero-shot abnormality annotation, anatomical and ontological organization, transfer to 221 CTPA-derived EHR phenotypes, and audit with clinical restriction.The evaluation uses 25,183 held-out patients, and CTPA was absent from image pretraining.
  • Organization: ACT’s fixed observation bank is anatomically organized, with Euclidean distance increasing as RadLex ontology separation increases.Observations from the external PMBB corpus occupied corresponding regions of the reference map.
  • Transfer: On CTPA, ACT’s concept-anchored CT representation exceeded CT-CLIP in macro AUROC across 221 phenotypes under both zero-shot scoring and matched linear probing.The comparison tests transfer to phenotypes from a CT domain absent from image pretraining.

Results

ACT outperformed CT-CLIP and other vision-language baselines across CT finding annotation, retrieval, and 221 EHR phenotypes. Probe inspection revealed widespread target-mismatched directions, while clinician-specified evidence restrictions redirected probes without reducing AUROC.

  • Finding annotation: 0.749 on CT-RATE, ACT attained the highest mean AUROC among five 2D and 3D vision-language baselines for zero-shot finding annotation.ACT’s mean AUROC was 0.689 on chest PMBB and 0.683 on abdominal PMBB.
  • Concept-conditioned retrieval: 46.9% on CT-RATE, ACT achieved the highest Recall@10 across all four cohorts, versus 43.0% for CT-CLIP and 32.8% for Merlin.Random-ranking floors were 1.6%, 7.8%, and 15.6% at Recall@1, @5, and @10, respectively.
  • EHR phenotyping: 0.651 versus 0.572 in zero-shot scoring and 0.709 versus 0.662 with linear probing, ACT’s macro AUROC exceeded CT-CLIP across 221 EHR phenotypes.The paired macro-AUROC differences were 0.079 in zero-shot scoring and 0.047 with linear probing.
  • Probe audit: 97 strings filled the 221 rank-1 positions, while only 389 distinct observations appeared across the union of all 221 top-5 lists.The phrase describing aortic and coronary calcification ranked first for 20 phenotypes, including osteoporosis and urinary tract infection.
  • Probe audit: 97 of 221 probes had vascular calcification or atherosclerosis as a lexical majority in their complete top-25 profiles, despite 71 probes achieving mean held-out AUROC of at least 0.75.Pulmonary air-space disease formed a lexical majority for 54 of 221 probes and 20 of the 71 high-performing probes.
  • Evidence restriction: 0.751 versus 0.741, rule-restricted probes had higher mean held-out AUROC than full-bank ACT references across 86 eligible phenotypes.Restricted-probe AUROC was higher in 55 phenotypes and lower in 31; examples placed largest weights on pleural effusion for CHF, a flat or collapsed IVC for hypovolemia, and active contrast extravasation for shock.

Discussion

ACT uses a report-derived observation vocabulary to connect CT annotation, phenotype prediction, model auditing and evidence restriction. Its audit shows that accurate phenotype probes can rely on nonspecific observations, while clinician-specified restriction preserves accuracy and redirects probes toward clinically associated findings.

  • Study framework: ACT carried one report-derived observation vocabulary through CT annotation, phenotype prediction, model audit and observation-bank restriction.The native volume-report model was trained on 38,317 patients and mined 376,194 observations.
  • Methodological contribution: ACT anchors each observation as a direction in a fixed text-embedding space, allowing the bank to span the full report vocabulary without concept annotation.This design avoids reserving a prelabelled latent coordinate for each concept.
  • Audit findings: 97 distinct strings occupied the leading positions of all 221 phenotype probes, and one aortic-and-coronary-calcification phrase led 20 probes.Those probes included osteoporosis, urinary tract infection and major depressive disorder, demonstrating that a shared top-ranked direction was not specific evidence for each diagnosis.
  • Limitations and future work: The alignment score measures global semantic correspondence because it omits scan-specific similarity and scan-dependent normalization, while correlated observation vectors make a leading direction represent a phrase neighborhood.Retaining scan-specific similarity is identified as a next experiment, alongside direct perturbation checks for counterfactuals.
  • Evidence restriction: Restricted probes matched the full-bank reference while using a median bank of under half a percent of observations.Retained directions included pleural effusion for CHF and a collapsed IVC for hypovolemia, separating a predictor’s semantic basis from its accuracy.
  • Limitations: The study’s external evidence came from one biobank, one public trauma dataset and one CTPA source, while PMBB outcomes were mined from reports and chest rules were calibrated against CT-RATE silver labels.The 221-phenotype criterion used test-set prevalence, and most multiplicity was left unadjusted, making those comparisons descriptive.

Methods

ACT was developed from paired chest and abdominal CT volumes and radiology reports, using impressions for contrastive pretraining and findings to construct an observation vocabulary. Its concept-anchored representation maps native image embeddings through 376,194 report-derived observations for downstream phenotype probes, with evaluation scans held out from pretraining.

  • Data and supervision: 72,640 paired scan-report rows from 38,317 patients combined CT-RATE non-contrast chest CT and Merlin contrast-enhanced abdominal CT.CT-RATE contributed 47,146 rows from 20,000 patients, while Merlin contributed 25,494 rows from 18,317 patients.
  • Data and supervision: The impression supported contrastive pretraining, whereas the findings section supplied organ-system observations for ACT’s vocabulary.Findings sections were processed with Qwen3.5-35B-A3B to extract descriptive observations as JSON lists after normalization and deduplication.
  • Data splits and labels: Pretraining excluded evaluation scans through predefined CT-RATE and Merlin partitions, while phenotype labels came from same-visit INSPECT CTPA diagnosis codes mapped to hierarchical PheWAS phenotypes.The data-split design reserved CT-RATE validation for checkpoint selection and kept all Merlin partitions outside evaluation.
  • Model and training: ACT encoded each volume with a DINOv2 ViT-B/14-based architecture and jointly optimized image, report, fusion, projection, and temperature parameters using symmetric CLIP InfoNCE.The volume representation fused independently encoded axial slices with a two-layer Transformer and projected the result into a 768-dimensional image embedding.
  • Concept-anchored representation: The concept-anchored representation weighted 5,120-dimensional observation embeddings by cosine similarities between the normalized CT embedding and all M = 376,194 bank observations.This fixed construction applies a linear map of rank at most 768 before normalization, and downstream phenotype probes act on the resulting embedding.
  • Evidence restriction: Clinician-specified evidence rules retained 86 non-empty phenotype observation sets and excluded 135 phenotypes with no matching observation.For each retained phenotype, the corresponding rows of the native and F2LLM observation-embedding matrices were recomputed.

Data availability

The study’s datasets are available through designated repositories and platforms, subject to repository access terms, registration, or completion of data-use agreements.

  • Data availability: CT-RATE is available from Hugging Face subject to the repository’s access terms.The passage provides the dataset repository link.
  • Data availability: The Merlin abdominal CT dataset is available through AIMI after completion of its data-use agreement.The passage provides the AIMI dataset link.
  • Data availability: RSNA-2023 is available to registered users through its corresponding abdominal trauma detection competition.The passage identifies the competition platform as Kaggle.

Code availability

ACT code is publicly available and includes the full pipeline for CT preprocessing, model training, representation construction, phenotype probing, auditing, restriction, analysis, and figure generation.

  • Code availability: ACT code is publicly available at https://github.com/peterhan91/ACT and covers the complete analysis pipeline.The release includes native CT preprocessing, volume-report model training and inference, observation-bank extraction and embedding, concept-anchored CT representation construction, phenotype probing, probe-observation auditing, observation-bank restriction, statistical analysis, and figure generation.
  • Code availability: The release also provides a versioned software environment, prompt files, and phenotype-rule manifests.

ACT Supplementary Information … A.3 Selection rules for supplementary audit displays

The supplementary methods describe ACT’s native 3D volume-report architecture, uncertainty estimation for AUROC analyses, and selection rules for auditable display panels. Displays emphasize shared rank-1 observations, clinician-reviewed relevance, and paired full-bank versus refined audits.

  • A.1 Architecture of ACT’s native 3D volume-report model: ACT resizes each CT volume to 160 × 224 × 224 slices, height, and width, embeds slices with DINOv2 ViT-B, and aggregates them with a Transformer.The Transformer fuses a learnable classification token with slice embeddings into the volume representation.
  • A.2 Uncertainty estimation for supplementary AUROC tables: Zero-shot finding AUROCs use empirical full-cohort AUROCs for eligible cohort-finding-model combinations with complete row-aligned predictions.95% confidence intervals use 1,000 patient-clustered bootstrap resamples, reusing draws across eligible models and findings within each cohort.
  • A.2 Uncertainty estimation for supplementary AUROC tables: RSNA-2023 evaluation reconstructed the recorded 70/10/20 seed-42 patient split and lexicographic series ordering, yielding 943 series in 629 reconstructed patient clusters.The reconstruction was paired with the retained 943 × 9 outcome matrix and checked against retained evaluation ordering.
  • A.2 Uncertainty estimation for supplementary AUROC tables: CTPA phenotype-probe AUROCs are mean test AUROCs over 20 matched probe fits, with two-sided Student t confidence intervals across fits.ACT-versus-CT-CLIP paired differences use shared patient-clustered bootstrap resamples for the overall comparison.
  • A.2 Uncertainty estimation for supplementary AUROC tables: Observation-bank restriction AUROCs recompute one AUROC per probe within 1,000 shared patient-clustered resamples of 2,612 held-out scans from 2,223 patients.Models remain fixed within resamples, and reported intervals use the 2.5th and 97.5th percentiles.
  • A.3 Selection rules for supplementary audit displays: The shared-direction panels count distinct rank-1 observation strings across 221 phenotypes; the most common covers 20 phenotypes and the next covers 18.Because both describe atherosclerotic vascular calcification, the second panel instead displays biliary ductal dilatation, covering 15 phenotypes.
  • A.3 Selection rules for supplementary audit displays: Display-relevance review classifies selected observations as direct target evidence or clinically related but non-defining or qualified context.Uncertain or temporal phrases may remain rejected while appearing in light blue when clinically related to the phenotype.
  • A.3 Selection rules for supplementary audit displays: Paired audits compare full-bank natural mean ranks 1 to 5 with five highest bootstrap-mean-weight observations from restricted probes, using eight higher- and lower-AUROC examples.The eight examples come from 86 eligible phenotypes and exclude the eight already shown in Figure 6; displayed top-five profiles retain their raw form.

B Extended Results … PMBB query

The extended results report zero-shot finding annotation and cross-modal retrieval across multiple CT datasets, alongside analyses of ACT’s observation-bank composition and cross-corpus nearest neighbours. PMBB queries are represented through affirmative and explicitly negated observations compared across corpora.

  • B.1 Native-model annotation and retrieval: Supplementary Tables 1–2 report zero-shot finding annotation on the CT-RATE test set using AUROC with 95% percentile confidence intervals.The estimates come from 1,000 shared patient-cluster bootstrap resamples, with n+ denoting positive test scans.
  • B.1 Native-model annotation and retrieval: Supplementary Tables 3–8 report zero-shot finding annotation on PMBB chest, PMBB abdomen, and RSNA-2023 abdominal-trauma datasets.The tables divide results across two model panels and use AUROC with 95% percentile confidence intervals from 1,000 shared patient-cluster bootstrap resamples.
  • B.1 Native-model annotation and retrieval: Supplementary Figure 2 evaluates ACT’s native 3D model and five baselines on cross-modal concept retrieval across CT-RATE, chest PMBB, abdominal PMBB, and RSNA-2023.Concept-to-image retrieval uses Recall@1/5/10 in 64-candidate pools, while image-to-concept retrieval uses Recall@1/3/5 after per-finding z-scoring.
  • B.2 Observation-bank composition and cross-corpus neighbours: Supplementary Figure 3 qualitatively summarizes the lexical content of ACT’s report-derived observation bank through word clouds covering its 16 largest keyword-assigned categories.Category counts appear in parentheses, and word size reflects frequency among distinct observation strings.
  • B.2 Observation-bank composition and cross-corpus neighbours: The observation-bank analysis includes category counts and lexical summaries, with gastrointestinal abbreviated as GI and word-cloud colours described as decorative.The figure focuses on the lexical composition of report-derived observations rather than quantitative model performance.
  • PMBB query: Supplementary Table 9 compares nearest affirmative and explicitly negated observations for eight representative PMBB queries spanning four chest and four abdominal observations.The table reports three same-finding neighbours from CT-RATE and Merlin ranked by Euclidean distance between ℓ2-normalized 5,120-dimensional F2LLM observation vectors.
  • PMBB query: The PMBB nearest-neighbour analysis distinguishes affirmative observations (+NN) from explicit textual negations (−NN).The negations are textual and are not independently verified findings.

hepatic mass

The section highlights cystic pancreatic lesions as the principal reported observations, including lesions in the pancreatic tail and head and involvement of the pancreatic duct. Reported lesion sizes range from 24 mm to 5.2 x 4.6 cm, with associated d values from 0.730 to 0.877.

  • hepatic mass: Cystic pancreatic lesions are identified in the abdomen.
  • hepatic mass: 2 cystic lesion with an ap diameter of 28 mm in the tail of the pancreas has d = 0.870.
  • hepatic mass: 3 cystic lesion with 24 mm diameter at the pancreatic head has d = 0.877, while larger cystic lesions measure 5.2 x 4.6 cm and 3.9 x 3.0 cm with d = 0.730.
  • hepatic mass: Maintained pancreatic duct involvement of cystic lesions has d = 0.735.

B.3 Complete CTPA-EHR phenotype performance

Supplementary Table 10 reports matched linear-probe performance across all 221 INSPECT phenotypes, using mean AUROC and fit-variation confidence intervals from 20 matched probe fits. It also provides phenotype-positive counts among 2,612 test scans and identifies the larger model point estimate per phenotype.

  • Performance evaluation: 221 phenotypes are evaluated with matched linear probes using mean AUROC across 20 fits.The table reports two-sided 95% t-based confidence intervals around these estimates.
  • Performance evaluation: 2,612 test scans define the evaluation cohort, with n+ reporting each phenotype’s positive count and remaining scans treated as negative.The reported intervals quantify fit variation rather than patient-sampling uncertainty.
  • Performance evaluation: Boldface marks the model with the larger point estimate for each phenotype.This comparison is based on the model point estimates, not the confidence intervals.

B.4 Complete observation-bank restriction performance · B.5 Probe-observation audit: discordance and nonspecific context

Across 86 restriction-eligible INSPECT phenotypes, ACT performance was evaluated with full-bank and rule-restricted probes, while audits examined shared top-ranked observations and distinguished direct target evidence from clinically related context. Paired audits further compared full-bank with refined probes for phenotypes showing higher or lower refined AUROC.

  • B.4 Complete observation-bank restriction performance: 86 restriction-eligible INSPECT phenotypes were evaluated with full-bank and rule-restricted ACT probes using held-out scans and clustered bootstrap confidence intervals.The evaluation used 2,612 scans from 2,223 patients and 1,000 shared patient-cluster bootstrap resamples.
  • B.4 Complete observation-bank restriction performance: The restriction analysis recomputed AUROC for fixed full-bank and rule-restricted probes within each bootstrap resample and summarized one value per column.All scans from each sampled patient received the same multiplicity during resampling.
  • B.5 Probe-observation audit: discordance and nonspecific context: 20 phenotypes shared the top-ranked observation “calcific atherosclerotic changes in the thoracic aorta and coronary arteries,” despite lacking a common clinical target.The figure orders these phenotypes by mean held-out test AUROC and displays their observation ranks one through five.
  • B.5 Probe-observation audit: discordance and nonspecific context: Full-bank profiles separated direct target evidence from clinically related but non-defining or qualified context across bronchiectasis, pneumonia, coronary atherosclerosis, and pulmonary collapse with emphysema.Each phenotype panel shows its five highest-ranked observations, with colour distinguishing direct evidence from related context.
  • B.5 Probe-observation audit: discordance and nonspecific context: Paired audits compared full-bank and refined top-five observations for cardiomegaly, osteoporosis NOS, other aneurysm, and acute pulmonary embolism or infarction among phenotypes with higher refined AUROC.The panels present the full-bank audit first and the refined audit second, with mean ±1 s.d. and individual estimates.
  • B.5 Probe-observation audit: discordance and nonspecific context: Paired audits likewise compared full-bank and refined top-five observations for diseases of pancreas, cancer of bronchus or lung, osteoarthrosis, and diaphragmatic hernia among phenotypes with lower refined AUROC.The panels use the same prior-to-later audit layout, reporting mean ±1 s.d. and individual estimates.
Loading 2608.25948v1…