Source-linked AI summary
A radiographic world model for clinical reasoning and evidence generation
Suyang Xi, Songtao Hu, Shansong Wang, Mojtaba Safari, Luke del Balzo, Ehsan Ul Karim, Mingzhe Hu, Kuo Zhang, Tonghe Wang, Ralph R. Weichselbaum, Xiaofeng Yang
TL;DR
Medical imaging AI commonly separates diagnostic representation learning from image generation, despite both depending on the same underlying radiographic state. MedDream jointly learns these capabilities through a shared radiographic representation and shows broad diagnostic transfer alongside clinically grounded evidence generation and targeted synthetic augmentation.
Problem
Rare findings, complex anatomy, and underrepresented patient groups limit the clinical examples available for diagnostic learning and evidence construction.
Method
MedDream jointly optimizes image–text alignment and masked latent generation through a shared visual pathway, with progressive masking providing complementary diagnostic and generative supervision.
Results
Across held-out and pretraining-excluded evaluations, MedDream retained disease, severity, and spatial information while supporting report-conditioned generation with stronger pathology consistency and downstream utility.
Takeaways & Limitations
The findings support radiographic world models as a framework for interpreting, simulating, and constructing clinically grounded radiographic evidence.
Takeaways & Limitations
The study is restricted to examination-level chest radiography, and generated images were not exhaustively assessed for all conditioned findings, localization, or unsupported abnormalities.
Abstract
from arXiv · showhide
Medical imaging artificial intelligence (AI) is commonly developed as separate mappings from radiographs to diagnostic outputs or from clinical descriptions to generated images, although both arise from the same underlying radiographic state. A world-model formulation instead seeks to learn an internal representation of this state that can support both clinical readout and conditional simulation of radiographic observations. Here we introduce MedDream, a radiographic world model that learns a shared continuous latent state from paired chest radiograph-text observations for diagnostic reasoning and report-conditioned evidence generation. MedDream was pretrained on 2.65 million leakage-controlled chest radiograph-text pairs curated from 4.40 million candidates. Across eight clinical datasets and two independent reader cohorts, MedDream outperformed leading diagnostic and generative comparators. For diagnostic reasoning, MedDream showed strong generalization across disease recognition, label-scarce adaptation, severity assessment, and localization, while MedDream-supported review increased mean resident concordance with independent radiologist consensus from 56.3% to 63.0%. For evidence generation, MedDream produced radiographs that preserved clinically relevant pathology and improved downstream performance on held-out real data, with synthetic augmentation increasing external VinDr-CXR macro-AUROC from 76.4% to 81.4%. More importantly, conditioning generation on prespecified subgroup performance gaps enabled targeted evidence construction, increasing weighted F1 by 3.1 percentage points in Asian patients, whereas matched-volume unguided augmentation decreased it by 2.3 points. These findings establish radiographic world models as a path toward medical AI that learns clinically meaningful internal states for interpreting, simulating, and constructing evidence for clinical use.
Introduction
Rare and atypical radiographic findings remain difficult to interpret because comparable examples are scarce, while diagnostic representation learning and image generation have largely developed separately. MedDream addresses this gap by jointly learning diagnostic and report-conditioned generative functions through a shared visual pathway.
- Motivation: Rare or atypical findings are difficult to interpret because comparable clinical examples are scarce and anatomical variation can confound disease-related changes.
- Motivation: Long-tailed imaging cohorts underrepresent rare phenotypes and complex anatomy, allowing models to rely on unstable cues in sparsely represented settings.
- Research gap: Diagnostic representation learning and medical image generation have largely developed as separate capabilities.
- Research gap: Diagnostic and generative pretraining impose competing masking demands: mild masking preserves clinical cues, whereas stronger masking improves recovery of anatomy and disease-specific appearance.
- MedDream: MedDream jointly optimizes image–text alignment and masked latent generation over shared continuous visual tokens, using progressive masking to support both diagnostic transfer and image synthesis.
- Evaluation: The study evaluates MedDream on chest radiography using approximately 4.4 million candidate pairs and complementary diagnostic-readout and evidence-generation tracks.
Results
MedDream transferred diagnostic representations across held-out diseases, label-scarce settings, severity assessment, localization, and reader-supported review. Its report-conditioned synthetic radiographs also improved real-data performance, annotation efficiency, quantitative prediction, subgroup matching, and targeted subgroup outcomes.
- Diagnostic transfer: 81.2% mean AUROC across 14 NIH ChestX-ray14 findings exceeded Ark+ at 80.3% and MedCLIP at 78.3%.The evaluation used frozen-encoder linear probing on an excluded hold-out test set.
- Diagnostic transfer: 86.3% mean AUROC across 23 VinDr-CXR endpoints exceeded Ark+ at 85.2% despite complete pretraining exclusion.Performance remained strong for lower-prevalence findings, including interstitial lung disease at 85.1% AUROC.
- Diagnostic transfer: At 5% labeled ChestDR data, MedDream achieved 64.8% mean AUROC, supporting transfer under long-tailed label scarcity.The evaluation varied labeled training data from 5% to 100%.
- Clinical interpretation: MedDream-supported review increased concordance with radiologist consensus across all three residents while reducing each reader’s mean absolute ordinal error.Quadratic-weighted kappa increased for every reader, from 0.243 to 0.482, 0.868 to 0.935, and 0.799 to 0.852.
- Synthetic evidence: External VinDr-CXR macro-AUROC increased from 76.4% with real-only training to 81.4% with 2× MedDream augmentation.Synthetic radiographs inherited weak labels from paired real training studies and produced dose-dependent gains.
- Synthetic evidence: MedDream synthetic pretraining improved AUROC by a median 2.4 percentage points across class–budget combinations, with larger gains under limited real data.At 5% and 10% real-data budgets, gains were 4.0 and 4.9 percentage points, respectively.
- Subgroup evidence: Metadata-guided augmentation increased weighted F1 by 3.1 percentage points in Asian patients, whereas unguided augmentation decreased it by 2.3 points.Matched real–synthetic subgroup pairs also had lower median Wasserstein distance than unmatched pairs: 0.051 versus 0.141.
Discussion
MedDream jointly learns diagnostic representations and report-conditioned generation through a shared radiographic state, supporting interpretation and evidence construction when real examples are scarce. Its utility extends beyond visual plausibility, but limitations include scope restricted to chest radiography, incomplete generation assessment, possible residual data overlap, and retrospective single-system reader studies.
- Shared radiographic state: MedDream jointly optimizes image–text alignment and masked latent generation through a shared visual pathway, coupling diagnostic readout with report-conditioned evidence generation.The shared pathway is intended to use available clinical evidence while constructing plausible radiographic evidence where comparable observations are scarce.
- Optimization strategy: Joint optimization exceeded the alignment-only specialist by 2.1–3.5 percentage points across five diagnostic and localization benchmarks under matched budgets.It also improved on the generation-only specialist in distributional fidelity, pathology-profile correlation, and downstream external utility.
- Complementary clinical structure: The shared representation retained disease semantics, disease burden, spatial structure, report-conditioned generation, and alignment-guided trajectory selection across held-out and pretraining-excluded evaluations.These findings indicate that joint training preserved discriminative clinical information while enabling clinically grounded generation.
- Evidence construction: Synthetic evidence improved internal and cross-dataset performance when real annotations were limited, but its utility depended on how evidence was selected and allocated.Metadata-guided selection improved prespecified subgroup–pathology settings and reduced selected abnormal-finding omissions on held-out real cases compared with unguided augmentation.
- Limitations: The study evaluates examination-level chest-radiographic state rather than longitudinal patient-state dynamics, and generated radiographs were not exhaustively assessed for all conditioned findings, localization, or unsupported abnormalities.Reader-study practice effects were not fully excluded, and the composite workflow did not isolate the contributions of its individual components.
- Limitations: Residual overlap between pretraining and evaluation data cannot be fully excluded, while human reader studies were retrospective and conducted within a single health system.The literature-mined corpus retained an estimated 10.7% residual rate of non-compliant pairs, and perceptual hashing may miss transformed variants.
Methods
The methods assemble a large, leakage-controlled chest-radiography image–text resource from public clinical repositories and literature-mined captions, then screen, deduplicate, normalize, and audit the retained pairs. The corpus is broad but imperfectly curated, with residual non-compliant pairs and incomplete protection against transformed duplicates.
- Resource construction: Approximately 4.40 million image–text pairs combined 3,475,100 PMC-CXR image-caption pairs with approximately 0.93 million public clinical chest-radiography pairs.A leakage-controlled subset was derived for MedDream pretraining.
- Clinical repositories: MIMIC-CXR contributed 377,110 chest radiographs linked to 227,827 free-text reports, using findings or impression sections as paired text.Only the official training split was used for pretraining.
- Clinical repositories: CheXpert Plus, PadChest, BRAX, and IU X-Ray supplied additional report- or label-paired chest radiographs and were screened for cross-source near-duplicates against evaluation cohorts.These datasets were not used in downstream evaluation and were included in full after screening.
- Leakage control: NIH ChestX-ray14 contributed only its official training and validation partitions totaling 86,524 images, while the 25,596-image test set was excluded from pretraining.Training labels were converted into natural-language radiological descriptions using rule-based templates.
- Literature mining: PMC-CXR candidates were retrieved from PubMed Central using chest-radiography terms and radiology-focused MeSH terms, then filtered by image and metadata criteria.The corpus broadened visual and linguistic pretraining coverage through literature-mined figure-caption pairs.
- Quality audit: A 2,000-pair manual audit assessed CXR precision, caption–image alignment, and absence of non-compliant content across BiomedCLIP similarity quartiles.This audit was part of validating the retained corpus.
- Corpus limitations: PMC-CXR comprised approximately 1.80 million image-caption pairs after exclusions but retained an estimated 10.7% residual rate of non-compliant pairs.Perceptual hashing may miss transformed variants, so the resource is characterized as large-scale but imperfectly curated.
- Preprocessing and deduplication: Perceptual hashing used a Hamming-distance threshold of 8 for deduplication, with patient-level exclusions where identifiers were available and study- or image-level exclusions otherwise.Images were resized to 256×256 px and normalized to zero mean and unit variance.
Datasets for diagnostic evaluation
Diagnostic evaluation spans in-domain recognition, broader and fine-grained labels, long-tailed scarcity, few-shot rare findings, localization, generation-related cohorts, severity regression, and demographic subgroup rebalancing. The datasets use explicit held-out partitions and patient- or study-level separation to limit leakage across evaluation settings.
- Recognition: NIH ChestX-ray14 evaluation used 86,524 training and 25,596 test images across 14 pathological categories, with the training partition overlapping pretraining and the test partition excluded.The official benchmark label space was retained without prevalence-based disease categorization.
- Recognition: VinDr-CXR contributed 18,000 radiographs with 22 local and 6 global labels; evaluation used 15,000 training and 3,000 test images across 23 thoracic endpoints.The dataset was excluded from pretraining, and multi-annotator labels were aggregated by majority vote.
- Localization: RSNA Pneumonia provided 30,000 frontal radiographs with lung-opacity and pneumonia labels and was excluded from pretraining.It supplied a larger single-disease localization setting.
- Label scarcity: ChestDR contains 4,848 images with 19 long-tailed thoracic disease labels, including linear-probe settings using 5%, 10%, 25%, 50%, and 100% of training data.The 5% setting contained approximately 49 total training images shared across all 19 labels.
- Few-shot rare findings: CXR-LT evaluated four rare findings with a two-way k-shot protocol using k positive and k negative support examples, where k ranged from 1 to 5.The findings were hydropneumothorax, round atelectasis, infarction, and bulla.
- Localization: MS-CXR contains 1,162 image-sentence pairs across eight cardiopulmonary findings, with training used for localization heads and validation used solely for segmentation-threshold selection.The test annotations were excluded from training, model selection, threshold selection, and trajectory selection.
- Generation and augmentation: Generation and synthetic-augmentation experiments used separate MIMIC-CXR p19 cohorts, with 3,500 studies for impression-conditioned generation and 18,634 studies for synthetic augmentation and pretraining.These cohorts were separated at the patient level, and synthetic images were conditioned only on training-set impressions.
- Severity and subgroup evaluation: Severity regression used RALO with 1,898 training and 475 test images scored by radiologists, while demographic rebalancing used a separate p19 cohort of 11,273 training and 2,684 test images.RALO is independent of MIMIC-CXR, and demographic experiments used structured metadata for targeted sampling and subgroup evaluation.
Datasets for clinical reader studies
MedDream uses a shared latent radiographic representation to support clinical readout and report-conditioned image generation. Its design combines clinical alignment, masked latent diffusion, and internally selected generation trajectories.
- MedDream architecture: MedDream learns diagnostic reasoning and report-conditioned image generation over shared continuous visual tokens rather than separate model pathways.The same pretrained visual pathway supports both functions.
- Latent representation: A pretrained VAE converts images into spatial continuous latent tokens, while an encoder derives the shared radiographic representation from visible tokens and buffer tokens.Clinical text is excluded from the image encoder, and the resulting representation is reused for clinical readout and conditional generation.
- Clinical alignment: Clinical image–text alignment organizes the visual representation around disease findings, anatomical locations, and imaging appearances.Image and text projections are trained with symmetric contrastive objectives.
- Latent generation: Masked latent generation reconstructs missing continuous latent values by predicting injected noise under clinical text conditioning.The diffusion head operates within a masked latent generation framework rather than directly generating pixels.
- Training strategy: A progressive masking curriculum reconciles semantic alignment, which benefits from lower masking, with generation, which benefits from recovering latent regions from limited context.The joint objective gates alignment and diffusion losses across masking regimes.
- Inference strategy: Semantic trajectory selection ranks partially recovered latent candidates in the shared representation space before completing and decoding the selected trajectory.The default procedure uses 368 recovery steps versus 256 for single-trajectory generation, a 43.8% increase.
General evaluation protocol
MedDream’s evaluations used leakage-controlled held-out data and frozen-encoder protocols to isolate pretrained representation quality. Diagnostic, generation, and synthetic-data experiments each assessed performance on real held-out evidence.
- Evaluation safeguards: Held-out test data excluded from corresponding pretraining, fine-tuning, and synthetic-generation cohorts.Benchmark data used for pretraining were restricted to designated development partitions, with evaluation confined to held-out test partitions.
- Diagnostic evaluation: Frozen encoders isolated diagnostic information in pretrained representations from downstream task-specific model capacity.Only lightweight classifiers, regression heads, or segmentation heads were trained unless otherwise specified.
- Metrics: Classification used AUROC, average precision, F1 score, and Matthews correlation coefficient, with macro scores averaged across disease labels.Severity tasks additionally used accuracy, macro F1, weighted F1, and macro AUROC; segmentation used intersection-over-union and Dice score.
- Synthetic evidence evaluation: Synthetic images were evaluated for fidelity, diversity, pathology consistency, plausibility, and downstream utility on held-out real test sets.Subgroup analyses were performed on held-out real data.
Frozen-encoder diagnostic representation transfer
Diagnostic representation transfer used frozen image encoders and lightweight probes across disease recognition, label efficiency, severity assessment, and lesion localization. The protocol covered held-out benchmarks with task-specific splits and metrics.
- Protocol: Frozen image features supported linear probes across classification, severity, and segmentation tasks.Features were standardized using training-split statistics before downstream fitting and evaluation.
- Disease recognition: NIH ChestX-ray14 evaluated common thoracic disease recognition using 86,524 development images and 25,596 held-out test images.The development partition contributed to pretraining, while the official test partition was reserved for final evaluation.
- Label efficiency: ChestDR measured label efficiency at 5%, 10%, 25%, 50%, and 100% of training data using macro AUROC.The dataset contained 19 disease labels with a long-tailed class distribution and repeated random trials when feasible.
- Severity assessment: RALO assessed four-category radiographic burden using accuracy, macro F1, weighted F1, and macro AUROC.Radiologist opacity scores were mapped to none or trace, mild, moderate, and severe categories with subject-level splits.
- Localization: MS-CXR and RSNA Pneumonia evaluated lesion-level grounding with abnormality segmentation and intersection-over-union and Dice metrics.MS-CXR used eight abnormality classes, while RSNA data used patient-level 70%/10%/20% splits.
Impression-conditioned image generation
Impression-conditioned generation was evaluated on independent MIMIC-CXR test radiographs, with image quality assessed across fidelity, diversity, pathology consistency, and expert-rated plausibility. Real and synthetic feature distributions were also compared in a shared embedding space.
- Evaluation set: MIMIC-CXR p19 evaluation used 3,500 independent frontal radiographs with prompts derived from report impression or conclusion sections.Reference images were excluded from prompt preparation and generation to prevent prompt–reference leakage.
- Generation setup: Images were generated at 512 × 512 resolution using 256 autoregressive iterations, 100 diffusion steps, and classifier-free guidance scale 3.0.The final checkpoint was refined at the same resolution used for comparison with ChexGen and MINIM.
- Quality assessment: Generation quality combined CLIP-FID and XRV-FID, intra-prompt structural similarity, pathology-profile correlations, and blinded expert ratings.Four experts rated realism, acquisition plausibility, and prompt-level consistency on five-point Likert scales.
- Feature-space comparison: Real and synthetic classifier-derived embeddings were projected with t-SNE, with density overlap quantified by the overlapping coefficient.The coefficient used the summed pointwise minimum of normalized real and synthetic densities on a shared grid.
Synthetic-data augmentation and synthetic pretraining
Synthetic-data experiments tested whether generated radiographs improved classifiers and quantitative prediction when paired with real training data. Comparisons used matched labels, controlled training procedures, and multiple augmentation or real-data fractions.
- Synthetic augmentation: Synthetic counterparts were generated from the same impression text as real MIMIC-CXR training images and appended with their paired weak labels.The real corpus contained 18,634 deduplicated frontal studies split at the patient level.
- Training controls: Real-only baselines used matched total gradient updates to control for larger augmented training sets.DenseNet-121 classifiers were trained with AdamW, learning rate 1 × 10^-4, weight decay 1 × 10^-4, and 40-epoch cosine annealing.
- Synthetic pretraining: Synthetic pretraining initialized classifiers with model-specific generated images before fine-tuning on matched real-data fractions.Experiments used 5%, 10%, 25%, 50%, or 100% of available real studies with identical ImageNet-1k initialization.
- Quantitative prediction: RALO quantitative prediction used 1,898 real training images, a fixed 475-image test set, and 1×, 2×, or 5× synthetic augmentation scales.Performance was measured with mean absolute error, root mean squared error, and Pearson correlation coefficient.
Metadata-guided subgroup evaluation and synthetic rebalancing
The study evaluated whether metadata-guided synthetic augmentation could align subgroup pathology distributions and improve performance on held-out real images. Demographic groups were defined from structured metadata, and rebalancing was compared with real-only and unguided augmentation.
- Subgroup construction: Patients were stratified into 22 single-axis and intersectional gender, race and age groups for subgroup analysis.Race included White, Black or African American, and Asian categories; age groups were 19–40, 41–65 and 66+ years.
- Distributional evaluation: Subgroup distributional alignment used averaged one-dimensional Wasserstein distances across nine-dimensional pathology-probability vectors, with lower values indicating closer matching.A fixed DenseNet-121 classifier estimated probabilities for nine pathologies from real and synthetic images.
- Evaluation design: The rebalancing experiment trained on 11,273 real MIMIC-CXR p19 images and tested on 2,684 held-out real images with 14 CheXpert-style pathology labels.Splits were defined at the patient level to prevent patient overlap between training and testing.
- Synthetic rebalancing: Parity-targeted augmentation added synthetic samples according to each group’s deficit relative to the largest real training group, scaled by the augmentation factor.At 1×, the procedure equalized real-plus-synthetic sample counts across the selected demographic axis; sampling with replacement was used when necessary.
- Outcome measures: Subgroup performance was measured with weighted F1, while false-negative-rate comparisons required at least 10 positive test cases per subgroup–pathology pair.Subgroups with fewer than 30 test samples were excluded from weighted-F1 analysis, and changes were reported relative to real-only baselines.
Controlled optimization-strategy ablations.
The ablation study controlled data, optimization, evaluation, and generation settings to isolate the contribution of optimization strategy. It compared single-objective, sequential, and joint configurations under matched protocols and budgets.
- Controlled training: All single-model ablations used the same 266,000 image–text pairs, architecture, initialization, optimizer, schedule, batch size and optimization budget.The subset represented 10% of the leakage-controlled pretraining corpus.
- Evaluation controls: Diagnostic and localization comparisons used identical frozen-encoder protocols and downstream evaluation pipelines across configurations.Generation-capable models also shared report-derived prompts, sampling seeds, guidance scale, temperature, autoregressive iterations and diffusion steps.
- Statistical analysis: Confidence intervals generally used bootstrap resampling at the patient, image or study level according to available identifiers.Intervals were reported for classification, localization, prediction, synthetic-augmentation and subgroup-performance endpoints.
- Statistical analysis: Matched held-out model comparisons used two-sided paired bootstrap tests, with Benjamini–Hochberg adjustment across labels, models, subgroups and subgroup–pathology pairs.The correction families varied by experiment, including diagnostic transfer, severity classification, localization and synthetic augmentation.
- Task-specific tests: Generation metrics used Wilcoxon signed-rank tests for per-case MS-SSIM and pathology-profile correlation, while severity-reader changes used McNemar tests.FID was treated as a cohort-level distributional metric rather than a per-case comparison.
- Repeated subsampling: Few-shot results averaged across 20 random support-set samples, whereas ChestDR label-fraction experiments used up to five random trials when sufficient positives were available.Rare-label settings with inadequate support were reported using available runs only.
- Data-budget controls: Synthetic augmentation and synthetic pretraining were compared on fixed held-out cohorts under matched data budgets, including an external VinDr-CXR test set.Synthetic-pretraining gains were measured against ImageNet initialization at the same real-data budget.
Implementation details
MedDream is a large continuous-token masked autoregressive model for latent chest-radiograph generation, using a frozen CXR-specific medical VAE to encode and decode images.
- Model architecture: MedDream contains approximately 636 million trainable parameters, including a 201.5 million-parameter encoder.Its latent image-modelling design follows prior continuous-token autoregressive image-generation work.
- Latent image modelling: A frozen CXR-specific medical VAE encodes chest radiographs into latent representations and decodes generated latents back into image space.The VAE uses the MedVAE architecture.
Supplementary Note
The supplementary materials collect diagnostic, generation, subgroup, baseline, training, and reader-study specifications. They also report evaluation conventions and the model’s shared diagnostic-generation setup.
- Supplementary tables: Tables S1–S3 cover diagnostic transfer and localization, generation and synthetic-data utility, and subgroup rebalancing statistics, respectively.Their captions identify the principal evaluation domains summarized in the supplementary material.
- Representation baselines: Table S4 lists frozen diagnostic-encoder baselines and indicates whether each model supports diagnostic transfer and report-conditioned radiograph synthesis.It also records training-corpus information where available.
- Generative baselines: Table S5 compares generative medical-image baselines by diagnostic-representation support, image-generation support, training corpora and conditioning inputs.MedDream is described as providing both capabilities through a jointly optimized shared visual pathway.
- Training overview: MedDream pretraining used approximately 2.65 million image–text pairs from six public chest-radiograph datasets and PMC-CXR, followed by high-resolution refinement on MIMIC-CXR patient groups p10–p18.The MIMIC-CXR p19 subset was excluded from pretraining.
- Diagnostic adaptation: Extended Data Fig. 2 compares frozen-encoder ROC and precision–recall curves for selected VinDr-CXR labels against MedCLIP and AFLoc.The evaluation targets fine-grained diagnostic transfer.
- Severity evaluation: Extended Data Fig. 3 reports RALO severity classification by multiple metrics, severity-wise F1, and quantitative-prediction scaling with 1×, 2× and 5× synthetic augmentation.Comparators include MedDream, BiomedCLIP and AFLoc.
- Reader study: Extended Data Fig. 4 describes a controlled reader study in which three readers reassessed 100 held-out frontal radiographs with model-estimated severity, confidence, and severity-matched synthetic references.Independent radiologist consensus served as the reference.
- Omission adjudication: Extended Data Fig. 5 uses blinded two-phase adjudication to compare abnormal-finding omissions between baseline and rebalanced models.A false negative was a reviewer-confirmed finding reported as absent; for “no finding,” it represented a false-positive alert.