Source-linked AI summary
ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs
Sebastián Andrés Cajas Ordóñez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
TL;DR
Medical VLMs can answer from radiology reports, making it difficult to determine whether they use the image. ModaLens uses paired image swaps with fixed reports and questions to measure report-conditioned image sensitivity, finding fewer answer changes when the report is available under this protocol. The study cautions that report-derived labels and readout choices constrain conclusions about visual correctness.
Problem
It is difficult to tell whether a medical VLM uses the image when a radiology report already supplies an answer.
Method
ModaLens swaps each case image while keeping the report and question fixed, comparing paired output changes across MIMIC-CXR trials.
Results
20.94% of trials changed without the report versus 4.26% with it, a paired difference of 16.7 points [15.6, 17.7] under an explicit answer instruction.
Takeaways & Limitations
Under this protocol, report availability reduces measured image-swap sensitivity, and the direction persists across evaluated controls and model lineages.
Takeaways & Limitations
Report-derived labels and readout limitations constrain conclusions about visual correctness and can miss shifts that do not cross the binary decision boundary.
Abstract
from arXiv · showhide
A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a paired image-swap audit, measures how report availability changes image sensitivity: MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, all 14 questions per case (13 finding-specific and one composite), each image replaced by one from another study, usually of the same patient, with question and report fixed. Under an explicit answer instruction, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, a paired increase of 16.7 points (patient-clustered 95 percent CI 15.6 to 17.7), so report availability reduces image-swap sensitivity under this protocol; the original prompt with a lowercase first-token readout gives 4.70 percent against 17.07 percent, and substitutions also move continuous answer scores where the binary prediction does not change. The labels are derived from reports, which limits conclusions about visual correctness; the direction replicates in two further model lineages. Code, the exact prompts and a run record for every number are at https://github.com/criticaldata/MODALENS.
1 Introduction
ModaLens addresses whether report-conditioned medical VLMs actually use images by holding the question and report fixed while swapping images. It measures paired output changes to separate image sensitivity from answer accuracy.
- Motivation: A report may describe a prior study, contain an error, or mention a finding that has resolved, so report-image disagreement can obscure image use.The audit measures output movement when the image changes while the report remains fixed.
- Motivation: ModaLens varies report availability while applying the same image substitution twice, measuring paired answer flips and margin changes.This design targets image-swap sensitivity rather than visual correctness.
- Related work: Earlier work documented text-only answering and visual omissions, while related medical VQA research treated image dependence as a debiasing target.ModaLens extends this line by manipulating report availability rather than merely omitting reports.
2 Methods
The study estimates image-swap sensitivity by replacing each case image while fixing its report and question, using paired MIMIC-CXR trials and clustered bootstrap intervals. Its all-14 design evaluates 13 findings plus a composite question, with substitutions usually drawn from the same patient and prompts tested in both block orders.
- Counterfactual design: Each case is rendered with its original image and a discordant substitute, changing only the image while holding report and question fixed.The estimand is sensitivity to the image intervention, not grounding of the queried finding.
- Cohort and units: The cohort is the MIMIC-CXR test split, with one frontal source image and one substitute per case; the estimand averages over trials.Alternative weighting and filtering choices yield paired differences between 10.6 and 14.3 points.
- Question design: The primary design asks all 14 questions per case: 13 finding-specific questions plus one composite acute-cardiopulmonary-finding question.A secondary one-question design uses the case’s own primary labeled finding or a composite no target.
- Image substitution: Substitutes are chosen once per case, usually from another frontal study of the same patient with a different CheXpert-positive set; view and acquisition time are not matched.Cross-patient substitutes are used for 19 single-study patients, and no pair exceeds the specified SSIM threshold of 0.9.
- Prompt and readout: The model receives an image block and a report-question text block, with text-first and image-first orders tested while the question template has no answer-format instruction.The report is not identified as historical or potentially unreliable in the main prompt.
- Statistics: Intervals use patient-clustered percentile bootstraps, retaining all trials from sampled patients; exploratory per-finding, category, and layer analyses are uncorrected for multiplicity.The primary all-14 analyses use 10,000 bootstrap draws, while other analyses use stated resampling procedures.
3 Experimental setup
The experiments use MedGemma-27B, an instruction-tuned Gemma 3 release, with controlled hardware and software environments. The study records exact package versions and follows a documented chronology for the readout experiments.
- Model and environment: Experiments use MedGemma-27B, an instruction-tuned release built on Gemma 3, with 62 decoder layers and residual width 5376.Runs use bfloat16 and one NVIDIA H200 or H100 under Python 3.11, PyTorch 2.10, and Transformers 5.3.
- Reproducibility: The run records include exact package versions and the experimental chronology begins with the lowercase first-token readout under the plain prompt.The explicit answer-instruction analysis followed the initial plain-prompt measurement.
4 Results
With the report fixed, replacing the image changes MedGemma’s generated answer far less than when the report is absent, while continuous margins still move without binary flips. The effect persists across readouts, question designs, model lineages, and robustness checks, but report-derived labels limit visual-correctness conclusions.
- 4.1 Report availability and image sensitivity: 20.94% without the report versus 4.26% with it, a paired difference of 16.7 points [15.6, 17.7] under the explicit answer instruction.The original prompt gives 17.07% versus 4.70%, a 12.37-point difference, and remains a sensitivity analysis.
- 4.1 Report availability and image sensitivity: The pooled effect varies substantially across findings: five findings flip at or below 4.1% without the report, while nine range from 16.8% to 32.0%.The five low-flip findings have image-only specificity at or near zero, and atelectasis and support devices do not flip on any trial.
- 4.1 Report availability and image sensitivity: 2.307 without the report versus 0.690 with it for mean absolute margin change, and 82.9% of trials move more without the report.Continuous margins change even when the binary prediction does not: 33.0% of label-unchanged trials move by more than 0.5 with the report.
- 4.3 Secondary analyses: The analysis uses report-derived labels, so agreement or accuracy comparisons do not establish visual correctness.Mechanistic steering and probing were negative, and the paper makes no localization claim.
- 4.1 Report availability and image sensitivity: The report effect replicates across Qwen3.5-9B, Qwen3.5-27B, and LLaVA-NeXT on Mistral-7B, with paired differences of 13.3, 19.60, and 12.75 points.The 4B model shows the same interaction with a 15.85-point paired difference, and substitute redraws change the one-question effect by at most 0.34 points.
- 4.4 Image sensitivity across decoder layers: Report-token attention blocking raises image-swap sensitivity when begun early, but later interventions reach baseline and do not localize all downstream report use.From layer 0, masking report attention raises flips from 1.50% to 7.75%, versus 12.00% with no report; single-layer blocking never exceeds 2.50%.
5 Limitations
The audit’s conclusions are constrained by report-derived labels, readout choices, substitution sampling, prompt order, and limited model and dataset coverage.
- Report-derived labels measure consistency with the report rather than visual correctness, and no independent image annotations were obtained.The audit measures output movement under image swaps, but cannot determine whether following the report was wrong.
- The binary readout misses margin shifts that do not cross its decision boundary, while report-absent rates vary by readout choice.Under the paper’s tokens the report-absent one-question rate is 10.3%, versus 17.4% on token families; the report-present rate does not vary.
- Unmatched substitutions can share the asked finding, lowering flip rates even for image-conditioned models, and prompt order changes effect magnitude.The report effect remains significant but is +7.8 points with text first versus +12.4 points with image first on identical trials.
- The headline rests on one model family and one snapshot, although the interaction’s direction and significance replicate across three model families.The evaluated families include MedGemma, Qwen3.5, and LLaVA-NeXT; replication establishes report anchoring, not image competence on the one-question set.
6 Conclusion
ModaLens finds that report availability changes measured image sensitivity in paired MIMIC-CXR image swaps. The effect persists across readouts, controls, generated-answer scoring, and additional model families, but independent image annotations are needed to assess visual correctness.
- 16.7 points: removing the report raises generated-answer image-change rates from 4.26% to 20.94% across 44,786 paired trials.The patient-clustered 95% interval for the paired difference is [15.6, 17.7] under an explicit answer instruction.
- The report effect persists across evaluated controls, generated-answer scoring, three model families, and two model sizes within two families.On six prespecified finding-specific questions, the effect ranges from 13.1 to 20.9 points in all three model families.
- Independent image annotations are needed to determine when report-conditioned behavior produces visually correct answers.The labels used by the audit are report-derived, so the measured effect concerns image sensitivity rather than visual correctness.
A Datasets, cohort and questions
The study uses MIMIC-CXR for the report-availability comparison and pairs image-swap analyses with multiple question designs and image-distance measurements. Its supporting datasets and agreement analyses provide controls and characterization rather than headline report comparisons.
- MIMIC-CXR is the only dataset pairing each image with a free-text report and carries every headline number.VQA-RAD and SLAKE serve as an image-decisive control and question-type breakdown, while OmniMedVQA and ProbMed appear only in supplementary tables.
- 3,199 discordant image pairs characterize the MIMIC-CXR cohort, with 3,180 same-patient pairs used for time-gap analyses.The pair composition table counts CheXpert positive-status differences between source and substitute studies.
- The primary design asks 14 questions per case: 13 CheXpert finding-specific questions plus one composite acute-cardiopulmonary-finding question.The secondary one-question design asks only the primary labelled finding or a no-target composite for cases without positive findings.
- With the report absent, flip rates rise from 8.5% in the most similar embedding-distance quartile to 14.5% in the most distant quartile.The corresponding increase is +6.0 points [2.5, 9.5]; with the report present, the rise is from 0.6% to 1.9%.
- Table 10 reports pooled observed agreement and chance-adjusted kappa for five findings, correcting for marginal imbalance rather than grounding.Intervals are patient-clustered bootstrap intervals over 293 patients.
D Retrospective chronology analysis of stale reports
The chronology analysis retrospectively compares same-patient substitutions that are later or earlier than the shown report. Historical or reliability framing does not measurably restore image sensitivity under this design, but the analysis cannot establish stale-report harmlessness.
- The historical instruction is a retrospective framing intervention, not a test of reasoning about staleness, because temporal consistency was not enforced.Chronology is reconstructed from signed study gaps after the model receives no chronology information.
- A stronger warning that the report may be wrong yields a generated-answer flip rate of 1.50% against the plain framing.Format compliance rises to 99.94% and 99.87% on the two image sides, while the flip-rate shift is +0.53 points [0.12, 0.97].
- 1.72% versus 1.74%: under historical framing, image-swap flip rates are nearly identical when the substitute is later versus earlier than the report.The difference is −0.02 points [−1.03, +1.06], with no detected difference between strata.
- The analysis does not show that stale reports are harmless, because no report was rewritten and independent dated image validation was unavailable.It is the nearest available stale-report condition in these data, not a direct test of real stale reports.
E Sensitivity analysis using manually annotated report labels
Manual report annotations closely agree with automated labels, supporting robustness to label extraction while not validating image findings.
- 0.943 agreement [0.938, 0.947] between automated and manual report labels supports sensitivity robustness to automated label extraction.The comparison covers 4,074 trials from 291 pairs and 90 patients with manual labels on both studies.
- Manual labels assess report-label extraction rather than visual correctness because the annotator read the report, not the image.
F Other models: MedGemma-4B, a second family and a third
The report-presence interaction replicates across Gemma, Qwen, and LLaVA-NeXT models, including prevalence-controlled questions, while image competence remains limited on this question set.
- MedGemma-4B: 15.85 points [13.45, 18.25] is the MedGemma-4B paired increase, with 3.16% flips with the report versus 19.01% without it.Its all-14 increase is 15.1 points [13.8, 16.3], positive across all 14 findings.
- Questions selected independently of the source label: 13.1 to 20.9 points is the report effect across models under six prevalence-selected finding-specific questions asked independently of source labels.The fixed prevalence rule avoids the one-question shortcut, while original-image balanced accuracy is above chance only for Qwen models and near chance for LLaVA-NeXT.
G Sensitivity of the primary effect to pairing and label choices
The primary report-presence effect persists across label policies, question subsets, substitute resampling, margins, text controls, and validated readouts, though its magnitude varies.
- Label choices: −0.50 points [−1.03, 0.02] is the changed-minus-unchanged flip-rate difference under the paper’s label policy, with similar results under alternative uncertain-label treatments.Label-changing trials therefore do not account for the report-presence interaction by themselves.
- Substitute resampling: 8.91, 8.66 and 9.00 points are the paired effects across three substitute-image sampling seeds, differing by at most 0.34 points.The flip rate varies in which cases change, but not materially in the size of the report effect.
- Continuous answer scores: The absolute margin change is larger without the report on 82.9% of trials, with aligned changes generally larger without reports across findings.Examples include pleural effusion at 0.52 versus 3.94 and edema at 0.32 versus 2.45.
- Readout validation: Under the validated instructed readouts, the all-14 paired difference is 16.1 points for lowercase logits, 16.7 for token families, and 16.7 for generated answers.The plain headline prompt instead gives 4.70% versus 17.07%, a 12.37-point difference.
J Modality ablation on both designs
Modality ablations show that report access changes image sensitivity across question designs and prompts, while surrogate-label and layerwise analyses qualify what the measured effect means.
- Balanced accuracy: −0.218 [−0.224, −0.212] is the pooled balanced-accuracy interaction, with a macro-averaged interaction of −0.128 [−0.141, −0.115].These values summarize the all-14 modality cells over 44,786 trials.
- Classifier surrogate: When the classifier surrogate predicts a queried-finding change, no-report flips occur three times as often as when it predicts the same label for both images.The report-derived split shows no comparable separation: 16.91% versus 17.68% without the report.
- Prompt and modality ablations: The generated-answer flip rate is 1.38% with the report versus 10.28% without it under the default one-question design, and 1.06% versus 8.28% under an alternative question template.The alternative template improves report-absent balanced accuracy from 0.479 to 0.621 without changing the report-present value of 0.908.
- Abstention: An explicit abstention option still yields an 18.3% no-report flip rate versus 1.1% with the report, a paired difference of 17.2 points [15.2, 19.3].The no-report condition never produced Uncertain, while the report-present condition did so on 1.5% of cases.
- Layer-wise readouts and attention interventions: The layerwise projection does not track the model’s answer until layer 46 of 62, and shuffled pairing reproduces its settling-depth difference.Attention knockout instead restores image sensitivity when blocking report influence begins shallow, but not when it begins deep.
N.1 The projection and its calibration
The logit-lens projection is calibrated against the model’s forced-choice output before interpreting layerwise behavior. It becomes informative only late, while exploratory projections and interventions do not support mechanistic claims.
- Calibration: The projection clears constant-predictor baselines only from layer 46 of 62, with balanced agreement of 0.909 with the report and 0.760 without it.Layer 46 was the shallowest informative depth in all 2,000 patient-clustered draws and all eight arms.
- Calibration: Below layer 46, the projection’s signal is limited and non-persistent, peaking at balanced agreement 0.648 at layer 25 before returning near the 0.500 constant baseline.It emits a single constant answer at 37 of 61 projected layers.
- Calibration: Layer-25 readout disagreement is prompt-specific: 0.2000 with the report versus 0.0067 without it, so the peak supports no network-level claim.The projection’s output-layer value equals the Table 4 flip rate, but intermediate behavior is not evidence of where decisions are made.
- Report-token attention knockout: Question-and-instruction blocking also restores flips over layers 22 to 27, confounding interpretation because the model may lose the question rather than recover image use.The report-specific difference is separated from this control in the earlier layer range.
- Report-token attention knockout: Report-token blocking raises image-swap sensitivity only when initiated at layers 0 to 20; later initiation diminishes the measured effect.This identifies when direct access to report-token positions affects the response, not where all downstream report use occurs.
- Limits: The report-token intervention has a ceiling and can over-restore: layer 22 reaches 14.00% versus 12.00% without a report, while patching gives a null result.The answer position still reads report tokens directly, and the wrong-donor patch reaches 14.07% versus 12.25% for the correct donor.
P Robustness of the main result
Robustness analyses preserve the main report-presence effect across prompt, input, decoding, corruption, swap-severity, and agreement analyses. These checks vary the protocol and report complementary behavior rather than replacing the primary estimate.
- Protocol robustness: The full robustness rerun changes prompt, input order, resolution, report section, decoding, and image content, with none of these variants entering the main-text estimates.The supplied passage describes these as protocol checks rather than primary analyses.
- Modality ablation: Modality ablation on 3,199 MIMIC-CXR cases gives accuracy 0.931 with both inputs, versus 0.913 report plus question and 0.773 image plus question.The image adds 0.018 in this report-derived-label evaluation.
- Prompt and input order: The modality-order pilot differs from the full-cohort rerun: the latter reports 2.84% text-first against 1.41% image-first.The pilot tables are superseded by the full-cohort rerun for this comparison.
- Additional robustness: Decoding-noise, blank-image, graded-corruption, and condition-specific agreement analyses provide complementary controls for answer stability and image sensitivity.The supplied table descriptions define these analyses but do not provide their cell values.
Q Other datasets and models
Additional datasets and model families test whether the paper’s image-sensitivity patterns extend beyond the primary MIMIC-CXR setup. Results support replication of the report effect in other lineages but do not establish a transferable mechanistic depth claim.
- Other datasets: Under the swap protocol, OmniMedVQA accuracy falls from 0.579 on the original image to 0.457 on the substituted image, a drop of 0.122 across 427 swaps.All 427 swaps changed the correct answer, so accuracy is scored against each condition’s ground truth.
- Mechanistic analyses: The supplied lens and patching analyses are descriptive: activation patching restores original answers most often late, but the main text draws no mechanistic claim.The projection tables distinguish model outputs from projected readouts and report settling properties separately.