Source-linked AI summary

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

Benlu Wang, Yifan Zhang, Jiaqing Yu, Chin Siang Ong, Juncheng Huang, Zhuohao Li, Zhenyu Zhang, Arman Cohan, Hong Yu, Zonghai Yao

arXiv:2609.03261v2cs.CVcs.CL

TL;DR

Medical multimodal benchmark scores may credit shortcut routes as well as intended image-grounded reasoning, making route-level validity evidence necessary. The paper audits six datasets with modality ablations and matched repairs, finding substantial non-image performance and measurable effects from three option-form cue families. It also constructs MEDQA-MM, a 1,000-item shortcut-mitigated subset, while retaining residual-risk and validation caveats.

  • Problem

    A single medical multimodal MCQ score does not reveal whether correctness came from intended image-grounded evidence or benchmark-preserved shortcut cues.

  • Method

    The paper audits six datasets using prompt- and image-side screens, modality ablations, matched repairs preserving the medical target and answer key, and expert review.

  • Results

    Across 13 open-model configurations, full-input accuracy is 62.63%, versus 53.96% text-only and 29.71% options-only; removing three cue families lowers accuracy by 6.58, 3.50, and 4.77 percentage points.

  • Takeaways & Limitations

    Medical image-reasoning claims require route-level evidence, and MEDQA-MM is shortcut-mitigated rather than shortcut-free.

  • Takeaways & Limitations

    Several audit components are model-judged or rule-screened rather than exhaustively clinician-validated, and sample-based review cannot guarantee every item.

Abstract

from arXiv · show

A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.

1 Introduction

Medical multimodal MCQ accuracy can reflect intended image-grounded reasoning or benchmark-preserved shortcut routes, so scores alone may overstate what models demonstrate. The paper audits these routes and reports behavioral evidence from ablations, matched repairs, and a shortcut-mitigated subset.

  • Route ambiguity: A benchmark score records correctness but not whether the answer came from a construct-relevant image clue or a shortcut cue.Shortcut cues include option form, non-visual context, visible image text, markup, and other preserved artifacts.
  • Construct: The intended construct integrates medical visual evidence, clinically relevant text, and answer-option semantics, while shortcuts are item-level construct-irrelevant signals.Clinical context, devices, demographic facts, or annotations are not automatically shortcuts; validity depends on the specific item.
  • Route ambiguity: Reasoning inflation names the risk of interpreting benchmark accuracy as direct evidence of image-grounded reasoning when multiple answer routes are available.The term does not claim that models never use images; it identifies an interpretation problem for unaudited scores.
  • Approach: The study combines six-dataset audits, modality ablations, and safe matched repairs that preserve the medical target and answer key.This evidence ladder separates candidate cues from stronger behavioral evidence.
  • Results: 62.63% full-input accuracy coexists with 53.96% text-only and 29.71% options-only accuracy across 13 open-model configurations.The results show that answer selection can remain substantially accurate without full multimodal input.
  • Results: 6.58, 3.50, and 4.77 percentage-point drops follow removal of length-gap, absolute/conspicuous, and spatial/prepositional cues, respectively.These matched repairs provide behavioral evidence that the targeted option-form cues affect model accuracy.
  • Results: MEDQA-MM contains 1,000 shortcut-mitigated items, with text-only and options-only accuracy reduced to 5.21% and 12.33%.The authors describe the subset as shortcut-mitigated rather than shortcut-free.

2 Related Work

Prior work shows that benchmark scores can be inflated by item-writing flaws, annotation artifacts, language priors, and medical image-side cues. The paper brings these concerns together in a data-centric audit-and-repair framework.

  • MCQ cueing: MCQ studies identify option length, absolute wording, paired options, grammatical mismatch, and implausible distractors as answer cues.These item-writing concerns motivate auditing answer choices in multimodal medical benchmarks.
  • Dataset artifacts: NLP and VQA research shows that high accuracy can arise from annotation artifacts, shallow heuristics, language priors, or answer-set regularities.Such findings establish that aggregate scores may not isolate the intended capability.
  • Medical multimodal evaluation: Medical image benchmarks add visible labels, arrows, overlays, devices, acquisition markers, and care-process artifacts as potential shortcut sources.These image-side cues make score interpretation especially important in medical evaluation.
  • Benchmark repair: A data-centric workflow detects candidate cues, repairs them when safe, validates medical validity, and documents residual uncertainty.The approach treats benchmark quality as dependent on construction and maintenance, not only model design.

3 Method

The method builds an evidence ladder that moves from screened shortcut candidates to ablation, repair, and expert validation while preserving medical answer validity. It audits prompt and image channels, constructs targeted repair sets, and evaluates multiple input routes.

  • 3.1 Construct and Evidence Ladder: The audit distinguishes screened, ablation-supported, repair-supported, and expert-validated candidates as progressively stronger evidence levels.Final item disposition depends on cue status, repair fidelity, validity, and residual risk.
  • 3.2 Shortcut Scope: Text-side screens cover option structure, wording, lexical overlap, and semantic compatibility, while image-side screens cover text, markup, hints, devices, and acquisition artifacts.These screens feed targeted ablations, repairs, and expert review.
  • 3.7 Human Review and Shortcut-Mitigated Subset: The six-dataset workflow audits option form, non-visual context, and image-side hints before retaining medically validated items in MEDQA-MM.Table 1 organizes paired analyses and treats image-text and device/material counts as curation risks unless paired evidence is available.
  • 3.3 Option-Form Repairs: 306 length-gap, 101 spatial/prepositional, and 101 absolute/conspicuous repair pairs use minimal edits that preserve the stem, image, answer key, option order, and one-best-answer structure.The common metric is ∆text = Accedited − Accoriginal, so negative values indicate reduced accuracy after cue removal.
  • 3.4 Modality Ablations: The evaluation compares full input, text-only, options-only, and Image+Options settings to test whether full-input scores isolate image-grounded reasoning.Diagnostic options-only and Question+Options guessers are stress tests rather than deployable systems or leaderboard baselines.
  • 3.5 Image-Channel and Natural-Cue Interventions: Image-hint interventions produce 200 paired original/edited cases, while natural-cue interventions retain 302 cases after neutralizing context without changing the image, options, key, or core medical content.The intervention metrics use modified minus original, and bad-use measures harmful reliance on removed cues.

4 Experiments

The experiments cover six medical multimodal MCQ datasets and report accuracy, accuracy changes, uncertainty intervals, and expert-review counts under repeated evaluations. The canonical audit contains 7,706 examples.

  • Datasets: 7,706 examples constitute the canonical audit across six medical multimodal MCQ datasets.The datasets include AMBOSS, JAMA Clinical Challenge, MMMU Health/Medicine, MedThinkVQA, MedXpertQA-MM, and NEJM Image Challenge.
  • Metrics and Uncertainty: Accuracy and accuracy change are reported in percentage points, with random accuracy defined as the mean item-level chance rate, 1/K.Repeated generations are averaged within item×configuration×condition cells before uncertainty estimation.

5 Results

Across six datasets, full-input accuracy mixes image reasoning with option-form, no-image text, and image-channel routes. Modality ablations, matched repairs, and targeted interventions show that these shortcuts can materially affect predictions, while mitigation weakens non-visual routes.

  • Option-Form Shortcuts Create Answer Priors: Options-only prediction exceeds the item-level random baseline on every audited dataset, and matched repairs reduce accuracy by 6.58, 3.50, and 4.77 percentage points.The repaired cue families are length-gap, absolute/conspicuous, and spatial/prepositional wording, respectively.
  • Destination-Specific Evidence for Length and Spatial Cues: Targeted option edits increased selection of the edited distractor by 9.47 percentage points for length-gap pairs and preserved a +2.70-point spatial contrast after exclusions.These destination shifts are more specific than generic option-identity-agnostic difficulty, though target-specific wording or plausibility changes may still contribute.
  • Question Text and Options Create Strong No-Image Routes: 62.63% full-input accuracy coexists with 53.96% text-only and 29.71% options-only accuracy across the 13-configuration panel.The comparison demonstrates substantial no-image answerability, although full input is usually strongest.
  • Image-Channel Cues Are Common, But Reliance Is Model-Dependent: Image-side screening found 5,645 text residuals, 4,970 markups or overlays, 2,642 demographic/context features, and 955 device/material cues across 13,940 images.These prevalence counts establish curation risk rather than model reliance; paired interventions show reliance varies by model.
  • Visible Image-Text Leakage Is Rare But Direct: Image-text leakage and device/material cues are treated as direct data-quality risks, but their limited repairability prevents using them as aggregate performance results.The audit identifies 65 high-risk image-text cases with only 15 directly repairable, while device/material screening confirms 9 shortcut-risk cases among 1,347 candidates.
  • Shortcut Mitigation Weakens Non-Visual Routes: MEDQA-MM reduces text-only and options-only accuracy to 5.21% and 12.33%, while retaining 26.59% full-input accuracy and 28.47% Image+Options accuracy.The 1,000-item subset is shortcut-mitigated rather than shortcut-free, and Image+Options exceeding full input cautions against over-reading one setting.

6 From Two-Aha Framing to Route-Aware Measurement

Full-input accuracy is a route-mixed measurement because correct answers may arise from visual evidence or benchmark-preserved cues. The audit therefore combines item-level cue judgment, modality diagnostics, matched repairs, and residual-risk documentation.

  • Route-Aware Measurement: Full-input accuracy can mix image-grounded reasoning with option form, no-image text, visible labels, annotations, devices, or natural context.Strong text-only, options-only, or Image+Options results indicate that a reported gain may reflect residual cues rather than improved visual reasoning.
  • Evidence Ladder: The audit treats cue presence separately from cue dependence by linking cue use to correctness changes after removal or neutralization.Its evidence ladder progresses from candidate detection to rationale-level use, paired behavior change, and validation.
  • Item-Specific Validity: Medical context has no fixed validity status because the same device, annotation, or demographic fact may be target evidence in one item and an unintended route in another.Clinically necessary context is retained and documented, whereas nonessential predictive context is repaired, flagged, or excluded.
  • Benchmark Repair: The workflow detects cues, minimally repairs them when safe, evaluates matched pairs, validates edits, documents residual risk, and discards unsafe repairs.Blanket deletion can discard useful material or introduce new confounds.

7 Conclusion

The paper argues that medical multimodal benchmarks should not be treated as direct measures of image-grounded clinical reasoning without shortcut audits. It unifies construct-validity analysis across modalities, quantifies shortcut effects, and builds a shortcut-mitigated subset.

  • Conclusion: Medical multimodal benchmarks require shortcut auditing before their scores are interpreted as direct evidence of image-grounded clinical reasoning.The framework connects MCQ cueing, NLP/VQA artifacts, and medical imaging shortcuts.
  • Conclusion: The study quantifies shortcut prevalence and no-image signal across six datasets, identifies three behaviorally consequential text shortcut families, and constructs MEDQA-MM.The three families are length-gap, absolute/conspicuous, and spatial/prepositional cues.

Limitations

The audit’s strongest behavioral claims are limited to three option-form shortcut families, while other analyses remain primarily curation-risk evidence. Validation, automated repairs, and public release also face important scope and compliance constraints.

  • Behavioral evidence is strongest for three option-form shortcut families; image-text and device/material analyses remain primarily curation-risk evidence.
  • Model-judged or rule-screened audits and sample-based expert validation cannot guarantee every screened or repaired item is error-free.GPT-generated edits may also alter wording, fluency, plausibility, or difficulty beyond the targeted cue.
  • MedQA-MM is built from existing datasets, so public release requires source-license, image-rights, item-level validation, and redistribution checks.

Ethics Statement

The paper frames benchmark auditing and release as requiring privacy, consent, licensing, and clinician-review safeguards. Its appendices document evidence routes, validation status, denominators, and the final inventory underlying the shortcut-mitigated benchmark.

  • Ethics Statement: Repaired subsets must respect source licenses, consent constraints, and privacy requirements, while demographic or clinical-context edits require clinician review.The work is an audit of benchmark reliability, not clinical advice or a deployment evaluation.
  • Documentation: Appendix materials separate text- and option-side analyses, image- and context-side evidence, benchmark construction, prompts, taxonomy, and representative cases.Table 3 maps main-text claims to appendix evidence, and Table 4 reconciles denominators.
  • Inventory and Release: The canonical six-dataset audit underlies the main inventory, while the final MEDQA-MM benchmark contains 1,000 examples after additional filtering and device/material cue removal.Feature-specific screens use slightly different totals because of option normalization or source variants.

A.3 Detailed Experimental Settings

The experiments combine four input settings, diagnostic guessers, and matched repairs to test option-form shortcuts across six datasets. Accuracy is aggregated across fixed configuration panels, while repair analyses report paired modified-minus-original changes with controls.

  • Modality and panel settings: 62.63% full-input accuracy exceeded 53.96% text-only and 29.71% options-only accuracy across the 13-configuration panel.The four-setting panel also includes Image+Options; configuration-level accuracies are averaged with equal weight.
  • Diagnostic guessers: The diagnostic guessers use options-only or Question+Options inputs, with unique-item splits and evaluation identifiers excluded from training.The options-only model receives no question text or image, while the Question+Options model receives no image pixels, captions, rationales, or explanations.
  • Length-gap repair: The length-gap intervention expands distractors while preserving the correct option, option count and order, and one-best-answer structure.The final 306 pairs were evaluated in original-versus-edited conditions with 20 repeated runs.
  • Spatial/prepositional repair: The spatial/prepositional repair set contains 101 validated pairs, including a 46-case distractor-only subset and a nested 39-case length-control subset.The 39-case subset excludes seven newly unique-longest edits and is not a spatial reverse intervention.
  • Repair evaluation: Matched-repair tables report signed accuracy changes as modified minus original, with negative values indicating reduced accuracy after cue removal.The spatial 39-case analysis is nested within the 46-case clean subset, which is nested within the 101-pair repair set.
  • Length-gap controls: Figure 7 combines forward and reverse length-gap interventions with a 21-item text-identical Options-only control and a natural margin screen.The control remains near zero, while the natural-screen bins count records rather than verified unique items.

B.5 Additional Text-or-Options Ablation Results

Additional analyses treat lexical and option-form associations as observational screens and evaluate three accepted repair cohorts with paired repeated-run comparisons. The text-or-options analysis is narrower than the main modality lattice but supplies item-level union diagnostics.

  • Observational screens: Lexical clang and length-as-unique-longest associations are observational screens used to motivate controlled edits, not causal estimates.The reported canonical odds ratios are 1.97 and 1.33, respectively.
  • Repair cohorts: The three accepted repair cohorts contain 306 length-gap, 101 spatial/prepositional, and 101 absolute/conspicuous pairs evaluated over 20 repeats per condition.Construction and validation criteria were applied before paired evaluation.
  • Text-or-options analysis: Text-or-options is defined as the union of cases solved in either options-only or text-only settings.Table 24 reports this narrower auxiliary analysis for item-level union diagnostics.
  • Image-side analyses: Tables 27 and 28 report model- and family-level image-channel results after image-side screening and paired intervention construction.Only retained paired cases enter original-versus-edited model evaluation.

C.3 Natural Clinical-Context Cue Details

Natural clinical-context analyses retain 302 repairable cases and distinguish cue presence from harmful conversion through paired neutralization and rationale-based evidence. The study separately audits visible image text and device/material cues, with repairability constrained by medical validity.

  • Natural-cue construction: The natural-cue set contains 302 paired cases whose clinical-context cues can be deleted, generalized, or neutralized while preserving the image, options, key, and core medical content.Cue categories include demographic, social, geographic, occupational, access, and reproductive-context information.
  • Model-level natural-cue behavior: Family patterns are mixed: GPT-5.4 drops while GPT-5.4-mini improves, Qwen3.5 shows the strongest harmful reliance, and Qwen3-VL often mentions cues without high harmful conversion.One MedGemma configuration frequently mentions cues but shows a weak aggregate drop.
  • Interpretation boundary: Cue counts are not a demographic-bias benchmark; a clinical-context cue is problematic only when nonessential to the intended image-grounded construct and contributing to answer selection.The categories describe shortcut-like behavior in a selected neutralization set rather than dataset prevalence.
  • Visible image-text leakage: Visible image-text leakage is screened through overlap with correct-answer content, adjudication, and repairability review, but many cases are not safely editable.When text names or strongly cues the answer, masking, cropping, or neutral relabeling is preferred; otherwise the item should be removed or redesigned.
  • Device and material cues: Device/material shortcuts can narrow answers through treatment history rather than intended clinical or imaging evidence, but confirmed cases are rare and difficult to repair safely.Removing hardware or implants may alter the medical image or create artifacts, so unsuitable items should be excluded from the shortcut-mitigated benchmark.
  • Shortcut-mitigated benchmark: MEDQA-MM contains 1,000 hard, image-dependent examples after confirmed device/material cue cases were removed, with documented residual risk rather than a shortcut-free guarantee.The benchmark is therefore presented as shortcut-mitigated, not shortcut-free.

D.2 MEDQA-MM Error-Analysis Protocol

The error-analysis protocol reviews sampled MEDQA-MM errors with mutually exclusive failure labels and expert validation of repairs, while treating aggregate modality differences as underdetermined.

  • Error-analysis protocol: The analysis samples 100 incorrect predictions from each of GPT-5.4, MedGemma-27B, and Qwen3.5-122B-A10B configurations.GPT-5.6-Sol judged whether question-stem content overrode decisive image evidence.
  • Error-analysis protocol: 36 of 70 development-split errors were labeled visual-understanding or reasoning/integration failures, but the proportions are descriptive rather than held-out estimates.The five mutually exclusive categories were visual-understanding, reasoning/integration, item/data issue, response/scoring failure, and shortcut/anchoring bias.
  • Interpretation: Neither the error sample nor its taxonomy identifies a single cause of the aggregate Image+Options–Full difference.Stem interference appeared in a minority of sampled errors and may explain some cases where removing the stem improves performance.
  • Expert validation: Expert review covers 150 option-form repairs, 50 natural-cue neutralizations, and 46 image-hint edits, assessing cue relevance, medical validity, answer preservation, and mitigation quality.The review pool includes two senior clinicians and three medically trained annotators.
  • Dataset context: The audit reports a 1,000-item hard, image-dependent MEDQA-MM subset after removing device/material cue cases.Figure 8 averages configuration-level accuracies within each base model when multiple inference modes are evaluated.

F.2 Image-Side, Multimodal, and Dataset-Level Cues

The taxonomy covers image-side, acquisition, device, dataset, and interface cues, then assigns evidence-based handling rather than treating every detected cue as an automatic exclusion.

  • Image-side cues: Image-side risks include diagnostic labels, anatomy labels, arrows, regions marked by shapes or overlays, enhancement artifacts, devices, and care-process context.Mitigation depends on whether the cue defines the task or bypasses intended visual reasoning, with expert validation when edits may damage interpretation.
  • Acquisition and dataset cues: Acquisition and source artifacts include modality or projection markers, protocol giveaways, scanner style, answer-class imbalance, repeated templates, near duplicates, option count, and translation artifacts.Suggested controls include compatible option sets, cross-source validation, source-stratified splits, deduplication, majority baselines, and language balancing.
  • Evidence tiers: Candidate cues establish possible presence, whereas behaviorally useful cues require ablation, matched repair, rationale, or prediction-change evidence suggesting model use.Validated repairs additionally require expert judgment that the cue is shortcut-like, the edit is medically valid, and mitigation improves or partially improves performance.
  • Evidence tiers: Unresolved cues are residual risks that should carry explicit risk flags or move to an appendix rather than support shortcut-free claims.Cases without controlled post-edit evaluation remain qualitative only.
  • Scope: The broader taxonomy is intended as a reusable checklist for future releases, while headline experiments emphasize categories with stronger current evidence.The main categories include option-form repairs, non-visual solvability probes, visual-text and annotation cues, device/material cases, natural-context neutralization, and MEDQA-MM construction.

G.1 Sociodemographic Cue Sensitivity

Paired counterfactuals show that adding sociodemographic descriptors can overturn image-grounded diagnoses, while modality ablations expose device and image-text shortcuts that support answer selection. These cases distinguish controlled behavioral evidence from qualitative shortcut risk and legitimate visual findings.

  • Sociodemographic cue sensitivity: Adding a single identity descriptor flipped the diagnosis from the image-grounded gold answer to a demographic-linked distractor in a paired evaluation.The model changed from pulmonary epithelioid haemangioendothelioma to sarcoidosis after the stem added that the patient self-identifies as Black.
  • Sociodemographic cue sensitivity: An uninsured-patient descriptor similarly shifted a text-only answer from sarcoidosis to tuberculosis by invoking socioeconomic disadvantage.The original stem was extremely short, so the added descriptor and options carried disproportionate weight.
  • Sociodemographic cue sensitivity: A Spanish-interpreter descriptor changed the answer from echinococcosis to amebic liver abscess by serving as an unsupported endemicity proxy.No travel, immigration, animal-exposure, or sanitation history was added.
  • Device and material cues: Image-plus-options ablation recovered the keyed osteogenesis-imperfecta figure through visible orthopedic hardware, despite the full multimodal answer being wrong.The audit classifies this as device-shortcut risk because treatment hardware can act as a severity prior rather than evidence of the collagen disorder itself.
  • Image-text leakage: The IVC-filter case separates legitimate device recognition from image-text leakage: embedded brand text offered a stronger lookup route than morphology.Removing the image caused the text-only answer to become wrong, while the original rationale cited both filter morphology and the visible “OPTEASE” brand.
  • Image-text leakage: The embedded table in another case contained the exact gold answer string, and the image-plus-options model used it successfully while text-only input failed.The OCR audit marked this as a complete correct-option leak, showing that the visual text alone could recover the answer without the clinical stem.
Loading 2609.03261v2…