Source-linked AI summary
A Source-Grounded Framework for Constructing and Evaluating Progressive Multimodal Diagnostic Dialogues from Clinical Case Reports
Yufan Wang, Rui Yang, Yi Liu, Yi Lin, Yifan Peng
TL;DR
Most medical benchmarks use fixed inputs or endpoint answers, while interactive diagnostic agents combine evidence selection with evidence interpretation. This paper constructs progressive multimodal diagnostic dialogues from case reports under shared staged evidence disclosure and evaluates diagnoses, reasoning, and image findings; reference dialogues largely preserved article-derived evidence, whereas two MLLMs showed substantially lower alignment, especially for image findings.
Problem
Most multimodal medical benchmarks evaluate fixed inputs or endpoint answers, while interactive diagnostic agents conflate evidence selection with evidence interpretation.
Method
The paper constructs source-grounded progressive multimodal diagnostic dialogues from case reports and evaluates diagnoses, diagnostic reasoning, and image findings under controlled evidence disclosure.
Results
Reference dialogues largely preserved article-derived diagnoses and evidence, whereas two evaluated MLLMs showed substantially lower alignment, especially for image findings.
Takeaways & Limitations
Fluent clinical reasoning is not necessarily evidence-grounded, supporting controlled evaluation of multimodal evidence interpretation.
Takeaways & Limitations
The study used predefined prompts, 24 publicly available internal medicine case reports, and two evaluated MLLMs, limiting evaluation of autonomous behavior and generalizability to routine clinical records.
Abstract
from arXiv · showhide
Clinical diagnosis requires progressive integration of patient history, physical examination, laboratory findings, medical images, and diagnostic-informative tests. However, most multimodal medical benchmarks evaluate fixed inputs or endpoint answers, while fully interactive diagnostic agents conflate evidence selection with evidence interpretation. We present a source-grounded framework to construct progressive multimodal diagnostic dialogues from case reports and an evaluation strategy for assessing MLLMs on final diagnosis, diagnostic reasoning, and image-finding interpretation. Evaluation on 24 internal medicine case reports showed that our framework can accurately convert case reports into reference dialogues, achieving a diagnosis F1 of 0.99 and a reasoning-quality score of 4.79 out of 5. Evaluation on two frontier MLLMs (o4-mini and Claude Haiku 4.5) achieved reasoning-quality scores of 2.75 and 2.50, respectively, with substantially lower diagnosis, reasoning, and image-finding F1 scores. The results demonstrate that fluent responses do not necessarily reflect evidence-grounded clinical reasoning and highlight the utility of the proposed framework for evaluating multimodal diagnostic reasoning.
1. Introduction
Clinical diagnosis unfolds through progressive integration of multimodal evidence, but existing evaluations either use fixed inputs or conflate evidence selection with interpretation. The framework addresses these gaps by constructing source-grounded dialogues and evaluating diagnosis, reasoning, and image findings separately.
- Clinicians progressively integrate history, examination, laboratory findings, medical images, and diagnostic-informative tests as cases unfold.
- Fully interactive diagnostic agents jointly select questions or tests, interpret returned evidence, and update differential diagnoses, making failure attribution difficult.
- Case reports support controlled evaluation because they connect patient-level clinical evidence, medical figures, diagnostic reasoning, and final diagnoses, while raw conversion risks information leakage.
- Correct diagnosis alone may conceal omitted evidence, unsupported image findings, source contradictions, or invalid intermediate inferences, while exact response matching is too brittle.
- The framework extracts article-derived references, aligns them with images, applies progressive evidence disclosure, and evaluates atomic findings alongside case-level reasoning quality.
2. Related Work
Prior multimodal benchmarks established image-question answering and open-ended biomedical visual generation, but largely used fixed inputs. This work combines component-level alignment with holistic reasoning assessment for diagnostic dialogue.
- VQA-RAD, SLAKE, LLaVA-Med, and Med-Flamingo primarily measure responses to predefined or fixed biomedical visual inputs.
- Open-ended diagnostic dialogue requires evaluation beyond exact matching or final-answer accuracy because retained clinical content does not guarantee sound diagnostic synthesis.
- Atomic-item alignment localizes specific evidence errors, while case-level reasoning-trace evaluation assesses the overall quality of diagnostic reasoning.
3. Methods
The study uses a source-grounded workflow to transform case reports into structured references and progressive multimodal dialogues, then evaluate both reference and MLLM-generated dialogues.
- The workflow comprises computer-interpretable representation extraction, progressive dialogue generation, and multifaceted evaluation.
- Article-derived references are structured source-supported information containing clinical elements, figure-linked evidence, and diagnostic assessment references.
- Reference dialogues and MLLM-generated dialogues are distinguished by whether they are generated from article-derived references or by evaluated multimodal models.
3.2. Input Requirements
The framework relies on PMC full-text packages with associated media and admits cases only when their structure, images, diagnoses, and clinical content support progressive dialogue construction and evaluation.
- PMC Open Access article packages provide JATS-compliant full-text metadata together with associated media files, including images.
- Eligible cases require parseable XML, at least one retrievable image, an extractable final diagnosis, and sufficient clinical and diagnostic content.
- For multi-figure cases, preserving original image order maintains alignment between visual evidence, article narrative, and diagnostic reasoning.
3.3. Computer-interpretable representation extraction
The framework extracts source-grounded, computer-interpretable case representations from clinical case reports. It separates clinical elements, figure-linked evidence, and diagnostic assessment references while avoiding unsupported inference and image-pixel interpretation.
- The extractor used a CARE-informed schema covering patient information, clinical findings, timeline, diagnostic assessment, interventions or outcomes, and diagnostic rationale.
- Each non-empty field had to contain source-supported text, while unreported information was left empty and new medical content was prohibited.
- The representation organized information into clinical case elements, figure-linked evidence, and diagnostic assessment references.
- Routine laboratory findings were separated from key diagnosis-informative findings such as immunohistochemistry, molecular, or genetic results.
- Figure-linked evidence included article-stated modality, clinical findings, and diagnostic significance, without inspecting or interpreting image pixels.
- The extractor retained the article-reported final diagnosis and source-supported diagnostic-process evidence without rewriting or extending the clinical reasoning.
3.4. Progressive Dialogue Construction
Progressive dialogues disclose clinical evidence in a staged workflow that mirrors diagnostic practice. History, examination, routine tests, images, and withheld diagnosis-informative findings precede final diagnostic synthesis.
- Dialogues were constructed from computer-interpretable case representations in a saved dialogue question-answer format.
- The protocol represented history, physical examination, laboratory testing, imaging, and downstream diagnostic actions as sequential clinical information nodes.
- The initial turns provided clinical history, followed by physical examination findings and routine laboratory results.
- Highly diagnosis-informative laboratory, molecular, microbiologic, genetic, pathology, or immunohistochemistry results were withheld before image review to avoid revealing the diagnosis.
- Images were presented sequentially in source-article order with raw images and non-diagnostic metadata, while prompts requested observations distinct from interpretation.
- After image review, key diagnostic findings were disclosed and the final turn required reasoning and diagnosis using the complete clinical and diagnostic information.
3.5. Atomic Item Evaluation
Atomic item evaluation decomposes references and model outputs for one-to-one matching across diagnoses, reasoning, and image findings. Strict and lenient matches distinguish clinically essential alignment from broader or less precise overlap.
- An LLM-as-judge workflow decomposed article-derived references and dialogue outputs into target-specific atomic items for final diagnosis, reasoning, and image findings.
- True positives were matched generated items, false negatives were missing reference items, and false positives were unsupported extra generated items.
- One-to-one matching allowed each item to be used once and returned strict matches, lenient matches, and contradictions.
- Strict matches preserved clinically essential meaning, whereas lenient matches shared the diagnostic or visual evidence core despite differences in precision or completeness.
- Diagnosis items preserved modifiers including etiology, organism, anatomical site, subtype, complication, severity, and causal relationships.
- Reasoning items captured clinically meaningful support, exclusion, differentiation, localization, characterization, or explanation steps, including incorrect or unsupported model claims.
- Image-finding items represented clinically meaningful visual observations, while diagnostic impressions such as schwannoma were excluded from visible-image units.
3.6. Case-Level FVCU Scoring
Case-level FVCU scoring evaluates diagnostic reasoning beyond item alignment. Judges assess factuality, validity, coherence, utility, and an overall reasoning-quality score using structured evidence summaries.
- Diagnostic reasoning quality was evaluated at the case level using criteria adapted from prior reasoning-trace evaluation work.
- The judge reviewed reference and generated diagnoses, reasoning, and matched, missing, extra, or contradictory evidence before assigning scores.
- Each FVCU dimension and the separate holistic overall reasoning-quality score used a 1-to-5 scale.
- FVCU dimensions: Factuality measured support for case-specific clinical, imaging, laboratory, pathology, microbiology, molecular, and genetic claims.
- FVCU dimensions: Validity measured whether diagnostic inferences were clinically reasonable given the evidence presented in the reasoning.
- FVCU dimensions: Coherence assessed internal consistency and logical organization, while utility assessed coverage, prioritization, and synthesis of case-defining evidence.
4. Experiments
The experiments evaluated reference-dialogue preservation and two MLLMs under identical progressive evidence-disclosure conditions across 24 internal medicine case reports and diverse medical images.
- Experimental corpus: 24 publicly available internal medicine case reports formed the experimental corpus, with a median of 3.5 images per case.Images spanned radiology, angiography, pathology, cytology, endoscopy, ophthalmic imaging, echocardiography, and electrocardiography.
- Evaluation pipeline: GPT-4.1 performed source-grounded extraction, reference-dialogue generation, and item decomposition, while GPT-5.1 performed semantic matching and FVCU scoring.
- Test models: o4-mini and Claude Haiku 4.5 received identical staged prompts, source-image order, and evidence-disclosure schedules.
- Evidence control: Neither evaluated model received figure captions, article-derived findings, diagnostic reasoning references, or the final diagnosis.Key diagnosis-informative findings were disclosed only during final synthesis.
- Metric aggregation: Final-diagnosis and reasoning metrics were macro-averaged over 24 cases, image-finding metrics over 86 evaluable image turns, and FVCU over 24 cases.
5. Results and Discussion
Reference dialogues closely preserved article-derived diagnoses and retained core reasoning and image evidence, while evaluated MLLMs showed substantially weaker alignment, especially for image findings. The study also identifies distinct sources of unmatched items and reports limitations concerning scope, annotation independence, and LLM-based evaluation.
- Reliability of Reference Dialogue Construction: 0.99 strict F1 measured reference-dialogue preservation of article-reported final diagnoses, alongside 0.99 precision and 1.00 recall.
- Reliability of Reference Dialogue Construction: 0.69 reasoning F1 and 0.67 image-finding F1 showed that reference dialogues preserved most article-derived evidence, with additional or differently segmented items.Their overall FVCU score was 4.79 out of 5.
- Performance of Existing Multimodal LLMs: 0.35 and 0.25 final-diagnosis strict F1 were obtained by o4-mini and Claude, respectively, with strict precision of 0.32 and 0.20.
- Performance of Existing Multimodal LLMs: 0.28 and 0.24 lenient reasoning F1, and 0.17 and 0.11 lenient image-finding F1, were obtained by o4-mini and Claude, respectively.Image-finding precision was 0.12 for o4-mini and 0.06 for Claude, indicating unsupported visual claims.
- Performance of Existing Multimodal LLMs: 4.54 and 4.29 coherence scores contrasted with utility scores of 2.29 and 2.12 for o4-mini and Claude, respectively.The reported pattern identifies evidence coverage and factual grounding as the main limiting factors despite coherent narratives.
- Error Analysis: 0.35 and 0.38 extra-generated-item rates for reference reasoning and image findings mainly reflected added explanatory wording or different item granularity.These unmatched items were not necessarily clinically implausible.
- Error Analysis: One reference ECG turn omitted extracted annotation details such as color markers for P waves and QRS complexes during concise dialogue reformulation.Core atrioventricular-block and junctional-escape findings were preserved.
- Error Analysis: 0.89 and 0.94 image-finding extra-generated-item rates for o4-mini and Claude indicate many unsupported visual claims beyond missed source findings.
6. Conclusion
The paper presents a source-grounded framework for evaluating evidence integration in progressive multimodal diagnostic dialogues. Reference dialogues largely preserved source evidence, whereas evaluated MLLMs showed lower alignment, particularly for image findings, despite high coherence.
- The framework provides a controlled basis for evaluating multimodal evidence interpretation and developing adaptive diagnostic agents and clinician-facing evaluation protocols.