Source-linked AI summary
A visual large language foundational model for medical image recognition using clinician-oriented social media
Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin
TL;DR
Medical multimodal LLMs lack large-scale VQA data that captures clinical reasoning and explicit image–text alignment. The study converts de-identified clinician-authored cases into ThoughtMed-1M and trains FOLTMed, which outperforms state-of-the-art models across diverse medical VQA benchmarks and response-generation metrics.
Problem
Medical multimodal LLMs lack large-scale VQA resources that provide clinically grounded reasoning and explicit image–text alignment.
Method
The study fine-tuned Gemma4-31B-PT on ThoughtMed-1M, a medical VQA dataset constructed from clinician-authored online cases.
Results
FOLTMed achieved superior performance across multiple modalities and task categories, with 3%–5% improvements over state-of-the-art models across BERTScore, FActScore, and AlignScore.
Takeaways & Limitations
ThoughtMed-1M and FOLTMed provide a scalable paradigm for clinically grounded multimodal medical AI research.
Takeaways & Limitations
Performance may require additional data curation and targeted training for underrepresented modalities and task types.
Abstract
from arXiv · showhide
Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared on clinician-oriented social media. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical logic and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M test set. It outperformed state-of-the-art models by 3--5% across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.
Author Information
The study was conceived, conducted, analyzed, and supervised by the listed authors, with contributions spanning data collection, model development, validation, clinical expertise, and manuscript preparation.
- L.H., Y.X., Y.Z. and J.L. conceived the study.
- L.H. and T.N. conducted data collection, data analysis, model development and model validation.
- L.H., T.N., Y.Z. and J.L. drafted the manuscript and supplementary materials, and all authors critically reviewed and approved the final version.
- Clinical experts contributed clinical expertise and clinical evaluation, while computational resources and hardware support were provided by K.C. and T.N.
- Correspondence is addressed to Trung Nguyen, Yan Zhuang and Jiangli Lin.
Main
The paper addresses the lack of large-scale medical VQA resources that connect images with clinical knowledge, structured reasoning, and explicit alignment. It introduces ThoughtMed-1M and FOLTMed as a clinically grounded data-and-model pipeline built from clinician-authored online cases.
- Motivation: Large-scale paired image–text data are needed to encode clinical knowledge, structured reasoning, and explicit image–text alignment.
- Motivation: Existing medical VQA resources are often closed-ended, short-answer, or restricted to a single modality, organ system, or task.These limitations constrain supervision for image–text alignment, spatial grounding, and structured diagnostic reasoning.
- Data source: Clinician-oriented platforms provide de-identified medical images and expert descriptions spanning diverse diseases and imaging modalities.
- Contribution: ThoughtMed-1M is a large-scale medical VQA dataset covering diverse diseases and imaging modalities with anatomical localization and clinically grounded reasoning.Its answers explicitly embed anatomical bounding boxes and stepwise clinical reasoning for image–text alignment, spatial grounding, and medical reasoning.
- Contribution: FOLTMed is a medical multimodal LLM developed for image interpretation, cross-modal reasoning, and question answering using ThoughtMed-1M.
- Contribution: The proposed paradigm transforms heterogeneous clinician-authored online cases into scalable resources for medical multimodal model development and evaluation.
Results
ThoughtMed-1M was assembled from large collections of clinician-authored cases and filtered into over one million VQA pairs. FOLTMed performed strongly across 42 medical VQA datasets and generated factually reliable, aligned responses, while underrepresented categories remained weaker.
- Dataset construction: 1,018,472 high-quality VQA pairs comprise the final ThoughtMed-1M dataset after quality control and filtering.
- Dataset construction: ChatGPT 5.2 achieved the highest overall VQA-generation mean score of 8.05 (95% CI: 7.95-8.15) and was selected for large-scale generation.It outperformed Claude Sonnet 4.5, Qwen3-VL-Max, and DeepSeek-V4 in the reported generation setting.
- Benchmark evaluation: 0.854 macro accuracy was FOLTMed’s best overall performance across 42 medical VQA datasets, exceeding Gemma4-31B-IT by over 4%.FOLTMed also exceeded MedGemma-27B’s macro accuracy of 0.807 and ChatGPT-5.2’s macro accuracy of 0.824.
- Benchmark evaluation: 84.93% anatomy-identification accuracy gave FOLTMed a 5.78% advantage over the second-best model.The reported advantage was associated with richer image descriptions and explicit bounding-box grounding in ThoughtMed-1M.
- Benchmark evaluation: 17 first-place rankings across 42 datasets gave FOLTMed the strongest overall ranking profile among the comparison models.It ranked in positions 2–3 on 21 datasets and positions 4–5 on 4 datasets.
- Limitations: Underrepresented modalities and task types may require additional data curation and targeted training to improve performance.
Discussion
The study presents ThoughtMed-1M and FOLTMed as a scalable approach for clinically grounded medical multimodal learning, while identifying data imbalance, residual errors, and deployment limits. ThoughtMed-1M supplies structured reasoning and image–text alignment, and FOLTMed shows strong performance across diverse medical tasks.
- Dataset and model: 1,018,472 high-quality VQA pairs comprise ThoughtMed-1M, with detailed clinical reasoning and structured medical supervision.The dataset was constructed from clinician-oriented public resources using LLM-assisted generation and quality control.
- Dataset and model: FOLTMed is a general medical LLM fine-tuned from Gemma4-31B-PT for diverse imaging modalities and clinical question types.The model was evaluated across 42 datasets spanning multiple modalities and task categories.
- Performance: 3%–5% improvements across BERTScore, FActScore, and AlignScore indicated stronger semantic fidelity, factual precision, and contextual alignment than SOTA models.Clinician-based assessment further evaluated clinical accuracy and practical utility on the ThoughtMed-1M test set.
- Limitations: Most task-level accuracies remained within 80–90%, so FOLTMed is not suitable as an independent clinical diagnostic standard.The authors position the dataset and model as a foundation for future research rather than a directly deployable clinical decision-making system.
- Dataset and model: ThoughtMed-1M incorporates structured reasoning, diagnostic decision-making processes, and bounding-box grounding for explicit image–text alignment.These annotations are designed to support medical knowledge, spatial grounding, and anatomically grounded visual reasoning.
Methods
The authors built ThoughtMed-1M from clinician-verified medical cases, images, and discussions, then used clinician-designed prompting and quality control to generate image-grounded VQA pairs.
- Case collection: Clinician-verified case information, imaging findings, and expert discussions provided factual grounding for VQA generation.The dataset drew on Radiopaedia and figure1.com cases and associated clinician discussions.
- Case collection: The pipeline extracted the first frame of each image series as the representative two-dimensional image for VQA generation.This choice followed the standardized plane displayed for each series.
- VQA generation: Prompting required questions and answers to use factual case information, image-specific features, and clinically valid reasoning without cross-image contamination or fabricated details.When correspondence was limited, the model generated simpler image-based questions; otherwise, it used step-wise diagnostic logic and spatial descriptions.
- VQA generation: The framework generated observational, analytical, diagnostic, symptom-imaging, and multi-step reasoning questions, with structured reports allowed only when evidence was sufficiently complete.Outputs were required to be traceable, image-specific, and free of hallucinated content.
- Quality control: A clinician-designed refinement and selection pipeline preserved ThoughtMed-1M pairs with stable factual grounding and clinically reliable reasoning.The resulting pairs were characterized as having high factual fidelity, strong image-text alignment, and meaningful reasoning depth.
Model construction and training strategy
FOLTMed was constructed by adapting a multimodal backbone through staged text and VQA training, then using reinforcement fine-tuning to improve clinical reasoning and spatial localization.
- Model initialization: FOLTMed used pretrained Gemma4-31B-PT as its multimodal backbone, with Gemma4-31B-IT serving as the comparison baseline.Full-parameter fine-tuning adapted the backbone to medical vision-language tasks.
- Domain-adaptive pre-training: Two-stage domain-adaptive pre-training used a Medical Domain Corpus followed by MIMIC-III to build biomedical knowledge and clinical-language competence.The first corpus established medical conceptual coverage, while MIMIC-III supplied realistic clinical communication and reasoning patterns.
- Domain-adaptive pre-training: Causal language modeling trained the model to predict each token from preceding context, aligning representations with biomedical and real-world clinical text.The objective targeted domain-specific terminology, linguistic structures, and reasoning conventions.
- Supervised fine-tuning: Supervised fine-tuning first used established medical VQA datasets and then ThoughtMed-1M to learn multimodal task patterns, clinical reasoning, and image-text alignment.Training instances conditioned answers on prompts, structured clinical information, and visual-encoder image embeddings.
- Supervised fine-tuning: ThoughtMed-1M supplied broad modalities, clinical narratives, spatial localization, and multi-step reasoning beyond basic disease identification.The dataset was described as having higher factual density, deeper reasoning structure, and more rigorous image-text alignment than existing public medical VQA datasets.
- Reinforcement fine-tuning: Reinforcement fine-tuning with GRPO ranked groups of candidate responses by relative reward, increasing higher-quality responses and suppressing lower-quality ones.This stage was introduced after SFT to address incomplete clinical concepts and inaccurate or incomplete bounding boxes.
- Reinforcement fine-tuning: GRPO-based refinement improved spatial localization accuracy, reasoning completeness and correctness, hallucination control, and robustness in complex medical VQA scenarios.The reported gains followed reinforcement fine-tuning under a composite reward.
Model validation
FOLTMed was evaluated against open and proprietary models on broad medical VQA benchmarks and curated clinical cases using accuracy, automated consistency, and clinician judgments.
- Benchmark comparison: The evaluation compared FOLTMed with representative open-source, domain-specific, general-purpose, and proprietary LLMs.Gemma4-31B-IT was used as the direct comparison baseline.
- Benchmark comparison: OmniMedVQA provided a comprehensive testbed spanning 42 sub-datasets, diverse imaging modalities, disease categories, and question types.The open-access subset was used for closed-ended medical multiple-choice evaluation.
- Automated evaluation: Automated evaluation used BERTScore, FActScore, and AlignScore to measure semantic similarity, factual precision, and contextual alignment.Responses were compared with reference answers and clinical evidence from curated Radiopaedia cases.
- Clinical-case evaluation: The clinical-case evaluation used 120 Radiopaedia cases yielding 2,158 VQA pairs across imaging modalities, anatomical regions, and clinical conditions.Questions followed the same generation protocol as the main dataset.
- Clinical-case evaluation: Ten experienced clinicians conducted multi-round questionnaire evaluations covering clinical accuracy, relevance, reasoning quality, completeness, and practical utility.This assessment simulated clinical interactions with each model.
- Evaluation design: The combined evaluation design assessed correctness under standardized benchmarks alongside usefulness, reliability, and reasoning quality in clinical scenarios.Accuracy-based and clinician-based evaluations provided objective and practical perspectives.
Statistical information
Model validation combined accuracy metrics, semantic and factual consistency measures, clinician ratings, and uncertainty estimates to assess performance across benchmark and clinical settings.
- Accuracy evaluation: Macro- and micro-accuracy evaluated closed-ended clinical decision-making performance on OmniMedVQA.Multiple-choice accuracy was the primary benchmark metric.
- Automated consistency: BERTScore, FActScore, and AlignScore quantified semantic similarity, factual precision, and contextual alignment, respectively.ChatGPT-5.2 served as the evaluator model for FActScore and AlignScore.
- Clinician evaluation: Clinicians rated model responses on a 0–10 scale across clinical accuracy, relevance, reasoning quality, completeness, and practical utility.Mean scores and standard errors were reported for the evaluation dimensions.
- Uncertainty estimation: Bootstrap resampling produced 95% confidence intervals and interquartile ranges to characterize rating uncertainty and distributional variability.These estimates supplemented the reported clinician-rating summaries.
Data Availability
The study uses publicly available medical datasets and plans to release ThoughtMed-1M through Figshare upon publication. Original images and posts are not redistributed because of licensing restrictions.
- Original medical images and case-post content are not redistributed to comply with Radiopaedia and figure1.com licensing policies.Users remain responsible for complying with applicable terms of use, licensing restrictions, and institutional or regulatory requirements.
- ThoughtMed-1M’s complete dataset, including generated VQA pairs and source-case URLs, will be released through Figshare upon publication.The repository is currently private and will become publicly accessible upon publication.
- The study uses publicly available datasets including PMC-VQA, SLAKE, Medical Domain Corpus, MIMIC-III, and OmniMedVQA.The OmniMedVQA subset includes all 42 validation datasets evaluated in the experiments.
Code Availability
The source code and FOLTMed model weights are hosted in a private repository and are intended for later public release.
- Source code and FOLTMed model weights are available through the ThoughtMed GitHub repository.The repository is currently private, with access planned for editors and reviewers upon request during peer review.
Ethics declarations
The study used de-identified public clinical data and anonymous clinician participants, with institutional review board approval.
- The study used de-identified clinical data from publicly available online sources and anonymous clinicians for evaluating the dataset and model outputs.
- The study protocol was approved by the Institutional Review Board of Guang’anmen Hospital, China Academy of Chinese Medical Sciences.The approval number was 2026-031-KY.
Competing interests
The paper presents workflows for curating clinician-authored cases, generating clinically grounded VQA pairs, evaluating candidate language models, and training FOLTMed.
- Data curation: Radiopaedia cases and figure1.com discussion posts were curated to construct ThoughtMed-1M training and test data.The cited text specifies training data collected through 26 December 2025 and test cases collected from 1 March to 1 May 2026.
- VQA generation: Clinical presentations, imaging reports, and expert discussions ground VQA pairs that capture image findings, diagnostic reasoning, and symptom–imaging associations.Generated answers include bounding-box annotations linking textual reasoning to relevant image regions.
- Model selection: Clinician ratings compare Claude-Sonnet-4.5, Qwen3-VL-Max, and DeepSeek-V4 across overall quality, answer quality, clinical utility, and relevance and accuracy.These ratings provide complementary evidence for selecting the final LLM used to construct ThoughtMed-1M.
- Evaluation: Heatmaps summarize six comparison models across imaging modalities and clinical question types using BERTScore, FActScore, and AlignScore.For each model, separate heatmaps summarize modality-specific and question-type-specific performance.