Source-linked AI summary
MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain
Sourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha, Amlan Chakrabarti, Saptarsi Goswami, Priti Singh
TL;DR
Clinical AI lacks benchmarks that jointly test multilingual, multimodal, and time-series-grounded reasoning. MMTClinic addresses this gap with a 30,000-pair benchmark and evaluation across clinical tasks, models, languages, and modalities, revealing distinct performance patterns and persistent limitations.
Problem
Existing clinical benchmarks do not adequately unify clinical text, images, physiological time series, multilingual evaluation, and reasoning tasks.
Method
MMTClinic combines text, medical images, and multivariate physiological signals in 30,000 MCQ and open-ended QA pairs across five languages, evaluating 13 models on ICU tasks.
Results
Open-source models perform well on text-based time-series reasoning, while proprietary models excel in complex multimodal settings; vision-language reasoning and regional-language performance remain challenging.
Takeaways & Limitations
MMTClinic provides a clinically validated benchmark for assessing multimodal, multilingual, and time-series reasoning in medical AI.
Takeaways & Limitations
The benchmark uses only five Indian languages and a single 2012 PhysioNet/CinC ICU data source, limiting broader linguistic and population generalization.
Abstract
from arXiv · showhide
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of clinically reliable and linguistically inclusive medical AI systems remains a significant challenge, primarily due to the lack of multimodal, multilingual, and time-series-grounded benchmarks that reflect the complexity of real-world clinical scenarios. To fill this gap, we present MMTClinic, a benchmark designed to evaluate large language models (LLMs) on complex reasoning and question-answering tasks involving clinical time-series. MMTClinic combines text, medical images, and multivariate physiological signals and includes 30,000 QA pairs (15,000 multiple choice questions (MCQs) and 15,000 open-ended questions) across five languages: English, Hindi, Bengali, Marathi, and Tamil. These questions cover three important clinical tasks---mortality prediction, heart rate forecasting, and SOFA score estimation. We evaluate 13 state-of-the-art LLMs in zero-shot, few-shot, and chain-of-thought settings. Our evaluation reveals notable differences in model performance across tasks, languages, and modalities, highlighting current limitations in clinical reasoning capabilities. MMTClinic provides a valuable resource for advancing multilingual, multimodal, and time-series-aware medical AI research. The dataset will be made publicly available on successful acceptance of the work.
1 Introduction
MMTClinic addresses gaps in clinical AI benchmarks by combining multilingual clinical text, images, and physiological time series for multimodal question answering and reasoning. It evaluates diverse models across ICU tasks, languages, modalities, and prompting settings with expert-refined data.
- Existing clinical AI studies largely focus on text or text–image inputs, leaving complex clinical time-series modalities underexplored.
- MMTClinic integrates multivariate clinical time series with textual and visual modalities for MCQ and open-ended reasoning tasks.
- 30,000 QA pairs span five languages: English, Hindi, Bengali, Marathi, and Tamil.
- The benchmark covers mortality prediction, heart-rate forecasting, and SOFA-score estimation using ICU data from the first 48 hours.
- MMTClinic evaluates 13 models across two QA formats, four modality settings, three ICU tasks, and three prompting modes.
- Approximately 70% of samples from each task and modality were independently reviewed by four certified medical experts, while linguists checked multilingual translations.
2 Related Works
Prior work leaves separate gaps in temporal, multimodal, multilingual, and open-ended clinical reasoning evaluation. MMTClinic unifies these dimensions in one benchmark across five languages and multiple clinical input types.
- Existing LLMs often fail to effectively handle temporal physiological signals or multivariate time-series inputs needed for clinical reasoning.
- Multimodal clinical approaches commonly combine text with static images such as radiographs rather than modeling multivariate temporal trends.
- MCQ-heavy evaluation may not adequately assess nuanced clinical reasoning beyond standardized question answering.
- MMTClinic combines multivariate clinical time series, textual and visual modalities, open-ended reasoning, and five-language evaluation.
3 Construction of MMTClinic
MMTClinic is constructed from ICU time-series data through preprocessing, multimodal question generation, translation, and expert validation. Its benchmark provides MCQ and open-ended reasoning data across multilingual clinical tasks and modality settings.
- Dataset dimensions: MMTClinic contains 30,000 questions in two formats: multiple-choice and open-ended reasoning.
- Dataset dimensions: The dataset supports five languages: English, Hindi, Bengali, Marathi, and Tamil.
- Data collection and preprocessing: Source data comprise ICU measurements recorded during patients’ first 48 hours, resampled hourly into numerical and visual representations of six physiological signals.
- Data collection and preprocessing: Preprocessing removes patients with more than 35% missing values, imputes remaining values, and uniformly resamples irregular signals to one-hour resolution.
- Question generation: The pipeline generated 15,000 MCQ and 15,000 open-ended reasoning pairs in English, then translated them into four additional languages.
- Validation: About 70% of samples from each task category were reviewed independently by four medical experts, alongside native-linguist evaluation of translations.
- Validation: Only 3.8% of Tamil and Marathi samples, 2.1% of Bengali samples, and 1% of Hindi samples required further refinement after scoring.
4 Experimental Section
MMTClinic evaluates diverse language models across text, time-series, and image settings using multiple prompting strategies, revealing strong task-, modality-, and language-dependent differences. Open-source models lead on text and temporal reasoning, while proprietary and multimodal models are comparatively stronger on image-related tasks.
- Evaluation setup: 13 language models are evaluated across zero-shot, few-shot, and chain-of-thought settings spanning text, time-series, and image inputs.The analysis covers four task-modality settings and compares model performance across clinical tasks, languages, and modalities.
- Text and time-series settings: In text and time-series reasoning, Qwen 3-235B reaches around 76% accuracy, whereas Qwen 2.5-7B remains below 10%.DeepSeek-R1-LLaMA-8B achieves about 71%, while proprietary models fall to the low-30% range.
- Image-based settings: Adding images causes a notable performance drop in MCQs, with GPT-4.1-nano leading at roughly 46% and vision models generally scoring lower.Qwen 2.5-VL models score 30–32%, while other vision-language models score around 26–30%.
- Image-based settings: Reasoning with text, time-series, and images is the most challenging setting, where Qwen 2.5-VL 7B slightly exceeds GPT-4.1-nano at 33% versus 32%.Larger multimodal models and Gemini 2-Flash fall into the low-to-mid-20% range.
- Overall comparison: Open-source models outperform on text-based MCQs and reasoning, whereas proprietary models perform better on image-based MCQs and multimodal reasoning.Multimodal models handle image-based tasks best, proprietary models show stronger multilingual consistency, and open-source models excel in text-only reasoning.
- Language effects: English generally leads multilingual performance, while reasoning accuracy declines most sharply in languages such as Tamil.Top models lose 15–20% in Tamil and Marathi for text-time-series reasoning, and all model tiers perform worst on reasoning tasks.
- Prompting performance: DeepSeek-R1 consistently performs best across multimodal and multilingual clinical tasks with minimal reliance on chain-of-thought prompting.It reaches up to 98.36% accuracy, while only a small subset of models effectively exploits chain-of-thought prompting.
5 Error Analysis
The error analysis distinguishes well-structured, clearly framed questions from questions weakened by limited subjectivity, insufficient contextual grounding, and semantic overlap.
- Questions answered correctly tend to be well-structured and clearly framed.
- Incorrectly answered questions more often involve limited subjectivity, insufficient contextual grounding in data, or semantic overlap.
6 Conclusion:
MMTClinic benchmarks multilingual, multimodal clinical reasoning across text, physiological time-series, and visual data. Its evaluation finds modality- and language-specific weaknesses alongside differing strengths between open-source and proprietary models.
- MMTClinic evaluates multimodal and multilingual reasoning across text, physiological time-series, and visual data in five languages.
- Experiments on 13 state-of-the-art models show that some open-source models handle time-series reasoning well, while proprietary models excel in complex multimodal settings.
- Consistent performance drops in vision-language reasoning and regional languages reveal persistent challenges in multimodal integration and cross-lingual generalization.
- The reported gaps are presented as reasoning limitations rather than dataset artifacts, supporting MMTClinic as a foundation for clinically grounded multimodal AI research.
7 Limitations
MMTClinic’s limitations concern language coverage, data-source diversity, visual richness, translation reliability, evaluation metrics, and the use of discrete answer formats.
- Coverage is limited to five Indian languages, which may not generalize to other global or low-resource languages.
- The benchmark relies solely on the 2012 PhysioNet/CinC ICU dataset, limiting evaluation to one institution’s patient population and measurement protocols.
- Visual inputs are restricted to line-plot images of time-series trends rather than richer modalities such as radiographs or CT scans.
- Machine translation may introduce subtle phrasing errors or biases despite expert review of regional-language translations.
- Accuracy-only evaluation does not capture calibration, reasoning depth, or clinical safety.
- Discrete outputs enable scalable objective comparisons but do not directly measure explanation quality or interpretability.
8 Ethical Considerations
The study describes compliance safeguards for proprietary-model use and emphasizes that the PhysioNet 2012 data were de-identified and openly accessible.
- Proprietary models were used according to each provider’s access guidelines, and no personally identifiable or protected health information was accessed or shared.
- The evaluations used the de-identified, openly accessible PhysioNet 2012 dataset to support adherence to HIPAA compliance and data privacy standards.
9 Appendix
The appendix documents benchmark construction, validation, evaluation settings, and model comparisons. It also records limitations involving explanation quality, deployment scope, and generalization.
- Validation: MMTClinic uses clinically grounded questions whose answers come from PhysioNet data, with expert review supporting their clinical relevance and linguistic quality.Medical and linguistic evaluations reported mean ratings from 4.2 to 4.5.
- Limitations: The benchmark prioritizes discrete answer correctness, leaving free-form rationales, faithfulness, and explanation quality for future extensions.This choice reflects subjectivity and scalability challenges in evaluating explanations.
- Evaluation settings: Evaluations compare models under identical inference settings across tasks, languages, modalities, and prompting conditions, including zero-shot, few-shot, and chain-of-thought setups.Models operate in inference-only mode without additional fine-tuning.
- Limitations and use: Error analysis associates incorrect answers with long-range temporal integration, multimodal grounding, limited contextual grounding, and semantic overlap.The benchmark is intended as a controlled evaluation resource rather than a clinical decision-support system.
- Results: DeepSeek-R1 leads across the multilingual reasoning comparisons, while GPT-4.1-nano follows and shows stronger chain-of-thought gains in Hindi and Bengali.DeepSeek-R1’s advantage is reported as 8 to 20%, while GPT-4.1-nano gains 6 to 8% in Hindi and Bengali.
9.7 Instructions for the Doctor for Validation
The validation protocol combines medical and linguistic expertise with checks for clinical relevance, ground-truth correctness, temporal consistency, ambiguity, and reasoning soundness.
- Validation reviewers: Four critical-care physicians with 8 to 12 years of ICU experience conducted the clinical validation.Linguistic validation was performed by native-speaking language experts.
- Validation outputs: Figures 15(a) and 15(b) visualize model performance across five languages, modalities, and few-shot versus chain-of-thought settings.The figures use radar-based comparisons for MCQ and reasoning tasks.
- MCQ validation: MCQ validation checks clinical realism, agreement with clinical data, temporal alignment, and absence of ambiguity.Reviewers verify that the correct option matches the CSV or multivariate plot.
- Reasoning validation: Reasoning-question validation checks pathophysiological soundness, data justification, completeness, and consistency with SOFA and mortality logic.Reviewers are instructed to avoid conclusions that exceed the available evidence.
9.8 Results
The results section presents empirical comparisons across languages, tasks, modalities, models, and prompting strategies. It includes both detailed tables and example predictions.
- Multilingual results: Tables 8–13 report model performance across Bengali, English, Hindi, Marathi, and Tamil multimodal inputs, including reasoning results.The tables organize results by language and task type.
- Prediction examples: Figure 16 provides examples of MCQ and reasoning predictions generated by the evaluated language models.The figure illustrates model outputs rather than aggregate performance.
- Prompting results: Tables 19–23 compare few-shot and chain-of-thought accuracy across languages, models, and multimodal reasoning settings.The comparisons include both MCQ and reasoning tasks.
9.9 Prompts
The prompts generate MCQ and open-ended reasoning questions from ICU time-series, outcomes, and optional line-plot images. They target mortality, heart-rate forecasting, and SOFA estimation using multi-step temporal reasoning.
- SOFA estimation: Example SOFA questions aggregate blood-pressure averages, respiratory-rate exceedances, and low-temperature events to infer severity.The prompts ground these judgments in physiological trends from the first 48 hours.
- MCQ prompts: MCQ prompts require two questions per subtask, four answer options, one supported correct answer, and three plausible distractors.The prompts prohibit reuse of prior examples or questions.
- Task design: The prompts define three subtasks: in-hospital mortality prediction, heart-rate forecasting over the first 48 hours, and SOFA severity estimation.Mortality uses outcome labels, heart-rate questions use summary statistics or windows, and SOFA questions use physiological deviations.
- Heart-rate forecasting: Example heart-rate questions compare HR changes with glucose events or compute mean, maximum, minimum, and time-window statistics.One example tests whether HR rises by more than 20% within 8 hours after Glucose >180.
- Reasoning prompts: Reasoning prompts require at least two parameters or timestamps, multi-hop numeric logic, and brief answers based only on the supplied patient files.The multimodal version adds line-plot images alongside CSV time-series and outcome data.
- Mortality prediction: Example mortality questions combine hypotension and tachycardia episodes with the in-hospital outcome label.The examples use thresholds such as SysABP < 90, DiasABP < 60, and HR > 100.