Source-linked AI summary

Can Multimodal Large Language Models Understand OCT?

Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li, Zihan Nie, Kailin Jiang, Yuntao Du, Weiye Song

arXiv:2607.16609v1cs.CVcs.CL

TL;DR

Existing OCT benchmarks do not adequately separate visual perception, medical cognition, and clinical reasoning. OCT-Bench evaluates these capabilities across 20 tasks and 20 MLLMs, finding that reliable OCT understanding remains challenging, with performance dropping from perception to reasoning.

  • Problem

    Existing OCT evaluations often conflate visual perception, medical cognition, and clinical reasoning, limiting informative assessment of where models fail.

  • Method

    OCT-Bench follows the clinical interpretation workflow, organizing OCT understanding into three dimensions, nine capability groups, and 20 fine-grained tasks.

  • Results

    62.0% overall accuracy was achieved by the best model, with performance dropping markedly from perception to reasoning and limited gains from scale or medical adaptation.

  • Takeaways & Limitations

    OCT-Bench enables comprehensive evaluation of MLLM capability bottlenecks across the progression from visual perception to clinical reasoning.

Abstract

from arXiv · show

Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.

Introduction

OCT-Bench addresses the lack of comprehensive OCT evaluation by modeling interpretation as a hierarchical process from perception through cognition to clinical reasoning. Its evaluation shows that current MLLMs remain unreliable, with performance declining substantially at more advanced capability levels.

  • OCT interpretation requires integrating specialized medical knowledge for disease analysis, clinical judgment, and decision-making beyond visual recognition.
  • Existing benchmarks often reduce OCT understanding to disease classification or isolated visual question answering, overlooking the hierarchical cognitive process of real-world interpretation.
  • Its clinically grounded taxonomy spans three dimensions, nine capability groups, and 20 fine-grained tasks for identifying deficiencies at different interpretation stages.
  • 62.0% overall accuracy is the best model result, while the highest perception score reaches 75.8% and the highest reasoning score only 42.9%.The benchmark evaluates 20 representative MLLMs across proprietary, open-source general-purpose, and medical-domain model families.
  • OCT-Bench contains 10,076 questions over 4,137 OCT images from seven public datasets, enabling evaluation beyond coarse disease classification.

Related Work

Prior work has advanced general-purpose multimodal model development and evaluation, but existing benchmarks are poorly suited to ophthalmic OCT understanding. OCT-Bench addresses this gap by evaluating fine-grained visual perception, medical cognition, and clinical reasoning.

  • Multimodal large language models: Proprietary models such as GPT-4o, Gemini, and Grok demonstrate strong multimodal capabilities, while open-source models improve reproducibility and transparency.The cited open-source models include LLaVA, Phi-3-Vision, InternVL, Qwen2.5-VL, and Gemma 3.
  • Multimodal benchmarks: General-purpose benchmarks advance MLLM evaluation but primarily target natural scenes, limiting assessment of ophthalmic anatomy, retinal lesions, and clinical significance.The cited benchmarks include MMBench, MME-RealWorld, MathVista, and SEED-Bench.
  • OCT understanding: OCT understanding requires recognition of fine-grained retinal structures and lesions alongside medical knowledge for clinical reasoning.OCT-Bench evaluates these capabilities across three progressive levels: visual perception, medical cognition, and clinical reasoning.

OCT-Bench

OCT-Bench models OCT understanding as a clinical workflow spanning Perception, Cognition, and Reasoning, organized into nine capability groups and 20 fine-grained tasks. Its construction combines systematic data and knowledge curation, task-driven question generation, expert quality control, and coverage and reliability analyses.

  • Capability taxonomy: OCT-Bench organizes OCT interpretation into Perception, Cognition, and Reasoning, further divided into nine capability groups and 20 fine-grained tasks.The taxonomy follows the workflow from visual evidence to medical concepts and clinical decision-making.
  • Capability taxonomy: Perception evaluates extraction of image attributes, retinal structures, reflectivity patterns, quantitative information, and spatial relationships from OCT images.These capabilities assess whether models capture fine-grained visual evidence as the basis for medical understanding.
  • Capability taxonomy: Cognition links visual information to anatomy, pathology, disease states, and functional impacts, while Reasoning supports disease assessment, treatment planning, and prognosis management.Separating Reasoning from Perception and Cognition distinguishes visual grounding failures from higher-level inference failures.
  • Benchmark construction: The benchmark pipeline comprises five stages: data collection, task design, medical knowledge collection, visual question answering generation, and expert quality control.Images come from seven public datasets, while cognition and reasoning tasks use evidence-based guidelines, expert consensus, and authoritative ophthalmic references.
  • Benchmark validation: 10,076 expert-verified multiple-choice questions comprise OCT-Bench after automated and expert checks for quality, answer uniqueness, image–text consistency, medical correctness, visual answerability, relevance, and ambiguity.The benchmark spans 10 disease categories, 2 types of region recognition, 5 types of retinal layer recognition, and 10 types of lesion recognition.

Experiment

Across 20 MLLMs evaluated zero-shot on OCT-Bench, the best overall accuracy was 62.0%, while performance declined from perception to reasoning and remained task-dependent. Medical specialization and scaling improved selected capabilities but did not ensure reliable, comprehensive OCT understanding.

  • Evaluation Setup: 20 representative MLLMs were evaluated under a unified zero-shot setting without fine-tuning or in-context demonstrations.The comparison covered proprietary, general-purpose open-source, and medical-domain models.
  • Overall Performance: 62.0%: GPT-5.4-mini achieved the best overall accuracy, followed by Gemini-2.5-flash at 60.5% and Hulu-Med-32B at 58.4%.All models exceeded the 25% random-guess baseline, but the best model still answered nearly 38% of questions incorrectly.
  • Capability-Level Results: 75.8% to 42.9%: the best dimension-level accuracy declined from Perception through Cognition at 64.2% to Reasoning, a 32.9-point drop.GPT-5.4-mini led Perception, but strong visual perception did not necessarily translate into clinical reasoning.
  • Model Comparison: Domain specialization and model scaling produced uneven benefits rather than uniform superiority across capabilities.Gemini-2.5-flash led Cognition, a medical-domain model led Reasoning, and scaling Lingshu from 7B to 32B increased Cognition by 14 points while Perception decreased slightly and Reasoning improved by 1.2 points.
  • Fine-Grained Tasks: 97.9% versus 23.4%: average accuracy was high for modality perception (T01) but low for subtle morphological description (T03), while disease diagnosis (T17) averaged 30.1%.Reasoning also remained weak: treatment planning (T19) and follow-up adjustment (T20) averaged 33.8% and 33.9%, with best scores of 56.0% and 50.9%.

Case Study

A representative Region Identification case shows that model performance varies substantially in identifying the choroid beneath the retinal layers. GPT-5.4-mini and Gemini-2.5-flash answer correctly, while Grok-4-fast and many open-source models confuse it with other ocular structures.

  • Region Identification: The highlighted region lies beneath the retinal layers and corresponds to the choroid.This example comes from the Region Identification task.
  • Region Identification: GPT-5.4-mini and Gemini-2.5-flash correctly identify the highlighted region as the choroid.
  • Region Identification: Grok-4-fast misidentifies the choroid as the retina.
  • Region Identification: General-purpose open-source models show greater variation, frequently confusing the choroid with the retina, vitreous, or cornea.

Conclusion

OCT-Bench evaluates OCT understanding through 20 fine-grained tasks organized across Perception, Cognition, and Reasoning, using expert-verified questions from 4,137 images. Across 20 MLLMs, reliable understanding remains limited, with reasoning, retinal structures, diagnosis, and clinical grounding posing major challenges.

  • Benchmark contribution: OCT-Bench contains 10,076 expert-verified multiple-choice questions built from 4,137 OCT images across seven public datasets.The benchmark follows the clinical interpretation workflow and spans 20 fine-grained tasks.
  • Benchmark contribution: The benchmark organizes OCT understanding into three dimensions: Perception, Cognition, and Reasoning.These dimensions support comprehensive evaluation of overall performance and capability bottlenecks.
  • Evaluation findings: 62.0% overall accuracy is achieved by the best model, with performance dropping markedly from perception to reasoning.The result comes from evaluations of 20 representative multimodal large language models.
  • Evaluation findings: Models consistently struggle with fine-grained retinal structures, disease diagnosis, and image-grounded clinical reasoning.These weaknesses indicate that reliable OCT understanding remains challenging for current models.
  • Evaluation findings: Larger model scale and medical-domain adaptation provide limited gains across the evaluated models.The conclusion reports that neither factor consistently resolves the benchmark’s capability challenges.
Loading 2607.16609v1…