Source-linked AI summary

Volumetric Radiology AI in the Era of Multimodal Large Language Models

Zanting Ye, Shengyuan Liu, Xin Liu, Chenhui Wang, Zhisong Wang, Jiashuai Liu, Zipei Wang, Cheng Wang, Wentao Pan, Mengjie Fang, Di Dong, Mohammad Salmanpour, Arman Rahmim, Yu Gu, Yong Xia, Hongming Shan, Yixuan Yuan, Yefeng Zheng, Lijun Lu

arXiv:2608.20549v1cs.AI

TL;DR

Volumetric radiology AI must reconcile full-volume, quantitative, and contextual evidence with MLLMs that often receive selected views, compressed representations, or reports. This review synthesizes model-level representations and system-level agents across more than 200 publications, using Claim–Design–Validation to assess their clinical support. It concludes that native volumetric modeling and agentic workflows should be matched to task requirements, with traceability, aligned validation, and human oversight supporting clinical credibility.

  • Problem

    Volumetric interpretation requires spatial, quantitative, and contextual evidence that may be lost when examinations are reduced to selected 2D images, compressed representations, or report-derived text.

  • Method

    The review organizes more than 200 publications around volumetric representations, multimodal models, agentic systems, clinical applications, evaluation, and Claim–Design–Validation alignment.

  • Results

    Native volumetric modeling is essential when interpretation depends on full-volume spatial relationships, while selected 2D views can suffice when decisive evidence is confined to limited images; agents add tools, memory, and iterative orchestration.

  • Takeaways & Limitations

    Clinical credibility depends on preserving task-relevant volumetric evidence, making system behavior traceable, aligning validation with claims, and defining human oversight.

  • Takeaways & Limitations

    The review closes its literature window in July 2026, and several rapidly developing agentic systems are preprints or early reports rather than evidence of deployment readiness.

Abstract

from arXiv · show

Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual representations, or report-derived text. Reliable volumetric radiology AI therefore requires representations that preserve task-relevant three-dimensional (3D) information and systems that can access, verify, and integrate this information across clinical workflows. In this Review, we examine more than 200 publications through July 2026. We organize the literature around volumetric representation and multimodal understanding at the model level, agentic orchestration at the system level, and their links to clinical applications and evaluation. We review volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction. We distinguish settings in which selected 2D views or report-mediated reasoning may suffice from those that warrant native volumetric modeling. We also introduce a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation. Across the literature, native volumetric modeling and agentic capabilities depend on the spatial, quantitative, contextual, and workflow requirements of the intended task. Clinical credibility requires faithful volumetric representation, traceable system behavior, claim-aligned validation, and clearly defined human oversight in realistic workflows.

1 Introduction

Volumetric radiology AI must preserve and access evidence distributed across 3D space, acquisition settings, and clinical context. This review connects model-level representations, agentic workflow systems, clinical applications, and claim-aligned evaluation.

  • Clinically relevant radiological evidence may depend on full-volume continuity, quantitative measurements, prior examinations, and clinical context.
  • Most current MLLMs use selected 2D images, compressed representations, or report-derived text, which may be insufficient when findings depend on the complete examination.
  • Agentic systems extend MLLM inference by acquiring evidence iteratively through tools, context retrieval, prior-examination comparison, memory, and human handoff.
  • The review links volumetric representation and multimodal interpretation with agentic orchestration, clinical applications, and evaluation requirements.
  • The synthesis examines when 2D or report-mediated reasoning suffices, when native volumetric modeling is necessary, and how systems preserve, retrieve, verify, and use spatial evidence.
  • The Claim–Design–Validation framework assesses whether stated capabilities are matched by technical designs and validation settings that support the intended conclusions.

2 Preliminaries

Volumetric radiology evolved from task-specific pipelines toward learned representations, language-centered interfaces, and agents that coordinate tools and clinical evidence. Across these stages, reliable image grounding remains essential because text-based processing can lose spatial and longitudinal information.

  • Traditional radiomics pipelines combined acquisition, reconstruction, registration, region delineation, handcrafted features, and downstream statistical or machine-learning models.
  • Handcrafted feature pipelines were sensitive to segmentation, scanner protocol, reconstruction, voxel spacing, standardization, and classifier choices, limiting robustness and transferability.
  • Supervised 3D networks learned volumetric features but generally remained task-specific, label-dependent, and unable to transfer capabilities automatically across tasks or institutions.
  • LLMs support language understanding, reporting, question answering, explanations, retrieval, and contextual reasoning rather than replacing image interpretation.
  • Reducing CT volumes to a few sentences or slices can lose subtle findings, longitudinal correspondence, and lesion-level information, motivating spatially grounded language interfaces.
  • Radiology agents plan steps, invoke tools, inspect image regions, preserve provenance, compare prior examinations, and expose intermediate observations for review.

3 Foundation Models in Volumetric Radiology

Volumetric foundation models provide reusable representations and interfaces for multiple downstream tasks, but their progression must preserve 3D anatomical and spatial information while making it accessible to clinical language.

  • Foundation models are pretrained on broad volumetric, visual, textual, or multimodal medical data and adapted across multiple downstream tasks.
  • Volumetric radiology models must acquire reliable 3D representations, connect them to clinical semantics, and expose visual information in a form that supports reasoning and communication.
  • The model landscape progresses functionally from volumetric self-supervised learning to vision–language alignment and then to multimodal large language models.

3.1 Volumetric Self-Supervised Representation Learning

Volumetric self-supervised learning uses unlabeled 3D scans and pretext objectives to learn transferable anatomical and spatial representations before linguistic supervision. The surveyed objectives exploit volumetric structure through masking, restoration, deformation, contrastive invariance, and related strategies.

  • Motivation: Volumetric SSL learns transferable anatomical representations from unlabeled radiology scans rather than relying on predefined task labels.This makes SSL particularly suitable for the foundation-model layer, while supervised and weakly supervised approaches remain important for task-specific systems.
  • Self-Supervised Objectives: 3D masked image modeling reconstructs missing voxels, encouraging long-range spatial modeling across the volume.The approach extends masked autoencoders to volumetric medical imaging.
  • Self-Supervised Objectives: Contrastive learning maximizes similarity between differently augmented views while distinguishing diverse anatomical structures in latent space.Positive and negative design requires care because crops can share tissues and abnormalities may be spatially sparse.
  • Role in the Foundation-Model Stack: Together, SSL frameworks establish the visual perception required for multimodal integration by acquiring anatomical and spatial priors before linguistic supervision.These objectives convert intrinsic 3D structure into reusable supervision for later alignment and MLLM integration.
  • Representative Strategies: The surveyed self-supervised signals restore disrupted anatomy, reconstruct masked or deformed structures, enforce geometric or contrastive invariance, and scale across dimensions, modalities, and datasets.Representative designs include handcrafted transformations, foreground-aware masking, and hierarchical reconstruction across structural scales.

3.2 Vision–Language Alignment

Vision–language pre-training maps volumetric visual representations into a clinically addressable semantic space. The literature progresses from global volume–report matching toward alignment that preserves lesions, anatomy, report hierarchy, uncertainty, and medical knowledge.

  • Purpose: Vision–language pre-training acts as a semantic bridge that makes radiological visual features addressable through clinical language.It converts visual features into a language-addressable representation space rather than serving merely as another generic training strategy.
  • Alignment Objectives: CLIP-style alignment jointly optimizes vision and text encoders so paired images or volumes and reports become mutually retrievable.The approach uses paired clinical text and symmetric image-to-text and text-to-image contrastive objectives.
  • Alignment Objectives: The image-to-text direction treats each visual study as a query for its paired report, while the text-to-image direction retrieves the paired study from batch candidates.The final CLIP-style objective averages the two directional losses.
  • Representative Approaches: The literature organizes alignment methods by correspondence level: global, fine-grained or lesion-aware, and knowledge-enhanced or hierarchical.These categories model increasingly explicit spatial correspondence and clinical semantics, including anatomy, lesions, negation, uncertainty, synonyms, and report hierarchy.
  • Representative Approaches: 3D methods scale volume–report alignment, add spatial or anatomy-level constraints, and introduce hierarchical or knowledge-enhanced objectives to address dense volumes and clinical language.Examples include large 3D volume–report corpora, weak axial-depth anchors, multi-level semantic losses, and soft-weighted clinically structured similarity targets.
  • Downstream Role: The resulting text-aligned feature spaces provide the semantic substrate for later MLLMs to support language-mediated volumetric interpretation.The field is moving from global matching toward representations that preserve disease concepts, anatomical regions, lesions, and report hierarchy.

3.3 Multimodal Large Language Models

Volumetric radiology MLLMs integrate 3D visual representations with generative language models for interactive interpretation. The central design challenge is preserving volumetric information while making visual tokens usable by the language model.

  • Model capabilities: Volumetric radiology MLLMs extend interpretation beyond retrieval and classification to report generation, multi-turn question answering, localization, uncertainty communication, and instruction following.The section frames this expansion as an integration problem built on self-supervised learning backbones and vision–language pre-training.
  • Model comparison: Table 1 compares representative medical MLLMs by input dimensionality, visual backbone, generative backbone, parameter scale, training recipe, and vision–language projector.Its input-dimensionality grouping describes how images or volumes are presented to the language model, rather than providing a complete encoder taxonomy.
  • Vision encoders: Visual encoder dimensionality determines the visual tokens available to the language model and whether through-plane relations are preserved.The main alternatives are 2D, native 3D, and hybrid 2D/3D encoders, with different semantic and anatomical pre-training priors.
  • Vision encoders: 2D-native systems provide transferable semantic priors but may weaken continuity, lesion extent, and volumetric morphology when studies are reduced to selected views or slice sequences.Native 3D models instead accept the volume as the primary visual object; CT2Rep conditions report generation on full-volume representations rather than key slices.
  • Vision encoders: Later systems connect volumetric self-supervised or vision–language priors to generative MLLMs through 3D spatial aggregation, disease-aware attention, and prototype memory.Examples include E3D-GPT, Dia-LLaMA, and other systems that transfer or specialize 3D representations for generation.
  • Vision encoders: Domain-specialized 3D systems indicate that suitable visual encoders depend strongly on imaging modality and anatomy.Brain-focused systems illustrate specialization for brain CT or MRI settings.

3.3.2 Vision-Language Interfaces: Projection, Resampling, and Fusion

Vision-language interfaces determine how dense volumetric information is selected, compressed, ordered, and grounded before reaching the generative backbone. The literature spans lightweight projectors, query-based compression, structured fusion, region-aware extraction, and routing mechanisms, with training increasingly organized around volumetric preservation and spatial grounding.

  • Interface role: Vision-language interfaces are the bottleneck between dense volumetric information and the generative token stream, deciding how visual tokens are selected, compressed, ordered, grounded, and inserted.Their role extends beyond dimensional matching to clinically relevant information preservation.
  • Projection and resampling: Lightweight 2D projectors provide alignment baselines, but they do not solve the harder problem of selecting and preserving clinically relevant 3D information.MLP and transformer-decoder projections, as well as perceiver resampling with gated cross-attention, mainly address vision-language alignment.
  • Projection and resampling: Query-based compression converts variable-size 2D/3D scans into fixed visual prefixes while reducing volumetric tokens through learned latent queries.RadFM combines a shared 3D ViT with a perceiver and learned latent queries for this purpose.
  • Fusion and region awareness: Hybrid and region-aware interfaces fuse 2D/3D streams or align texture, geometry, masks, and global context with region-level descriptions.These designs make the interface a fusion or extraction module rather than a single projection.
  • Routing and scheduling: Recent interfaces act as mixers, schedulers, or routers, using multi-scale mixing or prompt-conditioned mixture-of-experts projection to select information across spatial and modality levels.These mechanisms target fine spatial, semantic, and multiparametric information transfer.
  • Training and generation: Generative backbones must be assessed within the complete encoder-interface-generation pipeline because clinical value depends on spatial grounding, visual-token delivery, and the training objective.Training curricula increasingly reorganize pretraining and instruction tuning around volumetric information preservation and spatial grounding.

4 Agentic Systems in Volumetric Radiology

Agentic systems extend volumetric radiology MLLMs from fixed-context prediction to iterative workflow interaction. Their four linked capability families select evidence, perform grounded operations, retain context, and coordinate outputs with procedures and human review.

  • System architecture: Agentic systems add iterative evidence acquisition, planning, tool use, and context management to single-pass volumetric MLLM inference.They may augment perception directly or use language models as controllers that acquire volumetric information indirectly.
  • Functional taxonomy: The functional taxonomy comprises reasoning and planning, tool-augmented perception, memory and dynamic context management, and workflow interaction or multi-agent collaboration.These families correspond to the operations required for distributed volumetric analysis.
  • Reasoning and planning: Reasoning and planning select anatomy, slices, series, or phases, determine whether more evidence or a tool call is needed, and synthesize observations into study-level conclusions.Policy-constrained acquisition can request another phase when current evidence is judged insufficient.
  • Tool-augmented perception: Tool-augmented perception converts clinical intent into explicit spatial or quantitative operations such as segmentation, registration, measurement, retrieval, SUV conversion, and dose simulation.Traceable tool outputs connect image data to grounded clinical claims.
  • Memory and context: Memory integrates evidence across slices, series, analysis steps, retrieved cases, and longitudinal examinations while managing limited active context.The review distinguishes working, longitudinal patient, and retrieval-augmented memory.
  • Workflow interaction: Workflow interaction coordinates models, tools, intermediate artifacts, and human review across reporting, dosimetry, treatment planning, and other multi-step procedures.Reported evaluations include report quality, expert preference, dose agreement, processing time, plan quality, constraint satisfaction, and reproducible logs, while few studies assess tool selection or failure detection.
  • Integrated workflow: Together, the four capability families form an iterative loop linking evidence selection, grounded outputs, cross-step memory, clinical coordination, and human oversight.This extends or complements volumetric MLLMs at the workflow level.

5 Clinical Applications and Evaluation in Volumetric Radiology

Clinical applications should be organized by the claim a volumetric radiology system makes, with validation matched to the output, evidence requirements, and operational setting. Across application families, technical feasibility is more common than external robustness, workflow benefit, or integrated validation of error propagation and recovery.

  • Clinical evaluation should begin with the intended claim and match validation to the type of output and strength of the claim.
  • Internal benchmarks establish technical feasibility, external studies assess robustness, reader or workflow studies examine usability, and prospective evaluation supports deployment-oriented claims.
  • Diagnostic interpretation and reporting: Diagnostic interpretation requires joint assessment of spatial predictions, diagnostic labels, and language outputs, because textual similarity alone cannot establish volumetric grounding or clinical correctness.
  • Diagnostic interpretation and reporting: Current diagnostic resources broaden evaluation beyond text similarity, but calibration, reader or workflow studies, and external validation across institutions and acquisition settings remain limited.
  • Cross-task evaluation and validation: The principal cross-task gap is limited evaluation of whether upstream failures are detected, propagated, or corrected before affecting downstream decisions.
  • Validation requirements differ by application: prognostic claims need temporally defined cohort evidence, planning claims need end-to-end technical and workflow validation, and cross-task claims need coherent outputs and intermediate actions.

6 Discussion and Future Directions

The field is progressing from 2D-adapted task models toward native 3D foundations, multimodal interpretation, and agentic workflows, while remaining heavily CT-centric. Future directions emphasize acquisition-aware representations, workflow-centered human-AI collaboration, embodied and self-evolving systems, and traceable clinical-grade validation.

  • Volumetric radiology AI follows a progression from 2D adaptation to native 3D foundation models and ultimately agentic systems.
  • After 2024, 3D and hybrid MLLMs expanded rapidly, while advanced agentic workflow systems became concentrated in 2025–2026.
  • The literature remains heavily CT-centric, although sequence-aware MRI, PET or nuclear-medicine, and multimodality resources are diversifying the field.
  • Research agenda: The proposed agenda couples native volumetric representations with scanning physics, collaborative human-AI workflows, world models, self-evolving technologies, embodied infrastructure, and claim-specific validation.
  • Volumetric and acquisition-aware intelligence: Volumetric models should preserve native spatial continuity, compress data according to clinical salience, and represent acquisition context alongside anatomy.
  • Workflow-centric clinical intelligence: Future systems are expected to move from static perception toward token-level spatial grounding, multi-step orchestration, and bounded human-AI workflows.
  • Scope: The review focuses on 3D volumetric radiology and does not comprehensively cover 2D medical imaging, non-imaging medical language models, or legal, economic, and reimbursement issues.

7 Conclusion

Volumetric radiology AI is advancing from 2D-adapted models toward native volumetric representations and agentic workflows. Clinical credibility depends on preserving relevant evidence, tracing system actions, aligning validation with claims, and defining human oversight.

  • Native volumetric modeling becomes essential when interpretation depends on full-volume spatial relationships, whereas selected 2D views can suffice when decisive evidence is confined to limited images.
  • Agentic systems couple volumetric perception with tools, memory, and iterative workflow orchestration to acquire and verify evidence beyond a single forward pass.
  • Clinical credibility requires alignment among the intended claim, preserved evidence, system actions, and supporting validation, with traceability and clearly defined human oversight.
Loading 2608.20549v1…