Source-linked AI summary

The Role of Mixed and Augmented Reality in Medical Visualization: Literature Review and A Context-Aware Taxonomy

Xinrui Zou, Mingxu Liu, Hongchao Shu, Ruixing Liang, Mathias Unberath, Alejandro Martin-Gomez

arXiv:2608.27644v1cs.GR

TL;DR

Medical imaging is difficult to fully interpret when 3D anatomy is presented through conventional 2D displays, and AR/MR implementations can introduce perceptual inconsistencies. The paper conducts a systematic literature review and proposes a six-dimensional, context-aware taxonomy; it identifies recurring design trends and persistent challenges in modality integration, perceptual design, and soft-tissue adaptability.

  • Problem

    Conventional 2D displays make clinicians mentally reconstruct spatial relationships in 3D medical data, while AR/MR requires careful design to avoid perceptual inconsistencies.

  • Method

    The paper conducts a systematic literature review and organizes medical AR/MR visualization systems into six dimensions linking design choices with clinical tasks.

  • Results

    The review identifies common trends and gaps, including persistent challenges in modality integration, perceptual design, and adapting visualization to soft-tissue dynamics.

  • Takeaways & Limitations

    The taxonomy supports transparent reporting, cross-context comparison, and identification of underexplored visualization combinations in medical AR/MR.

  • Takeaways & Limitations

    The taxonomy is descriptive rather than evaluative and does not establish which visualization choices are inherently more effective.

Abstract

from arXiv · show

The discovery and evolution of medical imaging technologies have enabled non-invasive visualization of internal anatomy that has be- come essential for supporting diagnosis, monitoring, and treatment. However, because medical imaging relies on complex physical processes and contrast mechanisms for image formation, imaging alone is not sufficient to enable humans to leverage the resulting in- formation fully. In addition, traditional methods to visualize the resulting information use two-dimensional displays to present three-dimensional anatomical structures. The introduction of Augmented and Mixed Reality (AR/MR) technologies offers an opportunity to provide valuable paradigms for medical imaging visualization, allowing users to observe, explore, and interact with anatomical infor- mation in more spatially intuitive ways. However, naive implementation without careful design considerations can lead to perceptual inconsistencies, potentially compromising utility and effectiveness. In this paper, we present a structured taxonomy of medical AR/MR visualization strategies aimed at providing clearer insight into how visualization design varies across clinical use cases. The taxonomy organizes techniques based on four core design components: image modality, data dimensionality, display technology, and clinical application. In addition, we introduce two critical dimensions that are often overlooked in the literature: visualization anchoring (the spatial relationship between virtual content and the physical world), and perceptual awareness (the use of visual cues to support spatial interpretation). Together, these components form a comprehensive taxonomy, offering a detailed framework for selecting appropriate visualization techniques in medical applications.

1 INTRODUCTION

Medical imaging enables non-invasive visualization of internal anatomy, but conventional 2D presentation makes 3D spatial relationships difficult to interpret. The review proposes a context-aware AR/MR taxonomy to organize visualization strategies and support clinical integration.

  • Medical image visualization transforms imaging data into visual representations that support anatomical interpretation, diagnosis, planning, and image-guided procedures.
  • Although CT and MRI provide volumetric 3D datasets, clinicians have traditionally viewed medical data through 2D films and monitors.
  • Reliance on 2D displays requires clinicians to mentally reconstruct spatial relationships among anatomical structures.
  • AR/MR integrates virtual content with the physical environment, offering more spatially aligned visualization than isolated immersive environments for anatomy-dependent procedures.
  • The proposed taxonomy classifies medical AR/MR visualization by modality, dimensionality, display technology, clinical context, anchoring, and perceptual awareness.
  • Visualization anchoring describes virtual content’s spatial relationship to the real world, while perceptual awareness uses visual cues to support spatial interpretation.

2 RELATED WORK

Existing reviews commonly classify medical AR/MR by clinical usage or general system design, but provide limited systematic analysis of how clinical objectives shape visualization choices. This paper addresses that gap by linking design dimensions to clinical context.

  • Few studies systematically examine how clinical objectives influence spatial integration and perceptual support in AR/MR visualization.
  • Prior reviews emphasize applications across specialties and tasks, including orthopedics, plastic surgery, education, and surgical guidance.
  • Existing reviews often prioritize system performance, feasibility, and workflow integration over visualization design.
  • The taxonomy links image modality, dimensionality, and display technology to medical use cases while incorporating anchoring and perceptual awareness.

3 DATA COLLECTION

The review used a systematic search of peer-reviewed studies across major medical and technical databases, followed by independent screening, eligibility assessment, and structured extraction by multiple authors.

  • The search covered peer-reviewed original studies in PubMed, IEEE Xplore, and Scopus that described medical imaging representation, rendering, or AR/MR integration.
  • Search reporting followed applicable PRISMA 2020 items, with database-specific strings and filters documented in supplementary material.
  • The search included studies published before February 5, 2026, while filters excluded reviews and meta-analyses where available.
  • Two reviewers independently screened titles and abstracts before full-text eligibility assessment and exclusion of studies outside the review scope.
  • Four authors contributed to full-text extraction, with each study reviewed by two authors and disagreements resolved through discussion.

4 TAXONOMY

The taxonomy provides a structured framework for analyzing medical AR/MR visualization across core data and display choices, clinical context, spatial anchoring, and perceptual support. It is intended to guide context-aware design and technique selection.

  • The taxonomy organizes medical AR/MR visualization systems across application, image modality, dimensionality, display technology, anchoring, and perceptual support.
  • The four core dimensions specify what data is used, how it is visualized, where it is displayed, and the clinical context.
  • Visualization Anchoring: Visualization anchoring specifies whether virtual content is placed in-situ, inter-situ, or off-situ relative to the physical environment.
  • Perceptual Awareness: Perceptual awareness addresses depth perception, spatial alignment, scene understanding, and visual consistency through visual-perception cues.

5 IMAGE MODALITY

Clinical context determines which imaging modality is presented through AR/MR. CT and X-ray commonly depict rigid bone, MRI supports soft tissue, ultrasound updates deformable or moving anatomy, and multimodal imaging combines structures.

  • Clinical context determines the imaging modality selected for AR/MR visualization.
  • CT and X-ray are commonly used to visualize rigid bony anatomy.
  • MRI is commonly used when visualizing soft-tissue anatomy.
  • Ultrasound can provide intraoperative updates when anatomy deforms or moves.
  • Multimodal combinations can visualize both rigid and deformable structures.
  • AR/MR can place two-dimensional radiographic projections into three-dimensional spatial context during bony-anatomy procedures.This can reduce mental effort and improve scene understanding and coordination during guidance.

6 DIMENSIONALITY

Dimensionality captures the spatial and temporal properties of imaging data and the representation strategies used in AR/MR. It constrains rendering and registration choices, while 4D adoption remains limited by acquisition, processing, and workflow demands.

  • Dimensionality describes the spatial and temporal dimensions of imaging data and its AR/MR representation.
  • Dimensionality shapes whether systems use slices or projections, segmented surfaces, or volume rendering.
  • 4D imaging can represent time-varying anatomy, including respiratory motion and cardiac cycles.
  • Broader 4D adoption is limited by acquisition burden, real-time processing demands, and workflow integration overhead.
  • Two-dimensional images serve as direct anatomical references in AR/MR, including X-rays, 2D ultrasound, and CT/MRI slices.

7 DISPLAY TECHNOLOGY

Display technology affects how augmented information is perceived and integrated into clinical workflows. Device selection reflects sterility, hands-free operation, operative-field visibility, team viewing, latency, portability, and deployment cost.

  • Display technology shapes both perception of augmented information and its fit within clinical workflows.
  • Procedural device selection is influenced by sterility, hands-free operation, line-of-sight, team-shared viewing, and latency tolerance.
  • Hand-held devices prioritize portability and low deployment cost for point-of-care viewing and education.
  • Hand-held devices require manual holding and interaction, which conflicts with hands-free and sterile requirements in many intraoperative workflows.
  • Spatial displays use fixed displays that can be viewed by the clinical team.

8 VISUALIZATION ANCHORING

Visualization anchoring classifies how imaging-derived content is positioned relative to the patient and operative scene. The taxonomy distinguishes off-situ, inter-situ, and in-situ displays, which differ in spatial registration, alignment demands, obstruction, and workflow risks.

  • 8 VISUALIZATION ANCHORING: The taxonomy defines off-situ, inter-situ, and in-situ anchoring according to increasing spatial relationship with the patient and operative scene.
  • 8.1 Off-situ: Off-situ content is separate from its physical source and requires users to mentally map it back to the patient.External monitors, floating 3D models, and fixed-screen virtual panels are examples.
  • 8.1 Off-situ: Off-situ visualization maintains workflow compatibility but increases cognitive load through mental alignment.
  • 8.2 Inter-situ: Inter-situ visualization presents augmented content outside the immediate physical space while retaining a visual reference to real-world objects.
  • 8.2 Inter-situ: Inter-situ methods can reduce visual obstruction while preserving spatial alignment.Mirror-based approaches and duplicated AR are examples.
  • 8.3 In-situ: In-situ visualization overlays virtual content directly on anatomy, supporting real-time spatial alignment and reducing hand–eye coordination demands.
  • 8.3 In-situ: In-situ effectiveness depends on accurate registration and consistent handling of occlusion and depth cues.Misalignment can cause incorrect spatial interpretation, while poor cue integration can make content appear detached from the body.
  • 8 VISUALIZATION ANCHORING: More than 50% of surveyed systems employed in-situ visualization, whereas approximately 78% of reviewed medical-education literature adopted off-situ visualization.

9 PERCEPTUAL AWARENESS

Perceptual awareness in medical AR/MR addresses how users interpret spatial relationships between virtual and real objects, not merely whether content is registered. The taxonomy groups visual support into pictorial, kinetic, and illustrative cues.

  • Accurate registration alone is insufficient when missing occlusion or weak depth cues can cause misjudgments of anatomy or tool placement.
  • Pictorial cues: Occlusion conveys depth by making partially blocked objects appear behind the occluding object, helping locate hidden anatomical structures.Improper occlusion can make correctly registered internal content appear outside the body.
  • Pictorial cues: Texture gradients encode depth through reduced texture or thinner vessel strokes for farther structures, but texture alone has limited perceptual impact.Combining texture with chromadepth may enhance spatial perception.
  • Pictorial cues: Illumination effects and reflections provide image-based depth and shape information, with augmented mirrors offering alternative views of anatomical overlays.
  • Kinetic cues: Kinetic cues use motion and viewpoint changes to provide additional spatial information, including observer-driven focus modulation and dynamic depth representations.
  • Illustrative cues: Illustrative cues intentionally encode spatial relationships through color mapping or numeric annotations, including chromadepth and direct depth measurement.These encodings may depend on user familiarity with the scheme.

10 APPLICATION

The review organizes AR/MR visualization studies by clinical task and, for intraoperative systems, by surgical area. Across applications, visualization choices reflect workflow demands, anatomy, imaging needs, and spatial integration constraints.

  • Studies are first categorized by clinical task because workflow demands determine AR/MR use and visualization approaches, then intraoperative systems are stratified by specialty.
  • Education: Educational systems emphasize intuitive exploration of pre-generated CT or MRI models rather than real-time imaging or precise anatomical alignment.
  • Education: AR/MR education grounds anatomy learning in physical bodies, manikins, or clinical tools, supporting spatial reasoning and communication.
  • Training and Simulation: Training and simulation systems remain scarce but support interaction with physical instruments, anatomical models, and task environments.
  • Training and Simulation: Procedural rehearsal uses realistic operational constraints to train task steps, tool coordination, line-of-sight management, and interpretation of imperfectly aligned overlays.Anchoring and registration become part of the skill being trained, while inter-situ setups can support progressive learning.
  • Preoperative Planning: Preoperative planning commonly reconstructs CT or MRI into rotatable, measurable, and annotatable 3D surface models; more than half use off-situ integration.
  • Preoperative Planning: Planning systems support inspection, measurement, and comparison, but explicit attention to perceptual awareness remains limited.
  • Intraoperative Use: Most intraoperative systems use in-situ CT-based 3D surface overlays on monitors or OST-HMDs, making depth ambiguity, occlusion, and clutter important concerns.Only a minority report perceptual cues.

11 DISCUSSION AND FUTURE DIRECTIONS

The discussion identifies persistent underemphasis of perceptual and context-sensitive visualization design, alongside challenges from modality differences and deformable anatomy. It proposes descriptive, context-aware heuristics to guide reflection rather than prescribe universal standards.

  • Visualization Underprioritized in System Design: Many systems use generic surface models and standard rendering, while few incorporate contour enhancement, occlusion handling, or focused views for spatial interpretation.
  • Visualization Underprioritized in System Design: Cue selection affects depth judgment and response time, while display and interaction design influence the mental effort needed to transform image information into action.
  • Visualization Underprioritized in System Design: The review observes that perceptual cues declined from 25.9% to 8.8% as OST-HMD use increased from 31.0% to 59.6% between 2016–2020 and 2021–2026.The authors caution that reporting bias, terminology differences, shifting priorities, or insufficient detail may contribute to this temporal pattern.
  • Deformable Tissue and Dynamic Anatomy: Deformable tissues can shift, stretch, or compress, causing augmented content to diverge from current anatomy and compromising accuracy, especially for organs such as the liver or lungs.
  • Deformable Tissue and Dynamic Anatomy: NeRF and 3DGS offer view-consistent anatomical reconstruction and support real-time or deformable modeling, but clinical integration remains limited by data, computation, calibration, and robustness requirements.
  • Context-Aware Design Heuristics: The taxonomy is descriptive rather than evaluative, so its observed associations do not establish which visualization choices are inherently more effective.
  • Context-Aware Design Heuristics: The proposed heuristics prompt designers to preserve modality-specific strengths, align display and interaction with workflow, and use perceptual cues when depth or registration ambiguity threatens performance.
  • Context-Aware Design Heuristics: The heuristics provide a scaffold for reflection, not definitive design benchmarks or universal standards.

12 CONCLUSION

The paper introduces a six-dimensional taxonomy of medical AR/MR visualization based on a systematic literature review. It identifies persistent challenges in modality integration, perceptual design, and adaptation to soft-tissue dynamics, while positioning the taxonomy as support for more visualization-centered development.

  • The taxonomy classifies medical imaging AR/MR systems by application, image modality, dimensionality, display technology, visualization anchoring, and perceptual support.
  • The review examines how these design choices relate to clinical tasks and identifies common trends and gaps.
  • Persistent challenges remain in modality integration, perceptual design, and adaptability to soft-tissue dynamics.
  • The taxonomy is intended to support transparent reporting, cross-context comparison, and identification of underexplored design combinations.
Loading 2608.27644v1…