Source-linked AI summary

Operational digital twin clinics enable task-based evaluation of embodied AI

Xinyuan Wu, Jingrao Zhang, Mengdi Xu, Henry K. Chu, Mingguang He, Danli Shi

arXiv:2608.21416v1cs.ROcs.AI

TL;DR

Embodied AI requires realistic, robot-testable clinical environments, but existing simulators and scalable site-specific digital-twin construction remain limited. This paper converts routine ophthalmic images into editable, simulator-ready digital twins and evaluates reconstruction, contact geometry, perturbations and policies. The results support operational validity as an intermediate layer between offline development and physical deployment.

  • Problem

    Embodied AI must be evaluated in variable clinical environments, while current simulators and digital-twin methods often lack robot-relevant geometry or are difficult to scale.

  • Method

    The study uses single-image 3DGS, local editing, simulator conversion, mesh grounding and task-specific anchors to evaluate 39 ophthalmic clinic scenes.

  • Results

    Routine images preserved workspace structure, while mesh grounding, perturbation testing and closed-loop evaluation revealed task- and embodiment-specific robot feasibility.

  • Takeaways & Limitations

    Operational digital twins shift clinical scene reconstruction from visual realism alone toward task-centred robot-facing evaluation before physical testing.

  • Takeaways & Limitations

    The evaluation covered a limited set of ophthalmic and optometric scenes, and temporal variation, human-robot interaction and force-dependent device operation remained outside scope.

Abstract

from arXiv · show

Embodied artificial intelligence (AI) must be tested in the clinical environments where it will operate, but building realistic, robot-testable settings is costly and difficult to scale. Here we show that routine clinic images can be transformed into operational digital twins for task-based evaluation of embodied AI. Using 39 ophthalmic clinic scenes, we converted single photographs into editable, simulator-ready environments and assessed reconstruction quality, room-scale geometry, mesh grounding, multi-robot feasibility, perturbation sensitivity and closed-loop policy performance. The reconstructed scenes preserved workspace structure, while local editing enabled controlled device reconfiguration. Device meshes, collision proxies and semantic anchors converted visual reconstructions into contact-aware simulation scenes. Across three robot embodiments, shared task targets showed different patterns of reachability and contact feasibility. Small device translations and rotations produced task-specific changes in contact margins that were not captured by visual similarity alone. Digital-twin trajectories also supported local policy learning and closed-loop evaluation. These findings establish operational validity as a key principle for clinical digital twins and provide an intermediate layer between offline development and physical deployment of embodied AI in healthcare.

Introduction

Clinical embodied AI needs robot-testable environments that preserve task-relevant clinic structure, but existing simulators and digital-twin construction methods remain limited. The study proposes operational digital twins from routine ophthalmic images and evaluates their use for staged robot-facing assessment.

  • Introduction: Outpatient clinical robots must navigate variable rooms, equipment, approach paths and task-specific contact surfaces near patients and staff.Existing diagnostic evaluation on images or records does not capture these physical interaction requirements.
  • Introduction: Current clinical simulators often lack scene geometry needed for localization, motion planning, contact reasoning and policy assessment.Most are designed for human training, procedural rehearsal or visual demonstration rather than robot learning and evaluation.
  • Introduction: Site-specific healthcare digital twins remain difficult to scale because manual modelling is time-consuming, while multi-view scanning can disrupt workflow and raise privacy concerns.A practical outpatient digital twin should be inexpensive, preserve context, permit local reconfiguration and provide robot-queryable geometry.
  • Introduction: Visual realism alone is insufficient: operational usefulness depends on preserving spatial relationships, support surfaces and device geometry required for defined tasks.The paper frames robot-facing evaluation, rather than visual perfection, as the key criterion for image-derived scenes.
  • Introduction: Across 39 scenes, the workflow generated neural scene representations, simulator environments and clinician-supervised evaluations spanning editing, contact testing, perturbations and policy assessment.The workflow uses routine clinic images and stages evaluation before physical deployment.

Image-derived clinic reconstruction

Single-image reconstructions preserved clinic appearance, room-scale geometry and dominant scene structure across 39 routine photographs. Local editing enabled targeted device changes while largely preserving the surrounding clinic context.

  • Image-derived clinic reconstruction: 39 routine clinic photographs were reconstructed with single-image 3DGS, preserving room layout, equipment silhouettes, tabletop devices and coarse depth ordering.The scenes included examination rooms, waiting areas and corridor-like spaces.
  • Image-derived clinic reconstruction: 0.908 ± 0.005 mean SSIM, 0.034 ± 0.001 LPIPS, 0.969 ± 0.006 CLIP similarity and 31.99 ± 0.30 PSNR characterized high observed-view reconstruction quality.Circular render trajectories were generally stable, with greater variation mainly in highly occluded views.
  • Image-derived clinic reconstruction: 8.54 ± 0.24 mm ground-plane RMSE and 13.77 ± 1.23° dominant-surface normal error indicated retained room-scale geometry.LiDAR-referenced depth agreement yielded AbsRel 0.0869 ± 0.0063 and δ<1.25 0.930 ± 0.010, with residuals concentrated near thin structures and occlusion boundaries.
  • Image-derived clinic reconstruction: 61.5% of scenes met all three prespecified conversion criteria: visual similarity, cyclic alignment and structural preservation.Mean symmetric Chamfer similarity was 0.773 ± 0.013, cyclic alignment similarity 0.718 ± 0.013 and structural preservation 0.878 ± 0.006.
  • Editable clinic configurations: 32 scenes underwent image-space editing that removed or replaced visible device regions while preserving surrounding rooms and non-target equipment.The editing workflow was designed for local clinical reconfiguration before reconstruction.
  • Editable clinic configurations: 84.4% of edited scenes stayed within the prespecified 0.15 scene-quality preservation boundary.Unmasked-region SSIM was 0.977 versus masked-region SSIM 0.526, while depth renderings remained stable outside edited regions.

Mesh-grounded contact geometry

Mesh grounding supplemented reconstructed clinic context with device geometry, collision proxies and semantic anchors for contact-aware robot evaluation. The resulting scenes supported reachability and contact assessment without substantially obscuring the environment.

  • Mesh-grounded contact geometry: Clinical-device meshes, conservative collision proxies and semantic task anchors added task-relevant interaction geometry to selected reconstructed rooms.The reconstructed scene retained site-specific context, while meshes supplied device geometry for screen and handle contact-proxy tasks.
  • Mesh-grounded contact geometry: 4.76 ± 0.36% average visible collider occupancy across seven mesh-grounded scenes remained below 7% in every scene.Collision geometry was incorporated without obscuring the reconstructed environment.
  • Mesh-grounded contact geometry: 52 of 58 filtered reachability traces formed the dominant interaction cluster.The cluster provided the basis for contact-surface characterization and operational scene assessment.
  • Mesh-grounded contact geometry: The dominant contact surface had 47.1 mm in-plane spread, 5.7 mm normal-direction spread and 9.6 mm plane-fit RMSE.The normal-to-in-plane dispersion ratio was 12.2%, supporting access and contact assessment.
  • Mesh-grounded contact geometry: Mesh grounding converted image-derived scenes into operational environments that could support robot access and contact assessment.The figure describes this step as linking reconstructed context to robot-queryable geometry for contact-proxy evaluation.

Embodiment-specific feasibility

Shared clinical targets produced embodiment- and scene-specific contact feasibility, while small device-pose changes shifted task success boundaries. Digital-twin trajectories also supported policy learning and closed-loop simulation evaluation.

  • Embodiment-specific feasibility: SO101 achieved 99.3% handle-contact and 98.3% screen-contact feasibility, compared with lower rates for Kinova Gen3 and Franka Panda.Kinova Gen3 achieved 79.0% and 94.2%, while Franka Panda achieved 85.7% and 93.7%, respectively.
  • Embodiment-specific feasibility: Feasibility varied across scene–robot combinations, with Gen3 handle feasibility reaching 58.0% in scene 1 and Franka handle feasibility reaching 75.0% in scene 3.Gen3 screen feasibility was 84.0% in scene 2.
  • Contact-margin diagnostics: Among failed trials, mesh not reached accounted for 59.2%, mesh out of reach for 21.5% and safety-gated penetration for 15.1%.These failure modes varied with robot embodiment, base placement, end-effector geometry, reach envelope and collision constraints.
  • Perturbation sensitivity: Screen-contact success was more configuration sensitive than handle contact under ±20 mm translations and ±5° yaw rotations.Franka Panda screen success fell from 100% in most conditions to 0% under +20 mm local-y translation, whereas handle success stayed at or above 75%.
  • Policy learning and evaluation: Digital-twin trajectories supported joint-only, target-conditioned, diffusion and residual behaviour-cloning policies for screen-touch and handle-contact tasks.The workflow included trajectory collection, offline rollout and closed-loop simulation evaluation.
  • Policy learning and evaluation: Target-conditioned behaviour cloning reduced one-step action MAE relative to joint-only behaviour cloning for screen contact from 2.85 to 1.92 mm and for handle contact from 1.80 to 1.58 mm.Residual behaviour cloning achieved the lowest reported offline screen final error at 33.3 mm and a handle final error of 3.7 mm.
  • Policy learning and evaluation: Closed-loop simulator success reached 73.1% for screen touch and 69.2% for handle contact.Errors decreased during rollout, although failures remained associated with target non-reachability and incorrect screen-contact locations.

Discussion

The study positions operational digital twins as a task-centred extension of visual clinic reconstruction. Its layered representation supports interaction-critical simulation while retaining important scope limits around generalization and real-world autonomy.

  • Discussion: Routine clinic images can become operational digital twins when combined with simulator conversion, mesh grounding and task-specific anchors.This extends single-image reconstruction beyond visual scene representation toward evaluation near patients, staff and equipment.
  • Discussion: The paper shifts emphasis from visual realism alone to task-centred operational validity.The relevant criterion is whether reconstructed scenes support robot-facing evaluation.
  • Discussion: Unlike many existing clinical simulators designed for training or visual demonstration, this framework targets robot learning and evaluation in equipment-dense outpatient environments.Existing simulators often lack scene-level geometry for localization, motion planning, contact reasoning and policy assessment.
  • Discussion: Across matched scenes, shared contact anchors produced different feasibility margins across robot embodiments, and screen contact was more sensitive to device-pose perturbations than handle contact.These patterns connect the discussion to the measured embodiment- and task-specific boundaries.
  • Discussion: A hybrid representation preserves reconstructed room context while concentrating explicit interaction geometry around task-relevant regions.Device meshes, colliders, articulation structures and semantic anchors provide the interaction-critical layer.
  • Discussion: Local image-space editing enables matched variants of the same clinic context while preserving surrounding room and non-target equipment.This supports controlled variation in target-device regions for outpatient layouts and equipment configurations.
  • Limitations: The evaluation covered a limited set of ophthalmic and optometric scenes, so generalizability across clinical sites remains to be assessed.The framework also excludes temporal variation, human-robot interaction and force-dependent device operation.
  • Limitations: The findings establish operational feasibility within the tested conditions but do not determine autonomous performance, safety or clinical utility in routine care.The conclusion frames the workflow as an intermediate layer rather than evidence of routine-care deployment.

Methods

The methods used a staged single-image reconstruction workflow to create editable scene representations and export them for visual, geometric and simulator-based analyses.

  • Methods: The study evaluated single-image scene reconstruction, local reconfiguration, contact-aware task definition, multi-arm feasibility and policy evaluation.The workflow was organized as a staged evaluation of operational digital twins for outpatient ophthalmic clinics.
  • Dataset: The dataset contained 39 RGB routine photographs from real-world outpatient ophthalmic clinic environments.Scenes included examination rooms, waiting areas, treatment bays and more open clinic layouts.
  • Scene reconstruction: Each source photograph was reconstructed with a single-image 3DGS pipeline to generate an editable scene representation.The representation was treated as an operational scene substrate rather than a complete architectural model.
  • Scene reconstruction: Reconstructions were exported for observed-view rendering, ordered orbit rendering and downstream simulator conversion.These outputs supplied a shared visual and geometric basis for fidelity, conversion, editing and robot-facing analyses.

Image-space scene editing

Image-space editing removed selected ophthalmic instruments and reconstructed the edited scenes under a matched workflow. Quality metrics indicated localized edits with substantial preservation outside the target region.

  • Image-space scene editing: Instrument-removal masks were generated for device-containing scenes using image-guided segmentation and AI-assisted, scene-specific mask generation.Inpainting approximated the local background before reconstruction was rerun.
  • Image-space scene editing: Edited images were processed with the same reconstruction pipeline as unedited source images, producing paired raw and edited reconstructions.The matched workflow enabled comparison of scene preservation after local device removal.
  • Simulator conversion: Edited scenes were converted into simulator-compatible assets with clinical-device meshes, conservative collision proxies, pivots and semantic task anchors.The simulator environment combined reconstructed site-specific context with task-relevant device geometry for contact-proxy tasks.

Multi-view coherence assessment

The study evaluates reconstructed clinic scenes through trajectory-based visual stability, conversion fidelity, editing preservation, and geometry of interaction regions. These measures distinguish viewpoint consistency, structural preservation, and task-relevant scene properties.

  • Cross-view stability: Adjacent-view SSIM, PSNR, LPIPS and DISTS measured structural and perceptual stability along ordered circular render trajectories.Consecutive frames came from the same orbit and camera-angle progression, isolating viewpoint change within each trajectory.
  • Conversion fidelity: Three descriptive scores quantified simulator-conversion fidelity: symmetric Chamfer similarity, cyclic alignment similarity and structural preservation.The first two assess cross-format visual correspondence and circular trajectory alignment, respectively.
  • Conversion fidelity: 0.72, 0.65 and 0.85 were the prespecified cutoffs for visual similarity, cyclic alignment and structural preservation, respectively.The structural-preservation cutoff corresponds to a smoothness gap of ≤ 0.15 and was used for cohort stratification.
  • Editing preservation: Editing preservation separated intended target-region change from preservation of the surrounding scene.Masked-region scores described edited-area recovery and residual artifacts, while unmasked-region scores reflected surrounding-scene preservation.
  • Operational scene grounding: Clinical-device meshes were scaled, aligned, and assigned conservative collision geometry, pivots and task anchors after simulator instantiation.These representations supported stable simulator interaction and contact-proxy queries.
  • Operational scene grounding: PCA summarized dominant interaction-manifold geometry using dispersion, plane-fit RMSE and normal-to-in-plane dispersion ratio.The measures characterize spread around contact centroids, deviation from the task plane and planar concentration.

Multi-arm contact-feasibility probing

Multi-arm inverse-kinematics probing queried shared task intent across three robot embodiments in mesh-integrated clinic scenes. The evaluation combined contact-proxy endpoints, task-specific probe metrics, perturbations, and local policy formulations.

  • Multi-arm probing: SO101, Kinova Gen3 and Franka Panda were queried across matched operational scene variants using robot-specific kinematic adapters and constraints.Each adapter specified active joints, limits, end-effector frame and base-placement convention.
  • Endpoint definitions: Task-specific contact-proxy success was the primary IK endpoint, while final positional error served as a secondary anchor-deviation diagnostic.Screen and handle probing used signed residual, in-plane error or mesh-proximity readouts against configured contact margins.
  • Perturbation assay: Seven instrument-pose conditions tested sensitivity to baseline, ±20 mm local x and y translations, and ±5° yaw rotations.Translations and rotations were applied in instrument-local coordinates around the instrument root.
  • Perturbation assay: 20 jittered trials were used for each robot-task-condition after resetting the instrument and recomputing screen and handle anchors.Independent conditions therefore began from the nominal pose.
  • Endpoint definitions: Strict task-native success was reported separately from task-proxy success to distinguish full simulator completion from contact-relevant geometric evidence.For screen contact, task-proxy success was the union of point, mesh-touch and center-touch success; handle contact used contact-proxy success.
  • Policy formulation: Four local policy families were evaluated: joint-only, target-conditioned, diffusion, and residual behavior cloning.They predicted or refined low-dimensional end-effector displacement commands in simulation rather than serving as general-purpose autonomous policies.

Offline and closed-loop rollout evaluation

Policy evaluation progressed from supervised prediction to offline rollout and closed-loop simulation, with task-specific success gates. Physical testing used a non-patient SO101 platform in limited replay and online-execution settings.

  • Simulation evaluation: Evaluation used one-step supervised validation, offline rollout and closed-loop simulation rollout.Offline rollout applied predictions to stored states without runtime rendering or contact stepping, whereas closed-loop rollout reset SO101 in the policy scene.
  • Success gates: 10 mm defined closed-loop screen point-alignment success, while handle success required bbox distance ≤ 5 mm and probe distance ≤ 35 mm.These task-specific gates evaluate local simulated contact behavior and are not equivalent physical error metrics.
  • Physical testing: Physical testing used a non-patient SO101 platform for clinic-side replay and simplified tabletop handle-policy execution.The planned cohort included three handle and three screen replay trajectories, each repeated five times, plus five handle-policy executions.
  • Physical testing: Online execution used live encoder-derived joint signals, runtime URDF forward kinematics, target-relative observations and numerical-Jacobian joint commands.Replay converted six commanded joint angles from radians to servo ticks before publication.
  • Analysis framing: Scene-impact preservation and conversion were analyzed descriptively using score differences, thresholded proportions and continuous scene-level metrics.The analysis goal was fidelity assessment rather than intervention-versus-control inference; statistics were reported as mean ± SEM unless otherwise stated.

Funding

The paper reports support from the RCSV seed fund Eye Robot for Autonomous Clinic: Prototype Development. It also records conflict-of-interest status and resource-access information.

  • Funding: D.S. and M.H. disclose review and publication support from the RCSV seed fund Eye Robot for Autonomous Clinic: Prototype Development.The listed grant identifier is P0057912 from PolyU.
  • Funding: The disclosed funding is identified as supporting review and publication of the work.The disclosure names D.S. and M.H. as the supported authors.
  • Disclosures: The authors declare no conflicts of interest.
Loading 2608.21416v1…