Source-linked AI summary
MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception
Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Aladağ, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker
TL;DR
Real surgical multi-view reconstruction lacks rigorous evaluation benchmarks because clinical telerobots provide only one stereo viewpoint and suitable synchronized data with accurate geometry and poses are scarce. MV-dVRK combines three exposure-synchronized stereo viewpoints with validated static references, dynamic sequences, and controlled sparse-view comparisons. Registered stereo is strongest with two endoscopes, while geometry-refined multi-view methods become best with three and achieve 66–67% coverage within 1 mm at six views.
Problem
Real surgical multi-view reconstruction lacks rigorous evaluation benchmarks because clinical telerobots provide one stereo viewpoint and suitable synchronized data with accurate references are scarce.
Method
MV-dVRK combines three exposure-synchronized stereo viewpoints, validated static reference geometry and poses, sparse-view benchmarking, and synchronized dynamic sequences.
Results
Geometry-refined methods achieve 66–67% coverage within 1 mm at six synchronized views, compared with 28–44% for feed-forward models, while the best paradigm changes with viewpoint count.
Takeaways & Limitations
The benchmark identifies registered stereo as competitive with two endoscopes and geometry-refined multi-view reconstruction as strongest with a third viewpoint.
Takeaways & Limitations
The static benchmark covers eight scenes from two porcine specimens and three surgical tasks, so measured effects should not be assumed to transfer unchanged to new anatomies, procedures, or in-vivo conditions.
Abstract
from arXiv · showhide
Large-scale training and refined optimization techniques have greatly improved sparse multi-view 3D reconstruction. Despite their relevance to surgery, such methods have never before been rigorously evaluated on real endoscopic images. Current clinical telerobots deploy a single stereo camera inside the patient, making multi-viewpoint data extremely rare. This paper presents MV-dVRK, the first ex-vivo surgical dataset to combine multiple exposure-synchronized stereo viewpoints with accurate surface geometry and camera poses. The static subset of the benchmark provides dense SfM reference geometry, validated against an industrial 3D scanner, together with ground-truth camera poses and sparse-view test sets. We use MV-dVRK to systematically compare zero-shot monocular, stereo, multi-stereo, and multi-view 3D reconstruction methods as the number of viewpoints increases. With two endoscopes, multi-stereo reconstruction achieves the highest coverage. With a third viewpoint, optimization-based multi-view methods perform best, covering 67% of ground-truth surface points within a 1 mm tolerance and recovering highly accurate relative camera poses. By contrast, feed-forward foundation models cover only 43% of the ground-truth surface in the same setting. MV-dVRK also includes ten dynamic sequences spanning multiple surgical tasks, with increasing kinematic complexity and tissue deformation, providing a basis for future research in multi-viewpoint surgical perception. The project is available at: https://mv-dvrk.is.mpg.de.
1. Introduction
MV-dVRK addresses the scarcity of rigorous sparse multi-view 3D reconstruction benchmarks for real surgical imagery. It combines synchronized viewpoints, validated geometry, and systematic evaluation to study how reconstruction changes with viewpoint count.
- Accurate geometric reconstruction can support free-viewpoint rendering, geometry-aware robotic assistance, and autonomy in robot-assisted minimally invasive surgery.
- Clinical telerobots provide only a single stereo viewpoint, leaving unseen anatomy ambiguous despite methods for tissue deformation and occlusion.
- Sparse multi-view surgery remains challenging because few cameras, large baselines, limited overlap, and scarce suitable benchmarks constrain evaluation.
- MV-dVRK provides three exposure-synchronized stereo viewpoints, validated reference geometry and poses, sparse-view test sets, and synchronized dynamic sequences.
- The benchmark compares monocular, stereo, multi-stereo, and multi-view reconstruction as synchronized input viewpoints increase.
2. Related work
Prior surgical reconstruction datasets and methods face limitations in viewpoint diversity, synchronization, geometric ground truth, and sparse-view robustness. MV-dVRK is positioned against these gaps across stereo, monocular, feed-forward multi-view, and geometry-refined reconstruction approaches.
- Most robotic-surgery datasets provide one viewpoint and trade realistic data against accurate geometric ground truth.
- Scanner, structured-light, and computed-tomography ground truth can be costly, time-consuming, disruptive, or unavailable for large in-vivo collections.
- Existing multi-view datasets include software-synchronized streams or relatively small inter-viewpoint baselines, limiting precise geometric evaluation in some settings.
- Stereo methods provide metric depth from calibrated pairs, while monocular methods target geometry recovery from single endoscopic video and temporal accumulation.
- MV-dVRK uniquely combines independently moving viewpoints, exposure synchronization, and reference geometry and camera poses across static and dynamic subsets.
- Single-camera reconstruction of deforming and unseen anatomy remains fundamentally ill-posed.
- Feed-forward models regress geometry and camera parameters from sparse unposed views, whereas geometry-refined methods jointly optimize structure and poses under sparse, wide-baseline conditions.
3. The MV-dVRK dataset
MV-dVRK extends the dVRK with independently controlled and fixed stereo endoscopes to collect synchronized static and dynamic surgical data. Its static subset supports geometric benchmarking, while its dynamic subset captures camera motion and tissue deformation.
- Data acquisition platform: The MV-dVRK platform extends the dVRK with a second independently controlled stereo endoscope and a fixed stereo endoscope, yielding three synchronized stereo viewpoints.
- Data acquisition platform: A shared hardware trigger exposes all six sensors simultaneously, with clock alignment achieving microsecond-level timestamp agreement across cameras.
- Dataset collection and composition: Data were recorded during mock surgery on freshly excised abdominal organs from two pigs while surgeons performed cholecystectomy, tubular dissection, and small-bowel anastomosis.
- Dataset collection and composition: The static subset contains eight reference scenes spanning three surgical tasks and varying tissue appearances, robot configurations, and camera geometries.
- Dataset collection and composition: Each static scene contributes 10 synchronized keyframes, producing 80 test instances with three stereo viewpoints each.
- Dataset collection and composition: The dynamic subset contains ten long synchronized multi-view sequences with surgical manipulation, tool–tissue interaction, camera motion, tissue deformation, and dVRK kinematics.
4. Benchmark design
The benchmark uses dense SfM reconstructions as validated references and evaluates sparse synchronized inputs across viewpoint counts. It measures coverage, geometric error, and relative camera-pose accuracy after prediction-to-reference alignment.
- Benchmark protocol: Each test instance contains three synchronized stereo viewpoints, and experiments vary the available viewpoints from one to three.
- Reference reconstruction: The reference combines dense SfM geometry and camera poses from static scenes with independent industrial-scanner validation.
- Reference reconstruction: Reference reconstructions use dense views, while evaluated methods receive sparse synchronized images, enabling pixel-aligned predicted-to-reference 3D comparisons.
- Reference reconstruction: The benchmark uses 105 retained reference images per scene, with the third stationary endoscope sampled less densely because repeated frames add little viewpoint diversity.
- Prediction-to-reference registration: Predictions are aligned before evaluation, and absolute metric-scale recovery is not an evaluation target.
- Prediction-to-reference registration: Fitted registration estimates a robust Sim(3) transformation from direct E1L pixel correspondences using LO-RANSAC and Umeyama fitting.
- Metrics: The benchmark evaluates coverage, mean 3D error, and relative pose error, including translation in millimeters and rotation in degrees.
- Statistical analysis: Statistical comparisons average metrics within each of eight scenes before paired scene-level analysis and report confidence intervals and sign-consistency counts.
5. Benchmark analysis
The benchmark compares reconstruction paradigms as synchronized viewpoints increase, revealing a crossover from registered stereo to geometry-refined multi-view methods. Performance depends on cross-view support, while static benchmark scope limits transferability.
- Paradigm comparison: Registered multi-stereo is most accurate with two viewpoints, whereas a third viewpoint favors jointly optimized multi-view reconstruction.Depth Anything 3 + GGPT achieves the highest coverage and lowest E1L mean error in the three-view comparison.
- Single-view reconstruction: 27 percentage points: FoundationStereo increases coverage over two-frame Endo3R across all eight scenes.After 50 monocular frames, stereo still retains a 15-point coverage advantage.
- Multi-stereo reconstruction: 14 percentage points: registering a second stereo reconstruction increases coverage, followed by a further 7.2-point gain from a third viewpoint.Cross-endoscope registration can introduce duplicated structures and inconsistencies despite accurate local stereo geometry.
- Multi-view reconstruction: 1.9 mm: adding a third viewpoint reduces geometry-refined methods’ E1L mean error, a 54% relative improvement across eight scenes.The E2L reduction is more variable, with a mean of 2.1 mm and a 95% CI of [−0.05, 4.2].
- Paradigm comparison: 66–67%: geometry-refined methods cover this share within 1 mm at six synchronized views, versus 28–44% for feed-forward models.The crossover indicates that explicit geometric optimization benefits most when sufficient cross-view support is available.
- Role of view overlap: Reduced co-visibility is associated with increased reconstruction error because partially overlapping views leave regions without cross-view geometric constraints.Under partial overlap, Depth Anything 3 + GGPT relies more on its learned prior for inaccurate 3D predictions.
- Implications and limitations: MV-dVRK suggests complementary roles for calibrated stereo and global multi-view optimization, while practical deployment requires lower-latency matching and optimization.The benchmark contains eight static scenes from two porcine specimens and three surgical tasks, and its effect sizes may not transfer unchanged to new settings.
6. Conclusion
MV-dVRK combines synchronized multi-view stereo data with validated reference geometry and supports controlled sparse-view reconstruction evaluation. Results show that geometry-refined multi-view methods benefit strongly from a third viewpoint, while dynamic sequences extend the benchmark to motion and deformation.
- MV-dVRK combines three exposure-synchronized stereo viewpoints with validated reference geometry and camera poses for controlled evaluation on real robotic surgery data.
- Geometry-refined multi-view methods benefit strongly from a third viewpoint and achieve the best coverage and reference-view accuracy.
- Synchronized dynamic sequences extend MV-dVRK to camera motion and tissue deformation, supporting future multi-viewpoint surgical 3D perception research.