Source-linked AI summary

ORBIT++: Benchmarking SfM in the Wild with 360° Video

Sara Sabour, Linyi Jin, Richard Tucker, Amir Hertz, Marcus Brubaker, Saurabh Saxena, Junhwa Hur, Andrea Tagliasacchi, Deqing Sun, David J. Fleet, Richard Szeliski, Noah Snavely

arXiv:2608.22039v1cs.CV

TL;DR

Challenging real-world SfM lacks reliable benchmarks, limiting evaluation of camera pose estimation on complex motion and dynamic scenes. ORBIT uses 360° videos to recover robust trajectories and construct difficult perspective clips; experiments find that current methods completely fail on at least 36% of clips.

  • Problem

    Challenging and realistic camera-pose benchmarks are scarce, because existing evaluations are largely synthetic or simple.

  • Method

    ORBIT curates difficult 360° videos, estimates camera poses and point clouds using omnidirectional SfM, then crops and reprojects challenging regions into perspective-view benchmark clips.

  • Results

    At least 36% of clips cause complete camera-pose failure for every evaluated method, while the ground-truth pipeline achieves ATE 0.07 ± 0.04m and RPE-R 0.13 ± 0.07 on 360Loc.

  • Takeaways & Limitations

    ORBIT provides a diverse real-world testbed for diagnosing SfM failure modes and measuring progress on difficult camera-pose estimation.

Abstract

from arXiv · show

Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes. Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard to gauge real-world progress or to pinpoint where improvements are most needed. To address this gap, we introduce a new benchmark for evaluating camera pose estimation. Our key insight is to leverage online panoramic 360° video as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery. The panoramic nature of these videos provides richer visual context for tracking camera motion, even when parts of the view are affected by blur, motion, or dynamic objects. After tracking camera motion across full 360° videos, we crop and reproject selected portions to generate perspective-view clips that serve as our benchmark, called ORBIT. Experiments show that COLMAP, as well as recent optimization-based and feed-forward SfM methods struggle to accurately estimate camera poses on our benchmark. Hence, ORBIT provides a valuable testbed where researchers can meaningfully measure progress on truly challenging, real-world SfM problems.

1. Introduction

Current SfM methods can fail on complex dynamic real-world videos, while reliable ground-truth benchmarks for these conditions are scarce. ORBIT addresses this gap by using 360° videos to build challenging perspective clips with robust camera trajectories, revealing substantial weaknesses in existing methods.

  • SfM methods can fail on videos containing moving objects, complex motion, and specular surfaces, but benchmarks reflecting these challenges are scarce.
  • High-quality ground-truth trajectories are difficult to obtain for in-the-wild videos, and COLMAP fails on many such videos despite its frequent use as a reference.
  • 360° videos provide broader visual context and known intrinsics, helping camera-motion estimation when parts of a scene contain difficult content.
  • ORBIT curates difficult 360° videos, estimates pseudo-ground-truth poses with a custom 360° SfM method, and renders perspective crops selected for baseline difficulty.The benchmark contains 300 evaluation videos, each up to roughly 30 seconds long.
  • Every evaluated method completely fails on at least 36% of ORBIT clips, while challenge categories support comparison of failure modes.The evaluated methods include COLMAP, MegaSaM, and VGGT.

2. Related Work

Classic SfM and SLAM methods support reliable reconstruction and efficient tracking but lack learned priors for challenging cases. Existing dynamic-video datasets often rely on estimated poses that are difficult to verify, whereas ORBIT uses full-field-of-view information for ground-truth construction.

  • COLMAP registers unordered images through feature matching and incremental pose and structure estimation, while ORB-SLAM tracks ordered video frames and builds maps incrementally.
  • Classic methods lack learned priors to guide estimation in challenging cases.
  • DynPose-100K and SpatialVID provide camera annotations, but their use of pose-estimation systems such as COLMAP and MegaSaM makes accuracy difficult to verify.
  • ORBIT computes camera poses from 360° videos using the full field of view and generates evaluation videos through perspective projection.

3. Building ORBIT

ORBIT is constructed from online 360° videos through pose estimation, validation, and perspective reprojection. The resulting benchmark contains diverse clips with inherited camera-position estimates and varied viewing conditions.

  • The pipeline identifies challenging online 360° clips, estimates poses with a customized 360° SfM system, validates them by cross-validation, and reprojects them into perspective videos.
  • ORBIT comprises 308 clips selected from 77 unique 360° videos after pose estimation, filtering, and cropping.
  • The benchmark clips contain 150 to 1,000 frames at 30 fps and use either 16:9 or 4:3 aspect ratios in portrait or landscape formats.
  • Source videos are screened for at least 1K resolution, suitable camera motion, and camera continuity so reprojection can produce at least VGA-compatible perspective videos.

3.2. Ground truth trajectory estimation

ORBIT estimates camera trajectories from full 360° videos using a rig-based SfM pipeline, then validates and rescales them before using them as ground truth. The pipeline exploits multiple perspective cube-face views and cross-validates trajectories on independent projections.

  • The pipeline estimates per-frame camera poses and a 3D point cloud from 360° video while jointly using information across the input.
  • Each equirectangular frame is reprojected into four rigidly attached perspective cube-face views to avoid anisotropic projection and wraparound difficulties.The faces share a projection center and fixed relative orientations.
  • ORB-SLAM2 initializes the rig geometry, after which incremental SfM and bundle adjustment recompute cross-face correspondences and jointly refine camera poses and 3D points.The initialization uses the first k=32 ORB-SLAM2-derived poses.
  • Ground-truth trajectories are approximately metrically rescaled using monocular metric depth, accepting clips only when frame-scale variability satisfies σS < µS.The accepted clip scale is the mean frame scale µS.
  • The trajectories are cross-validated against independent ORB-SLAM2 runs on nonfrontal perspective crops using aligned ATE and RPE, and clips failing thresholds are culled.The frontal cube face is excluded because it initialized the 360° estimates.

3.3. Deriving benchmark video clips

ORBIT derives difficult perspective benchmark clips by perturbing viewpoint orientation and field of view within the full 360° capture. The resulting clips retain transformed ground-truth camera poses from the original trajectories.

  • Perspective crops can use arbitrary rotational trajectories and fields of view, allowing benchmark videos to mimic regular-camera motion.
  • When an independent perspective projection fails verification, its cube face is selected as the initial benchmark view to maximize trajectory difficulty.Otherwise, rendering starts from the frontal cube face.
  • Each synthesized viewpoint combines intentional low-frequency motion with optional high-frequency jitter assigned at none, medium, or large shake amplitude.Low-frequency patterns include waypoints, scan, cinematic, roll, and orbit.
  • The field of view is sampled uniformly from [30°, 120°], spanning telephoto-like narrow views to ultra-wide-angle views.
  • The resulting perspective videos inherit ground-truth camera poses from the 360° trajectories after rotations are transformed according to the per-frame perturbations.

3.4. Challenges faced by SfM methods

ORBIT includes diverse challenges that disrupt feature tracking, motion assumptions, or visibility of static scene structure. These include weak texture and lighting, rapid or rotational motion, dynamic and crowded scenes, egocentric captures, and fluid textures.

  • Nearly all clips contain moving objects in some or all frames, while classic SfM is especially affected by unstable or moving keypoints.
  • Approximately 12% of clips have low-texture static regions, and another approximately 12% contain dark, low-contrast scenes.Examples include snow, desert, underwater scenes, concrete walls, sky, nighttime settings, and caves.
  • Approximately 25% of clips involve high-speed cameras or jittery motion that can violate SLAM displacement or smoothness assumptions.
  • Large camera rotations and jitter are deliberately rendered, with fully feedforward methods such as VGGTLong and MonST3R more prone to error under large rotations.
  • Crowded scenes obscure static regions, egocentric captures create competing world and camera-object frames, and fluid water textures produce noisy matches.Crowded scenes comprise about 16% of clips.

3.5. Failure modes

ORBIT measures failure explicitly because some methods output inaccurate poses while others produce incomplete or no trajectories. Its benchmark videos span diverse dynamic content, rotation patterns, fields of view, and aspect ratios.

  • Methods differ in whether they always output a camera estimate or can drop frames or completely fail when reliable feature matches are insufficient.
  • A clip is counted as a complete failure when no camera estimate is generated or ATE and RPE exceed their clip-specific thresholds.
  • Example benchmark clips vary in dynamic-object presence, rotation patterns, fields of view, and aspect ratios.
  • Average algorithm success rates are reported in the final column of Table 1.

4. Results

ORBIT evaluations show that current SfM methods remain unreliable on challenging clips, with substantial failure and method-dependent performance variation. Challenge-specific analyses reveal distinct weaknesses across optimization-based and feedforward approaches, supporting ORBIT as a diagnostic benchmark.

  • 4.2. Evaluations: ATE, RPE-T, and RPE-R exhibit large standard deviations, so average errors are not robust for ORBIT's catastrophic-failure regime.The evaluation therefore also reports median ATE and failure rate.
  • 4.2. Evaluations: 36%: Every evaluated method completely fails on at least this fraction of ORBIT clips.Failure is defined using stringent trajectory-error criteria, while the benchmark also reports median and mean error metrics.
  • 4.2. Evaluations: Feedforward methods such as MonST3R and VGGT-Long rarely match ground-truth trajectories as well as bundle-adjustment methods such as COLMAP and MegaSaM.Sweeping the ATE threshold exposes different failure-rate distributions between these method families.
  • 4.2. Evaluations: High per-clip standard deviation indicates that different methods perform very differently on the same clip, consistent with distinct method-specific failure modes.Many clips have at least one method with relatively small ATE even when their average ATE is high.
  • 4.3. Challenge Analysis: Lack of texture challenges COLMAP and MegaSaM, while high-speed motion challenges VGGT-Long and RoMo masking helps most categories except texture-poor clips.MegaSaM is especially robust for crowds and improves over COLMAP on egocentric-object clips.
  • 4.3. Challenge Analysis: VGGT-Long underperforms MegaSaM and COLMAP on most rotation patterns but outperforms COLMAP on original patterns, indicating possible training-distribution bias.The analysis suggests diversifying camera rotations during feedforward-method training as a mitigation.
  • 4.3. Challenge Analysis: ORBIT's challenge diversity enables failure-mode comparison and diagnosis across SfM methods.The benchmark combines challenge-category analyses with per-clip performance variation to support method evaluation.

5. Conclusions

ORBIT addresses the lack of challenging, realistic camera-pose benchmarks by providing a diverse real-world SfM benchmark derived from 360° videos. Evaluations reveal a substantial gap between current state-of-the-art methods and ground truth, while future extensions are planned.

  • ORBIT is a new real-world SfM benchmark derived from 360° videos, addressing the lack of challenging and realistic evaluation benchmarks.
  • The benchmark includes diverse indoor and outdoor scenes with moving objects, loops, straight paths, and varied camera rotations and translations.
  • Current state-of-the-art methods evaluated on ORBIT show a large performance gap with ground truth, leaving many challenges open.
  • ORBIT is intended to evolve, with planned focal-length variations enabling future intrinsics-prediction challenges.

A. Method Parameter Details

The benchmark-generation pipeline applies verified 360° trajectory estimates, configurable rotation perturbations, and varied fields of view to synthesize challenging perspective clips. Its rotation model combines intentional low-frequency motion with optional handheld noise.

  • Metric-scale estimation subsamples frames by 10× and limits each cube face to 100 visible points, while depth uses 90° FOV cube faces resized to 512 × 512.
  • A rig-based trajectory is verified against cube-face trajectories only when it satisfies all specified verification conditions, including RPEr(g, e) < 0.2 and t(g, e) < 1.0.
  • Each synthesized clip either retains its original viewpoint or receives a randomly selected rotation pattern.
  • Applied rotations combine a low-frequency intentional motion component Li with an optional high-frequency handheld-noise component Ni, using an initial viewing rotation matrix B.
  • The low-frequency component uses five modes, with sampled pitch and roll distributions, while horizontal field of view is sampled from [30°, 120°].
  • Figure 9 reports the number of benchmark clips assigned to each rotation pattern.

B. Benchmark Method Configurations

The benchmark evaluates classic, learned, feed-forward, hybrid, and SLAM-based methods using specified configurations. These include detailed COLMAP settings, default or adapted neural-model settings, and explicit ORB-SLAM2 failure criteria.

  • COLMAP uses exhaustive matching, specified initialization-image targets, bundle-adjustment tolerances of 1e-6, and the SIMPLE RADIAL camera model.
  • Table 3 lists randomly selected rotation modes and their sampled parameter ranges for benchmark synthesis.
  • MegaSaM uses Depth Pro for monocular depth estimation, relative depth for sky prediction and metric-depth alignment, and default remaining hyperparameters.
  • RoMo+MegaSaM combines RoMo with SAM 2 features, RAFT optical flow, SAM 2 refinement, batch size 16, learning rate 2e-2, and trusted-frame ratio threshold 0.5.
  • MonST3R is adapted to longer sequences through 224-pixel resizing and windowed inference with window size 50 and overlap ratio 0.3.
  • VGGT-Long uses chunk size 60, overlap 30, loop chunk size 20, and dense-alignment IRLS parameters δ = 0.1, maximum iterations 5, and tolerance 1e-9.
  • ORB-SLAM2 uses up to 4,000 features per frame and counts a sequence as failed when fewer than 30 frames are tracked or two consecutive frames are dropped.

C. Trajectory Comparisons

Trajectory-comparison figures juxtapose sample frames with ground-truth and estimated trajectories, with ground truth shown in blue. Feed-forward methods rarely match perfectly, whereas bundle-adjustment methods can match many frames but also diverge dramatically.

  • Figures 10 and 11 show sample frames and trajectory graphs comparing ground truth with each method’s output on ORBIT clips.
  • Feed-forward methods such as MonST3R and VGGT-Long almost never produce perfectly matching trajectory estimates.
  • Bundle-adjustment methods such as COLMAP and MegaSaM perfectly match many frames but can go completely out of bounds on others.
  • Ground truth trajectories are shown in blue in the trajectory comparisons.
Loading 2608.22039v1…