Source-linked AI summary

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

Haosong Peng, Hao Li, Jiaqi Chen, Yuhao Pan, Runmao Yao, Yalun Dai, Fushuo Huo, Fangzhou Hong, Zhaoxi Chen, Haozhao Wang, Dingwen Zhang, Ziwei Liu, Wenchao Xu

arXiv:2605.27367v2cs.CV

TL;DR

SpatialBench addresses limited evaluation of spatial foundation models by providing a deterministic, cross-paradigm benchmark spanning diverse domains and input densities. Across 41 models, it finds that current models are not yet all-round players, while full-context attention, bounded-memory strategies, and targeted in-domain data address different performance gaps.

  • Problem

    Existing evaluations have narrow paradigm and domain coverage with arbitrary frame sampling, limiting assessment of spatial foundation models under domain shifts, varying input densities, and hardware constraints.

  • Method

    SpatialBench uses deterministic density-aware sampling, cross-paradigm comparisons, and diverse datasets to evaluate spatial foundation models across domains, densities, and reconstruction task suites.

  • Results

    Current spatial foundation models are not yet all-round players; full-context attention maximizes accuracy, bounded-memory models enable long-horizon scalability, and data quality outweighs data volume.

  • Takeaways & Limitations

    Targeted domain-specific data curation can close embodied viewpoint gaps, with DA-Next improving over DA3-Giant by +47%/+59% in depth estimation and +3.1%/+5.5% in pose estimation on sparse/medium inputs.

  • Takeaways & Limitations

    Dense evaluation imposes a maximum frame budget for feasibility, and evaluating 41 models across more than 100 scenes per model is time-consuming.

Abstract

from arXiv · show

While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round players capable of generalizing robustly across diverse downstream tasks, arbitrary viewpoints, shifting scene domains, varying input densities, and specific hardware constraints? Answering this overarching question requires a holistic assessment, yet current models are mainly evaluated on specific domains for which they were specifically designed or trained. Such evaluations are intrinsically limited by narrow paradigm coverage, limited scene domains, and arbitrary frame sampling, making it fundamentally difficult to assess their true generalization capabilities. To address this gap, we present SpatialBench, a cross-paradigm, domain-diverse benchmark for spatial foundation models with deterministic sampling. SpatialBench features unprecedented scale and rigorous deterministic design, comprising 19 datasets and 546 scenes across 5 diverse spatial domains. It comprehensively evaluates 41 models across 6 paradigms on 5 task suites under 4 different input density settings. Our extensive evaluation reveals that current models are not yet all-round players, and uncovers crucial insights for future advancement. Specifically, we demonstrate that full-context attention maximizes accuracy while bounded-memory strategies unlock long-sequence scalability. Moreover, our empirical evaluations in challenging embodied and egocentric tasks demonstrate that strict domain alignment and high data quality are far more critical to performance than simple dataset scaling. Furthermore, to address the largest data gap identified in our analysis, we go beyond evaluation by introducing a large-scale dataset, DA-Next-5M, and a strong baseline model, DA-Next, pushing the boundaries of spatial representation learning.

1. Introduction

SpatialBench addresses limited, nonstandardized evaluations by comparing spatial foundation models across paradigms, domains, input densities, and tasks. Its analyses show that current models remain short of all-round generalization, while targeted domain data improves embodied-view performance.

  • Motivation: Spatial foundation models must remain reliable across scene shifts, sparse-to-dense inputs, and hardware memory constraints in real-world applications.These demands motivate evaluation beyond standard reconstruction benchmarks.
  • Motivation: Existing benchmarks provide narrow paradigm coverage and inconsistent protocols, using different splits, subsets, and frame selections.This makes direct comparison and assessment of generalization difficult.
  • Benchmark: SpatialBench evaluates models with deterministic sampling across four density regimes and unifies comparisons across geometric tasks.The protocol includes single-frame, sparse, medium, and dense settings.
  • Findings: Full-context attention maximizes accuracy, whereas bounded-memory strategies enable long-horizon scalability at the cost of geometry estimation accuracy.The benchmark analysis identifies a trade-off between accuracy and scalable sequence processing.
  • Findings: DA-Next gains +47%/+59% in depth estimation and +3.1%/+5.5% in pose estimation over DA3-Giant on sparse/medium inputs.The gains are reported for embodied and egocentric domain data curated in DA-Next-5M.

2. Related Work

Spatial foundation models evolved toward online, memory-based, and test-time adaptive reconstruction, while existing benchmarks broadened task coverage without fully standardizing cross-domain and cross-paradigm evaluation.

  • Spatial foundation models: Early systems reframed pose-free reconstruction as dense pointmap prediction, reducing reliance on optimization-heavy pipelines while retaining alignment or post-processing components.DUSt3R and MASt3R established this direction.
  • Long-sequence models: Recent online, streaming, SLAM-based, and test-time training models target continuous reconstruction from realistic video streams.These methods use temporal memory, recurrent states, compact historical context, or inference-time adaptation.
  • Long-sequence models: The field increasingly emphasizes memory management, temporal consistency, dynamic-scene robustness, and long-range geometric alignment.These concerns extend beyond accurate single-shot reconstruction.
  • Related benchmarks: E3D-Bench supports multi-task cross-method comparison but lacks comparisons across diverse domains and model paradigms.Model-specific suites likewise remain tied to individual studies and protocols.
  • Related benchmarks: SpatialBench adds deterministic, tag-aware evaluation across input densities, viewpoint types, scene dynamics, and foundation-model paradigms.Its design directly targets the coverage and standardization gaps identified in prior benchmarks.

3. SpatialBench Design

SpatialBench standardizes heterogeneous 3D datasets into deterministic scene-density evaluations spanning diverse environments, viewpoints, dynamics, and task suites. Its regimes balance viewpoint coverage, temporal continuity, and feasibility under model memory limits.

  • Dataset: SpatialBench combines heterogeneous datasets covering scene categories, capture conditions, and viewpoint configurations for systematic evaluation.Figure 2 summarizes scene-category counts, data sources, and frame densities.
  • Dataset: The benchmark normalizes each scene into RGB frames, metric depth, camera-to-world poses, and intrinsics, then fixes frame indices for every scene-density pair.This decouples data ingestion from evaluation and gives methods identical inputs.
  • Data curation: DROID wrist-view sequences use stereo depth, initial MapAnything poses, SAM3 dynamic masks, and masked depth-photometric bundle adjustment for pose refinement.The masks exclude grippers and interacting objects from optimization under a static-background assumption.
  • Multi-density regimes: SINGLE fixes one reproducible frame, while SPARSE greedily selects a small voxel-covering set that promotes viewpoint diversity.Sparse selection is formulated as weighted set cover over scene voxels.
  • Multi-density regimes: MEDIUM favors overlapping views with a length-adaptive budget, whereas DENSE preserves temporal continuity for online long-horizon reconstruction.Dense evaluation imposes a maximum frame budget to remain feasible across methods.
  • Tasks and metrics: The evaluation covers 31 methods, 41 variants, six paradigms, and five general tasks using unified metrics for pose, trajectory, depth, reconstruction, and prior-enhanced prediction.The task metrics include pose accuracy, trajectory errors, depth errors, reconstruction scores, and two prior-input settings.

4. Depth-Anything-Next

Depth-Anything-Next addresses egocentric and wrist-view gaps using DA-Next-5M, a large-scale 3D dataset designed around challenging embodied-intelligence viewpoints. The model extends DA3 with end-to-end absolute-scale prediction and optional camera-token guidance.

  • DA-Next-5M Dataset: DA-Next-5M contains 5.5M high-quality frames across 22K scenes, primarily from egocentric and wrist-view perspectives.These viewpoints involve high motion dynamics, frequent occlusions, and ultra-close-range capture.
  • DA-Next-5M Dataset: DA-Next-5M provides image sequences, metric depths, camera intrinsics, and extrinsics, with domain randomization applied to simulation data.Randomized factors include background appearance, object size, color, and wrist camera placement.
  • Model Architecture: DA-Next extends DA3 by predicting absolute scene scale end-to-end from frame sequences.The architecture uses patch, camera, and scale tokens within a transformer encoder.
  • Model Architecture: Additional scale tokens learn scene-level metric scale, while optional GT camera tokens provide auxiliary geometric guidance.Without auxiliary camera information, the model uses learnable camera tokens instead.
  • Training: Training uses a five-component multi-task loss covering depths, depth gradients, ray maps, points, and scale.Ground-truth signals are canonicalized to the first-camera coordinate frame and normalized by a per-scene scale factor.
  • Training: The model is trained mainly on DA-Next-5M while retaining 11 general 3D datasets for joint training.Training follows DA3-Giant with 41 transformer blocks and uses ground-truth camera input with probability p=20%.

5. Findings: How to Train Your Best Spatial Foundation Models?

SpatialBench shows that model design, memory strategy, training-data quality, and domain alignment jointly determine spatial foundation model performance. Full-context models maximize accuracy, bounded-memory methods enable long-sequence processing, and targeted egocentric and wrist-view data addresses major OOD failures.

  • Benchmark scope: 41 models are evaluated across five input settings, with depth, pose, trajectory, and point-cloud metrics reported on SpatialBench.The benchmark spans six reconstruction paradigms and summarizes performance across single-frame, sparse, medium, dense, and average settings.
  • Memory and accuracy: Full-context models define the accuracy upper bound under the same input budget, with DA3-Giant and π3 achieving the lowest depth errors among compared paradigms.Global attention jointly resolves cross-view correspondences and scene-level consistency, strengthening reconstruction accuracy under bounded inputs.
  • Memory and accuracy: Bounded-memory streaming, chunk-wise, and TTT models maintain flatter memory growth and process sequences that full-context models cannot complete, trading depth precision for scalability.These methods are better suited to long-horizon or resource-constrained deployment, while full-context models remain preferable when bounded-input accuracy is the priority.
  • Training data: Data volume correlates positively with performance, but carefully curated pseudo-GT supervision can outperform larger, noisier training mixtures at comparable dataset scales.DA3 achieves the highest composite scores among feed-forward models without using the largest training corpus.
  • Domain alignment: Ego-view and especially wrist-view sequences are the dominant OOD failure modes because current training mixtures lack real wrist-view data and contain sparse egocentric coverage.The gap affects the full evaluated method pool rather than a single weak model.
  • Domain alignment: DA-Next improves over DA3-Giant on egocentric and wrist-view data, reducing sparse-input depth AbsRel from 0.095 to 0.050 and medium-input AbsRel from 0.086 to 0.035.AUC@30 also improves by +3.1% on sparse inputs and +5.5% on medium inputs after training on DA-Next-5M.

6. Conclusion

SpatialBench evaluates spatial foundation models across diverse domains, input densities, and reconstruction suites, finding that current systems are not yet all-round players. The benchmark also introduces DA-Next-5M and DA-Next to address a major data gap.

  • 41 models across 6 paradigms reveal gaps in domain generalization and input-density robustness.
  • DA-Next-5M and the DA-Next baseline target the most significant data gap identified by the benchmark.

A. SpatialBench Data Curation Pipeline

SpatialBench curates heterogeneous real and simulated data through dataset-specific processing and a unified depth-cleaning pipeline. The curation emphasizes reliable poses, depth, masks, and point clouds for downstream evaluation.

  • SpatialBench processes Xperience, DROID, and simulation datasets, then applies unified depth-map post-processing to selected data.
  • DROID annotations are rebuilt from stereo depth because raw data contain noisy geometry and unreliable camera calibration.
  • S2M2 is selected for wrist-view sequences, after which SAM3 masks and bundle adjustment support globally aligned point-cloud reconstruction.
  • Xperience contributes head-mounted egocentric sequences with SLAM poses and metric depth estimated from rectified stereo pairs.
  • Robot manipulation sequences are collected from RLBench, Robo Colosseum, RoboTwin, and RoboLab simulators.
  • A five-stage cleaning pipeline creates per-frame validity masks by clipping depth, removing flying points and isolated regions, filtering edges, and masking sky.

B.1. Benchmark Composition

SpatialBench aggregates diverse datasets and evaluates models under deterministic single-frame, sparse, medium, and dense input regimes. It combines coverage-aware sampling with pose, depth, trajectory, and point-cloud metrics, while documenting challenging prior-enhanced failures.

  • 19 source datasets span indoor and outdoor scenes, static and dynamic sequences, three viewpoint regimes, real sensors, and simulation.
  • Sparse regime: Sparse sampling greedily selects a frame set that maximizes voxel coverage under a frame budget.
  • Medium regime: Medium sampling uses a length-adaptive frame range to avoid over-pruning short sequences and over-sampling long ones.
  • Dense regime: Dense evaluation preserves temporal continuity while limiting scenes to approximately 500 evenly spaced frames when trajectories are longer.
  • Evaluation metrics: Camera pose and trajectory quality are measured with relative rotation and translation accuracy, AUC, ATE, and relative pose errors.
  • Evaluation metrics: Depth metrics use median-scale alignment by default, while point-cloud quality is assessed with threshold-based precision and recall and Chamfer-based measures.
  • Failure analysis: Prior-enhanced models fail under object-centric scenes, extreme no-overlap views, and out-of-distribution wrist-view sequences.

C. The Collection of DA-Next-5M

DA-Next-5M draws on diverse egocentric, interaction, and robotic manipulation data sources. Its source collection spans real-world activity datasets and multiple simulation benchmarks with varied tasks, perturbations, embodiments, and competencies.

  • ADT contains egocentric RGB-D activity sequences captured with Meta Project Aria smart glasses.
  • HOI4D provides large-scale egocentric human-object interaction sequences with RGB-D, camera poses, hand poses, object poses, and segmentation.
  • RLBench supplies simulated manipulation tasks with expert demonstrations and observations from over-the-shoulder stereo and eye-in-hand cameras.
  • Robo Colosseum evaluates manipulation generalization across environmental perturbations including appearance, lighting, distractors, physics, and camera pose.
  • RoboTwin contributes bimanual manipulation across 50 tasks and 5 heterogeneous robot embodiments.
  • Robolab contributes photorealistic, physically accurate simulation scenes organized across visual, procedural, and relational competencies.

D.1. Model Architecture

DA-Next follows a pretrained DA3-Giant architecture and predicts depth, rays, scale, and confidence maps with a multi-term geometric objective. Training uses heterogeneous multi-view data, pose conditioning, canonicalized supervision, and varied input resolutions.

  • Architecture: DA-Next uses a 40-block pretrained DA3-Giant Transformer with DINOv2 ViT-Giant features and DualDPT prediction heads.Multi-scale features are extracted from layers 19, 27, 33, and 39; the ray and depth heads output two-dimensional predictions.
  • Outputs: The model predicts ray maps, depth maps, global scale, and confidence maps end to end.The predicted rays encode camera origins and viewing directions, while scale is used to recover metric geometry during inference.
  • Normalization: Ground-truth geometry is canonicalized relative to the first frame and normalized by scene scale before loss computation.At inference, the predicted scale is multiplied back to restore metric geometry; quantile filtering uses τ=0.98.
  • Data and Training: DA-Next trains on 18 datasets spanning normal, egocentric, and wrist views, synthetic and real scenes, and indoor and outdoor environments.Batches contain 18 frames sampled from multiple scenes, with frame selection based on camera-pose distance and a fixed random seed of 42.

E.1. Do More Input Frames Always Lead to Better Results?

More frames do not uniformly improve reconstruction: sparse inputs often lack correspondence, medium inputs help, and dense inputs can add redundancy and memory pressure. Thus, reconstruction quality depends on task and model rather than increasing strictly with input density.

  • Input Density: Sparse inputs often hinder geometric estimation because large baselines and limited overlap provide insufficient visual correspondence.
  • Input Density: Increasing inputs from sparse to medium consistently improves depth and camera-pose metrics across paradigms.The improvement is associated with a moderate level of multi-view overlap.
  • Input Density: Dense inputs do not uniformly improve reconstruction and can degrade feed-forward performance when redundant frames increase memory pressure.Near-duplicate frames may add little useful geometric constraint while increasing computational demands.
  • Takeaway: Bounded reconstruction tasks have an optimal input-density range, so selecting density is as important as selecting the model.Too few frames underspecify geometry, whereas excessive inputs can introduce redundancy and hurt performance.

E.2. How to Select the Right Spatial Foundation Model for Your Task?

SpatialBench rankings vary across Indoor, Outdoor, Driving, Wrist-view, and Ego-view domains and across input densities. Training-data coverage and architectural operating regimes therefore matter when selecting a model for deployment.

  • Ranking Protocol: SpatialBench ranks methods using weighted Average scores that aggregate AbsRel, AUC@30, ATE, and F-Score across five domain groups.The groups are Indoor, Outdoor, Driving, Wrist-view, and Ego-view.
  • Cross-Domain Selection: Model rankings are highly inconsistent across domain groups and input-density settings.The reported factors are training-data coverage and architectural choices favoring different operating regimes.
  • Cross-Domain Selection: Streaming methods may excel on dense, extended trajectories while underperforming on sparse multi-view inputs.
  • Takeaway: No single model dominates across all domains and input regimes, making target-domain coverage the first selection or fine-tuning consideration.The benchmark identifies domain coverage in training data as the primary predictor of per-domain ranking.

F. The Complete SpatialBench Results

SpatialBench reports detailed results across input regimes, datasets, and metric-depth evaluation. The supplied passages describe the organization and conventions of these results rather than individual numerical outcomes.

  • Aggregate Results: The benchmark provides regime-specific aggregate results for single-frame, sparse, medium, and dense settings.These are reported in Tables 9–12 and complement the aggregated results in Table 1.
  • Per-Dataset Results: Per-dataset metric breakdowns are provided for datasets including 7-Scenes, ADT, DROID, DTU, ETH3D, HiRoom, and KITTI Odometry.Additional per-dataset tables cover Lingbot-Depth, NRGBD, OmniWorld, RLBench, Robolab, RoboTwin, Xperience, ScanNet++, Tanks and Temples, TUM RGB-D, Virtual KITTI, and Waymo.
  • Metric Depth: Table 32 compares native metric-depth predictions from six metric-scale-capable methods across single-frame, sparse, and medium settings.The table also reports an average across the three settings and highlights the best three results in each column.
  • Reading the Tables: Result tables use colored highlighting for the best, second-best, and third-best values, while dense results mark out-of-memory and timeout cases.Average values for affected rows may be excluded from per-column ranking, and some datasets lack dense-regime evaluation.

G. Limitations

SpatialBench has several practical and methodological limitations, including evaluation cost, hardware scope, untuned method-specific hyperparameters, and incomplete coverage of newly released models.

  • Evaluating 41 models across more than 100 scenes per model under the dense regime is time-consuming.The authors note that parallel evaluation across multiple GPUs can mitigate this cost.
  • All evaluations use H200 GPUs with 141 GB VRAM, so behavior may differ under larger-memory configurations such as B100 or B200.
  • Some methods may require task- or scene-specific hyperparameter tuning for optimal performance, but such tuning is outside the benchmark’s scope.The benchmark instead uses recommended configurations from each method’s official codebase for consistency.
  • SpatialBench cannot guarantee complete coverage of all existing methods because newly released models continue to be open-sourced.The authors commit to integrating and evaluating new methods as they become available.
Loading 2605.27367v2…