Source-linked AI summary

WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence

Xiangyu Han, Mengyu Yang, Jiaqi Li, Bowen Chang, Ziyu Chen, Hexu Zhao, Rahul Kumar Agrawal, Anthony Rodriguez, Fiona Hua, Marco Pavone, Chen Feng, Yiming Li

arXiv:2607.06838v1cs.CV

TL;DR

City-scale spatial intelligence lacks continuous, long-range real-world visual-spatial data. WildCity provides such a dataset and an urban-tailored reconstruction baseline with closed-loop simulation, establishing a practical testbed for realistic city-scale reconstruction.

  • Problem

    City-scale spatial intelligence lacks continuous, long-range real-world visual-spatial data capturing the scale and complexity of urban environments.

  • Method

    WildCity combines autonomous-fleet multimodal observations with an urban-tailored reconstruction baseline and integrates reconstructed environments into a closed-loop simulator.

  • Results

    WildCity provides a practical testbed for studying real-world city-scale reconstruction under realistic noise, uncertainty, and long-range spatial extent.

  • Takeaways & Limitations

    The dataset and baseline support investigation of simulation-ready urban digital twins and city-scale spatial intelligence.

  • Takeaways & Limitations

    Real-world uncertainty from dynamic objects, illumination changes, and pose noise remains a fundamental challenge for city-scale street-view reconstruction.

Abstract

from arXiv · show

Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI build spatial representations at a comparable scale? Although recent foundation models have advanced scene reconstruction and embodied intelligence, scaling to entire cities remains an open challenge, primarily due to the lack of city-scale data. To bridge the gap, we introduce WildCity, a real-world multimodal dataset collected by autonomous fleets traversing complex urban environments. Our dataset includes 18 trajectories, each averaging 83.7 kilometers in length, and preserves the core challenges of in-the-wild perception, e.g., dynamic objects, lighting variations, and imperfect camera poses. We further establish an urban-tailored reconstruction baseline and convert the reconstructed environments into a closed-loop simulator. Beyond the dataset and baseline, we systematically analyze the key challenges on the path to simulation-ready urban digital twins: scalability, extrapolation, and uncertainty. Ultimately, WildCity aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition. Project page: https://han-xiangyu.github.io/Wild-City/

1 Introduction

WildCity addresses the lack of continuous, long-range real-world visual-spatial data for city-scale spatial intelligence by introducing a multimodal dataset and testbed collected across complex urban environments. It also provides an urban-tailored reconstruction baseline, a closed-loop simulator, and analysis of scalability, extrapolation, and uncertainty.

  • Motivation: Current spatial-intelligence models are mainly demonstrated in small-scale scenarios and struggle when spatial coverage or video duration increases.Examples include a single room, synthetic apartment, or city block.
  • Motivation: Existing datasets do not provide continuous, long-range real-world visual-spatial observations: synthetic data has a sim-to-real gap, while real benchmarks use short clips of isolated urban scenes.This shortage makes building city-scale spatial representations infeasible.
  • WildCity dataset: WildCity is a real-world multimodal dataset and testbed collected by autonomous vehicle fleets traversing complex urban environments.It targets city-scale spatial intelligence and supports reconstruction, rendering, and embodied AI.
  • WildCity dataset: Over 1,500 km of traversed roads and 2.5 hours per log characterize WildCity’s city-scale coverage and long sensory streams.The dataset covers distinct functional zones and includes dynamic objects, lighting variations, motion blur, and imperfect camera poses.
  • Contributions: WildCity combines an urban-tailored reconstruction baseline with a closed-loop simulator and analyzes scalability, view extrapolation, and data uncertainty.The reconstructed environments support downstream embodied tasks such as end-to-end autonomous driving and motivate simulation-ready urban digital twins.

2 Related Works

Prior work improves large-scale 3D reconstruction and urban perception, but remains constrained by computational cost, long-horizon drift, limited representations, and the lack of real-world city-scale digital twins. WildCity addresses this gap with multi-city, long-route, surround-view data for reconstruction and interactive spatial reasoning.

  • City-Scale Street-View Datasets: 83.7 kilometers average and 1507.1 kilometers total distinguish WildCity as a real-world dataset combining six-city coverage, long continuous traversals, surround-view sensing, and city-scale route length.The comparison table reports WildCity as a city-scale dataset, while “∼” denotes partial or limited support for other datasets.
  • Large-Scale 3D Reconstruction: Learned 3D reconstruction methods remain limited in multi-kilometer urban settings by long-horizon pose drift and restricted fidelity of sparse point-based representations.The passage identifies DUSt3R, VGGT, and longer-horizon extensions as promising alternatives that still face these limitations.
  • City-Scale Street-View Datasets: Existing urban perception and navigation works still lack real-world, city-scale digital twins for interactive spatial reasoning and evaluation.WildCity provides the sensory foundation for constructing such twins and supporting future city-scale spatial memory and localization research.

3 WildCity Dataset

WildCity is a multimodal, multi-city dataset built from autonomous-fleet trajectories using calibrated surround-view cameras, LiDAR, GPS, and IMU. It provides long-horizon urban coverage with synchronized measurements, standardized trajectory segmentation, and semantic masks for reconstruction.

  • Sensor platform: The sensor suite combines a roof-mounted LiDAR, six surround-view RGB cameras, an IMU, and GPS with rigorous calibration into the ego-vehicle coordinate frame.Three cameras face forward with narrow angles, while three wide-angle cameras cover lateral and rear views.
  • Geographic diversity: The fleet operates in six U.S. cities spanning diverse urban styles, geographic layouts, climates, traffic patterns, and street-level scenarios.Scenarios include dense intersections, narrow local streets, multi-lane arterials, and high-speed parkways.
  • Semantic preprocessing: Semantic masks for ground, sky, and potentially movable objects support road regularization, sky modeling, and moving-object filtering during reconstruction.SAM3 text prompts generate the masks, while 3D tracking cuboids remove stationary instances such as parked vehicles and standing pedestrians.
  • Trajectory organization: Trajectory processing uses 0.5 m keyframe spacing, contiguous 5 km chunks, and sub-trajectories ranging from 50 m to 5 km.Segments share a city-level coordinate system, allowing direct concatenation into longer routes without additional registration.
  • Dataset scale: 3.01M keyframes provide synchronized surround-view RGB, LiDAR, GPS, and IMU measurements across approximately 40.18 km2 per city.Each driving log lasts about 2.5 hours and covers 83.7 km.

4 WildCity Method

WildCity builds a city-scale reconstruction pipeline around 3D Gaussian Splatting, jointly optimizing scene and multi-camera rig poses while adding sky, ground, and extrapolation-specific components. It further supports scalable training and closed-loop embodied interaction in the reconstructed urban environment.

  • 3D Gaussian Splatting: 3D Gaussian Splatting represents the scene with anisotropic Gaussians parameterized by opacity, position, rotation, scale, and view-dependent color, rendered through depth-ordered alpha blending.Projected Gaussians overlapping each pixel contribute according to their opacity and view-dependent color.
  • Rig pose optimization: Jointly optimizing ego poses and rig extrinsics with scene parameters preserves the rigid multi-camera setup and regularizes poses toward localization and calibration estimates.The resulting camera pose is T_t,c = T^ego_t T^rig_c, with rendering and pose-distance losses optimized together.
  • Sky model: A view-dependent MLP models the sky separately and composites its color with the rendered Gaussian image, decoupling sky appearance from geometry and reducing far-depth floaters.The sky is treated as an infinite background using the rendered Gaussian opacity.
  • Ground regularization: Ground regularization stabilizes underconstrained road geometry by penalizing local height variation and encouraging ground Gaussians to be vertically aligned and sufficiently opaque.The method combines distance, alignment, and opacity losses for ground Gaussians.
  • Extrapolated view post-repair: Progressive render-repair-augment training uses Difix3D+ to repair extrapolated renders, adds repaired views to the training set, and re-optimizes the Gaussians to reduce extrapolation artifacts.The cycle repeats progressively under sampled extrapolated viewpoints.
  • Multi-GPU training and closed-loop simulation: Multi-GPU training shards Gaussian parameters across devices and synchronizes statistics and gradients, enabling scaling to billion-level primitives and supporting closed-loop urban interaction.With Alpamayo 1, the simulator demonstrates smooth navigation and plausible multi-step reasoning from rendered observations.

5 Experiments

Experiments evaluate WildCity reconstruction across local and long-horizon urban sequences using image-quality and depth metrics, established baselines, and qualitative rendering comparisons. Results emphasize that simulation-ready city-scale reconstruction requires structural stability, pose consistency, and coherent off-trajectory views beyond image similarity alone.

  • Experimental Setup: Evaluation spans Ann Arbor-0.5k and Atlanta-5k using PSNR, SSIM, LPIPS, and Depth L1 on static regions with valid LiDAR depth.The sequences share the same sensor setup; depth is measured in meters.
  • Baselines: Baselines include 3DGS, H-3DGS, CityGaussianV2, VGGT-Long, and VGGT-Long+CityGS.The comparison covers classic, scalable hierarchical, city-scale blockwise, and feed-forward long-sequence approaches.
  • Quantitative Results: 23.14 dB PSNR and 6.62 m D-L1 are achieved by our method on the 2.5 km trajectory, producing the best overall performance.The method prevents structural degradation under sparse-overlap conditions, whereas image-similarity optimization alone does not ensure simulation-ready 3D structures.
  • Analysis: Partitioning and hierarchy improve large-scene scalability but do not resolve geometric ambiguity from narrow, long street-view trajectories.Reliable city-scale reconstruction additionally requires structural constraints and pose consistency.
  • Qualitative Results: Our method produces sharper textures and cleaner edges than baselines, which often over-smooth scenes, lose thin structures, or show texture and geometric inconsistencies.These differences are observed in both large street scenes and slender structures such as poles and distant façades.
  • Extrapolation: Simulation-ready urban digital twins require stable geometry and coherent appearance for views beyond the recorded trajectory, not only high-quality in-trajectory renderings.Figure 7 compares in-trajectory views with extrapolated views sampled outside the trajectory and also shows diffusion-based artifact repair.

6 Discussion

The discussion identifies scalability, view extrapolation, and real-world uncertainty as coupled obstacles to simulation-ready city-scale reconstruction. It argues that future evaluation should prioritize geometric correctness and simulation utility alongside 2D fidelity.

  • Performance scalability: Rendering quality degrades steadily with trajectory length, while peak VRAM usage and optimized Gaussian count increase.Longer routes also accumulate appearance variation and pose inconsistency, preventing current methods from scaling gracefully to continuous city-scale scenes.
  • View extrapolation: Current baselines often fail on extrapolated views, producing broken geometry, floaters, and unstable depth.Difix3D+ improves visual quality and partially repairs artifacts, but remains costly and can hallucinate content when the base reconstruction is weak.
  • Data uncertainty: Real-world uncertainty from dynamic objects, illumination changes, and pose noise substantially impacts reconstruction fidelity and geometric consistency.The baseline addresses dominant sources with sky modeling, rig-aware pose optimization, and ground regularization.
  • Data uncertainty: Removing ground regularization destabilizes road geometry, removing the sky model introduces background artifacts, and removing rig pose optimization causes multi-view misalignment.These ablations show that modules improve structural consistency and simulation usefulness even when 2D image metrics may slightly worsen.
  • Future evaluation: City-scale reconstruction remains limited by incomplete 3D priors, long-horizon pose inconsistency, and weak extrapolation under sparse observations.The discussion therefore calls for evaluation beyond 2D fidelity toward geometric correctness and simulation utility.

7 Conclusion

WildCity is a real-world city-scale dataset and practical testbed for street-view reconstruction, built from long-horizon surround-view RGB-LiDAR observations in unconstrained urban environments. It also provides an urban-tailored reconstruction baseline and closed-loop interaction in a reconstructed city digital twin.

  • Conclusion: WildCity combines long-horizon surround-view RGB-LiDAR observations from unconstrained urban environments into a real-world city-scale dataset for street-view reconstruction and beyond.The dataset is designed to support reconstruction under realistic noise and uncertainty.
  • Conclusion: The authors establish an urban-tailored reconstruction baseline on WildCity and enable closed-loop interaction in the reconstructed city digital twin.Together, these components form a practical testbed for studying real-world city-scale reconstruction.

A Implementation Details · A.1 Data Processing

WildCity’s data-processing pipeline preserves cross-modal timing, compensates LiDAR motion, separates dynamic foregrounds, and organizes long routes into reconstruction-ready evaluation units. These choices aim to maintain geometric consistency while supporting progressively larger spatial-scale evaluations.

  • A.1 Data Processing: Keyframes use front-camera timestamps, with each other sensor stream matched to its closest valid original-timestamp measurement rather than resampling.The sensors operate asynchronously at different frequencies, including cameras and LiDAR at 10 Hz, IMU at 100 Hz, and GPS/localization at 2 Hz.
  • A.1 Data Processing: Ego poses are interpolated from the localization stream at each keyframe timestamp.The resulting timestamp mismatch between reference images and associated LiDAR or pose measurements is typically below 50 ms.
  • A.1 Data Processing: LiDAR sweeps are deskewed by transforming every point from its acquisition time to the common keyframe reference time.This improves temporal alignment between LiDAR geometry and camera observations, yielding cleaner projected depth and more reliable supervision and evaluation.
  • A.1 Data Processing: Masks cover ground, sky, and potentially movable objects using SAM3 text-prompt initialization refined with onboard 3D tracking cuboids.The dynamic mask projects only tracked cuboids whose speed exceeds 1 m/s, excluding slower instances such as parked vehicles or standing pedestrians.
  • A.1 Data Processing: Dynamic-mask construction separates static structure from dynamic foregrounds, making the released masks more suitable for reconstruction, evaluation, and downstream simulation.Tracked instances below the speed threshold are excluded from the dynamic mask.
  • A.1 Data Processing: Routes are partitioned into contiguous 5 km chunks according to cumulative traveled distance.The release protocol makes long-horizon city logs tractable for reconstruction.
  • A.1 Data Processing: Released sub-trajectories span 50 m, 250 m, 500 m, 1 km, 2.5 km, and 5 km for evaluation across progressively more challenging spatial scales.These lengths are provided on top of the route chunks.

A.2 Our Method Implementation

The method extends a 3DGS training pipeline with urban-specific modifications, including separate sky modeling, rigid-rig pose optimization, bounded Gaussian refinement, and distributed multi-GPU training.

  • Sky model: The pipeline uses a lightweight view-dependent MLP for sky modeling instead of 3D Gaussians, with shared default optimization settings except for scene-dependent training steps and Gaussian capacity.The sky output uses directional encoding and appearance embeddings, while scene-dependent adjustments include total training steps and Gaussian capacity budget.
  • Sky model: At test time, the method replaces unavailable image-index appearance embeddings with the mean embedding over the training set.The appearance embedding has dimension R16, and a sigmoid produces the output RGB values.
  • Rig pose optimization: Pose optimization jointly estimates each keyframe’s ego pose and camera extrinsics in rigid-rig mode from training start, using learning rate 10^-3 and regularization weight 10^-5.The rigid factorization is intended to prevent small localization errors from accumulating into surround-view multi-view inconsistency.
  • Gaussian capacity and multi-GPU training: The method caps capacity at 30M primitives, densifies from 10k through 60k iterations every 100 steps, and trains long sequences on two H200 GPUs by sharding Gaussian parameters.Each refinement step adds 0.5% of the current primitive count, and spherical harmonics use degree 1.

A.3 Baseline Implementation · B Additional Experiments

The baselines use official or public implementations with only WildCity-specific data-interface adaptations, shared inputs, and method-specific training pipelines. Experiments preserve released urban reconstruction procedures while using strong, comparable datacenter GPU configurations.

  • A.3 Baseline Implementation: Official or publicly released codebases are used whenever available, with adaptations limited to WildCity data loading and preprocessing interfaces.All methods share the same camera intrinsics, poses, and train/validation split.
  • A.3 Baseline Implementation: 3DGS uses single-stage end-to-end optimization, H-3DGS uses base-stage and post-optimization training, and CityGS methods use multi-stage pipelines.CityGS pipelines include coarse optimization, spatial partitioning, block-wise trimming, merging, and final evaluation.
  • A.3 Baseline Implementation: 50k iterations are used for 3DGS on Ann Arbor-0.5k and 400k iterations on Atlanta-5k, with urban-specific components disabled.Disabled components include sky modeling, depth supervision, dynamic masking, and curriculum strategies.
  • A.3 Baseline Implementation: H-3DGS uses 50k base plus 15k post-optimization iterations on Ann Arbor-0.5k, and 30k base plus 15k post-optimization iterations on Atlanta-5k.These settings follow its released hierarchical pipeline.
  • A.3 Baseline Implementation: CityGS-based methods retain the released 4×4 spatial partition with block-wise parallel optimization and the downstream schedule when using aligned VGGT-Long reconstruction.The default COLMAP initialization is replaced by the aligned VGGT-Long reconstruction.
  • A.3 Baseline Implementation: All runs use comparable high-memory H100/H200-class datacenter GPUs rather than an artificially fixed device budget.The baselines are therefore reported under strong and stable configurations within their natural settings.

B.1 Evaluation Across Cities · B.2 Effect of the Number of Gaussian Primitives · B.3 Extrapolation Offset Study

Across cities, reconstruction robustness persists under a shared evaluation protocol, but urban morphology and scene complexity affect difficulty. Increasing Gaussian capacity improves quality while increasing computational cost, and off-trajectory rendering degrades for all methods as lateral offset grows.

  • B.1 Evaluation Across Cities: The evaluation tests the same method and protocol across cities with distinct urban characteristics, including dense downtowns, residential grids, arterial roads, and high-speed corridors.A cross-city setting keeps initialization and hyper-parameters identical to assess robustness across urban morphologies.
  • B.1 Evaluation Across Cities: Cities with denser intersections, stronger occlusions, and more heterogeneous layouts are more challenging than cities with regular roads and cleaner visibility.The findings indicate that morphology and operational complexity, not route length alone, shape city-scale reconstruction difficulty.
  • B.1 Evaluation Across Cities: Most cities achieve PSNR above 25 dB on 2.5km sequences despite variation from urban morphology and scene complexity.Atlanta is highlighted as more difficult because of highly textured street scenes and frequent dynamic objects.
  • B.2 Effect of the Number of Gaussian Primitives: 3M, 4M, 5M, and 6M Gaussian primitives define Density Levels 1–4 for Scale 0.1k, while larger scales use 7M–13M and 15M–30M ranges.These point-count settings are used to study Gaussian primitive numbers across scales.
  • B.2 Effect of the Number of Gaussian Primitives: Increasing Gaussian primitive density raises PSNR and steadily increases SSIM, while also increasing memory usage and optimization cost.The experiment holds the rest of training fixed and reports rendering quality, depth accuracy, peak GPU memory, and training time.
  • B.2 Effect of the Number of Gaussian Primitives: Gaussian growth can quickly become a bottleneck for city-scale logs because higher capacity improves quality only up to a point while substantially raising computational cost.This trade-off clarifies the practical cost of scaling current pipelines to large urban scenes.
  • B.3 Extrapolation Offset Study: At 1 m, 3 m, and 5 m lateral offsets, the visual quality of every method degrades as viewpoints move away from the recorded trajectory.The study uses these larger offsets to make degradation more visually explicit on long street-view trajectories.
  • B.3 Extrapolation Offset Study: CityGS and H-3DGS increasingly show blurred structures, broken ground surfaces, unstable sky boundaries, and floating artifacts as the offset grows.The qualitative comparisons use recorded-view references and off-trajectory renderings across Figures 11–15.
Loading 2607.06838v1…