Source-linked AI summary

GrandTour: A Legged Robotics Dataset in the Wild for Multi-Modal Perception and State Estimation

Turcan Tuna, Jonas Frey, Frank Fu, Katharine Patterson, Tianao Xu, Maurice Fallon, Cesar Cadena, Marco Hutter

arXiv:2602.18164v3cs.RO

TL;DR

Legged-robot research lacks a large public dataset capturing the multimodal sensing, environmental variation, and locomotion disturbances needed for state estimation and perception. GrandTour introduces such a dataset on an ANYmal-D with the Boxi payload, synchronized multimodal sensors, and high-precision ground truth, and releases it with tools and processed outputs for broad evaluation and learning applications. Its main scope is broad but geographically concentrated and lacks several modalities and annotations.

  • Problem

    Public legged-robot datasets remain insufficient in scale, sensing modalities, and environmental diversity for developing and benchmarking state estimation, perception, and navigation methods.

  • Method

    GrandTour collects synchronized and calibrated LiDAR, RGB, depth, inertial, GNSS, and proprioceptive data on an ANYmal-D equipped with the Boxi payload, paired with high-precision ground truth.

  • Results

    GrandTour provides 49 sequences across indoor, urban, and natural environments and supports systematic state-estimation benchmarks and diverse learning-based applications.

  • Takeaways & Limitations

    The open-access dataset, tools, and processed outputs provide a broad resource for multimodal fusion, SLAM, high-precision state estimation, perception, and navigation research.

  • Takeaways & Limitations

    GrandTour is geographically limited to Switzerland, lacks revisits across seasons, weather, and times of day, and currently omits radar, FMCW LiDAR, event cameras, and semantic labels.

Abstract

from arXiv · show

Accurate state estimation and multi-modal perception are prerequisites for autonomous legged robots in complex, large-scale environments. To date, no large-scale public legged-robot dataset captures the real-world conditions needed to develop and benchmark algorithms for legged-robot state estimation, perception, and navigation. To address this, we introduce the GrandTour dataset, a multi-modal legged-robotics dataset collected across challenging outdoor and indoor environments, featuring an ANYbotics ANYmal-D quadruped equipped with the Boxi multi-modal sensor payload. GrandTour spans a broad range of environments and operational scenarios across distinct test sites, ranging from alpine scenery and forests to demolished buildings and urban areas, and covers a wide variation in scale, complexity, illumination, and weather conditions. The dataset provides time-synchronized sensor data from spinning LiDARs, multiple RGB cameras with complementary characteristics, proprioceptive sensors, and stereo depth cameras. Moreover, it includes high-precision ground-truth trajectories from satellite-based RTK-GNSS and a Leica Geosystems total station. This dataset supports research in SLAM, high-precision state estimation, and multi-modal learning, enabling rigorous evaluation and development of new approaches to sensor fusion in legged robotic systems. With its extensive scope, GrandTour represents the largest open-access legged-robotics dataset to date. The dataset is available at https://grand-tour.leggedrobotics.com on HuggingFace (ROS-independent), and in ROS formats, along with tools and demo resources.

1 Introduction

GrandTour addresses the scarcity of comprehensive public datasets for legged-robot autonomy by providing a large, multimodal, open-access resource with synchronized sensing and high-precision ground truth. It supports systematic evaluation of state estimation, SLAM, perception, and sensor-fusion methods across realistic environments.

  • Motivation: Existing legged-robot datasets remain sparse or limited in scale, sensing modalities, and environmental diversity.These limitations often lead researchers toward small-scale scenario-specific data or simulation when developing mapping, state-estimation, and perception-to-action methods.
  • Motivation: Legged-robot autonomy faces disturbances from locomotion dynamics and harsh, unstructured environments that degrade sensing and estimation.Intermittent contacts, foot slippage, rapid body rotations, poor illumination, occlusions, tight spaces, slippery terrain, airborne particles, and textureless surfaces complicate reliable perception and state estimation.
  • Applications: GrandTour benchmarks LiDAR-inertial, visual-inertial, LiDAR-only, LiDAR-inertial-visual, and legged-robot-specific odometry across challenging terrains.These benchmarks support systematic evaluation of multimodal fusion and drift in state estimation and SLAM.
  • Applications: The dataset supports learning-based applications including sim-to-real transfer, real-to-sim reconstruction, physical-parameter learning, navigation foundation models, vision-language evaluation, and self-supervised representation learning.It also provides cross-view image pairs for visual-geometry tuning and relocalization in legged robots.
  • Dataset: GrandTour captures 49 sequences across indoor facilities, urban outdoors, and natural terrains with synchronized, calibrated multimodal sensing and high-precision ground truth.The platform combines multi-LiDAR, RGB and depth cameras, IMUs, GNSS, and full proprioception for benchmarking under realistic conditions.
  • Release: GrandTour is released openly with tools, documentation, and intermediate or processed outputs such as odometry, deskewed point clouds, terrain maps, and occupancy maps.These resources are intended to facilitate adoption and targeted algorithm development.

2 Related Work

Prior multimodal datasets established the value of synchronized sensing and accurate ground truth, but most target wheeled, aerial, handheld, or narrowly scoped platforms. GrandTour extends this dataset lineage with legged-specific sensing, rich proprioception, broad environments, and survey-grade localization.

  • Reference datasets: KITTI, EuRoC MAV, and TUM VI demonstrated the importance of synchronized multimodal data and high-quality ground truth for robotic perception and odometry.These datasets primarily target wheeled, aerial, or handheld platforms rather than legged locomotion.
  • Multimodal perception: Multimodal datasets combine complementary sensors because no single modality remains reliable across all environments and failure modes.LiDAR, visual, thermal, radar, event, inertial, and GNSS sensing can address lighting, weather, sparse texture, and dynamic-scene challenges.
  • Legged datasets: Existing legged-focused datasets such as TAIL, TAIL-Plus, DiTer, and FusionPortable provide valuable sensing and terrain coverage but remain comparatively specialized.Their platforms, terrains, mission structures, or modality combinations address narrower settings than a general-purpose legged benchmark.
  • GrandTour: GrandTour combines multiple LiDARs, RGB/depth cameras, high-grade IMUs, rich proprioception, dual RTK-GNSS, and survey-grade ground truth across urban, industrial, and natural settings.The dataset includes day/night operations and adverse weather conditions.

3 System overview

GrandTour combines the Boxi perception payload with an ANYmal quadruped to provide a compact, rigidly mounted, time-synchronized multimodal sensing platform. Its suite spans LiDAR, RGB and depth imaging, IMUs, proprioception, and multiple compute interfaces.

  • Platform: Boxi integrates exteroceptive and proprioceptive sensors in a rigid, thermally stable monoblock housing mounted on ANYmal.The payload includes two rotating LiDARs and ten RGB cameras, while the full suite weighs about 7.1 kg and draws roughly 120 W.
  • Cameras: The dataset includes seven depth-capable cameras: one ZED2i stereo RGB-D camera and six ANYmal-mounted Intel RealSense D435i cameras.These cameras provide complementary depth observations from the payload and robot body.
  • ANYmal sensing: ANYmal contributes six depth cameras, a VLP16 LiDAR, an IMU, 12 joint encoders, commanded velocity, and intermediate locomotion outputs.Recorded outputs include elevation maps, state-estimation data, and locomotion-policy hidden states.
  • Calibration: All sensors are rigidly mounted with known extrinsic transforms and share a common timing source to support multimodal fusion.The calibration methodology builds on procedures developed for the Boxi payload.
  • System architecture: The platform uses heterogeneous compute units and interfaces connected through a shared UbiSwitch Ethernet device.Sensors connect through interfaces including USB 3.1, GMSL2, RJ45, and I²C.

Proprioceptive

GrandTour addresses temporal and spatial calibration across heterogeneous sensors using multiple synchronization mechanisms and calibration procedures. These establish aligned camera, IMU, LiDAR, prism, and robot-frame measurements for multimodal sensing and ground-truth construction.

  • Proprioception: GrandTour combines high-rate IMUs and joint encoders with hardware or software-triggered timestamps for proprioceptive sensing.The IMU suite includes STIM320 at 500 Hz, ADIS16475-2 and TDK ICM40609 at 200 Hz, while ANYmal provides 12 joint positions, velocities, and torques.
  • Time synchronization: Sensor synchronization ranges from nanoseconds with IEEE 1588v2 PTP to milliseconds with custom kernel timestamping, except for the USB-buffered ZED2i camera.The NovAtel SPAN CPT7 acts as the time grandmaster, with PTP and NTP distributing timing across devices and compute units.
  • Camera timing: Rolling-shutter cameras are hardware-timestamped at exposure midpoints, while the CoreResearch camera uses PTP and exposure alignment across attached cameras.The ZED2i relies on a USB serial buffer and therefore has unreliable timestamps.
  • Camera calibration: Camera calibration estimates intrinsics and inter-camera extrinsics using April-Grid observations, a pinhole model, and lens-specific distortion models.Extrinsic optimization uses static intervals because the ten camera shutters do not trigger simultaneously.
  • Cross-modal calibration: IMU extrinsics are calibrated against front-facing global-shutter cameras, while LiDAR-camera calibration uses intensity alignment that can exploit partial target observations.This data-efficient formulation is particularly useful for the narrow-field-of-view Hesai sensor.
  • Ground-truth calibration: A custom camera-to-prism procedure calibrates the Boxi frame against Leica total-station measurements by optimizing a nine-degree-of-freedom state.The procedure uses static poses around a calibration target and excites all translational and rotational degrees of freedom.

4 Dataset overview

GrandTour comprises 49 missions across diverse environments and conditions, with accessible HuggingFace and ROS releases, calibrated sensor data, and high-precision fused ground truth.

  • 49 missions span natural and man-made environments, with durations of 3–7 minutes and distances up to several hundred meters.
  • Recorded conditions include day and night, sunshine, clouds, rain, sand, snow, and gravel.
  • The dataset offers HuggingFace access in Zarr and JPEG formats, with calibration metadata, individual-mission downloads, and ROS-independent processing scripts.
  • A ROS Bag release preserves robotics compatibility through per-stream files, LZ4 compression, and conversion scripts for ROS 2 Jazzy and Ubuntu 24.04.
  • Ground truth combines GNSS/INS processing, total-station measurements, and high-grade inertial sensing; Holistic Fusion restricts reference spans when only IMU dead reckoning would remain.

B ROS Data

The ROS data release organizes sensor streams and derived products for selective retrieval, while accompanying maps and fusion outputs support inspection and evaluation.

  • ROS topics include camera images and metadata, LiDAR odometry, undistorted point clouds, transforms, and map-related outputs.
  • The HuggingFace Zarr layout stores multi-array topics, timestamps, sequence identifiers, calibration metadata, and processed assets such as Gaussian splats and meshes.
  • Holistic Fusion standard-deviation estimates increase when MS60 line of sight is obstructed, while the fused trajectory follows absolute position measurements when available.
  • Users can select total-station position measurements for sparse translation evaluation or dense IE-TC six-degree-of-freedom poses when GNSS coverage is good.

5 Data Collection

GrandTour missions are collected by teleoperating an ANYmal with Boxi while coordinating total-station tracking, safety interventions, and reference-prism checks.

  • Each deployment uses a Leica MS60 total station and a robot teleoperator, with the robot traversing regions that vary in GNSS availability and total-station visibility.
  • Site preparation defines safe, diverse trajectories and positions the total station and reference prism for coverage and stability checks.
  • Recording begins with sensor validation, reference-prism measurement, and a total-station lock when the robot is within line of sight.
  • The teleoperator follows behind the robot, may intervene for safety, and relocates when tracking is lost, sometimes causing pauses for re-locking.
  • Mission completion includes a second reference-prism measurement to validate station stability before recording stops.

6 Collected Data and Derived Outputs

GrandTour records raw proprioceptive and sensor streams and provides derived odometry, motion-compensated point clouds, INS trajectories, and curated mapping outputs.

  • Online leg odometry uses ANYmal IMU and joint-estimation data while also estimating foot-contact states.
  • Post-processing supplies motion-compensated LiDAR clouds, real-time and tightly coupled INS trajectories, and Direct LiDAR-Inertial Odometry outputs.
  • Table 5 catalogs raw sensor streams, while Table 6 catalogs derived outputs and their descriptions.
  • Data curation includes filtering the robot body from LiDAR point clouds and preparing message formats for downstream use.

7 Dataset applications

GrandTour evaluates localization and state-estimation methods across difficult missions, revealing strong dependence on environment, sensing configuration, and failure recovery. The results support detailed per-mission analysis rather than relying only on aggregate rankings.

  • Benchmark missions: CON-4 and ARC-7 expose long-term drift, clutter, low light, smoke, and limited re-observation challenges for localization methods.CON-4 is a large, cluttered construction-site mission, while ARC-7 combines dark indoor–outdoor transitions, smoke, and difficult visual conditions.
  • LiDAR odometry: Traj-LO achieves the best overall average rank among LiDAR-odometry methods but degrades substantially on the confined ARC-2 mission.It leads SPX-2 and EIG-1 and also leads selected metrics on SNOW-2 and CON-4; several methods outperform it on ARC-2.
  • LiDAR-inertial-(visual) odometry: Coco-LIC and FAST-LIVO2 rank best overall, yet no single LiDAR-inertial or visual method dominates across all missions.The best method varies by mission, with different approaches achieving the best ATE or RTE on SPX-2, EIG-1, CON-4, ARC-7, and ARC-2.
  • Cross-method comparison: Many methods cluster closely on easier missions, so per-mission errors and failure cases are more informative than aggregate ranks alone.Differences among strong methods can be comparable to reported standard deviations, while a single mission may drive rank changes.
  • Cross-method comparison: Method-specific advantages are difficult to isolate on non-tailored real-world tests, especially when implementations contain many configuration and tuning choices.The benchmark includes methods with distinct operational properties, but feature-rich systems can be difficult to configure consistently across missions.
  • Multi-LiDAR-inertial: CTE-MLO is the most consistent multi-LiDAR method, while other approaches vary substantially across environments and some fail on ARC-2.CTE-MLO has the best average rank and lowest ATE on all reported missions, whereas RESPLE-MLO and FAST-LIO-MULTI fail on ARC-2.
  • Multi-LiDAR-inertial: Multi-LiDAR sensing can improve coverage and odometry robustness, but accuracy gains depend on the method, mission, and LiDAR characteristics.Different sensor noise and coverage properties affect how effectively a method benefits from additional LiDARs.
  • Visual-inertial odometry: Across visual-inertial missions, illumination, feature scarcity, initialization, and long trajectories produce divergence, drift, or widely varying absolute errors.Bright snow harms feature extraction, dynamic motion can impair initialization, large loops accumulate drift, and dark ARC intervals cause multiple pipelines to fail.

7.2 Perception

GrandTour supports perception research through synchronized multimodal sensing, benchmark references, and integrations spanning mapping, geometry, terrain understanding, and neural scene representations.

  • Contact perception: Figure 12 compares onboard binary contacts with DCE learned contacts and probabilities over a 10 s SNOW-2 window, marking mismatches by leg.The top row encodes signed differences: +1 for DCE-only contact and −1 for missed DCE contact.
  • Multimodal perception: GrandTour integrates perception tools for online mesh generation, elevation mapping, probabilistic occupancy mapping, and physical terrain-parameter estimation.These integrations connect LiDAR, depth, RGB, semantic, and proprioceptive information to downstream perception tasks.
  • Geometric understanding: Accurate LiDAR–camera synchronization enables ground-truth metric geometry for training and evaluating monocular depth estimation in natural environments.The same setup can be extended to stereo depth estimation.
  • Dynamic-scene processing: The dataset includes dynamic-point filtering based on pretrained image segmentation models, with masks available for volumetric mapping.Removing measurements associated with dynamic semantic classes supports more accurate volumetric maps.
  • Neural scene representations: GrandTour provides converters for NerfStudio, enabling neural scene-representation research with sparse and dense depth measurements.The converters directly transform GrandTour data into NerfStudio-compatible formats.

7.3 Locomotion and Navigation

GrandTour supports locomotion and navigation by providing realistic robot trajectories, terrain information, and visual data for planning, imitation, visual navigation, and real-to-sim transfer.

  • Vision-language navigation: NaviTrace uses GrandTour HDR images to build a visual question-answering benchmark for generating navigation behavior from annotated key frames and image traces.Human annotations specify navigation instructions and coordinate traces for selected frames.
  • Motion planning: An MPPI planner combines local traversability analysis with a goal position to output safe actions toward that goal.Black-box optimization can tune planner parameters against recorded reference behavior to reduce trajectory error.
  • Visual navigation: GrandTour’s dense geometric information enables arbitrary-goal path generation with MPPI, increasing navigation-trajectory diversity for visual navigation.Synthetic paths are provided on HuggingFace, and state-of-the-art visual navigation performance was reported with comparatively small training data.
  • Real-to-sim transfer: GaussGym performs real-to-sim transfer by integrating GrandTour deployment environments into simulation through neural radiance fields and scene reconstruction.This uses recorded environments to enhance simulation capabilities.

8 Discussion and limitations

GrandTour is broad but remains geographically and modally bounded, and its current annotations limit some supervised perception uses.

  • Geographic scope: GrandTour’s geographic scope is limited to Switzerland, potentially biasing architectural, vegetation, and ambient-lighting coverage.The limitation concerns how representative the dataset may be beyond its collection region.
  • Temporal coverage: The dataset lacks revisits of the same environments across seasons, weather patterns, or times of day.This restricts direct evaluation of changing conditions within identical scenes.
  • Modalities and annotations: GrandTour currently lacks radar, FMCW LiDAR, event cameras, semantic labels, and dynamic-object segmentation masks.The paper identifies community-contributed annotations as a future extension direction.

9 Conclusion and future work

GrandTour is a comprehensive, diverse multi-modal legged-robotics dataset with high-precision ground truth and tools supporting benchmarking across real-world missions. Its planned extensions target broader environments, robot platforms, and mapping evaluation.

  • Conclusion: GrandTour contains 49 sequences spanning diverse environments, illumination conditions, and weather types, exceeding existing legged-robot dataset scope and diversity.The dataset provides diverse sensor modalities, motion characteristics, and high-precision ground truth.
  • Conclusion: Centimeter-level GNSS-based poses and millimeter-level, time-synchronized total-station positions provide high-precision reference trajectories with accurate extrinsic calibration.GNSS-derived poses are available when GNSS is available; Leica MS60 measurements provide the higher-precision positions.
  • Conclusion: Benchmarking 63 state-estimation methods across six missions reports per-mission errors, ranks, and failure cases while exposing robustness, initialization, and tuning limits.The evaluated methods span LO, LIO/LIVO, multi-LiDAR(-inertial), VIO/RGB-D, and KIO categories under real-world legged motion and sensing conditions.
  • Conclusion: Open-source tools and post-processed outputs support sensor fusion, depth estimation, real-to-sim transfer, cross-modal learning, foundation-model training, traversability assessment, and navigation research.The dataset is intended to facilitate straightforward data access and analysis.
  • Future work: Planned extensions add extreme-perception missions, longer deployments, more environments and robot types, and Leica RTC-360 scans for mapping-quality evaluation.The planned robot expansion includes bipeds and wheeled-legged robots.
  • Future work: These expansions are intended to help bridge controlled laboratory experiments and real-world legged-robot deployment scenarios.The stated future direction positions GrandTour as an ongoing benchmark and resource for robotics and computer-vision research.
Loading 2602.18164v3…