Source-linked AI summary
Are We Ready for Service Robots? The OpenLORIS-Scene Datasets for Lifelong SLAM
Xuesong Shi, Dongjiang Li, Pengpeng Zhao, Qinbin Tian, Yuxin Tian, Qiwei Long, Chunhao Zhu, Jingwei Song, Fei Qiao, Le Song, Yangquan Guo, Zhigang Wang, Yimin Zhang, Baoxing Qin, Wei Yang, Fangshi Wang, Rosa H. M. Chan, Qi She
TL;DR
Service-robot SLAM must remain localized and reuse maps across repeated operations in changing environments, but conventional evaluation largely assumes short, fresh-start runs. The paper releases OpenLORIS-Scene, repeated real-world indoor sequences with multimodal sensors, and introduces separate robustness and accuracy metrics; existing systems face the included scene changes and the benchmark exposes their localization shortcomings.
Problem
Most SLAM systems are designed and evaluated for a single fresh-start operation, while service robots need persistent map reuse and localization across changing environments.
Method
The paper builds OpenLORIS-Scene from repeated real-world indoor recordings with people and scene changes, and proposes metrics that evaluate localization robustness separately from accuracy.
Results
Existing SLAM systems find the datasets' scene changes challenging, and the proposed metrics separate robustness and accuracy when benchmarking localization.
Takeaways & Limitations
OpenLORIS-Scene provides a testbed for identifying SLAM shortcomings and assessing the maturity of algorithms for real-world service-robot deployment.
Abstract
from arXiv · showhide
Service robots should be able to operate autonomously in dynamic and daily changing environments over an extended period of time. While Simultaneous Localization And Mapping (SLAM) is one of the most fundamental problems for robotic autonomy, most existing SLAM works are evaluated with data sequences that are recorded in a short period of time. In real-world deployment, there can be out-of-sight scene changes caused by both natural factors and human activities. For example, in home scenarios, most objects may be movable, replaceable or deformable, and the visual features of the same place may be significantly different in some successive days. Such out-of-sight dynamics pose great challenges to the robustness of pose estimation, and hence a robot's long-term deployment and operation. To differentiate the forementioned problem from the conventional works which are usually evaluated in a static setting in a single run, the term \textit{lifelong SLAM} is used here to address SLAM problems in an ever-changing environment over a long period of time. To accelerate lifelong SLAM research, we release the OpenLORIS-Scene datasets. The data are collected in real-world indoor scenes, for multiple times in each place to include scene changes in real life. We also design benchmarking metrics for lifelong SLAM, with which the robustness and accuracy of pose estimation are evaluated separately. The datasets and benchmark are available online at https://lifelong-robotic-vision.github.io/dataset/scene.
I. INTRODUCTION
Lifelong SLAM addresses persistent localization and map reuse when service robots revisit regions across changing real-world conditions. OpenLORIS-Scene provides repeated indoor-scene data and separate robustness-oriented metrics to support this research.
- Motivation: Most SLAM systems assume a single operation from a fresh start, whereas service robots must reuse persistent maps day after day.Long-term deployment also requires retaining spatial knowledge and coordinate consistency despite scene changes and uncontrolled factors.
- Definition: Lifelong SLAM means building and maintaining a persistent regional map while continuously localizing the robot across operations, even when the environment changes.The map must be reused between operations rather than merely saved and loaded.
- Challenges: Key challenges include changed viewpoints, altered objects, illumination shifts, dynamic objects, and degraded or miscalibrated sensors.These factors can arise from human activity, natural conditions, mechanical stress, temperature, or dirty and wet lenses.
- Dataset contribution: OpenLORIS-Scene fills a lack of public datasets and benchmarks for unifying efforts toward practical lifelong SLAM systems.The data are collected with commodity sensors on a wheeled robot in typical indoor environments, with ground-truth poses from motion capture or high-accuracy LiDAR.
- Dataset contribution: The datasets include real-world scenes with people and multiple sequences per scene, capturing illumination, viewpoint, and human-activity-driven scene changes.Color-image examples show approximately the same place in different sequences after the scene changed.
- Benchmark contribution: A rich sensor combination supports comparisons across input types, while new correct-rate metrics evaluate localization robustness separately from accuracy.The benchmark explicitly treats robustness as a central concern rather than leaving it only implicit in accuracy metrics.
II. RELATED WORKS
Prior SLAM benchmarks provide useful sensor and environmental variation but generally do not unify lifelong evaluation around realistic scene changes and localization benchmarking. OpenLORIS-Scene extends this direction with real-world data and new performance metrics.
- Existing benchmarks: Common SLAM datasets such as TUM RGB-D, EuRoC MAV, KITTI, and TUM VI support evaluation across particular aspects and sensor types.Public datasets have generally been used to justify algorithm effectiveness in selected aspects.
- Existing benchmarks: OpenLORIS-Scene provides aligned RGB-D-IMU data and odometry, addressing a stated lack of public datasets combining RGB-D and IMU data.Odometry is widely used in industry but is often absent from public datasets.
- Scene-change datasets: Synthetic rendering could model scene changes in principle, but realistic changes occurring in natural life are difficult to reproduce.This motivates collecting real-world scene-change data for lifelong SLAM evaluation.
- Scene-change datasets: The COLD database contains real-world visual variation from weather, illumination, and human activities, while change-detection datasets are not designed for SLAM or ground-truth camera poses.OpenLORIS-Scene follows a related real-world collection principle with different sensor setups.
- Benchmarking direction: The work contributes new data and performance metrics to ongoing efforts toward unified SLAM benchmarking and automatic parameter tuning.Its benchmark is positioned as an extension of unified evaluation efforts.
III. OPENLORIS-SCENE DATASETS
OpenLORIS-Scene is designed as a real-world practicality testbed for lifelong SLAM in service-robot scenarios. It uses wheeled robots, commodity sensors, typical indoor scenes with people, and calibrated, synchronized multimodal data.
- Design principles: The datasets target real-world practicality by approximating service-robot scenarios in typical indoor scenes with people.Commercial wheeled robots equipped with commodity sensors collect the data.
- Design principles: Rich data types enable comparisons among SLAM methods that accept different kinds of inputs.All provided data are calibrated and synchronized.
A. Robots and Sensors
The collection platform combines complementary cameras, IMUs, wheel odometry, and LiDAR on wheeled robots. These sensors support multiple SLAM input modalities and provide calibrated measurements and ground-truth trajectories.
- Cameras and IMUs: A RealSense D435i supplies RGB-D images and IMU measurements, while a RealSense T265 supplies stereo fisheye images and IMU measurements.The setup supports monocular, stereo, RGB-D, and visual-inertial SLAM; each device synchronizes its IMU with its images.
- Odometry: Wheel encoder-based odometry is provided because it is widely available on wheeled robots.The Segway odometry fuses wheel encoders with a chassis IMU using proprietary filtering algorithms.
- Ground truth: Ground-truth trajectories come from an OptiTrack motion-capture system and Hokuyo LiDAR on the Segway, or RoboSense RS-LiDAR-16 on the Gaussian robot.The ground-truth sensing equipment is mounted or positioned near the cameras.
- Data organization: Table I organizes the dataset's available data types for comparing algorithms with different inputs.The surrounding dataset description identifies RGB-D, stereo fisheye, IMU, wheel odometry, and LiDAR modalities.
- Sensor naming: D435i denotes its color camera, whereas T265 denotes its left fisheye camera.These labels distinguish the camera streams associated with the two RealSense devices.
B. Calibration
The datasets combine calibrated, synchronized multi-sensor recordings with repeated sequences designed to expose lifelong SLAM challenges, including viewpoint, illumination, object, and human-induced changes. They contain 22 sequences spanning 2244 seconds across five indoor scenes.
- Calibration: Each non-camera sensor is calibrated against both cameras, enabling extrinsic consistency checks; most calibration errors stay below 1 cm in translation and 2° in rotation.Odometry calibration has a larger 7 cm translation error.
- Synchronization: Hardware synchronization covers RGB-D and IMU measurements within each RealSense device, while software synchronization aligns data across devices.The alignment includes RealSense D435i, RealSense T265, LiDAR, MCS, and odometer data.
- Synchronization: Synchronization uses controlled back-and-forth motion in a static, feature-rich area to reduce effects from SLAM and measurement noise.Only this controlled segment is used for synchronization.
- Scenes and Sequences: The benchmark contains five scenes with 2–7 sequences each, manually selected and clipped to represent major lifelong SLAM challenges.Some changes were deliberately introduced but were chosen to reflect changes likely over longer periods.
- Scenes and Sequences: The scenes include changed viewpoints, illumination, moved furniture, altered household objects, dynamic people, and changed supermarket goods.These variations occur across office, corridor, home, cafe, and market recordings.
- Dataset Size: The released data comprise 22 sequences with an accumulated duration of 2244 seconds.
E. Ground-truth
Ground-truth robot poses are provided for every sequence in a persistent scene map, using motion-capture data for the office and 2D laser SLAM for other scenes.
- Ground-truth Sources: Ground-truth poses are supplied for all sequences in each scene within a persistent map.
- Ground-truth Sources: Office ground truth comes from a motion-capture system covering all sequences in a persistent coordinate system at 240 Hz, with outliers removed.
- Ground-truth Sources: For the other scenes, a 2D laser SLAM method builds a full scene map and localizes the robot using each sequence frame's laser scan.
IV. BENCHMARK METRICS
The benchmark separates pose-estimation correctness from accuracy because failures and mismatched poses are especially severe under lifelong scene changes. It evaluates correctness per pose and robustness over trajectories, including tracking and re-localization behavior.
- Correctness: ATE and absolute orientation error (AOE) determine whether each estimated pose is correct against its ground-truth pose.
- Correct Rate: Correct Rate (CR) measures correctness over the full sequence time span, while Correct Rate of Tracking (CR-T) excludes initialization and re-localization time.The validity duration of a correct estimate is controlled by δ.
- Parameters: Thresholds ε and φ define acceptable translational and orientation errors, while δ determines how long a correct pose estimate remains valid.For common room or building-sized data, the authors suggest meter-scale ε and δ around one second.
- Re-localization: Correctness Score of Re-localization (CS-R) evaluates both re-localization correctness and the time required to recover localization.An immediate correct re-localization gives CS-R = 1, and the score decreases as re-localization takes longer.
B. Accuracy Metrics
Accuracy is measured using ATE and RPE statistics computed only from pose estimates that pass a correctness threshold, preventing incorrect results from affecting accuracy assessment.
- Accuracy Evaluation: ATE and RPE evaluate pose-estimation accuracy over one or more trajectories after incorrect estimates are excluded.
- Accuracy Evaluation: C0.1-RPE RMSE is the root mean square relative pose error for estimates selected by an ATE threshold of 0.1 meter.
- Accuracy Evaluation: Separating correctness filtering from accuracy statistics distinguishes wrong poses from inaccuracies among correct estimates.
V. EXPERIMENTS
The datasets and proposed metrics are evaluated across diverse open-source SLAM systems, showing that robustness and accuracy vary by scene and should be assessed separately.
- Algorithms and setup: The experiments test ORB-SLAM2, DSO, DS-SLAM, VINS-Mono, InfiniTAM, and ElasticFusion across the available sensor modalities.The algorithms represent feature-based, direct, dynamic-object-aware, visual-inertial, dense, and globally consistent mapping techniques.
- Per-sequence evaluation: Per-sequence evaluation aligns estimated and ground-truth trajectories before computing ATE RMSE, with interpolation providing exact timeline matches.DSO additionally uses optimal scaling, while scene-level statistics weight CR∞ by sequence duration and ATE RMSE by pose-estimate count.
- Metrics: CR∞ measures successful localization and ATE RMSE measures accuracy; larger CR∞ indicates greater robustness, whereas smaller ATE RMSE indicates greater accuracy.Figure 2 visualizes successful initialization and tracking with blue dots and lines.
- Results: VINS-Mono shows the best robustness among tested algorithms, although it fails to initialize in some low-light corridor sequences.Most algorithms track successfully in office but struggle in other scenes, especially featureless, low-light corridors.
- Results: Wheel odometry provides reliable tracking even in large scenes, supporting its consideration in practical service-robot SLAM design.The odometry data are evaluated alongside the SLAM algorithms in Figure 2.
- Metrics discussion: ATE and CR∞ exhibit a consistent negative correlation for similar algorithms, so evaluating only one metric can misrepresent performance.Longer tracking can accumulate more error, while high CR∞ can still accompany erroneous trajectories when the ATE threshold is inappropriate.
B. Lifelong SLAM Evaluation
Lifelong evaluation feeds sequences from the same scene consecutively to test re-localization under changing conditions. Re-localization is difficult, especially for viewpoint and illumination changes, while the proposed metrics retain known limitations on large scenes.
- B. Lifelong SLAM Evaluation: Sequences from each scene are fed consecutively, requiring algorithms either to re-localize or align a fresh map with the previous map.DSO and ElasticFusion are excluded because the evaluated implementations do not support re-localization.
- Results: Most algorithms completely fail to re-localize in the second through fifth home sequences, demonstrating the difficulty of lifelong re-localization.Incorrect pose estimates are identified using ATE thresholds of 1, 3, or 5 meters and an AOE threshold of 30°.
- Metrics discussion: The metrics can produce false alarms at the initial and final parts of some large-scene sequences because accumulated drift distorts the aligned trajectory.The authors suggest larger ATE thresholds for large scenes and further refinement of the accuracy judgment method.
- Factor analysis: Controlled office tests indicate that changed viewpoints and illumination are the most difficult factors for re-localization.The office sequences are designed to disentangle viewpoint, illumination, changed things, and dynamic-object effects.
VI. CONCLUSION
The paper introduces OpenLORIS-Scene datasets and metrics for benchmarking lifelong SLAM in changing indoor environments. The authors report that these conditions challenge existing systems and may support broader long-term scene-understanding research.
- VI. CONCLUSION: OpenLORIS-Scene captures day-night shifts, human-activity changes, viewpoint changes, moving people, poor illumination, and blur in real-world scenes.These factors were challenging for existing SLAM systems.
- VI. CONCLUSION: The proposed metrics evaluate localization robustness and accuracy separately for lifelong SLAM.The benchmark is intended to help identify shortcomings and encourage more robust localization designs.
- VI. CONCLUSION: The datasets are positioned as a testbed for assessing the maturity of future SLAM algorithms for real-world service-robot deployment.The supported scope is lifelong localization in changing environments.
- VI. CONCLUSION: With suitable annotation, the data may support incremental-learning benchmarks and exploration of spatio-temporal modeling for long-term scene understanding.The paper presents these as possible extensions beyond SLAM.