Source-linked AI summary

Socially Compliant Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Navigation

Haresh Karnan, Anirudh Nair, Xuesu Xiao, Garrett Warnell, Soeren Pirk, Alexander Toshev, Justin Hart, Joydeep Biswas, Peter Stone

arXiv:2203.15041v2cs.ROcs.CVcs.LGeess.SY

TL;DR

Social-navigation imitation learning is hindered by the lack of large-scale, real-world demonstrations of socially compliant behavior. The paper introduces SCAND, a multimodal dataset collected from two robot morphologies, and evaluates imitation-learned policies. The dataset supports socially compliant behaviors and distinguishes multiple compliant navigation strategies.

  • Problem

    Large-scale datasets of socially compliant robot-navigation demonstrations in the wild are lacking, limiting available evidence for imitation learning.

  • Method

    SCAND collects human-teleoperated demonstrations with multimodal sensors, joystick commands, odometry, and labeled social interactions on two morphologically different robots.

  • Results

    Imitation-learned policies trained on SCAND generate socially compliant behavior, while a classifier identifies demonstrators with 74.48% accuracy and participants rate the SCAND agent safer and more socially compliant than move_base.

  • Takeaways & Limitations

    SCAND provides demonstrations and annotations for studying socially compliant navigation and multiple socially compliant navigation strategies.

  • Takeaways & Limitations

    SCAND was collected in Austin and may encode regional norms such as staying to the right of the road or overtaking pedestrians from the left.

Abstract

from arXiv · show

Social navigation is the capability of an autonomous agent, such as a robot, to navigate in a 'socially compliant' manner in the presence of other intelligent agents such as humans. With the emergence of autonomously navigating mobile robots in human populated environments (e.g., domestic service robots in homes and restaurants and food delivery robots on public sidewalks), incorporating socially compliant navigation behaviors on these robots becomes critical to ensuring safe and comfortable human robot coexistence. To address this challenge, imitation learning is a promising framework, since it is easier for humans to demonstrate the task of social navigation rather than to formulate reward functions that accurately capture the complex multi objective setting of social navigation. The use of imitation learning and inverse reinforcement learning to social navigation for mobile robots, however, is currently hindered by a lack of large scale datasets that capture socially compliant robot navigation demonstrations in the wild. To fill this gap, we introduce Socially CompliAnt Navigation Dataset (SCAND) a large scale, first person view dataset of socially compliant navigation demonstrations. Our dataset contains 8.7 hours, 138 trajectories, 25 miles of socially compliant, human teleoperated driving demonstrations that comprises multi modal data streams including 3D lidar, joystick commands, odometry, visual and inertial information, collected on two morphologically different mobile robots a Boston Dynamics Spot and a Clearpath Jackal by four different human demonstrators in both indoor and outdoor environments. We additionally perform preliminary analysis and validation through real world robot experiments and show that navigation policies learned by imitation learning on SCAND generate socially compliant behaviors

I. INTRODUCTION

Social navigation requires robots to adjust their paths around other agents, but existing demonstrations lack large-scale, naturally occurring socially compliant behavior. SCAND addresses this gap with multimodal teleoperated demonstrations collected across robots and environments.

  • Social navigation involves recognizing and reacting to other agents’ objectives while adjusting the robot’s path and projecting signals that help reciprocal interaction.
  • Demonstration data can support Learning from Demonstrations and help researchers understand human navigation around autonomous robots.
  • Existing social-navigation datasets contain limited interactions in constrained settings or focus exclusively on indoor navigation, omitting naturally occurring behaviors.
  • Imitation learning avoids explicitly defining complex social-navigation rules, yet large-scale datasets of socially compliant demonstrations in the wild remain unavailable.
  • 8.7 hours, 138 trajectories, and 25 miles of multimodal demonstrations were collected from two morphologically different robots, including lidar, joystick, odometry, visual, and inertial data.

II. RELATED WORK

Related work applies learning to robot navigation through imitation, reinforcement, and inverse reinforcement learning. Existing social-navigation approaches rely on simulation, restricted scenarios, or modeled experts rather than broad real-world demonstrations.

  • Learning-based robot-navigation methods address adaptive planner parameters, viewpoint invariance in demonstrations, and end-to-end autonomous driving.
  • Tai et al. use simulation and the social force model to generate demonstrations, then train a social-navigation policy with Generative Adversarial Imitation Learning.
  • Reinforcement-learning approaches have shown real-world results but remain limited to specific social scenarios and require simulated episodic learning.
  • Inverse reinforcement learning has been used to learn cost functions for socially compliant navigation from human demonstrations.

B. Datasets for Social Navigation

Simulation provides fast, controllable social-navigation data collection, but simulated environments do not reproduce naturally occurring real-world interactions.

  • Simulated social environments enable fast data collection for social navigation.
  • Simulation can specify the number and locations of humans, room structure, objects, and interactions between people and objects.
  • Simulated platforms are limited because they lack natural real-world interactions.

2) Real-world Datasets for Robot Navigation:

Real-world robot datasets support perception and navigation research, but prior teleoperated datasets generally lack explicit social-compliance demonstrations and navigation strategies. SCAND adds multimodal, labeled demonstrations across different robot morphologies.

  • 2) Real-world Datasets for Robot Navigation:: Prior real-world datasets include laser scans, odometry, localization, LiDAR, RGBD, GPS, and IMU data for long-term robot navigation and perception.
  • 2) Real-world Datasets for Robot Navigation:: Teleoperated demonstrations in prior datasets are not explicitly socially compliant, limiting their use for studying socially compliant navigation strategies.
  • 2) Real-world Datasets for Robot Navigation:: JRDB contains 64 minutes from 54 indoor and outdoor trajectories and focuses on perception tasks, whereas SCAND focuses on the navigation component of social navigation.
  • 2) Real-world Datasets for Robot Navigation:: SCAND provides joystick commands, multimodal sensor data, and labels for twelve social interactions across naturally occurring scenarios.
  • 2) Real-world Datasets for Robot Navigation:: Data from the legged Spot and wheeled Jackal supports investigation of how robot morphology affects navigation and induced social interactions.

III. DATA COLLECTION PROCEDURE

SCAND collects socially compliant, human-teleoperated navigation demonstrations in campus environments using two morphologically different robots and rich multimodal sensing. The dataset includes natural social scenarios, joystick commands, and interaction annotations focused on navigation rather than human detection or tracking.

  • Collection setting: Four human demonstrators teleoperated Jackal and Spot robots across 138 trajectories on the University of Texas at Austin campus.Data collection included sidewalks, roads, lawns, and building interiors with people present during high-traffic periods.
  • Collection procedure: The human operator followed behind each robot at an average distance of two meters while teleoperating with a joystick.This procedure was used for every trajectory in SCAND.
  • Social scenarios: The demonstrations cover naturally occurring social navigation scenarios, including street crossings, narrow doorways, crowds, vehicle interactions, and stationary queues.Figure 2 pairs RGB imagery with lidar and Spot side-camera views for five example scenarios.
  • Sensor suite: SCAND records Velodyne lidar, joystick commands, odometry, camera visuals, and 6D inertial information from both robots.Jackal additionally provides stereo imagery and wheel odometry, while Spot provides five body-mounted monocular cameras and visual odometry.
  • Annotation scope: SCAND provides navigation-relevant multimodal data but no labeled annotations for human detection or tracking.The dataset instead includes demonstrator velocity commands and labels for 12 social interactions in every trajectory.

B. Labeled Annotations of Social Interactions

SCAND annotates trajectories with textual labels for social interactions and supports analyses of navigation strategies and learned planning. Its multimodal inputs include robot sensors, joystick commands, and interaction-related information.

  • Labeled Annotations of Social Interactions: Each SCAND trajectory is annotated with textual captions selected from twelve predefined social-interaction labels.The labels describe interactions occurring along the path and are intended to support studies of specific real-world navigation scenarios.
  • Labeled Annotations of Social Interactions: The analysis asks whether socially compliant navigation has multiple strategies and whether SCAND can support jointly learned global and local planners.These questions are addressed through demonstrator classification and behavior cloning.
  • Labeled Annotations of Social Interactions: The dataset’s multimodal information includes lidar, monocular and RGB imagery, inertial sensing, and joystick commands issued during demonstrations.Figure 3 presents sensors on the Jackal and Spot robots, while joystick commands accompany the sensor streams.
  • Labeled Annotations of Social Interactions: The demonstrator classifier uses ten-second sensor-observation sequences to predict the identity of the human demonstrator.The behavior-cloning agent uses a related architecture with global- and local-planner heads instead of a classifier head.

A. Demonstrator Classification

The paper hypothesizes that a given environment can support more than one socially compliant navigation strategy.

  • A. Demonstrator Classification: The authors hypothesize that multiple socially compliant navigation strategies can exist in the same scenario.The question concerns whether different strategies are possible for navigating socially in an environment.

1) Approach and Implementation:

The study tests whether navigation style reveals which demonstrator drove a trajectory by training a neural classifier on two people navigating the same route.

  • Approach and Implementation: Sixteen trajectories from two demonstrators on the same Speedway road route were used, with twelve for training and four for validation.The classifier was trained to identify the demonstrator from navigation data.
  • Approach and Implementation: The classifier processes ten-second sequences containing lidar BEV images, relative robot positions, future trajectories, inertial values, and joystick commands.A convolutional encoder handles lidar images, while fully connected networks process other observations and produce the classification output.
  • Results and Conclusion: 74.48% accuracy was achieved when classifying the demonstrator on the held-out test set.The paper contrasts this with a 50% random-guessing success rate and interprets the result as evidence of distinguishable navigation styles.

1) Approach and Implementation:

The approach trains behavior-cloning global and local planners on SCAND and evaluates them against move_base through trajectory comparison and human trials. The learned policies more closely reproduce socially compliant demonstrations and receive higher social-compliance and safety ratings, while simple scenarios leave room for stronger imitation methods.

  • Approach and Implementation: Behavior Cloning jointly trains end-to-end global and local planners from SCAND demonstrations.The global planner predicts a future socially compliant trajectory, while the local planner predicts demonstrated forward and angular velocities.
  • Results: Fig. 5 contrasts demonstrated, move_base, and learned trajectories as a human crosses the robot’s path.The learned path closely follows the socially compliant demonstration, whereas move_base turns toward the human’s future state.
  • Results: 0.26 average Hausdorff distance separates the learned global planner’s predicted trajectory from the demonstrated path, versus 1.25 for move_base.The evaluation uses a held-out test set and compares predicted or global paths with the demonstrator’s future path.
  • Results: Fourteen participants evaluated static and dynamic human-robot scenarios for perceived social compliance and safety.The static scenario used a stationary human, while the dynamic scenario involved the robot and human moving toward each other’s starting positions.
  • Results: SCAND’s policy scored higher than move_base for social compliance and safety in the human evaluation.Mean social-compliance scores were 4.39 versus 2.86, and safety scores were 4.71 versus 2.89; both comparisons were statistically significant by one-way ANOVA.
  • Limitations: The BC agent handles simple social-navigation scenarios, but more sophisticated scenarios in SCAND may require better imitation-learning algorithms.This is identified as a limitation of the demonstrated validation scope.

V. ANTICIPATED USE CASES

SCAND supports future work on benchmarking, real-to-sim transfer, trajectory analysis, and inverse reinforcement learning. Its broad scenario variety does not guarantee coverage of infrequent interactions, and single-city collection may introduce regional norms that affect generalization.

  • Generalization: Infrequent novel interactions may limit generalizability to unseen situations.The paper proposes representation learning with SCAND as one promising direction for improving generalization.
  • Generalization: Collection in Austin may encode regional norms such as staying right or overtaking pedestrians from the left.The paper calls for algorithms, evaluations, and metrics flexible enough to accommodate different local norms.
  • Benchmarking and Transfer: SCAND could augment simulation benchmarks with human-robot interaction trajectories and support real-to-sim transfer.The paper identifies both directions as future work because benchmarking is outside its scope.
  • Learning Applications: SCAND enables future trajectory prediction, trajectory classification, and inverse reinforcement learning for large-scale socially compliant cost-function learning.These directions extend earlier work based on human-only or robot-only trajectories.

VI. CONCLUSION

The paper introduces SCAND, a large-scale multimodal dataset of socially compliant mobile-robot navigation demonstrations. It uses the dataset to study navigation strategies and train behavior-cloning planners, with human trials indicating higher perceived social compliance and safety than move_base.

  • Conclusion: SCAND contains 8.7 hours, 138 trajectories, and 25 miles of demonstrations collected on two morphologically different robots.It also includes multimodal sensory streams and labeled annotations of social interactions.
  • Conclusion: A neural-network demonstrator-classification task shows that SCAND contains more than one socially compliant navigation strategy.The conclusion presents this as the first of several questions addressed with the dataset.
  • Conclusion: Behavior cloning on SCAND learns socially compliant global and local planners for mobile-robot navigation.The conclusion reports this as evidence that the dataset supports learning both planning levels.
  • Conclusion: Human trials on two social-navigation scenarios find the behavior-cloned local planner relatively more socially compliant and safe.The conclusion summarizes the validation of the local planner through human evaluation.
Loading 2203.15041v2…