Source-linked AI summary
INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps
Wei Zhan, Liting Sun, Di Wang, Haojie Shi, Aubrey Clausse, Maximilian Naumann, Julius Kummerle, Hendrik Konigshof, Christoph Stiller, Arnaud de La Fortelle, Masayoshi Tomizuka
TL;DR
Behavior-related research needs interactive motion data covering complex scenarios and different driving cultures, but existing datasets have limited diversity, interaction completeness, and semantic map coverage. The paper constructs the INTERACTION dataset from drone and traffic-camera recordings across countries, combining diverse and critical behaviors with complete semantic maps. It provides data for motion prediction, planning, behavior modeling, imitation, representation learning, interaction extraction, and social behavior generation.
Problem
Behavior-related research requires interactive motion data, but existing datasets provide limited scenario diversity, interaction completeness, and semantic map coverage.
Method
The paper constructs the INTERACTION dataset from drone and traffic-camera recordings across countries, with diverse scenarios, complex behaviors, critical situations, and semantic maps.
Results
The dataset contains diverse international driving scenarios, adversarial and cooperative interactions, near-collisions and slight collisions, complete interaction entities, and semantic maps.
Takeaways & Limitations
The dataset supports research on motion prediction, planning, behavior modeling, imitation, representation learning, interaction extraction, and social behavior generation.
Abstract
from arXiv · showhide
Behavior-related research areas such as motion prediction/planning, representation/imitation learning, behavior modeling/generation, and algorithm testing, require support from high-quality motion datasets containing interactive driving scenarios with different driving cultures. In this paper, we present an INTERnational, Adversarial and Cooperative moTION dataset (INTERACTION dataset) in interactive driving scenarios with semantic maps. Five features of the dataset are highlighted. 1) The interactive driving scenarios are diverse, including urban/highway/ramp merging and lane changes, roundabouts with yield/stop signs, signalized intersections, intersections with one/two/all-way stops, etc. 2) Motion data from different countries and different continents are collected so that driving preferences and styles in different cultures are naturally included. 3) The driving behavior is highly interactive and complex with adversarial and cooperative motions of various traffic participants. Highly complex behavior such as negotiations, aggressive/irrational decisions and traffic rule violations are densely contained in the dataset, while regular behavior can also be found from cautious car-following, stop, left/right/U-turn to rational lane-change and cycling and pedestrian crossing, etc. 4) The levels of criticality span wide, from regular safe operations to dangerous, near-collision maneuvers. Real collision, although relatively slight, is also included. 5) Maps with complete semantic information are provided with physical layers, reference lines, lanelet connections and traffic rules. The data is recorded from drones and traffic cameras. Statistics of the dataset in terms of number of entities and interaction density are also provided, along with some utilization examples in a variety of behavior-related research areas. The dataset can be downloaded via https://interaction-dataset.com.
I. INTRODUCTION
Behavior-related autonomous-driving research depends on comprehensive, accurate motion data, but existing public datasets restrict research through limited scenario diversity, interaction complexity, criticality, cultural coverage, and map information.
- Accurate prediction and comprehensive understanding of other road users are required for fully autonomous driving in complex scenarios.
- Motion datasets support prediction, behavior modeling, motion representation, imitation learning, and human-like behavior generation.
- Public datasets such as NGSIM and highD facilitated behavior-related research but restricted it through limited diversity, complexity, criticality, map information, and interaction-entity completeness.
- Existing public datasets are concentrated largely on highway scenarios, excluding many highly interactive settings such as roundabouts and unsignalized intersections.
- Datasets from different countries are needed to represent cultural differences in driving styles, preferences, risk tolerance, and traffic-rule understanding.
- Existing scenarios are often simple and structured, with explicit right-of-way, cautious behavior, little social pressure, and few aggressive or irrational decisions.
4) Criticality of the situations:
The paper motivates an international dataset designed to address sparse critical interactions, incomplete surrounding-entity coverage, and missing semantic maps in existing motion data.
- 4) Criticality of the situations:: Critical near-collision situations are sparse in existing datasets, limiting research on behavior prediction and other difficult problems.
- Semantic maps: Semantic maps containing references, lanelet connections, and traffic rules provide features required for motion prediction and planning and support generalization to other scenarios.
- 6) Completeness of interaction entities:: Complete motions of surrounding entities are needed to model, predict, and imitate interactive vehicle behavior, but onboard sensors suffer from occlusions and limited field of view.
- Dataset construction: The proposed dataset is collected by drones and traffic cameras to emphasize international coverage, complete interactions, critical situations, and semantic maps.
- Diverse and international: It includes diverse scenarios from different countries, including roundabouts, signalized and unsignalized intersections, and urban or highway merging and lane changes.
- Complex and critical: The dataset contains aggressive or irrational behavior, social-pressure effects, near-collisions, slight collisions, and motions of entities that may influence driving behavior.
- Existing datasets: NGSIM and highD provide useful vehicle-motion data, but their scenarios are limited and interactions at signalized intersections are rare.
NGSIM [13]
Existing motion datasets provide useful data but differ in sensing setup and coverage, creating limitations in interaction completeness, scenario diversity, repetition, and map information.
- Onboard-sensor datasets include surrounding motions from LiDAR and cameras or many data-collection vehicles from GPS.
- Onboard sensors can capture diverse, long-duration scenarios and partially recover ego-vehicle-view occlusions.
- GPS-based fleets lack recordings of surrounding vehicles or pedestrians without GPS devices, making actual interactions difficult to determine.
- LiDAR- and camera-based datasets cannot guarantee inclusion of all surrounding objects affecting behavior, especially when sensors have limited field of view.
- Large-area data collection can produce few repetitions at the same location, making multimodal driving behavior difficult to learn for prediction or planning.
- Most motion datasets lack map information, while Argoverse provides relatively rich physical and partially semantic map information.
- The proposed dataset offers more diverse, complex, and critical scenarios, full-semantic HD maps, and superior interaction-entity completeness compared with three existing datasets.
III. FEATURES OF THE DATASET
The INTERACTION dataset covers diverse, highly interactive scenarios across road types, traffic controls, and countries, with detailed examples of merging, roundabouts, intersections, and international comparisons.
- A. Diversity: The dataset includes urban zipper merging, highway ramp merging and lane changes, five roundabouts, unsignalized intersections, and an unprotected left turn at a signalized intersection.Figure 2 identifies the scenario categories and their traffic-control settings.
- A. Diversity: The highway scene contains both zipper merging across two lanes and forced merging across three lanes, requiring vehicles to change lanes.The upper lanes merge into one, while the lower lanes merge into two and impose lane changes.
- A. Diversity: The dataset includes an extremely busy 7-way roundabout with one yield branch and six stop branches, alongside roundabouts governed entirely by yield or stop signs.Vehicles enter simultaneously at relatively high speeds, producing intensive interactions.
- A. Diversity: It also contains busy all-way-stop intersections where multiple vehicles compete by inching forward, plus intersections combining stop-controlled branches with right-of-way traffic.The examples include a 9-lane all-way-stop intersection and a busy all-way-stop T-intersection.
- B. Internationality: Motion data spans three continents and four countries— the US, China, Germany, and Bulgaria—while all included countries drive on the right.The paper notes remarkable distinctions in driving culture despite this shared traffic-side convention.
- B. Internationality: Comparable roundabouts from the US, Germany, and China share yield-based nominal rules, while German and Chinese zipper-merge scenarios retain similar rules and heavy-traffic speeds.The roundabout examples lack stop signs, and the zipper scenarios remain comparable despite differing road contexts.
C. Complexity
The dataset emphasizes dense, highly interactive driving behavior, including cooperative and adversarial motions, negotiations, violations, near-collisions, and slight collisions across complex scenarios.
- C. Complexity: Strongly interactive vehicle pairs can appear every few seconds in selected ramps, entrances, and stop-controlled intersections.Examples include the ramp in ZS, entrance branches in FT, all-way-stop intersections in EP and MA, and a two-way-stop intersection in GL.
- C. Complexity: Unstructured roads without explicit lane restrictions enable irrational and dangerous insertions, while drivers negotiate near-simultaneous arrivals through inching or acceleration.In GL, a vehicle inserted between two stopped vehicles despite lacking explicit road structure; in MA and EP, drivers negotiated around ambiguous right-of-way.
- C. Complexity: Right-of-way violations create adversarial interactions, including a vehicle entering a roundabout and forcing the vehicle with right-of-way to stop and yield.The example occurs in FT, where V0 violated the rule while V1 was already in the roundabout.
- C. Complexity: Social pressure and delayed entry for vehicles without nominal right-of-way further increase behavioral complexity.Vehicles may wait minutes to enter and pass, with queues and honking contributing to impatience and aggressive behavior.
- C. Complexity: Critical interactions include extremely low time-to-collision-point situations, emergency swerves, and slight collisions.Examples include a near-collision in GL requiring an emergency swerve and a slight collision involving conflicting right-turn predictions.
E. Semantic Map
The dataset provides semantic lanelet2 maps alongside motion data, combining physical road geometry with reference paths, lane connectivity, traffic rules, and right-of-way information.
- E. Semantic Map: The semantic maps include physical layers, reference paths, lanelets and their connections, turn directions, traffic rules, and right-of-way information.The information is organized in a consistent format and toolkit for dataset users.
- E. Semantic Map: The construction pipeline combines drone video stabilization and map alignment with object detection, tracking, and trajectory smoothing.Raw 4K video at 30 Hz is downsampled to 10 Hz; Faster R-CNN, Kalman filtering, and an RTS smoother are used in processing.
- E. Semantic Map: Lanelet2 maps encode road borders, lane markings, traffic signs, lane courses, and regulatory elements such as right-of-way and speed limits.The physical layer is used to create atomic lanelets, which form the basis for regulatory elements.
B. Motions from Traffic Camera Data
Traffic-camera motion data are processed into entity trajectories and analyzed with semantic lanelet2 maps, enabling interpretation of movements in structured road environments.
- B. Motions from Traffic Camera Data: Traffic-camera processing detects vehicles and pedestrians using 2D bounding boxes, instance masks, and instance types.The pipeline begins with a state-of-the-art object detector applied to each frame.
- B. Motions from Traffic Camera Data: Detections are associated into tracks using mask-overlap tracking combined with a visual tracker that compensates for missed detections.The association combines an Intersection-over-Union tracker with a visual tracker.
- B. Motions from Traffic Camera Data: Ground-plane trajectories are estimated and smoothed with a pin-hole camera observation model, uncertainty handling, and a bicycle process model for vehicles.The bicycle model captures vehicle kinematic constraints while incorporating pixel-level measurement uncertainty.
- B. Motions from Traffic Camera Data: Lanelet2 maps provide centimeter-accurate representations of road geometry and traffic regulations for trajectory analysis.The maps encode road borders, lane markings, traffic signs, lanelets, and regulatory elements.
- B. Motions from Traffic Camera Data: Combining trajectories with maps helps explain deceleration near junctions based on right-of-way and interacting traffic participants.The maps support reasoning about why vehicles respond differently to the same junction layout.
- B. Motions from Traffic Camera Data: The dataset covers roundabouts, unsignalized and signalized intersections, merging, and lane-change scenarios.The paper reports vehicle trajectories across multiple locations and scenario categories.
B. Metrics for Interactive Behavior Identification
The dataset identifies interactive vehicle behavior using minimum time-to-conflict-point difference and waiting period metrics, revealing dense, highly critical interactions.
- IPV counts interaction pairs per vehicle using rule-based extraction across different spatial representations of vehicle paths.The metric is used to represent interactive-behavior density.
- ΔTTCPmin measures relative vehicle states when paths share a conflict point without a forced stop.It covers both static crossing or merging points and dynamic merging points derived from actual trajectories.
- TTCPmin ≤3 s defines an interaction, with traveling time computed from each vehicle’s speed and distance to the conflict point.The interaction period determines the start and end times used in the metric.
- C. Distribution of Interactivity: 13,375 interactive vehicle pairs were identified, and INTERACTION contains more intensive interactions with ΔTTCPmin ≤1 s than highD and NGSIM.The comparison plots vehicle counts and normalized density against ΔTTCPmin.
- C. Distribution of Interactivity: Across scenarios, the dataset has high vehicle density at ΔTTCPmin ≤1 s and waiting periods greater than 3 s.These distributions summarize interactive trajectories over different driving scenarios.
VI. UTILIZATION EXAMPLES
The dataset is designed to support behavior-related research, and utilization examples demonstrate prediction with interactive trajectories and semantic maps.
- The dataset provides utilization examples spanning motion prediction, imitation learning, planning, clustering, interaction extraction, and human-like behavior generation.These examples illustrate the dataset’s intended research applications.
- High-density interactive trajectories and HD semantic maps support both learning-based and planning-based motion prediction approaches.The maps and trajectories are described as useful for prediction in intensive-interaction situations.
- Motion/trajectory prediction: A WAE-based prediction method outperformed VAE, auto-encoder, and GAN models on position and yaw-angle RMSE and MAE in the reported comparison.The experiment used motion data from the FT scenario.
- Motion/trajectory prediction: The utilization section reports prediction-accuracy comparisons from method [36].
- Motion/trajectory prediction: HD semantic maps enabled a CVAE and inverse-reinforcement-learning planner to predict both irrational and rational vehicle behavior.The method defined learning features in the Frenet frame and reported improved generalization.
B. Imitation Learning
The dataset supports imitation learning and decision-making or planning validation in interactive scenarios, including cases involving uncertain intentions and right-of-way violations.
- B. Imitation Learning: Imitation learning uses semantic HD maps and surrounding-vehicle states as features to imitate human driving in the FT roundabout scenario.The framework extends the fast integrated learning and control method from.
- C. Validation of Decision and Planning: Dataset replay motions are suitable for testing planners when surrounding entities are independent of the ego vehicle’s motions.Examples include the ego vehicle lacking right-of-way or others violating traffic rules.
- C. Validation of Decision and Planning: A planner tested in the FT roundabout decelerated to avoid a collision with a vehicle entering despite the autonomous vehicle having right-of-way.The result is shown as a bird’s-eye-view simulation screenshot.
- C. Validation of Decision and Planning: A DBN-based predictor returned P(exit)=0.626 when another roundabout vehicle’s intention was unclear, supporting uncertainty-aware decision and planning.The framework combined integrated decision-making, sample-based planning, and probabilistic intention prediction.
- C. Validation of Decision and Planning: Long-term planning for the yielding case guaranteed a full stop in the worst case, while the non-conservatively defensive strategy preserved planned acceleration when appropriate.The reported speed profiles correspond to screenshots of the FT roundabout situations.
D. Motion Clustering and Representation Learning
The dataset supports motion-pattern discovery and human-like behavior modeling by combining map-based trajectory features with interactive driving data.
- D. Motion Clustering and Representation Learning: X-means clustered trajectories using vehicle-motion features represented in the Frenet frame from map information.The resulting clusters are visualized with the map and interacting vehicles’ longitudinal positions and speeds.
- D. Motion Clustering and Representation Learning: Different interactive motions were separated while similar motions were clustered, supporting extraction of motion patterns.The results are shown in trajectory, state-coordinate, and PCA feature-space views.
- Interaction extraction: The dataset was used to learn interaction relationships between pairs of agents through an interaction-frame extraction method.Examples include interacting car pairs in the FT scenario.
- F. Human-like Decision and Behavior Generation: A CPT-based human-behavior model outperformed a TTCP model and matched a neural-network model with less training data and better interpretability.All three models were learned and tested using FT roundabout data.
- VII. CONCLUSION: The dataset’s diverse, critical, and semantically mapped scenarios support representation learning, interaction extraction, and human-like behavior generation.The conclusion also lists prediction, imitation, decision-making, and planning as supported research areas.