Source-linked AI summary

PandaSet: Advanced Sensor Suite Dataset for Autonomous Driving

Pengchuan Xiao, Zhenlei Shao, Steven Hao, Zishuo Zhang, Xiaolin Chai, Judy Jiao, Zesong Li, Jian Wu, Kai Sun, Kun Jiang, Yunlong Wang, Diange Yang

arXiv:2112.12610v1cs.CVcs.RO

TL;DR

Autonomous-driving 3D perception requires high-precision, richly annotated real-world data, while testing disruptions reduced available road-test data. PandaSet addresses this gap with an openly licensed multimodal dataset, detailed annotations, challenging environments, and baselines across three perception tasks.

  • Problem

    3D perception depends on high-quality, precise real-world annotations, but autonomous-driving datasets face limited scene diversity and reduced road-test data availability.

  • Method

    PandaSet combines mechanical spinning and forward-facing LiDARs with cameras, multi-sensor-fusion annotations, diverse driving scenes, and a development kit.

  • Results

    28 annotation classes, 37 semantic-segmentation labels, and baselines for LiDAR-only detection, LiDAR-camera fusion detection, and point-cloud segmentation are provided.

  • Takeaways & Limitations

    PandaSet offers an openly accessible resource intended to help researchers and developers accelerate safe autonomous-vehicle deployment, including long-range and full-360° annotations.

Abstract

from arXiv · show

The accelerating development of autonomous driving technology has placed greater demands on obtaining large amounts of high-quality data. Representative, labeled, real world data serves as the fuel for training deep learning networks, critical for improving self-driving perception algorithms. In this paper, we introduce PandaSet, the first dataset produced by a complete, high-precision autonomous vehicle sensor kit with a no-cost commercial license. The dataset was collected using one 360° mechanical spinning LiDAR, one forward-facing, long-range LiDAR, and 6 cameras. The dataset contains more than 100 scenes, each of which is 8 seconds long, and provides 28 types of labels for object classification and 37 types of labels for semantic segmentation. We provide baselines for LiDAR-only 3D object detection, LiDAR-camera fusion 3D object detection and LiDAR point cloud segmentation. For more details about PandaSet and the development kit, see https://scale.com/open-datasets/pandaset.

I. INTRODUCTION

PandaSet addresses the need for precise, diverse real-world data for autonomous-driving 3D perception by offering an open multimodal dataset with broad sensing, annotation, and baseline support.

  • Motivation: High-quality annotated data is needed because imprecise sensor measurements can limit the performance of back-end perception algorithms.
  • Motivation: Scene diversity is important because varying lighting, traffic, road conditions, vegetation, human behavior, and unfamiliar objects challenge real-world driving.The paper presents PandaSet partly to address reduced road-test data during COVID-19.
  • Dataset contribution: PandaSet provides a complete high-precision sensor kit covering a 360° field of view, combining mechanical spinning and forward-facing LiDARs.It is described as the first open-source dataset with both LiDAR types and free research and commercial use.
  • Dataset contribution: 28 annotation classes and 37 semantic-segmentation labels are provided, with annotations created through multi-sensor fusion for accuracy and precision.
  • Dataset contribution: PandaSet covers metropolitan scenes including traffic, pedestrians, construction, hills, and lighting conditions from daytime through nighttime.The scenes target challenging conditions relevant to full level 4 and 5 driving autonomy.
  • Dataset contribution: The paper provides baselines for LiDAR-only detection, LiDAR-camera fusion detection, and LiDAR point-cloud segmentation, along with a development kit.

II. RELATED WORK

The paper situates PandaSet among open-source autonomous-driving datasets that combine camera and LiDAR data with 3D annotations, extending an established benchmark lineage.

  • Context: Recent machine-learning approaches have driven progress in 3D perception and autonomous-driving applications.
  • Dataset comparison: Table I is presented as a comparison of current autonomous-driving datasets.
  • Dataset comparison: The related-work comparison considers open-source datasets containing both camera-collected and LiDAR-collected data together with 3D annotations.
  • Dataset comparison: KITTI, launched in 2012, is identified as a pioneering benchmark collected for an autonomous-driving platform.Its listed sensors include two stereo camera systems, a mechanical spinning LiDAR, and GNSS/IMU.

III. PANDASET DATASET

PandaSet collects synchronized multimodal data with a vehicle-mounted sensor suite and presents representative views of camera annotations and fused LiDAR point clouds.

  • A. Data Collection: The collection vehicle carries six cameras, two LiDARs, and one GNSS/IMU device; five cameras cover 360° and the LiDARs provide 200 m and 300 m ranges.The forward-facing LiDAR is intended to support longer-range 3D object detection for high-speed driving scenarios.
  • A. Data Collection: Figure 2 shows projected 3D bounding boxes and semantic-segmentation annotations across six camera images, alongside a combined annotated point cloud in 3D view.Each visualization grid is 20 m long.
  • A. Data Collection: Table II provides detailed sensor specifications for the collection suite.

SENSOR SPECIFICATIONS

PandaSet organizes synchronized camera and LiDAR measurements into frame-based records, using trigger-based exposure timing and GPS/PTP time synchronization.

  • Each frame encapsulates image data or a LiDAR point-cloud scan cycle.The mechanical spinning and forward-facing LiDARs both scan at 10Hz.
  • A trigger board exposes each camera when the mechanical LiDAR scans across its field-of-view center, aligning captured objects in time.
  • Image timestamps combine exposure-trigger time with exposure duration, while point-cloud timestamps mark LiDAR scan completion.Point-level timestamps are also provided for each point-cloud frame.
  • PTP synchronizes the two LiDARs, and GPS provides the time source for the entire sensor suite.

B. Sensor Calibration

PandaSet combines calibrated multimodal sensing with diverse scenes and high-quality annotations, including synchronized imagery, LiDAR data, and broad object-label coverage.

  • PandaSet’s sample data overlays camera images with point clouds from both mechanical spinning and forward-facing LiDARs.Sequential LiDAR scans are also aligned in the sample visualization.
  • The 103 eight-second scenes include urban traffic, construction, uncommon objects, varied terrain, and lighting from day through night.
  • All 103 scenes have 3D bounding-box annotations, while 76 scenes also have semantic point-cloud annotations for both LiDARs.Annotations are produced at 10Hz.
  • The dataset’s sample images represent varied lighting conditions and road environments.

E. Dataset Statistics

PandaSet emphasizes long-range, dense, and diverse annotation coverage, extending object labeling to 300 meters and including full 360° views and rare classes.

  • 300 meters is the maximum stated annotation distance, enabled by the long-range mechanical spinning and forward-facing LiDARs.The paper describes this range as significantly farther than most other datasets.
  • Figure 5 reports instances per frame for seven traffic-participant classes, with PandaSet-Front restricted to the front-facing camera field of view for KITTI consistency.
  • Figure 7 summarizes total object instances by class and total semantic LiDAR points by segmentation class.
  • PandaSet provides object annotations in full 360° view rather than only frontal view, supporting higher per-frame label density than KITTI.
  • The dataset includes rare classes such as motorized scooters, rolling containers, animals, smoke, and car exhaust.These categories are presented as resources for addressing long-tail real-world driving scenarios.

IV. BASELINE EXPERIMENTS

PandaSet baselines cover LiDAR-only detection, LiDAR-camera fusion detection, and LiDAR point-cloud segmentation using fixed sequence-based training and test splits.

  • The baseline experiments evaluate LiDAR-only 3D object detection, LiDAR-camera fusion 3D object detection, and LiDAR point-cloud segmentation.
  • The first 50 frames of each sequence are used for training and the remaining 30 frames for testing.This split is selected as a tradeoff between robustness and evaluation uniformity.
  • The experiments draw on 103 detection sequences and 76 segmentation sequences.

A. LiDAR-only 3D object detection

PandaSet’s LiDAR-only detection baseline retrains PV-RCNN separately for its two LiDAR configurations and evaluates performance with standard AP and class-specific IoU thresholds.

  • Model setup: PV-RCNN was retrained for cars, pedestrians, and cyclists to align with the KITTI 3D object-detection benchmark.Separate models account for the differing coverage of the mechanical spinning and forward-facing LiDARs.
  • Evaluation: 0.7 IoU is used for cars, while 0.5 IoU is used for pedestrians and cyclists in the 3D test.Test results are evaluated using average precision with 11 recall positions.
  • Evaluation: Examples with ≤5 LiDAR points are designated as LEVEL 2, the more challenging detection category.LEVEL 1 and LEVEL 2 follow Waymo Open dataset difficulty definitions for single-frame 3D detection.

B. LiDAR-camera fusion 3D object detection

The LiDAR-camera fusion baseline combines image-derived semantic information with LiDAR points for 3D object detection, using a forward-facing camera field of view and KITTI-aligned inputs.

  • Model setup: PointPainting is re-implemented with DeepLabv3+ image segmentation and PointRCNN 3D detection.DeepLabv3+ supplies per-pixel class scores that are fused with point-cloud data.
  • Data and evaluation: Only object examples visible to the forward-facing, long-focus camera are used for training and inference.The LiDAR input is a 50m-range clipped point cloud from the mechanical spinning LiDAR.
  • Data and evaluation: The fusion baseline uses the same AP benchmark and IoU thresholds described for LiDAR-only 3D object detection.

C. LiDAR point cloud segmentation

PandaSet’s point-cloud segmentation baseline uses RangeNet53 and maps its original 37 output classes into 14 primary autonomous-driving classes.

  • Baseline: RangeNet53 provides the segmentation baseline without postprocessing.The publicly released implementation is used with the network unchanged.
  • Label mapping: 37 original output classes are merged into 14 primary classes for autonomous driving.At inference, each point receives the class with the highest score.

V. CONCLUSION

PandaSet contributes a freely licensed multimodal dataset combining mechanical spinning and forward-facing LiDARs, with detailed collection and annotation information for autonomous-driving research.

  • Contribution: PandaSet is presented as the first open-source dataset combining mechanical spinning and forward-facing LiDARs with free research and commercial use.
  • Impact: The release aims to help researchers and developers accelerate the safe deployment of autonomous vehicles.
  • Future work: Future plans include evaluation metrics, a public leaderboard, and map information for the dataset.
Loading 2112.12610v1…