Source-linked AI summary

The Multi Vehicle Stereo Event Camera Dataset: An Event Camera Dataset for 3D Perception

Alex Zihao Zhu, Dinesh Thakur, Tolga Ozaslan, Bernd Pfrommer, Vijay Kumar, Kostas Daniilidis

arXiv:1801.10202v2cs.RO

TL;DR

Event-based 3D perception lacks the broad labeled data available for traditional cameras, limiting testing and development. The paper addresses this gap with a diverse synchronized stereo event-camera dataset, calibrated multimodal measurements, and accurate pose and depth references for evaluating new methods.

  • Problem

    Event-based 3D-perception research lacks the wealth of labeled data available for traditional cameras for testing and development.

  • Method

    The paper constructs a synchronized stereo event-camera dataset across handheld, aerial, driving, and motorcycle settings, combining events, grayscale images, IMU, lidar, motion capture, and GPS.

  • Results

    The dataset provides stereo event streams with synchronized grayscale and IMU data, plus ground-truth pose and depth images for the cameras.

  • Takeaways & Limitations

    The dataset is intended to provide a common standard for evaluating and comparing event-based methods across diverse 3D-perception settings.

Abstract

from arXiv · show

Event based cameras are a new passive sensing modality with a number of benefits over traditional cameras, including extremely low latency, asynchronous data acquisition, high dynamic range and very low power consumption. There has been a lot of recent interest and development in applying algorithms to use the events to perform a variety of 3D perception tasks, such as feature tracking, visual odometry, and stereo depth estimation. However, there currently lacks the wealth of labeled data that exists for traditional cameras to be used for both testing and development. In this paper, we present a large dataset with a synchronized stereo pair event based camera system, carried on a handheld rig, flown by a hexacopter, driven on top of a car and mounted on a motorcycle, in a variety of different illumination levels and environments. From each camera, we provide the event stream, grayscale images and IMU readings. In addition, we utilize a combination of IMU, a rigidly mounted lidar system, indoor and outdoor motion capture and GPS to provide accurate pose and depth images for each camera at up to 100Hz. For comparison, we also provide synchronized grayscale images and IMU readings from a frame based stereo camera system.

I. INTRODUCTION

Event cameras offer low-latency, asynchronous sensing and high dynamic range, but existing robotics algorithms and labeled datasets are poorly suited to their measurements. This paper introduces a synchronized stereo event-camera dataset with diverse trajectories, calibrated sensors, and accurate pose and depth ground truth.

  • Motivation: Event cameras register log-intensity changes asynchronously, enabling tens-of-microseconds latency and over 130 dB dynamic range.Traditional cameras typically have tens-of-milliseconds latency and about 60 dB dynamic range.
  • Motivation: Most modern robotics algorithms assume synchronous measurements, while event data lacks intensity information on its own.These differences require new algorithms rather than direct reuse of traditional-camera methods.
  • Impact: The dataset is intended to support realistic evaluation, machine-learning training, and development of event-based methods.The authors make the full dataset available online.
  • Contributions: The dataset provides synchronized stereo event cameras with accurate ground-truth depth and pose.It is presented as the first dataset combining these properties for stereo event cameras.
  • Contributions: Sequences span handheld, hexacopter, car, and motorcycle platforms across varied speeds, illumination levels, and environments.The rig combines event cameras with calibrated lidar, IMUs, and frame-based images.

II. RELATED WORK

Earlier event-camera datasets provide useful multimodal measurements, but are predominantly monocular and often restrict accurate pose or depth ground truth to small indoor settings. This work instead offers stereo sequences with ground-truth pose and depth across indoor and outdoor environments.

  • Existing datasets: Existing datasets include monocular event cameras paired with RGB-D, motion capture, lidar, wheel encoders, pan-tilt units, or vehicle measurements.Their modalities and ground-truth sources vary substantially across datasets.
  • Existing datasets: Prior datasets provide indoor pose or depth ground truth, but coverage is often limited to small environments or short-displacement sequences.Some datasets lack outdoor sequences or significant displacement with ground-truth information.
  • Gap: Most earlier datasets provide only monocular sequences, with 6DoF pose ground truth limited to small indoor environments and few depth-ground-truth sequences.The paper identifies this as the central gap addressed by its dataset.
  • This work: This work provides stereo sequences with ground-truth pose and depth images in varied indoor and outdoor settings.The broader coverage is intended to support more meaningful comparisons between event-based methods.

B. Event Based 3D Perception

Event-based 3D-perception research spans stereo depth, feature tracking, visual odometry, and SLAM, but evaluations often rely on small, paper-specific datasets. The dataset packages synchronized multimodal measurements and reference pose and depth to improve comparability.

  • Stereo depth: Stereo depth methods use spatial, temporal, cooperative, epipolar, ordering, polarity, and orientation-based matching constraints.Several approaches specifically adapt cooperative stereo methods to asynchronous point-based event measurements.
  • Odometry and SLAM: Event-based odometry and SLAM work includes feature tracking, visual and visual-inertial odometry, angular-velocity estimation, and up-to-scale mapping.Some systems fuse events with depth sensors or other measurements.
  • Evaluation gap: Most methods are evaluated on small datasets generated for individual papers, making performance comparisons difficult.The paper seeks more extensive ground truth for more meaningful evaluation.
  • Dataset contents: Each sequence includes left and right DAVIS events, APS grayscale images, IMU measurements, VI Sensor data, and Velodyne point clouds.The measurements are provided in ROS bag format.
  • Ground truth: The dataset supplies reference poses for the left DAVIS camera and reference depth images for both left and right cameras.These references support stereo depth and broader 3D-perception evaluation.

A. Sensors

The dataset uses a calibrated, synchronized stereo event-camera rig augmented with lidar, GPS, and a synchronized frame-based stereo system. These sensors provide event, grayscale, IMU, and reference measurements across the platform configurations.

  • Sensor package: The rig includes left and right DAVIS cameras, APS grayscale imaging, IMU measurements, and calibrated extrinsics between sensors.The sensor axes and inter-sensor transformations are documented through calibration.
  • Event stereo cameras: The two mDAVIS-346B cameras form a horizontally mounted, timestamp-synchronized stereo pair with a 10cm baseline and 346x260-pixel resolution.The cameras provide up to 50fps APS output and approximately 87 degrees of horizontal field of view.
  • Reference sensors: A Velodyne Puck LITE lidar is rigidly mounted above the stereo cameras to provide overlapping, dense depth reference measurements.The lidar’s vertical field of view fully overlaps the stereo DAVIS rig.
  • Reference sensors: Outdoor configurations include a GPS device for an additional latitude-and-longitude reference.GPS is mounted away from the rig to reduce interference from USB 3.0 cables.
  • Comparison camera: A synchronized stereo VI Sensor with IMU is included for comparison with frame-based methods, although it is mounted upside down and its transform is provided.The VI Sensor is used alongside the DAVIS cameras in the sensor package.

B. Sequences

The dataset spans handheld, aerial, driving, and motorcycle-mounted sequences across indoor and outdoor environments and varied illumination. It combines motion-capture and lidar-SLAM references while recording sequence-level motion and event statistics.

  • Aerial sequences: The hexacopter sequences use separate indoor Vicon and outdoor Qualisys motion-capture arenas, each providing millimeter-accuracy poses at 100Hz.The indoor arena measures 26.8m × 6.7m × 4.6m, while the outdoor arena measures 30.5m × 15.3m × 15.3m.
  • Sequence environments: Sequences cover the full sensor rig in indoor and outdoor environments, including indoor scenes with and without external lighting, with pose and depth from lidar SLAM.The setup is carried through a loop to test high-dynamic-range scenarios.
  • Sequence statistics: Table II reports total time, distance, maximum linear and angular velocity, and mean event rate for sequences from each vehicle.The table also marks unavailable VI-Sensor data and failures of right DAVIS grayscale images.
  • Visual examples: Sample sequences show blue and red events overlaid on images from indoor and outdoor scenes captured during day and evening.The examples illustrate event overlays across different illumination conditions.

3) Outdoor Driving:

Outdoor driving sequences cover moderate-speed sedan runs and high-speed motorcycle runs, with complementary lidar, IMU, GPS, and ground-truth pose or depth products. Cartographer and LOAM are used for different reference-generation strengths.

  • Sedan sequences: Sedan sequences reach 12 m/s across West Philadelphia neighborhoods during day and evening, including direct sun, with lidar-map depth and lidar-odometry/GPS pose.These sequences include varied illumination and outdoor driving conditions.
  • Motorcycle sequences: Motorcycle sequences reach 38m/s and provide GPS longitude, latitude, and relative velocity.The DAVIS stereo rig and VI Sensor are mounted on the motorcycle handlebar.
  • Reference generation: Cartographer fuses lidar sweeps and IMU data into loop-closed 2D poses, while LOAM generates dense 3D local maps and camera depth images.The lidar-derived pose is transformed into the left DAVIS frame using calibration.
  • Reference generation: Cartographer is preferred for longer trajectories because loop closure reduces global drift, whereas LOAM produces better-aligned local maps.The 2D-pose assumption is considered valid because the driven roads mostly have a consistent grade.

A. Ground Truth Pose

Ground-truth pose is obtained from motion capture when available and otherwise from lidar- and IMU-based odometry, with GPS used outdoors for comparison. Lidar maps are also projected into calibrated cameras to generate depth images.

  • Pose sources: Motion-capture sequences provide body-frame poses at 100Hz with millimeter-level accuracy.The measured pose is defined for each time t in the motion-capture world frame.
  • Pose sources: Outdoor pose estimation uses Cartographer to fuse lidar sweeps and IMU data into a loop-closed 2D body pose with minimal drift.Raw GPS readings are also provided for outdoor scenes.
  • Pose validation: 4.7m is the overall average position error between Cartographer and GPS across outdoor driving sequences.The reported error is consistently around 5m for the sample Car Day 2 sequence, while a spike near 440 seconds is attributed to GPS error.
  • Pose transformation: The calibrated body-to-DAVIS transform estimates each left-DAVIS pose relative to the first left-DAVIS pose in a sequence.The formulation uses the body pose and the fixed extrinsic transform between the body and left DAVIS frames.
  • Depth generation: Lidar point clouds are transformed into local maps using LOAM poses, projected into each DAVIS camera, and reduced to the closest point per pixel for depth maps.Points outside image bounds are discarded, and rectified depth can also be converted into raw distorted depth images.

V. CALIBRATION

The dataset calibrates camera, IMU, lidar, and motion-capture relationships, and provides calibration results for reproducibility. Calibration is repeated when collection conditions or the sensing payload change.

  • Intrinsic, stereo, camera–IMU, camera–lidar, and motion-capture transformations are calibrated across the sensing system.The calibration results are provided in YAML form.
  • Calibration uses Kalibr, Camera and Range Calibration Toolbox, manual refinement, and CamOdoCal for different sensor relationships.
  • Ground-truth examples include full and local maps, depth images with events overlaid, and GPS–Cartographer trajectory comparisons.
  • Calibration is repeated each day and whenever the sensing payload is modified to compensate for rig changes.Raw calibration data is also available on demand for researchers performing their own calibration.

A. Camera Intrinsic, Extrinsic and Temporal Calibration

Camera calibration estimates intrinsics and inter-camera extrinsics from AprilTag grids, calibrates camera–IMU transformations, and synchronizes timestamps across sensors.

  • Camera Intrinsic, Extrinsic and Temporal Calibration: AprilTag-grid calibration estimates each camera’s focal length, principal point, distortion parameters, and inter-camera extrinsics.
  • Camera Intrinsic, Extrinsic and Temporal Calibration: Temporal offset is calibrated by maximizing cross-correlation between gyroscope angular-velocity magnitudes from the left DAVIS and VI Sensor IMUs.VI Sensor timestamps are modified to compensate for the estimated offset.
  • Camera Intrinsic, Extrinsic and Temporal Calibration: Camera–IMU transformations are estimated from calibration sequences with the rig moving in front of the AprilTag grid.The procedures separate calibration stages to optimize each individual calibration.

C. Motion Capture to Camera Extrinsic Calibration

Motion-capture poses are transformed into camera poses through hand-eye calibration, while lidar–camera alignment is refined using projected depth and event images.

  • Motion Capture to Camera Extrinsic Calibration: Motion-capture model poses require an additional calibration because the model frame is not aligned with a camera frame.
  • Motion Capture to Camera Extrinsic Calibration: Static poses of the left DAVIS and mocap model are used to solve for the transform from the camera frame into the model frame.
  • Motion Capture to Camera Extrinsic Calibration: CamOdoCal solves the hand-eye calibration with a linear method followed by nonlinear refinement.
  • Motion Capture to Camera Extrinsic Calibration: Lidar–camera calibration initially produced up to five pixels of projected-depth error and was manually refined for rotation and time offset.Translation was fixed using CAD-model values, and lidar timestamps were compensated for the offset.

VI. KNOWN ISSUES

Known issues include depth errors on moving objects, possible timing offsets, and polarity imbalance in indoor flying sequences. The dataset is intended as a standard resource for evaluating event-based methods.

  • Known Issues: Depth-map generation assumes static scenes and may produce errors of up to two meters on moving objects such as other cars.Moving-object points are typically rare, and future filtering could classify and omit them.
  • Known Issues: Motion-capture and GPS timestamps rely on host-computer synchronization, while lidar timestamps may lag measurements because of its spin rate.
  • Known Issues: Indoor flying sequences show a positive-to-negative event ratio approximately 2.5–5x higher than usual, especially over speckled floors.The cause is unknown, so researchers using event polarities are advised to account for the imbalance.
  • Known Issues: The dataset provides stereo event-camera sequences across vehicles and environments with ground-truth 6DoF pose and depth images.The authors hope it will support standardized evaluation and comparison of new event-based methods.
Loading 1801.10202v2…