Source-linked AI summary

Ultimate SLAM? Combining Events, Images, and IMU for Robust Visual SLAM in HDR and High Speed Scenarios

Antoni Rosinol Vidal, Henri Rebecq, Timo Horstschaefer, Davide Scaramuzza

arXiv:1709.06310v4cs.CVcs.RO

TL;DR

Traditional visual-inertial systems struggle with motion blur, limited dynamic range, and low light, while event-only methods provide little information during limited motion. This paper tightly fuses events, standard frames, and inertial measurements, improving average position accuracy by 85% over frames-plus-IMU and 130% over events-plus-IMU pipelines.

  • Problem

    Traditional visual-inertial pipelines remain limited by camera motion blur, low-light failure, and restricted dynamic range in difficult sensing conditions.

  • Method

    The paper introduces a tightly coupled state-estimation pipeline fusing event measurements, standard intensity frames, and inertial measurements for real-time mobile-robot operation.

  • Results

    85% average position-accuracy improvement over frames-plus-IMU and 130% over events-plus-IMU were achieved on the Event Camera Dataset.

  • Takeaways & Limitations

    The hybrid pipeline supported the first reported closed-loop quadrotor flight using an event camera and robust tracking across low-light, fast-motion, and limited-motion conditions.

Abstract

from arXiv · show

Event cameras are bio-inspired vision sensors that output pixel-level brightness changes instead of standard intensity frames. These cameras do not suffer from motion blur and have a very high dynamic range, which enables them to provide reliable visual information during high speed motions or in scenes characterized by high dynamic range. However, event cameras output only little information when the amount of motion is limited, such as in the case of almost still motion. Conversely, standard cameras provide instant and rich information about the environment most of the time (in low-speed and good lighting scenarios), but they fail severely in case of fast motions, or difficult lighting such as high dynamic range or low light scenes. In this paper, we present the first state estimation pipeline that leverages the complementary advantages of these two sensors by fusing in a tightly-coupled manner events, standard frames, and inertial measurements. We show on the publicly available Event Camera Dataset that our hybrid pipeline leads to an accuracy improvement of 130% over event-only pipelines, and 85% over standard-frames-only visual-inertial systems, while still being computationally tractable. Furthermore, we use our pipeline to demonstrate - to the best of our knowledge - the first autonomous quadrotor flight using an event camera for state estimation, unlocking flight scenarios that were not reachable with traditional visual-inertial odometry, such as low-light environments and high-dynamic range scenes.

I. INTRODUCTION

Traditional visual-inertial systems struggle with high-speed motion and high-dynamic-range scenarios, while event and standard cameras offer complementary strengths and limitations. The paper proposes the first pipeline fusing events, standard frames, and IMU measurements, enabling onboard quadrotor flight in difficult scenarios.

  • Motivation: Visual-inertial odometry pipelines struggle with high-speed motions and high-dynamic-range scenarios because traditional cameras suffer from motion blur and low dynamic range.These limitations impede ego-motion estimation for applications including augmented or virtual reality and autonomous robot control.
  • Event-camera advantages: Event cameras transmit asynchronous per-pixel brightness changes, offering microsecond-scale latency, 140 dB dynamic range, and immunity to motion blur.Each event carries the brightness-change location, time, and sign; standard cameras provide intensity frames at a fixed framerate.
  • Event-camera limitations: Event-only state estimation must reconstruct environmental intensity information by combining many events into semi-dense or dense representations.The cited approaches reconstruct either a semi-dense depth map or a dense depth map with intensity values.
  • Sensor complementarity: Standard cameras directly provide intensity values but fail in low light, blur during fast motion, and offer only 60 dB dynamic range with over- or under-exposed areas.Their synchronous exposure across the sensor contributes to motion blur during fast motions.
  • Contribution and novelty: The paper proposes the first state-estimation pipeline fusing events, standard frames, and an IMU, and the first quadrotor system using this combination for difficult scenarios.The system uses only onboard sensing and computing; combining all three modalities and event-camera quadrotor flight were identified as open or underexplored problems.

Contributions · II. RELATED WORK

The paper contributes a real-time hybrid state-estimation pipeline that tightly couples events, standard frames, and inertial measurements, evaluates its accuracy and tractability, and demonstrates reliable autonomous quadrotor flight in challenging conditions. Related work spans standard-camera VIO architectures, event-camera state estimation, complementary sensing, event-IMU fusion, and event-camera quadrotor control.

  • Contributions: The authors introduce a pipeline that fuses events, standard frames, and inertial measurements for robust and accurate state estimation.It extends with standard frames and adds improvements aimed at real-time use on mobile robots.
  • Contributions: The proposed approach improves state-estimation accuracy while keeping computational load tractable.The paper evaluates this tradeoff quantitatively.
  • Contributions: The method enables onboard state estimation for an autonomous quadrotor that flies reliably during low-light scenes and fast motions.These experiments highlight event cameras’ potential for robust state estimation.
  • II. RELATED WORK: Standard-camera VIO is organized into filtering, fixed-lag smoothing, and full smoothing according to how many camera poses are estimated.Filtering estimates only the latest state, fixed-lag smoothing retains a recent window while marginalizing older states, and full smoothing estimates the entire state history.
  • II. RELATED WORK: Event-camera state estimation progressed from restricted tasks such as rotational motion and planar SLAM to 6-DOF pose estimation using only events.This development followed the introduction of commercial event cameras in 2008.
  • II. RELATED WORK: Prior event-frame pipelines using complementary sensors omit inertial measurements and rely on sharp, correctly exposed intensity frames as event-alignment templates.Consequently, they work only when the standard frames are of good quality.
  • II. RELATED WORK: Prior event-IMU methods include continuous-time fusion, feature tracking with iterative Expectation-Maximization, and a real-time visual-inertial odometry pipeline based on motion-compensated event frames.The continuous-time approach is not suited for real-time usage because updating spline parameters for every event requires expensive optimization.
  • II. RELATED WORK: Earlier event-camera quadrotor control tracked 6-DOF motion during a high-speed flip, but only in an artificial scene containing a known black square on a white wall.The cited work illustrates an early application of event cameras to robot control.

III. HYBRID STATE ESTIMATION PIPELINE

The proposed pipeline extends an event-camera-and-IMU state estimator by adding a standard camera that supplies fixed-framerate intensity frames. The paper focuses on the approach’s differences from and evaluates the improved pipeline on the Event Camera Dataset.

  • Pipeline design: The pipeline builds on while extending its event-camera-and-IMU setup with a standard camera.The added modality provides intensity frames at a fixed framerate.
  • Pipeline design: The method description emphasizes differences from needed to incorporate standard frames.
  • Evaluation: The improved pipeline is evaluated on the Event Camera Dataset.

A. Overview … 3) Synthesis of Motion-Compensated Event Frames:

The pipeline combines motion-compensated virtual event frames, standard-frame tracks, and IMU measurements in a shared optimization framework. Event windows are synchronized to standard-frame timestamps, adapt to event rate, and are collapsed into synthetic frames using inertial motion compensation and scene-depth estimates.

  • A. Overview: The approach synthesizes virtual event frames from spatiotemporal event windows, detects and tracks features with FAST and Lucas-Kanade, and triangulates reliable 3D landmarks.The resulting camera trajectory and landmark positions are estimated from these tracks.
  • A. Overview: The method maintains feature tracks from virtual event frames and standard frames in parallel, then jointly optimizes camera poses using events, images, and IMU data.This extends event-frame tracking by feeding heterogeneous feature tracks into one optimization module.
  • 1) Coordinate Frame Notation:: The sensor is a rigidly mounted hybrid of an event camera, a standard camera, and an IMU represented relative to an inertial world frame.Coordinate transformations use homogeneous matrices with rotational components in SO(3).
  • 2) Spatio-temporal Windows of Events:: Each event window is created when a standard frame arrives at timestamp tk, synchronizing the event window to the standard-camera frame.The window contains a fixed number of events selected around the frame timestamp.
  • 2) Spatio-temporal Windows of Events:: The duration of each fixed-count event window varies inversely with the event rate, so its temporal extent automatically adapts to scene activity.Figure 2 illustrates this adaptation with N = 4 events per window.
  • 3) Synthesis of Motion-Compensated Event Frames:: Every event window is collapsed into a synthetic event frame by drawing events after correcting each event’s position for its individual timestamp motion.The corrected positions are transferred to a reference event-camera frame.
  • 3) Synthesis of Motion-Compensated Event Frames:: Motion compensation uses inertially integrated incremental camera transformations, calibrated event-camera projection, and scene depth estimated by interpolating reprojected landmarks.The depth estimate is obtained through 2D linear interpolation on the image plane.
  • 3) Synthesis of Motion-Compensated Event Frames:: N is adjusted according to scene texture; the quadrotor experiments used N = 20 000 events per frame.The event count per frame is therefore an implementation parameter rather than a fixed universal setting.

4) Feature Tracking:

The pipeline detects FAST features in virtual event and standard frames, tracking them independently with KLT. Reliably triangulated candidates become persistent 3D landmarks for continued tracking, with parameters shared across sensors.

  • Feature detection and tracking: FAST features are extracted from virtual event frames and standard camera frames, then tracked independently across each sensor’s frames using KLT.This produces two independent sets of feature tracks.
  • Landmark initialization: Candidate features are tracked over multiple frames until reliable triangulation converts them into persistent 3D landmarks.Landmarks are triangulated linearly and tracked through subsequent frames.
  • Implementation details: The same detection and tracking parameters are used for motion-compensated event frames and standard frames.The FAST threshold is 50, with pyramidal KLT using 2 pyramid levels and 24 × 24-pixel patches.
  • Implementation details: A 32 × 32-pixel bucketing grid keeps features evenly distributed across each sensor’s image plane.Bucketing is applied independently to each sensor’s image plane.

5) Visual-inertial Fusion through Nonlinear Optimization: · 6) Additional Implementation Details:

The pipeline jointly optimizes event-camera, standard-camera, and inertial errors over a bounded keyframe and sliding-window state set. Initialization and a zero-velocity prior address startup calibration and almost-still motion without explicit sensor switching.

  • 5) Visual-inertial Fusion through Nonlinear Optimization:: The cost function jointly contains weighted reprojection errors from event and standard cameras plus an inertial error term.The formulation combines both visual modalities and inertial measurements in one optimization problem.
  • 5) Visual-inertial Fusion through Nonlinear Optimization:: Landmark measurements use information-matrix weighting, while inertial residuals represent differences between predicted and actual states.The IMU prediction uses standard kinematics and bias models, with a multiplicative minimal orientation error.
  • 5) Visual-inertial Fusion through Nonlinear Optimization:: Optimization uses M keyframes and a sliding window of the last K frames rather than all observed frames.Between frames, sensor-state prediction is propagated with IMU measurements, and Google Ceres performs the optimization.
  • 5) Visual-inertial Fusion through Nonlinear Optimization:: The formulation avoids explicit switching between standard and event cameras by naturally using the best available sensing modalities.This is achieved through the joint optimization formulation.
  • 6) Additional Implementation Details:: Initialization assumes the sensor is static for one or two seconds to estimate pitch, roll, and gyroscope and accelerometer biases.The pipeline collects inertial measurements during this static phase.
  • 6) Additional Implementation Details:: When event rates fall below a threshold, a strong zero-velocity prior is added to force the sensor to remain still.Experiments used a threshold on the order of 10^3 events/s with event-rate windows of 20 ms.
  • 6) Additional Implementation Details:: The implementation details report accuracy evaluations for frames, events, and IMU against event-and-IMU and frame-and-IMU baselines.Additional evaluation compares the proposed combination against [13], which uses events and IMU.

B. Evaluation

The proposed pipeline is evaluated on the Event Camera Dataset using aligned trajectory errors and comparisons across sensor combinations and prior work. It reports that jointly using frames, events, and IMU improves accuracy across almost all datasets, while distinguishing its event-plus-IMU setup from the state-of-the-art implementation.

  • The evaluation uses Event Camera Dataset sequences containing ground-truth tracking, extremely fast motions, and very high dynamic range, while excluding rotational-only and non-inertial datasets.
  • Estimated and ground-truth trajectories are aligned with a 6-DOF SE3 transformation over the interval from second 3 to second 8.Mean position and yaw errors are computed as percentages of total traveled distance; pitch and roll are omitted because gravity makes their errors comparable across pipelines.
  • The proposed pipeline is evaluated with three combinations: standard frames, events, and IMU; events plus IMU; and frames plus IMU.These combinations quantify the accuracy gained by jointly using events and frames with IMU compared with either modality alone alongside IMU.
  • The events-plus-IMU pipeline in Table I differs from the state-of-the-art pipeline in Table II because their event-frame rates and shared parameter values differ.These implementation differences explain the different results reported for Table I (E + I) and Table II (state-of-the-art E + I).
  • Using frames and events is more accurate on almost all datasets, and the work reports the first Event Camera Dataset results using all three sensor modalities.

IV. QUADROTOR FLIGHT WITH AN EVENT CAMERA … B. Flight Experiments

The paper demonstrates autonomous quadrotor flight using a hybrid frame-and-event pipeline in challenging illumination and motion conditions. The platform combines custom hardware, onboard computing, exposure control, and cascaded flight controllers for three experiments.

  • IV. QUADROTOR FLIGHT WITH AN EVENT CAMERA: The authors built an autonomous quadrotor to evaluate their hybrid frame-and-event pipeline in real-world challenging conditions.The paper separates the platform description from the specific in-flight experiments.
  • B. Flight Experiments: The flight platform and experiments are presented as evidence that the system can fly a quadrotor autonomously in challenging conditions.The experiments are intended to demonstrate the practical potential of the hybrid frame-and-event pipeline.
  • A. Aerial Platform: The quadrotor uses a DJI frame, RCTimer motors, AR drone propellers, and a PX4FMU autopilot assembled from off-the-shelf and custom 3D-printed components.These components constitute the aircraft platform shown in Fig. 4(a).
  • 1) Platform:: Onboard computation runs on an Odroid XU4 with a 2.0 GHz quad-core processor using Ubuntu 14.04 and ROS.A DAVIS 240C sensor with a 70° field-of-view lens is mounted on the quadrotor’s front.
  • 1) Platform:: The platform includes an open-source proportional auto-exposure algorithm that controls mean image intensity toward 70 in the experiments.The algorithm is designed to regulate the mean image intensity to a desired value.
  • 1) Platform:: Cascaded controllers stabilize the quadrotor and follow reference trajectories through high-level position and attitude control and low-level body-rate control.The high-level controller runs on the Odroid and sends desired body rates to the PX4 low-level controller.
  • B. Flight Experiments: The flight experiments test autonomous operation while switching indoor lights on and off, performing fast circles in a low-lit room, and hovering.These conditions target abrupt illumination changes, very low light, and fast motion.

2) Control: · 1) Switching the light off and on, in flight: · 2) Fast Circles in a Low-lit Room:

The autonomous quadrotor experiments show that the hybrid pipeline maintains robust state estimation when standard frames become unusable because of darkness or motion blur. Events support reliable tracking during fast, low-light circular flight, while trajectory noise reflects spatial differences in illumination.

  • 2) Control:: The pipeline provides robust state estimation when standard frames are completely black or severely motion-blurred, because the events remain usable.These conditions arise when the light is off or the quadrotor moves rapidly; the events are left unaffected.
  • 1) Switching the light off and on, in flight:: With the room light switched off during autonomous circular flight, standard frames become completely black and useless for state estimation.Residual window light is sufficient for the event camera, although the resulting events are noisier.
  • 1) Switching the light off and on, in flight:: The light-switching experiment demonstrates that events still carry enough information to support reasonable flight-state estimation in near darkness.The quadrotor autonomously flies circles while the only remaining illumination comes from the windows.
  • 2) Fast Circles in a Low-lit Room:: The fast-circle experiment uses a 1.2 m radius at a top linear velocity of 1.68 m/s in a 1.0 m-high, low-lit room.The commanded angular velocity is set to 1.4 rad/s, and the resulting image-plane optical flow is approximately 340 pixels/s.
  • 2) Fast Circles in a Low-lit Room:: At speeds below 1.2 m/s, standard frames remain sufficiently sharp for the pipeline to track features in both standard and event frames.The experiment begins with moderate speed before increasing along the circular trajectory.
  • 2) Fast Circles in a Low-lit Room:: As speed increases, severe motion blur reduces tracked standard-frame features, whereas motion-free virtual event frames preserve reliable feature tracks.The comparison is illustrated by the top and bottom rows of the corresponding feature-track examples.
  • 2) Fast Circles in a Low-lit Room:: The estimated fast-circle trajectory follows the desired trajectory, but its right side is slightly noisier than its left side.This matches the room’s illumination, with the left side more illuminated than the right, consistent with the quantitative experiments in section III-B.

3) Hovering: · V. CONCLUSIONS

The hybrid events–images–IMU pipeline remains stable during hovering, where event-only estimation drifts, and combines robust state estimation with demonstrated onboard quadrotor flight. On the Event Camera Dataset, it improves accuracy over both events-plus-IMU and standard-frames-plus-IMU baselines.

  • 3) Hovering:: During hovering, the events-only pipeline produces a drifting state estimate.Hovering represents a near-no-motion condition encountered by a drone.
  • 3) Hovering:: Using images together with event frames, the drone maintains its position without noticeable drift while hovering.The experiment directly contrasts this behavior with the events-only pipeline.
  • 3) Hovering:: Standard-frame features remain trackable during hovering, whereas event-frame features are lost.The difference is attributed to successful feature tracking on standard frames and feature loss on event frames.
  • 3) Hovering:: Vibrations during hovering frequently change motion direction, shortening event-stream feature tracks and increasing drift.The event camera’s tracked-feature appearance can change drastically with motion direction, unlike standard cameras.
  • V. CONCLUSIONS: The proposed contribution is a hybrid pipeline fusing events, standard frames, and inertial measurements for robust and accurate state estimation.The conclusion identifies these three sensing modalities as the basis of the pipeline.
  • V. CONCLUSIONS: 130 % accuracy boost was achieved over using only events plus IMU on the Event Camera Dataset.The reported comparison is against the events-plus-IMU configuration.
  • V. CONCLUSIONS: 85 % accuracy boost was achieved over using only standard frames plus IMU on the Event Camera Dataset.The reported comparison is against the standard-frames-plus-IMU configuration.
  • V. CONCLUSIONS: The pipeline enabled onboard state estimation on a computationally-constrained quadrotor and the first closed-loop flight of a quadrotor using an event camera.The paper qualifies this as the first such flight to the best of its knowledge.

APPENDIX

The appendix provides detailed comparisons of pipeline performance on the Event Camera Dataset across three sensor configurations: standard frames, events, and IMU; events and IMU; and images and IMU.

  • Figure 8: Figure 8 compares Event Camera Dataset pipeline performance using standard frames, events, and IMU; events and IMU; and images and IMU.The configurations are denoted Fr, E, and I for standard frames, events, and IMU.
  • Figure 9: Figure 9 continues the detailed Event Camera Dataset comparison across the same three sensor configurations.The compared configurations are standard frames, events, and IMU; events and IMU; and images and IMU.
Loading 1709.06310v4…