Source-linked AI summary
Lidar for Autonomous Driving: The principles, challenges, and trends for automotive lidar and perception systems
You Li, Javier Ibanez-Guzman
TL;DR
Autonomous-vehicle perception needs accurate information about nearby entities, motivating LiDAR alongside camera- and radar-based systems. This review analyzes automotive LiDAR hardware, ranging and scanning technologies, perception algorithms, and their challenges and trends. It concludes that LiDAR offers precise measurement, camera fusion addresses recognition weaknesses, and future work targets physical accuracy and semantic estimation.
Problem
Autonomous vehicles need accurate perception of nearby vehicles, pedestrians, and other entities, while LiDAR technology choices and perception approaches remain diverse and unsettled.
Method
The paper reviews LiDAR ranging principles, components, scanning methods, commercial solutions, and LiDAR-processing algorithms spanning model-based and deep-learning approaches.
Results
The review finds that LiDAR is highly precise for measurement, camera fusion remedies its recognition weakness, and RangeNet++ results are impressive.
Takeaways & Limitations
Future directions emphasize extracting more accurate physical information and exploiting LiDAR potential for semantic estimation as new technologies progress.
Abstract
from arXiv · showhide
Autonomous vehicles rely on their perception systems to acquire information about their immediate surroundings. It is necessary to detect the presence of other vehicles, pedestrians and other relevant entities. Safety concerns and the need for accurate estimations have led to the introduction of Light Detection and Ranging (LiDAR) systems in complement to the camera or radar-based perception systems. This article presents a review of state-of-the-art automotive LiDAR technologies and the perception algorithms used with those technologies. LiDAR systems are introduced first by analyzing the main components, from laser transmitter to its beam scanning mechanism. Advantages/disadvantages and the current status of various solutions are introduced and compared. Then, the specific perception pipeline for LiDAR data processing, from an autonomous vehicle perspective is detailed. The model-driven approaches and the emerging deep learning solutions are reviewed. Finally, we provide an overview of the limitations, challenges and trends for automotive LiDARs and perception systems.
I. INTRODUCTION
The introduction positions LiDAR as a precise ranging sensor increasingly used with cameras in autonomous-vehicle perception, while surveying competing technologies and processing approaches. The paper reviews LiDAR hardware, perception algorithms, challenges, and future trends.
- Motivation: LiDARs actively illuminate surroundings and estimate ranges from returned laser signals, providing precise physical information for autonomous-vehicle perception.Their outputs support object detection, classification, tracking, and intention prediction.
- Open question: More than 20 companies are developing distinctive LiDAR systems, but which type will dominate autonomous driving remains uncertain.Current high-level autonomous vehicles commonly use LiDAR despite its high cost and moving parts.
- Motivation: LiDAR perception outputs span physical descriptions, semantic categories, and predicted object intentions for autonomous navigation.These levels include pose, velocity, shape, object categories, and likely behavior.
- Motivation: Cameras and LiDARs complement each other because cameras are weak at distance estimation, whereas LiDAR is less satisfactory for object recognition.Their combination supplies more precise physical and semantic information for perception.
- Related approaches: Traditional LiDAR processing is computation-friendly and explicable, while deep learning methods provide stronger semantic-information capabilities.The review covers both model-based processing and emerging data-driven approaches.
- Scope of the review: The review classifies automotive LiDAR technologies, surveys classic and deep-learning perception algorithms, and discusses open challenges and future trends.Its hardware coverage spans ranging principles, components, representative products, and manufacturers.
II. LIDAR TECHNOLOGIES
A LiDAR combines a laser rangefinder with a scanning system to measure distances across directed viewing angles. The review distinguishes direct ToF and coherent FMCW ranging and describes the resulting point-cloud outputs.
- Range measurement: The transmitter emits a near-infrared laser, and reflected energy returns through the scanner to a photodetector for filtering and distance estimation.Signal processing compensates for changes in reflected energy and conditions between transmitter and receiver.
- Outputs: LiDAR outputs include 3D point clouds representing scanned environments and intensities representing reflected laser energies.These outputs provide geometric and returned-energy information for downstream processing.
- System architecture: A LiDAR system consists of a laser rangefinder and a scanning system that directs beams across the field of view.The rangefinder includes a transmitter, photodetector, optics, and signal-processing electronics.
- Scanning system: The scanner steers laser beams at different azimuths and vertical angles, defining the direction of each measurement.The directions are denoted by φ_i and θ_i, with i indexing the pointed beam direction.
- Review framework: The technology review first explains ranging and scanning limitations, then classifies LiDARs and examines commercially available automotive systems.This structure connects operating principles to sensor field of view and product comparison.
- Ranging principles: Direct-detection rangefinders measure pulsed-laser time of flight, whereas coherent FMCW systems infer distance and velocity through Doppler effects.The signal modulation determines the operating principle of the laser rangefinder.
1) LiDAR power equation:
LiDAR received power depends on transmitted energy, range, target reflectance, transmission losses, and system efficiency. ToF estimates range from signal delay, while FMCW methods derive distance and velocity from frequency processing.
- Power model: Received optical power is modeled from transmitted pulse energy, receive-aperture area, range, target reflectance, transmission, and system efficiency.The reflected signal is captured by receiving optics and converted into an electrical signal by a photodetector.
- Power model: Adverse weather scatters and absorbs photons, increasing transmission loss and reducing target reflectivity so less energy reaches the receiver.Fog, rain, dust, and snow therefore weaken received signals.
- Power model: Received power decreases quadratically with range, making objects hundreds of meters away orders of magnitude darker than objects tens of meters away.Increasing transmitter power is constrained by IEC 60825 eye-safety requirements, motivating improvements in optics, detectors, and signal processing.
- Time-of-flight (ToF): ToF LiDAR calculates range from the time difference between transmitted and received laser signals, using light speed and propagation-medium refractive index.ToF systems are prevalent because of their simple structure and signal processing, but transmit power limits maximum range.
- Coherent detection: FMCW LiDAR mixes a chirped local oscillator with the reflected signal to generate an intermediate frequency whose processing yields distance and velocity.Unlike ToF, FMCW measures both quantities simultaneously and can reduce interference from sunlight and other laser sources, but requires high-quality coherent lasers.
B. Laser Transmission and Reception
Automotive LiDAR transmitters use pulsed diode, fiber, and semiconductor laser technologies with different beam, power, cooling, repetition-rate, and integration properties. These trade-offs affect range, resolution, cost, and vehicle packaging.
- Laser sources: ToF LiDAR uses pulsed amplitude-modulated laser signals, commonly generated by pulsed laser diodes or fiber lasers.The transmitter and receiver electronics influence laser rangefinder performance and cost.
- Semiconductor lasers: VCSELs produce circular beams and facilitate two-dimensional arrays, whereas EELs produce elliptical beams that require additional beam-shaping optics.VCSEL arrays can increase resolution, but VCSEL range is shorter because of power limits.
- Diode sources: 905 nm diode sources are cost-effective and compatible with silicon detectors, but have limited pulse repetition rate, lower peak power, and possible cooling requirements.Stacked diode sources also face heat-dissipation challenges and eye-safety constraints as emitted power accumulates.
- Fiber lasers: Fiber lasers provide higher output power, strong pulse repetition and beam quality, and can route beams to multiple sensor locations, but their bulk complicates vehicle integration.Their packaging can result in non-compact systems.
2) Laser wavelength:
Automotive LiDAR wavelength selection balances atmospheric transmission, eye safety, detector compatibility, and cost. NIR remains mainstream, while 1550 nm permits higher eye-safe power but requires costlier detectors and faces additional atmospheric and efficiency constraints.
- Wavelength selection: Laser wavelength selection must jointly consider atmospheric windows, eye-safety requirements, and cost.The commonly used bands are 850–950 nm NIR and 1550 nm SWIR.
- SWIR: 1550nm lasers allow higher eye-safe maximum power than 850–950nm lasers, enabling potentially larger range.This advantage comes with higher detector cost because InGaAs photodiodes are required.
- SWIR: 1550nm detection uses InGaAs photodiodes whose efficiency is lower than mature silicon detectors for NIR lasers, while atmospheric water absorption is stronger.These factors constrain the practical advantages of the longer wavelength.
- Photodetectors: Photodetector choice is closely tied to laser wavelength because photosensitivity depends on the wavelength of received light.Common detector types include PIN photodiodes, APDs, SPADs, and SiPMs.
- Photodetectors: APDs multiply photocurrent and provide higher internal current gain, around 100, and higher SNR than PIN photodiodes.SPADs can reach a gain of 10^6 and detect extremely weak light over long distances.
- Photodetectors: CMOS fabrication enables integrated SPAD arrays that can increase LiDAR resolution while reducing cost and power consumption.SiPMs combine independently operating SPAD microcells to estimate instantaneous photon-flux magnitude.
C. Scanning system
Automotive LiDAR scanning systems steer laser beams through mechanical or solid-state approaches, trading field of view, robustness, cost, and range. Mechanical spinning is established, while MEMS, flash, and OPA designs reduce or eliminate moving parts with distinct constraints.
- Mechanical Spinning: Mechanical spinning LiDAR uses motor-controlled rotating mirrors or prisms to steer beams and create a large field of view.Nodding-mirror and polygonal-mirror systems are the main conventional types.
- Mechanical Spinning: Mechanical spinning provides high SNR over a wide FOV, but its rotating mechanism is bulky and fragile under vehicle vibration.
- MEMS Micro-Scanning: MEMS LiDAR uses chip-integrated mirrors driven by electromagnetic and elastic forces, with single- or dual-axis scanning in resonant or programmed modes.Programmed scanning can dynamically change the FOV and path to focus on critical regions.
- OPA (Optical Phased Array): OPA LiDAR steers beams through optical phase modulators without moving components, but no commercial product was yet available.
D. Current Status of Automotive LiDAR
Automotive LiDAR has progressed from mechanically spinning products toward solid-state, alternative-wavelength, and FMCW technologies. Current development emphasizes mass production, lower cost, robustness, and operation in adverse weather.
- Commercial adoption: Mechanical spinning LiDAR entered mass-produced vehicles, including Audi’s A8 equipped with Valeo’s Scala for automated driving functions.Scala is a four-layer mechanical spinning LiDAR, and the A8 was reported to achieve L3 automated driving functions subject to legislation.
- Technology trends: Companies are developing solid-state scanning systems to reduce cost and improve robustness.Innoviz, Continental, and Quanergy are identified as developers of such systems.
- Technology trends: FMCW LiDAR attracted automotive manufacturer interest, while representative startups Strobe and Blackmore were acquired by Cruise and Aurora.
- Challenges: Adverse weather increases transmission loss and weakens object reflectivity, motivating 1550 nm LiDAR because its higher transmission power is expected to improve harsh-weather performance.
III. LIDAR PERCEPTION SYSTEM
LiDAR perception converts sensor outputs into hierarchical object descriptions through a classic pipeline of detection, tracking, recognition, and motion prediction. The reviewed detection methods emphasize ground filtering, clustering, and range-view processing.
- Pipeline overview: A traditional LiDAR perception pipeline comprises object detection, tracking, recognition, and motion prediction.The perception system interprets the environment into physical, semantic, and intention-related object descriptions.
- Object Detection: Object detection typically filters ground points and clusters non-ground points into object candidates.Detection supplies initial physical information such as object position, while later stages add heading and speed.
- Object Detection: Polar-grid approaches discretize 3D points and can lose raw measurement information.
- Range View: Spherical-coordinate processing represents each LiDAR return by range and fixed or scan-determined angles, naturally filling a range image.For the Velodyne UltraPuck, laser-beam elevation is fixed while azimuth depends on scanning time and motor speed.
- Range View: Range-view processing provides fast access to neighboring points and supports efficient ground segmentation and clustering.One cited 32-beam LiDAR method reached 4 ms on an Intel i5 processor.
B. Object Recognition
LiDAR object recognition uses engineered global or local features followed by supervised classification. The review describes multiple feature and classifier families, while real-time requirements constrain feature complexity.
- Recognition pipeline: Object recognition furnishes semantic information such as pedestrian, vehicle, truck, tree, and building classes.
- Recognition pipeline: Recognition commonly extracts compact object descriptors and applies pretrained classifiers to predict object categories.The process is supervised and relies on ground-truth datasets such as KITTI.
- Feature extraction: Features are broadly divided into global descriptors for whole objects and local descriptors for individual points.Global examples include size, radius, moments, intensity, and PCA-based shape features; local examples include salience features and Spin Images.
- Feature extraction: Spin Images describe point neighborhoods by distances to a normal-defined line or plane, with object-level variants using a central point.
- Classification: Real-time requirements constrain the complexity of recognition features.
- Classification: SVM with an RBF kernel is described as the most popular classifier because of its speed and accuracy, alongside Naive Bayes, KNN, Random Forest, and GBT.Evidential neural networks are noted for handling unknown classes encountered in practice.
C. Object Tracking
LiDAR object tracking combines state estimation, data association, and shape modeling to maintain object identities and physical states over time. Traditional filters remain practical, while maneuver and recurrent models extend prediction beyond fixed kinematics.
- LiDAR multiple object tracking maintains object identities and estimates physical states such as trajectories, poses, and velocities.
- A typical tracker combines a Bayesian single-object state estimator with data association that assigns new detections to existing tracks.
- Kalman-filter variants, including KF, EKF, and UKF, form a popular toolbox for LiDAR tracking under specified motion assumptions.
- IMM filters run multiple motion models in parallel, while particle filters handle cases that do not meet Gaussian-linear assumptions.
- Particle filters require many particles in high-dimensional state spaces, so the KF family is more popular in real-time perception systems.
- LiDAR tracking must model detection shapes as well as positions, but simple 2D bounding boxes are insufficient for general objects.
- Maneuver recognition and recurrent models address the limits of fixed kinematic assumptions for long-term prediction; LSTM outperformed traditional machine-learning methods in driver-intention classification.
E. Emerging Deep Learning Methods
Deep learning now supports LiDAR ground segmentation, detection, tracking, recognition, and point-wise semantic segmentation. Its effectiveness is constrained by LiDAR sparsity, physical sensing limits, and the availability of annotated datasets.
- Deep neural networks can implement ground segmentation, object detection, tracking, recognition, and semantic segmentation in LiDAR perception.
- CNNs detect vehicles from LiDAR points represented in bird’s-eye-view and can combine range-image and BEV features with camera detections.
- LiDAR’s physical limits make pedestrian detection difficult; DENFIDet achieved 52.40% average precision on the KITTI benchmark when reported.
- Deep tracking first generates detection proposals from LiDAR data and images, then estimates tracks by scoring detection associations.
- PointNet enabled semantic segmentation of 3D point clouds, but LiDAR sparsity with distance reduces its suitability for autonomous-driving scenarios.
- Limited massive annotated datasets initially kept LiDAR semantic-segmentation methods from deployment, while SemanticKITTI and RangeNet demonstrated improving performance and speed.
IV. CONCLUSION AND FUTURE DIRECTIONS
The review concludes that automotive LiDAR provides highly reliable physical measurements but weaker semantic descriptions, motivating sensor fusion and continued algorithmic development. It identifies cost, reliability, range, weather, resolution, and integration constraints while highlighting deep learning and evolving hardware as future directions.
- The review covers LiDAR technologies, their operating principles, perception pipelines, challenges, and future trends for autonomous driving.
- Automotive LiDARs face constraints involving cost, safety and reliability standards, measurement distance, adverse weather, image-level resolution, and size.
- Candidate solutions vary laser wavelength, scanning method, and ranging principle, so the review considers future market dominance difficult to predict.
- LiDAR is the most precise of camera or radar for measuring range, making its estimates of object positions, headings, and shapes highly reliable.
- LiDAR’s semantic description is limited by poor resolution and its role as a distance-measuring rather than contextual sensor; camera fusion addresses recognition weakness.
- Precise LiDAR physical information strengthens intention prediction, while applying deep learning to LiDAR 3D data is identified as an important future direction.
- SemanticKITTI and RangeNet++ mark progress beyond the annotated-3D-dataset bottleneck, with the review reporting impressive RangeNet++ results.