Source-linked AI summary

Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

Meda Lazar, Sourab Sridhar, Shashwata Gupta, Alexandra Tripcea, Varun Ravi, Senthil Yogamani

arXiv:2608.20874v1cs.CVcs.RO

TL;DR

Traffic-sign perception must generalize across regions, detect small distant signs, and remain temporally stable under nonlinear perspective effects. The paper combines camera and LiDAR sensing with nonlinear-motion tracking and semantic attributes, achieving a 0.49% Object Miss Ratio across 221,068 evaluation sequences.

  • Problem

    Traffic-sign detection remains limited by regional appearance variability, long-range sensing challenges, and fragile tracking during vehicle approach.

  • Method

    The framework fuses LiDAR depth and reflectance with camera features, uses dual motion models for nonlinear perspective transformations, and classifies operational semantic attributes.

  • Results

    0.49% Object Miss Ratio is achieved across 2,500+ hours of driving data in the large-scale evaluation.

  • Takeaways & Limitations

    The pipeline provides region-invariant traffic-sign perception with temporal consistency and operational context for downstream planning.

  • Takeaways & Limitations

    Initial 3D detection was constrained by sensitivity to calibration drift and sensor desynchronization, motivating a 2D multimodal detector.

Abstract

from arXiv · show

Reliable traffic sign detection is a prerequisite for the global deployment of autonomous driving systems, where regulatory compliance and road safety depend on perceiving signs correctly across regions, ranges, and weather conditions. Despite recent progress, vision-based methods continue to face three fundamental limitations: poor cross-regional generalization due to high diversity across countries, degraded performance on small-object detection at long ranges (traffic signs occupy as little as $10{\times}10$ pixels at 200m), and fragile temporal tracking under the strongly non-linear perspective distortion that occurs as a vehicle approaches a sign. In this paper, we address the problem of robust, long-range, region-agnostic traffic sign perception by combining camera and Light Detection and Ranging (LiDAR) sensing. We present a multi-modal detection framework whose Intensity-Aware Deformable Fusion module aligns retro-reflective LiDAR cues with camera features, anchoring detection on geometric invariants rather than region-specific visual appearance. We further introduce a dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach, substantially improving temporal consistency over linear motion assumptions. Additionally, we develop a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedness, and road relevance, providing actionable context to downstream planning. Extensive evaluation on our dataset, spanning 60+ countries and 2,500+ hours of driving data, shows that the proposed pipeline achieves an Object Miss Ratio (OMR) of 0.49% across 221,068 evaluation sequences, demonstrating globally generalizable traffic sign perception in commercial-grade autonomous driving systems.

I. INTRODUCTION

Traffic-sign detection remains difficult because global sign diversity, adverse conditions, and limited multimodal benchmarks undermine reliable long-range perception. The paper addresses these gaps with LiDAR-camera fusion, nonlinear-motion tracking, semantic attributes, and large-scale validation.

  • Traffic-sign detection supports regulatory compliance and road safety but remains challenging across diverse categories, ranges, and environmental conditions.
  • Public benchmarks are geographically narrow, often image-only, and insufficiently representative of global regulatory standards and real-world deployment conditions.
  • LiDAR-camera fusion anchors detection on physical sign geometry rather than region-specific visual appearance, targeting geographic generalization.
  • The dual motion model explicitly represents nonlinear perspective transformations, improving temporal stability and reducing track fragmentation against linear models.
  • Semantic attributes estimate occlusion, readability, relevance, and embeddedness to filter non-actionable detections for downstream planning.
  • 0.49% Object Miss Ratio is achieved across 2,500+ hours of driving data on a large-scale dataset spanning diverse global regions.

B. LiDAR-Camera Fusion for Object Detection

The paper situates its approach against camera-only detection, tracking-by-detection, and general-purpose multimodal fusion. It evaluates a larger, geographically broader dataset with synchronized multimodal annotations and long-range coverage.

  • Camera-only traffic-sign detectors remain vulnerable to environmental variation, while tracking-by-detection methods dominate online multi-object tracking.
  • The method uses 2D-centric LiDAR-camera fusion chosen for tolerance to long-range cross-sensor misalignment and exploits traffic-sign retro-reflective signatures.
  • The tracker replaces a single constant-velocity assumption with parallel constant-acceleration and constant-jerk models for vehicle-approach dynamics.
  • The Qualcomm dataset spans 60+ countries and contains over 142 million annotated objects for diverse deployment conditions.
  • Synchronized 2D and 3D annotations and high-fidelity 3D labels extend to 200 m, supporting multimodal long-range localization.

B. Evaluation Process and Metrics

The evaluation uses COCO-style detection metrics alongside standard statistical measures for tracking and semantic attribute classification.

  • Average Precision averages detection results over IoU thresholds from 0.50 to 0.95 in 0.05 increments.
  • AP50 and AP75 report precision at IoU thresholds of 0.50 and 0.75, respectively.
  • Scale-specific AP evaluates small, medium, and large objects using thresholds below 322 px, 322–962 px, and above 962 px.
  • Accuracy is computed from true positives, true negatives, false positives, and false negatives.

3) End-to-End Efficacy:

The paper develops a 2D camera–LiDAR framework after identifying scalability and reliability problems in an initial 3D approach, especially for sparse long-range sign returns and sensor misalignment.

  • The initial 3D detector failed to scale because traffic signs provide only a handful of LiDAR returns beyond 100 m.Sparse returns make stable 3D bounding-box regression difficult without substantial visual support.
  • Calibration drift and sensor desynchronization caused 3D LiDAR proposals to detach from corresponding sign imagery.The resulting spatial shear caused fusion with empty 3D space and substantially reduced precision.
  • The proposed 2D-centric detector combines LiDAR depth and intensity with high-resolution imagery to address extreme-range visual degradation.
  • The complete pipeline detects signs, links detections into tracks, and classifies semantic attributes after post-processing.
  • Temporal point-cloud accumulation increases density, projects points into depth and intensity maps, and supplies a geometric anchor for visual detection.Median filtering suppresses depth noise, while max-pooling preserves peak retro-reflectivity.

2) Intensity-Aware Deformable Fusion:

Intensity-Aware Deformable Fusion dynamically aligns camera and LiDAR features using geometric and reflective cues, while the tracker models nonlinear perspective changes with complementary motion filters.

  • 2) Intensity-Aware Deformable Fusion: Intensity-Aware Deformable Fusion predicts sampling offsets from joint geometric-intensity representations rather than visual features alone.
  • 2) Intensity-Aware Deformable Fusion: The fusion mechanism pulls RGB features toward high-reflectivity regions to provide self-correcting cross-modal alignment.
  • 2) Intensity-Aware Deformable Fusion: Geometric and intensity attention jointly emphasize structurally consistent, semantically relevant reflective regions while suppressing non-sign objects.
  • 2) Intensity-Aware Deformable Fusion: The fused features pass through a SwinTransformer and multi-scale transformer detection head for fine-grained features and contextual cues.
  • 1) Dual Motion Models: The tracker runs third-order constant-jerk and second-order constant-acceleration filters in parallel to capture different perspective regimes.
  • 1) Dual Motion Models: Adding the second-order and then third-order filters produced incremental improvements over the baseline tracker.

2) Multi-Stage Data Association:

The association stage combines spatial, shape, and overlap cues to improve matching when blur or tiny distant targets make IoU unreliable. A dual-filter tracker reconciles second- and third-order predictions using posterior covariance and propagates a stable track for semantic labeling.

  • 2) Multi-Stage Data Association:: A composite matching cost combines Mahalanobis distance, shape similarity, and IoU for robust association.The weights are tuned to support tracking when detection confidence or IoU alone is insufficient.
  • 2) Multi-Stage Data Association:: Second- and third-order filters independently update each active track at every frame.
  • 2) Multi-Stage Data Association:: Overlapping predictions retain the estimate with lower posterior covariance trace, while non-overlapping predictions remain independent observations.
  • 2) Multi-Stage Data Association:: The unified track supports consistent assignment of semantic attributes across challenging sequences containing occlusions, distant signs, and perspective variation.

V. SEMANTIC ATTRIBUTE CLASSIFICATION

The semantic attribute layer augments 2D sign localization with operational context. It models occlusion, readability, embeddedness, and relevance so downstream systems can prioritize actionable detections and assess visibility-related performance separately.

  • V. SEMANTIC ATTRIBUTE CLASSIFICATION: Each detection is represented as a 2D bounding box paired with a composite semantic attribute vector.
  • V. SEMANTIC ATTRIBUTE CLASSIFICATION: The attribute vector encodes occlusion, readability, embeddedness, and relevance at sign level.These properties distinguish visibility, interpretability, hierarchy, and operational significance.
  • V. SEMANTIC ATTRIBUTE CLASSIFICATION: The classification layer filters irrelevant or unreadable signs and reduces downstream motion-planning input.
  • V. SEMANTIC ATTRIBUTE CLASSIFICATION: Occlusion classification uses cropped tracked regions, LoRA-adapted DINOv2 features, and five levels from 0% to 100%.
  • V. SEMANTIC ATTRIBUTE CLASSIFICATION: Occlusion filtering separates visibility-limited instances from algorithmic failures when evaluating 2D detection reliability.

C. Embedded Attribute

The embedded attribute identifies signs physically nested within larger panels and stabilizes this relationship across tracks. Relevance filtering then uses spatial, distance, and orientation constraints to distinguish actionable signs from operationally inert detections.

  • C. Embedded Attribute: Embedded signs are modeled explicitly to reduce redundant detections and misrouted downstream processing.
  • C. Embedded Attribute: A candidate is embedded only when its box is contained, its 3D centroid is sufficiently close, and its orientation is consistent with the parent.
  • C. Embedded Attribute: Once triggered in one frame, the embedded attribute propagates across the entire track as a static property.This prevents attribute flickering caused by occlusion or noise.
  • C. Embedded Attribute: Relevance filtering combines drivable-area membership, a 100 m operational horizon, and geometric orientation checks.The detector reports signs to 200 m, while relevance marks the tighter actionable planning horizon.
  • C. Embedded Attribute: Evaluation projects LiDAR-frame annotations onto images and evaluates objects through 200 m without filtering.

B. Tracker

The tracker is evaluated under challenging viewpoint and occlusion conditions, where conventional methods improve precision but lose recall. Higher-order kinematic models recover recall, with the third-order model achieving the strongest reported tracking result, while DINOv2 with LoRA performs best among tested occlusion classifiers.

  • B. Tracker: Conventional trackers achieve higher precision than the single-frame detector but lower recall under perspective shifts.OC-SORT and BoostTrack++ report precision of 0.70 and 0.72, with recall of 0.68 and 0.69, versus detector recall of 0.73.
  • B. Tracker: 0.74 Recall and 0.71 Precision are achieved by the third-order kinematic model.Adding jerk to the state improves recall over the single-frame detector’s 0.73.
  • B. Tracker: The third-order model addresses non-linear image-plane motion caused by changing perspective during vehicle approach.
  • C. Occlusion Classification: DINOv2 + LoRA achieves 0.92 recall and 0.93 precision for occluded signs, plus 0.74 recall and 0.72 precision for non-occluded signs.
  • C. Occlusion Classification: At 75% occlusion, DINOv2 + LoRA improves recall from 0.13 to 0.36 versus frozen DINOv2.
  • C. Occlusion Classification: DINOv2 + LoRA outperforms ViT-Base in recall at 0% and 100% occlusion, with values of 0.74 vs. 0.71 and 0.59 vs. 0.51.

D. Readability Classification

The readability classifier evaluates whether traffic signs are readable, not readable, or visible from the back, using camera-only and multi-modal approaches. The section reports the classifier’s performance and embedded-sign examples.

  • Readability is categorized as Readable, NotReadable, or BackOfSign to determine which signs should influence the vehicle’s ego-path.
  • The section compares a vision-only readability approach with multi-modal fusion using LiDAR reflectivity to resolve visual ambiguities.
  • Figures 8 and 9 show correctly classified NotEmbedded edge cases and embedded signs, respectively, with embedded signs marked by green boxes.
  • The embedded-attribute classifier achieves 0.99 recall and 0.99 precision on the full evaluation set containing embedded and non-embedded signs.

F. Relevance Classification

Relevance classification combines lane-distance and point-inside constraints to identify signs associated with a driving lane. The broader evaluation also examines LiDAR configurations, temporal accumulation, and performance across diverse dataset conditions.

  • At a fixed 5 m threshold, the point-inside constraint supplements lane distance by requiring the sign to lie within a lane-derived polygon.
  • The relevance classifier reports class-wise precision and recall for Relevant and NotRelevant assignments under lane-distance and combined rules.
  • The proposed all-LiDAR configuration reaches 0.60, while camera plus LiDAR reaches 0.64 in the reported 3D detection results.
  • Multi-sweep LiDAR accumulation improves overall AP from 0.55 to 0.60 and raises 150–200 m AP from 0.14 to 0.20.
  • Absolute AP values differ between Zenseact and Qualcomm because their geographic coverage, sign distributions, and annotation densities differ.

H. Comparison of Camera-Only and Camera+LiDAR 2D Detector

The 2D detector remains primarily image-based while adding LiDAR depth and intensity as complementary geometric and reflectivity cues. The broader evaluation reports robustness across illumination, road types, weather, and global driving variability.

  • LiDAR depth and intensity maps improve the image-domain detector by supplying complementary geometric and reflectivity information.
  • The largest AP gain occurs for small objects, increasing by +0.04 from 0.47 to 0.51, compared with +0.02 for medium and +0.01 for large objects.
  • The dataset captures long-tail global variability across 60+ countries, including differing sign shapes, LED displays, and multilingual content.
  • The pipeline achieves a total Object Miss Ratio of 0.49% across 221,068 evaluation sequences representing over 2,500 hours of driving data.
  • Object Miss Ratio varies by illumination and road type, with Night reaching 0.17% versus 0.67% during Day and urban sequences performing best.
  • Fog is the most challenging weather condition at 0.72% miss ratio because it degrades camera contrast and range while increasing LiDAR backscatter noise.
  • Rare weather categories such as thunderstorm and hail contain very few sequences, so their miss ratios are indicative rather than statistically conclusive.
Loading 2608.20874v1…