Source-linked AI summary

HMS-SCP: Task-Oriented Multi-Scale Semantic Communication for V2X Cooperative Perception

Chun-Yeow Yeoh, Chee Keong Tan, Joanne Mun-Yee Lim, Heng-Siong Lim

arXiv:2608.14603v2cs.NIcs.AIcs.CVcs.IT

TL;DR

Dense urban V2X cooperative perception must balance intermediate-fusion bandwidth demands with reliable long-range detection under realistic wireless impairments. HMS-SCP transmits task-relevant multi-scale semantics through JSCC and, across OPV2V and DAIR-V2X, outperforms the closest method under severe Rayleigh fading by 42.7% and 23.7%, respectively, beyond 50 m.

  • Problem

    Existing cooperative-perception methods inadequately address realistic channel impairments and single-scale semantic encoding, limiting bandwidth-efficient long-range perception in V2X.

  • Method

    HMS-SCP dynamically selects task-relevant features across hierarchical scales and maps them directly to complex-valued JSCC symbols, using multi-scale semantic redundancy for communication.

  • Results

    42.7% and 23.7% far-field gains over the closest competing method on OPV2V and DAIR-V2X, respectively, under severe Rayleigh fading beyond 50 m.

  • Takeaways & Limitations

    HMS-SCP provides consistent hierarchical performance gains across simulated and real-world datasets while balancing semantic richness and communication overhead across V2X settings.

  • Takeaways & Limitations

    The current framework uses a uniform spatial compression ratio across hierarchical scales; adaptive scale-wise allocation remains future work.

Abstract

from arXiv · show

Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending sensing coverage beyond occlusions and mitigating blind spots. While critical for autonomous driving and safety, practical deployments often rely on bandwidth-efficient late fusion. Recently, intermediate fusion has emerged as a promising approach for an optimal bandwidth-accuracy trade-off. However, in dense urban environments, cumulative bandwidth demands can overwhelm network capacity, potentially compromising safety-critical Cooperative Intelligent Transport Systems (C-ITS) functions. To alleviate these problems, this paper proposes Hierarchical Multi-Scale Semantic-Aware Cooperative Perception (HMS-SCP), a robust noise-resilient and bandwidth-efficient framework for task-oriented semantic communication in cooperative perception. HMS-SCP employs a spatial importance predictor to identify task-relevant grid elements at each scale, which are then directly mapped into complex-valued symbols for Joint Source-Channel Coding (JSCC). Unlike prior methods that rely on high-dimensional symbol projections for robustness, HMS-SCP exploits structural semantic redundancy across multiple scales to enhance resilience against channel noise, while maintaining an ultra-low symbol rate. This design significantly reduces bandwidth consumption and mitigates network congestion in high-density vehicular environments. Extensive evaluations on the simulated OPV2V and real-world DAIR-V2X datasets demonstrate that HMS-SCP effectively prevents performance collapse under severe Rayleigh fading and extreme compression ratio, maintaining high-confidence far-field detection with a real-time latency of below 16~ms, well within the safety-critical thresholds for dynamic V2X environments.

I. INTRODUCTION

Cooperative perception is increasingly central to V2X-enabled autonomous driving, but bandwidth constraints and single-scale semantic encoding limit reliable detection of distant actors. HMS-SCP addresses this gap by jointly optimizing hierarchical semantic representations for improved transmission efficiency and noise resilience.

  • Motivation: 1.19 million lives are claimed globally each year by road traffic accidents, motivating ITS development for transportation safety.WHO identifies road traffic accidents as the leading cause of death among people aged 5–29 years.
  • Cooperative perception: V2X communication extends ITS beyond driver assistance toward cooperative perception, prediction, decision making, and control for autonomous driving.Conventional vehicular perception remains constrained by ego-centric sensor suites.
  • V2X constraints: ETSI ITS-G5 typically uses 10 MHz channels with nominal physical-layer data rates up to 6 Mbps, while C-V2X supports flexible 10–20 MHz bandwidths.These systems provide the connectivity foundation for vehicular communication but impose finite bandwidth resources.
  • Semantic communication: Semantic communication replaces dense feature-tensor exchange with importance-aware selection and compact intermediate representations for more communication-efficient, channel-resilient cooperative perception.This approach targets vehicular networks affected by severe and time-varying wireless impairments.
  • Problem: Existing semantic encoding is generally restricted to a single-scale feature map, limiting perception of distant actors represented at smaller feature-hierarchy scales.The resulting limitation contributes to a shorter safety horizon under stringent 6–12 Mbps V2X bandwidth constraints.
  • HMS-SCP contribution: HMS-SCP jointly optimizes hierarchical semantic representations and replaces traditional channel-coding redundancy with multi-scale semantic redundancy.The framework is designed to improve transmission efficiency and noise resilience while targeting the bandwidth-accuracy trade-off needed for high-speed autonomous navigation.

II. RELATED WORKS · A. V2X based Cooperative Perception

V2X cooperative perception addresses the occlusion, sparsity, and viewpoint limits of single-vehicle BEV perception by aggregating complementary observations across distributed agents. Intermediate fusion and semantic communication improve bandwidth efficiency, but existing methods remain limited by channel assumptions, insufficient multi-scale exploitation, and single-scale representations.

  • A. V2X based Cooperative Perception: Single-vehicle BEV perception remains constrained by line-of-sight occlusions, sparse long-range sensing, and viewpoint limitations.
  • A. V2X based Cooperative Perception: V2V- and V2I-enabled cooperative perception aggregates complementary observations across agents to mitigate individual-sensor unreliability in challenging environments.
  • A. V2X based Cooperative Perception: Intermediate fusion has emerged as a dominant paradigm because it balances perception accuracy and communication efficiency.
  • A. V2X based Cooperative Perception: Where2comm selectively transmits perceptually critical regions using a spatial confidence map, reducing communication bandwidth while maintaining high perception performance.
  • A. V2X based Cooperative Perception: Existing intermediate-fusion frameworks largely assume ideal or quasi-static communication channels and use uniform allocation that overlooks multi-scale importance distributions in BEV representations.
  • A. V2X based Cooperative Perception: These methods fail to exploit the hierarchical structure of BEV representations for minimal sufficient transmission and lack robustness in safety-critical scenarios.
  • A. V2X based Cooperative Perception: Semantic communication improves cooperative-perception efficiency by prioritizing task-relevant information, including JSCC with importance maps to extract critical LiDAR semantic features.
  • A. V2X based Cooperative Perception: Existing semantic cooperative-perception architectures are predominantly single-scale, prioritizing regions within one feature representation while overlooking hierarchical perception-backbone features.

III. METHODOLOGY … D. Noise-Resilient Architectural Design

HMS-SCP presents a hierarchical semantic-communication framework that selectively transmits task-critical multi-scale features through a JSCC codec, then reconstructs and fuses them for cooperative perception. Its noise-resilient design combines sparse complex-valued transmission, power-constrained signaling, and stabilized convolutional decoding under wireless impairments.

  • A. Semantic Communication based Cooperative Perception: HMS-SCP selectively encodes and transmits multi-scale backbone features through a semantic communication pipeline for robust, communication-efficient cooperative perception.The framework includes rate-controlled feature selection and a JSCC-based semantic codec that maps selected features directly into channel symbols.
  • B. Cooperative Perception Problem Formulation: Each agent identifies task-relevant spatial grid elements at every scale, transmits sparse semantic representations, and reconstructs them for fusion with local features.The ego-agent combines reconstructed multi-scale features from collaborators with its own features through a collaborative perception network.
  • B. Cooperative Perception Problem Formulation: The global transmission ratio R is tunable to available bandwidth B, while hierarchical scales support dynamic channel-use allocation for nearby and distant objects.Finer scales handle nearby objects and coarser scales provide global context for distant objects, with allocations based on their contribution to the perception objective.
  • C. HMS-SCP Framework: The HMS-SCP architecture comprises multi-scale feature extraction, rate-controlled selection, semantic coding, collaborative fusion, and object detection.LiDAR point clouds are converted into hierarchical BEV features at 2×, 4×, and 8× downsampled resolutions before semantic selection.
  • C. HMS-SCP Framework: A learnable saliency predictor selects the top-k spatial locations at each scale, imposing a sparsity constraint that directly controls the communication budget.The selected features are masked, encoded separately by scale-specific semantic encoders, and mapped into channel symbols under JSCC.
  • C. HMS-SCP Framework: Per-scale attention fusion emphasizes reliable complementary observations while suppressing features corrupted by occlusion, misalignment, or channel distortion.The fused representations are upsampled to a common high-resolution grid, concatenated, and reduced before detection; joint optimization combines detection and reconstruction losses.
  • D. Noise-Resilient Architectural Design: The noise-resilient codec uses stacked multi-scale convolutions, GN, PReLU, sparse complex-valued symbols, power normalization, and symmetric decoding under AWGN or Rayleigh fading.With Cout = 1, each selected location maps to a single complex-valued symbol; scale clipping limits normalization during early training or highly sparse activation.

E. Task-Oriented Training Strategy

HMS-SCP uses a two-stage optimization strategy: first establishing semantic reconstruction, then jointly adapting transmission and perception components for task-oriented accuracy while preserving pretrained geometric features. The second stage also incorporates channel fading and spatial compression during training to improve robustness to transmission distortions.

  • Stage 1: Feature Reconstruction: Stage 1 jointly trains the semantic codec and spatial importance predictor to minimize the multi-scale JSCC reconstruction loss.Training starts from a pretrained model without the semantic codec or spatial importance predictor.
  • Stage 2: Task-Oriented Joint Optimization: Stage 2 freezes the feature-extraction backbone and jointly fine-tunes the hierarchical codec, importance predictor, fusion, and detection architecture for perception accuracy.Freezing preserves learned geometric representations and prevents catastrophic forgetting of geometric features.
  • Stage 2: Task-Oriented Joint Optimization: Fusion and neck layers adapt to channel-induced distortions so the synthesized BEV representation is conditioned for perception accuracy rather than structural similarity.The objective shifts from signal-level reconstruction to task-oriented perception accuracy.
  • Stage 2: Task-Oriented Joint Optimization: Rayleigh fading channels and spatial compression ratios are incorporated during Stage 2 using a methodology consistent with Stage 1.This exposes task-oriented optimization to transmission conditions represented by the wireless channel and spatial compression process.

IV. EXPERIMENTS · A. Experimental Setup

The experiments evaluate HMS-SCP across simulated and real-world cooperative perception datasets, using multi-scale implementation settings and transmission ratios spanning extreme to moderate spatial reduction. Training employs staged optimization with detection prioritized under low-SNR conditions.

  • A. Experimental Setup: HMS-SCP is evaluated on simulated OPV2V and real-world DAIR-V2X datasets covering V2V and V2I collaboration, respectively.OPV2V evaluation uses a maximum collaboration radius of 70 m.
  • A. Experimental Setup: The framework is implemented in OpenCOOD with PointPillar as the LiDAR feature encoder and a voxel grid size of (0.4 m, 0.4 m).
  • A. Experimental Setup: The experiments compare different semantic-aware cooperative perception frameworks using channel resilience, importance prediction, complex-valued symbols, and single- versus multi-scale design dimensions.The comparison terminology defines Ch. as channel-resilient, Imp. as importance predictor, Sym. as complex-valued symbols per selected feature dimension, and SS/MS as single-/multi-scale.
  • A. Experimental Setup: Three hierarchical scales provide fine geometric, object-level semantic, and broad scene-context representations for cooperative perception.The finest scale supports small or distant object detection, while additional scales increase semantic-coding complexity.
  • A. Experimental Setup: R ∈{0.01, 0.05, 0.10} configures 100×, 20×, and 10× spatial reduction regimes, with the predictor selecting top-k percentile spatial features.Each selected feature dimension retains a single complex-valued symbol bottleneck.
  • A. Experimental Setup: Adam training uses learning rates of 0.0002 for stage 1 and 0.0001 for stage 2, with µ = 0.05 in stage 2 to prioritize detection loss.Reconstruction loss serves as a mild regularizer for physically meaningful latent representations under low-SNR conditions.

B. Comparative Performance Analysis

HMS-SCP is evaluated with AP@0.5 and AP@0.7 against Where2Comm and SComCP under varying SNR conditions on OPV2V and DAIR-V2X. It maintains near-upper-bound accuracy and robust far-field perception under severe Rayleigh fading and extreme compression.

  • Quantitative Evaluation: HMS-SCP maintains stable AP performance and graceful degradation at 0 dB under Rayleigh fading, unlike non-channel-resilient baselines and Where2Comm (MS).Where2Comm (MS) exhibits a pronounced cliff effect below 6 dB, despite high performance from 6 to 20 dB.
  • Quantitative Evaluation: 1.99%, 0.75%, and 0.73% are HMS-SCP’s AP@0.5 performance gaps from the upper bound on DAIR-V2X at R ∈ {0.01, 0.05, 0.10}, respectively.At R = 0.05, the 0.75% gap corresponds to recovery of 99.25% of upper-bound detection performance.
  • Reliability Analysis: 12.7% and 23.7% are HMS-SCP’s AP@0.5 and AP@0.7 gains over SComCP at the 100 m horizon on DAIR-V2X, respectively, at an identical 1% spatial sampling ratio.R = 0.01 indicates 100× compression, with 1 complex-valued symbol per feature.
  • Reliability Analysis: 19.7% and 42.7% are HMS-SCP’s AP@0.5 and AP@0.7 gains over SComCP on OPV2V at the identical 1% spatial sampling ratio.The results associate hierarchical semantic representations with a more resilient information bottleneck against channel-induced distortions.
  • Qualitative Evaluation: Beyond 50 m at 0 dB, baseline far-field detections degrade noticeably, whereas HMS-SCP preserves semantic details and produces no false-positive predictions compared with SComCP.The qualitative evaluation covers OPV2V and DAIR-V2X under severe Rayleigh fading at R = 0.01.

C. Ablation Studies

Under 0 dB SNR Rayleigh fading and R = 0.01, ablations show that hierarchical scales and fusion are especially important for far-field cooperative perception. Scale 1 neighbor transmission and high-resolution fusion most strongly affect distant-object detection, while HMS-SCP remains consistent across OPV2V and DAIR-V2X.

  • Ablation setup: At 0 dB SNR Rayleigh fading and R = 0.01, ablations test transmission suppression and fusion ablation across hierarchical feature levels.The experiments evaluate both datasets, with emphasis on far-field perception.
  • Transmission suppression: Suppressing Scale 1 neighbor transmission drops far-field AP@0.5 from 0.8170 to 0.6078 for OPV2V and from 0.6344 to 0.5037 for DAIR-V2X.The collapse confirms that high-resolution neighbor semantics are critical for occluded or distant-object detection, whereas Scale 3 suppression has negligible impact.
  • Fusion ablation: Disabling Scale 1 fusion reduces OPV2V far-field AP@0.5 to 0.5932, while full HMS-SCP yields absolute gains of 13.6% and 1.13% on DAIR-V2X configurations.These results identify high-resolution cooperative fusion as particularly important for fine-grained spatial cues and distant objects.
  • Safety-critical metrics: Suppressing Scale 2 transmission reaches far-field Prec@0.5 of 0.7986 for OPV2V and 0.6031 for DAIR-V2X, but Rec@0.5 falls to 0.8463 and 0.6760, respectively.The precision increase therefore trades off against recall in the safety-critical 50–100 m range.
  • Cross-dataset consistency: Consistent hierarchical dependencies across simulated OPV2V and real-world DAIR-V2X confirm HMS-SCP’s architectural adaptability and stable performance gains under per-dataset evaluation.The conclusion holds across diverse cooperative perception settings.

D. Bandwidth and Communication Efficiency Analysis

HMS-SCP achieves communication efficiency through 1% spatial sampling and one-symbol-per-location semantic mapping. Compared with SComCP, it uses substantially fewer symbols while communicating many more spatial anchors for improved distant-vehicle coverage.

  • Transmission design: HMS-SCP fixes the transmission ratio at R = 0.01, sampling 1% of spatial grid points at each hierarchical scale.Each selected location is mapped to exactly one complex-valued symbol.
  • Comparison with SComCP: Approximately 3,072 symbols per perception cycle result from SComCP transmitting 12 selected semantic features with 256 symbols per feature.SComCP consumes nearly 7 times more bandwidth than HMS-SCP while communicating fewer spatial locations.
  • Spatial coverage: 444 vs. 12 or 331 vs. 12 spatial anchors are relocated by HMS-SCP relative to SComCP through its 1-to-1 semantic mapping.The denser spatial representation helps capture distant vehicles that sparse models may overlook.
  • Communication footprint: 331 transmitted elements constitute HMS-SCP’s remarkably small communication footprint.This footprint supports a balance between transmission volume and perceptual reliability.

E. Complexity and Latency Analysis

HMS-SCP is evaluated for practical real-time V2X deployment using learnable-parameter complexity and inference latency on OPV2V and DAIR-V2X. It achieves low latency and sufficient throughput for high-speed collaborative 3D detection in dynamic urban environments.

  • Complexity: 8.70M learnable parameters comprise the proposed HMS-SCP architecture.The complexity analysis quantifies computational overhead in terms of learnable parameters.
  • Latency: 15.73 ms end-to-end inference latency is achieved on OPV2V, corresponding to approximately 63 FPS.Latency is measured on the OPV2V benchmark.
  • Latency: 12.17 ms end-to-end inference latency is achieved on DAIR-V2X, corresponding to approximately 82 FPS.Latency is measured on the DAIR-V2X benchmark.
  • Practical Feasibility: The results confirm sufficient throughput for high-speed collaborative 3D detection in dynamic and complex urban environments.This conclusion follows the reported latency and throughput evaluation.

V. CONCLUSION

HMS-SCP is a hierarchical multiscale semantic-aware framework for task-oriented semantic communication in V2X collaborative 3D object detection. Experiments on OPV2V and DAIR-V2X show consistent gains over SOTA baselines, including substantial far-field improvements under severe Rayleigh fading.

  • Framework: HMS-SCP dynamically identifies critical semantic features across hierarchical feature levels using a spatial importance predictor.The selected features are directly mapped to complex-valued symbols via a JSCC codec.
  • Experimental validation: HMS-SCP consistently outperforms SOTA baselines on simulated OPV2V and real-world DAIR-V2X under AWGN and Rayleigh fading channels.These experiments evaluate the proposed framework across both simulated and real-world cooperative-perception datasets.
  • Far-field performance: 23.7% on DAIR-V2X and 42.7% on OPV2V: HMS-SCP outperforms the closest competing method under severe Rayleigh fading beyond the 50 m far-field horizon.The result confirms the effectiveness of hierarchical semantic transmission for long-range cooperative perception.
Loading 2608.14603v2…