Source-linked AI summary
Cooperative ISAC for Joint Localization and Velocity Estimation in Cell-Free MIMO Systems
Zihuan Wang, Vincent W. S. Wong, Robert Schober
TL;DR
Cooperative cell-free ISAC must gather global sensing information without incurring the heavy fronthaul cost of forwarding raw signals. The paper therefore uses locally encoded and quantized AP observations with CPU-side fusion, and reports lower overhead with maintained sensing performance.
Problem
Centralized sensing requires high-dimensional raw signals from each AP, creating substantial fronthaul signaling overhead, while prior two-phase methods can suffer error propagation.
Method
D-VQVAE uses distributed AP encoders and codebooks to locally compress and quantize sensing signals, forwarding codeword indices to a CPU decoder for joint estimation.
Results
The proposed D-VQVAE outperforms baseline schemes in sensing accuracy while reducing fronthaul signaling overhead by 99% versus centralized sensing.
Takeaways & Limitations
Local compression and quantization let the CPU fuse distributed sensing information while using substantially less fronthaul data.
Abstract
from arXiv · showhide
In this paper, we explore a cooperative integrated sensing and communication (ISAC) framework that utilizes orthogonal frequency division multiplexing (OFDM) waveforms. Under the control of a central processing unit (CPU), multiple access points (APs) collaboratively perform multistatic sensing while providing communication service in a cell-free multiple-input multiple-output (MIMO) system. Achieving high sensing accuracy requires the collection of global sensing information at the CPU, which can lead to significant fronthaul signaling overhead due to the feedback of the sensing signals from each AP. To tackle this issue, we propose a collaborative processing scheme in which the APs locally compress and quantize the received sensing signals before forwarding them to the CPU. The CPU then aggregates the information from all APs to estimate the location and velocity of the targets. We develop a distributed vector-quantized variational autoencoder (D-VQVAE) to enable an end-to-end implementation of this scheme. D-VQVAE consists of distributed encoders at the APs to locally encode the received sensing signals, codebooks for quantizing the encoded results, and a decoder at the CPU for location and velocity estimation. It effectively reduces the amount of data transmitted from each AP to the CPU while maintaining a high sensing accuracy. We employ a collaborative learning-assisted scheme to train D-VQVAE in an end-to-end manner. Simulation results show that the proposed D-VQVAE network outperforms the baseline schemes in sensing accuracy and reduces fronthaul signaling overhead by 99% when compared with the centralized sensing approach.
I. INTRODUCTION
Cooperative ISAC extends cell-free MIMO with multistatic sensing, but collecting raw sensing signals centrally creates substantial fronthaul overhead. The paper addresses this trade-off with distributed preprocessing, quantization, and CPU-side fusion for joint target localization and velocity estimation.
- Motivation: OFDM-based ISAC jointly supports communication and sensing, using reflected signals to estimate target range and velocity.OFDM provides high data rates, Doppler tolerance, and no range-Doppler coupling, while communication data introduce phase shifts that sensing must account for.
- Motivation: Cell-free cooperative ISAC distributes APs across the coverage area to provide multiview sensing and coordinated communication through a CPU.Distributed APs improve spatial diversity, sensing range, and access reliability compared with a single monostatic BS.
- Research gap: Existing localization and velocity-estimation methods often use two phases, allowing errors in range, angle, and radial velocity to propagate.The paper’s earlier CNN approach directly estimated location and velocity but required high-dimensional sensing signals at the CPU.
- Proposed approach: The proposed collaborative scheme splits sensing between receive APs and the CPU: APs compress, extract features, and quantize locally, while the CPU fuses the results.This creates a trade-off between fronthaul signaling overhead and sensing performance relative to fully distributed and centralized approaches.
- Proposed approach: D-VQVAE uses distributed AP encoders and codebooks together with a CPU decoder to convert reflected sensing signals into quantized representations for target estimation.Only quantized codeword indices are forwarded, reducing the data sent over fronthaul links while preserving sensing information.
D. Paper Structure and Notations
The paper defines a cell-free MIMO cooperative ISAC system with distributed transmit and receive APs, OFDM communication signals, reflected target echoes, and CPU-based estimation. It then specifies the target variables, array geometry, and notation used throughout the model.
- Paper structure: The paper is organized around the cooperative ISAC system model, D-VQVAE design, collaborative training, performance evaluation, and conclusions.The system model is introduced first, followed by the proposed network, training procedure, evaluation, and conclusion.
- Notation: The notation section defines vector, matrix, complex-number, expectation, diagonalization, and norm conventions used in the equations.It also specifies conjugation, transpose operations, real and imaginary parts, and the imaginary unit.
- System model: Transmit APs send OFDM signals to users and targets, while receive APs collect the reflected sensing signals for CPU processing.The system contains N transmit APs and M receive APs connected to the CPU through synchronized fronthaul links.
- Target variables: The CPU estimates the location and velocity vectors of all Q point-like targets in a 2D coordinate system.The joint target state is represented by ψ, which collects all target locations and velocities.
- Array geometry: Transmit and receive APs use uniform linear arrays whose orientations are known by the CPU.The model defines angle of departure and angle of arrival relative to the array geometry, along with antenna spacings and carrier wavelength.
A. Signal Model
The signal model uses OFDM transmission from distributed APs, centralized MMSE beamforming, and received signals processed through standard RF and Fourier operations. Perfect CSI and fixed transmit beamforming are assumed for the sensing analysis.
- OFDM Transmission: OFDM symbols are transmitted across Ns subcarriers and Ts intervals, with cyclic-prefix insertion after IDFT processing.The transmit power at each AP is constrained by P.
- OFDM Transmission: Each transmit AP precodes user data using per-user beamforming vectors before jointly serving the communication users.The transmit signal combines the assigned users’ symbols through the AP-specific precoders.
- Communication Reception: The users’ received subcarrier signals combine AP transmissions through the communication channels with additive noise.Reception includes down-conversion, ADC, cyclic-prefix removal, and DFT processing.
- Beamforming Assumption: The CPU assumes perfect CSI and uses centralized MMSE beamforming to mitigate multiuser interference.The analysis focuses on target sensing under a fixed transmit beamforming design.
C. Sensing Model
The sensing model describes multistatic target echoes collected by distributed receive APs, whose delays, Doppler shifts, angles, and pathloss encode target geometry and radial motion. The paper seeks target location and velocity estimates from these multiview signals while noting practical modeling assumptions.
- Echo Formation: Transmit signals reflected by targets are collected at receive APs after sampling and DFT processing.The model assumes a line-of-sight path between each transmit or receive AP and target.
- Geometry: Bistatic range is the sum of the transmit-target and target-receive distances, dn,m,q = dn,q + dm,q.The pathloss model uses this bistatic range for the transmit AP, target, and receive AP path.
- Sensing Information: The received sensing signal contains angle, range, and radial-velocity information through AoA, delay, and Doppler terms.Radial velocities are defined relative to the corresponding transmit and receive APs.
- Estimation Objective: The objective is to estimate target locations and velocities from sensing signals gathered by multiple receive APs.The paper contrasts distributed approaches with centralized deep-learning sensing that directly maps echoes to these estimates but requires high fronthaul signaling overhead.
- Modeling Assumptions: The sensing-channel model neglects multipath contributions for simplicity, while simulations evaluate their impact on sensing performance.This assumption bounds the nominal channel model rather than excluding multipath from evaluation.
III. D-VQVAE FOR COOPERATIVE ISAC-ASSISTED LOCALIZATION AND VELOCITY ESTIMATION
D-VQVAE splits cooperative sensing between receive APs and the CPU: APs compress, extract, and quantize sensing features, while the CPU fuses them for localization and velocity estimation. The architecture uses distributed encoders, shared codebooks, and CPU decoding to reduce fronthaul signaling.
- Collaborative Processing: Receive APs locally compress sensing signals, extract sensing-related features, and quantize the results before forwarding them to the CPU.The CPU fuses information from all receive APs for target localization and velocity estimation.
- A. Encoder Design: The encoder aggregates normalized real and imaginary sensing components across subcarriers and OFDM intervals into a 3D input tensor.The tensor preserves space, frequency, and time structure in the received sensing signals.
- A. Encoder Design: 3D CNNs downsample the sensing tensor and extract joint space-frequency-time features, using real and imaginary parts as two input channels.This design targets sensing-related feature extraction from the reflected signals.
- B. Codebook-Based Vector Quantization: Each AP applies codebook-based vector quantization to convert continuous latent features into discrete latent features.A codebook contains Nc codewords of dimension D, and each continuous feature is mapped to a codeword.
- B. Codebook-Based Vector Quantization: Only selected codeword indices are sent over the fronthaul, reducing signaling overhead while the CPU reconstructs discrete latent vectors.The CPU receives the indices and uses the recovered features for downstream target estimation.
C. Decoder Design
The decoder combines quantized features from all receive APs with transmitted OFDM signals to estimate target locations and velocities at the CPU. The design processes these inputs through space-frequency-time feature extraction and fully connected output layers, while the overhead analysis contrasts compressed indices with distributed and centralized alternatives.
- Decoder architecture: The CPU reconstructs each receive AP’s discrete latent feature matrix and uses it with the transmit signals to estimate target locations and velocities.The reconstructed features are aggregated across receive APs before decoding.
- Decoder architecture: Concatenated receive-AP feature tensors are processed across the space-frequency-time domain using residual 3D CNN layers.The tensors are formed by reshaping each AP’s latent matrix and concatenating them along the first dimension.
- Decoder architecture: The decoder separately processes transmit OFDM signals, using their real and imaginary parts as two convolutional channels before feature fusion.The transmit tensor aggregates signals across transmit APs, subcarriers, and OFDM symbol intervals.
- Decoder architecture: After residual feature extraction and linear projection, concatenated features pass through fully connected layers that output location and velocity vectors for all targets.The output dimensions are 2Q for locations and 2Q for velocities.
- Signaling overhead: Centralized sensing sends high-dimensional raw sensing tensors to the CPU, whereas the proposed scheme forwards only quantized codeword indices after local encoding.The proposed overhead per receive AP is NbL bits, with Nb = log2 Nc.
- Signaling overhead: The proposed compressed representation preserves useful sensing information while reducing signaling relative to centralized sensing.The text frames the method as a trade-off between fully distributed and centralized sensing.
IV. COLLABORATIVE LEARNING-ASSISTED TRAINING OF THE D-VQVAE NETWORK
Collaborative learning jointly trains AP-side encoders and codebooks with the CPU-side decoder using sensing inputs, transmit signals, and target labels. The objective combines estimation accuracy with commitment of encoder outputs to discrete codewords, while codebooks are updated by EMA.
- Training data: Training samples pair transmit OFDM signals and reflected sensing signals with target location and velocity labels.Each receive AP observes only its own reflected sensing signals, while the CPU has the transmit signals and labels during training.
- Training framework: The collaborative framework jointly optimizes distributed encoders, codebooks, and the CPU decoder.Training is performed between the CPU and receive APs.
- Optimization objectives: The estimation loss minimizes mean-squared discrepancy between ground-truth and estimated target locations and velocities.This loss updates the distributed encoders and decoder.
- Codebook updates: Codebooks are updated with an exponential moving average using feature vectors assigned to each codeword.The update uses per-codeword assignment counts and assigned continuous feature vectors.
- Optimization objectives: A commitment loss penalizes mismatch between encoder outputs and selected codewords to stabilize discrete representations.The stop-gradient operator keeps codewords fixed while the encoder is updated through this term.
B. Training Procedure
The training procedure alternates local AP encoding and quantization with CPU decoding and backpropagation. Gradients are propagated back to update AP encoders, while codebooks are updated locally and later made available for online execution.
- Alternating training procedure: Each training step samples data, then receive APs encode reflected sensing signals, quantize features, send discrete features to the CPU, and update codebooks locally.The AP-side operations run in parallel.
- Alternating training procedure: The CPU decodes the AP-derived discrete features together with sampled transmit signals to produce estimated target locations and velocities.The decoder output contains one estimated location-and-velocity set for each input in the batch.
- Backpropagation: CPU-side backpropagation minimizes estimation and commitment losses and updates decoder parameters with Adam.Gradients are computed from estimated results and sampled labels.
- Online execution: During online execution, APs send quantized feature indices to the CPU, whose decoder estimates target location and velocity.The online algorithm performs encoding, transformation, quantization, index forwarding, and CPU decoding.
- Backpropagation: The CPU sends discrete-feature gradients back to the corresponding APs, which copy them to encoder outputs and update encoders locally with Adam.This completes the AP-side backpropagation stage.
- Training schedule: The alternating steps repeat for R training steps per epoch and continue for E epochs before the trained codebooks and decoder are used online.After training, each receive AP sends its trained codebook to the CPU.
V. PERFORMANCE EVALUATION
The simulations evaluate D-VQVAE in a two-transmit-AP, two-receive-AP cell-free MIMO topology over a 100 × 100 m2 area. AP positions are fixed across channel realizations, while users and targets are randomly distributed, with most generated samples used for training.
- Simulation setup: The simulated cell-free MIMO system contains N = 2 transmit APs and M = 2 receive APs in a 100 × 100 m2 coverage area.Each AP is configured with 16 antennas unless otherwise specified.
- Simulation topology: Transmit APs are placed at (25, 0) and (75, 0), while receive APs are placed at (25, 100) and (75, 100).The topology fixes AP locations across channel realizations.
- Modeling assumptions: The modeled setting assumes APs and targets share a horizontal plane and APs use ULAs without elevation diversity.The network can be extended to 3D by adding z-axis location and velocity outputs.
- Simulation topology: Users and targets are randomly distributed within the simulation area for different channel realizations.The topology figure describes this distribution alongside the fixed AP placement.
- Target configuration: The evaluation considers Q = 2 targets with x- and y-direction velocities between −20 m/s and 20 m/s.The targets and users are assumed to be randomly distributed in the area.
- Dataset and training: The dataset contains 20,000 channel-realization samples, with 16,000 used for offline training and 4,000 for online testing.The learning rate is set to 10^-4, and samples are normalized during training.
A. Baselines and Benchmark
The evaluation compares D-VQVAE with monostatic, fully distributed, and other DNN-based sensing schemes across operating conditions. Cooperative and DNN-based processing improve robustness and estimation accuracy, including under multiple targets and limited observations.
- The benchmark includes monostatic sensing, CS-based fully distributed sensing, and MUSIC-based distributed sensing schemes.
- Cooperative ISAC significantly improves sensing performance over monostatic single-BS sensing by exploiting multiview observations from distributed APs.The monostatic scheme is more susceptible to noise, whereas distributed observations provide more reliable information.
- DNN-based schemes jointly extract space-frequency-time features across APs, improving robustness when target signals overlap or noise is present.This contrasts with fully distributed schemes, whose APs process sensing signals independently.
- As the number of targets increases, conventional monostatic and fully distributed schemes experience significant performance degradation.The comparison is reported for RMSE in location and velocity estimation versus target count Q.
- Fewer OFDM symbols mainly degrade velocity estimation through reduced Doppler resolution, while D-VQVAE mitigates this effect by extracting distributed time-domain features.Location estimation is less sensitive to the number of OFDM symbols because it relies on spatial correlations.
- Increasing receive antennas improves localization through higher spatial resolution but has little impact on velocity estimation.Velocity estimation primarily depends on Doppler shifts rather than additional spatial resolution.
- The estimated target location errors remain below 1 m, while estimated velocity is close to the ground truth in the considered cell-free MIMO system.
- A learning rate of 10^-4 achieves slightly lower training loss than 10^-5, while 10^-3 may not converge to a desirable trained result.The selected learning rate is therefore 10^-4.
C. Signaling Overhead and Online Execution Runtime
The signaling evaluation compares fully distributed, centralized, and D-VQVAE fronthaul costs, together with online execution runtime. D-VQVAE uses quantized latent representations to reduce communication while retaining sensing performance.
- Centralized sensing requires 67 Mbit per update, whereas fully distributed sensing requires 192 bits per receive AP.
- D-VQVAE achieves good performance with a codebook size Nc = 32, corresponding to a 5-bit codebook and 82 kbit fronthaul overhead.Larger codebooks provide only small additional performance improvements while increasing complexity.
- D-VQVAE reduces signaling by downsampling the sensing signals by a factor of 4 in each dimension before quantization.
- The encoder and codebook-based quantization complete in only a few milliseconds, indicating low computational load at each receive AP.Runtime is evaluated for batch size B = 20 on a CPU-GPU computing server.
D. Effect of Multipath Components
The robustness evaluation examines clutter, multipath, and residual clock offsets. Increasing clutter raises RMSE, but D-VQVAE remains satisfactory under high clutter variance and robust to phase errors from asynchronous clocks.
- Multipath and clutter: Increasing the number of clutter scatterers gradually increases RMSE because additional scatterers introduce interference and reduce sensing signal-to-noise ratio.The evaluation varies clutter variance and scatterer count across non-line-of-sight channels.
- Multipath and clutter: Under high clutter variance ˜χ^2 = 1, D-VQVAE still achieves satisfactory sensing performance.
- Clock asynchronism: D-VQVAE remains robust to phase errors caused by unsynchronized clocks when trained with varying clutter levels and timing and frequency offsets.
- Clock asynchronism: The network exploits delay and Doppler phase-slope patterns across subcarriers and successive symbols despite random AP-specific phase shifts.With sufficient OFDM symbols, the overall impact of residual timing and carrier-frequency offsets on accuracy is small.
- The conclusion reports 99% lower fronthaul signaling overhead than centralized sensing while maintaining higher robustness as the number of sensed targets varies.
APPENDIX A PARTIAL DERIVATIVES OF ρi,m[t]
Appendix A derives the partial derivatives of the received sensing signal with respect to target position, velocity, and channel parameters used in the Fisher-information analysis.
- The LoS sensing channel is rewritten using target-dependent phase and amplitude terms before differentiating the received signal.
- The appendix derives partial derivatives of the received signal with respect to target position components.
- It also derives the corresponding derivatives with respect to target velocity components and channel reflection coefficients.