Source-linked AI summary

Pruning the Pilots: Deep Learning-Based Pilot Design and Channel Estimation for MIMO-OFDM Systems

Mahdi Boloursaz Mashhadi, Deniz Gunduz

arXiv:2006.11796v3cs.ITeess.SPstat.ML

TL;DR

Pilot overhead makes wideband massive-MIMO channel estimation difficult, especially when short pilots leave estimation underdetermined. The paper jointly learns frequency-aware pilot design and downlink channel estimation with neural networks, adding attention and pruning-based pilot reduction. The reported scheme outperforms LMMSE estimation while reducing overhead through non-uniform pilot allocation across subcarriers.

  • Problem

    Pilot lengths smaller than antenna counts make massive-MIMO channel estimation severely underdetermined, while simple LS, LMMSE, and orthogonal FFT pilots perform poorly.

  • Method

    A neural network jointly designs frequency-aware pilots and estimates channels, using dense layers, convolutional layers, and pruning to allocate pilots non-uniformly.

  • Results

    The proposed NN-based pilot design and channel estimation scheme outperforms LMMSE estimation, while pruning effectively reduces pilot overhead through non-uniform subcarrier allocation.

  • Takeaways & Limitations

    Learning channel statistics and correlations enables joint pilot design and estimation with reduced pilot transmission across subcarriers.

Abstract

from arXiv · show

With the large number of antennas and subcarriers the overhead due to pilot transmission for channel estimation can be prohibitive in wideband massive multiple-input multiple-output (MIMO) systems. This can degrade the overall spectral efficiency significantly, and as a result, curtail the potential benefits of massive MIMO. In this paper, we propose a neural network (NN)-based joint pilot design and downlink channel estimation scheme for frequency division duplex (FDD) MIMO orthogonal frequency division multiplex (OFDM) systems. The proposed NN architecture uses fully connected layers for frequency-aware pilot design, and outperforms linear minimum mean square error (LMMSE) estimation by exploiting inherent correlations in MIMO channel matrices utilizing convolutional NN layers. Our proposed NN architecture uses a non-local attention module to learn longer range correlations in the channel matrix to further improve the channel estimation performance. We also propose an effective pilot reduction technique by gradually pruning less significant neurons from the dense NN layers during training. This constitutes a novel application of NN pruning to reduce the pilot transmission overhead. Our pruning-based pilot reduction technique reduces the overhead by allocating pilots across subcarriers non-uniformly and exploiting the inter-frequency and inter-antenna correlations in the channel matrix efficiently through convolutional layers and attention module.

I. INTRODUCTION

FDD massive MIMO channel estimation faces substantial pilot overhead and underdetermined estimation when pilot length is smaller than the antenna count. The paper proposes a neural-network scheme that jointly designs frequency-aware pilots and estimates channels, with pruning to reduce overhead.

  • Motivation: Large antenna and user counts make FDD downlink pilot and CSI-feedback overhead significant, motivating efficient pilot design and channel estimation.In FDD, users estimate downlink channels from broadcast pilots and feed CSI back to the base station.
  • Motivation: Pilot lengths smaller than the number of antennas make channel estimation severely underdetermined, causing LS, LMMSE, and orthogonal FFT pilots to perform poorly.This motivates methods that exploit additional channel structure for more efficient estimation.
  • Prior approaches: Sparse and low-rank compressive-sensing approaches overlook statistical correlations beyond those patterns and impose computationally demanding iterative processing on users.The paper contrasts these assumptions and computational costs with data-driven neural-network approaches.
  • Proposed approach: The proposed FDD massive MIMO-OFDM scheme jointly designs pilots and estimates channels using dense and convolutional layers without requiring covariance matrices or other prior channel assumptions.Dense layers design frequency-aware pilots, while convolutional layers exploit inherent MIMO-OFDM channel correlations.
  • Proposed approach: A non-local attention module learns longer-range channel correlations, extending the convolutional architecture's ability to exploit channel structure.The architecture is designed to learn channel statistics from a dataset and use them for joint pilot and estimator optimization.
  • Pilot reduction: NN pruning reduces pilot overhead by allocating pilots non-uniformly across subcarriers, using fewer pilots where convolutional layers can reconstruct channels from inter-frequency correlations.The paper reports improved estimation NMSE for the same time-frequency resources allocated uniformly across subcarriers.
  • Results: Extensive simulations report that the proposed scheme outperforms ideal LMMSE channel estimation and recent neural-network-based methods, with an ablation study examining pilot design and channel estimation.The introduction also reports improvement over simple frequency-independent pilot schemes.

II. SYSTEM MODEL

The system models downlink channel estimation in an FDD massive MIMO-OFDM grid with pilots transmitted over selected time-frequency resources. Because the antenna dimension makes the estimation problem severely underdetermined, the paper jointly learns pilot signals and channel estimation from channel structure rather than relying solely on conventional assumptions.

  • The FDD system uses a base station with N antennas and OFDM over M subcarriers to serve a single-antenna user.
  • Nearby subcarriers and antennas exhibit correlations arising from shared propagation paths, gains, and angular characteristics.
  • A pilot block of size L × M estimates a channel over a T × M time-frequency grid, with L ≤ T and a channel assumed constant across the grid.
  • For each subcarrier, the received pilot signal is modeled as ym = Pmhm + nm, where Pm contains the transmitted pilots and nm is independent complex Gaussian noise.
  • A large pilot length increases overhead and complexity, while massive antenna dimensions make the estimation equation severely underdetermined.
  • The proposed approach jointly designs frequency-aware pilots and estimates the downlink channel using a data-driven neural network that learns dataset statistics and reduces pilot overhead.

III. NN-BASED PILOT DESIGN AND CHANNEL ESTIMATION

The proposed end-to-end neural architecture combines frequency-specific fully connected branches, convolutional layers, and non-local attention for joint pilot design and channel estimation. Training learns the pilots and estimator from data, while pruning enables non-uniform pilot allocation and the reported gains are strongest with shorter pilots or lower SNR.

  • Fully connected branches jointly model pilot transmission and produce initial channel estimates for each subcarrier.
  • Convolutional layers exploit local inter-subcarrier and inter-antenna correlations to improve reconstruction accuracy beyond the initial estimates.
  • The non-local attention module targets longer-range channel correlations and further improves NMSE, particularly with fewer pilots and at lower downlink SNRs.
  • The dense layers are trained end-to-end with convolutional processing to jointly optimize pilot signals and channel estimation under an empirical MSE objective.
  • The method is data-driven: it learns channel statistics from training data instead of assuming prior channel statistics, providing a tractable empirical-MSE approximation to general MMSE estimation.
  • The neural estimator outperforms LMMSE, with larger improvements at shorter pilot lengths and lower SNR values.
  • Pruning less significant dense-layer neurons enables non-uniform pilot allocation across subcarriers and further reduces MSE at the same pilot overhead.

IV. PILOT ALLOCATION BY NN PRUNING

The paper reduces downlink pilot overhead by pruning least-significant neurons in fully connected pilot-design layers, producing non-uniform pilot allocation while limiting reconstruction-MSE degradation.

  • Pilot transmission consumes downlink resources, motivating a pruning-based reduction of pilot overhead.
  • The proposed method prunes neurons from fully connected reduction layers, with each neuron corresponding to a pilot signal.
  • Pruning enables non-uniform pilot allocation, sending fewer pilots on subcarriers whose channels can be reconstructed by later convolutional layers.
  • An l1-regularized magnitude-based approach pushes pilot-related weight measures toward zero before removing the least-significant neurons.
  • The pilot mask is initialized with all ones and updated through a linear schedule of 10 pruning steps, each removing S/10 of the pilots.
  • The saliency choice is heuristic, although the authors motivate magnitude-based pruning because low-weight pilot locations receive lower SNR and contribute least to NMSE.

V. SIMULATION RESULTS

The simulations use COST 2100 channel realizations for indoor and outdoor scenarios, with datasets and training settings specified for evaluating normalized MSE.

  • Training and testing use the COST 2100 geometry-based stochastic channel model, which reproduces MIMO channel statistics across time, frequency, and space.
  • The evaluation considers an indoor picocellular scenario at 5.3 GHz and an outdoor rural scenario at 330 MHz.
  • Networks are trained for 110000 steps using batch size 100 and Adam, with NMSE used as the performance measure.

A. NN vs. LMMSE performance

The NN-based estimator outperforms LMMSE across tested SNRs and pilot lengths, while attention provides additional gains at low SNR and can halve pilot overhead for a target NMSE.

  • The LMMSE comparison assumes users have perfect knowledge of covariance matrices estimated from the training dataset; estimating them from a coarse LS estimate degrades LMMSE performance.
  • The NN-based approach outperforms LMMSE for all tested SNR and pilot-length values.
  • The attention module further improves performance specifically at low SNR.
  • For a target NMSE of ∼−15dB at 5dB downlink SNR, NN+Attention reduces pilot length from 16 to 8, a 50% overhead reduction.
  • The fully connected subnetwork achieves almost the same NMSE as LMMSE, and its learned parameters are close to the LMMSE solution.

B. Impact of the scattering environment

The NN-based estimator improves reconstruction NMSE in both scattering environments, with a larger improvement over LMMSE in the outdoor scenario.

  • The NN-based approach achieves a more significant performance improvement over LMMSE in the outdoor scenario than in the indoor scenario.
  • At L = 8, N = 32, and M = 256, NN+Attention improves reconstruction NMSE over LMMSE by 5.49dB outdoors and 9.36dB indoors.

C. Impact of channel SNR

Performance depends on the relationship between training and test SNR, while frequency-aware and attention-enhanced pilot designs improve channel reconstruction and reduce pilot requirements.

  • Training–test SNR mismatch: NMSE degrades when test SNR falls below the network’s training SNR, while performance saturates when test SNR exceeds it.A network trained on SNRs uniformly sampled from [−5, 10]dB performs satisfactorily across the whole SNR range.
  • Frequency-aware pilot design: The frequency-aware approach improves performance by learning subcarrier-specific statistics and structures through different pilots across subcarriers.SP denotes same pilots, whereas DP denotes different pilots.
  • Frequency-aware pilot design: 33.3% reduction in pilot overhead is achieved by reducing pilot length from 12 to 8 at a target NMSE of −12.9dB and downlink SNR of 10dB.The comparison is between the proposed different-pilot approach and same pilots over all subcarriers.
  • Attention-enhanced design: At downlink SNR −5dB, DP+Attention reduces pilot length from 16 to 8 and improves reconstruction NMSE by 1.39dB versus SP.The comparison covers NMSE curves for SP, DP, and DP+Attention with L = 16 and L = 12.

E. Ablation study

The ablation study compares neural and conventional pilot-design and channel-estimation combinations. The fully neural scheme performs best, while attention captures longer-range correlations beyond convolutional processing.

  • Compared schemes: Four schemes compare neural pilot design and estimation, NN-designed pilots with LMMSE, FFT pilots with NN+Attention, and FFT pilots with LMMSE.The proposed configuration is denoted “NN+NN+Attention,” while the conventional baseline is “FFT+LMMSE.”
  • Ablation results: The fully NN-based scheme significantly outperforms all three alternatives in reconstruction NMSE.The NN-based pilot design is identified as producing a remarkable improvement.
  • Ablation results: The proposed attention approach exploits long-range channel correlations that remain after convolutional layers capture local correlations.These long-range correlations further improve reconstruction accuracy relative to simpler NN-based estimation.
  • LMMSE comparison: Extended LMMSE requires a covariance matrix of size NM × NM and becomes infeasible as both antenna and subcarrier dimensions grow.Practical comparisons therefore use subcarrier-wise LMMSE, which exploits inter-antenna correlations within each subcarrier.
  • Data dependence: With limited training data, extended LMMSE degrades significantly, whereas other approaches remain almost unchanged across smaller dataset sizes.With abundant data, extended LMMSE slightly outperforms the NN-based approach; the proposed NN+Attention method nevertheless remains very close while reducing parameter complexity by 134M.
  • Complexity and accuracy: The proposed NN+Attention estimator achieves reconstruction NMSE very close to extended LMMSE while reducing parameter complexity by a factor of 134M.This comparison is reported for L = 16, N = 32, and M = 256.

G. Computational complexity

The paper compares deployment complexity through trainable parameters and FLOPs. Convolutional and attention components maintain parameter counts independent of channel dimensions, while attention has computational cost comparable to extended LMMSE.

  • Training and deployment: Offline training for the proposed NN+attention estimator with L = 16, N = 32, and M = 256 took 2 hours through 110000 steps on an NVIDIA GEFORCE RTX 2080 Ti GPU.After training, the network is deployed for channel estimation.
  • FLOP model: The proposed NN estimator’s forward-pass FLOPs include fully connected expansion, convolutional, and attention operations.The corresponding orders are O(LNM), O(NMρ^2st), and O(N^2M^2q), respectively.
  • FLOP comparison: The NN+attention and extended LMMSE methods require the same order of FLOPs, whereas removing attention significantly reduces NN FLOPs.The attention module therefore preserves estimation capability at a computational order comparable to extended LMMSE.

H. Pilot pruning

The pruning scheme reduces pilot overhead by removing less significant dense-layer neurons and allocating pilots non-uniformly across subcarriers. It preserves channel-estimation and symbol-error performance under the evaluated settings.

  • Pruning results: The pruning scheme reduces pilot overhead while the NMSE degrades slowly as sparsity increases.This supports reducing pilot transmission without significantly degrading channel-estimation accuracy.
  • Overhead reduction: For L = 12, pruning saves up to 25% of time-frequency resources while degrading NMSE by only 1.07dB.The scheme allocates pilots non-uniformly along subcarriers rather than using a uniform rectangular or periodic pattern.
  • Overhead reduction: With the same pilot overhead, L = 12 and S = 50% achieves NMSE −15.35dB versus −19.43dB for L = 16 and S = 25%.These settings occupy the same number of time-frequency pilot resources.
  • Allocation comparison: Pruning 50% of a 12 × 256 pilot block improves NMSE by 2.23dB over a 6 × 256 block with the same pilot overhead.Starting from a larger pilot length and pruning to the available budget improves NMSE under fixed resources.
  • Practical trade-off: Using L = 12 slightly increases receiver estimation delay because channel estimation waits for all 12 resource blocks.This is the stated practical trade-off of the larger starting pilot length.
  • Non-uniform allocation: The NN allocates fewer pilots to subcarriers that can be accurately interpolated and more pilots to other subcarriers using statistical CSI correlations.For the evaluated setting, more resources are saved for data transmission in subcarrier ranges (30-110) and (150-230).
  • SER performance: The pruning scheme improves SER while using the same number of pilot resources for all S ≤50% with L = 8 and L = 12.The evaluation uses 4-QAM, zero-forcing equalization, and maximum-likelihood detection at SNR = 10dB.

VI. CONCLUSION

The paper combines frequency-aware dense layers, convolutional layers, and attention to design pilots and estimate downlink channels in FDD massive MIMO-OFDM. NN pruning reduces pilot overhead through non-uniform subcarrier allocation while preserving efficient channel reconstruction.

  • The proposed network uses dense layers for frequency-aware pilot design and convolutional layers to exploit inherent channel-matrix correlations for channel estimation.
  • Figure 9 compares periodic pilot removal with NN-optimized pruning for L = 8, S = 25% and L = 12, S = 50%, marking saved data-transmission resources in black.
  • An attention module exploits long-range correlations in the channel matrix beyond the local correlations captured by conventional convolutional layers.
  • NN pruning gradually removes less significant neurons from dense layers to reduce pilot overhead and save time-frequency resources for data transmission.
  • The proposed NN-based pilot design and channel estimation scheme outperforms LMMSE estimation.
  • The pruning-based technique allocates pilots non-uniformly, allowing fewer pilot transmissions on subcarriers that subsequent convolutional layers can reconstruct satisfactorily.
Loading 2006.11796v3…