Source-linked AI summary
Bayes-Optimal Joint Channel-and-Data Estimation for Massive MIMO with Low-Precision ADCs
Chao-Kai Wen, Chang-Jen Wang, Shi Jin, Kai-Kit Wong, Pangan Ting
TL;DR
Very low-precision ADCs reduce massive-MIMO hardware cost and power but make channel acquisition difficult and can require long training. The paper develops Bayes-optimal JCD estimation with a GAMP-based implementation, derives large-system performance through scalar AWGN decoupling, and validates the analysis by simulation. The results support efficient performance evaluation and identify system-design observations, while the practical algorithm may remain too computationally expensive for commercial systems.
Problem
Low-precision ADCs reduce hardware cost and power but make channel-state acquisition difficult, while JCD performance in quantized MIMO systems was not clearly understood.
Method
The paper uses Bayes-optimal JCD inference, implements it with a modified BiG-AMP algorithm, and analyzes its large-system behavior using the replica method.
Results
The analysis characterizes channel and data MSEs and SERs through scalar AWGN channels, and Monte-Carlo simulations confirm the analytical results.
Takeaways & Limitations
The analytical results enable quick evaluation of quantized MIMO performance and provide observations relevant to system design.
Takeaways & Limitations
The GAMP-based JCD algorithm may still have computational complexity too high for affordable commercial systems.
Abstract
from arXiv · showhide
This paper considers a multiple-input multiple-output (MIMO) receiver with very low-precision analog-to-digital convertors (ADCs) with the goal of developing massive MIMO antenna systems that require minimal cost and power. Previous studies demonstrated that the training duration should be {\em relatively long} to obtain acceptable channel state information. To address this requirement, we adopt a joint channel-and-data (JCD) estimation method based on Bayes-optimal inference. This method yields minimal mean square errors with respect to the channels and payload data. We develop a Bayes-optimal JCD estimator using a recent technique based on approximate message passing. We then present an analytical framework to study the theoretical performance of the estimator in the large-system limit. Simulation results confirm our analytical results, which allow the efficient evaluation of the performance of quantized massive MIMO systems and provide insights into effective system design.
I. INTRODUCTION
Massive MIMO can improve communication capacity but creates substantial ADC cost and power challenges. This work studies joint channel-and-data estimation with very low-precision ADCs to address difficult CSI acquisition and long training requirements.
- Massive MIMO uses hundreds or thousands of base-station antennas to serve many users over shared time-frequency resources.
- ADC hardware complexity and power consumption increase exponentially with bits per sample, motivating 1–3-bit ADCs for quantized MIMO systems.
- Coarse quantization reduces effective measurements and makes CSI acquisition harder than in unquantized systems.
- One-bit quantized MIMO may require training sequences approximately 50 times the number of users to achieve sufficient channel information.
- The paper proposes Bayes-optimal JCD estimation because it yields minimum mean square errors for channels and data symbols.
- A modified BiG-AMP method, called the GAMP-based JCD algorithm, approximates the Bayes-optimal estimator for quantized MIMO systems.
III. BAYES-OPTIMAL JCD ESTIMATION
The Bayes-optimal JCD estimator computes posterior means for channels and data from quantized observations and known pilots. Its SISO interpretation connects quantized posterior estimation with scalar AWGN-based updates.
- JCD estimates both the channel H and payload data Xd from all quantized observations eY given the known pilot matrix Xt.
- Bayesian posterior-mean estimates of H and X minimize the corresponding Bayesian mean square errors.
- For known pilots, the pilot estimate is exact and its mean square error is zero.
- The estimator treats quantized observations through posterior mean and variance calculations, with separate real and imaginary processing for complex signals.
- The SISO quantized-channel expressions specialize to one-bit quantization and converge to the unquantized-channel expressions as B →∞ and ∆→0.
- In the joint unknown-product case Z = HX, direct posterior computation is intractable because the required marginalization involves high-dimensional integrals.
B. GAMP-Based JCD Algorithm
The GAMP-based JCD algorithm makes Bayes-optimal joint estimation tractable by approximating high-dimensional message passing for the bilinear channel-data model.
- Direct marginal posterior calculations are intractable because they require high-dimensional integrals over channel and data variables.
- The factorization of the posterior motivates message passing between factor nodes representing quantized observations and variable nodes representing H and X.
- AMP and GAMP replace infeasible sum-product computations with tractable approximations for marginal posterior estimation.
- The proposed algorithm adapts BiG-AMP to quantized observations and known pilot symbols.
- Intermediate estimates of Zt = HXt and Zd = HXd provide auxiliary variables and associated variances during the iterations.
- The algorithm computes posterior moments, residual terms, and scalar-channel observations before updating estimates of the data and channel.
C. Nonlinear Steps
The nonlinear steps transform quantized observations into posterior updates for the product, data, and channel variables. These updates use Gaussian approximations and the appropriate signal priors.
- At each iteration, estimates of Znt, Xkt, and Hnk separately act as estimators over a bank of scalar channels.
- The Z update uses the quantized likelihood and a Gaussian prior to compute posterior means and variances.
- The H and X updates replace the quantized output distribution with a Gaussian distribution, making their updates equivalent to AWGN-channel estimation.
- The data update uses a square QAM constellation and its uniform prior when the payload symbols are uniformly distributed.
- The nonlinear expressions implement the GAMP-based JCD algorithm and were evaluated using the open-source GAMPmatlab software suite.
IV. PERFORMANCE ANALYSIS
The paper analyzes Bayes-optimal JCD estimation through average free entropy in the large-system limit. The replica method facilitates this analysis despite not being mathematically rigorous.
- Performance-analysis framework: Asymptotic MSEs for data and channels are obtained as saddle points of the average free entropy.The analysis reduces performance characterization to finding the average free entropy.
- Large-system limit: The large-system limit sends N, K, and T to infinity while keeping α, β, βt, and βd fixed and finite.The paper denotes this regime simply by K →∞.
- Analytical challenge: Computing the expected logarithm of the marginal likelihood is the major difficulty in evaluating the free entropy.The paper rewrites the free-entropy expression to address this expectation.
- Replica method: The replica method moves the expectation inside the logarithm by first evaluating integer moments and then extending the result to positive real orders.The method comes from statistical physics and has been successful in difficult statistical-physics and information-theory problems, although it is not mathematically rigorous.
A. Parameters of Proposition 1
The analytical framework represents data and channel estimation through scalar AWGN channels, whose parameters determine asymptotic MSEs and error rates. It then specializes these relations to perfect CSIR, pilot-only estimation, and quantized-system design trade-offs.
- Scalar-channel representation: The asymptotic data and channel MSEs are associated with scalar AWGN channels whose equivalent SNRs are ˜qXd and ˜qH.Posterior means and MSEs for Xd and H are computed from these scalar channels.
- Pilot phase: For the known pilot phase, mseXt = 0 and the pilot-phase mutual information is zero because the pilot matrix is known.The corresponding pilot-phase scalar channel is therefore omitted from subsequent performance discussions.
- Decoupling principle: In the large-system limit, Bayes-optimal JCD decouples the quantized MIMO input-output relationship into scalar AWGN channels for both data symbols and channel responses.This extends the decoupling principle to joint estimation of data and channel response.
- Special cases: With perfect CSIR, the training phase is unnecessary, whereas the pilot-only scheme estimates H from pilot observations before detecting payload data.The pilot-only analysis exchanges the roles of H and Xt relative to the perfect-CSIR case.
- JCD gain: In the pilot-only scheme, the second term in ˜qH for JCD is the gain from data-aided channel estimation.Reducing mseH increases qH and therefore increases the effective data-channel SNR ˜qXd = αqHχd.
- Training and quantization trade-offs: For high SNR and βt ≫1, doubling training length increases pilot-only mseH by 6 dB, equivalent to losing one ADC bit.The paper uses this relation to frame a trade-off between training length and ADC word length.
- Training and quantization trade-offs: At a target mseH = −30dB, a 4-bit receiver requires βt = 4, while reducing to 1 bit increases the required training length eightfold.The paper presents this comparison as motivation for using JCD to exploit payload data during channel estimation.
A. Accuracy of the Analytical Results
Simulations largely validate the analytical performance predictions for Bayes-optimal JCD estimation, while revealing step-size, CSIR, convergence, and implementation boundaries.
- Simulation validation: 10,000 channel realizations were used to compare simulated SER and channel-MSE results with the analytical expressions.The GAMP-based JCD algorithm used tolerance ϵ = 10−8 and a maximum of 100 iterations.
- Simulation validation: 3-bit quantization incurs a 1.29 dB loss and 2-bit quantization a 2.95 dB loss relative to the unquantized target at SER = 10−3.These losses are reported for the Bayes-optimal JCD estimator.
- Optimal step-size: The Bayes-optimal estimator’s SER-optimal step size differs substantially from the Gaussian-distortion-optimal value.For B = 2, the Gaussian-distortion-optimal step size is approximately 0.7128, whereas SER optimization produces a different value.
- Optimal step-size: The optimal step size varies slightly across input constellations and system settings but decreases as SNR increases, indicating that SNR is its main determinant.The examined inputs include QPSK, 16QAM, 64QAM, and Gaussian signals.
- Effects due to the Absence of CSIR: The Bayes-optimal JCD estimator substantially improves over pilot-only estimation in the 1-bit and unquantized cases, while the no-CSIR gap is larger for 1-bit quantization.Increasing training length can reduce the 1-bit gap, but the improvement remains limited under practical coherence-time constraints.
- System-design implications: Pilot-only channel-estimation MSE decreases by approximately 6 dB for each added ADC bit or each doubling of training length.The conclusion identifies this scaling as support for JCD, particularly in quantized MIMO systems.
APPENDIX A: DERIVATIONS OF (23) AND (24)
The appendix derives the posterior mean and variance estimators in (23) and (24) by evaluating quantized-output likelihood integrals and differentiating their Gaussian-weighted expressions. The replica calculation then applies replica symmetry to obtain fixed-point equations and scalar AWGN-channel representations.
- Derivation of (23) and (24): The derivation starts from the quantized likelihood for an output bin and evaluates the denominator of the posterior estimator in (17).For b=1, the lower boundary is r0=−∞, eliminating the second term in the likelihood expression.
- Derivation of (23) and (24): Differentiating the Gaussian-weighted likelihood integral with respect to the posterior mean parameter yields the expression used to obtain (23).The marginal posterior mean follows after multiplying both sides by 1/C.
- Derivation of (23) and (24): The posterior variance expression (24) is obtained by differentiating the corresponding expression twice and substituting the result and (23) into (67).The derivation treats eY≤0 explicitly and extends to eY>0 analogously.
- Replica calculation: Under the replica-symmetric assumption, the saddle-point optimization reduces to scalar parameters associated with channel and data overlaps.The auxiliary Gaussian variables are represented through replica-symmetric covariance forms.
- Replica calculation: Setting derivatives of the replica-symmetric free entropy to zero produces the fixed-point equations and connects them to the scalar AWGN channels in (36).At τ→0, the derivation gives ˜cH=0, ˜cXo=0, cH=E{|H|2}, and cXo=E{|Xo|2}.
APPENDIX C: PROOF OF PROPOSITION 2
The proof establishes the joint-moment characterization required for Proposition 2 by introducing replica variables, constructing a generalized free entropy, and showing that its derivatives recover the moments of the true and estimated channel and data entries.
- Joint-moment characterization: The proof targets convergence of joint moments for selected entries of the channel, data, and their estimates.The limiting joint distribution is considered for (Hnk, Xd,kt, bHnk, bXd,kt).
- Joint-moment characterization: A generalized free entropy is introduced so that derivatives with respect to auxiliary parameters generate the desired joint moments.The construction inserts exponential perturbations involving functions of the channel and data variables.
- Conclusion: The derivatives of the generalized free entropy exactly provide the joint moments of H, Xd, bH, and bXd.Consequently, the joint moments of interest are uniquely determined by (86) through the Carleman theorem.
APPENDIX D: PROOF OF PROPOSITION 3
The appendix analyzes the high-SNR behavior of the channel MSE as the training-to-user ratio grows. A quantizer-dependent constant determines the resulting asymptotic scaling.
- High-SNR asymptotics: The derivation considers the limiting case of infinite SNR.The noise variance is taken to the high-SNR limit before evaluating the asymptotic behavior.
- High-SNR asymptotics: As βt→∞, the effective channel parameter ˜qH diverges, and a Taylor expansion gives 1−mseH≈1−1/˜qH.This yields the corresponding asymptotic approximation for mseH.
- High-SNR asymptotics: The channel MSE scales asymptotically as mseH≈(βtcB)^−2, where cB depends on the quantizer.The equivalent dB expression uses CB=−20 log10(cB), with values obtained numerically.
APPENDIX E: A GENERALIZATION OF PROPOSITION 1
The appendix extends Proposition 1 to users with different large-scale fading factors by defining user-specific scalar AWGN channels and corresponding fixed-point parameters and MSEs.
- Generalized model: The generalization addresses users having different large-scale fading factors σ2hk.The proof follows the same steps as Appendix A, with details omitted.
- Scalar AWGN channels: User-specific scalar AWGN channels are defined for the channel and data variables in the generalized setting.The effective noises are standard complex Gaussian, while Hk follows PHk≡NC(0,σ2hk).
- Fixed-point characterization: The analysis averages user-dependent quantities using ⟨ak⟩K=1/K∑k=1K ak.This average is applied over the set of user-indexed values.
- Fixed-point characterization: The generalized fixed-point equations determine qXo,k, qHk, ˜qXo,k, and ˜qHk, which characterize the user-specific MSEs and mutual informations.The MSEs mseHk and mseXd,k correspond to the channel and data scalar channels, while mseXt,k=0.