Source-linked AI summary
Achievable Information Rates for Fiber Optics: Applications and Computations
Alex Alvarado, Tobias Fehenberger, Bin Chen, Frans M. J. Willems
TL;DR
Fiber-optical communication needs reliable, FEC-relevant metrics and practical AIR computation for coded-system design. The paper reviews MI and GMI, compares their predictions with coded systems, and develops computation methods and approximations for multidimensional AWGN channels. It concludes that AIRs are versatile design metrics, while their predictive use is limited by realistic FEC decoding and modeling assumptions.
Problem
Fiber-optical system design requires information-theoretic performance metrics, but realistic coded systems and nonlinear channels complicate AIR evaluation and prediction.
Method
The paper reviews MI and GMI for coded modulation, compares AIR predictions with LDPC and polar codes, and presents numerical integration, Gauss-Hermite quadrature, and ready-to-use approximations.
Results
1.53 dB is the high-SNR gap between the MI envelope and multidimensional AWGN capacity for the reported high-cardinality QAM results.
Takeaways & Limitations
AIRs are versatile design metrics for comparing coded optical systems and guiding modulation, DSP, decoding, and nonlinear-compensation choices.
Takeaways & Limitations
MI and GMI assume capacity-achieving FEC with ideal maximum-likelihood decoding and do not capture error floors from practical codes and suboptimal decoding algorithms.
Abstract
from arXiv · showhide
In this paper, achievable information rates (AIR) for fiber optical communications are discussed. It is shown that AIRs such as the mutual information and generalized mutual information are good design metrics for coded optical systems. The theoretical predictions of AIRs are compared to the performance of modern codes including low-parity density check (LDPC) and polar codes. Two different computation methods for these AIRs are also discussed: Monte-Carlo integration and Gauss-Hermite quadrature. Closed-form ready-to-use approximations for such computations are provided for arbitrary constellations and the multidimensional AWGN channel. The computation of AIRs in optical experiments and simulations is also discussed.
I. INTRODUCTION AND MOTIVATION
Fiber-optical systems need bandwidth-efficient coded modulation as traffic growth strains available capacity. The paper presents AIRs, especially MI and GMI, as FEC-related metrics for comparing coded systems, modulation, DSP, and transmission techniques.
- Motivation: Fiber capacity constraints motivate information-theoretic analysis and bandwidth-efficient coded modulation with forward error correction.The paper describes trellis-coded modulation, BICM, and MLC as key coded-modulation approaches.
- AIRs as design metrics: AIRs measure reliably transmissible information bits per symbol and are inherently related to FEC performance.The paper contrasts AIRs with pre-FEC BER, SER, Q-factor, and EVM as performance metrics.
- Applications: AIRs support fair comparisons of constellations, DSP, decoding, nonlinear-compensation techniques, and post-FEC error prediction.The paper uses AIRs to compare coded-system throughput and performance predictions.
- AIR selection: MI is the relevant AIR for nonbinary coded modulation and MLC with multistage decoding, whereas GMI applies to binary FEC and MLC with parallel decoding.The paper considers equally likely symbols and notes that GMI can be extended to nonuniform symbols.
- Code comparisons: For polar-coded systems, MI predicts MLC-MSD throughput and GMI predicts BICM performance, with finite-length gaps caused by code suboptimality.The cited comparison uses polar codes with finite block length and targets post-FEC bit error probability below 10^-4.
- Optical-system design: For optical transmission, AIR comparisons predict reach differences and show that binary LDPC codes follow GMI rather than MI when modulation formats differ.Single-channel DBP with 64QAM offers an approximately 1100 km reach increase over EDC at 8 bit/sym in the cited simulation.
- Post-FEC prediction: Normalized MI and GMI provide decoding thresholds for SD-FEC, with normalized MI predicting NB-LDPC post-FEC SER and normalized GMI predicting binary-LDPC post-FEC BER.The paper reports these predictions across the cited NB-LDPC and binary-LDPC examples.
B. Limitations
AIR-based predictions are limited by asymptotic coding assumptions, decoder suboptimality, interleaver assumptions, and mismatched or approximate LLRs.
- Coding and decoding assumptions: AIRs assume capacity-achieving FEC with ideal maximum-likelihood decoding, whereas practical codes use suboptimal low-complexity algorithms.For LDPC codes, the sum-product algorithm is given as an example of practical suboptimal decoding.
- Finite-length effects: Finite codeword lengths create coding gaps between AIR predictions and practical polar or LDPC performance.The paper suggests density evolution or finite-blocklength bounds when finite-length constraints matter.
- Finite-length effects: Polar-code successive-cancellation decoding can be capacity-achieving yet perform poorly at short and moderate block lengths.The gap can decrease with longer codes or list decoding.
- Interleaving assumptions: The concatenated FEC analysis assumes ideal interleavers, while the reported LDPC results use a new random interleaver for every transmitted codeword.This makes the prediction more accurate but less practically relevant than fixed or absent interleaving.
- Information-metric assumptions: Mismatched or approximate LLRs, including max-log computation, can produce lower AIRs and motivate LLR correction.The paper connects LLR correction with improving decoder performance.
III. INFORMATION-THEORETIC ELEMENTS AND AIRS
The paper frames coded optical communication using multidimensional symbols, channel models, and AIRs tailored to different soft-decision coded-modulation structures. It defines the coding and channel concepts needed to interpret MI and GMI analyses.
- Coded modulation: The paper focuses on MLC and BICM with soft-decision FEC, while distinguishing hard-decision and soft-decision information passed to the decoder.These structures are treated as practically relevant coded-modulation formats.
- Channel model: The optical channel may retain intersymbol and interpolarization interference, but the information-theoretic model uses a memoryless demapper that ignores residual temporal correlation.The resulting information-theoretic channel is interpreted as an average channel.
- Signal model: Symbols are uniformly drawn multidimensional constellation points with N complex dimensions and cardinality M = 2^m.Received symbols have the same multidimensional representation, and symbol differences are defined as vectors.
- AIR selection: The paper considers four soft-decision coded-modulation structures: NB-CM, MLC-MSD, BICM, and MLC-PDL.MI is the AIR for NB-CM and MLC-MSD, whereas GMI is the AIR for BICM and MLC-PDL.
- AIR foundations: An achievable rate is a rate supported at a given block length and error probability, while channel capacity is the largest AIR attainable with vanishing error as block length grows.The paper subsequently focuses on MI and GMI as AIRs.
C. Mutual Information
Mutual information characterizes the reliable rate of symbol-wise decoding on memoryless channels and also serves as an AIR for NB-CM and MLC-MSD.
- The symbol-wise maximum-likelihood receiver selects codewords using the channel law for the observed symbol sequence.The transmitted bits are mapped to symbols, and decoding operates on the resulting sequence of channel observations.
- Reliable transmission with symbol-wise decoding is possible when the coded-modulation rate does not exceed the mutual information I(X; Y).The channel-coding theorem identifies MI as the largest achievable rate for a memoryless channel under this decoding model.
- The MI for discrete constellations and equally likely symbols is expressed through an expectation or multidimensional integral over channel observations.The integral form uses the conditional channel density and applies to arbitrary multidimensional memoryless channels.
- MI is an achievable information rate for both nonbinary coded modulation and multilevel coding with multistage decoding.The MLC-MSD interpretation follows from the chain rule of mutual information.
- Polar-coded MLC-MSD code rates are designed to match the bit-wise conditional mutual informations.This rate matching is reported in the paper’s implementation discussion and Table II.
D. Generalized Mutual Information
Generalized mutual information provides an achievable rate for bit-wise receivers whose decoding metrics do not use the full symbol-wise channel dependence. Its interpretation relies on independent bits and bit-wise metrics, with known scope limitations.
- Bit-wise decoding first computes L-values and then applies one or more binary soft-decision decoders, unlike symbol-wise maximum-likelihood decoding.The bit-wise rule can be represented as mismatched decoding because its metric is not generally matched to the symbol-wise channel.
- The GMI is an achievable rate for bit-wise receivers such as BICM and MLC-PDL.These receivers ignore dependencies between the bits within a symbol, making MI generally unsuitable as their achievable rate.
- The GMI expression is general for arbitrary decoding metrics and symbol distributions, but the simplified form assumes independent bits and bit-wise receiver metrics.The independence assumption holds for uniform symbol distributions considered in the paper.
- Under the stated assumptions, GMI can be expressed as a sum of unconditional bit-wise mutual informations.These bit-wise mutual informations can be used to select code rates in MLC-PDL.
- The GMI has not been proven to be the largest achievable rate for BICM, and the largest achievable rate with a bit-wise decoder remains open.The paper nevertheless states that GMI predicts the performance of coded modulation with capacity-approaching soft-decision FEC decoders.
E. MI and GMI for AWGN Channel
For the multidimensional complex AWGN channel, the paper develops general MI and GMI expressions and compares them for QAM under Gray labeling. At high SNR, GMI closely approaches MI.
- The multidimensional AWGN model uses independent complex Gaussian noise components with total variance σ_z^2.The noise variance per complex dimension is σ_z^2/N, while transmitted power determines the SNR.
- The paper presents general expressions for MI and GMI for the multidimensional AWGN channel.These expressions apply to the channel model with discrete multidimensional input constellations.
- The rate penalty I − G caused by the suboptimal bit-wise decoder is known to be small for Gray-labeled constellations.The paper states the general inequality I ≥ G and connects the penalty to bit-wise decoding.
- For Gray-labeled QAM, GMI is very close to MI at high SNR.Figure 8 plots MI, GMI, and AWGN capacity versus SNR for the complex AWGN channel.
- The AWGN MI and GMI curves are calculated using the numerical integration method described in the paper.The example uses MQAM constellations and the BRGC labeling rule for GMI.
F. LLR-based GMI
The paper relates GMI to the L-values used by bit-wise decoders and contrasts exact LLR computation with the max-log approximation. Approximate or mismatched LLRs can reduce achievable rate.
- Exact L-values are computed from bit-conditioned constellation likelihoods, while the max-log approximation reduces their computational complexity.For square QAM over AWGN, max-log L-values become piecewise linear in the received symbol, simplifying implementation.
- GMI can be interpreted as a sum of bit-wise mutual informations between code bits and their L-values.This interpretation applies for equally likely symbols and general LLR calculations or channels.
- When L-values are calculated exactly from the channel model, I(B_k; Y) equals I(B_k; L_k).Under other L-value approximations, this equality does not generally hold.
- Max-log L-values can incur an achievable-rate loss, although correction strategies can recover the loss under certain conditions.LLR scaling and other correction methods can also improve decoding performance in mismatched scenarios.
- LLRs mismatched to the channel or computed with an approximation result in rate loss.The paper describes LLR correction strategies as a way to improve achievable rate and FEC-decoder performance.
IV. COMPUTATION METHODS FOR AIRS
The paper presents Monte-Carlo methods for approximating MI and GMI, including multidimensional AWGN formulas and experimental estimation procedures. It also identifies when these methods are preferable and notes limitations for circularly symmetric noise models and mismatched LLRs.
- Monte-Carlo Integration: Monte-Carlo integration approximates MI and GMI for AWGN channels, while nonlinear-channel AIR computation is addressed separately.The method replaces integrals with finite sums and extends to multidimensional AWGN channels.
- Monte-Carlo Integration: The multidimensional AWGN approximations apply to any multidimensional constellation but assume circularly symmetric Gaussian noise.Correlated-noise generalizations and experiments without the circular-symmetry assumption are cited separately.
- Monte-Carlo Integration: Using D = 10^4 samples, MI was computed for 227 two-dimensional complex constellations in a few hours on a standard computer.The resulting curves support comparisons among modulation formats with different constellation sizes.
- Monte-Carlo Integration: Experimental MI estimation uses noise-variance estimation, transmitted-symbol subtraction, and inner-sum evaluation from received samples.For non-AWGN optical channels, this produces an AIR under a mismatched metric.
- Monte-Carlo Integration: For GMI with max-log or mismatched LLRs, minimizing over s is mandatory for an information-theoretically precise rate calculation.Using the simplified expression without this minimization yields a rate lower than the true one.
B. Gauss-Hermite Quadrature
Gauss-Hermite quadrature provides deterministic approximations of MI and GMI for multidimensional AWGN channels. It is especially effective for low-dimensional constellations, where it can be faster than Monte-Carlo integration.
- Gauss-Hermite Quadrature: The quadrature order J controls the trade-off between computation speed and accuracy, and the approximation becomes exact as J →∞.Nodes and weights are numerically available for different J values.
- Gauss-Hermite Quadrature: After digital signal processing, experimental AIR estimation uses samples whose conditional means match the transmitted symbols.The processing includes filtering, equalization, synchronization, matched filtering, and sampling.
- Gauss-Hermite Quadrature: Gauss-Hermite quadrature approximates MI and GMI for multidimensional AWGN channels using multidimensional quadrature nodes.The method generalizes the one-dimensional rule to complex multidimensional integrals.
- Gauss-Hermite Quadrature: Gauss-Hermite quadrature is generally faster for real and two-dimensional complex constellations, whereas Monte-Carlo suits higher dimensions or unknown channel laws.Quadrature also avoids randomness from sampled integration.
- Gauss-Hermite Quadrature: For MQAM constellations up to M ≤ 65536 = 2^16, J = 10 quadrature computes MI curves reaching the ultimate shaping-gain gap at high SNR.These constellations support a maximum spectral efficiency of 32 bit/sym.
V. MI AND GMI FOR THE NONLINEAR OPTICAL CHANNEL
The nonlinear optical fiber channel has memory, so MI and GMI from memoryless models can be loose lower bounds on achievable rates. The paper relates channel capacity to multidimensional mutual information while identifying a bandwidth-related limitation.
- Nonlinear Optical Channel: The nonlinear optical channel has memory, making the previously discussed MI and GMI potentially loose lower bounds on achievable rates.The channel law is represented by a conditional PDF over transmitted and received sequences.
- Nonlinear Optical Channel: Under information-stability assumptions, channel capacity is defined through multidimensional MI maximized over input distributions satisfying a power constraint.The capacity expression is based on a discrete-time channel model.
- Nonlinear Optical Channel: The discrete-time capacity expression has units of bit per symbol or bit per channel use and does not directly determine capacity in bit/s/Hz.It does not account for bandwidth or possible bandwidth expansion at high powers.
- Nonlinear Optical Channel: For i.i.d. symbols and a suboptimal receiver that ignores residual memory after DSP, memoryless MI is the largest achievable rate for that receiver.This rate also lower-bounds the capacity of the channel with memory.
B. Lower Bounds on the MI and GMI
When the nonlinear channel law is analytically unknown, the paper derives MI and GMI lower bounds using auxiliary channel models and estimates them with Monte-Carlo samples. The Gaussian auxiliary model is presented as a practical approximation, with stated scope and future extensions.
- Lower Bounds on the MI and GMI: When the nonlinear channel law is unknown analytically, MI can be lower-bounded using an arbitrary auxiliary conditional PDF.The bound is tightest when the auxiliary PDF is close to the true channel law and becomes equality when they are identical.
- Lower Bounds on the MI and GMI: A circularly symmetric Gaussian auxiliary channel is a common practical choice, although it may be suboptimal without a better non-Gaussian model.The model uses a Gaussian noise variance parameter.
- Lower Bounds on the MI and GMI: Analogous auxiliary-channel lower bounds are formulated for GMI and approximated through Monte-Carlo integration.The estimates converge to the corresponding bounds as the number of conditional samples increases.
- Lower Bounds on the MI and GMI: The Gaussian noise assumption is experimentally supported for long-haul dispersion-unmanaged optical systems.This supports its use as an approximation in that setting.
- Lower Bounds on the MI and GMI: The methodology can be generalized to non-AWGN channels and channels with memory, while finite-blocklength and code-universality analyses remain future work.The paper specifically suggests designing multidimensional constellations and coded modulation for nonlinear channels as a future application.
APPENDIX A PROOF OF THEOREM 1
The appendix derives the stated expressions by substituting earlier equations, expanding complex-vector norms, and exploiting circularly symmetric Gaussian statistics. It then uses these expressions to obtain the MI and GMI formulas and complete the proof.
- Proof derivation: Earlier equations are substituted into the theorem’s expressions to derive intermediate formulas for the relevant quantities.The derivation uses (19) in (7), (18), and (17).
- Proof derivation: The norm expansion follows from the inner-product identity for complex vectors, with complex conjugation handled through the Hermitian-transpose representation.The appendix invokes ∥z∥2 = zHz and the corresponding expansion of ∥z1+z2∥2.
- Proof derivation: Integration is changed from y to zi, after which the zero-mean circularly symmetric Gaussian statistics allow the complex conjugate and index i to be dropped.The resulting statistics are independent of i, yielding the stated expression and completing that proof step.
- Proof conclusion: The MI and GMI expressions are obtained in (21) and (22), and the final proof applies (38) to equations (63) and (64).The appendix explicitly identifies these equations as the final steps of the derivation.