Source-linked AI summary

Information Theoretic Proofs of Entropy Power Inequalities

Olivier Rioul

arXiv:0704.1751v2cs.IT

TL;DR

The paper addresses the gap that Shannon’s entropy power inequality is not derived from basic entropy or mutual-information properties in existing information-theoretic proofs. It unifies prior proof ingredients and gives a mutual-information proof, then extends the approach to several EPI variants.

  • Problem

    Existing information-theoretic proofs of the EPI rely on Fisher information or MMSE representations derived from de Bruijn’s identity rather than only basic entropy or mutual-information properties.

  • Method

    The paper unifies prior proofs around data processing applied to a covariance-preserving linear transformation and integration along a continuous Gaussian-perturbation path.

  • Results

    The paper gives a brief EPI proof through a mutual-information inequality that relies only on basic mutual-information properties, replacing earlier Fisher-information and MMSE inequalities.

  • Takeaways & Limitations

    The same ideas are generalized to linear transformations, dependent variables, covariance constraints, and Costa’s concavity inequality for entropy power.

Abstract

from arXiv · show

While most useful information theoretic inequalities can be deduced from the basic properties of entropy or mutual information, up to now Shannon's entropy power inequality (EPI) is an exception: Existing information theoretic proofs of the EPI hinge on representations of differential entropy using either Fisher information or minimum mean-square error (MMSE), which are derived from de Bruijn's identity. In this paper, we first present an unified view of these proofs, showing that they share two essential ingredients: 1) a data processing argument applied to a covariance-preserving linear transformation; 2) an integration over a path of a continuous Gaussian perturbation. Using these ingredients, we develop a new and brief proof of the EPI through a mutual information inequality, which replaces Stam and Blachman's Fisher information inequality (FII) and an inequality for MMSE by Guo, Shamai and Verdú used in earlier proofs. The result has the advantage of being very simple in that it relies only on the basic properties of mutual information. These ideas are then generalized to various extended versions of the EPI: Zamir and Feder's generalized EPI for linear transformations of the random variables, Takano and Johnson's EPI for dependent variables, Liu and Viswanath's covariance-constrained EPI, and Costa's concavity inequality for the entropy power.

CNRS LTCI

The paper addresses the exceptional status of Shannon’s entropy power inequality by replacing conventional Fisher-information and MMSE-based arguments with a simpler mutual-information proof. It unifies earlier proofs around data processing and continuous Gaussian perturbations, then extends the approach to several generalized EPIs.

  • Earlier proofs: Earlier information-theoretic proofs rely on Fisher information or MMSE, typically through de Bruijn’s identity or related integral representations.These approaches include Stam and Blachman’s Fisher information inequality and Guo, Shamai, and Verdú’s MMSE representation.
  • Motivation: The EPI is an important information-theoretic inequality used in channel-capacity and rate-distortion converses, but it has not followed directly from basic entropy or mutual-information properties.Its applications include additive-noise channels and several Gaussian coding problems.
  • New proof: The new EPI proof uses only elementary mutual-information properties and requires neither de Bruijn’s identity nor Fisher information or MMSE.It is presented as comparatively simpler and shorter than conventional arguments.
  • Consequences: The approach handles vector variables as easily as scalar variables and yields a mutual-information inequality with independent interest.The paper also gives a simple generalized de Bruijn identity based on the relationship between relative entropy and Fisher information.
  • Unified view: The paper unifies existing proofs through two ingredients: data processing inequalities and integration along a path of continuous Gaussian perturbations.The new proof uses the same ingredients in a more expedient form.
  • Extensions: The ideas extend to linear transformations, dependent variables, covariance-constrained formulations, and Costa’s entropy-power concavity inequality.The paper states that these extensions retain reliance on basic entropy and mutual-information properties, with further generalizations in some cases.

B. Definition of the differential entropy

Differential entropy is defined through an integral that may fail to exist in the generalized sense, so the paper establishes conditions ensuring it is well defined. The section also relates entropy power to Gaussian vectors with matching entropy or covariance properties.

  • Differential entropy is defined by the integral −∫p(x) log p(x) dx when its positive and negative parts are not both infinite.
  • A density can make differential entropy undefined because both positive and negative parts of the defining integral may be infinite.The paper gives a heavy-tailed density example with this behavior.
  • If E log+ 1/p(X) is finite, then h(X) is well defined and satisfies −∞ ≤ h(X) < +∞.Finite first or second moments are identified as sufficient special cases.
  • When X lacks a Lebesgue density, the paper sometimes extends the definition by setting h(X) = −∞.This includes distributions assigning positive mass to singletons.
  • Entropy power is characterized through Gaussian vectors with matching entropy, while its covariance bounds distinguish Gaussian and white cases.The first bound has equality iff X is Gaussian; the second has equality iff X is white.

II. EARLIER PROOFS REVISITED

The section develops Fisher-information and MMSE data-processing inequalities, then relates them through Gaussian-noise estimation and least-squares estimation.

  • Fisher information is defined from the score, the log-derivative of the density, while the score is linear exactly for Gaussian variables.
  • The paper assumes sufficiently smooth densities with sufficient decay so Fisher informations exist, possibly with value J(X) = +∞.
  • For a Markov chain θ → X → Y, data processing states that transforming X into Y cannot improve estimation of θ.
  • Equality in the estimation data-processing inequalities occurs when the optimal estimators based on X and Y coincide, equivalently when Y is sufficient relative to X.
  • The complementary Fisher-information/MMSE relation applies when independent Gaussian noise is added, with a white-noise specialization obtained by taking the trace.

C. Proofs of the FII via Data Processing Inequalities

The paper unifies three proofs of the Fisher information inequality as applications of data processing to a covariance-preserving linear transformation, with equality characterized by Gaussian inputs.

  • Three existing proofs of the Fisher information inequality are presented as variations on one data-processing argument.
  • The transformation maps independent variables (X_i)_i to their weighted sum Y, preserving the relevant covariance structure.
  • Applying data processing to the transformation yields the Fisher information inequality, and the Stam and Blachman proofs are equivalent.
  • The MMSE proof adds independent white Gaussian noise to each variable and becomes equivalent to the Fisher-information proof through the complementary Fisher information–MMSE relation.
  • Equality holds in the Fisher information inequality exactly when all variables with nonzero coefficients are Gaussian with identical covariances.

D. De Bruijn’s Identity

De Bruijn’s identity is presented as a general relation between entropy and Fisher information, proved through local expansions of mutual information under additive perturbations.

  • De Bruijn’s identity connects differential entropy and Fisher information and is conventionally used to derive the EPI from the Fisher information inequality.
  • The paper gives a simpler proof for independent X and finite-covariance perturbation Z, without requiring Z to have a density.
  • The proof writes I(X + θZ; Z) as an entropy difference and obtains the identity through a second-order divergence expansion.
  • The factor 1/2 in de Bruijn’s identity arises from the second-order Taylor-expansion factor in the definition of Fisher information.
  • For non-Gaussian perturbations, the general identity remains valid, but the positive-t extension relies on Gaussian stability and cannot generally be extended to non-Gaussian Z.
  • The identity also yields local interpretations: Fisher information measures sensitivity to additive noise, while Gaussian inputs minimize first-order mutual-information sensitivity in additive channels.

5) Applications:

The paper applies its information-theoretic framework to mutual-information saddlepoint properties, generalized de Bruijn identities, and the EPI with explicit equality conditions.

  • A Gaussian random vector with the same second moments as X minimizes I(X + Z; Z) among inputs when Z is independent Gaussian noise.
  • This saddlepoint proof uses divergence data processing and does not require the EPI, making it less involved than earlier scalar proofs.
  • The paper emphasizes that existing information-theoretic EPI proofs integrate Fisher-information or MMSE inequalities along a continuous Gaussian-perturbation path.
  • Along the standardized Gaussian path, the relevant Fisher-information function is nonincreasing, so f(0) ≥ f(∞) = 0 yields the EPI.
  • Equality in the generalized EPI requires the participating random vectors to be Gaussian with identical covariances; in the classical form, their covariances must be proportional.

2) Integral Representations of Differential Entropy:

Earlier EPI proofs represent entropy through integrals along Gaussian-perturbation paths, while the new proof works directly with mutual information and derives the EPI from an MII.

  • Earlier integral representations: Earlier proofs rewrite differential entropy as an integral involving Gaussian perturbations and Fisher information or MMSE.The representations use paths such as X + √t Z or √t X + Z and are equivalent under changes of variables.
  • Earlier integral representations: The EPI follows immediately from the corresponding Fisher information inequality when entropy is represented through de Bruijn’s identity.
  • Mutual-information proof: The new approach proves a mutual information inequality for independent random vectors with finite covariances and normalized coefficients.Its proof uses data processing for a covariance-preserving linear transformation and Gaussian perturbations.
  • Mutual-information proof: The mutual information inequality implies the EPI and is established using basic mutual-information and entropy properties, including Gaussian smoothing identities.The argument uses identities connecting I(X + Z; Z) with entropy differences and technical lemmas for Gaussian perturbations.
  • Mutual-information proof: The proof establishes the required inequality by showing a perturbation-dependent function is non-increasing and has value zero at the zero-noise limit.

B. Insights and Discussions

The paper notes that its mutual-information approach is motivated by the familiarity of Shannon mutual information and its data-processing theorem.

  • B. Insights and Discussions: The mutual-information formulation is presented as more familiar to readers than the corresponding Fisher-information data-processing theorem.

1) Relationship to Earlier Proofs:

The new proof shares the data-processing and Gaussian-perturbation structure of earlier EPI proofs but applies these ingredients directly to mutual information, avoiding de Bruijn’s identity, Fisher information, and MMSE.

  • Relationship to Earlier Proofs: The new proof uses these ingredients directly in mutual-information terms rather than through Fisher information or MMSE.
  • Relationship to Earlier Proofs: The paper identifies data processing under a covariance-preserving linear transformation and integration over continuous Gaussian perturbations as common ingredients of earlier proofs.
  • Relationship to Earlier Proofs: The mutual information inequality implies both the EPI and the Fisher information inequality through de Bruijn’s identity.
  • Limitations and scope: The paper notes scope boundaries involving equality analysis, moment assumptions, and unresolved extensions to integer-valued random variables.The method does not easily settle equality in the MII, finite covariance assumptions might be weakened, and its contribution to integer-valued cases remains open.
  • Relationship to Earlier Proofs: Equality in the mutual information inequality holds exactly when the nonzero-weight random vectors are Gaussian with identical covariances.This yields the corresponding necessity condition for equality in the EPI, although that condition is not evident from mutual-information properties alone.

7) On the EPI for Discrete Variables:

The paper discusses discrete entropy-power analogues and generalized linear-transformation inequalities, while emphasizing that the classical continuous EPI does not transfer directly to discrete variables.

  • On the EPI for Discrete Variables: The EPI also holds for independent discrete random vectors when differential entropies are replaced by entropies.
  • On the EPI for Discrete Variables: The classical entropy-power expression does not hold generally for discrete random vectors, with deterministic variables providing a counterexample.
  • On the EPI for Discrete Variables: Discrete analogues use different structures, including the Poisson distribution’s role for integer-valued variables and modulo-2 addition for binary variables.
  • Generalized linear transformations: The paper presents equivalent generalized EPI forms for linear transformations AX of independent variables under a full-row-rank matrix A.
  • Generalized linear transformations: The generalized mutual-information proof uses Gaussian perturbations, data processing, and Sato’s inequality before recovering the Zamir–Feder EPI.
  • Generalized linear transformations: The generalized result extends to vector-valued entries when A has orthonormal rows and the variables have finite variances.

V. TAKANO AND JOHNSON’S EPI FOR DEPENDENT VARIABLES

The paper extends the EPI to dependent random vectors through a perturbation condition expressed using mutual information, and shows that this condition is stronger than Takano’s and Johnson’s conditions.

  • General dependent-vector condition: A perturbation condition on dependent random vectors implies both the mutual information inequality (MII) and the EPI.The condition requires that adding suitably weighted Gaussian perturbations makes the variables more dependent for every t > 0.
  • Strength and application: The new condition is stronger because it yields an EPI valid for any choice of coefficients, not only the original unweighted sum.This stronger form is noted as potentially relevant to blind separation of dependent components.
  • Scalar characterization: For scalar variables, the perturbation condition is equivalent to a matrix inequality involving the joint score and Fisher information.The equivalence follows by differentiating mutual information with respect to the perturbation magnitude and applying de Bruijn’s identity.
  • Relation to prior conditions: The matrix condition implies both Takano’s and Johnson’s conditions for two dependent variables.Restricting the coefficient vector and minimizing the resulting quadratic forms recovers the two earlier criteria.

VI. LIU AND VISWANATH’S COVARIANCE-CONSTRAINED EPI

The paper develops covariance-constrained mutual information and entropy power inequalities using only basic mutual information properties, then recovers Gaussian optimality under covariance constraints.

  • Constrained Gaussian optimality: Liu and Viswanath’s constrained optimization problem still admits a Gaussian solution when Cov(X) ≤ C.The covariance matrix C is assumed positive definite.
  • Covariance-constrained inequalities: The paper gives explicit covariance-constrained MII and EPI forms for independent random vectors with proportional covariance structures.The derivation uses Gaussian perturbations and mutual-information identities rather than Fisher-information inequalities.
  • Mutual information inequality: The central mutual information inequality bounds I(a1X1 + a2X2 + Z; Z) by the sum of two individual mutual informations, up to o(α).The coefficients satisfy the normalization condition a1^2 + a2^2 = 1.
  • From MII to EPI: Integrating the resulting monotonicity relation yields the covariance-constrained MII and then the generalized EPI.The proof follows the same path-integration strategy used for the basic EPI.
  • Recovery of the optimization result: The covariance-constrained EPI recovers the Gaussian optimizer for the Liu–Viswanath maximization problem.The proof compares an arbitrary vector with a Gaussian vector having the same covariance as the restricted optimum.

VII. COSTA’S EPI: CONCAVITY OF ENTROPY POWER

The paper proves a generalized Costa inequality: entropy power is concave under addition of an arbitrary independent Gaussian vector when either summand is Gaussian.

  • Costa’s inequality: Costa’s EPI states that N(X + √t Z) is concave in t when Z is white Gaussian.The paper notes that this concavity is stronger than Shannon’s EPI for the same setting.
  • Generalization: The new theorem extends concavity to an arbitrary, not necessarily white, Gaussian random vector.The result applies when either X or Z is Gaussian.
  • Mutual-information proof: The proof reduces the claim to a mutual information inequality and uses Gaussian perturbations along a continuous path.The argument introduces independent Gaussian copies and derives the required second-derivative inequality from mutual-information relations.
  • Open extension: The authors leave open whether the proof extends to Costa’s matrix-valued generalization, where t is replaced by an arbitrary positive semi-definite matrix.This is stated as an unresolved extension of the method.

VIII. OPEN QUESTIONS

The paper surveys further entropy-power generalizations and derives several associated inequalities, while identifying limits of the mutual-information proof strategy.

  • Subset generalizations: Generalized EPI results extend the classical inequality to arbitrary collections of subsets of independent variables or vectors.Balanced collections can be assumed by adding singleton subsets until each index appears equally often.
  • Existing proof techniques: Existing proofs of the generalized inequalities rely on Fisher-information or MMSE integrations and may require an additional variance drop lemma.These techniques establish the generalized EPI and related inequalities through continuous perturbation paths.
  • Open limitation: A direct proof of the generalized mutual information inequality remains unavailable within the paper’s approach.The authors suggest that such an extension may require generalized data-processing or Sato-type inequalities.
  • Reverse derivation: The paper derives Fisher-information convexity from a mutual information inequality, obtaining a shorter proof than earlier direct arguments.The derivation divides by t and takes the limit t → 0 using de Bruijn’s identity.
Loading 0704.1751v2…