Source-linked AI summary

A Bivariate Measure of Redundant Information

Malte Harder, Christoph Salge, Daniel Polani

arXiv:1207.2080v3cs.ITphysics.data-an

TL;DR

The paper tackles the difficulty of defining redundant information about a third variable when existing measures can confuse redundancy with synergy or overlook information content. It introduces a geometric bivariate measure based on projections in probability-distribution space, proves its redundancy properties and non-negative decomposition, and demonstrates applications to examples and transfer entropy. The paper also identifies remaining limits concerning the separation of source and mechanistic redundancy and extension beyond two variables.

  • Problem

    Existing redundancy measures do not fully capture shared information about a third variable, motivating a formalism that distinguishes redundancy from synergy and respects desired properties.

  • Method

    The paper defines a bivariate redundancy measure using geometric projections of conditional outcome distributions and proves its required axiomatic properties.

  • Results

    The measure follows several redundancy intuitions, supports a non-negative mutual-information decomposition, and can decompose transfer entropy without the identified fake state-dependent transfer-entropy issue.

  • Takeaways & Limitations

    The measure serves as a replacement for the bivariate version of minimal information and emphasizes source versus mechanistic redundancy.

  • Takeaways & Limitations

    The current measure is restricted to the bivariate case, and separating mechanistic and source redundancy when they coexist remains unresolved.

Abstract

from arXiv · show

We define a measure of redundant information based on projections in the space of probability distributions. Redundant information between random variables is information that is shared between those variables. But in contrast to mutual information, redundant information denotes information that is shared about the outcome of a third variable. Formalizing this concept, and being able to measure it, is required for the non-negative decomposition of mutual information into redundant and synergistic information. Previous attempts to formalize redundant or synergistic information struggle to capture some desired properties. We introduce a new formalism for redundant information and prove that it satisfies all the properties necessary outlined in earlier work, as well as an additional criterion that we propose to be necessary to capture redundancy. We also demonstrate the behaviour of this new measure for several examples, compare it to previous measures and apply it to the decomposition of transfer entropy.

I. INTRODUCTION

The paper addresses the difficulty of measuring information shared by multiple variables about a third outcome, because existing approaches can conflate redundancy with synergy. It proposes a new bivariate formalism intended to satisfy established axioms, add an identity criterion, and support non-negative information decomposition.

  • Motivation: Interaction information cannot distinguish redundant information from synergy, making it unsuitable for precise multivariate information decomposition.It can capture both effects in the same quantity.
  • Motivation: Redundant information concerns information about a third variable that is shared by multiple input variables.The paper distinguishes this from mutual information between two variables.
  • Related work: Existing redundancy and synergy measures lack consensus and do not fully explain multivariate information in terms of atomic information quantities.The paper discusses total correlation, interaction information, and interaction complexity as related but insufficient approaches.
  • Related work: Williams and Beer’s decomposition captures redundancies and synergies across subsets, but it relies on a redundancy measure whose adequacy remains contested.The decomposition can also be applied to transfer entropy.
  • Contribution: The paper proposes a geometric bivariate measure, compares it with existing measures, and proves that it satisfies required axioms while preserving non-negativity.The authors also propose an additional criterion for redundancy.

B. Redundancy Axioms

The paper reviews redundancy axioms and introduces an identity property to ensure that redundancy reflects shared information content rather than merely equal information amounts. It then motivates a geometric construction using conditional distributions and information projections.

  • Redundancy axioms: The minimal-information measure is non-negative and satisfies the three established redundancy axioms.These axioms imply non-negativity and an upper bound by each source’s mutual information with the outcome.
  • Failure of minimal information: Minimal information assigns 1 bit of redundancy when independent binary X and Y jointly copy into Z=(X,Y), contradicting the intuition that their information is not shared.X and Y independently describe different components of Z.
  • Failure of minimal information: The issue is that X and Y can provide equal information amounts while inducing different posterior distributions and therefore different information content.The paper proposes separating these contributions geometrically in the space of distributions over Z.
  • Additional axiom: The identity property requires redundancy for a copying mechanism to equal the mutual information between the input variables.For multivariate measures, monotonicity then bounds redundancy by the minimum pairwise mutual information.
  • Geometric construction: The proposed measure uses information projections of conditional distributions onto convex closures formed from the other source’s conditional distributions.The construction operates in the probability-distribution space over Z and uses Kullback–Leibler divergence.

2. Projective Information

Projected information is constructed by projecting conditional distributions onto the convex closure generated by another variable. It measures shared information through the resulting Kullback–Leibler divergence geometry.

  • Definition: Projected information projects conditionals of one variable onto the convex closure of the other variable’s conditional distributions.The projection is defined through a distance-minimization construction.
  • Interpretation: The construction measures information shared by X and Z that can be expressed through information shared by Y and Z.The projection is illustrated for binary input variables.
  • Properties: Projected information is well-defined, finite, and non-negative.These properties hold even though the projection itself need not be unique.
  • Properties: Non-unique projections do not affect projected information because all minimizing solutions have the same relevant KL divergence.The proof rewrites projected information as a difference of KL divergences.
  • Properties: The projected-information construction yields a non-negative redundancy quantity through its KL-divergence-based formulation.The supplied derivation states the resulting quantity is non-negative.

3. Definition of Bivariate Redundancy

The proposed bivariate redundancy measure is defined as the minimum of two projected-information terms, using geometric projections of conditional distributions. The paper establishes that it satisfies the required redundancy axioms and is a good candidate measure.

  • Ired(Z; X, Y ) is defined as the minimum of the two projected information terms.The minimum is taken after correcting for distributional changes in different directions by projecting the conditionals.
  • The geometric construction projects conditional distributions in the probability-distribution space using Kullback-Leibler divergence.The projection is formulated as a constrained KL-divergence minimization over distributions associated with the other source variable.
  • The measure is symmetric and satisfies self-redundancy, monotonicity, and the proposed identity property.The paper introduces lemmas and propositions to establish the required axioms, including monotonicity and identity.
  • The measure is bounded by each source’s mutual information with the target, including Ired(Z; X, Y ) ≤ I(Z; X).The bound follows from the non-negativity of KL-divergence and supports the measure’s redundancy interpretation.
  • Monotonicity requires that adding variables to one source cannot decrease the measured redundancy.The paper concludes Ired(Z; X, Y ) ≤ Ired(Z; X, (Y, W)).
  • The paper concludes that Ired is a good candidate for measuring redundancy with respect to a target variable.

IV. COMPARISONS

The paper introduces examples to examine the behavior of the bivariate redundancy measure and its comparison with related measures.

  • The examples evaluate the bivariate measure through several redundancy calculations.

A. Relation to Minimal Information

Compared with minimal information, Ired generally avoids the overestimation associated with ignoring directional changes in probability-distribution space. The resulting PI decomposition remains non-negative and separates redundant, unique, and synergistic information.

  • Imin generally overestimates redundancy relative to Ired, especially as the dimension of Z increases.The paper attributes this tendency to the error from not accounting for directionality in higher-dimensional distribution spaces.
  • The decomposition can be visualized as a PI-diagram whose regions represent redundant, unique, and synergistic information.
  • PI-atoms decompose mutual information into redundant, unique, and synergistic contributions for the source variables and target Z.The bivariate decomposition contains four atoms: redundancy, two unique-information terms, and synergy.
  • The sum of the four PI-atoms equals the mutual information I(Z; X, Y ).
  • Using Ired, the bivariate information decomposition is shown to be non-negative.Non-negativity of redundancy and its upper bounds by I(X; Z) and I(Y; Z) ensure non-negative unique terms, while a further lemma establishes non-negative synergy.

C. Examples

The paper examines the proposed bivariate measure on examples selected to test desired properties of redundancy and synergy measures.

  • The examples are drawn from cases previously discussed as tests for desired redundancy and synergy properties.

1. Copying - From Redundancy to Uniqueness

The copying example varies input correlation to exchange unique information for redundancy, while XOR isolates synergy and AND reveals mechanistic redundancy even with independent inputs.

  • Copying: λ ∈[0, 1] controls the correlation between binary inputs X and Y in the copying mechanism Z = (X, Y).Varying λ changes the outcome entropy while exchanging unique information for redundancy.
  • Copying: At λ = 0, X and Y are identical copies of W, yielding the redundant-information example with Ired(Z; X, Y) = Ired(W; X, Y).At the opposite extreme, λ = 1 makes the inputs independent and recovers the unique-information example.
  • XOR: The XOR gate yields purely synergistic information, with Ired(Z; X, Y) = Imin(Z; X, Y) = 0.The output is only known when both inputs are available.
  • AND: For the AND gate, independent inputs still produce Ired(Z; X, Y) = Imin(Z; X, Y) = 0.311278 because observing either input as 0 makes Z certain.This redundancy reflects shared information about the output distribution generated by the mechanism, rather than dependence between sources.
  • Mechanistic redundancy: The paper calls redundancy arising only from the mechanism mechanistic redundancy and distinguishes it from source redundancy already present in input mutual information.The paper does not give a rigorous separation because some cases do not make the distinction clear.

4. Summing Dice

The dice examples vary both source correlation and the summation mechanism, showing that redundant information depends on how inputs are combined as well as on source dependence.

  • Summing dice: For R = αD1+D2 with α ∈{1, 2, 3, 4, 5, 6}, redundancy decreases as the summation becomes more informative about the joint dice state.At α = 6, 6D1 + D2 is isomorphic to (D1, D2).
  • Composition of mechanisms: The composite RdnXor example produces one bit of redundant and one bit of synergistic information, matching Imin.It combines the redundant copy case with an XOR gate.
  • Composition of mechanisms: The RdnUnqXor example yields 1 bit in each partial-information term and 4 bits of total mutual information.The terms comprise 1 bit redundant, 1 bit synergistic, and 1 bit unique information per input.
  • Summing dice: With independent dice, the amount of redundancy depends on the summation mechanism, whereas at maximal source correlation all redundancy comes from source correlation.Figure 8 varies the source correlation λ and reports curves for α = 1 through 6.
  • Composition of mechanisms: The XorAnd composition differs from the corresponding earlier result because the AND component introduces mechanistic redundancy.The paper’s summary states that Ired captures the proposed redundancy concept and agrees with desired examples except where mechanistic redundancy appears.

D. Information Transfer

The paper applies Ired and Imin to transfer-entropy decompositions into state-independent and state-dependent components, finding that Ired avoids a misleading redundancy assignment in a constructed process.

  • Decomposition: Transfer entropy is decomposed into State Independent Transfer Entropy (SITE) and State Dependent Transfer Entropy (SDTE) using a partial-information decomposition.SITE is information uniquely from the source process; SDTE is source information dependent on the target process state.
  • First example: In the first coupled-process example, decompositions using Ired and Imin coincide across d, with d = 0 giving only state-independent transfer and d = 1 only state-dependent transfer.The example varies dependence on the previous state of X.
  • Second example: For the second process, the decompositions coincide for d ≤0.5, while at d = 1 Ired gives complete state-independent transfer entropy and Imin gives total state-dependent transfer entropy.At d = 0, the processes are independent and overall transfer entropy vanishes.
  • Second example: Imin incorrectly detects redundancy between independent Xt and Zt because it compares changes in different directions in distribution space.The parallel independent process Xt creates a dependency that does not exist in the process itself.
  • Projection analysis: Projected information supplies the redundancy calculation: for d ≤0.5, the relevant projections equal conditionals given Zt, so Iπ Yt+1 (Yt % Zt) = I(Yt+1; Zt).The one-dimensional distribution space of binary Yt+1 makes the projection comparison direct.
  • Second example: For the reduced transfer TZ→Y, the mechanism creates redundancy between independent Yt and Zt because both states alter the distribution of Yt+1 in the same direction.At d = 0.5, either current state contributes to predicting Yt+1; redundancy vanishes at d = 0 and d = 1.

1. Open Loop Controllability

The paper extends the connection between transfer-entropy decomposition and open-loop controllability from Imin to the new redundancy measure Ired.

  • Controllability: Perfect controllability requires that every initial state and final state have some control state c with p(x′|x, c) = 1.The paper uses this definition in relating controllability to transfer-entropy decompositions.
  • Open-loop controllability: Perfect open-loop controllability is equivalent to perfect controllability together with I(X; C) = 0.The cited theorem relates this condition to vanishing state-dependent transfer entropy under Imin.
  • Extension to Ired: The paper proves that the same equivalence holds when Ired is used as the underlying redundancy measure.Thus the relation between open-loop controllability and transfer-entropy decomposition transfers to the new measure.
  • Proof ingredients: The direct implication uses the established condition for perfect open-loop controllers to obtain zero state-dependent transfer entropy with Ired.The paper states that this proves perfect open-loop controllability implies perfect controllability with zero SDTE under Ired.
  • Proof ingredients: Under perfect controllability, Lemma 11 establishes Ired(X′; X, C) = 0.The proof uses the fact that the relevant convex closure reduces to the marginal distribution p(x′).

V. DISCUSSION

The paper introduces a bivariate redundant-information measure based on similarities in how observing each input changes the outcome distribution. It satisfies established redundancy properties, supports non-negative mutual-information and transfer-entropy decompositions, and highlights mechanistic versus source redundancy while retaining important scope limitations.

  • The new measure is conceptually motivated by similarities in the direction of change in the outcome distribution when either input is observed.
  • The construction satisfies redundancy properties from the literature and supports a non-negative decomposition of mutual information.
  • The measure decomposes transfer entropy, whereas using minimal information can detect fake state-dependent transfer entropy.
  • The definition emphasizes mechanisms by distinguishing source redundancy already present in inputs from mechanistic redundancy arising from the mechanism.Copying, Andgate, and 50:50-readout examples motivate this distinction, while the separation is not yet explicit.
  • A current limitation is that the measure is bivariate, while multivariate generalization has more complex geometry and no clear replacement for the identity property.Future work is also needed to determine whether mechanistic and source redundancy can be separated when they occur simultaneously.
Loading 1207.2080v3…