Source-linked AI summary

Quantifying unique information

Nils Bertschinger, Johannes Rauh, Eckehard Olbrich, Jürgen Jost, Nihat Ay

arXiv:1311.2852v2cs.IT

TL;DR

The paper proposes measures that decompose information about X in (Y,Z) into shared, unique, and synergistic components. It motivates these measures operationally, studies their properties, and compares them with alternatives, while noting limits of operational characterization and differences from Ired.

  • Problem

    Information about X in Y and Z can comprise shared, unique, and synergistic components, motivating measures that distinguish these contributions.

  • Method

    The paper formalizes unique information through decision problems and defines functions for unique, shared, and synergistic information over joint distributions.

  • Results

    The proposed functions are non-negative, and f SI is shown to satisfy f SI(X : Y ; Z) ≥ 0.

  • Takeaways & Limitations

    The measures provide a studied decomposition framework whose unique-information characterization can distinguish when unique information vanishes.

  • Takeaways & Limitations

    The operational idea distinguishes when unique information vanishes but does not quantify it.

Abstract

from arXiv · show

We propose new measures of shared information, unique information and synergistic information that can be used to decompose the multi-information of a pair of random variables $(Y,Z)$ with a third random variable $X$. Our measures are motivated by an operational idea of unique information which suggests that shared information and unique information should depend only on the pair marginal distributions of $(X,Y)$ and $(X,Z)$. Although this invariance property has not been studied before, it is satisfied by other proposed measures of shared information. The invariance property does not uniquely determine our new measures, but it implies that the functions that we define are bounds to any other measures satisfying the same invariance property. We study properties of our measures and compare them to other candidate measures.

1 Introduction

The paper frames information about X in (Y,Z) as shared, unique, and synergistic components, then proposes new measures based on pairwise marginal invariance. These measures are shown to be nonnegative and to bound other decompositions satisfying the same invariance property.

  • 1 Introduction: MI(X : (Y, Z)) is decomposed into shared, two unique, and synergistic information terms.The corresponding identities require the four terms to be nonnegative and recover MI(X : Y) and MI(X : Z).
  • 1 Introduction: Co-information equals shared information minus synergistic information, but does not by itself separate those contributions.The paper notes that positive co-information signals redundancy and negative co-information expresses synergy, while a satisfactory separation remains missing.
  • 1 Introduction: The proposed unique-information framework restricts measures to depend only on the marginal distributions of (X,Y) and (X,Z).The construction considers joint distributions sharing those two pair marginals and defines the functions using optimization over that set.
  • 1 Introduction: The paper defines f UI, f SI, and f CI as new measures of unique, shared, and complementary information, with the fourth decomposition term determined by the identities.Their motivation is an operational interpretation in which unique information should be extractable through decision problems.
  • 1 Introduction: The paper proves that f UI, f SI, and f CI are nonnegative and studies their further properties.It also compares the proposed functions with other information decompositions and discusses computational aspects and examples.

2 Operational interpretation

The paper gives unique information an operational meaning through decision problems and then quantifies it using distributions that preserve the pair marginals with X. Under this invariance assumption, the resulting functions provide bounds and characterize when synergistic information can vanish.

  • 2 Operational interpretation: Unique information is characterized by whether observing one variable can yield a strictly higher reward than observing the other in some decision problem.The decision problem consists of a prior over X, an action set, and a reward function.
  • 2 Operational interpretation: The operational criterion identifies when unique information vanishes but does not quantify its amount.The paper therefore seeks a measure depending only on the prior and the two conditional channels.
  • 2 Operational interpretation: The proposed optimization ranges over all joint distributions sharing the marginal distributions of (X,Y) and (X,Z).The discussion assumes full support for X; otherwise X can be replaced by its support.
  • 2 Operational interpretation: Under the invariance assumption, unique and shared information are constant across this set, while only complementary information depends on the joint distribution.The functions f UI, f SI, and f CI are obtained through minima and maxima over the set, which are well-defined by compactness and continuity.
  • 2 Operational interpretation: The functions f UI, f SI, and f CI bound any nonnegative continuous decomposition satisfying the decomposition identities and the invariance assumption.Equality holds when some compatible joint distribution has zero synergistic information, and conversely tightness implies such a distribution exists.
  • 2 Operational interpretation: The paper interprets the proposed functions as the unique measures for which compatible pair marginals guarantee a joint distribution with zero complementary information.Other invariant measures can permit nonvanishing complementary information to be inferred from the same pairwise data.

3 Properties

The proposed measures are non-negative, satisfy key bivariate consistency and symmetry properties, and characterize when shared, unique, or complementary information vanishes. They also recover expected values for structured distributions, while the shared-information measure does not extend to the multivariate PI lattice.

  • Characterization and positivity: The optimization problems defining f UI, f SI, and f CI are convex optimization problems on convex sets.The relevant mutual-information functions are convex, while co-information and conditional entropy are concave.
  • Characterization and positivity: f UI, f SI, and f CI are non-negative functions.f SI is bounded below by a non-negative co-information value, while f UI follows from non-negativity of mutual information and f CI by definition.
  • Vanishing shared and unique information: f UI vanishes exactly when the corresponding source has no unique information about X relative to the other source, according to the operational definition.The characterization uses a row-stochastic matrix relating the relevant pair marginals, and Corollary 7 states the operational equivalence.
  • The bivariate PI axioms: The measures satisfy symmetry and bivariate PI inequalities, including f SI(X : Y ; Z) ≤MI(X : Y ) and f UI(X : Y \ Z) ≥MI(X : Y ) −MI(X : Z).They also satisfy the unique-information consistency relation MI(X : Z) + f UI(X : Y \ Z) = MI(X : Y ) + f UI(X : Z \ Y ).
  • The bivariate PI axioms: If X is conditionally independent of Z given Y, then f SI(X : Y ; Z) = MI(X; Z) and f CI(X : Y ; Z) = 0.More generally, zero conditional mutual information implies both the corresponding unique and complementary information terms vanish.
  • Probability distributions with structure: For X = (Y, Z), the structured-distribution identities give f SI((Y, Z) : Y ; Z) = MI(Y : Z), f UI((Y, Z) : Y \ Z) = H(Y |Z), and f UI((Y, Z) : Z \ Y ) = H(Z|Y ).The measures also decompose independent component pairs additively, but f SI cannot be generalized to n = 3 within the PI lattice.

4 Comparison with other measures

The paper compares f SI with Ired and Imin, finding shared-information agreement in several special cases but also systematic differences. Both Ired and Imin satisfy the invariance assumption and therefore upper-bound f SI.

  • Ired and Imin satisfy the invariance assumption, yielding Ired ≥ f SI and Imin ≥ f SI.
  • Ired(X : Y ; Z) = 0 if and only if f SI(X : Y ; Z) = 0.
  • Theorem 22 shows that UIred vanishes exactly when f UI vanishes, so Ired is consistent with the paper’s operational criterion.
  • Under conditional independence of X and Z given Y, or of X and Y given Z, Ired equals f SI.
  • If (X,Y) and (X,Z) have the same marginal distribution, then Ired(X : Y ; Z) = f SI(X : Y ; Z).
  • Although f SI and Ired often agree, the dice example shows they are different functions, and Ired does not satisfy property (∗∗).

5 Examples

The examples compare f SI with Ired on standard cases and a parameterized dice system. They show broad agreement in paradigmatic examples, but different dependence on the dice parameters in the more complex case.

  • In the paradigmatic examples of Table 1, f SI agrees with Ired and with the intuitively plausible expected values.
  • For X = Y + αZ, the dice dependence is controlled by λ: λ = 0 gives complete correlation, while λ = 1 gives independence.
  • For α = 1, 5, and 6, and for λ = 0 and λ = 1, f SI and Ired agree; otherwise f SI ≤ Ired.
  • For small λ and α > 1, f SI depends only weakly on α, whereas Ired depends more strongly on α.
  • The paper does not determine which of the two parameter-dependence patterns is more intuitive.

6 Outlook

The outlook extends the paper’s bivariate decomposition conceptually to larger systems while identifying a fundamental obstruction to a direct multivariate PI-lattice generalization. The bivariate unique-information measure can nevertheless assess individual variables in larger systems.

  • The decomposition provides non-negative shared, unique, and synergistic terms and satisfies properties including the PI axioms and identity axiom.
  • For n > 2, there is no universal agreement that shared, unique, and synergistic information form a complete decomposition.
  • The PI-lattice framework decomposes MI(X : Y1, Y2, Y3) into 18 shared-information terms with specified interpretations.
  • The function f SI cannot be generalized to n = 3 within the PI lattice because the identity axiom conflicts with non-negativity.
  • For larger systems, bivariate f UI can measure Yi’s unique information relative to all other variables and assess its value when synergy is ignored.
  • Unique information cannot increase when additional variables are taken into account: f UI(X : Y \ (Z1, ..., Zk)) ≥ f UI(X : Y \ (Z1, ..., Zk+1)).
  • The authors position the operationally motivated measure as a starting point toward a general multivariate information decomposition.

A.1 The optimization domain ∆P

The optimization domain ∆P consists of joint distributions preserving the pair marginals of (X,Y) and (X,Z). It is a polytope that can be represented through conditional distributions for each positive-probability value of X.

  • Computing f UI, f SI, and f CI requires solving a convex optimization problem over ∆P.
  • The domain ∆P is the intersection of an affine space and a simplex, hence a polytope.
  • The affine map A preserves the pair marginals of (X,Y) and (X,Z), with ∆P = (P + ker A) ∩ ∆.
  • The defect of A, equal to dim ker A, is |X|(|Y| − 1)(|Z| − 1), and the γ vectors provide spanning and basis descriptions of ker A.
  • Because γ vectors for different x have disjoint supports, ∆P factors into simpler polytopes, although mutual information does not respect this product structure.
  • For each x with P(x) > 0, ∆P is represented by joint distributions of Y and Z whose marginals match P(Y|X = x) and P(Z|X = x).
  • The map from Q ∈ ∆P to its conditional distributions given X is a linear bijection, and Q(X = x,Y = y,Z = z) = P(X = x)Q(Y = y,Z = z|X = x).
  • When both conditional distributions P(Y|X = x) and P(Z|X = x) are point measures, ∆P is a singleton.

A.2 The critical equations

The section derives critical-point conditions for the optimization defining the measures and illustrates their behavior in the AND-example and an ill-conditioned binary example.

  • Critical equations: The optimization problems are characterized by vanishing directional derivatives along feasible directions γx;y,y′;z,z′.The derivative is evaluated in directions preserving the relevant pair marginals, and Q solves the problems exactly when the corresponding conditions hold.
  • Example 30 (The AND-example): In the AND-example, the measures assign 0 unique information, 3 shared information, and 1 synergistic information.The resulting decomposition contains only shared and synergistic information.
  • Example 31: The binary i.i.d. example makes the feasible set ΔP a square and shows that CoIQ can vary very little along one diagonal.Along that diagonal, X is independent of (Y, Z), corresponding to very low synergy.
  • Example 31: Although the optimizing distribution QP is unique in this example, finding it can be difficult for constrained numerical optimization.Mathematica’s FindMinimum does not always find the true optimum out of the box.
  • Example 31: The uniform distribution is the maximum of CoIQ at the square’s center, while the corners distinguish independence from XOR-type relationships.The dark corners represent X independent of Y and Z, whereas the light corners represent independent Y and Z with X equal to XOR or its negation.

A.3 Technical proofs

The technical proofs establish optimality through convexity, critical-point conditions, and product constructions, and verify a zero-shared-information result under specified independence assumptions.

  • Proof of Lemma 9: If MIQ0(Y : Z) = 0, all partial derivatives vanish at Q0, so Q0 solves the optimization problems and fSI(X : Y ; Z) = 0.The argument combines the co-information identity with the construction of Q0 and the critical-equation lemma.
  • Proof of Lemma 19: For independent triples, the product Q = Q1Q2 belongs to ΔP and solves the corresponding optimization problems for the combined variables.The construction combines solutions for (X1, Y1, Z1) and (X2, Y2, Z2).
  • Proof of Lemma 19: The proof that Q is optimal again uses the critical-equation characterization after constructing the product distribution.The relevant derivatives vanish at the constructed Q, which therefore solves the optimization problems.
  • Proof strategy: A critical point of the divergence restricted to a convex set is a global minimizer because information divergence is jointly convex.This convexity argument is used to prove optimality of the constructed distributions.
  • Proof of Lemma 21: Under the stated conditional-independence assumptions, PX is shown to be a critical point, yielding Iπ(X : Y ↘ Z) = 0 and Ired(X : Y ; Z) = 0.The proof uses independence of Y and Z to establish the required derivative condition.
Loading 1311.2852v2…