Source-linked AI summary

Exploration of synergistic and redundant information sharing in static and dynamical Gaussian systems

Adam B. Barrett

arXiv:1411.2832v2cs.ITq-bio.NC

TL;DR

The paper asks how joint information from two sources can be decomposed when Shannon theory does not uniquely specify redundancy, unique information, and synergy. It analyzes PID procedures for static and dynamical Gaussian systems, finding a common Gaussian reduction with frequent net synergy and implications for continuous time-series analysis.

  • Problem

    Shannon information theory does not provide enough equations to determine redundancy, unique information, and synergy separately in a PID of two sources and one target.

  • Method

    The paper analyzes previously proposed PID procedures on continuous Gaussian systems, including static variables and dynamical time-series pasts.

  • Results

    For a broad class of Gaussian systems, the analyzed PIDs reduce redundancy to the minimum source mutual information, while Gaussian systems frequently exhibit net synergy.

  • Takeaways & Limitations

    The resulting independent formulae enable synergy and redundancy to be studied among continuous time-series variables and in complex-system applications.

  • Takeaways & Limitations

    PIDs involving more than two source variables are beyond the paper’s scope, and synergistic complexity for many-element systems is left for future work.

Abstract

from arXiv · show

To fully characterize the information that two `source' variables carry about a third `target' variable, one must decompose the total information into redundant, unique and synergistic components, i.e. obtain a partial information decomposition (PID). However Shannon's theory of information does not provide formulae to fully determine these quantities. Several recent studies have begun addressing this. Some possible definitions for PID quantities have been proposed, and some analyses have been carried out on systems composed of discrete variables. Here we present the first in-depth analysis of PIDs on Gaussian systems, both static and dynamical. We show that, for a broad class of Gaussian systems, previously proposed PID formulae imply that: (i) redundancy reduces to the minimum information provided by either source variable, and hence is independent of correlation between sources; (ii) synergy is the extra information contributed by the weaker source when the stronger source is known, and can either increase or decrease with correlation between sources. We find that Gaussian systems frequently exhibit net synergy, i.e. the information carried jointly by both sources is greater than the sum of informations carried by each source individually. Drawing from several explicit examples, we discuss the implications of these findings for measures of information transfer and information-based measures of complexity, both generally and within a neuroscience setting. Importantly, by providing independent formulae for synergy and redundancy applicable to continuous time-series data, we open up a new approach to characterizing and quantifying information sharing amongst complex system variables.

1 Introduction

The paper addresses the inability of Shannon information theory to uniquely decompose joint information from two sources into redundancy, unique information, and synergy. It analyzes proposed PID procedures for Gaussian systems and develops tools for studying information sharing in continuous time series.

  • The PID problem: Shannon information theory does not uniquely determine the redundant, unique, and synergistic components of I(X; Y, Z).The standard equations determine only Whole-Minus-Sum net synergy, S(X; Y, Z) − R(X; Y, Z).
  • The PID problem: A complete PID requires an additional definition that fixes one of the four information-sharing quantities while preserving nonnegativity and source symmetry.Several distinct PID definitions have been proposed from different conceptions of redundancy and synergy.
  • Gaussian PID result: Positive net synergy can occur in jointly Gaussian systems even when the sources are uncorrelated.An example uses equally correlated sources with the target and zero correlation between the sources.
  • Gaussian PID result: For jointly Gaussian systems with a univariate target and arbitrarily dimensional sources, the proposed PID procedures reduce redundancy to min{I(X; Y), I(X; Z)}.The common minimum mutual information (MMI) PID then determines the remaining quantities through the PID equations.
  • Scope and applications: The framework extends analysis of information sharing beyond discrete systems to Gaussian continuous time-series variables and potential applications in complex systems.The paper discusses implications for information transfer, complexity measures, and neuroscience.

2 Notation and preliminaries

The paper establishes covariance-based notation and entropy/mutual-information identities for continuous Gaussian variables, then introduces dynamical past-state notation and illustrates Gaussian net-synergy examples.

  • Covariance notation: For continuous variables, covariance matrices and cross-covariance matrices describe within-variable and between-variable relationships.The notation uses Σ(X) for covariances and Σ(X, Y) for cross-covariances.
  • Covariance notation: The partial covariance Σ(X|Y) is defined as Σ(X) − Σ(X, Y)Σ(Y)^−1Σ(Y, X).For jointly Gaussian variables, it equals the covariance matrix of the conditional variable X|Y = y.
  • Information measures: Entropy quantifies uncertainty, while mutual information measures the average reduction in uncertainty about X from knowing Y.For continuous variables, the entropy expression is differential entropy under a density assumption.
  • Information measures: Mutual information is symmetric, and joint mutual information obeys the chain rule I(X; Y, Z) = I(X; Y|Z) + I(X; Z).The conditional mutual information is the expected mutual information between X and Y given Z.
  • Dynamical notation: For dynamical variables, the notation distinguishes the infinite past from a finite collection of p past states relative to time t.These past states provide the temporal source variables used in dynamical Gaussian analyses.
  • Gaussian examples: Figure 2 presents two Gaussian systems with positive net synergy: one has uncorrelated sources, and the other has an uncorrelated target-source pair.Variables are circles, and correlated pairs are connected by lines.

3 Synergy is prevalent in Gaussian systems

Jointly Gaussian systems frequently exhibit positive net synergy, including cases with uncorrelated sources. Its dependence on source correlation varies with source–target correlations and with the information measure used.

  • Positive WMS net synergy is a sufficient condition for positive synergy because PID synergy and redundancy are nonnegative.WMS is whole information minus the sum of individual informations, and equals synergy minus redundancy.
  • Net synergy can occur when equally target-correlated sources are uncorrelated, despite the absence of source correlation.This result is reported for a = c and b = 0.
  • Zero source correlation gives zero covariance-reduction WMS, whereas Shannon-information WMS can remain positive because reductions are combined through a concave logarithm.The covariance-based measure adds individual covariance reductions linearly in this case.
  • For equal positive source–target correlations, net synergy decreases with source correlation; for equal opposite correlations, it increases.The dependence is described for both Shannon information and the variance-reduction alternative.
  • With unequal source–target correlations, net synergy is U-shaped in source correlation and diverges near singular covariance limits.Near those limits, the target becomes completely determined by the two sources.
  • Net redundancy does not necessarily indicate highly correlated source variables.The paper uses the correlation-dependent Gaussian examples to establish this distinction.

4 Partial information decomposition on Gaussian systems

Several proposed PIDs coincide for jointly Gaussian systems with a univariate target under a marginal-dependence condition. They yield minimum-mutual-information redundancy, zero unique information for the weaker source, and synergy as its additional contribution.

  • Common structure: Previously proposed PIDs make redundancy and unique information depend only on (X,Y) and (X,Z), while synergy depends on the full joint distribution.This shared structural property is the basis for the Gaussian reduction.
  • Previously proposed PIDs: The Griffith–Bertschinger union-information construction minimizes joint information over alternatives preserving each source–target marginal.The alternative variables may differ in their relations to each other.
  • Previously proposed PIDs: Harder, Salge and Polani define redundancy through the closeness of conditional distributions of X given Y and Z.Their formulation uses a divergence from conditional distributions and linear combinations of them.
  • Common Gaussian PID: For jointly Gaussian variables with a univariate target and arbitrary-dimensional sources, the resulting unique PID sets redundancy to min{I(X;Y), I(X;Z)}.The paper terms this the MMI, or minimum mutual information, PID.
  • Common Gaussian PID: The weaker source has zero unique information, while synergy is the extra information it contributes when the stronger source is known.All previously proposed PIDs reduce to this MMI PID in the stated Gaussian case.
  • Scope: The common PID does not extend to multivariate targets because the required matrix equation need not have a solution.The paper leaves that more general case for future work.
  • MMI synergy: In the univariate Gaussian case, MMI synergy diverges at singular limits, reaches zero at b = a/c, and is positive between those values.Its dependence on source correlation is decreasing, increasing, or U-shaped according to the relative signs and magnitudes of a and c.

5 Dynamical systems

The paper analyzes information sharing in Gaussian MVAR systems through explicit two- and three-variable examples. Net synergy depends on history length, connectivity, and source correlation, with three-variable systems supporting persistent synergy.

  • Example 1: Synergistic two-variable system: Example 1 has positive net synergy for immediate pasts but zero net synergy for infinite pasts of X and Y about present X.The infinite-history result follows because restricted regression on one variable’s past can become infinite order.
  • Example 1: Synergistic two-variable system: In Example 1, infinite-history synergy equals redundancy, while the MMI PID gives the same synergy at infinite and one lag but less redundancy.Thus the complete histories contain less synergy relative to redundancy than the immediate pasts.
  • Example 2: An MVAR model with no net synergy: Example 2 has zero one-lag synergy and greater infinite-history redundancy than synergy because information from X’s past is mediated through Y.The model therefore exhibits negative net synergy, equivalent to positive net redundancy, for infinite pasts.
  • Example 3: Synergy between two variables influencing a third variable: Example 3 shows that two source histories can provide net synergy to X, with MMI synergy increasing with weaker-link strength and decreasing with source correlation.Synergy vanishes when the weaker connection α is zero or when the sources are perfectly correlated, ρ = 1.
  • Example 3: Synergy between two variables influencing a third variable: For uncorrelated equally connected sources, net synergy is approximately α4/2 for small α and grows approximately logarithmically for large α.The paper also notes that net synergy can become arbitrarily large as connection strength increases in the relevant model.

6 Transfer entropy

The paper shows that synergy changes how transfer entropy and Gaussian Granger causality should be interpreted. Pairwise transfer measures can include synergistic predictive information, and conditioning can alter transfer entropy through synergy.

  • Pairwise transfer entropy: Pairwise transfer entropy measures unique information from Y plus synergistic information jointly provided by the pasts of X and Y.This differs from the common interpretation that it isolates only Y’s contribution beyond X’s past.
  • Pairwise transfer entropy: Lagged mutual information instead combines Y’s unique information with redundant information shared by the pasts of X and Y.Consequently, transfer entropy need not be less than lagged mutual information when net synergy is present.
  • History length: In Example 1, transfer-entropy quantities based on infinite histories become equal because complete past histories have zero net synergy.The distinction between immediate- and infinite-history measures therefore depends on the system’s history-dependent PID.
  • Conditional transfer entropy: Conditional transfer entropy can be affected by synergy even with infinite lags.In Example 3, non-conditional and conditional transfer entropy differ through their respective redundancy and synergy terms.
  • Transfer entropy and Granger causality: For jointly Gaussian variables, transfer entropy is equivalent to the linear formulation of Granger causality.Granger causality is expressed through prediction improvements and residual-variance ratios.

7 Implications for measures of overall interactivity and complexity

The paper examines how synergistic information affects measures of overall transfer and complexity. It proposes synergistic complexity, contrasts it with causal density and global transfer entropy, and identifies scope limits for larger systems.

  • Synergistic complexity: Synergistic complexity measures how much joint information from source pairs exceeds the sum of their individual information contributions.The proposal treats whole-system information exceeding the sum of its parts as a complexity-related quantity.
  • Causal density: Causal density counts synergistic information multiple times while neglecting redundant information.In Example 3, its non-zero contributions include twice the synergy term.
  • Causal density: Despite this overcounting, causal density remains a sensible measure of overall novel predictive transfer in Example 3.It increases with connection strengths α and γ, decreases with source correlation ρ, and vanishes when both connections are zero or ρ approaches 1.
  • Global transfer entropy: Global transfer entropy weights unique, redundant, and synergistic information equally as average information flow from the entire system to individual elements.Unlike causal density, it is not selective for particular PID components.
  • Global transfer entropy: In Example 3, global transfer entropy increases with source correlation because positively correlated source fluctuations combine to produce larger target fluctuations.Thus it does not operationalize complexity as inhomogeneity of information sources in this example.
  • Scope and future work: The paper does not evaluate these complexity measures on systems with many more than three elements or PIDs with more than two source variables.Such comparisons and generalized measures are left for future research.

8 Discussion

The paper shows that Gaussian systems commonly exhibit non-trivial information sharing, including net synergy, and derives PID interpretations relevant to dynamical analyses and neuroscience. These results challenge simple links between redundancy and source correlation while enabling separate computation of synergy and redundancy in Gaussian time-series data.

  • 8 Discussion: Gaussian systems with linear interactions frequently exhibit net synergy, where joint information exceeds the sum of individual informations.Examples include correlated targets with uncorrelated sources and targets linked to one source when the sources themselves are correlated.
  • 8 Discussion: For a broad class of Gaussian systems, redundancy is the weaker source’s mutual information and is independent of source correlation.This result applies to jointly Gaussian systems with a univariate target and sources of arbitrary dimension under several previously proposed PIDs.
  • 8 Discussion: Synergy is the extra information contributed by the weaker source once the stronger source is known, and its dependence on source correlation can vary with correlation signs.Thus, source correlation alone does not determine whether Gaussian information sharing is redundant or synergistic.
  • 8 Discussion: In dynamical Gaussian systems, the pasts of two variables can provide synergistic information about a third variable’s present state.The paper demonstrates this with an MVAR model and connects the result to information-transfer analyses.
  • 8 Discussion: Net synergy does not necessarily indicate a nonlinear suppressor variable, because it also occurs in linear Gaussian systems.Suppressor-variable systems are non-Gaussian and remain outside this paper’s scope.
  • 8 Discussion: The MMI PID enables separate computation of redundancy and synergy in neurophysiological datasets when a Gaussian approximation is valid.This provides a more detailed basis for information-theoretic analyses of brain variables and continuous time series.
Loading 1411.2832v2…