Source-linked AI summary
Quantifying High-order Interdependencies via Multivariate Extensions of the Mutual Information
Fernando Rosas, Pedro A. M. Mediano, Michael Gastpar, Henrik J. Jensen
TL;DR
High-order statistical synergy is difficult to quantify with existing multivariate information measures, despite its relevance to complex systems. The paper introduces the O-information from dual perspectives of shared randomness and collective constraints, and uses it to distinguish redundancy- from synergy-dominated systems. The framework is applied to Baroque music, where Bach’s chorales are found to be strongly synergistic while Corelli’s pieces are redundant-dominated.
Problem
Quantifying high-order interdependencies and statistical synergy remains difficult, while existing measures have limited interpretability or theoretical scope.
Method
The paper defines the O-information as the difference between collective constraints and shared randomness, using a symmetric, model-agnostic framework for multivariate interdependencies.
Results
The O-information distinguishes redundancy- and synergy-dominated systems and identifies Bach’s chorales as strongly synergistic in a Baroque music case study.
Takeaways & Limitations
O-information complements measures of interdependency strength by revealing whether multivariate correlations are predominantly redundant or synergistic.
Abstract
from arXiv · showhide
This article introduces a model-agnostic approach to study statistical synergy, a form of emergence in which patterns at large scales are not traceable from lower scales. Our framework leverages various multivariate extensions of Shannon's mutual information, and introduces the O-information as a metric capable of characterising synergy- and redundancy-dominated systems. We develop key analytical properties of the O-information, and study how it relates to other metrics of high-order interactions from the statistical mechanics and neuroscience literature. Finally, as a proof of concept, we use the proposed framework to explore the relevance of statistical synergy in Baroque music scores.
I. INTRODUCTION
High-order interdependencies are important in complex systems but difficult to quantify with existing multivariate information measures. The paper addresses this challenge by introducing the O-information as a symmetric metric that distinguishes redundancy- and synergy-dominated systems.
- Motivation: High-order interdependencies structure complex systems, yet quantifying synergy is challenging when systems contain relatively few parts.The motivation spans brain activity, econometric indices, and gene interactions.
- Limitations of existing approaches: Existing metrics often have ad hoc definitions and few theoretical guarantees, while PID lacks a settled computation and scales super-exponentially.These limitations restrict the practical use of several neuroscience- and information-theory-based approaches.
- Proposed framework: The O-information compares multivariate interdependencies as shared randomness and collective constraints, selecting the more parsimonious perspective.It coincides with interaction information for three variables and provides a more meaningful extension for larger systems.
- Proposed framework: The O-information distinguishes redundancy-dominated systems from synergy-dominated systems with high-order patterns absent from low-order marginals.It is symmetric and does not require separating variables into predictors and targets.
- Paper scope: The paper concludes by outlining its analytical framework and applying it to Baroque music scores as a case study.The study also compares the O-information with other high-order interaction metrics.
2. Shared randomness
The paper decomposes multivariate information into collective constraints and shared randomness, then defines the O-information as their balance. Its sign indicates whether redundancy or synergy dominates, while its properties make the quantity intrinsic and focused on beyond-pairwise interactions.
- Shared randomness: Residual entropy measures information accessible only by measuring a specific variable, while binding entropy measures randomness shared by multiple variables.Binding entropy is also called dual total correlation, excess entropy, or binding information.
- O-information: The O-information compares collective constraints with shared randomness to identify which perspective gives a more parsimonious description of interdependencies.Positive values favor a shared-randomness description, whereas negative values favor collective constraints.
- Relation to interaction information: For three variables, the O-information equals interaction information, which measures the difference between synergy and redundancy.Interaction information is positive for redundancy-dominated examples and negative for synergy-dominated xor structures.
- Basic properties: The O-information is permutation-invariant and vanishes for two variables, so it captures intrinsic interactions beyond pairwise relationships.For more than three variables, it generally differs from co-information.
- Interpretation: Positive O-information defines redundancy-dominated systems, whereas negative O-information defines synergy-dominated systems.This sign convention is the paper’s operational criterion for classifying the dominant interaction characteristic.
B. Information decompositions
The paper represents set partitions of variables as a refinement lattice and organizes their relationships into a directed acyclic graph. Paths through this graph encode sequences of binary partitions.
- Partition lattice: A partition divides the variable indices into disjoint cells, and all such partitions form a lattice ordered by refinement.The unique source is the coarsest partition containing all indices in one cell.
- Partition graph: The directed acyclic graph has partitions as nodes and edges linking each partition to a partition that covers it.A path is an ordered sequence of covering relations between two partitions.
- Three-variable case: For three variables, the double diamond displays the possible binary-partition paths from the source to the two sink nodes.Each path corresponds to a decomposition of either C(X3) or B(X3).
2. Lattice decompositions of C(Xn) and B(Xn)
The lattice construction assigns information-theoretic weights to partition refinements, yielding path decompositions of total correlation and binding entropy. For three variables, these decompositions recover equivalent expressions involving mutual and conditional mutual information.
- Weight construction: Node and edge weights on the partition DAG are constructed from entropy, residual entropy, mutual information, and conditional mutual information.The entropy-based and residual-based constructions place finer partitions in opposite vertical directions.
- Path decompositions: Every path from the source partition to a sink provides a decomposition of C(Xn) or B(Xn).The two weight systems correspond respectively to mutual-information and conditional-mutual-information edge terms.
- Three-variable case: For n = 3, the graph has three source-to-sink paths, each isolating one variable or pair at the intermediate partition.All three paths yield the same C(X3) and B(X3) decompositions.
- Three-variable case: For three variables, B(X3) can be written as I(Xi; Xj, Xk) + I(Xj; Xk|Xi).The indices {i, j, k} are a permutation of {1, 2, 3}.
3. Lattice decomposition of Ω(Xn)
The paper uses partition-lattice paths to decompose the O-information into interaction-information terms, including decompositions valid under any variable ordering. Assembly paths provide an explicit sequential construction, while the lattice’s super-exponential growth motivates heuristic exploration.
- Lattice decomposition: The proposed edge weights can be negative and provide a decomposition of the O-information along every path from the source to sink partition.This extends the path-based decomposition framework beyond nonnegative weights.
- Lattice decomposition: The O-information can always be expressed as a sum of interaction-information terms involving three sets of variables.Consequently, it retains triple interaction information’s ability to reflect the balance between synergies and redundancies for systems of any size.
- Computational scope: Partition lattices grow super-exponentially with system size, so heuristic methods are needed to explore them.Assembly paths form a particularly useful sub-family of the complete set of source-to-sink paths.
- Assembly paths: Assembly paths sequentially separate variables or, viewed backward, assemble the system by connecting variables one at a time.For example, the path separates X_n, then X_{n−1}, and so on; its reverse connects X_1 and X_2, then X_3 to X_2.
- Assembly paths: Corollary 1 gives path-based decompositions of total correlation, binding entropy, and O-information using the variables X_k and X^k.These decompositions are constructed from the newly assigned edge weights.
- Ordering invariance: The decompositions remain valid under any relabeling or ordering of the system’s variables because they arise from the lattice construction.This ordering invariance is used in subsequent sections.
A. Characterising extreme values of Ω
The O-information has analytically characterized extremes that distinguish copy-like redundancy from xor-like synergy, while its sign and magnitude constrain interactions across subsystem scales.
- Extreme distributions: For binary systems, Ω(Xn) reaches n − 2 exactly for an n-bit copy and 2 − n exactly for an n-bit xor.The n-bit copy consists of identical fair-coin variables; the n-bit xor uses independent fair coins with the final variable equal to their modulo-2 sum.
- Extreme distributions: For equal-alphabet variables, the maximum Ω(Xn) = (n −2) log m occurs when variables copy one another, while the minimum Ω(Xn) = (2−n) log m occurs for an independent-uniform modulo-m sum.These bounds are attained by the corresponding redundant and synergistic constructions.
- Extreme distributions: Unlike interaction information, Ω remains consistently negative and decreases with n for an n-bit xor, avoiding interaction information’s oscillation between −1 and +1.The paper identifies the copy and xor distributions as unique extremes for Ω.
- Constraints across scales: The sign of Ω determines whether interaction constraints are lower or upper bounds, while larger |Ω| constrains progressively smaller groups of variables.Positive Ω forces dependencies among sufficiently large groups; negative Ω upper-bounds their correlation strength, with stronger absolute values affecting smaller scales.
- Constraints across scales: Fixing the interaction strength of one m-variable subset narrows Ω’s range from 2(n−2) to 2(n−2)−(m−1), and both bounds are tight when |Ω| ≥(n −m + 1) log |X|.For a binary symmetric channel, the pairwise constraint is C(X2) = I(X1; X2) = 1 −H(η), with the bounds illustrated in Figure 4.
- Independent and overlapping subsystems: Ω is additive across independent subsystems, so disjoint pairwise interactions suffice for Ω = 0, although redundant and synergistic subsystems can also cancel to zero.Overlapping pairwise interdependencies cannot be factorized in this way and may yield either positive or negative Ω.
A. High-order interactions in statistical mechanics
The section examines high-order spin interactions and compares TSE complexity with O-information as measures of multivariate dependence. High-order interactions make O-information more negative, while TSE tracks overall integration but conflates redundancy and synergy.
- High-order interactions in statistical mechanics: Systems with k-th order interactions were generated using random Hamiltonians with independently standard-normal interaction strengths and β = 0.1.The analysis tested how O-information changes with interaction order k.
- High-order interactions in statistical mechanics: Ω is usually near zero for k = 2 and becomes negative as k increases, matching traditional statistical-physics intuitions about high-order interactions.Figure 5 reports the same trend for ensembles of n = 5 spins with randomly generated Hamiltonians.
- Complexity and integration: The sum of total correlation and binding entropy provides a very accurate approximation of TSE complexity, with correlation consistently above 0.97 on random distributions.The approximation outperformed other proposed TSE approximations in the reported Monte Carlo simulations.
- Complexity and integration: TSE complexity has the same value for the extreme 3-bit copy and xor distributions, so it conflates redundancy with synergy.The comparison uses a linear mixture between copy and xor distributions; results are qualitatively similar for larger systems.
- Complexity and integration: TSE measures overall interdependency strength, whereas O-information indicates whether multivariate correlations are predominantly redundant or synergistic.The two quantities are presented as complementary aspects of multivariate dependence.
VI. CASE STUDY: BAROQUE MUSIC SCORES
The paper demonstrates its framework on Baroque musical scores, using four-part works by Bach and Corelli. The case study analyzes multivariate statistics of these scores as a brief practical application of O-information.
- VI. CASE STUDY: BAROQUE MUSIC SCORES: The study is presented as a brief demonstration of O-information’s value for practical data analysis.The section describes the data procedure and then reports the case-study results.
- VI. CASE STUDY: BAROQUE MUSIC SCORES: The case study examines four-voice Bach chorales and Corelli’s Opus 1 and 3–6, all from the Baroque period.The repertoire is characterized by elaborate counterpoint between melodic lines.
- VI. CASE STUDY: BAROQUE MUSIC SCORES: The analysis uses electronic scores with four melodic lines, selecting Major-mode pieces and preprocessing them with Music21.Bach scores provide soprano, alto, tenor, and bass; Corelli scores provide four string instruments.
2. Research questions and tools
The study measures harmonic interdependencies in Baroque scores using entropy, global O-information, and local O-information. Bach’s chorales are synergy-dominated, whereas Corelli’s pieces are redundant, with compositional practices offering a possible explanation.
- The analysis asks whether simultaneously played notes are redundant or synergistic, focusing exclusively on harmony and chords.
- Entropy measures each voice’s harmonic richness, while O-information determines the ensemble’s dominant multivariate behavior.The notes are modeled as a four-dimensional random vector over an alphabet of 13 notes, and computations use muts.
- Bach’s chorales have negative O-information, and every local O-information term is negative, indicating synergy beyond pairwise dependencies.The pairwise dependence between each pair of voices is comparatively smaller than the group-level dependencies.
- Corelli’s pieces have positive O-information, with positive local values for every pair except violins 1 and 2.The strongest local O-information occurs between viola and cello, indicating high redundancy between those parts.
- Corelli’s viola–cello redundancy is consistent with the basso continuo practice in scores originally written for two soloists and a bass line.The bass line could be interpreted by different bass instruments, corresponding here to viola and cello.
- The framework defines O-information as the difference between collective constraints and shared randomness, capturing the net balance between synergy and redundancy.It is symmetric, sums triple interaction informations, and reaches its maximum and minimum for an n-bit copy and xor, respectively.
Appendix A: Compatibility between Ωand prior work
The appendix relates O-information to the prior Ψ metric and formalizes their compatibility. Random binary-system calculations show good agreement, while O-information offers more theoretical properties with fewer terms.
- Convexity of Ψ(k) corresponds to synergy, whereas concavity corresponds to redundancy in the interdependency structure.Convexity indicates relatively independent small scales and correlated large scales; concavity indicates highly correlated small groups.
- The comparison with Ψ quantifies its convexity or concavity by measuring distance from the straight line joining Ψ(1) and Ψ(n).
- O-information and Ψ show good agreement in randomly generated binary systems of different sizes.The result confirms the analytic relationship between the metrics.
- O-information formalizes the intuitive notions behind Ψ while possessing more theoretical properties and requiring fewer calculated terms.
- The appendix proves monotonic behavior of R across ordered partitions by reducing the argument to elementary refinement steps.The proof follows paths through covering relationships between partitions.
Appendix E: Proof of Lemma 3
The appendix proves Lemma 3 by applying bounds on pairwise conditional and triple interaction information. Equality cases identify n-bit copies as the upper extremum and n-variate xor structures as the lower extremum.
- The proof applies bounds on conditional and triple interaction information to the relevant expressions for C, B, and O-information.The bounds are stated for all variable triples and are tight by construction.
- For an n-bit copy, C(Xn)=n−1 and B(Xn)=1, attaining the upper bound for O-information.The converse shows that equality forces every pair of variables to be identical Bernoulli variables.
- For an n-bit xor, C(Xn)=1 and B(Xn)=n−1, attaining the lower bound for O-information.The equality conditions imply joint independence of the relevant variables and deterministic functional relationships characteristic of xor.
- The proof of an intermediate O-information bound begins from Ω(Xn)=C(Xn−1)−B(Xn−1|Xn)≤C(Xn−1).The remaining inequality follows by applying Lemma 6.