Source-linked AI summary
Measuring and modelling correlations in multiplex networks
Vincenzo Nicosia, Vito Latora
TL;DR
The paper addresses the limited study of correlations that are intrinsic to multiplex networks rather than single layers. It measures activity and degree correlations in real-world multiplexes, then proposes models to reproduce or assess them. The results show non-trivial cross-layer correlations and heterogeneous node involvement, while existing modeling assumptions about universal node activity across layers are too simplistic.
Problem
Correlations across layers of multiplex networks have not been investigated as thoroughly as correlations in single-layer networks, despite potentially richer relationships among node properties.
Method
The authors analyze five real-world multiplex networks, introduce measures of activity and inter-layer degree correlations, and develop models including simulated-annealing algorithms for tunable correlations.
Results
Real-world multiplex networks are sparse and exhibit heterogeneous node activity, non-trivial inter-layer degree correlations, and strong correlations in node presence and involvement across layers.
Takeaways & Limitations
A multiplex network contains structural information beyond the sum of its layers, so single-layer aggregation cannot fully describe the observed cross-layer patterns.
Takeaways & Limitations
Many existing multiplex models assume every node is active on every layer and that all layers contain the same number of nodes, assumptions the authors deem too simplistic for real-world systems.
Abstract
from arXiv · showhide
The interactions among the elementary components of many complex systems can be qualitatively different. Such systems are therefore naturally described in terms of multiplex or multi-layer networks, i.e. networks where each layer stands for a different type of interaction between the same set of nodes. There is today a growing interest in understanding when and why a description in terms of a multiplex network is necessary and more informative than a single-layer projection. Here, we contribute to this debate by presenting a comprehensive study of correlations in multiplex networks. Correlations in node properties, especially degree-degree correlations, have been thoroughly studied in single-layer networks. Here we extend this idea to investigate and characterize correlations between the different layers of a multiplex network. Such correlations are intrinsically multiplex, and we first study them empirically by constructing and analyzing several multiplex networks from the real-world. In particular, we introduce various measures to characterize correlations in the activity of the nodes and in their degree at the different layers, and between activities and degrees. We show that real-world networks exhibit indeed non-trivial multiplex correlations. For instance, we find cases where two layers of the same multiplex network are positively correlated in terms of node degrees, while other two layers are negatively correlated. We then focus on constructing synthetic multiplex networks, proposing a series of models to reproduce the correlations observed empirically and/or to assess their relevance.
I. INTRODUCTION
The paper argues that multiplex networks require studying correlations across layers because node properties and roles can differ substantially between interaction types. It develops measures and models for these correlations and shows that real-world multiplex systems contain non-trivial cross-layer structure that aggregation can obscure.
- Motivation and contribution: Multiplex networks extend single-layer correlation analysis by relating node properties, activities, and degrees across different layers.The paper studies correlations between the same or different properties of the same node across layers, alongside standard within-layer degree correlations.
- Empirical findings: Real-world multiplex networks show strong correlations in node presence and involvement across layers, requiring these patterns to be considered in modeling.The authors report heterogeneous node-activity patterns and inter-layer degree correlations.
- Motivation and contribution: The study combines empirical analysis of five biological, technological, and social multiplex networks with correlation measures and synthetic models.The networks range from hundreds of nodes and two interaction types in C.elegans to millions of nodes and dozens of layers in IMDb.
- Why multiplex representations matter: C.elegans and BIOGRID exhibit different layer densities, incomplete node overlap, and weak correspondence between node roles across layers.In C.elegans, average degree is 13.9 in the synaptic layer and 3.7 in the gap-junction layer; in BIOGRID, only 9,738 nodes are non-isolated in both layers.
- Why multiplex representations matter: In BIOGRID, the overlap of top-L degree-ranked nodes remains much smaller than 10% up to approximately L = 600.A hub in one interaction layer therefore has a small probability of also being a hub in the other layer, according to the reported ranking overlap.
- Why multiplex representations matter: The paper concludes that aggregating layers into a single network can discard structural patterns and may miss multiplex-specific phenomena in network dynamics.The conclusion links the observed non-trivial layer structure to dynamical processes reported in prior work, without claiming that aggregation always fails.
III. CORRELATIONS OF NODE ACTIVITY
Multiplex networks represent each node across multiple interaction layers, enabling correlations in node properties and layer involvement to be studied beyond single-layer projections.
- A multiplex network is represented by one adjacency matrix per layer, with nodes replicated across layers.Each matrix records connections for the same node set at a particular interaction type.
- Node properties include a multidimensional multi-degree because a node can have different degrees, or be isolated, across layers.The degree component k_i^[α] records node i's degree at layer α.
- Node and layer activity provide basic quantities for characterizing involvement across multiplex layers.Activity captures whether nodes participate in layers and how many nodes are active in each layer.
- Airline multiplexes show heterogeneous node-activity distributions, with power-law fits whose exponents range from 1.8 to 2.3.The distributions indicate that the number of layers in which nodes are active can fluctuate substantially.
A. Node activity
Node activity records both how many layers include a node and which layers they are, revealing heterogeneous and highly uneven participation patterns.
- A. Node activity: A node is active at layer α when it has at least one edge there, and its node-activity B_i counts active layers.The activity vector b_i records the layer-by-layer presence or absence of the node.
- A. Node activity: Airport node-activity distributions follow power laws with exponents δ between 1.8 and 2.4.For δ < 3.0, fluctuations in the number of active layers are unbounded as the number of layers grows.
- A. Node activity: Most authors and actors are active in only one or a few layers, while a few outliers participate in almost all layers.The paper calls such highly active outliers multi-active hubs.
- A. Node activity: Node activity is not strictly correlated with total incident edges, so high activity and high degree need not coincide.Nodes may have many edges in few layers or relatively few edges across nearly all layers.
- A. Node activity: APS and IMDb activity-vector frequencies are power laws, with exponents 1.53 and 1.2, respectively.The most frequent patterns involve nodes active on just one or two layers, whereas unusual patterns are rarer.
- A. Node activity: For a fixed number of active layers, activity-pattern frequencies are heterogeneous and usually power-law distributed.In IMDb, actors active in exactly two genres concentrate in a few genre pairs, while some pairs are extremely rare.
B. Layer activity
Layer activity measures how many nodes participate in each layer and how much two layers overlap, revealing strong heterogeneity in layer importance and activity patterns.
- B. Layer activity: Layer activity N^[α] is the number of active nodes in layer α, equal to the number of nonzero entries in its activity vector.This quantity supports comparisons of layer size across the multiplex.
- B. Layer activity: Continental airline networks show power-law layer-activity distributions, with many layers containing no more than 10 active nodes and some containing several hundred.The largest layers contain roughly 10% to 30% of all nodes.
- B. Layer activity: Randomly removing an airline layer usually causes minor disruption, but removing a particularly large layer might break the system apart.The consequence follows from the highly uneven distribution of active nodes across layers.
- B. Layer activity: Pairwise multiplexity Q_α,β measures the fraction of nodes active in both layers and increases with similarity between their activity patterns.Its values range from 0 to 1.
- B. Layer activity: Most layer pairs share relatively few active nodes: airline pairs commonly overlap by less than 1%, although some reach 20%.APS and IMDb pairwise multiplexities are also usually below 20%.
- B. Layer activity: Normalised Hamming distance H_α,β measures activity differences between two layers on a 0-to-1 scale.Values near zero indicate similar activity vectors, while values near one indicate maximal separation under the definition.
- B. Layer activity: Airline layer pairs generally have large Hamming distances, whereas about 1% of pairs across all systems have distances below 0.05.The latter corresponds to large overlaps in node activity.
IV. MODELS OF NODE AND LAYER ACTIVITY
The section tests whether real-world node and layer activity patterns can be reproduced by randomized or growth-based multiplex models. Comparisons show that activity correlations and heterogeneous activity distributions require models preserving more than marginal activity counts.
- Models of node and layer activity: Random activity models provide null models for assessing whether observed node-activity patterns are statistically distinctive.HM fixes layer activity while randomizing active-node assignments; MDM preserves each node’s activity count but destroys cross-layer activity correlations, while MSM also randomizes activity stochastically.
- Models of node and layer activity: LGM models multiplex growth through the addition of layers with fat-tailed activity and preferential activation of already active nodes.The parameter A gives inactive nodes a non-zero activation probability, while Bi(t) measures how many existing layers already contain node i.
- Model comparisons: LGM reproduces the European airlines network’s pairwise multiplexity distribution and node-activity distribution more accurately than the other tested models.It approximates both the shape and slope of P(Qα,β), and better approximates P(Bi) than HM; MDM and MSM preserve node-activity distributions by construction.
- Model comparisons: MDM and MSM generate stepwise-constant activity-vector rank distributions in APS and IMDb, unlike the heterogeneous real-world patterns.Each step corresponds to vectors with the same number of non-null entries, that is, the same node-activity Bi.
- Conclusions: Real-world node activity across layers is heterogeneous and often correlated across layers, so single-layer analysis or aggregation gives only a partial picture.The result supports preserving multiplex activity structure when modelling these systems.
V. CORRELATION BETWEEN ACTIVITY AND DEGREE
This section examines how node activity across layers relates to multidegree structure using overlapping degree and participation coefficient. Activity is positively associated with both measures, but substantial fluctuations remain.
- Measures: Overlapping degree and participation coefficient provide complementary summaries of a node’s multidegree and layer involvement.Overlapping degree counts total incident edges, whereas participation coefficient captures how evenly those edges are distributed across layers.
- Multiplex cartography: APS contains many mixed hubs, whereas almost all IMDb hubs are truly multiplex according to participation-coefficient patterns.The diagrams show varied multiplex roles even though overlapping degree and participation coefficient identify different connectivity aspects.
- Activity and overlapping degree: Node activity is positively correlated with overlapping degree, although activity fluctuates substantially among nodes with the same overlapping degree.Nodes with many links tend to be active on more layers, but the average relationship has large deviations.
- Activity and participation: Node activity is also positively correlated with participation coefficient, with large fluctuations around the average relationship.More uniformly distributed edges across layers are associated with activity on more layers.
VI. INTER-LAYER DEGREE CORRELATIONS
The paper measures inter-layer degree correlations with Pearson, Spearman, and Kendall coefficients, finding assortative correlations in APS but both positive and negative correlations in IMDb. The choice of coefficient depends on the system.
- Correlation coefficients: Pearson, Spearman, and Kendall coefficients quantify correlations between the degree sequences of two layers using nodes active on both layers.Restricting averages to jointly active nodes avoids bias from relatively small multiplexity.
- Empirical patterns: APS inter-layer degree correlations are exclusively assortative, whereas IMDb contains both positive and negative correlations.This contrast is consistent across the three coefficient types, although each coefficient shows somewhat different behavior.
- Empirical patterns: IMDb’s Adult and Talk-Show layers are negatively correlated with nearly all other layers but positively correlated with each other.The paper links these patterns to limited overlap between actors in those genres and most other movie categories.
- Interpretation: Disassortativity can occur in social networks when their interactions are represented as separate multiplex layers rather than aggregated.IMDb provides the paper’s example of this multiplex-specific pattern.
- Methodological caveat: The most appropriate correlation coefficient may depend on the system, motivating inter-layer degree correlation functions as a more accurate alternative.Spearman and Kendall can capture some nonlinear ranking relationships, but coefficient choice remains system-dependent.
B. Inter-layer correlation functions
Inter-layer correlation functions summarize how a node’s degree in one layer varies with its degree in another. Their slopes distinguish assortative from disassortative relationships and reveal heterogeneous patterns across multiplex systems.
- Multi-degree distribution: The multi-degree distribution P(k) jointly records the degrees a randomly chosen node has across all M layers.For APS and IMDb, its Zipf plot has a power-law tail with exponent near −1.0, but most multi-degree vectors occur very rarely.
- Pairwise inter-layer correlations: Pairwise joint and conditional distributions describe the degrees of the same node in two layers and provide the basis for inter-layer correlation functions.These distributions use k[α] and k[β] to characterize degree relationships between layers α and β.
- Inter-layer correlation functions: The function k[β](k[α]) gives the average degree in layer β for nodes having degree k[α] in layer α, reducing fluctuations relative to full probability distributions.An increasing function indicates assortative inter-layer correlations, while a decreasing function indicates disassortative correlations.
- Empirical patterns: C.elegans, BIOGRID, and APS show increasing inter-layer correlation functions, indicating assortative degree relationships across the displayed layer pairs.The functions are fitted with k[β](k[α]) ∼ (k[α])^μ.
- Empirical patterns: IMDb contains assortative, disassortative, and uncorrelated layer pairs, including positive Drama-Western, negative Adult-Western, and uncorrelated Drama–Game Show relationships.The inter-layer correlation network encodes the sign and magnitude of μ with blue and red edge weights.
VII. MODELS OF INTER-LAYER DEGREE CORRELATIONS
The paper proposes two simulated-annealing models to reproduce observed pairwise inter-layer degree correlations.
- VII. MODELS OF INTER-LAYER DEGREE CORRELATIONS: Two models reproduce observed pairwise inter-layer degree correlations using simulated annealing.One tunes the Spearman rank correlation coefficient, while the other prescribes a correlation exponent.
A. Model for ρ
The model tunes node assignments between two layers to reach a prescribed Spearman correlation coefficient while preserving one-to-one node correspondence.
- A. Model for ρ: The model couples two equal-sized layer graphs through a one-to-one node assignment.Each node in one layer is connected to exactly one node in the other layer.
- A. Model for ρ: The cost function F(S) = |ρS − ρ∗| measures deviation from the target Spearman correlation.The assignment is optimized by minimizing this absolute difference.
- A. Model for ρ: Simulated annealing swaps endpoints of randomly selected inter-layer edges to modify the assignment.Favorable swaps are accepted, while unfavorable swaps may be accepted with finite probability to explore configurations.
- A. Model for ρ: The algorithm stops when F(S) < ε because the discrete assignment problem may prevent exact achievement of ρ∗.To reduce bias from multiplex sparsity, it is usually run only on nodes active in both layers.
- A. Model for ρ: Synthetic networks preserve original node-activity vectors while reassigning active-node degrees to match selected inter-layer rank correlations.For consecutive layer pairs, the procedure iteratively fixes assignments while proceeding through the multiplex.
- A. Model for ρ: Algorithm 1 can tune the magnitude and sign of correlations between any pair of real-valued node properties, not only degrees.It compares rankings induced by the selected node properties.
B. Model for k[β](k[α])
The second model tunes inter-layer assignments to reproduce a prescribed power-law degree-correlation function and evaluates its agreement with real multiplex data.
- B. Model for k[β](k[α]): The desired inter-layer correlation function is q(k) = ak^µ, with µ obtained empirically and a determined during modeling.Here k is the degree in layer α and q is the degree in layer β.
- B. Model for k[β](k[α]): Algorithm 2 swaps endpoints to minimize the difference between the actual q(k) and the desired power-law form.Favorable swaps are always accepted, while unfavorable swaps are accepted with exponentially decreasing probability.
- B. Model for k[β](k[α]): The coefficient a is adaptively updated from the best power-law fit as the algorithm proceeds.Updates occur every ta algorithm steps, where ta is user-set.
- B. Model for k[β](k[α]): The synthetic APS multiplex has a qualitatively similar µ distribution but can differ substantially in the actual exponent values.Because Algorithm 2 sets µ for only M − 1 layer pairs, the mismatch suggests correlations are not solely pairwise.
VIII. CONCLUSIONS
The study finds non-trivial correlations and activity patterns across real-world multiplex layers, supporting multiplex descriptions beyond aggregated single-layer networks.
- VIII. CONCLUSIONS: Real-world multiplexes are sparse, with few nodes active on more than one layer, yet their cross-layer involvement patterns are correlated.The evidence comes from biological, technological, and social systems spanning a wide range of sizes.
- VIII. CONCLUSIONS: Non-trivial multiplex patterns mean that aggregating layers into a single-layer network cannot fully describe the system.The paper links these patterns to dynamical processes that may be absent from single-layer projections.
- VIII. CONCLUSIONS: Many existing multiplex models assume every node is active in every layer and all layers contain the same number of nodes.The paper identifies these assumptions as too simplistic for real-world multiplex systems.
APPENDIX
The appendix describes the real-world data sets used to construct multiplex networks across biological, transportation, collaboration, and movie domains. It records layer definitions, network sizes, and selected statistical properties for these systems.
- Data sets: The study uses data sets whose multiplex layers represent distinct interaction types, while focusing on node-activity correlations rather than edge weights.Most data sets also permit edge weights measuring interaction strength, but the current work emphasizes node activity.
- Biological networks: The C. elegans neural network contains 281 neurons and around two thousand connections.It is described as the only fully mapped brain of a living organism.
- Biological networks: BIOGRID is modeled with two undirected, unweighted layers for physical and genetic protein interactions, using 54,549 nodes.The source database contains around 500,000 registered interactions across more than 40 species.
- Transportation networks: OpenFlight supplies continental airplane multiplex networks in which airports are nodes and layers capture finer-grained aerial-route information.Six continental networks are characterized by the number of layers and the power-law exponent η for non-isolated nodes per layer.
- Collaboration networks: The APS coauthorship multiplex assigns layers to collaborations among papers in 10 high-level PACS categories.Each layer has up to around 79,000 active nodes, with density varying across layers.
- Movie collaboration network: The IMDb multiplex represents actors as nodes, links co-acting pairs, and uses movie categories as genre-specific layers.The data set covers several million movies belonging to 30 genres, while the corresponding appendix table reports 28 layers.