Source-linked AI summary
Statistically validated networks in bipartite complex systems
Michele Tumminello, Salvatore Miccichè, Fabrizio Lillo, Jyrki Piilo, Rosario N. Mantegna
TL;DR
Projected networks can be difficult to interpret when heterogeneous relationships produce multiple link types and unclear community structure. The paper introduces statistically validated networks and applies minimalist community-detection and cluster-characterization approaches, finding refined partitioning and a network with different kinds of links.
Problem
The multi-link statistically validated network is difficult to analyze because it presents different kinds of links, making community detection non-obvious, while detected clusters still require interpretation.
Method
The paper uses statistically validated links with full control against a random reference and applies minimalist community detection plus attribute-based cluster characterization.
Results
The resulting network presents different kinds of links, and applying Infomap refines the natural partitioning of the network.
Takeaways & Limitations
Statistical validation supports informative analysis of heterogeneous network structure while retaining distinctions among link types and characterizing clusters through their attributes.
Abstract
from arXiv · showhide
Many complex systems present an intrinsic bipartite nature and are often described and modeled in terms of networks [1-5]. Examples include movies and actors [1, 2, 4], authors and scientific papers [6-9], email accounts and emails [10], plants and animals that pollinate them [11, 12]. Bipartite networks are often very heterogeneous in the number of relationships that the elements of one set establish with the elements of the other set. When one constructs a projected network with nodes from only one set, the system heterogeneity makes it very difficult to identify preferential links between the elements. Here we introduce an unsupervised method to statistically validate each link of the projected network against a null hypothesis taking into account the heterogeneity of the system. We apply our method to three different systems, namely the set of clusters of orthologous genes (COG) in completely sequenced genomes [13, 14], a set of daily returns of 500 US financial stocks, and the set of world movies of the IMDb database [15]. In all these systems, both different in size and level of heterogeneity, we find that our method is able to detect network structures which are informative about the system and are not simply expression of its heterogeneity. Specifically, our method (i) identifies the preferential relationships between the elements, (ii) naturally highlights the clustered structure of investigated systems, and (iii) allows to classify links according to the type of statistically validated relationships between the connected nodes.
I. CLUSTER DETECTION AND CHARACTERIZATION
The paper uses community detection to partition statistically validated and adjacency networks, then compares detected clusters with prior classifications to interpret them. Infomap identifies communities through probability flow, but repeated runs can produce slightly different partitions.
- I. CLUSTER DETECTION AND CHARACTERIZATION: Infomap partitions network vertices by minimizing the description length of random-walk flows through the network.The method applies to weighted and unweighted networks.
- I. CLUSTER DETECTION AND CHARACTERIZATION: Independent Infomap applications to the same network may produce slightly different vertex partitions.The authors run Infomap 103 times and select the best partition for each network.
- I. CLUSTER DETECTION AND CHARACTERIZATION: The partition depends on whether the network is weighted or unweighted and on how link weights are selected.The paper therefore discusses cluster detection separately for each network.
- I. CLUSTER DETECTION AND CHARACTERIZATION: Cluster interpretation compares Infomap partitions with a priori classifications such as movie genres or stock economic sectors.This comparison addresses the problem of interpreting detected clusters.
A. Cluster characterization
Cluster characterization tests whether system attributes are over-expressed in detected clusters relative to random selection. Hypergeometric p-values and Bonferroni correction account for cluster composition and multiple comparisons.
- A. Cluster characterization: For each cluster C and attribute Q, the method tests whether the observed count N_C,Q exceeds the count expected from random selection.The system contains N elements, and each cluster contains N_C elements.
- A. Cluster characterization: The hypergeometric distribution H(X|N, N_C, N_Q) models the number of attribute-Q elements expected in a randomly selected cluster.N_Q is the total number of system elements with attribute Q.
- A. Cluster characterization: An attribute is considered over-expressed when its p-value falls below the statistical threshold p_b.The p-value is associated with the observed number N_C,Q of attribute-Q elements in cluster C.
- A. Cluster characterization: Bonferroni correction sets the threshold to p_b = 0.01/N_A when testing all N_A attributes for each cluster.The correction controls the threshold across the multiple attribute comparisons.
A. Cluster detection
The paper constructs a multi-link statistically validated stock network and uses a pragmatic weighting scheme to detect communities. Correlated and anti-correlated relationships receive different treatment during community search.
- A. Cluster detection: The 500-stock network is a multi-link network representing different kinds of statistically validated relationships.The authors describe it as a new kind of network and note that links may have different natures.
- A. Cluster detection: Community detection removes anti-correlated evolution links L4 and L5, equivalently assigning them zero weight.The procedure distinguishes co-occurrences of correlated from anti-correlated evolution.
- A. Cluster detection: Remaining links are weighted according to whether their statistical validation is single or twofold.Twofold links receive weight 2, while onefold links receive weight 1.
- A. Cluster detection: The weighting and partitioning approach is pragmatic and heuristic, while a theoretically based method for multi-link networks remains desirable.The authors identify this as a limitation of the approach.
B. Cluster characterization
Infomap finds structured clusters in the Bonferroni stock network, and economic-sector characterization shows that most clusters correspond to specific subsectors. The categorized-return representation further indicates persistent interrelations among stock groups.
- B. Cluster characterization: The Bonferroni network contains 349 non-isolated stocks characterized by 81 subsectors, while the FDR network contains 494 non-isolated stocks characterized by 96 subsectors.Some stocks lack subsector information and retain their broader sector classification.
- B. Cluster characterization: 37 Bonferroni-network clusters contain 2–48 stocks, and subsector analysis finds 41 over-expressions across 31 distinct clusters.The clusters are shown in Fig. 4.
- B. Cluster characterization: The seven largest illustrated clusters correspond to specific economic subsectors including real estate operations, technology, energy, utilities, and financial insurance.The listed labels distinguish the detected clusters by economic subsector.
- B. Cluster characterization: Most clusters are described by a single economic subsector, with 1.32 over-expressed subsectors per cluster on average.The analysis tests stock classifications at both sector and subsector levels.
- B. Cluster characterization: Figure 5 categorizes stock excess returns over time, with green, red, and black indicating up, down, and null states.Stocks are ordered by clusters, which are separated by blue lines.
- B. Cluster characterization: The Services–real-estate and Technology clusters show opposite excess-return movements, while co-occurring states persist across the investigated period.The paper reports continuously observed interrelations among different stock groups.
III. NETWORK OF MOVIES
The movie network contains a large adjacency-network component that Infomap partitions into many clusters, while the Bonferroni network is much more fragmented. These partitions show some homogeneity in production country, language, and genre.
- The largest adjacency-network component contains 77,193 movies.
- Adjacency-network clusters show some homogeneity in production country, language, and genre.
- Infomap partitions the adjacency network into 3,967 clusters, with sizes decreasing from 1,478 to 2 movies.
- The Bonferroni network lacks a giant connected component; its largest component contains only 13% of linked movies.
- Infomap detects 2,782 Bonferroni clusters ranging from 577 to 2 movies, while the network has 2,456 connected components.
A. Community detection in weighted movie networks
The movie projection weights co-appearance links by the productivity of shared actors, while statistical validation controls links against a heterogeneity-aware null hypothesis. Weighted community detection refines the movie partition, and validated networks yield similar results to their unweighted counterparts.
- The method statistically validates links against a null hypothesis that accounts for heterogeneous movie relationship counts.
- Because actors play different numbers of movies, the procedure uses a heuristic based on Reference 6 rather than a rigorous computationally feasible treatment of this heterogeneity.
- A weighted link between movies is constructed from their shared actors, with each actor’s contribution depending on the total number of movies they played.
- Weighted statistically validated networks produce results very similar to the corresponding unweighted networks.
B. How link weights affect the community structure of networks
The study compares weighted and unweighted community partitions using normalized mutual information. Accounting for link weights makes the partitions increasingly similar in statistically validated networks, especially under Bonferroni correction.
- Normalized mutual information quantifies the difference between partitions obtained with and without link weights.
- Mutual information equals 1 for identical partitions and reaches 0.798, 0.913, and 0.976 for adjacency, FDR, and Bonferroni networks, respectively.
- Mutual information increases when statistically validated networks are considered.
- The Bonferroni network’s mutual-information value is very close to 1 under the most restrictive statistical requirements.
- The comparison includes adjacency, FDR, and Bonferroni movie networks in both unweighted and weighted forms.
C. Cluster size and inclusiveness
Cluster-size profiles differ between the adjacency and statistically validated movie networks: FDR and Bonferroni cluster sizes decay approximately as a power law. Bonferroni clusters typically correspond to strongly interconnected parts of larger FDR clusters, but this inclusiveness is not strict.
- Infomap cluster sizes decrease with rank differently in adjacency networks than in statistically validated networks.
- FDR and Bonferroni networks exhibit cluster-size decay versus rank that is well approximated by a power law.
- Bonferroni clusters usually correspond to the strongest interconnected parts of larger FDR clusters.
- FDR clusters generally correspond to sets of movies present in larger clusters of the weighted adjacency network.
- The observed inclusiveness is typical for most detected clusters but is not a strict inclusive relation.
D. Cluster characterization
The study characterizes Infomap clusters in weighted movie networks by testing over-expression across four IMDb classifications. Across network constructions, stricter statistically validated networks yield smaller clusters with more specific characterizations.
- Cluster characterization: Clusters are characterized by separately testing over-expression for production country, language, genre, and filming location.IMDb provides these four movie attributes, with unavailable information recorded as “not available.”
- Cluster characterization: Different networks show different over-expression profiles, with specificity higher for smaller clusters, especially in genre and filming location.Country and language over-expression is generally specific, with exceptions involving former Yugoslavia and the Soviet Union after country splits.
- Cluster characterization: Average over-expressions per cluster decrease from the weighted adjacency network to FDR and then Bonferroni networks, except for language between FDR and Bonferroni.For genre, the averages are 1.41, 1.34, and 1.33, respectively.
- Cluster characterization: The decrease is more pronounced for genre and filming-location characterizations, indicating greater specificity in cluster characterization.The networks are analyzed to illustrate changes in sensitivity and specificity across weighted constructions.
1. Largest cluster of the weighted adjacency network and the overlapping FDR clusters
The case studies compare movie communities from adjacency, FDR, and Bonferroni weighted networks. Statistically validated networks split broad movie clusters into smaller, more homogeneous groups, trading sensitivity for specificity.
- Largest cluster and overlapping FDR clusters: The largest adjacency-network cluster contains 6,523 movies, including 6,433 with USA as the over-expressed production country.Its over-expressed attributes include English and multiple genres and filming locations.
- Largest cluster and overlapping FDR clusters: FDR partitioning produces clusters more than one order of magnitude smaller than the adjacency cluster and more refined for USA movies.These clusters are usually characterized by USA production and English language, with genre and filming location providing major additional characterization.
- Indian movie cluster: Bonferroni clusters show still higher specificity, with one language characterization per cluster, although filming-location characterization is limited by missing Indian IMDb data.“Not available” is the most frequent filming-location over-expression in this case.
- Indian movie cluster: Across the Indian case, statistically validated networks describe communities that are smaller and more homogeneous than adjacency-network communities.This conclusion is stated for the considered movie classifications and the FDR and Bonferroni weighted networks.