Source-linked AI summary

Community detection for correlation matrices

Mel MacMahon, Diego Garlaschelli

arXiv:1311.1924v3physics.data-anphysics.soc-phq-fin.PMq-fin.RMq-fin.ST

TL;DR

The paper addresses how to identify emergent mesoscopic groups in multivariate time series when existing correlation filters lose information and network community detection uses inconsistent null models. It redefines modularity with random-matrix-theory-based null models and applies the resulting methods to financial data, revealing internally correlated, mutually anti-correlated communities and nested structure with hard cores and soft peripheries.

  • Problem

    Existing correlation filters are not explicitly designed to identify mesoscopic modules, while naively applying network community detection to correlation matrices introduces null-model inconsistency and bias.

  • Method

    The paper adapts modularity-based community detection by defining correlation-specific null models from random matrix theory, with extensions for global dependencies, multiresolution structure, and multifrequency behavior.

  • Results

    Financial applications identify communities that are internally positively correlated and mutually anti-correlated after noisy and market-wide dependencies are removed, including groups not reducible to sector taxonomy.

  • Takeaways & Limitations

    The methods provide a framework for extracting mesoscopic structural information from multivariate time series and reveal cross-sector market organization beyond standard industry classifications.

Abstract

from arXiv · show

A challenging problem in the study of complex systems is that of resolving, without prior information, the emergent, mesoscopic organization determined by groups of units whose dynamical activity is more strongly correlated internally than with the rest of the system. The existing techniques to filter correlations are not explicitly oriented towards identifying such modules and can suffer from an unavoidable information loss. A promising alternative is that of employing community detection techniques developed in network theory. Unfortunately, this approach has focused predominantly on replacing network data with correlation matrices, a procedure that tends to be intrinsically biased due to its inconsistency with the null hypotheses underlying the existing algorithms. Here we introduce, via a consistent redefinition of null models based on random matrix theory, the appropriate correlation-based counterparts of the most popular community detection techniques. Our methods can filter out both unit-specific noise and system-wide dependencies, and the resulting communities are internally correlated and mutually anti-correlated. We also implement multiresolution and multifrequency approaches revealing hierarchically nested sub-communities with `hard' cores and `soft' peripheries. We apply our techniques to several financial time series and identify mesoscopic groups of stocks which are irreducible to a standard, sectorial taxonomy, detect `soft stocks' that alternate between communities, and discuss implications for portfolio optimization and risk management.

I. INTRODUCTION

The paper addresses the open problem of identifying emergent mesoscopic modules in multivariate time series without prior information. It adapts community detection to correlation data by replacing network null models with random-matrix-theory-consistent formulations.

  • Mesoscopic organization consists of groups whose activity is more correlated internally than with other groups, but these modules are not generally evident from individual similarities.
  • Threshold and geometric filtering methods use arbitrary criteria and can discard information needed to identify emergent modules through iterative partitioning.
  • Naively replacing network data with cross-correlation matrices creates null-model inconsistencies and partition-search bias, especially when true communities differ in size.
  • The proposed approach redefines modularity using null models appropriate for time series and dictated by random matrix theory.
  • The paper adapts three modularity-based algorithms and extends them to hierarchically nested subcommunities and hard cores versus soft peripheries.
  • Financial applications identify cross-sector stock correlations and communities that are internally correlated and mutually anti-correlated after removing noisy and market-wide dependencies.

A. Asset Graphs

Asset Graphs, MSTs, PMFGs, and RMT each filter correlation data but are not by themselves designed to identify emergent mesoscopic modules. Their information retention and structural assumptions create corresponding limitations.

  • Asset Graphs: Asset Graphs retain correlations above a global threshold, making them robust to noise but potentially discarding weaker within-module correlations needed to detect mesoscopic groups.
  • Minimal Spanning Trees: MSTs retain N −1 correlations and produce hierarchical clustering, but assume the original correlations are well approximated by the filtered ultrametric structure.
  • Minimal Spanning Trees: MSTs discard weaker correlations and become progressively less faithful to the original matrix at higher taxonomic levels.
  • Planar Maximally Filtered Graphs: PMFGs retain more correlations than MSTs while imposing planarity, an approximating structure with no obvious natural basis for stocks or other time series.
  • Planar Maximally Filtered Graphs: MST- and PMFG-based hierarchies seek an approximating structure rather than optimizing groups whose internal correlations exceed their cross-group correlations.
  • Random Matrix Theory: RMT treats eigenvalues within the Marcenko-Pastur range as mostly random and larger eigenvalues as meaningful structure, but filtering alone does not resolve mesoscopic organization.
  • Random Matrix Theory: For the S&P 500 example, T = 2500 returns from N = 445 stocks produce an empirical maximum eigenvalue of about 175 versus a random-matrix maximum of approximately 2.

III. COMMUNITY DETECTION IN GRAPHS AND ITS INCONSISTENCY WITH CORRELATION MATRICES

Network community detection identifies dense clusters by maximizing modularity relative to a null model. The paper shows that standard network modularity is inconsistent with correlation matrices and motivates correlation-specific null models and multiresolution methods.

  • A. Community detection in networks: Community detection seeks relatively dense node clusters, and modularity optimization evaluates partitions against a community-free null model.
  • A. Community detection in networks: Modularity assigns nodes to the same community when observed connections exceed null-model expectations based on individual network characteristics.
  • A. Community detection in networks: The network null model controls degree or strength heterogeneity by comparing observed edges with expected edges under otherwise random topology.
  • A. Community detection in networks: Standard modularity cannot resolve communities below a typical scale, motivating multiresolution extensions with a resolution parameter.
  • B. Correlation-matrix inconsistency: The paper replaces network benchmarks with null models consistent with correlation-matrix properties and develops a theoretically consistent multiresolution approach without ad hoc parameters.

B. The inconsistency of modularity for cross-correlation matrices

Applying network modularity directly to cross-correlation matrices uses a null model inconsistent with correlation-matrix constraints, producing biased community searches. Appropriate correlation null models must account for positive semidefiniteness and realistic eigenvalue structure.

  • Replacing network data with correlations makes the null model depend on each series’ correlation with the aggregate signal Xtot, rather than only on direct pairwise correlation.Xtot is the total increment series and is not standardized even when individual series are standardized.
  • Under independence, the appropriate expected correlation matrix is the identity matrix I, while other realistic correlation-based null models are introduced later.
  • Network modularity assumes matrices constrained by row and column sums, whereas correlation matrices also require non-negative eigenvalues and therefore obey different structural constraints.The network null model permits matrices that are not valid correlation matrices.
  • The naïve null model has an extremely simple spectrum, with one nonzero eigenvalue and N−1 zero eigenvalues, irrespective of the original data.This spectrum is unlike the richer eigenvalue distributions expected for realistic correlation matrices.
  • Adding a resolution parameter does not repair the mismatch: the eigenvalues remain scaled versions of the same two-value spectrum and cannot reduce to the independent-series null model.The paper therefore motivates a different implementation of multiresolution community detection for correlations.

C. The bias produced by the na¨ıve approach

The naïve modularity produces a community-size-dependent bias in its expected correlations. This bias can make large communities difficult to detect and cannot generally be removed by the resolution parameter.

  • Equation (32) cannot produce off-diagonal zeros, so its expectation cannot equal the correct benchmark expectation for independent communities.With equally sized communities, it instead gives ⟨Cij⟩naive = φ/c for all i,j.
  • For equally sized communities, the naïve off-diagonal expectation has a single peak and zero standard deviation, making its constant offset irrelevant to modularity maximization apart from diagonal issues.
  • The standard deviation measures absolute bias and scales linearly with φ, whereas the coefficient of variation measures relative bias and is independent of φ.
  • The relative bias rises with community-size heterogeneity up to approximately two, then decreases in an extreme regime dominated by one giant community and very small communities.
  • The naïve expectation increases for pairs within larger communities, progressively biasing the community search toward community size.For the largest community, the expected internal correlation can exceed correlations between any pair of communities.
  • Reducing the absolute bias from around 0.3 to approximately 0.01 requires φ around 0.03, but this does not reduce the relative bias.

IV. REDEFINING COMMUNITY DETECTION METHODS FOR MULTIPLE TIME SERIES

The paper redefines modularity for multiple time series using random-matrix-theory null models, then extends the resulting methods to multiresolution and global-mode settings. These formulations target mesoscopic communities between unit-specific noise and system-wide dependence.

  • The paper introduces three modularity redefinitions based on random matrix theory and corresponding correlation-based counterparts of popular network community-detection algorithms.The methods are benchmarked on multiple test cases after their construction.
  • For homogeneous community sizes, the naïve null distribution has one peak at 0.125, whereas stronger heterogeneity produces 64 peaks and dominant values of 0.7536 and 0.3694.The benchmarks use N = 1000 time series and c = 8 communities with φ = 1.
  • The first null model uses the identity expectation for infinitely long independent time series, while the second incorporates finite-length noise through the random-matrix correlation component C(r).
  • A third formulation removes both random noise and a dominant positive global component, such as the financial market mode.This prevents a global mode from driving the trivial all-in-one-community partition.
  • The resulting definition is aimed at mesoscopic communities between microscopic unit-specific noise and macroscopic system-wide fluctuations.
  • Iteratively filtering the global mode within individual communities can reveal multiple hierarchical levels of community structure.This provides a multiresolution extension when nested communities are present.

4. A unified redefinition

The paper defines correlation-based modularities using RMT-consistent null models and adapts community-detection algorithms to maximize them. The optimal partitions yield positively correlated communities with negative residual correlations between communities.

  • Unified formulation: The three modularity definitions are expressed in a unified form with filtered correlation matrices indexed by l.The filtered matrices are denoted C^(l), with l selecting different filtering levels.
  • Unified formulation: The normalization divides summed intra-community filtered correlations by the variance of the total increment, thereby controlling for system volatility.This normalization is useful when comparing community structures across volatile systems or time windows.
  • Interpretation: Modularity values are meaningful only in relative terms because richer null models generally produce lower values.Consequently, absolute modularity magnitudes should not be compared directly with ordinary network modularity values.
  • Scope: The RMT-based definitions require large N and T with T > N, limiting application when long, approximately stationary time series are unavailable.This dimensionality condition is described as a curse of dimensionality.
  • Community structure: The authors reformulate three popular network algorithms and prove that the maximizing partition has positive intra-community and negative inter-community filtered correlations.Thus, the resulting communities are mutually anti-correlated after filtering.

C. Multiresolution community detection

The paper introduces recursive multiresolution analysis for correlation matrices, resolving nested subcommunities while separately accounting for community-specific common factors. Controlled benchmarks show accurate recovery despite strong noise and market-mode components.

  • Multiresolution method: The proposed multiresolution method recursively applies RMT-based null models to each detected community’s original correlation submatrix.Iteration stops automatically when no further subcommunities are resolved.
  • Multiresolution method: At each recursion, the noise component retains its node-specific interpretation, while the global mode becomes a community mode containing system-wide and community-specific factors.This distinguishes nested structure from common variation within the current community.
  • Scope: For small submatrices, randomly shuffled temporal increments are preferred over the asymptotic RMT spectrum because RMT becomes less reliable at low dimensionality.The alternative estimates the relevant eigenvalue bounds from the shuffled time-series spectrum.
  • Benchmarks: The nine benchmarks use eight heterogeneously sized communities with varying noise and market-mode strengths, including challenging regimes where both components exceed one.The benchmark construction includes local noise and global-signal factors.
  • Benchmarks: Even when µ and ν exceed one, filtered matrices retain clear block structure and the method recovers the correct partition.Variation of Information quantifies agreement between detected and true partitions, while maximum modularity decreases with stronger market mode but is less affected by noise.

V. THE MESOSCOPIC ORGANIZATION OF REAL FINANCIAL MARKETS

The financial application evaluates correlation-based community detection on major stock indexes and contrasts it with thresholded asset graphs and unfiltered network modularity. Standard baselines either miss the target mesoscopic structure or collapse all stocks into one community.

  • Financial data: The study applies three algorithms to S&P 500, FTSE 100, and Nikkei 225 stocks over approximately ten years, using GICS as an external taxonomy.The selected datasets contain 445 S&P stocks, 78 FTSE stocks, and 193 Nikkei stocks.
  • Method comparison: All three adapted algorithms produce very similar partitions after filtering each financial correlation matrix.The authors use this agreement as a preliminary consistency result before presenting the main findings.
  • Asset Graph baseline: The threshold-selection procedure assumes normally distributed log-returns and does not incorporate multiple-hypothesis-test corrections.The authors note that real log-return distributions are fat-tailed.
  • Asset Graph baseline: Thresholded asset graphs reveal strong within-industry links but cannot identify communities that are internally more correlated than with the rest of the market or mutually anti-correlated.A global threshold discards weaker but comparatively informative within-community correlations.
  • Unfiltered baseline: Ordinary network modularity and unfiltered correlation-based modularity both return one trivial community containing all S&P 500 stocks.The Louvain algorithm is used in both cases.

2. Na¨ıve application of community detection

Naïvely applying network community detection to correlation matrices yields trivial or biased partitions, whereas the proposed correlation-based null models reveal mesoscopic, internally correlated and mutually anti-correlated stock communities.

  • Naïve baseline: A single community spans the entire S&P 500 when ordinary network community detection treats the correlation matrix as a weighted network.The same trivial partition results from Q1 using the null model ⟨C⟩ = 1.
  • Why naïve methods fail: The single-community result arises from modularity inconsistency or an inadequate financial-correlation null model dominated by the systemic market mode.The market mode affects all stocks simultaneously, obscuring residual community structure.
  • Correlation-based method: Discounting random and market-wide correlations with Q3 enables detection of five mesoscopic communities in the S&P 500.The communities are identified from correlations between microscopic and macroscopic levels, with sector composition shown within each community.
  • Cross-market application: The method also detects communities in the FTSE 100 and Nikkei 225, while naïve detection places all stocks into one community.The maximized Q3 values for these markets are reported separately.
  • Economic composition: Some sectors concentrate within single communities, while sector subgroups can split across communities and correlate more strongly with stocks from different sectors.Examples include cross-sector Health Care and Information Technology associations across the S&P 500, FTSE 100, and Nikkei 225.
  • Main finding: The principal result is identification of correlated stock communities that are irreducible to standard sector taxonomy and anti-correlated with one another.This conclusion emphasizes quantitative structure extracted through appropriate null models and community detection.

D. Residually anti-correlated communities and portfolio optimization

The detected communities form mutually anti-correlated groups useful for constructing community-specific market indexes and informing portfolio risk analysis. Comparative tests show similar structures across algorithms, while multiresolution analysis reveals nested subcommunities and unresolved outliers.

  • Residual anti-correlation: Maximizing correlation-based modularity identifies groups whose residual correlations are negative between communities.This property is reported for the financial-market communities in the S&P 500, FTSE 100, and Nikkei 225.
  • Portfolio implications: Residual negative correlations between community indexes are desirable for risk management and portfolio optimization.The indexes move in opposition after accounting for overall market and purely random fluctuations.
  • Portfolio indexes: Community-specific indexes can be constructed by summing the time series of stocks within each community.For distinct communities A and B, their covariance is below the null-model expectation: Cov[ X̃_A, X̃_B ] < ⟨Cov[ X̃_A, X̃_B ]⟩_l.
  • Algorithm comparison: All three algorithms identify very similar community structures, making the reported results robust to changes in detection protocol.The maximized Q3 values and community counts are closely matched, and variation-of-information values between partitions are low.
  • Algorithm comparison: Small maximized Q3 values do not imply weak communities when the market mode is strong.The chosen normalization produces small modularity values even in the presence of well-defined communities.
  • Hierarchical structure: Multiresolution detection recursively reveals subcommunities that remain internally positively correlated and mutually residually anti-correlated.The hierarchy can be continued until no community can be split into two or more anti-correlated sets.
  • Hierarchical structure: Within the Information Technology community, subcommunities separate into Software, Semiconductor, and Internet Software & Services groups, with Consumer Discretionary exceptions.Amazon and Priceline appear in the Internet Software & Services subcommunity despite belonging to Consumer Discretionary.
  • Interpretive boundary: Some stocks remain outliers whose associations may reflect unexamined corporate relationships or coincidence, limiting immediate interpretation of their community membership.Suggested explanations include shared parent companies, sizable investments, and common board members.

VI. MULTIFREQUENCY COMMUNITY DETECTION

The authors test whether detected stock communities remain stable when returns are measured at different temporal resolutions. Modularity stays nearly constant, while partitions remain mostly similar and reveal stable cores alongside frequency-sensitive soft stocks.

  • Multifrequency analysis: Nine return resolutions from 1 minute to 2 days were analyzed using three modified community-detection algorithms and RMT-based null models.The resolutions included 1, 5, 10, 15, and 30 minutes, 1 hour, 0.5, 1, and 2 days.
  • Robustness across resolutions: Q3(⃗σ∗) remained almost constant across resolutions, while VI differed from the daily partition by no more than 10% in most cases.The daily partition has VI = 0 by construction and serves as the reference.
  • Robustness across resolutions: The results indicate that correlations between large stock groups are not strongly dependent on the chosen time-step resolution.The normalization controls for varying total-return volatility, focusing the comparison on stock relationships.
  • Hard and soft stocks: Co-occurrence heat maps show large hard-stock cores and smaller groups of soft stocks that shift between communities across resolutions.Soft stocks alternate particularly between Utilities, Health Care, and Consumer Staples, and between Consumer Discretionary and Financials.
  • Hard and soft stocks: This frequency dependence captures potentially overlapping community structure despite using a non-overlapping modularity partition.The distinction between hard and soft stocks arises from how communities change with the frequency of the original time series.

B. Six-month window

The paper examines how stock communities evolve across six-month windows and finds that community strength changes substantially during the financial crisis while composition remains comparatively stable. Heat maps further reveal persistent sectoral cores and more fluid affiliations among some groups.

  • Six-month temporal analysis: Six-month windows largely preserve homogeneous community structure, but both modularity and community composition show some fluctuations.The six-month analysis refines the two-year-window results over the same ten-year S&P period.
  • Six-month temporal analysis: Q3(⃗σ∗) shows a significant drop around the last half of 2007, whereas VI indicates that overall community composition remains relatively constant.The modularity decrease is interpreted in connection with reduced group-level coherence during the financial crisis.
  • Window-to-window similarity: The most recent 2009–2011 windows have slightly more similar communities and are closer to the structure measured over the full ten-year period.The comparison is based on pairwise VI values among six-month windows and against the total period.
  • Temporal coherence of communities: The ten-year stock-level co-occurrence map reveals strong, unwavering cores that remain constantly anti-correlated with one another.Core energy, IT, and financial stocks occupy separate communities, while Energy, Materials, and Utilities often co-occur.
  • Implications and scope: The study presents the modified community-detection framework as applicable to portfolio optimization, risk management, and future extensions such as overlapping or hierarchical communities.The authors describe the method as a proof of concept for adapting additional algorithms and null models consistent with correlation matrices.

Appendix A: Redefining community detection methods

Appendix A reformulates three network community-detection algorithms for correlation matrices, including a Potts-model implementation that optimizes the modified correlation-based modularity. The adaptation preserves the optimization logic while replacing network-specific quantities with correlation-consistent ones.

  • Method adaptations: The appendix reformulates three popular network algorithms to detect communities of correlated time series using the modified modularity function.The three algorithms are presented as distinct implementations of modularity maximization.
  • Potts algorithm: The Potts approach represents partitions as spin configurations and treats modularity as proportional to the negative energy of a spin glass.Optimization seeks the lowest Hamiltonian, corresponding to the highest modularity.
  • Potts algorithm: The Potts implementation uses simulated annealing to search over spin configurations, so it can operate with the redefined correlation-based modularity.Simulated annealing is approximate and may return a different solution on each run.
  • Correlation-specific adaptation: For correlation-based networks, the Hamiltonian must be adapted because every node is connected to every other node, eliminating the network non-link contribution.The resulting optimization is equivalent to minimizing a Hamiltonian whose relevant expression is opposite in sign to the modified modularity.

2. Modified Louvain method

The modified Louvain method preserves Louvain’s greedy aggregation strategy while redefining interactions and modularity for correlation data. Covariance-based renormalization makes successive coarse-graining steps consistent with the correlation-based objective.

  • Louvain procedure: The Louvain method is a fast greedy agglomerative algorithm that repeatedly improves modularity and aggregates nodes into larger units.It begins with singleton communities, moves nodes when modularity increases, and then builds a renormalized network.
  • Renormalization: The correlation-based adaptation requires renormalized interactions between communities to remain meaningful at each aggregation level.The authors construct these interactions from the covariance bilinearity of standardized time series.
  • Renormalization: A community’s renormalized time series is defined so that aggregated interactions correspond to covariances rather than correlations.The summed time series resembles an index fund for the grouped stocks, giving node aggregation a financial interpretation.
  • Modularity consistency: The modularity is invariant under renormalization, so coarse-grained Louvain steps remain consistent with the original node-level objective.Equation (A10) coincides with the modularity defined for individual nodes.
  • Modularity updates: The adapted method retains Louvain’s computational efficiency by calculating modularity gains for moving a node between communities at every aggregation level.The gain is based on the renormalized interaction between the moved node and the target community.

3. Modified spectral method

The modified spectral method adapts modularity-based spectral optimization to correlation networks by eigendecomposing a filtered correlation-based modularity matrix. Its largest-eigenvalue eigenvector determines recursive community bisections through the signs of its components, while retaining a nonzero normalization term absent in the network formulation.

  • 3. Modified spectral method: Spectral optimization recursively bisects a network by eigendecomposing its modularity matrix and maximizing modularity.The modularity matrix is defined as the observed adjacency matrix minus the null model.
  • 3. Modified spectral method: The leading eigenvector of C(l) yields the optimal bisection: nodes are assigned to communities according to the signs of its components.The partition can be represented by a vector with entries −1 and +1 for the two communities.
  • 3. Modified spectral method: For correlation networks, the modularity matrix is replaced by the filtered matrix C(l) constructed from one of the paper’s three null models.The original spectral procedure is adapted to this correlation-based modularity matrix.
  • 3. Modified spectral method: Unlike Newman’s network formulation, the correlation-based modularity matrix does not have rows summing to zero, so the normalization term is retained.This produces a matrix-form modularity objective that includes the additional term.
Loading 1311.1924v3…