Source-linked AI summary

Analysis of community structure in networks of correlated data

Sergio Gomez, Pablo Jensen, Alex Arenas

arXiv:0812.3030v2physics.soc-ph

TL;DR

The paper addresses how to detect communities in networks built from correlated data without discarding correlation signs or self-correlations. It reformulates modularity for directed, weighted, signed networks with self-loops, and applies the formulation to retail correlations in Lyon, where it improves agreement with the Chamber of Commerce classification.

  • Problem

    Standard correlation-network analyses and modularity treatments can mislead by discarding correlation signs, self-loops, or the probabilistic null-model semantics needed for general networks.

  • Method

    The paper extends modularity to directed, weighted, signed networks with self-loops, using separate probabilities for positive and negative links.

  • Results

    The generalized modularity recovers standard modularity without negative weights, is zero for a single-community partition, and is antisymmetric under weight-sign reversal.

  • Takeaways & Limitations

    The formulation permits correlated-data networks to be analyzed without symmetrizing links, removing autocorrelations, or replacing correlations with unsigned values.

Abstract

from arXiv · show

We present a reformulation of modularity that allows the analysis of the community structure in networks of correlated data. The new modularity preserves the probabilistic semantics of the original definition even when the network is directed, weighted, signed, and has self-loops. This is the most general condition one can find in the study of any network, in particular those defined from correlated data. We apply our results to a real network of correlated data between stores in the city of Lyon (France).

I. INTRODUCTION

Community structure reveals locally organized groups that global network statistics can obscure. For correlation-derived networks, modularity offers a network-based way to identify such clusters, but standard extensions do not fully address signed correlations and self-loops.

  • Community structure consists of densely connected groups of nodes separated by sparse or weak connections.
  • Studying communities helps elucidate network organization and may relate groups of nodes to functionality.
  • Modularity is a quality function that compares network partitions and is widely used because it combines accuracy with manageable computational cost.
  • Correlation networks are commonly built by thresholding correlations, retaining unsigned positive links, and removing self-loops.
  • The paper argues that correlation-network analysis must retain correlation signs and self-loops, then extends modularity to directed, weighted, and signed links.

II. GENERALIZATION OF MODULARITY

The paper generalizes modularity to preserve its probabilistic interpretation for directed, weighted, signed networks with self-loops, addressing failures of standard modularity on correlated data.

  • Standard modularity: Standard modularity compares within-community edge weight with a null network preserving node strengths, and larger modularity indicates greater deviation from that null case.The node strength is the sum of its connection weights.
  • Motivation: In a network whose within-group correlations are +1 and between-group correlations are −1, standard modularity is zero for every partition and cannot recover the communities.The example assigns state +1 to one community and −1 to the other, producing a correlation matrix with four constant blocks.
  • Signed modularity: Negative weights invalidate the single-probability interpretation of relative strength, so the formulation introduces separate probabilities for positive and negative links.The method separates positive and negative weights, strengths, and total strengths before combining their modularity contributions.
  • Signed modularity: The final modularity balances positive links that favor communities against negative links that oppose them, weighted proportionally to their respective total strengths.The generalized expression is antisymmetric under reversing all weight signs and reduces to standard modularity when negative weights are absent.
  • General case: The generalization extends the framework beyond undirected networks through substitutions for directed links while retaining the stated modularity properties.The paper explicitly identifies directed, weighted, signed links and self-correlations as the target general case.

III. COMPARISON WITH OTHER METHODS

The paper compares its generalized modularity with the Potts model and original modularity on a signed network containing two positive cliques joined by positive and negative edges.

  • Comparison with other methods: The new modularity succeeds in recovering two communities, whereas the original modularity and Potts model select an incorrect partition.The example uses two positive-link cliques connected by one positive and one negative edge.
  • Comparison with other methods: If |v| < 1, the Potts model joins both cliques because their inter-clique strength is 1 + v > 0.Its Hamiltonian omits the null-case term, so positive net strength between cliques is rewarded.
  • Comparison with other methods: When |v| exceeds the number of positive links, original modularity also chooses a single module containing all nodes.Although original modularity includes a null case, it was not designed for negative weights.
  • Comparison with other methods: The compared alternative modularity for positive and negative links is compatible with the current definition and equivalent when λ = γ = 1.

IV. APPLICATION TO A REAL NETWORK

The method is applied to a directed attraction–repulsion network of Lyon retail activities, where it identifies communities and interpretable retailer roles using signed, directional interactions.

  • Network construction: Attraction and repulsion coefficients compare observed neighborhood concentrations with random-placement reference concentrations.The logarithm of the actual-to-reference concentration ratio is positive for attractions and negative for repulsions.
  • Network construction: 11629 stores yield a directed network with 97 retail-activity nodes and 1131 links: 715 positive and 416 negative.Only coefficients significantly different from zero under Monte Carlo sampling enter the adjacency matrix.
  • Community detection: Equation (18) produces six communities, compared with four from Eq. (1) and nine predefined by the Lyon Chamber of Commerce.Rand, Jaccard, and NMI indices all show Eq. (18) better recovers the reference classification.
  • Retailer roles: The z-score compares each node’s internal strength with the mean and standard deviation for its assigned community.The analysis preserves link direction and sign, distinguishing positive and negative incoming and outgoing roles.
  • Retailer roles: In the largest 34-activity proximity-store community, z-scores identify attractive, repulsive, attracted, and repelled retailers.Sports facilities and funeral services show systematic attraction patterns, while gas stations are both most attracted and most repelled.

V. CONCLUSIONS

The paper proposes a modularity formulation for general networks and demonstrates its use on attraction–repulsion data from Lyon retail activities.

  • Conclusions: The formulation handles directed, weighted, signed links and self-loops while preserving modularity’s original probabilistic semantics.
  • Conclusions: It analyzes correlated-data networks without symmetrizing links, removing autocorrelations, or discarding correlation signs.
  • Conclusions: In Lyon, the new modularity outperforms the original definition against the Chamber of Commerce classification and motivates node roles based on direction and weight sign.
Loading 0812.3030v2…