Source-linked AI summary

Geographic constraints on social network groups

Jukka-Pekka Onnela, Samuel Arbesman, Marta C. González, Albert-László Barabási, Nicholas A. Christakis

arXiv:1011.4859v2physics.soc-phcond-mat.dis-nncs.SI

TL;DR

The paper asks how geography constrains social group formation beyond constraints on individual social ties. Using a mobile-phone network with known communication locations and algorithmically detected communities, it examines geographic span, spatial clustering, and the relation between topological and geographic positions. It finds that geography strongly shapes community structure: small communities are geographically tight, larger communities show a marked increase in span beyond about 30 members, and topological and geographic centrality are essentially uncorrelated.

  • Problem

    The paper addresses the lack of large-scale evidence on how geography shapes network communities beyond its effect on dyadic social interactions.

  • Method

    The authors analyze mobile-phone calls and texts with recorded tower locations, detect network communities using modularity maximization, and compare their geographic structure with null models.

  • Results

    Small communities are geographically tight, community span shows a large increase above about 30 members, and topological and geographic centrality within communities are essentially uncorrelated.

  • Takeaways & Limitations

    Network communities and dyadic ties are shaped by geography in distinct ways, so spreading models should account for simultaneous geographic and network constraints.

Abstract

from arXiv · show

Social groups are fundamental building blocks of human societies. While our social interactions have always been constrained by geography, it has been impossible, due to practical difficulties, to evaluate the nature of this restriction on social group structure. We construct a social network of individuals whose most frequent geographical locations are also known. We also classify the individuals into groups according to a community detection algorithm. We study the variation of geographical span for social groups of varying sizes, and explore the relationship between topological positions and geographic positions of their members. We find that small social groups are geographically very tight, but become much more clumped when the group size exceeds about 30 members. Also, we find no correlation between the topological positions and geographic positions of individuals within network communities. These results suggest that spreading processes face distinct structural and spatial constraints.

Introduction

Social groups can be represented as network communities, but geography may constrain their formation beyond the network structure of pairwise ties. Mobile-phone data make it possible to study this relationship at large scale, addressing how geographically grounded groups compare with larger groups enabled by modern society.

  • Introduction: Social groups reflect friendship, kinship, and work relationships and can be represented as communities within social networks.In network terms, communities are densely connected internally and sparsely connected to other nodes.
  • Introduction: Modern technology and political organization enable larger social units and group assembly across wider geographic ranges than older social groups.This raises whether traditional groups resemble large-scale groups possible in modern society.
  • Introduction: Mobile-phone data now support large-scale measurement of social interactions and geographic positions for millions of individuals, alongside algorithmic community detection.This feasibility creates an empirical basis for examining how community structure depends on actual communication locations.
  • Introduction: Previous studies found smaller communities more homogeneous in members’ billing postal codes, but large-scale studies using actual communication locations had not examined communities’ geographic properties.The paper positions its analysis as addressing this gap.
  • Introduction: Geography constrains social connections, but its effects on network communities beyond dyadic interactions remain unclear.The open question concerns how geographic constraints and opportunities shape group formation above their influence on individual pairs.

Results

Geography strongly constrains communication ties and the spatial organization of network communities. Small communities are geographically compact, while communities larger than about 30 members show a pronounced span increase that distance-decay null models cannot explain; topological and geographic centrality are unrelated.

  • Network and dyadic ties: 72.4 million calls and 17.1 million texts produced a 3.4-million-node network with 5.2 million weighted ties and average degree ⟨k⟩≈3.0.Locations of communication towers were used to assign individuals’ frequent geographic positions.
  • Network and dyadic ties: Five orders of magnitude: communication probability decreases as distance increases from 1 km to 1,000 km.Power-law exponents were approximately 1.58 for voice ties and 1.49 for text ties; average connected-node distances were 42 km and 51 km, respectively.
  • Network and dyadic ties: Tie strength is nearly flat with distance despite the strong distance dependence of tie existence.The figure reports only a minor decreasing trend for voice ties.
  • Community structure: Communities were detected by modularity maximization after combining voice and text ties into one network.Within communities, betweenness centrality and distance from the geographic center were converted to percentiles for comparison.
  • Community structure: 42,123 nodes from communities sized 10–1,000 were analyzed, with no relationship between topological and physical centrality.Both centrality measures were represented as percentiles to account for their dependence on community size.
  • Community geography: Large bump: geographic community span increases sharply for communities exceeding 30 nodes, after an earlier upward trend and leveling off.The empirical span is much smaller than the community-null span; for size ten, empirical span is about 50 km versus roughly 300 km for the size-ten null comparison.
  • Community geography: Distance-decay dyadic null models produce a smooth span curve and cannot explain the observed bump.The model incorporates P(d) ∼d^-1.5 and population concentration in cities while varying the scale parameter.
  • Community geography: 0.67 versus 0.29 spatial clusters per added member: growth is faster in small than large communities.An increase from 5 to 15 members adds 6.1 clusters, whereas 25 to 35 members adds 2.7; the nonlinear fit has exponent b3 = 0.43.

Discussion

Geography shapes social communities beyond its influence on individual ties: community span and clustering change systematically with group size, while network and geographic centrality are essentially uncorrelated. These findings challenge well-mixed assumptions and motivate models that jointly represent geographic and network constraints on spreading.

  • Discussion: P(d) ∼d−1.5 differs from a prior Belgian estimate of P(d) ∼d−2 for communication ties.The paper attributes possible differences to the datasets and distance measures used.
  • Discussion: Communities show a geographic-span bounce near 30 members: growth from 10 to 20 members raises span about 40%, versus about 100% from 30 to 40.The authors interpret this pattern as weakening geographic cohesion beyond roughly 30 members.
  • Discussion: Topological and geographic centrality within communities are essentially uncorrelated, indicating that community geography has properties distinct from pairwise interactions.The discussion distinguishes group-level geographic structure from the geographic constraints observed for dyadic ties.
  • Discussion: Observed ties and community locations do not follow well-mixed null models, so geography remains a compartmentalizing factor at both network and community levels.This qualifies the use of perfect-mixing assumptions in models of infectious disease or technology diffusion.
  • Discussion: Spreading models should account simultaneously for geographic and network constraints rather than treating population interactions as governed by only one structure.The paper identifies this joint treatment as a direction for modeling pathogen and information flow.

Methods

The study constructs anonymized mobile-phone interaction networks, filters interactions and participants, symmetrizes communication modes, and detects communities by modularity maximization. It uses the Louvain heuristic because maximizing modularity is NP-hard and the networks contain millions of nodes.

  • Methods: The dataset contains four weeks of anonymized calls and texts, retaining operator customers and individuals who made at least two calls.Sporadic interactions were filtered using a minimal reciprocation requirement for included ties.
  • Methods: Directed voice and text interactions were converted into a symmetric unweighted network before community detection.The two communication modes were effectively combined into one network.
  • Methods: Communities were detected by maximizing modularity, which compares within-group edge concentration with a degree-preserving chance expectation.The Newman-Girvan null model is represented by the k_i k_j/(2w) term.
  • Methods: The Louvain method was chosen because modularity maximization is NP-hard and the analyzed networks contain millions of nodes.Louvain provides a computationally efficient heuristic for this scale.
Loading 1011.4859v2…