Source-linked AI summary

Unified framework for measuring segregation resolves how social and geographical space jointly shape connections

Johannes Happenhofer, Sahil Loomba, Till Hoffmann, Sumeet Agarwal, Nick S. Jones

arXiv:2609.16469v1physics.soc-phcs.SI

TL;DR

The paper addresses the limited joint study of geographical and social segregation, including the lack of publicly available individual-level geosocial networks. It develops a unified counterfactual framework and an aggregated-data-compatible model, finding that social segregation predominates and that geographical distance is associated with stronger social homophily.

  • Problem

    Existing approaches often study geographical and social segregation separately, conflate their effects, and lack publicly available large-scale individual-level geosocial network data.

  • Method

    The paper compares observed or estimated geosocial network models with counterfactual null models and fits a joint geosocial intervening-opportunities model to aggregated Facebook, Census, Pew, and Meta data.

  • Results

    Social segregation predominates over geographical segregation across race, education, and income, while social homophily increases with geographical distance.

  • Takeaways & Limitations

    Separating geographical opportunities from social preferences reveals that the same total segregation can arise from different underlying mechanisms.

  • Takeaways & Limitations

    Geographical coordinate coarsening cannot increase the paper’s segregation measures and is expected to typically decrease them.

Abstract

from arXiv · show

Our understanding of how geographical and social segregation interact remains limited, as relatively few studies investigate them jointly, and existing approaches often lack a framework distinguishing geographical, social, and total segregation. Additionally, large-scale individually resolved geo-social network data are rarely publicly available. We address both. Conceptually, we develop a unified framework that measures segregation in geosocial networks by comparing network models to appropriate null models and recovers the Theil index, dissimilarity index, and network modularity as special cases. Empirically, we turn to privacy-preserving aggregated relational data (ARD): we combine the Facebook Social Connectedness Index for the US with US Census and Pew data, and introduce an ARD-compatible joint geosocial intervening-opportunities model to infer link probabilities between region--group cell pairs. Applying our segregation framework, we find that social segregation predominates over geographical segregation, with notable separation for White--Black, college-degree--no-degree, and high-income--low/middle-income across both segregation types. We find increasing social homophily with geographical distance and group-specific geographical connectivity patterns, suggesting that geographical segregation may affect cross-group connectivity not only directly but also by amplifying social segregation.

1 Introduction

Existing segregation measures often study geography and social ties separately, while marginal connectivity measures conflate their effects. The paper proposes a unified, model-based framework and applies it to privacy-preserving aggregated network data.

  • Existing measurement gaps: Traditional region-based geographical measures face arbitrary boundaries, including modifiable areal unit and checkerboard problems.These measures assume coarse regional distances and can assign identical segregation values to substantively different spatial arrangements.
  • Conceptual framework: Marginal group-to-group connectivity cannot identify social segregation because stronger geographical homophily can produce more within-group links without social homophily.Hypothetical societies A and B have different geographical homophily but absent social homophily, while B and C can share marginal mixing despite different geographical and social homophily.
  • Conceptual framework: The paper distinguishes geographical, social, and total segregation by defining counterfactual network models relative to the observed or estimated geosocial network.Geographical segregation reflects distance-related connectivity decay, whereas social segregation reflects group-based connectivity preferences.
  • Data challenge: Large-scale individual-level geosocial networks are usually unavailable because of privacy concerns and difficult statistical inference at millions of nodes and billions of edges.The paper therefore uses aggregated connection volumes between sets of nodes, preserving privacy while compressing the data.
  • Empirical strategy: The empirical analysis combines the Facebook Social Connectedness Index with Census, Pew, and Meta reach estimates to infer geosocial connectivity across US regions and demographic groups.The model specifies tie probabilities conditional on geosocial coordinates and a likelihood for the observed aggregated data.
  • Empirical strategy: A joint geosocial intervening-opportunities model accounts for population-size scaling and is reported to outperform simpler factored geographical-plus-social models.Its intervening opportunities are defined using latent joint distance in geosocial space rather than geographical distance alone.

2 Results

The paper introduces a unified framework that separates total, social, geographical, and traditional geographical segregation by comparing geosocial network models with counterfactual references. Applied to US Facebook ARD, the framework finds that social segregation dominates geographical segregation and that joint geosocial intervening-opportunities models best explain connectivity.

  • 2.1 Measuring geographical, social, and total segregation: The framework operationalizes total, social, geographical, and traditional geographical segregation through corresponding network models and recovers modularity, dissimilarity, and Theil indices as special cases.It also distinguishes homophilous from heterophilous contributions using a normalized f-divergence and homophily index.
  • 2.3 Joint geosocial intervening opportunities best explain SCI connectivity: Joint fuzzy intervening-opportunities models consistently fit SCI better than distance-based and hybrid models across 40 US states and multiple specifications.At small distances, fuzzy intervening opportunities can halve mean squared error relative to a standard intervening-opportunities baseline.
  • 2.4 Inferred connectivity reveals strong group homophily along race, education, and income: Inferred social distances and connection probabilities show pronounced homophily by race, education, and income, especially between White and Black, college-degree and no-degree, and affluent and other groups.These patterns remain visible under geographical counterfactuals but are strongest under the social segregation counterfactual.
  • 2.5 Social segregation predominates over geographical segregation: Social segregation is two to three times stronger than traditional geographical segregation, which is about twice as strong as SCI-induced geographical segregation across race, education, and income.Total segregation is only slightly stronger than social segregation, indicating that most inferred segregation reflects group-based connectivity preferences rather than group distributions alone.
  • 2.5 Social segregation predominates over geographical segregation: Social segregation is nearly entirely homophilous and predominantly within counties, whereas geographical segregation includes more heterophilous and between-county components.For education and income, traditional geographical segregation is weaker within rural than urban areas; race segregation is similar across urban and rural settings.
  • 2.5 Social segregation predominates over geographical segregation: Under a neighbourhood-based definition, traditional geographical segregation has no or weakly negative correlation with social segregation, suggesting local diversity need not indicate stronger same-group connectivity preferences.This contrasts with country-reference measures whose traditional geographical segregation is positively associated with social segregation.

3 Discussion

The paper unifies segregation measurement across geosocial networks and applies the framework to aggregated US connectivity data. It finds that social segregation predominates, while geographical distance and segregation are associated with stronger social homophily, subject to important inference and causal limitations.

  • The framework distinguishes total, social, geographical, and traditional geographical segregation using counterfactual network models and recovers several established indices as special cases.The framework also decomposes segregation into homophilous and heterophilous contributions and generalizes Theil-index decomposition properties.
  • Privacy-preserving aggregated relational data can support geosocial connectivity inference through a joint intervening-opportunities model that fits Facebook Social Connectedness Index data better than models separating geography and social space.The model uses regional variations in sociodemographic proportions and enables countrywide segregation analysis without individual-level network data.
  • Social segregation predominates over geographical and traditional geographical segregation for race, education, and income.Traditional geographical segregation is larger than the paper’s geographical segregation, while several group separations persist across all four segregation types.
  • Social homophily increases with geographical distance, and greater geographical race isolation in a ZCTA is associated with higher local social race homophily.These findings suggest that geographical segregation may shape same-group connectivity preferences beyond restricting cross-group opportunities, but the paper presents this as a hypothesis.
  • Inference remains limited because group-specific connection probabilities are latent, depend on estimated composition and Facebook usage, and cannot exclude unobserved confounding.The paper calls for validation with richer aggregated or individual-level network data and notes that Facebook friendship ties may not represent all forms of social connection equally.
  • The f-divergence framework is best suited to categorical sociodemographic variables because it does not incorporate ordering among variable values.
  • The cross-sectional framework quantifies segregation but does not establish whether geography strengthens social segregation, social homophily reinforces geography, or both co-evolve.The authors identify replication across countries and longitudinal causal analysis as priorities for future research.

4 Methods

The framework separates total, social, geographical, and traditional geographical segregation by comparing geosocial network models with counterfactual reference models. It combines f-divergence measures with a joint geosocial intervening-opportunities model to avoid density-induced degree artifacts.

  • 4.1 Segregation framework: Network-based f-divergences compare coordinate distributions among connected pairs with distributions under suitable reference models, extending traditional segregation measures.The construction is designed to recover established measures as special cases.
  • 4.1 Segregation framework: The framework defines total, social, geographical, and traditional geographical segregation through models that selectively retain or remove geographical and sociodemographic mechanisms.Total segregation comes from the model of interest; the other forms use counterfactual models.
  • 4.1.6 Recovery of established segregation measures: The regional isolation model recovers the multigroup Dissimilarity and Theil Information Theory indices, the squared coefficient of variation, Relative Diversity, nominal assortativity, and network modularity.These equivalences follow from specific generator choices and model constructions.
  • 4.3 Geosocial connectivity model: Intervening opportunities replace direct distance to account for geographical population density and avoid unrealistically large expected degrees in dense regions.The model measures separation using the number of individuals geographically closer to one person than another.
  • 4.3 Geosocial connectivity model: The joint IO model extends separate geographical and social components, while containing hybrid and distance-based models as special cases.The hybrid model can produce unrealistically low degrees for locally rare groups when group-specific density is omitted.
  • 4.4 Model specification: Marginal connection probabilities remain similar across three alternative sociodemographic factorizations, supporting the final estimates' qualitative robustness.The factorizations vary the ordering of Urban, Gender, Race, Age, Education, and Income variables.

Data availability

The study uses publicly available US demographic, Facebook-usage, Facebook reach, and Social Connectedness Index data sources.

  • Data availability: All study data were publicly available and combine US Census, Pew Facebook-usage, Meta Marketing API reach, and Social Connectedness Index sources.The demographic sources include the American Community Survey, 2020 Decennial Census, and Current Population Survey.

5 Extended Data

Extended analyses test model-fit robustness, count perturbations, predictive checks, Facebook-usage estimates, and fuzzy intervening opportunities. The results generally preserve qualitative connectivity patterns while identifying sensitivity to group counts, usage weights, and model assumptions.

  • Extended Data Fig. 1: Marginal connectivity patterns are qualitatively similar across models with constrained, agnostic, hybrid, and distance-based specifications.Estimated marginal degrees vary substantially for some age, education, and income groups and should be interpreted cautiously.
  • Extended Data Fig. 2: Substituting census counts preserves most connectivity patterns, but permuting census or Facebook counts weakens or removes social-connectivity signals and worsens SCI fit.Age patterns are especially sensitive to estimated Facebook usage.
  • Extended Data Fig. 3: Perturbing census and Facebook counts clarifies how count assumptions alter the implied social-segregation models.The extended-data analysis compares unperturbed counts with alternative exposure and spatial-composition perturbations.
  • Extended Data Fig. 6: Posterior predictive checks accurately reproduce marginal ARD patterns but incompletely capture residual spatial autocorrelation.The checks are shown for California, the largest state in the analysis.
  • Extended Data Fig. 7: Facebook usage is highest among ages 25–34 and 35–44, while API reach estimates overestimate these groups and underestimate older groups.The age-specific errors reflect a low true-negative rate for ages 25–34 and low true-positive rates for ages 55–64 and 65+.
  • Extended Data Fig. 8: Fuzzy intervening opportunities substantially reduce SCI prediction error at small geographical distances but have negligible effects at large distances.The reported tuning ranges favor ν in {3 km, 4 km} and k in {7, 10, 15} when local links are emphasized.

S1 Segregation framework

The supplementary framework expresses traditional and spatially weighted segregation as f-divergences, then generalizes them to geosocial networks through model-based edge-coordinate distributions. It supplies counterfactual null and maximum-segregation references and proves recovery of established indices.

  • S1.4–S1.5 Geosocial segregation measures: The supplementary framework defines no-segregation and maximum-segregation reference models and compares network-induced edge-coordinate distributions with these references.The maximum-segregation model is shown to maximize the relevant f-divergence over a suitable model class.
  • S1.8 Theoretical properties: The framework's measures cannot increase under deterministic coarsening of the coordinate space.This result provides a formal monotonicity property for aggregating coordinate distinctions.
  • S1.1 Traditional and spatially weighted segregation: Traditional and spatially weighted geographical segregation measures are represented as f-divergences between joint coordinate-group distributions and products of their marginals.The spatially weighted formulation addresses dependence on predefined regions.
  • S1.2 Network interpretation: A spatial connectivity kernel can be interpreted as a network model with constant expected degrees, while the Erdős–Rényi model supplies a null reference with marginal connection probability.The model compares expected edge fractions between coordinate pairs under the connectivity model and the reference model.

S1.3 Coordinate spaces and geosocial network models

The framework represents geosocial networks through coordinate spaces and connection probabilities, then constructs counterfactual models that isolate social, geographical, traditional geographical, and combined segregation.

  • Coordinate spaces and models: Geosocial network models jointly specify coordinate distributions and conditional connection probabilities for sampled individuals.Coordinates may be full, geo-egocentric, or purely social, with regional coordinates obtained by coarsening point-level locations.
  • Counterfactual models: The homogeneous sociodemographic model removes geographical concentration while retaining group-specific connectivity patterns across regions.It preserves observed connectivity conditional on geosocial coordinates while equalizing regional sociodemographic composition.
  • Counterfactual models: The social ambiphily model removes direct dependence of connectivity on group membership while preserving the observed joint distribution of regions and groups.Its resulting connection probability depends on region pairs but not on the social identities of either endpoint.
  • Counterfactual models: The regional isolation model preserves population distribution and network density while allowing connections only within predefined regions, isolating traditional geographical segregation.It is sensitive to group distributions across selected regions rather than actual geographical connectivity patterns.
  • Reference models: The framework defines random-mixing and maximal-segregation references for each segregation-defining model, including total, social, geographical, and traditional geographical cases.The choices M, MHB, MAB, and MRI correspond to those four segregation concepts, while different coordinate resolutions are related by marginalization and coarsening.
  • Reference models: Degree-configuration and social-isolation models respectively represent random mixing and complete within-group separation while preserving the same marginal geosocial connectivities.Their composition yields a counterfactual model with both geographical and social segregation removed, although the framework instead adopts random mixing as its no-segregation reference.

S1.5 Divergence-based segregation indices

Segregation is measured as an f-divergence from degree-constrained random mixing and normalized against degree-constrained social isolation, yielding comparable indices across segregation concepts.

  • Divergence-based indices: The unnormalized measure is the f-divergence between a segregation-defining model’s induced distribution and its degree-configuration reference model.The divergence is zero when the induced distribution equals the degree-configuration distribution.
  • Normalization: Normalizing by degree-constrained social isolation places segregation indices on a common scale while preserving the model’s coordinate distribution and marginal connectivities.The resulting choices distinguish total, social, geographical, and traditional geographical segregation according to the defining model.
  • Normalization: Degree-constrained social isolation maximizes the f-divergence from degree-constrained random mixing under fixed marginal coordinate distributions and connectivities.Under the stated admissibility conditions, the normalized social-coordinate measure lies between 0 and 1.
  • Reference-model choice: The combined homogeneous-ambiphily model removes both direct and geography-mediated social dependence, but the framework instead uses random mixing as its no-segregation reference.This distinction matters because degree-constrained random mixing controls for expected degree differences across geosocial coordinates.
  • Connections to established indices: For x log x, 1, and x^2 − 1 generators, the normalization constants correspond to the Theil Information Theory, Dissimilarity, and Pearson χ2 segregation indices.These generators therefore connect the general divergence framework to established index families.

S1.6 Recovery of traditional geographical segregation measures as special cases

Traditional geographical segregation indices emerge as exact special cases of the geosocial divergence framework under regional isolation and suitable null models.

  • Exact recovery: Traditional Theil, Dissimilarity, and squared-coefficient-of-variation indices are recovered exactly as special cases of the proposed geosocial measures.The recovery is established for traditional geographical segregation represented by the regional isolation model.
  • Reference-model dependence: Geographical segregation remains measurable even after restricting the representation to social coordinate space.The framework therefore retains geographical information through model-based comparisons beyond explicit geographic coordinates.
  • Exact recovery: On geo-egocentric coordinate space, comparing regional isolation with an Erdős–Rényi random graph yields traditional geographical indices from connection probabilities.The same construction extends to spatially weighted indices when a geographical connectivity kernel replaces regional isolation.
  • Normalization: The general normalization approach recovers normalized Theil, Dissimilarity, and χ2 indices rather than requiring separate index-specific normalizations.This provides a unified normalization principle across the traditional measures.
  • Reference-model dependence: Using the homogeneous sociodemographic model as reference recovers the geographical divergence only for selected generators, not for general f.For x log x or −log x, recovery holds up to a factor of 2; for general f, it does not.

S1.7 Choice of coordinate space and information loss

The framework uses geo-egocentric geosocial space as a practical compromise: it supports principled reference models and traditional-index recovery but loses sensitivity to detailed distance-dependent connectivity.

  • Choice of coordinate space: Geo-egocentric space is selected because degree-configuration and degree-constrained social-isolation models provide suitable reference models and principled normalization.This coordinate space is also less restrictive than full geosocial space for formulating the reference models.
  • Information retained and lost: Geo-egocentric measures remain sensitive to regional variation in social connectivity but are insensitive to the rate or irregularity of geographical-distance decay when marginal group connectivity is unchanged.Full geosocial measures remain sensitive to distance-dependent connectivity relative to the null model.
  • Choice of coordinate space: Traditional geographical segregation measures are recovered as exact special cases on geo-egocentric coordinate space.This is one reason the chosen representation supports continuity with established segregation indices.
  • Information retained and lost: Coarsening geographical coordinates cannot increase σ-based segregation values and is expected to typically decrease them.Regional aggregation therefore entails real information loss, even when full geographic detail is unnecessary for some applications.
  • Information retained and lost: Questions about group-specific geographical connectivity or how homophily changes with distance require full geosocial coordinate space.These questions cannot be answered from the coarser geo-egocentric representation alone.

S1.8 Conditional coordinate and joint coordinate–link distributions

The framework distinguishes conditional-coordinate and joint coordinate–link distributions, showing when their segregation measures coincide or become asymptotically equivalent. This supports using the simpler conditional-coordinate distribution as the default in common applications.

  • The joint coordinate–link distribution τ captures coordinates and connectivity, while σ provides the natural normalization and recovers traditional geographical segregation measures.
  • As marginal connectivity vanishes or approaches one, τ-based measures become asymptotically equivalent to σ- or ξ-based measures under smooth generators.
  • Df,τ equals ¯P Df,σ plus (1 − ¯P) Df,ξ when models share coordinate distributions and marginal connectivity.
  • For Total Variation distance, the three normalized measures coincide exactly, including the relation shown for Dg,τ, Dg,σ, and Dg,ξ.
  • For the χ2 divergence with an Erdős–Rényi null, exact equivalence also holds, while common smooth f-divergences motivate choosing σ as the simpler default.

S1.9 Homophily, heterophily, and network modularity

The section decomposes segregation departures into homophilous and heterophilous contributions and defines signed homophily. Under Total Variation distance, this framework recovers network modularity and nominal assortativity.

  • Signed homophily subtracts heterophilous from homophilous contributions, distinguishing the direction of departures from degree-constrained random mixing.
  • The τ- and σ-based decompositions are asymptotically equivalent in sparse networks, while their normalized Total Variation versions coincide exactly.
  • The framework recovers network modularity for arbitrary models and Relative Diversity under the regional isolation model.
  • For Total Variation distance, the normalized signed homophily measure equals nominal assortativity and the unnormalized measure equals generalized network modularity.
  • Unlike signed homophily, the Dissimilarity index counts both homophilous and heterophilous deviations toward segregation.

S1.10 Regional decomposability

The framework decomposes KL-divergence segregation into between-region and within-region components. It also clarifies that alternative decompositions can become uninformative when most connectivity information is assigned to one component.

  • The construction compares connectivity involving super-regions with countrywide sociodemographic connectivity while retaining finer regional structure conditionally.
  • The regional decomposition separates segregation between super-regions from weighted within-region components.
  • For the regional isolation model, the resulting decomposition exactly matches the traditional regional decomposition of the Theil index.
  • Alternative decompositions may place most connectivity information in the between component, making the resulting interpretation less appealing.

S1.11 Group decomposability

The framework also decomposes segregation across coarse sociodemographic groups and their finer within-group structure. Under regional isolation, this recovers the traditional group decomposition of the Theil index.

  • The group decomposition separates segregation between coarse sociodemographic groups from additional segregation within those groups.
  • Within-group components measure the additional impact of finer sociodemographic coordinates conditional on membership in a coarse group.
  • For the regional isolation model with an Erdős–Rényi null, the decomposition recovers the traditional group decomposition of the Theil index.

S1.12 Normalization of regional and group decompositions

The regional and group decompositions derive normalization constants by applying KL-divergence decompositions to the framework’s global normalization. Under the regional isolation model, these constants recover the standard Theil-index normalizations.

  • Normalization construction: The proposed between and within normalizations decompose the global entropy or KL-divergence using mutual information and conditional entropy terms.The entropy satisfies H(B) = I(S; B) + H(B|S), motivating between and within components.
  • Normalization construction: The between normalization is a KL-divergence between the model’s coarse group edge distribution and its degree-constrained isolation reference.The corresponding within normalization compares fine-group edge distributions conditional on the coarse group.
  • Normalization construction: The within normalization is likewise a KL-divergence between the model’s fine-group conditional edge distribution and the degree-configuration reference.These quantities are interpreted as conditional components of the global normalization rather than independently solved restricted problems.
  • Recovery of established measures: When the underlying model is regional isolation, the proposed group and regional normalizations recover the traditional Theil-index normalization constants.This establishes compatibility with established Theil decompositions while retaining the unified geosocial formulation.
  • Normalization caveat: The adopted normalization prioritizes comparability with the traditional Theil index because an alternative constrained optimization does not recover its regional-isolation normalizations.The alternative would preserve coarse-group and regional marginals, but the relevant fine-coordinate marginals generally differ.

S1.13 Additional results for the segregation framework

The additional results establish how coordinate coarsening affects the framework’s divergence measures. Monotonicity holds for the σ-based measures after normalization, but not generally for the τ- and ξ-based DSI-scaled measures.

  • Coordinate coarsening: For every generator and coordinate distribution, deterministic coordinate coarsening cannot increase the unnormalized f-divergence.This applies across the τ, σ, and ξ formulations.
  • Coordinate coarsening: For σ-based measures, the same coarsening inequalities remain valid after normalization by the divergence of the maximal-segregation model from the null.Thus σ-based normalized segregation is monotone under the considered coordinate coarsenings.
  • Coordinate coarsening: The normalized τ- and ξ-based measures need not be monotone because coarsening can reduce their maximal-segregation divergence proportionally more than the model divergence.The paper gives counterexamples for both formulations.
  • Counterexample: In the counterexample, the ξ-based divergence changes from approximately 0.0911 at full coordinates to 0.0997 after marginalizing the second individual’s region.The DSI-scaled ξ measure therefore increases under coordinate coarsening.
  • Counterexample: The τ-based divergence similarly changes from approximately 0.0809 to 0.0821 after coarsening, so its DSI-scaled measure can also increase.This demonstrates that the monotonicity guarantee is specific to σ-based normalization.

S2 Proofs for the segregation framework

The proofs establish the framework’s normalization, recovery, and coarsening properties through convexity, degree constraints, marginalization, and specialization to regional isolation. They also derive the stated relationships among geosocial, regional, and group-level models.

  • Normalization proof: The normalization proof uses convexity of the f-divergence generator to show that degree-constrained social isolation maximizes separation from the degree-configuration null.Equality is attained by the degree-constrained social-isolation model.
  • Normalization proof: The model equivalence-class argument verifies that the degree-configuration null and degree-constrained social-isolation model preserve the required marginals.These constraints support the maximal-divergence result.
  • Recovery proofs: Specializing the framework to a single region reduces the geosocial models and their constraints to the corresponding group-level models.This yields the group-level recovery result by applying the normalization theorem in the single-region case.
  • Marginalization proofs: The proofs relate geosocial measures to their regional and group marginals through marginalization of the coordinate distributions.The resulting identities establish the stated connections among the framework’s model formulations.
  • Recovery proofs: For the regional isolation model, the region-based kernel makes conditional group composition depend only on the region containing the geographical coordinate.Grouping fine geographical coordinates by region then recovers the standard regional segregation expressions.
Loading 2609.16469v1…