Source-linked AI summary

Social Structure of Facebook Networks

Amanda L. Traud, Peter J. Mucha, Mason A. Porter

arXiv:1102.2166v1cs.SInlin.AOphysics.soc-ph

TL;DR

The paper studies Facebook friendship networks at American colleges and universities, examining how user attributes organize ties and larger-scale communities. Using complete network snapshots and community comparisons, it finds that high school matters more at large universities, while residence is more important for community organization at some institutions than others.

  • Problem

    The paper examines how user attributes organize Facebook friendship networks and addresses the need to understand the social structure underlying these networks.

  • Method

    The study analyzes complete Facebook networks from 100 American colleges and universities using single-day snapshots, dyadic measures, regression models, and algorithmically detected communities.

  • Results

    High school plays a greater role in the social organization of large universities, while residence is much more important for community organization at some institutions than at others.

  • Takeaways & Limitations

    Microscopic and macroscopic perspectives provide complementary insights into the social organization of university Facebook networks.

  • Takeaways & Limitations

    The findings should be complemented by studies of corresponding real-life social networks because they have different properties from online social networks.

Abstract

from arXiv · show

We study the social structure of Facebook "friendship" networks at one hundred American colleges and universities at a single point in time, and we examine the roles of user attributes - gender, class year, major, high school, and residence - at these institutions. We investigate the influence of common attributes at the dyad level in terms of assortativity coefficients and regression models. We then examine larger-scale groupings by detecting communities algorithmically and comparing them to network partitions based on the user characteristics. We thereby compare the relative importances of different characteristics at different institutions, finding for example that common high school is more important to the social organization of large institutions and that the importance of common major varies significantly between institutions. Our calculations illustrate how microscopic and macroscopic perspectives give complementary insights on the social organization at universities and suggest future studies to investigate such phenomena further.

1. Introduction

The paper uses complete Facebook networks from 100 American colleges and universities to compare how user characteristics organize social structure at dyadic and community scales.

  • Motivation: Facebook friendships form reciprocated, undirected ties that often draw from users’ real-life social networks.The paper treats Facebook networks as imperfect proxies for offline social networks and emphasizes the limitations of that inference.
  • Motivation: SNSs provide unusually large social and demographic datasets for studying social organization at unprecedented size and detail.Their users also voluntarily reveal substantial personal information, although online users are a biased sample of the broader population.
  • Study scope: The study analyzes complete Facebook networks from 100 American colleges and universities in a single-day September 2005 snapshot.Because users needed .edu addresses and most friendships were within institutions, links across institutions are ignored.
  • Approach: For each institution, the authors compare gender, class year, major, high school, and residence with homophily and algorithmically detected community structure.This combines dyad-level analysis with comparisons between computed communities and partitions based on categorical user data.
  • Contribution: The 100 networks enable comparisons of how university social organizations differ and provide benchmark examples for community-detection computations.The authors argue that these networks can support comparisons of underlying university social networks, which they imperfectly represent.

2. Data

The dataset consists of complete, anonymized Facebook friendship networks from 100 institutions, restricted to within-institution ties and enriched with volunteered categorical attributes.

  • Data coverage: The data contain the complete set of users and friendship links for Facebook networks at 100 American institutions.Unlike survey or automated-sampling studies, complete networks avoid missing nodes and links that can affect graph analyses.
  • Network construction: The analysis treats ties between people at the same institution as 100 separate university network realizations.This restriction permits structural comparisons across institutions.
  • Network variants: Four network variants are considered for each dataset: Full, Student, Female, and Male largest connected components.The gender-specific networks are subsets of Full rather than Student.
  • User attributes: Available user attributes are gender, class year, high school, major, and residence, with Missing used when a characteristic is not volunteered.These categorical variables support comparisons between institutions.
  • Scope and assumption: The study assumes Facebook structural organization reflects offline social organization imperfectly, while recognizing that the degree of imperfection remains an important research issue.The paper’s conclusions apply directly to the September 2005 university Facebook networks and are expected to provide insight into real-world networks as well.

3. Methods

The authors combine local homophily measures and regression models with algorithmic community detection to compare microscopic ties and macroscopic organization across Facebook networks.

  • Community analysis: Communities are detected algorithmically and compared with partitions based on demographic labels to assess correspondence between local homophily and global organization.Assortativity alone may miss community-level effects when a shared attribute is locally predictive but not sufficiently prevalent.
  • Dyad-level methods: Assortativity coefficients quantify homophily for gender, major, residence, class year, and high school at the dyad level.For smaller networks, logistic regression estimates log-odds contributions for shared categorical values, while ERGMs add triangle terms for transitivity.
  • Dyad-level methods: Newman’s assortativity coefficient gives r = 0 for random mixing and r = 1 for perfectly assortative mixing.It is computed from a normalized mixing matrix whose entries count edges between categorical types.
  • Regression and ERGM specification: The models estimate tie propensity for users sharing residence, class year, major, or high school, while gender is handled through single-gender subnetworks.Missing values are not treated as matching categorical values.
  • Community analysis: High school is more dominant in the community structure of large institutions than small institutions, presumably because common-high-school pairs occur more frequently.The paper also compares community organization across Full, Student, Female, and Male networks.

4. Results

Across institutions and network subsets, class year is generally the strongest organizing characteristic, while residence, high school, and major become especially important in particular settings. Dyad-level models and community-level analyses often identify different patterns, showing that social organization depends on analytical scale.

  • For almost all institutions and network subsets, class year produces higher assortativity than the other demographic characteristics.
  • Residence provides the highest assortativity across all four network subsets at Rice, Caltech, and Oklahoma, and in selected networks at other institutions.
  • Dyad-level models often assign the largest coefficient to common high school, whereas community-level correlations usually emphasize class year, with Caltech illustrating the contrast.
  • Community structures are organized overwhelmingly by class year at most institutions, although residence dominates or strongly influences communities at Rice, Caltech, UCSC, Smith, Auburn, and Oklahoma.
  • High school becomes relatively more important in Student networks, with USF, Tennessee, UF, FSU, and GWU located closer to the High School vertex.
  • Male networks vary substantially: residence is especially important at Notre Dame and Michigan, high school influences several institutions, and major distinguishes Texas, Rutgers, and Illinois.

5. Conclusions

The paper compares Facebook networks across 100 American institutions using microscopic measures of homophily and macroscopic community detection. These complementary perspectives reveal institution-specific social organization and support a dataset ensemble for future network studies.

  • 5. Conclusions: The study examines Facebook friendship networks at 100 American institutions using both microscopic and macroscopic perspectives.Assortativity and regression coefficients probe local homophily, while algorithmic community detection provides a complementary larger-scale view.
  • 5. Conclusions: Community structure captures the importance of Caltech’s House system better than the analyzed alternative perspective.The House system is treated as an assumed ground truth in the Caltech networks.
  • 5. Conclusions: The 100 networks form a real-world ensemble that can support comparisons of network-formation models and studies of dynamic processes and community detection.Different adoption rates mean the single-time-point data may represent different stages in online social-network formation.
  • 5. Conclusions: Class year is often important, Houses are important at Caltech, and high school matters more at large than smaller institutions.Smaller institutions typically contain fewer pairs of people from the same high school.
  • 5. Conclusions: Female-only and male-only networks show differences in community structure, with women typically more likely to have friends within a common residence.Male-only community characteristics exhibit wider variation, motivating further investigation with other data and methodologies.
  • 5. Conclusions: Facebook networks imperfectly represent corresponding real-life social networks, so results require comparison with offline networks.Such studies are needed to quantify differences and distinguish causes of friendships from observed correlations.

Appendix A. Tables

The appendix documents network sizes, assortativity calculations, regression models, and community-detection comparisons across institutions and network subsets. It also provides the accompanying tables and specifies how categorical variables and model coefficients are evaluated.

  • Appendix A. Tables: Table A.1 reports node and edge counts for each Facebook network and its subsets.The characteristics include institution identity, Facebook identifier, and numbers of nodes and edges.
  • Appendix A. Tables: Assortativity values are reported for Gender, Major, Residence, Year, and High School across Full, Student, Female, and Male networks.Gender is calculated only for Full and Student networks.
  • Appendix A. Tables: Logistic regression and ERGM tables report edge or density, nodematch, and triangle-related coefficients for categorical attributes.The reported coefficients are statistically different from zero with p-values less than 1 × 10^-4.
  • Appendix A. Tables: Community-detection partitions are compared with categorical partitions using maximum Rand-coefficient z-scores.The appendix organizes networks according to which category gives the highest or second-highest z-score.
Loading 1102.2166v1…