Source-linked AI summary
Measuring large-scale social networks with high resolution
Arkadiusz Stopczynski, Vedran Sekara, Piotr Sapiezynski, Andrea Cuttone, Mette My Madsen, Jakob Eg Larsen, Sune Lehmann
TL;DR
The paper addresses the limited ability of siloed, single-channel datasets to capture social interactions that span communication modes and evolve over time. It presents the Copenhagen Network Study, a longitudinal, interactive, multi-channel study of about 1,000 participants using smartphones and related data sources. Early results and project perspectives emphasize that a single data stream rarely provides a comprehensive picture, while larger and longer studies impose substantial participant and operational costs.
Problem
Single-channel datasets provide limited evidence for understanding social networks because interactions span multiple communication channels and change over time.
Method
The Copenhagen Network Study collects longitudinal, high-resolution data from about 1,000 participants across smartphones, questionnaires, online social networks, communication, location, and background information.
Results
The project reports early data-analysis results showing that a single data stream rarely supplies a comprehensive picture of human interactions, behavior, or mobility.
Takeaways & Limitations
The study provides a multi-layered testbed in which data channels can be examined separately and jointly by researchers from different fields.
Takeaways & Limitations
The study population is sampled from a finite participant set, and the relatively small network has boundary links to the surrounding society.
Abstract
from arXiv · showhide
This paper describes the deployment of a large-scale study designed to measure human interactions across a variety of communication channels, with high temporal resolution and spanning multiple years - the Copenhagen Networks Study. Specifically, we collect data on face-to-face interactions, telecommunication, social networks, location, and background information (personality, demographic, health, politics) for a densely connected population of 1,000 individuals, using state-of-art smartphones as social sensors. Here we provide an overview of the related work and describe the motivation and research agenda driving the study. Additionally the paper details the data-types measured, and the technical infrastructure in terms of both backend and phone software, as well as an outline of the deployment procedures. We document the participant privacy procedures and their underlying principles. The paper is concluded with early results from data analysis, illustrating the importance of multi-channel high-resolution approach to data collection.
2 Introduction
The paper argues that shallow, siloed data cannot adequately capture human social networks, which span multiple communication channels and change over time. It presents the Copenhagen Network Study as an interactive, longitudinal, interdisciplinary testbed for collecting and jointly analyzing high-resolution data.
- Motivation: Human interactions span face-to-face communication, calls, messages, social networks, and email, so single-channel datasets are insufficient for understanding social networks.
- Motivation: The study is designed to adapt data collection when existing measurements lack the temporal resolution needed for particular research questions.
- Motivation: Longitudinal data are needed because human behavior, networks, and thinking change over months and years.
- Study design: The Copenhagen Network Study combines questionnaires, online social networks, and smartphone data to create multilayered views of individuals, networks, and environments.The deployments involved about 1,000 participants and support separate or joint analysis by researchers from different fields.
- Paper scope: The paper reviews related work, motivation, experimental planning, data collection, privacy and consent practices, initial results, and future directions.
3 Related Work
Computational social science uses large-scale digital traces to study individuals, groups, and societies. Its methods extend questions previously addressed mainly through self-reports or direct observation, while creating new data and privacy challenges.
- Computational social science studies individuals and groups using data such as phone records, GPS traces, transactions, webpage visits, emails, and social-network data.
- These data-driven analyses address dynamics in work groups, face-to-face interaction, mobility, and information spread.
- The related-work section reviews central data-collection and analysis methods alongside principles for privacy and data treatment.
3.1 Data collection
Prior large-scale behavioral studies collected data through call records, dedicated sensing platforms, and smartphones, but longitudinal rich-data deployments remain logistically difficult. Call records provide useful proxies while missing some face-to-face interaction.
- Call detail records contain phone calls and messages collected by mobile operators and can proxy mobility and social interaction.
- Face-to-face interaction may be missed by call records and online social-network data, motivating additional sensing channels.
- Sensor-driven collection systems range from dedicated software to complete platforms, including ContextPhone, SocioXensor, MyExperience, and Funf-related tools.
- Reality Mining collected data from 100 mobile phones for nine months, while Social fMRI followed 130 smartphone users for 15 months.
3.2 Data analysis
Related computational social science studies use communication, mobility, encounter, and mobile-sensing data to analyze predictability, social ties, temporal structure, information spread, health, and socioeconomic outcomes. Their datasets range from hundreds of thousands to millions of users and trips.
- Human mobility studies report high regularity and predictability, including a 93% upper bound of mobility predictability for most users.
- Face-to-face interaction data support analyses of social ties, organizational rhythms, meeting locations, and temporal scales in social networks.One cited study identified a natural four-hour time scale for face-to-face social networks.
- Communication studies examine weak ties, distance-dependent link probabilities, reciprocity, degree, active-tie limits, and social communication strategies.
- Mobile and face-to-face sensing has been applied to earthquake detection, relief estimation, fitness interventions, illness-related features, political opinions, app use, and virus spread.
- Network structure and communication patterns have been linked to movement during events and emergencies, localized information spread, job satisfaction, group work quality, and economic opportunities.
- Research on network influence highlights that homophily and influence can be confounded, prompting methodological criticism and statistical frameworks for separating them.
3.3 Privacy
The study treats participant data as sensitive and privacy as a fundamental requirement. It situates privacy protection within risks of re-identification, inference, and misuse, alongside technical approaches that trade data utility against protection.
- Participant data can reveal routines, friendships, identities, and network positions, making privacy protection fundamental.Risks include unauthorized analysis, re-identification from public datasets, record linkage, and inference from network structure.
- Personally Identifiable Information includes data such as names, addresses, education, employment, and financial status.De-identification can remove, aggregate, or transform such information, with a resulting privacy–utility trade-off.
- Noise perturbation can produce privacy-preserving statistics, while homomorphic encryption supports computation on encrypted data.These approaches aim to limit exposure of sensitive information during analysis.
- Information-flow controls, auditing, data expiration, and watermarking restrict, track, invalidate, or trace data use.Together, these mechanisms address sharing, access duration, and leak identification.
4 Motivation
The Copenhagen Networks Study is motivated by the limits of single-channel, unevenly sampled data for studying multiplex human networks. It therefore emphasizes deep, high-resolution, longitudinal measurement and controlled experimentation.
- Single-source studies can produce incomplete and biased views because social interactions span calls, face-to-face meetings, emails, and online platforms.The study frames multi-channel collection as necessary for interdisciplinary analysis of the same population and period.
- Uneven sampling can reduce data quality and introduce selection bias, including distorted estimates when analyses retain only high-frequency users.The study aims for evenly sampled, high-quality data without discarding participants, while recognizing this goal is difficult.
- High-resolution longitudinal data are needed because networks and behavior change across both months-to-years and short daily intervals.Low sampling frequency can miss short-term dynamics, while short observation periods can miss long-term change.
- Dense multiplex networks require models that account for time-varying active links, since information may traverse links active at successive times.The paper identifies understanding such dynamics as beyond current theory and needing methodological expansion.
- Controlled experiments and smartphone surveys support investigation of causality and confounding factors in network settings.Participants can be divided into sub-populations and exposed to distinct stimuli.
- Real-time processing enables participant-facing feedback and study of self-awareness, behavior change, and engagement.The study also uses dynamic network measurements to examine diffusion of behavior, information, and infectious disease across link types.
- Network structure matters for societal countermeasures because the appropriate responses differ with the details of that structure.
- Incomplete network data can make overlapping communities appear disjoint, so accurate full-network descriptions are needed to study sampling effects.This motivates high-quality, high-resolution datasets that can inform theories of incomplete data.
5 Data Collection
The Copenhagen Networks Study addresses single-modality data by combining questionnaires, Facebook, and smartphone sensing in a dense participant population. Its sensors collect at fixed intervals to reduce uneven sampling, while density also heightens privacy concerns.
- The study combines questionnaires, Facebook, and smartphone measurements of location, telecommunication, and face-to-face interactions.These sources provide context for building networks and interpreting social phenomena.
- Fixed-interval sensor collection mainly overcomes the uneven sampling problem associated with activity-dependent records.The study’s dense participant population helps address missing data but allows information about absent participants to be inferred from others.
- The paper describes technical challenges and solutions for multi-channel data collection in the 2012 and 2013 deployments.
5.1 Data Sources
The study collected complementary quantitative, qualitative, online, survey, mobile, and campus connectivity data across its deployments. These sources supported participant feedback, interdisciplinary analysis, and grounding of mathematical models.
- The deployments combined questionnaires, Facebook, mobile sensing, anthropological fieldwork, and campus WiFi data.
- The 2012 survey contained 95 optional questions on socioeconomic factors, working habits, and personality traits.
- The 2013 deployment asked 310 questions covering personality, self-esteem, narcissism, life satisfaction, locus of control, and loneliness.A customizable web application controlled question branching, progress saving, privacy, and data handling; participation was required.
- Facebook authorization was optional, and data were collected as friendship-graph or broader snapshots every 24 hours across deployments.
- Smartphone data were collected with a modified Funf framework on Android phones handed to participants.The deployments used Samsung Galaxy Nexus phones in 2012 and LG Nexus 4 phones in 2013.
- The 2013 deployment linked phones to students through OAuth2 registration and token-based data identification rather than device IMEI tracking.
- An anthropologist observed a randomly selected group of approximately 60 students from August 2013 to August 2014.Participant observation generated qualitative data about rationales underlying group formation.
- Qualitative data supplied participant feedback, increased interdisciplinary links between computational and social science, and grounded mathematical modeling.They helped improve services, relate observations to mobile-sensing data, and address why network changes occur.
5.2 Backend System
The backend evolved from a 2012 system focused on testing sensor-driven collection to an extensible 2013 framework for collection, sharing, and analysis.
- The 2013 backend was redesigned as an extensible framework for data collection, sharing, and analysis.
- The 2012 backend used Django, stored multi-source participant data in CouchDB, and exposed mobile sensing data through an API.
- 2012 performance was inadequate for real-time application access, mainly because of inefficient de-identification.
- The 2013 system used openPDS with platform, services, and applications layers, including identity management and user authorization.
5.3 Deployment Methods
The deployments used different enrollment and handout strategies as the study expanded, while emphasizing participant engagement, data quality, and operational learning.
- The two deployments pursued the same purpose but differed substantially in scale and participant enrollment and engagement methods.
- SensibleDTU 2012: Approximately 200 phones were available in 2012, so distribution focused on selected majors with sufficient enrollment to maximize social-network coverage and density.
- SensibleDTU 2012: The 2012 handout included participant instruction, informed consent, questionnaires, and a symbolic cash deposit intended partly to encourage phone care.
- SensibleDTU 2012: The 2012 process required considerable organization and missed students’ earliest interactions because phones were distributed 3–4 weeks into the semester.
- SensibleDTU 2013: The 2013 deployment distributed 1 000 phones and began recruitment earlier through university acceptance materials and online registration.
- SensibleDTU 2013: The 2013 process moved consent and questionnaires into online registration and pre-installed software to streamline handout, while participants completed phone setup.
- Recommendations: The authors recommend early engagement, age-appropriate communication channels, and clear malfunction procedures for large-scale handouts.
6 Privacy
The study treats privacy as alignment between participants’ understanding and actual data use, supported by tools for information, control, consent, and data access.
- Privacy is defined as the difference between what users understand and consent to regarding their data and what actually happens to it.
- The proposed privacy approach combines informing users, enabling control, and allowing them to verify whether changes have the expected effect.
- Participants rarely asked privacy questions and often accepted presented terms without thorough analysis, placing responsibility on researchers to increase awareness and empowerment.
- 2012 consent: The 2012 deployment recorded consent forms with participant identifiers, timestamps, and the full text shown to each user.
- 2013 privacy infrastructure: The 2013 system separated sensitive personally identifiable information from user data and supported distinct access levels for participants, researchers, and developers.
- 2013 privacy infrastructure: HTTPS, short-lived OAuth2 tokens, and versioned consent documents protected communications and preserved which consent text each user accepted.
- User access: Users could access their data through the research API and a data viewer designed to simplify querying.
- User engagement: The project engaged participants through blog posts, student presentations, and responses to privacy questions on Facebook.
7 Results
The study combines multiple behavioral channels and high-resolution longitudinal data to reveal temporal, spatial, and channel-specific patterns in human social interaction. Early results show that resolution and multimodal sensing materially shape the networks and mobility patterns observed.
- Overview: The 2012 deployment provides longitudinal results from face-to-face interactions, WiFi, location, calls, SMS, online networks, and participant traits.The paper presents these channels as complementary views of individuals, networks, and environments.
- Bluetooth and Social Ties: Face-to-face networks change substantially across the day and week, reflecting academic schedules, social patterns, and different interaction intensities.Participants meet in the morning, interact within study lines during classes, and connect across majors in the evening.
- Bluetooth and Social Ties: Larger temporal windows blur social interactions, while rescaled network distributions diverge across resolutions, suggesting distinct dynamics at different timescales.Degree distributions shift upward and edge-weight distributions shift downward as the temporal window grows.
- WiFi as Additional Channel for Social Ties: 98% of Bluetooth-derived face-to-face meetings included at least one common WiFi access point, while matching the strongest access point offered high recall and positive predictive value.The strongest-access-point measure is identified as a candidate proxy for face-to-face interactions, although further validation is required.
- Location and Mobility: Almost 90% of location samples had accuracy better than 40 meters, while the radius-of-gyration distribution peaked around 102 km and had a fat tail.The authors caution that the mobility density estimate suffers from few samples and suggest CDR-only studies may underestimate travel outside covered areas.
- Call & Text Communication Patterns: The number of unique SMS and voice-call contacts correlated at 0.75, but their average contact-set similarity was only 0.37, indicating distinct communication purposes.Calls and SMS also show different hourly patterns and differ from weekday face-to-face interactions, which are primarily university-related.
- Call & Text Communication Patterns: The three networks formed by voice calls, text messages, and face-to-face meetings provide very different views of participants’ social interactions.The online network likewise contains social-interaction information missed by the face-to-face network.
8 Perspectives
The study argues that future human-behavior research will increasingly use existing personal data, while requiring reusable infrastructure to address incomplete, costly, and difficult-to-link data. Its preliminary results show that single data streams rarely provide a comprehensive view, and scaling participation, duration, channels, and resolution remains expensive.
- 8 Perspectives: The scientific community should develop reusable solutions spanning privacy policies, deployment procedures, data-collection technologies, and analysis methods.The paper presents these as responses to the growing complexity and scale of social-system studies.
- 8 Perspectives: Single data streams rarely provide a comprehensive picture of human interactions, behavior, or mobility.The authors describe this result as preliminary because the project was intended to span multiple years.
- 8 Perspectives: Larger studies are becoming expensive as they increase participant numbers, duration, observed channels, or temporal resolution.Participant inconvenience includes reduced battery life, questionnaire burden, and reduced privacy; material incentives and study administration may become infeasible at scale.
- 8 Perspectives: Future studies are expected to access existing personal data, but linking multiple streams remains difficult because of data silos.The paper identifies mobility, social-network, and self-tracking data as examples of already available personal information.
- 8 Perspectives: Research testbeds such as the Copenhagen Networks Study are needed to examine how incomplete, user-controlled data affect results and to address replicability, reproducibility, and selection bias.The authors frame the study as a testbed for understanding the consequences of increasingly relying on existing personal data.