Source-linked AI summary

Behavioral calibration of mobile-phone GPS data for population-representative analyses

Nicolò Alessandro Girardini, Unchitta Kan, Eduardo López, Bruno Lepri, Lorenzo Lucchini, Simone Centellegher

arXiv:2609.01042v1physics.soc-phcs.CYcs.SI

TL;DR

Mobile phone mobility samples can misrepresent demographic and behavioral population distributions, limiting population-level inference. BePop combines census margins with time-use behavioral profiles to estimate person-level calibration weights. Across three metropolitan areas, it consistently improves agreement with representative behavioral distributions and changes downstream mobility indicators, although pandemic-period validation remains uncertain.

  • Problem

    Mobile phone data may be demographically and behaviorally unrepresentative, while existing calibration primarily addresses demographic or geographic composition.

  • Method

    BePop jointly calibrates mobile phone data using census-derived demographic margins and behavioral profiles from representative time-use surveys.

  • Results

    Across three metropolitan areas, BePop consistently improves agreement with representative distributions spanning time allocation, sequence structure, transitions, and motifs, while altering downstream mobility indicators.

  • Takeaways & Limitations

    Behavioral representativeness complements demographic calibration because demographically calibrated mobility data may still contain behavioral biases affecting downstream estimates.

  • Takeaways & Limitations

    Calibrated estimates depend on reference-data quality and temporal consistency in mobile phone observations, especially during abrupt behavioral change.

Abstract

from arXiv · show

Mobile phone mobility data have transformed the study of human behavior, but demographic and behavioral biases can compromise their representativeness and distort population-level inference. Existing calibration approaches primarily address demographic and geographic representativeness, leaving behavioral discrepancies largely uncorrected. Here we introduce the Behavioral Population (BePop) framework, which jointly calibrates mobility data to representative demographic and behavioral distributions using census data and time-use surveys. BePop embeds mobility sequences into behavioral profiles and estimates person-level weights that align both population composition and daily activity patterns. Across three U.S. metropolitan areas, the framework consistently improves agreement between GPS-derived mobility and representative behavioral distributions, including time allocation, activity transitions, and mobility motifs. Calibration also substantially alters downstream mobility indicators, demonstrating that behavioral biases can propagate into commonly used mobility measures. Our results establish behavioral representativeness as a critical complement to demographic calibration and provide a general framework for population-representative mobility inference.

Introduction

Mobile phone mobility data offer detailed behavioral measurements but can misrepresent populations because of demographic, geographic, measurement, and preprocessing biases. BePop extends calibration to jointly align mobility samples with representative demographic and behavioral distributions.

  • Mobile phone records enable high-resolution, longitudinal behavioral analysis and near-real-time monitoring.
  • Non-probability sampling and demographic differences in smartphone ownership, usage, and data sharing can make mobile phone samples unrepresentative.
  • GPS drift, incomplete points-of-interest data, and preprocessing differences can produce mobility assignments that diverge from real visit patterns.
  • BePop jointly calibrates socio-demographic composition and behavioral composition using census-derived population margins and representative time-use benchmarks.
  • Across three U.S. metropolitan areas, the evaluation tests alignment with ATUS benchmarks and examines effects on downstream mobility indicators and longitudinal analyses.
  • The framework is presented as flexible, compatible with weighted estimation and uncertainty quantification, and available through an open-source Python implementation.

Results

BePop assigns calibration weights by combining demographic targets with behavioral profiles derived from time-use data. Across metropolitan areas, it improves agreement with ATUS distributions and changes both behavioral and downstream mobility estimates, while pandemic validation remains uncertain.

  • The BePop calibration framework: BePop produces individual weights by jointly aligning mobile phone users with demographic distributions and behavioral distributions conditioned on demographic groups.
  • BePop in practice: an application to US metropolitan areas: The application uses longitudinal GPS traces from Phoenix, Boston, and New York with ACS demographic data and ATUS behavioral reference data.
  • BePop in practice: an application to US metropolitan areas: Daily mobility and ATUS records are represented as sequences of 48 half-hour bins, with each bin assigned its main activity.
  • Evaluating the effectiveness of BePop calibration: The evaluation compares time allocation, sequence structure, activity transitions, and daily behavioral motifs against ATUS distributions.
  • Evaluating the effectiveness of BePop calibration: Calibration improved time-use, activity-count, entropy, transition, and motif alignment, while reciprocity and turnover were already well aligned before calibration.
  • Evaluating the effectiveness of BePop calibration: 9.71% ± 0.74 reduction in mean Jensen–Shannon distance was observed across cities and behavioral dimensions, including 17.00% ± 0.16 for motif distributions.
  • BePop calibration on longitudinal mobility data: Behavioral-profile distributions showed reasonable ATUS agreement, with a dominant profile covering more than 70% of daily assignments on average.
  • BePop calibration on longitudinal mobility data: Pandemic-period improvement was 0.64% ± 3.86 relative to unweighted data, but large standard errors prevent determining whether the improvement persists.

Discussion

BePop addresses behavioral imbalances that demographic calibration alone does not directly correct. Across three metropolitan areas, it improves agreement with representative behavioral distributions and changes downstream mobility estimates, while embedding choice and data quality constrain performance.

  • Demographic calibration does not directly address behavioral imbalances caused by heterogeneous device usage.
  • BePop jointly aligns mobile phone data with population-representative socio-demographic and behavioral benchmarks using census margins and time-use profiles.
  • Across three metropolitan areas, calibration improves agreement in time allocation, sequence structure, transitions, and behavioral motifs.
  • BePop weighting changes radius-of-gyration and workplace co-location estimates before and during the COVID-19 pandemic.
  • Application-specific embeddings perform best for targeted behavioral dimensions, while richer embeddings may balance generality and calibration performance.
  • Calibration quality depends on representative reference data, stable behavioral targets, and sufficiently consistent mobile-phone observations.

Methods

BePop combines ACS demographic targets with ATUS behavioral profiles to calibrate GPS mobility sequences. The implementation maps both data sources into shared daily activity representations, assigns users to demographic and behavioral groups, and estimates corresponding adjustment factors.

  • ACS data provide age and income distributions for demographic assignment and CBSA-level calibration targets.
  • ATUS supplies nationally representative 24-hour activity diaries, respondent weights, and metropolitan-area behavioral references.
  • GPS traces are processed into stops, inferred home and work locations, POI-based activity categories, and shared 48-bin daily sequences.
  • The framework defines demographic strata using age and income, with alternative attributes possible when supported by the application or reference data.
  • Behavioral embeddings summarize activity durations, transitions, and sequence structure; cosine distances and K-Medoids produce four recurring behavioral profiles.
  • Mobile sequences are assigned to ATUS profiles through nearest-neighbor matching, after which demographic and within-stratum behavioral adjustment factors are estimated.
  • Longitudinal evaluation uses users observed before and after March 11, 2020, creating a persistent-user cohort that may introduce survivorship bias.

Author contributions statement

The authors divided the work across conception, preprocessing, experimentation, interpretation, writing, and critical manuscript revision.

  • The team shared responsibility for study conception, experiments, figure preparation, interpretation, manuscript writing, and substantive revision.

Data and Code availability

The study combines ACS, ATUS, and privacy-enhanced Cuebiq GPS data for three U.S. metropolitan areas. The materials describe survey construction, geographic and activity preprocessing, privacy safeguards, and data-access resources.

  • The findings data are accessible through Cuebiq’s Data for Good initiative, while ACS and ATUS are freely available and the implementation is open source.
  • ATUS time diaries record activities, durations, locations, and whether others were present over a 24-hour reporting period.
  • ATUS respondent weights account for survey design across weekdays and weekends to represent a typical week.
  • The behavioral reference uses ATUS records from 2004–2019 for Boston, New York, and Phoenix, totaling about 13,600 respondents.
  • ACS data provide age and income estimates at Census Block Group level and calibration marginals at Core-Based Statistical Area level.
  • Stops require at least five minutes within a 65-meter radius, are clustered with DBSCAN, and may be linked to OpenStreetMap POIs within 65 meters.
  • Cuebiq GPS data cover January–June 2020 in Boston, Phoenix, and New York, with approximately 760,000 consistently active users after filtering.

D. Constructing common activity space across ATUS and mobile phone data

The framework constructs a shared activity representation by temporally aligning ATUS diaries with GPS-derived stop sequences and aggregating locations into comparable categories.

  • ATUS and mobile-phone sequences are aligned through a mapping between survey locations and OSM-attributed stop locations.Travel and gaps are represented as Unspecified places in the aligned data.
  • Sensitive points of interest are excluded from the location mapping according to Cuebiq policies.
  • Individual points of interest are aggregated into one POI category to reduce unnecessary variability in sequence representations.The aggregation treats visits such as library stops as general POI visits.
  • Daily sequences follow a 4 AM–4 AM cycle divided into 30-minute activity intervals.
  • The shared representation preserves comparable activity categories while standardizing transportation and travel coding across datasets.

E. Constructing daily motifs

Daily mobility sequences are converted into structural motifs that capture transition organization independently of time spent, enabling comparisons between ATUS and mobile-phone routines.

  • Calibration outcome: BePop improves alignment for both embedding-level behavioral dimensions and higher-level structural representations.
  • Constructing daily motifs: Each daily sequence becomes a directed graph whose nodes are visited location categories and whose edges represent consecutive-location transitions.
  • Constructing daily motifs: Grouping isomorphic graphs yields structural motifs that preserve connectivity patterns while ignoring specific location identities.
  • Comparing motifs: The analysis compares prevalent motif distributions between ATUS and mobility datasets to assess recovery of daily routine structure.
  • Calibration procedure: Behavioral profiles and calibration weights are constructed through a two-stage procedure aligning ACS demographic and ATUS behavioral targets.

A. Part I: Design

The design combines CBSA-level demographic calibration with behavioral embeddings and clustered activity profiles, using ACS and ATUS reference distributions to construct calibrated mobility weights.

  • Calibration scope: Calibration is performed separately at the CBSA level, the geographic unit used for mobility-metric analysis.
  • Behavioral profiles: Activity sequences are embedded as vectors capturing durations, transitions, and overall sequence structure, then clustered into behavioral profiles using K-Medoids.
  • Behavioral profiles: Mobility sequences are assigned to ATUS-derived profiles by majority vote among k = 10 nearest neighbors, with ambiguous or outlier sequences left unassigned.
  • Demographic attributes: Age and household income define demographic strata, with age grouped into four life-course categories and income divided into CBSA-specific quartiles.
  • Demographic attributes: Mobility users receive probabilistic demographic assignments from ACS marginal distributions associated with their inferred home CBG.Age and income are sampled independently because joint CBG-level distributions are unavailable.
  • Cluster selection: The selected K = 4 has no clear optimum across sensitivity metrics, but improvements diminish beyond four clusters and medoid separation decreases.

B. Part II: Weight creation

BePop creates person-level weights by first aligning mobile-phone users with ACS demographic targets and then aligning behavioral profiles with ATUS distributions. The resulting weights preserve demographic alignment while calibrating behavioral composition, including treatment of users without assigned profiles.

  • Weight construction: BePop estimates calibration weights from demographic strata and behavioral profiles, using ACS population margins and ATUS behavioral shares as targets.The final user weight combines demographic and behavioral adjustment factors and scales representation to the target population.
  • ACS alignment: IPF alternately scales the sample joint distribution to available age and income margins until convergence.The weighted sample reproduces the ACS age and income margins exactly.
  • ATUS alignment: ATUS adjustment factors are calculated within each demographic stratum to align conditional behavioral-profile distributions with the ATUS reference.When both stages use the same stratification, the ACS and ATUS adjustments can be computed independently and in either order.
  • Joint calibration: The weighted mobility sample reproduces the ACS demographic distribution and the ATUS behavioral distribution simultaneously.The behavioral-profile distribution matches ATUS within every demographic stratum, while behavioral adjustment leaves ACS demographic alignment intact.
  • Unassigned users: Unassigned users receive zero behavioral adjustment, while their demographic weight is redistributed proportionally among assigned users in the same stratum.This preserves the stratum-level population constraint while excluding unassigned users from behavioral alignment.
  • Unassigned users: Across all three metropolitan areas, many unassigned users spent most of the day at an unspecified place and little or no time at home.These sequence patterns were treated as likely unrepresentative, motivating their exclusion from the behavioral alignment step.
  • Behavioral representations: Richer behavioral representations provided more robust general-purpose calibration, whereas tailored embeddings aligned best with their specific behavioral dimensions.Time Use-only embeddings achieved the strongest alignment for time-use metrics but lower alignment on other evaluated distributions.
  • Calibration stages: Behavioral-only calibration achieved alignment comparable to full calibration but could not recover the demographic composition of the CBSA.The limitation arises because behavioral-only weighting does not incorporate socio-demographic information.

C. Improvement in demographic representativeness by calibration stages

The calibration stages differ in which distributions they improve: ATUS-only adjustment targets behavioral alignment but does not reliably improve demographic representativeness, whereas BePop preserves ACS alignment while adding behavioral adjustment. The full framework therefore aligns both demographic and behavioral targets, with calibrated diagnostics close to their targets.

  • Calibration-stage comparison: Behavioral-only calibration cannot improve demographic alignment, whereas the full method aligns well with the demographic reference.The comparison evaluates the original sample, behavioral-only calibration, and BePop calibration against the reference population composition.
  • ACS target construction: IPF matches ACS age and income marginals, but its joint age-income cell values are outputs rather than externally calibrated joint targets.The ACS data provide marginal constraints rather than a ground-truth joint distribution.
  • Calibration-stage comparison: ATUS-only adjustment does not systematically improve demographic alignment and can move some joint age-income cells farther from the reference.BePop instead brings every joint age-income group to within 1% relative difference while retaining the demographic alignment after ATUS adjustment.
  • Calibration diagnostics: All calibrated statistics were within 1 percentage point of their targets.This diagnostic evaluates sample income, age, and cluster marginals before and after calibration.
  • Demographic diagnostics: The uncalibrated sample under-represented low-income individuals at 12.5% MAPE and over-represented high-income individuals at 12.8% MAPE.These discrepancies were assessed using MAPE for sample income, age, and cluster marginals before and after calibration.
  • Behavioral diagnostics: The framework reduced the home-work-errands pattern’s MAPE from 42.8% before calibration to 0.6% after calibration.The uncalibrated sample notably under-represented this behavioral pattern.
  • Weight diagnostics: In Phoenix, behavioral assignment excluded about 21.1% of users, while the effective final sample size was 90,519, approximately 65% of the original sample.Among assigned users, the mean and median final weights were 33.6 and 30.8, with a design effect of 1.22.
  • Replicate diagnostics: Final weights varied substantially within individuals across replicates, with around 72% of users having a coefficient of variation above 0.50.The paper emphasizes that weights target aggregate sample distributions rather than individual-level stability.

S6. CITY-WISE CALIBRATION RESULTS

Across Phoenix, Boston, and New York, BePop improves alignment between mobility-derived behavior and ATUS distributions. Alignment is closer in Phoenix and Boston than New York, whose scale and characteristics may require more tailored calibration.

  • All three cities show improved alignment with ATUS users’ behavior after calibration.
  • Phoenix and Boston achieve closer alignment than New York.
  • Reciprocity aligns less well than the other metrics in Phoenix.
  • Transition disalignment is reflected in the motif distributions.
  • New York’s larger population and metropolitan characteristics may require more specific calibration tailoring.

S7. LONGITUDINAL ANALYSIS WITH THE BEPOP FRAMEWORK

The longitudinal analysis fixes each user’s reference demographic and behavioral assignment while independently inferring daily behavioral profiles and calibrating daily observations. The resulting distributions remain representative of ATUS while responding to temporal behavioral shifts, including lockdown-related stay-at-home behavior.

  • The analysis evaluates whether selecting a single day per user supports longitudinal calibration while allowing behavioral shifts such as COVID-19 lockdowns.
  • The study selects 15,000 users per city with the broadest day coverage and retains these users across pre-lockdown and lockdown periods.
  • Reference demographic assignments, behavioral profiles, and associated weights remain fixed, while daily behavioral profiles are independently inferred and daily weights are computed for each observed population.
  • When behavioral patterns remain stable, weighted daily aggregation recovers a distribution close to the reference composition.
  • Unassigned users retain reference weights and comprise less than 10% of daily observations across cities.
  • Pre-pandemic profiles align highly with ATUS, while lockdown increases profile 1, corresponding to stay-at-home behavior.
  • BePop reduces JS distance by 12.81% ± 0.59 in Boston and 19.34% ± 0.37 in New York before the pandemic.

S8. SUPPORTING TESTS FOR BEHAVIORAL METRICS

Supporting tests compare BePop, demographic calibration, and uncalibrated mobility against ATUS across behavioral dimensions and motifs. BePop improves alignment, with the largest gains in time use, structural metrics, and motif distributions.

  • Demographic calibration produces KL and JS values nearly identical to the uncalibrated sample.
  • BePop significantly reduces JS distance and KL divergence across behavioral dimensions.
  • 9.71% ± 0.74: BePop’s relative reduction in combined JS distance versus the uncalibrated sample.
  • Transitions improve less than time use and structural metrics, potentially because their embedding representation is high-dimensional and may dilute behavioral information.
  • 17.00% ± 0.16: motif-distribution discrepancies are reduced despite motifs not being represented in the embedding space.
Loading 2609.01042v1…