Source-linked AI summary
Evidence for a Conserved Quantity in Human Mobility
Laura Alessandretti, Piotr Sapiezynski, Vedran Sekara, Sune Lehmann, Andrea Baronchelli
TL;DR
Human mobility research has struggled to reconcile repeated use of familiar locations with steadily growing exploration. Analyzing multi-year traces, the paper finds that individuals maintain a conserved set of familiar locations while their mobility patterns evolve smoothly.
Problem
Existing mobility models do not account for the evolution of individuals’ activity sets over time.
Method
The paper analyzes long-term mobility traces and introduces a limited-memory exploration and preferential return model incorporating activity-set evolution.
Results
The average individual location capacity remains constant over time, while individuals’ weekly net gain of locations equals zero.
Takeaways & Limitations
A conserved familiar-location capacity helps reconcile exploration with repeated visits and motivates mobility models that represent evolving activity sets.
Takeaways & Limitations
The EPR model does not account for the evolution of the activity set.
Abstract
from arXiv · showhide
Recent seminal works on human mobility have shown that individuals constantly exploit a small set of repeatedly visited locations. A concurrent literature has emphasized the explorative nature of human behavior, showing that the number of visited places grows steadily over time. How to reconcile these seemingly contradicting facts remains an open question. Here, we analyze high-resolution multi-year traces of $\sim$40,000 individuals from 4 datasets and show that this tension vanishes when the long-term evolution of mobility patterns is considered. We reveal that mobility patterns evolve significantly yet smoothly, and that the number of familiar locations an individual visits at any point is a conserved quantity with a typical size of $\sim$25 locations. We use this finding to improve state-of-the-art modeling of human mobility. Furthermore, shifting the attention from aggregated quantities to individual behavior, we show that the size of an individual's set of preferred locations correlates with the number of her social interactions. This result suggests a connection between the conserved quantity we identify, which as we show can not be understood purely on the basis of time constraints, and the `Dunbar number' describing a cognitive upper limit to an individual's number of social relations. We anticipate that our work will spark further research linking the study of Human Mobility and the Cognitive and Behavioral Sciences.
Methods
The study combines four mobility datasets spanning diverse populations and collection periods, using mobile-phone, WiFi/GPS, anonymized GPS, and GSM location traces with documented estimation procedures and privacy protections.
- Reality Mining dataset: 94 MIT subjects were monitored for nine months using an application that continuously logged cell-tower location data at fixed-rate sampling.The Reality Mining project was conducted from 2004–2005; 68 subjects were colleagues and 26 were incoming business-school students.
- Copenhagen Networks Study dataset: 851 Danish students’ positions were estimated from smartphone WiFi and GPS data, with location error below 50 meters in 95% of cases.The Copenhagen Networks Study ran between September 2013 and September 2015 and also collected calls, SMS activity, and survey data.
- MDC dataset: The Lausanne campaign used GSM data from approximately 185 volunteers because its sampling frequency exceeded that of the experiment’s GPS data.Data were collected between October 2009 and March 2011 from heterogeneous volunteers in the Lake Geneva region.
Data Availability Statement
Data access varies across the four datasets: CNS, MDC, and Lifelog data are restricted or unavailable publicly, while the Reality Mining Dataset is available from MIT’s Human Dynamics Lab.
- CNS data are not publicly available because of privacy, European Union regulations, and Danish Data Protection Agency rules.
- CNS access is limited to researchers meeting confidential-data criteria, signing a confidentiality agreement, and working under Copenhagen supervision.
- MDC data are licensed and not publicly available, but eligible institutions may request access from Idiap Research Institute.
- Lifelog raw data are unavailable publicly under Sony Mobile’s Privacy Policy, while derived supporting data can be requested from the corresponding authors.
- The Reality Mining Dataset is available from the MIT Human Dynamics Lab.
Correspondence … 1.1 Data pre-processing
The supplementary methods describe how four datasets were converted into location-interval records, using dataset-specific collection, location inference, cleaning, and temporal aggregation procedures. All datasets retained only intervals longer than 10 minutes.
- 1.1 Data pre-processing: All four datasets were converted into sequences of individuals’ pauses at locations, represented by user, interval start, interval end, and location.The data-preprocessing section describes collection and preprocessing for these records.
- 1.1 Data pre-processing: Intervals longer than 10 minutes were retained across all datasets.Dataset characteristics are summarized in the supplementary materials.
- Lifelog dataset: Lifelog stop-locations were inferred by grouping sequences of locations within dmax, with main-text results using dmax = 50m.The reported results also held for dmax = 30m.
- Lifelog dataset: Lifelog time-coverage changes were addressed by selecting users without coverage changes, while robustness also held after temporal down-sampling.Main-text results used method (a), and method (b) down-sampled weeks to the lowest weekly time-coverage.
- CNS mobility dataset: CNS mobility combined Wi-Fi sampled every ∼15s with high-spatial-resolution GPS, estimating access-point positions with error below 50 meters in 95% cases.Mobile or displaced access points were discarded.
- CNS mobility dataset: CNS locations were defined as connected components of an access-point distance graph, using d = 5m; most locations contained fewer than 10 access points.Dense areas such as university campuses contained up to ∼1000 access points in one location.
- MDC and RM mobility datasets: CNS observations were aggregated into 1 min bins by selecting the most likely location, whereas MDC and RM analyses used GSM data sampled every 60 seconds or without stated sampling frequency.MDC and RM data-collection procedures are referenced elsewhere in the supplementary notes.
1.2 Comparison with previous research
The datasets exhibit statistical properties consistent with prior human-mobility research, including a power-law rank-frequency distribution of location visits.
- The datasets display statistical properties consistent with previously analyzed human-mobility data.
- Visitation frequency is defined as the fraction of visits to a location and decreases with rank r as r−ζ, with ζ ∼1.
- The rank-frequency result is consistent with prior findings that f(r) ∝1/r.
1.3 Robustness Tests
Robustness tests show that location capacity remains temporally constant across location definitions and time windows, while activity sets evolve independently of measurement start times. Exploration remains sublinear, and randomized sequences yield higher capacities than observed behavior, indicating time constraints alone do not explain individual capacity.
- Capacity constancy: Location capacity remains constant over time across location definitions and window sizes W.This robustness is supported across all datasets and tested hypotheses cannot be rejected at α = 0.05.
- Capacity constancy: The weekly net gain of locations is zero, with σG,i/⟨Gi⟩ > 1 for a large majority of individuals across definitions and windows.Net gain compares locations added and removed over dt = 1 week.
- Capacity saturation: ⟨C⟩ ∼ 25 locations after accounting for differences in data collection, and individual capacities are distributed homogeneously around the mean.The average capacity saturates as the time-window W increases.
- Measurement robustness: Activity-set evolution is unaffected by the measurement start time, with J(t, γ) = J(γ), indicating equilibrium behavior in the monitored individuals.The similarity decay coefficient does not change substantially with starting time t.
- Exploration robustness: Location discovery grows sublinearly under alternative location definitions and independently of users’ age in the MDC dataset.The growth coefficient αi has a positive relation with age, with Pearson correlation ρ = 0.2 and p-value = 0.008.
- Time-constraint test: Randomized temporal sequences produce higher capacities than real sequences for both global and local randomization, with the distributions differing at p < α.The comparison uses 100 randomizations and Kolmogorov–Smirnov test statistics.
1.4 The EPR model with memory
Existing EPR variants reproduce the conserved size of individual capacity but fail to capture the evolution of activity sets. The paper therefore introduces a limited-memory EPR model that retains the EPR exploration strategy while restricting memory of past locations.
- Model limitations: EPR, d-EPR, r-EPR, and recency-based EPR reproduce the conserved size of individual capacity but do not account for activity-set evolution.These outcomes correspond to Figure 5A and Figure 5C, respectively.
- EPR mechanism: In EPR, a new location is explored with probability Pnew = ρS −γ, while a previously visited location is revisited with probability 1 −Pnew.S is the number of previously visited locations, and ρ and γ are model parameters.
- EPR mechanism: The EPR model scales time with the number of transitions as ∼n/β, where β is a model parameter.The model defines prior visits using mi(n), the total number of visits to location i before transition n.
- Model proposal: The limited-memory EPR model preserves the EPR exploration strategy while giving agents a limited memory M.Return behavior is based on visits to locations occurring at most M days before a transition.
- Model comparison: A comparison uses β = 0.8, ρ = 0.6, γ = 0.2, maps 1u = 1min, and sets memory to M = 200 days.The EPR-with-memory model is compared with the EPR model using parameters chosen in reference 4.
1.5 Additional measures
Mobility patterns evolve through gradual spatial displacement and substantial turnover in important locations, while the median radius of gyration remains constant over time. Location additions and dismissals are balanced and scale with location capacity, indicating greater routine instability for individuals with larger capacities.
- Spatial properties: The activity set’s center of mass changes position over time, suggesting individuals gradually displace their important locations.The average distance between r_cm(t) and r_cm(t+γ) increases with γ.
- Spatial properties: The median radius of gyration is constant in time, with a linear fit yielding a = 11 ± 8 Km.
- Location turnover: For all datasets, locations added and dismissed constitute more than nL,i/3 of important locations, confirming substantial activity-set turnover.
- Location turnover: For most users, nA,i ∼ nD,i, so the number of newly adopted locations approximately equals the number of dismissed locations.
- Location turnover: For most users, nA,i ∼ b ∗ Ci, with b = 1.85 for Lifelog, 2.00 for CNS, 1.61 for MDC, and 1.25 for RM.
- Location turnover: Location additions and removals are proportional to capacity, meaning individuals with larger capacity have more unstable routines.
2 Supplementary Figures
Supplementary analyses characterize dataset coverage and preprocessing, validate returner/explorer classifications, and test the robustness and temporal evolution of activity-set and location-capacity measures. They also examine seasonality and sensitivity to alternative location definitions.
- Data coverage and preprocessing: The datasets are evaluated for collection duration, weekly time coverage, temporal resolution, and the Lifelog dataset’s broad spatial coverage.Additional preprocessing analyses compare raw Lifelog data with downsampled and user-selected data.
- Returners and explorers: 75% of time-windows is the consistency threshold used to classify CNS participants as returners or explorers, covering about 64% of participants.The classification uses 20-week windows and distinguishes categories by whether r3_g/r_g exceeds or falls below 0.5.
- Returners and explorers: 75% of time-windows is the consistency threshold used to classify Lifelog users as returners or explorers, covering about 56% of users.The classification uses 20-week windows and the same r3_g/r_g threshold of 0.5.
- Activity set and capacity: Activity-set establishment, gain, normalized capacity, individual capacity, and window-size dependence are examined across the Lifelog, CNS, MDC, and RM datasets.Normalized capacity is computed as C_i/TC_i to account for differences in individual data collection.
- Seasonality: Seasonality analyses track weekly and monthly numbers of unique locations visited, including holiday and examination periods in the CNS data.The supplementary figures show these patterns for CNS and compare monthly variation across all four datasets.
- Temporal evolution and location definitions: Activity-set evolution is tested for invariance under time translation, cosine similarity across time lags, and sensitivity to alternative definitions of locations.The analyses also compare location discovery and activity-set overlap using power-law fits.
3 Supplementary Tables
The supplementary tables characterize the datasets and weekly locations, then document robustness tests for conserved capacity, discrepancies from randomized mobility, time allocation, and location-set evolution.
- Dataset characteristics: Table 1 defines dataset characteristics including individuals, temporal and spatial resolution, collection duration, and median weekly time coverage.For Lifelog trajectories, temporal resolution records location changes in motion, and weekly time coverage uses stop-locations exceeding 10 minutes.
- Location counts: Table 2 reports median total and unique locations visited per week across four datasets, highlighting unique weekly locations as the study’s relevant quantity.The caption notes that GSM-derived displacement totals vary substantially, whereas unique weekly locations are comparable after accounting for time coverage.
- Conservation of capacity: Table 3 tests conservation of capacity across location thresholds and window sizes using linear-fit, power-law, and individual-level hypothesis results.It reports the linear coefficient b and p-value for H0, the power-law coefficient β and p-value for H1, and percentages for Hj,k.
- Robustness tests: Tables 4 and 5 assess conservation robustness by reporting stable-capacity percentages and real-versus-randomized distribution discrepancies across thresholds, windows, randomization schemes, and datasets.Table 4 considers individuals with |Gi| < σGi, while Table 5 reports KS statistics and p-values for local and global randomization.
- Time allocation and evolution: Tables 6 and 7 examine time allocation and location-set evolution across location classes, thresholds, and window sizes using hypothesis-test p-values and average Jaccard similarity.Table 6 tests H0: b = 0 for different ΔT classes; Table 7 reports similarity between activity subsets separated by w, using W = 20 weeks.