Source-linked AI summary
Human Mobility: Models and Applications
Hugo Barbosa-Filho, Marc Barthelemy, Gourab Ghoshal, Charlotte R. James, Maxime Lenormand, Thomas Louail, Ronaldo Menezes, José J. Ramasco, Filippo Simini, Marcello Tomasini
TL;DR
Human mobility research needs quantitative accounts of individual and collective movement that can support modeling and applications. This survey synthesizes data sources, metrics, models, and applications across mobility scales, while noting important data and scope constraints. It presents human-mobility modeling as relevant to migration, transportation, urban studies, and epidemic spreading.
Problem
Quantitative human-mobility research must characterize movement across individuals, populations, spatial scales, and temporal scales for applications including transportation and epidemic spreading.
Method
The survey organizes data sources, metrics, scales, generative and phenomenological models, and applications from individual to population mobility, including multimodal movement.
Results
The survey synthesizes empirical findings and modeling approaches spanning intra-urban movement, transportation systems, migration, social mobility, and epidemic spreading.
Takeaways & Limitations
Human-mobility models and data provide a framework for studying movement regularities and applications across society, while raising questions about behavioral regularity and cognition.
Takeaways & Limitations
The survey does not cover spatial networks or the science of cities in detail, addressing them only when needed to explain human-mobility concepts.
Abstract
from arXiv · showhide
Recent years have witnessed an explosion of extensive geolocated datasets related to human movement, enabling scientists to quantitatively study individual and collective mobility patterns, and to generate models that can capture and reproduce the spatiotemporal structures and regularities in human trajectories. The study of human mobility is especially important for applications such as estimating migratory flows, traffic forecasting, urban planning, and epidemic modeling. In this survey, we review the approaches developed to reproduce various mobility patterns, with the main focus on recent developments. This review can be used both as an introduction to the fundamental modeling principles of human mobility, and as a collection of technical methods applicable to specific mobility-related problems. The review organizes the subject by differentiating between individual and population mobility and also between short-range and long-range mobility. Throughout the text the description of the theory is intertwined with real-world applications.
1. Introduction
Human mobility research examines how people move through space and time, from migration and daily trips to individual trajectories and population flows. This review traces foundational distance-based theories, organizes modern models and data across mobility scales, and connects them to applications.
- Motivation and History: Human mobility encompasses movement by individuals and groups across space and time, including both long-range migration and shorter, recurring daily trips.Daily trips support work, social, and leisure activities, while migration operates across broader spatial and temporal scales.
- Motivation and History: Distance emerged as a central constraint in early mobility theories, linking interaction between places to population size, intervening distance, and migration flows.Carey, Ravenstein, Stouffer, and Zipf developed influential formulations that treated distance and opportunities as different mechanisms shaping movement.
- Motivation and History: Ravenstein’s migration laws emphasized short-distance movement toward commercial and industrial centers, with migration flows propagating from nearby and more remote districts.Later additions included the greater role of immigration in town growth and the association between migration volume, transportation, and industrial development.
- Motivation and History: Stouffer’s intervening-opportunities model makes mobility depend indirectly on distance through the opportunities encountered between origin and destination.The number of migrants traveling a distance decreases as the cumulative number of intervening opportunities increases.
- Scope and Limitations: The review addresses a need to unify rapidly expanding, interdisciplinary human-mobility research and provide newcomers with shared concepts, metrics, models, and tools.Its organization spans data sources, measures, mobility models, applications, and future challenges, while excluding spatial networks except when needed for mobility concepts.
2. Data Sources
Mobility research draws on diverse data sources, from censuses and tax records to surveys, mobile-phone records, GPS traces, and online platforms. These sources support aggregate and individual mobility analysis, but differ in scale, temporal and spatial precision, accessibility, and representativeness.
- Census Data: Census data estimate commuting and internal migration flows from workplace and residence information collected in national surveys.United States county-to-county files provide residence-based or workplace-based commuting flows, with supporting spatial and sociodemographic data for model validation.
- Tax Revenue Data: Tax-return records estimate aggregate migration by matching successive-year filings and comparing mailing addresses.The IRS data identify migrants when relevant address information changes between current and prior-year returns.
- Surveys: Local travel surveys provide detailed trip purposes, transport modes, and timing, but sacrifice scale through fewer respondents than censuses.The Chicago Travel Tracker Survey collected detailed travel inventories from 10,552 households over one or two days.
- Mobile Phone Records: Call-detail records infer individual movements from cell-tower routing, while their mobility information depends on recording frequency and tower coverage.Irregular call patterns can confound temporal analysis, and tower footprints range from tens of meters in dense urban areas to a few kilometers rurally.
- GPS Data: Vehicle GPS traces support high-precision trajectory, velocity, distance, and trip measurements for urban traffic studies, despite positional inaccuracies.Signals are transmitted approximately every 2 km and when engines start or stop; temporal precision, instantaneous velocity, and distance are reported as high quality.
- Online Data: Online location data reveal mobility and social-interaction patterns, but require substantial cleaning and validation because users may not represent the general population.Only approximately 3% of users in one cited study had geotagging enabled, and Twitter users showed demographic differences from UK census data.
3. Metrics, Physics and Scales
The review characterizes human mobility through distance, temporal, visitation, individual-space, and aggregate-flow metrics. These measures reveal heavy-tailed movement, regular returns, distinct mobility types, and reproducible origin–destination patterns.
- Jump Lengths: Jump length Δr is the Euclidean distance between locations recorded at times t and t + dt, summarized by the probability distribution P(Δr).The distribution measures how likely a population member is to travel a given distance in a short time step.
- Jump Lengths: P(Δr) for dollar bills follows a power law with exponent β = 0.59 ± 0.02, independent of the population size of the entry point.Mobile-phone displacements show similar power-law behavior, while GPS-based data are described by a truncated power law.
- Jump Lengths: GPS-based jump-length distributions are constrained by a 500 km study-area length and by the physical distance individuals drive in one trip.These restrictions help explain differences between GPS-based and bank-note or mobile-phone dispersal distributions.
- Radius of Gyration: Individual movement combines logarithmic growth and saturation in the radius of gyration, consistent with frequent returns to a few highly visited locations.The observed saturation reflects regular travel patterns and recurrent movement between preferred locations.
- Radius of Gyration: The radius of gyration cannot identify how strongly particular locations determine mobility because it depends jointly on inter-location distances and visitation time or frequency.Distant home and work locations can produce a large radius, as can visits to many distant locations even when frequently visited locations are close.
- Mobility Types: Mobile users separate into returners, whose movement is dominated by recurrent preferred locations, and explorers, who wander among varying locations.For k-returners, k-radius of gyration approximates the characteristic distance for k ≥2; explorers have much smaller k-radius values.
- Frequented Locations and Motifs: Visitation frequency follows a Zipf law, with the probability of finding a user at rank L approximately P(L) ∼1/L.This captures the unequal frequency with which users visit locations.
- Frequented Locations and Motifs: ∼90% of recorded trips can be described by 17 daily networks, suggesting that these motifs capture regularities useful for modeling and simulating mobility.A daily network qualifies as a motif when it occurs more than 0.5% in the datasets.
3.2. Physics of mobility
The review explains how transportation hierarchies, congestion, energy expenditure, and multimodality shape relationships among distance, travel time, speed, and mobility efficiency.
- Distance, travel time, and effective speed: Apparent speed increases with travel distance approximately as a power law with exponent β ≈0.5.The dependence reflects transportation hierarchy and waiting times that decrease proportionally with trip distance.
- Distance, travel time, and effective speed: Longer trips produce effective acceleration because travelers progressively use faster roads before descending to slower roads near their destinations.This hierarchical cascade is represented by ⟨v⟩= v0 + at.
- Travel time budget: Congestion-related travel delay increases with urban population according to a power law exponent of 1.270 ± 0.067.This nonlinear increase contrasts with the hypothesis of an approximately constant daily travel-time budget.
- Energy expenditure: Daily travel-time distributions across transport modes can be described using canonical-like energy distributions, with a cutoff suppressing very small energy expenditures.The framework treats travel times and energies as measurable physical variables, while measurement errors and multimodal trips remain issues.
- The importance of multimodality: Multimodal travel introduces waiting and walking costs: short trips often spend more time waiting than riding, while complex networks add inter-layer delays.Time-respecting paths capture these delays more realistically than ideal quickest paths that assume perfect synchronization.
- The importance of multimodality: A small exponent μ ≈0.3 ± 0.1 indicates limited efficiency gains from increasing synchronization frequency between transport modes.The observed trend links larger frequency with better synchronization, but the small exponent makes inefficiency reduction difficult.
4. General Mobility Models
General mobility models distinguish individual mobility patterns from population flows and must account for mobility processes spanning broad spatial and temporal scales.
- General Mobility Models: Individual and population mobility require distinct modeling frameworks because mobility spans hundreds of meters to thousands of kilometers and hours to years.The distinction reflects differences in characteristic spatial and temporal scales and uncertainty in individual mobility.
4.1. Individual-Level (Random walks)
Individual-level mobility models describe trajectories through random walks and extensions that capture anomalous diffusion, preferential revisits, recency, and social influence. These models connect mathematical movement rules to empirical mobility regularities.
- Random walks: Random walks represent positions as sums of statistically independent displacements drawn from a probability distribution.The resulting position probability density determines spatial and temporal measures such as mean square displacement and higher moments.
- Brownian motion: R(t) ∼t1/2 characterizes ordinary Brownian diffusion, whose mean square displacement scales linearly with time.Brownian motion arises from independent, normally distributed increments and has zero mean displacement with variance proportional to time.
- Continuous-time random walks: R(t) ∼tα/β in ambivalent continuous-time random walks, with heavy-tailed jump lengths and waiting times producing sub- or super-diffusion.Empirical GPS, CDR, and dollar-bill data show power-law behavior in both distributions, with measured ranges 0.42 ≤α ≤0.8 and 0.31 ≤β ≤0.75.
- Behavioral and social models: Measured ζ = 1.2 ± 0.1 implies long-time saturation of movement, consistent with motion dominated by an individual’s most visited location.The model predicts MSD →Xmax for ζ > 1, where Xmax denotes the saturation point.
- Behavioral and social models: Preferential-return models combine exploration with revisitation, while recency-based refinements and social pressure improve reproduction of observed visitation profiles.The basic preferential-return model misses the broader distribution for recently visited locations; social pressure better reproduces re-visitation profiles measured in CDR data.
4.2. Population-Level
Population-level models estimate flows between regions using origin-destination matrices and spatial-interaction formulations. Gravity and radiation families differ in whether distance-related costs or intervening opportunities determine flows, and validation against empirical data remains essential.
- Population flows: Origin-Destination matrices organize aggregate mobility as flows between origins and destinations and can be represented as directed weighted networks.Individual trajectories aggregate straightforwardly into flows, but disaggregating flows generally requires relationships with static attributes of locations.
- Population flows: Spatial-interaction models infer flows from location attributes and pairwise variables such as population, distance, or travel time.Their differences primarily concern which variables enter the model and the functional forms used.
- Gravity and intervening-opportunity models: Gravity models assume trips decrease with distance, whereas intervening-opportunity models use the number of potential destinations between locations.The gravity formulation commonly combines origin and destination masses with a decreasing distance-deterrence function.
- Gravity and intervening-opportunity models: Travel cost may depend on scale, trip purpose, transportation mode, travel time, or monetary cost rather than distance alone.For commuting flows, the distance exponent is highly correlated with spatial scale.
- Constrained models: Constrained gravity models use observed outgoing or incoming flows, while non-constrained models estimate flows from indirect socioeconomic variables.The choice depends on the information available and the modeling objective.
- Radiation models: Radiation models avoid calibration parameters but are not robust to changes in spatial scale; extended versions add a scaling parameter to improve performance.The radiation formulation can use population or destination inflows to approximate opportunities, while the extended model controls the effect of intervening opportunities.
- Radiation models: Travel-distance distributions in commuting data are satisfactorily reproduced by the model described for job-density-based flows.The formulation uses job density as an opportunity-related quantity.
- Validation: Calculated flows must be validated against empirical evidence before practical application, accounting for spatial and temporal coverage limits.Observed data may cover a smaller region or shorter time window than the target application.
PAR LON USA
The review compares mobility models on commuting data from England and Wales, France, Italy, Mexico, Spain, the USA, London, and Paris, then examines multilayer approaches for multimodal transport.
- PAR LON USA: Commuting data from six countries and the cities of London and Paris were used to compare gravity, intervening-opportunities, and radiation-model variants.The models were evaluated under multiple levels of constraints.
- PAR LON USA: The common part of links (CPL) measures the proportion of links shared by simulated and observed networks, ranging from zero to one.It is null with no common link and equals one when the networks are topologically equivalent.
- PAR LON USA: For the selected datasets, the gravity model moderately outperformed the other models on the CPC metric, although the result may vary across datasets.CPC equals one for perfect agreement and zero when there is no overlap between data and model.
- Intermodality: Complete mobility descriptions require multimodal transport networks and transitions between transport modes, motivating multilayer or multiplex representations.These frameworks represent separate interaction networks as layers and can encode interlayer movement.
- Intermodality: Transport multilayers can combine bus, subway, and rail networks, while also representing alternative network spaces and traveler accessibility.The London example separates public transport into bus, subway, and rail layers, with the bus layer most used.
- Intermodality: Time-weighted multilayer networks support optimal-path calculations across complete transport systems while accounting for intermodality.They assign link weights according to trip duration and apply optimal-path algorithms across layers.
5. Selected Domains of Application
The review applies human-mobility models across pedestrian crowds, urban movement, transportation, and epidemic spreading. It contrasts modeling approaches, their empirical applications, and constraints arising from data, scale, and behavioral complexity.
- Pedestrian movement: Social force models represent individual pedestrians and reproduce directional stripes, intermittent door flow, and evacuation patterns.They apply analogues of Newton’s laws to each pedestrian and have been used for building, skyscraper, ship, and aircraft evacuation analysis.
- Pedestrian movement: High-density pedestrian turbulence exceeds the original social force model’s capabilities, motivating modified force terms such as the centrifugal force model.The limitation is documented in analyses of the 2006 Jamarat Bridge disaster, where individual trajectories became turbulent.
- Pedestrian movement: Cellular automata provide scalable crowd simulations, including simulations of crowds of size 10^5 with running time scaling linearly with the number of individuals.Their computational scalability comes with a trade-off relative to more detailed pedestrian representations.
- Intra urban mobility: Urban mobility studies link travel patterns to city morphology, built-environment factors, neighborhood, and residential segregation.Reported applications include London subway smart-card data, Spanish urban hotspots, Boston travel distances, and mobility segregation in Boston and Los Angeles.
- Intra urban mobility: At the city scale, human displacement distributions are more properly fitted by an exponential than a power-law distribution in the reviewed Foursquare analysis.The comparison covered 34 cities worldwide and used P(∆r) ∝exp(−∆r/r0) as the proposed form.
- Transportation and epidemic applications: Mobility models support transportation and epidemic applications, while model choice balances realism, data requirements, and analytical tractability.Agent-based models require highly detailed inputs, whereas metapopulation models offer greater tractability; GLEaM predicted global H1N1 arrival and prevalence peaks after calibration.
6. Conclusions
The survey synthesizes the state of human mobility modeling and connects its models to applications and open questions about behavioral regularities, social structure, and emerging transportation technologies.
- 6. Conclusions: The survey reviews current human mobility research, covering data, metrics, individual and population models, and selected applications.It presents the field's state of the art and links modeling approaches to real-world problems.
- 6. Conclusions: A central open question is why people return to frequently or recently visited locations and whether this regularity reflects memory, innate traits, or learning.
- 6. Conclusions: Mobility behavior may divide individuals into explorers and returners, raising questions about whether this dichotomy is an inherent property of society.
- 6. Conclusions: The field's future includes understanding how autonomous vehicles could transform transportation, mobility habits, society, the economy, and the environment.
- 6. Conclusions: Human mobility modeling offers insights relevant to real-world problems across society, including crime, urban planning, energy consumption, and social issues.
A. Modeling Frameworks and Algorithms
Implementing increasingly refined mobility models requires specialized computational, spatial, and statistical tools, many of which are widely available in popular programming languages.
- A. Modeling Frameworks and Algorithms: Refined mobility models require specialized tools such as efficient graph libraries, GIS capabilities, and statistical routines.
- A. Modeling Frameworks and Algorithms: Widely available libraries allow researchers to focus effort and grant resources on mobility research rather than implementing basic computational capabilities.
- A. Modeling Frameworks and Algorithms: The required tooling reflects the increasing complexity of models used to represent human traveling behaviors.
A.1. Modeling Frameworks
Computational mobility models simulate human movement through numerical, particle-based, or agent-based approaches, with complex system behavior emerging from smaller processes and interacting agents.
- A.1. Modeling Frameworks: Mobility models combine smaller blocks such as random-number generators, optimizers, and individual-level dynamics to produce macrolevel dynamics.A continuous-time random walk combines waiting-time and jump-length distribution models.
- A.1. Modeling Frameworks: Mobility models commonly rely on Monte Carlo methods, generating repeated random samples to approximate human displacements.
- A.1. Modeling Frameworks: Computational models simulate human mobility to validate analytical models by comparing generated data with real-world traces.
- A.1. Modeling Frameworks: Agent-based models represent interacting entities whose simple rules and microlevel interactions generate complex macroscopic regularities.
- A.1. Modeling Frameworks: Agents may operate autonomously, interact socially, react to their environment, and initiate actions to fulfill purposes.
- A.1. Modeling Frameworks: Agent-based simulations can span microlevel behaviors to macrolevel phenomena and support flexible agents with varied behaviors, rationality, and learning skills.
A.2. Algorithms
The algorithms section provides pseudocode for implementing selected mobility models and for sampling random numbers from specified probability distributions within an agent-based framework.
- A.2. Algorithms: The section gives pseudocode for algorithms implementing selected mobility models.
- A.2. Algorithms: It also provides pseudocode for generating random numbers sampled from specific probability distributions.
- A.2. Algorithms: The general framework for individual mobility is agent-based, as introduced in Section A.1.
A.2.1. Individual Mobility
Individual mobility models range from simple random walks and Lévy flights to models incorporating time variation, exploration, preferential return, recency, and social influence. These approaches provide simulation procedures while highlighting distributional and implementation limitations.
- Random Walks: Random walks simulate movement through successive random displacements, commonly by moving an agent one unit among four directions on a two-dimensional lattice.The process iterates until N steps have been completed.
- Lévy Flights: Lévy flights replace fixed or lattice-based jumps with heavy-tailed lengths, producing mostly short moves punctuated by occasional long jumps.In two dimensions, directions are selected uniformly and jump lengths are drawn from a power law.
- Implementation Considerations: Power-law implementations require parameter choices and may involve approximation or discretization errors, especially for small x.The continuous form is straightforward to invert, whereas the discrete form lacks a closed-form inverse.
- Continuous-Time Random Walks: Continuous-time random walks extend random-walk models by allowing variable waiting times between jumps, which can be drawn from power-law distributions useful for human mobility.The model separates jump generation from the waiting-time process.
- Human Mobility Models: The Exploration-Preferential Return model adds human-specific mechanisms missing from Lévy flight and CTRW models and uses power laws with exponential cutoffs for jump lengths and waiting times.Its construction includes exploration of new locations and preferential returns to previously visited locations.
- Human Mobility Models: The recency model modifies return decisions so that recently visited locations can be selected even when they have been visited only once or a few times.This produces a more even visitation-frequency distribution and spatial clusters farther from the starting location than the Individual Mobility model.
- Human Mobility Models: Social mobility models add a social step in which destination choice can follow social contacts or individual visitation preferences.The social-contact choice is controlled by probability α, while individual preference is used with probability 1 − α.
A.2.2. Population Mobility
Population mobility is represented through origin-destination matrices that count flows between spatial zones and can vary over time. Agent simulations normalize these matrices into movement probabilities and apply them to move agents between locations.
- Origin-Destination Matrices: An origin-destination matrix records the number of individuals moving from location i to location j across a set of spatial zones.Its entries T_ij form a directed, weighted network that is generally time-dependent.
- Origin-Destination Matrices: Origin-destination matrices differ from segment measurements because they capture complete movements between origins and destinations rather than counts through individual transport links.They are also described as extremely difficult and costly to obtain and measure.
- Agent Simulation: To simulate population flows, matrix counts T_ij are converted into relative movement probabilities P_ij for each time instant.The time-dependent matrix is normalized before the simulation iterates through origin-destination pairs.
- Agent Simulation: At each time step, an agent is moved from i to j when a random draw falls below the normalized probability P_ij.The procedure repeats across matrix elements until the simulated duration τ is reached.