Source-linked AI summary
Cross-checking different sources of mobility information
Maxime Lenormand, Miguel Picornell, Oliva G. Cantu-Ros, Antonia Tugores, Thomas Louail, Ricardo Herranz, Marc Barthelemy, Enrique Frias-Martinez, Jose J. Ramasco
TL;DR
Most mobility studies rely on single data sources, leaving uncertainty about source-related bias. This paper compares Twitter, census, and cell-phone data in Barcelona and Madrid across population distributions, temporal patterns, and mobility networks, finding comparable information at the studied spatial and temporal scales.
Problem
Most urban-mobility studies use a single data source, leaving uncertainty about how results may be biased by the source used.
Method
The study compares spatial and temporal population-density distributions and Origin-Destination mobility matrices from Twitter, census, and cell-phone data in Barcelona and Madrid.
Results
The three datasets yield comparable population-density and mobility patterns, including correlations near 0.9 for density profiles and above 0.97 for Madrid’s mobility coefficients of determination.
Takeaways & Limitations
The results support using the three data sources interchangeably for mobility analysis at the studied spatial and temporal scales, while accounting for their different resolution limits.
Abstract
from arXiv · showhide
The pervasive use of new mobile devices has allowed a better characterization in space and time of human concentrations and mobility in general. Besides its theoretical interest, describing mobility is of great importance for a number of practical applications ranging from the forecast of disease spreading to the design of new spaces in urban environments. While classical data sources, such as surveys or census, have a limited level of geographical resolution (e.g., districts, municipalities, counties are typically used) or are restricted to generic workdays or weekends, the data coming from mobile devices can be precisely located both in time and space. Most previous works have used a single data source to study human mobility patterns. Here we perform instead a cross-check analysis by comparing results obtained with data collected from three different sources: Twitter, census and cell phones. The analysis is focused on the urban areas of Barcelona and Madrid, for which data of the three types is available. We assess the correlation between the datasets on different aspects: the spatial distribution of people concentration, the temporal evolution of people density and the mobility patterns of individuals. Our results show that the three data sources are providing comparable information. Even though the representativeness of Twitter geolocated data is lower than that of mobile phone and census data, the correlations between the population density profiles and mobility patterns detected by the three datasets are close to one in a grid with cells of 2x2 and 1x1 square kilometers. This level of correlation supports the feasibility of interchanging the three data sources at the spatio-temporal scales considered.
I. INTRODUCTION
The study addresses whether mobility and population-density findings depend on using a single data source. It compares Twitter, cell-phone, and census-derived information across Barcelona and Madrid at multiple spatial and temporal scales.
- Most prior mobility and urban studies relied primarily on a single data source, raising concerns about source-dependent bias.
- The study compares spatial and temporal population-density distributions and Origin-Destination mobility matrices from Twitter, cell phones, and census data.
- The analysis focuses on the metropolitan areas of Barcelona and Madrid, where all three data sources are available.
- The comparison is designed to assess whether similar results can be obtained despite differences in data-source nature and resolution.
- The metropolitan areas are divided into regular square grid cells to compare activity and intra-city mobility.
A. Mobile phone data
Mobile-phone activity is reconstructed from anonymized call records assigned to BTS Voronoi areas, then aggregated over grid cells and day groups. The data also support commuting OD matrices from users with reliably recoverable home and work locations.
- 55 days of anonymized call records from September to November 2009 provide the mobile-phone dataset.
- Each call is assigned to a BTS Voronoi area, allowing activity to be estimated spatially and hourly.
- Mobile-phone users in each grid cell are estimated by weighting Voronoi-cell counts according to intersection area.The weighting uses the area shared by each Voronoi cell and grid cell relative to the Voronoi-cell area.
- The available days are grouped into four day categories, and average users are computed for each group and hour.
- The daily curves show peaks between noon and 3pm and between 6pm and 9pm, with higher user counts on weekdays than weekends.
- Commuting OD matrices use users with identifiable home and work locations and calls at home or work on more than 40% of study days.
B. Twitter data
The Twitter dataset comprises geolocated users in Barcelona and Madrid and provides hourly, day-grouped grid-cell counts. Twitter-based commuting matrices use stricter activity filtering and direct coordinate-to-grid assignment.
- The dataset contains 27,707 Barcelona users and 50,272 Madrid users who produced geolocated tweets during September 2012–December 2013.
- Twitter users are counted in each grid cell by hour and day group, analogously to the mobile-phone activity measures.
- Twitter commuting matrices require at least 100 weekday tweets per user because geolocated tweets are less frequent than calls.
- Geolocated tweets are assigned directly to grid cells using latitude and longitude, without an intermediate Voronoi-cell step.
- The census survey samples one fifth of the population and defines municipal OD flows from household and workplace municipalities.
III. RESULTS
Twitter and mobile-phone estimates show strong agreement in spatial activity distributions across Barcelona and Madrid. This agreement remains substantial when the grid resolution is increased from 2 km to 1 km.
- The spatial concentration areas inferred from Twitter and mobile-phone data are visually similar in the Barcelona comparison.
- 0.93 for Barcelona and 0.89 for Madrid are the average Pearson correlations across day groups and hours at the coarser resolution.
- 0.85 for Barcelona and 0.83 for Madrid are obtained when the grid side is reduced to 1 km.
B. Temporal distribution
The study normalizes activity across sources and compares temporal distributions by grid cell, day group, and hour. Mobile-phone data yields three robust temporal patterns, while Twitter reproduces them despite weak independent clustering.
- Normalization and clustering: Mobile-phone activity forms three clusters with average silhouette indices of 0.38 in Barcelona and 0.43 in Madrid.The clusters correspond to business, residential, and nightlife land uses.
- Temporal patterns: Business, residential, and nightlife clusters differ by weekday/weekend and time-of-day activity patterns.Business activity is concentrated on weekday mornings and afternoons, residential activity is stronger on weekends, and nightlife activity is highest at night, especially on weekends.
- Twitter comparison: Twitter data cannot independently recover meaningful clusters, with silhouette indices below 0.1 for fewer than 10 clusters in both cities.The authors attribute these low values probably to noise in the Twitter data.
- Twitter comparison: Twitter reproduces the three mobile-phone temporal patterns in Barcelona and Madrid across different scale values.The comparison assigns Twitter users to the clusters obtained from mobile-phone data.
C. Users’ mobility
The paper compares normalized origin–destination flows and travel-length distributions from Twitter and cell-phone data in Barcelona and Madrid. The sources show strong agreement, with discrepancies concentrated among low-flow or missing links.
- OD-flow comparison: Normalized Twitter and cell-phone OD flows show an overall Pearson correlation of approximately 0.9.The comparison uses commuter-normalized flows for links present in both networks.
- OD-flow comparison: Dispersion is higher for low-flow links, where statistical fluctuations have a stronger influence.Missing links typically contain only one commuter in the network where they appear.
- Missing links: Missing links are substantially weaker than the strongest links, with maximum weights 25 to 464 times lower and average weights 4 to 9 times lower.This indicates that cross-source differences mainly involve weak links.
- Travel-length comparison: Travel-length distributions are strongly similar between Twitter and cell-phone networks in both cities.
D. Census, Twitter and cell phone
The census comparison changes Twitter and cell-phone data to the municipal scale. Twitter and census flows correlate strongly in both cities, with Madrid showing the stronger agreement.
- Census comparison: The comparison uses municipal-level census data, requiring Twitter and cell-phone data to be rescaled from grids to municipality boundaries.
- Census comparison: Barcelona’s Twitter–census comparison gives R2 values of 0.8 and 0.9 for the two directional fits.The best-fit slope is 0.85, so the two directional coefficients of determination differ.
- Census comparison: Madrid achieves a Pearson correlation of approximately 0.99 and coefficients of determination above 0.97.
IV. DISCUSSION
The study compares mobility information from cell phones, Twitter, and census across spatial, temporal, and commuting dimensions. The sources produce similar density patterns and commuting structures, although Twitter requires longer integration times and has limitations for shorter-term mobility.
- The three data sources differ in their nature and the spatial and temporal scales at which mobility information is recovered.
- Commuting-distance distributions from Twitter and mobile phone data show strong similarity in Barcelona and Madrid for trips between different grid cells.
- Cell phone and Twitter data produce similar spatial and temporal density patterns, with Pearson correlation close to 0.9 in both cities.
- Similar temporal distribution patterns can be extracted from cell phone and Twitter datasets for different urban activity areas.
- At 1 or 2 km grid resolution, cell phone and Twitter data yield comparable Origin-Destination commuting networks.
- Twitter requires longer integration times than cell phone data to obtain similar commuting results and can struggle with shorter-term mobility.
APPENDIX
The preprocessing identifies and removes anomalous days before averaging mobile-phone activity patterns. Outliers include special days and days with incomplete data.
- Two types of outlier days are removed: special days and days with missing data for several hours.For Barcelona, one example is 11 October 2009, when data were unavailable from 5PM to 11PM.
Voronoi cells
The appendix models mobile-phone coverage with Voronoi cells intersecting the metropolitan area. It distinguishes cell locations relative to the metropolitan boundary, surrounding territory, and sea before estimating users in the intersections.
- Zero-user BTSs are removed, and Voronoi cells are computed for the remaining BTSs in the metropolitan area.
- Four Voronoi-cell types are distinguished according to whether cells lie inside the metropolitan area or intersect its surroundings and the sea.
- The number of users assigned to a metropolitan-area intersection is computed while accounting for the different Voronoi-cell types.The calculation uses each cell’s users and area together with its intersection with the metropolitan area.
- Intersections between Voronoi cells and the sea are removed under the assumption that calling users located at sea are negligible.
Origin-Destination matrices
The study converts mobile-phone BTS-level commuting flows into grid-level Origin-Destination matrices and compares flow structures across datasets. Supporting figures describe the temporal data, spatial partitioning, clustering, and cross-dataset flow comparisons.
- Origin-Destination transformation: A transition matrix is required to transform the BTS Origin-Destination matrix into a grid-cell Origin-Destination matrix.
- Origin-Destination transformation: The transition matrix contains grid-cell/BTS intersection areas and is column-normalized to represent proportions of BTS areas.
- Supporting analyses: Figure S1 contrasts eight-day mobile-phone activity periods without outliers against periods containing two outlier days in Barcelona.
- Supporting analyses: Figure S2 maps Barcelona’s metropolitan area and shows both BTS Voronoi cells and their intersections with the metropolitan area.
- Supporting analyses: Figure S3 plots average silhouette values against cluster count for AHC and k-means.
- Supporting analyses: Figures S4–S6 compare mobile-phone and Twitter temporal patterns for business, residential/leisure, and nightlife clusters in Madrid and Barcelona.
- Supporting analyses: Figures S7–S8 compare normalized non-zero commuting flows between Twitter, mobile-phone, and census datasets at grid-cell or municipality scales.