Source-linked AI summary
A New Insight into Land Use Classification Based on Aggregated Mobile Phone Data
Tao Pei, Stanislav Sobolevsky, Carlo Ratti, Shih-Lung Shaw, Chenghu Zhou
TL;DR
Urban land-use classification has largely emphasized physical characteristics rather than social functions. This paper uses aggregated mobile-phone activity patterns and call volumes with semi-supervised fuzzy c-means clustering, achieving a 58.03% detection rate in Singapore.
Problem
Urban land-use classification has limited use of information about social functions alongside physical characteristics.
Method
The paper combines hourly mobile-phone activity patterns and aggregated call volume through a weighted time series for semi-supervised fuzzy c-means land-use classification.
Results
58.03% detection rate was achieved for land-use classification in Singapore with β optimized to 0.75.
Takeaways & Limitations
Aggregated mobile-phone data can provide land-use information from the perspective of social function in Singapore.
Takeaways & Limitations
Land-use classification precision depends on cell-tower density, with low BTS density creating a risk of reduced information precision.
Abstract
from arXiv · showhide
Land use classification is essential for urban planning. Urban land use types can be differentiated either by their physical characteristics (such as reflectivity and texture) or social functions. Remote sensing techniques have been recognized as a vital method for urban land use classification because of their ability to capture the physical characteristics of land use. Although significant progress has been achieved in remote sensing methods designed for urban land use classification, most techniques focus on physical characteristics, whereas knowledge of social functions is not adequately used. Owing to the wide usage of mobile phones, the activities of residents, which can be retrieved from the mobile phone data, can be determined in order to indicate the social function of land use. This could bring about the opportunity to derive land use information from mobile phone data. To verify the application of this new data source to urban land use classification, we first construct a time series of aggregated mobile phone data to characterize land use types. This time series is composed of two aspects: the hourly relative pattern, and the total call volume. A semi-supervised fuzzy c-means clustering approach is then applied to infer the land use types. The method is validated using mobile phone data collected in Singapore. Land use is determined with a detection rate of 58.03%. An analysis of the land use classification results shows that the accuracy decreases as the heterogeneity of land use increases, and increases as the density of cell phone towers increases.
1. Introduction 65
Urban land-use classification often emphasizes physical characteristics, leaving social functions underused and making some heterogeneous classes difficult to distinguish. The paper proposes mobile phone data as a social-function-based information source and evaluates its applicability to urban land-use classification.
- Problem: Remote sensing methods based on spectral and textural information struggle to distinguish some heterogeneous land-use types, such as residential and commercial areas.Additional contextual, parcel, field, and expert information can improve inference but increases cost and delays updates.
- Problem: Remote sensing has progressed substantially, but classification still tends to emphasize physical characteristics while inadequately using social functions.This motivates incorporating information that reflects how people use different areas.
- Motivation: Mobile phone data can capture residents’ daily activities and indicate land-use social functions because activity routines differ across residential and business areas.For example, residential areas commonly show morning departures and evening returns, whereas business areas may exhibit the opposite pattern.
- Contribution: The paper evaluates mobile phone data as a potential new information source for urban land-use classification from the perspective of social function.Its stated objective is to verify the data source’s applicability and evaluate the results it produces.
2. Related work 107
Prior studies used aggregated mobile phone data to characterize residents’ activities and identify specific land use types, but did not fully address urban land use classification. This study targets limitations in representativeness, interpretability, and uncertainty by combining richer time-series information with a transparent semi-supervised approach.
- Existing applications: Prior research used aggregated mobile phone data to describe urban activity, estimate populations, identify social groups, detect events, and distinguish selected land use types.Applications included urban landscape description, population estimation, social-group identification, event detection, and differentiation of residential, business, and park areas.
- Research gap: Existing studies focused on identifying specific land use types rather than classifying urban land use comprehensively.This limitation motivated methods designed specifically for urban land use classification.
- Limitations of prior methods: Normalized calling patterns neglect total-volume differences between BTSs and may fail to distinguish heterogeneous areas containing commercial, residential, and recreational activities.Pattern-only methods capture temporal variation within a BTS but omit cross-BTS volume differences, limiting discrimination among mixed land uses.
- Limitations of prior methods: Toole et al.’s method combined normalized calling patterns and volume with supervised random forests, but remained difficult to interpret and used only average weekday and weekend patterns.The two-day representation neglected differences between weekdays and between weekends, despite activity differences across those days.
- Study contributions: The study constructs a linear combination of four-day call patterns and volume, proposes a semi-supervised inference scheme, and analyzes classification uncertainties.The resulting time series incorporates more mobile-phone characteristics, improves interpretability, and supports analysis of pattern and volume effects, BTS density, land-use heterogeneity, and fuzzy membership.
3. Semi-supervised fuzzy c-means (FCM) clustering method for urban land use 205
The method classifies urban land use by combining hourly mobile-phone activity patterns with total calling volume in a synthesized time series, then applying semi-supervised FCM clustering. Expert-labeled samples determine the pattern–volume weighting, while cluster numbers and land-use assignments are selected through validation and proximity.
- Method workflow: The five-step workflow meshes aggregated BTS data, synthesizes pattern and volume, estimates β from training samples, clusters with FCM, and assigns land-use labels.Training samples are selected using expert knowledge, and clustering results are post-processed into land-use types.
- Data preparation: Hourly BTS-level data are interpolated onto a mesh grid using Voronoi polygons, area-normalized volume density, and inverse distance weighting.The hourly values generated over each grid cell form the input time series.
- Synthesized time series: The synthesized series combines each cell’s hourly pattern X_i with range-transformed total volume Y_i through coefficient β, enabling comparison of their classification roles.The range transformation gives Y_i the same range as X_i.
- Weight estimation: β is optimized using known land-use samples from external information sources, minimizing misclassification against their true land-use types.Sample-group centers are calculated by averaging, and candidate β values are evaluated with an indicator-based objective function.
- Clustering and post-processing: FCM cluster numbers are chosen using a validation index, and each resulting cluster is assigned the land-use type with the closest sample-derived center.This preserves potentially specific cluster structures when mapped land-use categories combine multiple underlying types.
4. Aggregated mobile phone data from Singapore 291
The study uses a week of hourly aggregated call counts from more than 5,500 Singapore BTS towers to construct a 96-point land-use time series. It represents activity through a four-day temporal mode and validates clustering against a five-class urban planning map.
- Data: Hourly aggregated call counts from 5,500+ BTS towers were collected for one week in Singapore.The week ran from Monday 28 March to Sunday 3 April 2011.
- Feature construction: The input combines a normalized activity pattern with total call volume.The two components are linearly combined from the seven-day mobile phone timelines.
- Temporal representation: A four-day mode averages Monday–Thursday as an ordinary weekday while retaining Friday, Saturday, and Sunday separately.The weekday grouping reflects similar normalized patterns, whereas the remaining three days show significant differences.
- Temporal representation: The four-day mode forms a 96-point time series and produces the best classification result compared with two-day and seven-day modes.The comparison of detection rates confirms the superiority of the four-day processing choice.
- Validation and preprocessing: For validation, Singapore is divided into Residential, Business, Commercial, Open space, and Others using an urban planning map.Aggregated hourly data are interpolated to a 200 m × 200 m grid with IDW, generating 96 pattern layers and one volume layer.
5. Land use classification for Singapore 332
Aggregated mobile-phone time series combining call patterns and volume distinguish Singapore land-use types through fuzzy c-means clustering. The optimized weighting and resulting classification achieve an overall detection rate of 58.03%, with performance varying by land-use type and the relative importance of pattern versus volume.
- Time-series characterization: The time series combines a 96-point call-pattern component with call volume, and land-use types are characterized by both features.Residential areas show similar four-day patterns and medium volume, whereas Business areas show weekday high-thin patterns, low weekend patterns, and low volume.
- Classification performance: 58.03% is the overall detection rate for identifying all land-use types from the synthetic time series.This is close to the 54% detection rate reported for Toole et al. (2012).
- Classification performance: Open space, Residential, Business, Commercial, and Others are ordered from best to worst detection, with only the first three near or above 50%.Commercial and Others have detection rates below 50%, and some land-use types have misclassification rates above 30%.
- Pattern-volume contribution: Pattern information generally contributes more to classification than weighted volume, except for Commercial areas.The average pattern-to-weighted-volume distance ratio is 1.6471, while Commercial’s higher volume makes volume more important for separating it.
6. Comparison between classifications using different information
Using pattern or volume alone produced weaker land-use classifications than combining both information types. Pattern missed Commercial areas because of residential mixing, while volume missed Business regions because Business and Open space had similar volumes.
- Comparison of classification inputs: Pattern-only clustering generated five clusters but did not identify Commercial areas, whereas volume-only clustering generated four clusters and did not identify Business regions.The combined pattern-and-volume classification performed better overall.
- Comparison of classification inputs: 52.58% for pattern and 52.68% for volume were lower overall detection rates than the combination of pattern and volume.
- Pattern information: Only 9.89% of Commercial areas were correctly classified from pattern information, while 40.54% were mixed into Residential.This mixing explains why pattern information alone failed to identify Commercial land use.
- Volume information: Volume failed to detect Business land use because Business and Open space had very similar median values and ranges.Their similar volume distributions prevented separation using volume alone.
- Volume information: Volume-only classification identified only four land-use types because Business and Open space could not be separated by volume alone.
7. Discussion 500
The discussion attributes classification errors to mismatches between planned land use and observed social function, land-use heterogeneity, BTS information precision, and the fuzzy membership threshold. Error rates increase with land-use entropy and generally improve with BTS density, except for density 0, where detection reaches 60.56%.
- Sources of error: Four factors affect classification error: land-use definition versus mobile-derived function, land-use heterogeneity, BTS information precision, and the fuzzy membership threshold.The discussion identifies these as possible causes of classification errors.
- Land-use heterogeneity: Mixed social activities and heterogeneous areas can produce different classifications within one planned land-use type.Residential areas may mix commercial activity, while airport terminals are classified as Commercial and landing areas as Open space because their call volumes differ.
- BTS density: Low BTS density risks mixing calls from different land uses, whereas high density reduces interference and produces purer signals.The study calculates BTS volume from calls within Voronoi polygons, whose areas are larger when BTS density is low.
- BTS density: 60.56% detection occurs at BTS density 0, while detection generally increases with BTS density except at density 0.The high density-0 result is attributed to cells that are mostly Open space, where signals are purer.
- Land-use entropy: Error rate increases with land-use entropy because higher entropy indicates that more land-use types coexist within a cell.Average entropy is 0.42 for residential, 0.18 for business, 0.47 for commercial, 0.084 for open space, and 0.57 for others.
8. Conclusions and future work 582
The study used a synthesized mobile-phone activity time series and semi-supervised clustering to classify land use in Singapore, achieving a 58.03% detection rate with optimized β=0.75. Results show that combining activity pattern and volume is advantageous, while accuracy varies with land-use heterogeneity and BTS density; future work targets model improvement and additional data integration.
- Method: The study synthesized mobile-phone activity as a linear combination of a four-day pattern and aggregated-data volume, weighted by β.A semi-supervised clustering method used this synthesized time series to identify land-use types.
- Results: 58.03% detection rate was achieved for Singapore land-use classification with β set to its optimized value of 0.75.The optimized value was determined through a training process.
- Results: Combining pattern and volume produced better classifications than using either component alone, while the four-day mode outperformed the two-day and seven-day modes.The analysis also found that pattern was more important than volume for detecting most land-use types.
- Limitations and influencing factors: Mixed land use increases classification errors because it produces heterogeneous mobile-phone usage, whereas higher BTS density generally improves precision except where density is 0.Classification may perform well in areas with high BTS density and pure land-use types.
- Limitations and future work: Mobile-phone data reveal land-use social functions, but an overall detection rate below 60% makes them insufficient alone for urban land-use classification.Higher detection rates occur in areas with high BTS density, pure land use, and high fuzzy membership values.
- Future work: Future work proposes varying β spatially to capture different land-use characteristics and integrating remote-sensing data and POI into classification.These directions aim to improve the classification model and incorporate more information.