Source-linked AI summary

Automated identification and characterization of parcels (AICP) with OpenStreetMap and Points of Interest

Ying Long, Xingjian Liu

arXiv:1311.6165v3cs.CYcs.DB

TL;DR

The paper addresses the paucity of urban parcel data in China and the labor demands of conventional identification methods. It proposes using OSM road networks for parcel geometries and POIs for parcel characteristics, producing a dataset that approximates conventionally identified parcels while retaining caveats from open and crowd-sourced data.

  • Problem

    Urban parcel data are scarce in China, while conventional identification and characterization methods are labor intensive and resource consuming.

  • Method

    The paper presents an extensible framework using OSM to identify parcel geometries and POIs to characterize parcel-level land-use intensity, function, and mixing.

  • Results

    OSM and POIs could produce reasonably good approximations of parcels identified from conventional methods.

  • Takeaways & Limitations

    The resulting fine-scale urban parcel dataset has potential as a useful supplement for parcel data in China.

Abstract

from arXiv · show

Against the paucity of urban parcels in China, this paper proposes a method to automatically identify and characterize parcels (AICP) with OpenStreetMap (OSM) and Points of Interest (POI) data. Parcels are the basic spatial units for fine-scale urban modeling, urban studies, as well as spatial planning. Conventional ways of identification and characterization of parcels rely on remote sensing and field surveys, which are labor intensive and resource-consuming. Poorly developed digital infrastructure, limited resources, and institutional barriers have all hampered the gathering and application of parcel data in developing countries. Against this backdrop, we employ OSM road networks to identify parcel geometries and POI data to infer parcel characteristics. A vector-based CA model is adopted to select urban parcels. The method is applied to the entire state of China and identifies 82,645 urban parcels in 297 cities. Notwithstanding all the caveats of open and/or crowd-sourced data, our approach could produce reasonably good approximation of parcels identified from conventional methods, thus having the potential to become a useful supplement.

Identification and characterization of parcels

The paper addresses limitations of conventional parcel identification by combining OSM road networks for parcel geometries with POIs for parcel characteristics in an automatic, extensible framework.

  • Manual parcel identification is resource intensive, time consuming, and dependent on individual practitioners’ experience and technical proficiency.Experienced operators require 3–5 hours to identify and infer land use for 35–50 urban parcels covering one square kilometer.
  • Conventional methods can produce inconsistent datasets and often omit parcel-level attributes such as density, land use, and land use mix.Available Beijing parcel-density data covered only approximately 13.8% of the Beijing Metropolitan Area.
  • The method could produce reasonably good approximations of parcels identified by conventional methods, while remaining subject to open and crowd-sourced data caveats.The authors position the approach as a preliminary supplement for fine-scale urban parcel data in China.
  • Existing automated approaches may omit road space, impose heavy computational burdens, or subdivide predefined blocks rather than identify blocks from data.Other approaches may also neglect parcel characteristics or focus on small areas and developed countries with high-accuracy OSM data.
  • OSM provides street-network data across many cities, while POIs offer sub-parcel business information useful for inferring land use and urban functions.POIs also provide broad availability and high spatial and temporal resolution.
  • The proposed process uses OSM to delineate parcel geometries and POIs to infer parcel-level land-use intensity, function, and mixing.The framework is described as fully automatic, extensible, applicable to large geographic areas, and suitable for routine updates and free distribution.

Data

The study assembles administrative, road-network, POI, remote-sensing, and manually generated parcel data to identify and validate urban parcels across China.

  • The analysis covers 654 Chinese cities across five administrative levels and restricts the study area to legally defined urban land within city propers.Three locations—Sansha, Beitun, and Taiwan—were excluded because of OSM and POI data availability.
  • OSM road networks downloaded on October 5, 2013 are used to produce urban parcels, with an ordnance survey dataset providing a road-network comparison.The OSM dataset contains 481,647 road segments and 825,382 kilometers, respectively 8.0% and 31.5% of the ordnance survey map.
  • OSM captures only part of the ordnance survey data but covers most urban areas in China, especially large cities.The authors use visual overlay inspection and preliminary completeness checks to assess OSM road coverage.
  • The POI dataset contains 5,281,382 geo-tagged points aggregated from twenty initial types into eight broader categories.Commercial sites account for most POIs, followed by office building or space, transportation facilities, and government buildings.
  • POI counts support land-use density and mix analyses, while the framework can substitute other human-activity measures such as survey, mobility, or check-in data.POIs labeled ‘others’ are used for density estimation but excluded from land-use mix analysis because they are poorly classified.
  • Validation compares OSM-based parcels with DMSP/OLS and GLOBCOVER remote-sensing products and manually generated Beijing parcels, while accounting for differing spatial resolutions.The Beijing comparison uses BICP parcel data and floor-space information where available.

Methods

The method delineates parcel geometries from OSM road networks, selects urban parcels with constrained vector-based CA, and characterizes them using POI-derived density, dominant function, and land-use mix. It validates the outputs at parcel and regional levels, with detailed parcel-level validation available for Beijing.

  • Delineating parcel boundaries: Parcels are defined as continuously built-up areas bounded by roads, with parcel polygons formed by removing buffered road space.OSM roads are merged, cleaned, extended, buffered, and overlaid with administrative boundaries to assign parcels to cities.
  • Calculating density for all parcels: Land-use density is calculated from POI counts relative to parcel area and standardized from 0 to 1 for comparison.The density normalization uses each parcel’s raw density and the nationwide maximum density value.
  • Selecting urban parcels: A constrained vector-based CA model selects urban parcels using neighborhood urbanization, parcel size, compactness, and POI density.Each city has its own model, and selection continues until the constrained total urban area is reached.
  • Selecting urban parcels: The model’s logistic-regression parameters are estimated from 125,401 manually prepared Beijing parcels, including 57,817 urban samples.Parcel status is regressed on size, compactness, and POI density before parameters are incorporated into models for cities across China.
  • Validating parcel selection: The logistic regression achieves 74.2% precision, while Beijing validation reaches 78.6% accuracy by parcel counts.These results support applying the constrained CA model to identify urban parcels from generated parcels.
  • Inferring dominant urban function and land use mix for selected urban parcels: Urban function is assigned from a dominant POI type when it accounts for more than 50% of parcel POIs, while a mix index supplements this classification.Some parcels lack a dominant urban function; the mix index measures the degree of mixed land use using POI-type proportions.

Results

The method generated and characterized urban parcels across China, with results showing substantial agreement with conventional and alternative data sources. Coverage and correspondence were stronger in better-mapped, larger cities, while sparse road networks and coarse validation data constrained accuracy.

  • 82,645 of 232,145 generated parcels were labeled urban across 297 cities, covering 25,905 km2.
  • Urban parcel characterization used land use density, urban function, and land use mix degree.Density could be compared across parcels and cities using inferred and standardized attributes.
  • 58,915 urban parcels (71.3%) had dominant urban functions, while the average land mix degree was approximately 0.66.Dominant functions included residential, commercial, office, and government parcels.
  • 71.2% of OSM-based urban parcels overlapped BICP urban parcels in Beijing, indicating similar geographic distributions of land use activities.OSM-based parcels were generally larger because tertiary and more detailed roads were missing from OSM.
  • 58.1% of OSM-based urban parcel area overlapped ORDNANCE urban parcels, with ratios around 70% in MC, SPC, and OPCC cities and around 45% in FLC and CLC cities.The results further indicated better OSM completeness in big cities than in small cities.

Conclusions

The paper presents an extensible framework and dataset for automatically identifying and characterizing fine-scale urban parcels across 297 Chinese cities using OSM and POI data. The authors report that the resulting parcels can support planning, urban modeling, and integration of other spatial data, while emphasizing limitations from open-data quality, sparse road networks, POI simplification, and limited validation.

  • Conclusions: The framework uses OSM and POI data to automatically delineate parcels, identify urban parcels, and characterize parcel features.It incorporates a vector-based cellular automata model into urban-parcel identification.
  • Conclusions: The framework could extend parcel-data production to areas lacking conventional data sources and support routine updates where official parcel maps are infrequently updated.For Beijing, the authors state that official parcel data are generally updated every three years, whereas their approach could enable yearly updating.
  • Conclusions: The resulting dataset contains fine-scale urban parcels with detailed features for 297 Chinese cities and is intended for urban planning and modeling.The authors describe applications in periodic updates, vector-based urban models, and parcel-level urban expansion modeling.
  • Conclusions: Generated parcels can serve as spatial units for consolidating geotagged photos, transportation smart-card records, taxi trajectories, and mobile-phone traces.The authors also suggest that integrating such data could improve estimates of urban function, density, and land-use mix.
  • Conclusions: The method remains constrained by open-data caveats, sparse OSM roads, equal treatment of POIs with different sizes, and fine-scale validation limited to Beijing.The paper calls for caution about data completeness, accuracy, vandalism, temporal inconsistency, flexible taxonomies, and broader validation with real-world parcel data.
Loading 1311.6165v3…