Source-linked AI summary

Tourists' digital footprint in cities: comparing big data sources

Maria Henar Salas-Olmedo, Juan Carlos Garcia-Palomares, Javier Gutierrez

arXiv:1705.07951v1cs.CY

TL;DR

Urban tourists’ spatial behaviour is poorly documented by traditional sources, which lack detailed information on where visitors go. This paper compares three geolocated digital sources in Madrid and finds that combining them reveals complementary tourist activities and spatial patterns.

  • Problem

    Traditional tourism data lack detailed information about the places tourists visit in cities, while their digital footprints capture diverse activities.

  • Method

    The study compares tourist densities from Panoramio, Foursquare and Twitter in Madrid using spatial analysis, clustering and integrated activity-based tourist counts.

  • Results

    The three sources show complementary and partly redundant spatial patterns, with high tourist density in Madrid’s historic centre and activity-specific concentrations elsewhere.

  • Takeaways & Limitations

    Understanding urban tourist distribution requires multiple complementary data sources, supporting more informed public-policy and business opportunity identification.

  • Takeaways & Limitations

    Digital-footprint sources are biased because many tourists do not use or publicly share geolocated photography, Twitter or Foursquare data.

Abstract

from arXiv · show

There is little knowledge available on the spatial behaviour of urban tourists, and yet tourists generate an enormous quantity of data (Big Data) when they visit cities. These data sources can be used to track their presence through their activities. The aim of this paper is to analyse the digital footprint of urban tourists through Big Data. Unlike other papers that use a single data source, this article examines three sources of data to reflect different tourism activities in cities: Panoramio (sightseeing), Foursquare (consumption), and Twitter (being connected). Tourist density in the three data sources is compared via maps, correlation analysis (OLS) and spatial self-correlation analysis (Global Moran's I statistic and LISA). Finally the data are integrated using cluster analysis and combining the spatial clusters identified in the LISA analysis in the different data sources. The results show that the data from the three activities are partly spatially redundant and partly complementary, and allow the characterisation of multifunction tourist spaces (with several activities) and spaces specialising in one or various activities (for example, sightseeing and consumption). The case study analysed (Madrid) reveals a significant presence of tourists in the city centre, and increasing specialisation from the centre outwards towards the periphery. The main conclusion of the paper is that it is not sufficient to use one data source to analyse the presence of tourists in cities; several must be used in a complementary manner.

1. INTRODUCTION

The introduction frames Big Data as a valuable complement to traditional tourism statistics because tourists leave digital footprints across diverse urban activities. The paper therefore compares Panoramio, Foursquare, and Twitter to identify tourist presence in Madrid using tourists—not individual digital records—as the unit of analysis.

  • Motivation: Big Data complements surveys, hotel records, and museum admissions by providing abundant information about tourists’ locations and behaviour in cities.Traditional official sources do not offer detailed information on the places tourists visit.
  • Motivation: Tourists’ digital footprints span sightseeing, consumption, and Internet connectivity, producing information that can reveal different types of urban spaces.Photo-sharing data capture sightseeing, while establishments and social-network activity capture consumption and connectivity.
  • Research aim: The paper compares three geolocated data sources in Madrid: Panoramio for sightseeing, Foursquare for consumption, and Twitter for Internet activity.The study uses these platforms to identify tourist presence according to different activities.
  • Contribution: The study contributes by combining three data sources and treating tourists as the unit of analysis rather than analysing photos, check-ins, and tweets directly.The data are processed to count tourists in each city location according to each source.
  • Paper structure: The paper proceeds from literature review through data and methodology to results, conclusions, and further research directions.Sections 2–6 cover these stages in sequence.

2. RELATED LITERATURE Sightseeing: photo-sharing services · Consumption: Foursquare check-ins

Urban tourists leave digital traces through geolocated photo-sharing and consumption-related activities. Prior research uses these sources to identify tourist presence, movements, urban functions, and tourism-related patterns, although Foursquare has been less studied in tourism.

  • 2. RELATED LITERATURE Sightseeing: photo-sharing services: Sightseeing generates geolocated digital footprints through photo-sharing services including Instagram, Flickr, and Panoramio.Panoramio focused on georeferenced images of places or landscapes shared by users.
  • 2. RELATED LITERATURE Sightseeing: photo-sharing services: Panoramio was particularly suited to sightseeing analysis because it centered on georeferenced photographs of places and landscapes.Its images were viewable through the Panoramio website, Google Earth, and Google Maps.
  • 2. RELATED LITERATURE Sightseeing: photo-sharing services: Photo-sharing data have supported tourism research on social events, tourist numbers, tourist presence, and common tourist trajectories.These applications demonstrate the use of digital photographs to study both tourist volume and spatial movement.
  • 2. RELATED LITERATURE Sightseeing: photo-sharing services: Photo-sharing services have also been used to propose or assess tourist routes.This extends their role from observation toward evaluating tourism itineraries.
  • Consumption: Foursquare check-ins: Consumption-related activities such as shopping and restaurants generate digital trails through activities and bank-card payments.Bank-card transactions have been little used in tourism studies because accessing these databases is difficult.
  • Consumption: Foursquare check-ins: Foursquare provides an alternative source for analysing consumption activities when bank-card transaction data are difficult to access.The supplied passage introduces Foursquare as an alternative, but its remaining discussion is truncated.
  • Consumption: Foursquare check-ins: Foursquare research has examined venue distributions, functional urban areas, movement patterns, area popularity, traffic conditions, and trade areas.Despite these broad urban applications, few studies have used Foursquare data specifically in tourism.
  • Consumption: Foursquare check-ins: Tourism applications of Foursquare remain limited, with Ferreira et al. identified as one exception.The supplied passage does not provide the details of Ferreira et al.'s analysis because the text is truncated.

Being connected: Twitter · 3. DESCRIBING AND PRE-PROCESSING THE DATA · Panoramio

Twitter provides globally covered, freely available real-time geolocated traces for studying places visited and urban mobility, but tourism applications remain scarce within cities. The Panoramio dataset contains 307,062 Madrid photographs, with users classified as tourists or residents using activity duration.

  • Being connected: Twitter: Twitter is widely used because it offers global coverage and free, real-time access to geolocated tweets.Each geolocated tweet records the place and time of sending.
  • Being connected: Twitter: Processing tweets by user identifier approximates the places visited by each user.
  • Being connected: Twitter: Social-network activity can proxy changing city population densities, mobility patterns, and dominant activity types.Twitter daily-use profiles distinguish business, leisure/weekend, nightlife, and residential activity.
  • Being connected: Twitter: Tourism studies using Twitter are scarce and generally compare visitor spatial behaviour between cities rather than analysing patterns within cities.
  • 3. DESCRIBING AND PRE-PROCESSING THE DATA: The Madrid Panoramio dataset comprises 307,062 geolocated photographs uploaded between 2006 and 2014.Records include coordinates, photograph-owner ID, URL, and upload date.
  • Panoramio: Panoramio users were classified as tourists or residents using user IDs and photograph dates.Users photographing in Madrid for more than one week per year were attributed to residence.

Twitter and Foursquare

The section contrasts geolocated tweets and Foursquare check-ins as digital traces of tourists in Madrid. Their temporal patterns differ, with Foursquare activity distributed across the day and Twitter activity concentrated in the evening.

  • Twitter: The Twitter dataset comprises geolocated tweets sent from Madrid between 2012 and 2014, including coordinates, user IDs, languages, timestamps, and message content.Tourists were identified as users who tweeted for a week or less per year.
  • Foursquare: Foursquare provides check-ins at restaurants, nightlife spots, shops, and other places of interest, with 50 million community users.Check-ins are private by default but can become publicly accessible when users share them via Twitter.
  • Temporal distribution: Foursquare check-ins are more evenly distributed throughout the day than tourist tweets.The check-ins cover restaurants, shops, and other places, whereas tweets are temporally more concentrated.
  • Temporal distribution: 18–21 hours is the main concentration period for tourist tweets, suggesting Twitter use is predominantly related to accommodation locations.This contrasts with the more evenly distributed temporal pattern of Foursquare activity.

4. METHODOLOGY

The methodology quantified and standardized tourists’ spatial presence across census tracts using Panoramio, Foursquare, and Twitter data. It then compared, mapped, spatially analyzed, and integrated the three datasets to characterize tourist activity patterns.

  • Data preparation: Tourists’ photos, Foursquare check-ins, and tweets were assigned to census tracts and aggregated by user ID to count single tourists per source.This produced the number of single users in each census tract for each data source.
  • Data preparation: Tourist densities were calculated per census tract to account for differences in tract size and reduce the Modifiable Areal Unit Problem.Density addressed the tendency of larger tracts to register more tourists.
  • Data preparation: Visitor-density data were linearly rescaled to 0 to 1,000 to remove differing source ranges and enable comparison.The rescaling was applied to density data for each data source.
  • Comparative analysis: Rescaled tract data supported density maps, descriptive statistics, bivariate OLS comparisons, and standardized-residual maps across data-source pairs.The coefficient of determination measured shared variation between each pair of sources.
  • Integration: K-means clustering integrated the three sources to characterize census tracts by tourist activities, maximizing within-group similarity and between-group difference.The methodology also used the resulting spatial information to distinguish activity patterns across tracts.
  • Spatial analysis: Global Moran’s I and Local Moran’s I (LISA) were calculated separately for each source using inverse-distance weighting with a 500 m radius.Spatial autocorrelation related each location to neighboring locations rather than treating locations independently.

5. RESULTS · Tourist density maps and descriptive statistics

Tourist density patterns vary by activity: sightseeing and consumption concentrate strongly in central Madrid, whereas Twitter activity is more dispersed. Twitter detects more tourists overall, while Foursquare shows the greatest spatial concentration.

  • Tourist density maps and descriptive statistics: Panoramio reveals high visitor concentration in Madrid’s historic centre and along the north–south Paseo de La Castellana axis.The densest areas correspond to sightseeing locations including Plaza de Cibeles, Puerta del Sol, Plaza Mayor, the Royal Palace, and others.
  • Tourist density maps and descriptive statistics: The three maps use rescaled data and identical intervals, enabling comparison of sightseeing, consumption, and connectedness across census tracts.
  • Tourist density maps and descriptive statistics: Foursquare density is particularly concentrated in the historic centre, with additional concentrations around the Golden Mile, AZCA shopping centre, and Real Madrid Stadium.
  • Tourist density maps and descriptive statistics: Twitter density is more dispersed, remaining high in the historic centre and along Paseo de la Castellana while spreading across many census tracts.
  • Tourist density maps and descriptive statistics: Twitter detects substantially more tourists than Panoramio or Foursquare, which justifies rescaling the three variables.
  • Tourist density maps and descriptive statistics: Foursquare has the highest coefficient of variation, confirming that consumption activities exhibit the greatest spatial concentration.

Comparison between data sources: OLS analysis

OLS analysis found medium positive association between Twitter and Foursquare tourist densities, while Panoramio had low correlations with both, indicating greater complementarity. Residuals showed that each source over- or underrepresented tourists in distinct urban areas.

  • Method: The analysis calculated adjusted r^2 between each pair of data sources and used standardised regression residuals to locate their greatest spatial differences.The residuals were mapped to identify where distributions diverged most.
  • Results: Twitter-Foursquare tourist densities showed a medium positive correlation, whereas Panoramio-Twitter and Panoramio-Foursquare correlations were low, indicating greater complementarity.The results suggest that Twitter and Foursquare distributions overlap more than either does with Panoramio.
  • Results: Panoramio exceeded expected tourist counts relative to Foursquare and Twitter in the historic centre and major sightseeing spots, but fell below expectations around the city centre.Higher-than-expected locations included football stadiums, the bullring, Retiro Park, Torres Kio, Cuatro Torres, and the Temple of Debod.
  • Results: Foursquare exceeded expected tourist counts relative to Twitter in central, shopping, business, transport, and stadium locations, but fell below expectations in less central areas.Higher-than-expected locations included the historic centre, Salamanca, AZCA, La Vaguada, Real Madrid Stadium, and Atocha station.

Types of spaces according to tourist activities: cluster analysis

K-means clustering of rescaled activity densities classified census tracts into six tourist-space groups. The groups distinguish sightseeing-, consumption-, Internet-, and low-tourist areas with different spatial distributions and activity intensities.

  • Clustering method: K-means clustering using rescaled densities classified census tracts by tourist presence in sightseeing, consumption, and Internet activities into six groups.Table 3 reports group means and standard deviations, while Figures 5 and 6 summarise group characteristics and spatial distribution.
  • Sightseeing spaces: Groups 6 and 2 predominated in sightseeing-related tourism, with Group 6 showing very high Panoramio tourist counts and high Twitter and Foursquare counts.Group 6 included the bullring, Puerta de Alcalá, Glorieta de Atocha, and Plaza de España; Group 2 represented lower-intensity sightseeing spaces.
  • Consumption spaces: Groups 5 and 3 predominated in consumption-related tourism, including commercial areas of the historic centre and retail districts such as Golden Mile-Salamanca and AZCA.Group 5 included Gran Vía and Puerta del Sol, with very high Twitter and Panoramio tourist counts.
  • Internet-activity spaces: Group 1 represented Internet-activity spaces with fewer tourists and occupied the limits of the historic centre.The passage identifies this group as blue in the spatial classification.
  • Low-tourist spaces: Group 4 represented low-tourist, generally nontourist spaces corresponding mainly to peripheral census tracts.The passage identifies this group as yellow in the spatial classification.

Analysis of spatial autocorrelation

Using inverse-distance weighting with a 500 m threshold, the analysis found positive spatial autocorrelation across all three data sources. Cross-referencing high-high clusters identified multifunctional tourist areas combining sightseeing, consumption, and internet connectivity.

  • Analysis of spatial autocorrelation: Using IDW with a 500 m distance threshold, Global Moran’s I was positive for Panoramio, Foursquare, and Twitter.The positive values indicate spatial autocorrelation in all three datasets, with lower values for Panoramio than for the other sources.
  • Analysis of spatial autocorrelation: Foursquare’s high-high census tracts formed a single cluster spanning Madrid’s historic centre and Salamanca district, whereas Panoramio revealed several clusters.The Panoramio clusters included the historic centre, Real Madrid stadium, and Torres Kio-Cuatro Torres.
  • Analysis of spatial autocorrelation: Jointly cross-referencing high-high clusters from the three sources classified census tracts by the combination of tourist activities present.A tract included in all three high-high clusters indicates, within 500 m, high densities of sightseeing, consumption, and internet connectivity opportunities.

6. CONCLUSIONS

The study compares Panoramio, Foursquare, and Twitter to map tourists’ urban activities in Madrid. Its conclusions emphasize that multiple complementary data sources are needed, while recognizing biases in digital footprints and their relevance for public policy.

  • Conclusions: The study tracks Madrid tourists’ digital footprints through Panoramio photographs, Foursquare check-ins, and Twitter interactions representing sightseeing, consumption, and connectivity.The analysis examines where tourists carry out these activities rather than relying on a single data source.
  • Conclusions: A single data source is insufficient because tourists engage in different activities across different urban spaces.All three sources show high tourist density in Madrid’s historic centre, which contains monuments, shops, hotels, and restaurants.
  • Conclusions: Panoramio and Twitter, and Panoramio and Foursquare, show little distributional similarity, whereas Foursquare and Twitter exhibit some similarity in spatial patterns.This comparison indicates that the data sources capture partly different dimensions of tourists’ presence.
  • Conclusions: Digital-footprint measures are biased because many tourists neither upload photographs nor use Twitter or Foursquare, and photographs may not represent all monuments.The passage also notes that photography restrictions, particularly in museums, limit photographic coverage.
  • Conclusions: Spatial knowledge of urban tourists is important for public policies aimed at improving tourist experience in high-prevalence spaces.Possible actions include pedestrian-only streets, wider pavements, expanded free-WiFi public spaces, and additional tourist information points.
Loading 1705.07951v1…