Source-linked AI summary

A Survey of Location Prediction on Twitter

Xin Zheng, Jialong Han, Aixin Sun

arXiv:1705.03172v2cs.SIcs.IR

TL;DR

Location prediction on Twitter must handle incomplete, noisy, and ambiguous location information across users, tweets, and textual mentions. This survey synthesizes the field by defining three prediction tasks, organizing methods around Twitter content, network, and context, and reviewing related problems. It concludes that content is central across tasks, the network is especially important for home-location prediction, and methods must account for noisy social connections and task-specific distinctions.

  • Problem

    Twitter location information is incomplete and noisy, while mentioned locations also face variability and ambiguity in short, casual tweets.

  • Method

    The survey defines home, tweet, and mentioned location prediction, reviews evaluation metrics, and organizes prior methods by their use of Twitter content, network, and context.

  • Results

    The survey concludes that all three prediction problems rely heavily on tweet content, while Twitter network information plays a key role in home-location prediction.

  • Takeaways & Limitations

    Location prediction methods should distinguish task-specific signals and account for noisy Twitter friendships when using network information for home-location inference.

  • Takeaways & Limitations

    The survey includes only a small portion of point-of-interest recommendation studies because of its scope.

Abstract

from arXiv · show

Locations, e.g., countries, states, cities, and point-of-interests, are central to news, emergency events, and people's daily lives. Automatic identification of locations associated with or mentioned in documents has been explored for decades. As one of the most popular online social network platforms, Twitter has attracted a large number of users who send millions of tweets on daily basis. Due to the world-wide coverage of its users and real-time freshness of tweets, location prediction on Twitter has gained significant attention in recent years. Research efforts are spent on dealing with new challenges and opportunities brought by the noisy, short, and context-rich nature of tweets. In this survey, we aim at offering an overall picture of location prediction on Twitter. Specifically, we concentrate on the prediction of user home locations, tweet locations, and mentioned locations. We first define the three tasks and review the evaluation metrics. By summarizing Twitter network, tweet content, and tweet context as potential inputs, we then structurally highlight how the problems depend on these inputs. Each dependency is illustrated by a comprehensive review of the corresponding strategies adopted in state-of-the-art approaches. In addition, we also briefly review two related problems, i.e., semantic location prediction and point-of-interest recommendation. Finally, we list future research directions.

1 INTRODUCTION

This survey addresses Twitter location prediction across home, tweet, and mentioned locations, motivated by Twitter’s broad, real-time data and the incompleteness and noise of available location information. It organizes prior work around Twitter content, network, and context, while reviewing related research and future directions.

  • Scope: The survey focuses on predicting users’ home locations, tweet locations, and locations mentioned in tweets.These location types connect online activity with real-world events and support applications including public health monitoring, local recommendations, regional topic summarization, and emergency identification.
  • Challenges: Twitter location information is incomplete and often inaccurate: only 21% of users in one U.S. dataset reported residential cities, 5% gave home coordinates, and just 0.77% or 0.4% of tweets had attached location information in two studies.Self-declared home information may also be noisy or invalid.
  • Challenges: Twitter’s casual language introduces acronyms, misspellings, and special tokens, making location recognition and disambiguation more challenging than in formal documents.Users may nevertheless reveal location through local landmarks, events, dialects, slang, or other location-relevant words.
  • Motivation: Twitter’s worldwide coverage and real-time freshness have made location prediction an active research area despite tweets’ noisy, short, and context-rich nature.Twitter location prediction has also appeared as a shared task in the 2nd Workshop on Noisy User-generated Text.
  • Survey organization: The survey systematically reviews how Twitter content, network, and context support the three location prediction problems, then briefly covers semantic location prediction and point-of-interest recommendation.It also reviews evaluation metrics and discusses future research directions.

2 PROBLEM OVERVIEW

The survey models Twitter location prediction through three input types—tweet content, Twitter network, and tweet context—and distinguishes home, tweet, and mentioned location tasks. It defines task representations and ground-truth choices while outlining recognition and disambiguation for mentioned locations.

  • Inputs: Twitter provides three major input sources for geolocation problems: short noisy tweets, the user network, and heterogeneous contextual information.Context includes timestamps, geo-tags, and user-profile information.
  • Twitter network: Twitter friendship is unidirectional, and following does not necessarily indicate real-life friendship, although online mentions can provide clues to real-life relationships.The survey therefore treats following and mentioning actions uniformly when describing the Twitter network.
  • Home location prediction: Home locations are users’ long-term residential addresses and may be represented as administrative regions, geographical grids, or geographical coordinates.The appropriate granularity should be fixed for evaluation, while profile information and aggregated geo-tags may be combined to improve ground-truth coverage.
  • Tweet location prediction: Tweet location is the place where a tweet is posted, usually represented by a point of interest or coordinates derived from tweet geo-tags.Unlike home locations, tweet locations are generally based on tweet geo-tags and are not usually represented as administrative regions or grids.
  • Mentioned location prediction: Mentioned location prediction includes recognizing text fragments that refer to locations and disambiguating those fragments against entries in a location database.Both administrative regions and points of interest may be represented, with BIO or BILOU labeling commonly used for recognition.

2.3 Twitter Inputs for Location Prediction Problems

Location prediction on Twitter uses distance-based or token-based evaluation metrics, depending on whether locations are represented geographically or as discrete symbols. These metrics capture error, tolerance, ranking, exact correctness, and cases where systems abstain.

  • Distance-based metrics: Distance-based metrics represent predicted and ground-truth locations by geographical coordinates and measure their separation.Error Distance is the distance between the predicted and ground-truth locations.
  • Distance-based metrics: Mean and Median Error Distance aggregate prediction errors across users or tweets, with the median less sensitive to wildly inaccurate predictions.Mean Squared Error instead squares each Error Distance and is used by relatively few studies.
  • Token-based metrics: Token-based metrics treat locations as discrete symbols such as countries, cities, grids, or POIs, enabling broader usage despite ignoring geographical information.Exact Accuracy counts only predictions that coincide with the ground truth.
  • Token-based metrics: Ranking-based Acc@k counts a prediction list as correct when the ground-truth location appears among its top-k results.This preserves useful alternatives that ordinary Accuracy would discard.
  • Abstentions and recognition: Precision, Recall, and F1 evaluate systems that may abstain, including mentioned-location recognition where fragment boundaries must exactly match the ground truth.F1 is the harmonic mean of Precision and Recall.

3 HOME LOCATION PREDICTION

Home location prediction estimates where Twitter users live to support location-dependent applications. The task is difficult because profile locations are optional and often absent or noisy.

  • Motivation: User home locations support local content recommendation, location-based advertising, public health monitoring, and public opinion polling.Most studies predict home locations at the city level, sometimes at the state or country level.
  • Challenge: Optional profile completion makes Twitter users’ home locations mostly absent or noisy, motivating research on prediction.

3.1 Inference based on Tweet Content

Content-based home location prediction exploits words associated with places, but must distinguish genuinely local language from broadly used Twitter vocabulary. Existing methods are organized into word-centric and location-centric approaches.

  • Location-indicative content: Tweet content can reveal home locations through place-associated words, dialects, and references such as “Houston Rockets,” “howdy,” and “phillies.”
  • Method classes: Word-centric methods estimate p(l|w), whereas location-centric methods model the probability of generating a tweet d at location l, p(d|l).
  • Locality filtering: Content-based prediction must identify words with strong locality because generic terms such as “downtown” and “OMG” occur across Twitter.

Identifying Local Words

Local-word identification is necessary because frequent but location-irrelevant terms can make home-location predictions effectively random. Research addresses this through unsupervised statistical measures and supervised geographical classification.

  • Motivation: Location-irrelevant words can outnumber genuinely local words, so indiscriminate use may drive home-location predictions toward random results.Examples of broadly used terms include “downtown” and “OMG,” while “howdy” and “phillies” are more locally indicative.
  • Research direction: A substantial body of research focuses on identifying local words through either unsupervised or supervised methods.
  • Unsupervised methods: Unsupervised approaches identify local words using statistics computed directly from data, including Kernel Density Estimation and Ripley’s K statistic.These methods respectively smooth term occurrences spatially and measure geographical deviation.
  • Supervised methods: Supervised approaches formulate local-word identification as classification after fitting each word’s geographical distribution with a spatial variation model.The model represents a word using a geographical center, center frequency C, and dispersion ratio α.

Modeling Spatial Word Usage

Spatial-word approaches predict home locations by identifying locally informative words and modeling how their usage varies geographically. The survey contrasts probabilistic, classification, retrieval, and deep-learning strategies for this task.

  • Probabilistic approaches: Probabilistic methods model the conditional distribution of users’ home locations given their tweet contents, then decompose it to make predictions.These models focus on local words and spatial word usage.
  • Probabilistic approaches: Directly estimating P(l|w) from corpus counts is inferior because unobserved words in less populated locations may still be relevant.This sparsity problem motivates smoothing techniques.
  • Probabilistic approaches: Laplace, state-level, and grid-based neighborhood smoothing address sparse spatial word distributions, with Laplace smoothing assigning every location positive probability.Laplace smoothing does not incorporate geographical information, whereas the other methods use spatial structure.
  • Alternative approaches: Classification methods represent users with local-word statistics and predict among candidate locations as labels, including multinomial Naive Bayes and hierarchical ensembles.One approach uses the top 10,000 words ranked by CALGARI scores as term-frequency features.
  • Alternative approaches: Information-retrieval methods treat locations as pseudo-documents of resident tweets and retrieve the most similar location for a user.Grid-based language models are used to represent locations and compare pseudo-documents.
  • Alternative approaches: Recent deep-learning work encodes chronologically ordered messages with recurrent models and attention, while applying a similar process to context.The attention mechanism produces a global message representation emphasizing important information.

3.2 Inference Based on Twitter Network

Twitter-network methods infer home locations from relationships among users, extending direct friendship models toward social closeness, influence, and global inference. The survey emphasizes that network-connected users violate the independence assumption of typical prediction tasks.

  • Friendship-based methods: Friendship-based methods assume nearby friends tend to share home locations, while global inference addresses dependencies among multiple unknown user locations.Local inference uses one- or two-hop friendship or mentioning; global methods account for interlinked users jointly.
  • Friendship-based methods: A Facebook-derived model fits friendship probability as P(ui, uj are friends |dist(ui, uj) = x) = a(b + x)−c, with c = 1 producing a good fit.Under this model, friendship probability is inversely proportional to home distance, and locations are inferred by maximizing observed friendship-link probability.
  • Social-closeness-based methods: Twitter friendship is an imperfect proxy for home proximity: its friendship-distance distribution is bimodal, with peaks around 10 miles and far away.Subsequent studies therefore estimate social closeness rather than relying only on online friendship.
  • Social-closeness-based methods: Mentions and conversations provide social-closeness signals, and studies construct mention graphs or use unidirectional mentions when bidirectional mentions are too rare.These methods optimize or infer locations using interaction relationships beyond following links.
  • Influence-based methods: Influence can be negatively associated with social closeness because users may follow geographically distant celebrities, making social influence distinct from real-life familiarity.This motivates separating influence-driven following from proximity-related relationships.
  • Influence-based methods: Influence-based models represent each user’s influence as a location-centered bivariate Gaussian and learn unknown locations and influence scopes by Maximum Likelihood Estimation.Landmark users with many friends in a small region are another proposed cue for home-location prediction.
  • Global inference: Network-based home-location prediction is technically non-independent, yet many reviewed methods still perform local inference from one- or two-hop relationships.The survey identifies implementation problems that remain even when friendship and social-closeness features are carefully designed.

3.3 Inference based on Tweet Context

Tweet-context methods use posting time and self-declared profile information as contextual signals for home-location prediction. The reviewed approaches include time-zone classification and continuous geo-tag representations.

  • Contextual signals: Posting time and self-declared profiles, including locations and time zones, are the main tweet-context information used for home-location prediction.These signals are categorized as tweet context in the survey.
  • Posting time: Time-zone classifiers represent users by distributions of tweet-posting times after binning GMT days into equal-length slots.Time shifts in users’ posting-time distributions reveal their time zones and can support home-location prediction.
  • Geo-tags: Neural-network models with mixture density networks convert two-dimensional geo-tags into continuous vector spaces for input.

3.4 Summaries and Discussions

The survey summarizes home-location prediction as relying on both tweet content and Twitter network information, with additional context also contributing. A large systematic comparison evaluates competing methods using two ground-truth constructions.

  • Summary: Home-location methods rely equally on tweet content and Twitter network information, while content methods are word-centric or location-centric.Network methods model dependencies through friendship and interactions, including global inference approaches.
  • Systematic comparison: A systematic comparison evaluates methods from prior work on 1.3 billion tweets, 15 million users, and 26 million following relationships.The comparison uses both self-declared home locations and aggregated geo-tags as ground truth.

4 TWEET LOCATION PREDICTION

Tweet location prediction estimates where individual tweets were posted, using content, topics, network signals, and temporal or other contextual information. Compared with home-location prediction, it generally targets finer-grained and more dynamic locations, especially points of interest.

  • Task definition: Tweet location prediction typically receives one tweet rather than a user’s full tweet history, creating a distinct input setting from home-location prediction.The two tasks share content-based techniques but differ in available evidence.
  • Tweet content: Tweet-location methods model spatial word and n-gram usage, often smoothing distributions with Gaussian or Gaussian-mixture models.Location-centric alternatives rank locations using language-model likelihoods or KL-divergence.
  • Classification: Classification approaches assign tweets to discretized cells, fine-grained POIs, or location types, but small grids create data-sparsity problems.Some methods use Gaussian kernels for cell priors and word-conditionals, while others combine content with social relationships.
  • Topic and interest models: Geo-topic models represent geographical variation through latent topics, regions, and user interests that connect tweet content with locations.Examples include location-varied topics, location functions such as eating or shopping, and work or home regions.
  • Context and comparison: Tweet locations are usually predicted at POI-level granularity and are more dynamic than home locations, while Twitter-network use is less common.Temporal context can use a tweet’s timestamp, and dynamic models can incorporate friends’ real-time locations and historical trajectories.
  • Context and comparison: Home-location prediction commonly uses classification and Twitter network evidence, whereas tweet-location prediction relies more heavily on tweet content.The survey notes that the two tasks are not always clearly separated when users’ tweets and geo-tags are handled generically.

5 MENTIONED LOCATION PREDICTION

Mentioned location prediction identifies and resolves locations named in tweets, where short, noisy text intensifies recognition and disambiguation challenges. The survey reviews content-based pipelines, gazetteers, collective coherence, contextual signals, and joint recognition-linking approaches.

  • Task and motivation: Mentioned locations are revealed through names in tweets and require preprocessing before subsequent location analysis.They may describe restaurants, shopping malls, cinemas, parades, or disaster-related places.
  • Challenges: Tweet noise and brevity worsen entity variability and ambiguity, making location recognition and linking harder than in well-formatted documents.Informal forms can remove standard clues such as capitalization, “street,” and “at.”
  • Recognition and linking: Mention recognition and disambiguation generally use tweet content in a pipeline, with lexical boundaries helping recognition and surrounding words helping resolve references.Some studies instead couple the two components so disambiguation feedback can correct recognition outputs.
  • Recognition methods: Location-specific recognizers commonly use gazetteers such as Geonames or Foursquare, sometimes combining noun-phrase extraction with fuzzy matching.General tweet NER systems may rebuild the full pipeline for informal text rather than directly applying formal-document tools.
  • Disambiguation methods: Disambiguation resolves name collisions by exploiting location hierarchies, coherence among mentions, distances among POIs, or user-level interests.Examples include parent-child and sibling relations, collectively chosen POIs, and a user’s living city or entity interests.
  • Task distinctions: Mentioned locations depend heavily on tweet content and somewhat on network and context, but they do not necessarily indicate where a tweet was posted.A tweet can mention a place visited previously or planned for the future.
  • Empirical analysis: Human-annotated comparisons report that retrained StanfordNER outperforms the other competitors on disaster-related Twitter data.The survey also describes comparisons involving StanfordNER, OpenNLP, Yahoo! PlaceMaker, TwitterNLP, GeoLocator, and UnlockText.

6 OTHER RELATED PROBLEMS

The survey briefly contrasts semantic location prediction and point-of-interest recommendation with Twitter-based location prediction. These related problems differ in their targets, ground truths, data sources, and typical solution frameworks.

  • Semantic location prediction: Semantic location prediction targets places discussed in tweets, which may differ from the locations where tweets are posted.A tweet can discuss New York while its author is currently in Japan.
  • Semantic location prediction: Semantic-location studies often model latent user locations and restaurant-specific language, evaluating against manually annotated tweets.The subjective definition of semantic location makes annotation necessary.
  • Semantic location prediction: Manual ground truth is more laborious to obtain than geo-tags, and reviewed semantic-location evaluations contain only hundreds or thousands of annotated tweets.The survey identifies this annotation burden as one reason the problem attracts less attention.
  • Point-of-interest recommendation: POI recommendation suggests places users have not visited, using historical check-ins, ratings, comments, time, and current location rather than requiring a new tweet.The task is distinct from predicting a location already connected to a user or tweet.
  • Point-of-interest recommendation: POI recommendation generally uses collaborative filtering, with friendship, content, and context drawn mainly from location-based social networks.Twitter can serve as a network or data source when platform APIs restrict access to check-in histories.
  • Scope: The survey’s POI-recommendation discussion is limited to clarifying connections and differences with Twitter-based location prediction, not providing an extensive review.It explicitly refers readers to broader recommendation surveys.

7 CONCLUSION AND FUTURE WORK

The survey organizes Twitter geolocation around home, tweet, and mentioned locations, emphasizing content, network, and context as complementary inputs. It identifies data sparsity and insufficient joint modeling as priorities for future work.

  • The survey reviews home-location, tweet-location, and mentioned-location prediction, and relates them to document geolocation, semantic location prediction, and POI recommendation.
  • All three prediction problems rely heavily on tweet content, with word-centric and location-centric approaches forming the main categories for home and tweet location prediction.Word-centric methods model local-word usage, while location-centric methods construct pseudo-documents or classifiers for locations.
  • Mentioned-location systems address noisy content with sophisticated features and gazetteers, while collective disambiguation and joint recognition-disambiguation optimization address information scarcity.
  • Twitter network information is especially important for home-location prediction, motivating social-closeness modeling and global inference because users’ predictions can depend on one another.
  • Future work should combine Twitter properties with neural models, jointly model content, context, and network, and address sparsity through images, cross-platform information, or auxiliary knowledge.
Loading 1705.03172v2…