Source-linked AI summary
140 Characters to Victory?: Using Twitter to Predict the UK 2015 General Election
Pete Burnap, Rachel Gibson, Luke Sloan, Rosalynd Southern, Matthew Williams
TL;DR
Twitter election forecasting lacked consistent methods and faced population-bias concerns. This paper develops and applies a transparent baseline model to the UK 2015 General Election, finding a likely hung parliament with Labour gaining the most seats while underestimating SNP support.
Problem
Twitter election forecasts have produced mixed findings amid concerns about inconsistent methods, population bias, and insufficient attention to sentiment and existing party power.
Method
The paper builds a transparent baseline using pre-election Twitter data, sentiment scores, false-positive correction, and constituency-level seat projections based on 2010 results.
Results
The forecast indicated a likely hung parliament with Labour gaining the most seats, but significantly underestimated SNP seats.
Takeaways & Limitations
Twitter can serve as an electoral forecasting tool across national contexts when corrective steps address sentiment, party power, and related biases.
Takeaways & Limitations
Future applications need tweet geolocation and corrections for demographic bias in the Twitter sample.
Abstract
from arXiv · showhide
The election forecasting 'industry' is a growing one, both in the volume of scholars producing forecasts and methodological diversity. In recent years a new approach has emerged that relies on social media and particularly Twitter data to predict election outcomes. While some studies have shown the method to hold a surprising degree of accuracy there has been criticism over the lack of consistency and clarity in the methods used, along with inevitable problems of population bias. In this paper we set out a 'baseline' model for using Twitter as an election forecasting tool that we then apply to the UK 2015 General Election. The paper builds on existing literature by extending the use of Twitter as a forecasting tool to the UK context and identifying its limitations, particularly with regard to its application in a multi-party environment with geographic concentration of power for minor parties.
Introduction
Twitter election forecasting has produced mixed findings and methodological criticism, motivating a transparent baseline model for the UK’s multiparty 2015 election.
- Related literature: Prior Twitter election studies produced mixed findings and lacked consensus on forecasting methods.Critics emphasized the need for genuine pre-election forecasts, population-bias adjustments, existing party-power controls, and sentiment analysis.
- Related literature: Later studies found Twitter relationships with electoral outcomes after accounting for candidate, district, incumbency, and party-popularity factors.Other work modeled leader sentiment dynamically across social-media platforms before the 2010 UK election.
- Contribution: The paper proposes a basic, transparent Twitter forecasting model following the KISS principle used in agent-based modeling.The baseline is intended to be extended in later work.
- Contribution: The model issues a genuine forecast using data collected no later than one month before election day.It also incorporates tweet sentiment and predicts seats by accounting for existing parliamentary representation and party power in constituencies.
Methodology
The study collected party- and leader-related tweets, scored and filtered their sentiment, corrected selected false positives, and converted estimated vote shares into constituency-level seat projections.
- Data collection: 13,899,073 tweets were collected through Twitter’s streaming API between 28 November 2014 and 9 March 2015.Tweets were selected using party and leader names, with case-insensitive matching.
- Sentiment analysis: Sentiment software assigned each tweet positive and negative scores from -5 to +5.Tweets containing multiple search terms were removed because the system could not reliably identify which entity received the sentiment.
- Sentiment analysis: Tweets scoring below -1 were removed, while scores from -1 to +5 were summed into party and leader sentiment scores.The summed values were intended to capture overall sentiment magnitude rather than merely counting positive mentions.
- Data cleaning: Human annotation found that 78.9% of “Labour” tweets concerned the Labour Party, compared with 19.4% of “Greens” tweets concerning the Green Party.The Green Party error rate largely reflected tweets about the Australian Green Party, so weighting was applied.
- Seat projection: Vote shares were converted into seat forecasts by applying national swing from 2010 results to each UK constituency.The party with the maximum projected constituency value was assigned each seat.
Discussion and Conclusions
The forecast indicated a hung parliament with Labour gaining the most seats, but it underestimated SNP support because the model lacked tweet geolocation.
- Results: The forecast’s likely outcome was a hung parliament with Labour gaining the most seats.Predictions for other parties were described as within an expected range, while SNP seats were significantly underestimated.
- Limitations: Without geolocating tweets, the model underestimates regionally concentrated support such as that of the SNP.The calculation assumes individuals are randomly distributed across the UK, which is unsuitable for a majority regionalist party.
- Limitations: Only around 1% of tweets remain after applying precise-location geocoding, so the study retained its larger sample while accepting diluted SNP support.
- Contribution: The authors present the analysis as meeting at least three core criteria for a minimally acceptable forecast.
- Future work: Future applications need geolocation methods and corrections for demographic bias in the Twitter sample.The authors report that related work was underway for the next UK General Election.