Source-linked AI summary
The Geography of Happiness: Connecting Twitter sentiment and expression, demographics, and objective characteristics of place
Lewis Mitchell, Kameron Decker Harris, Morgan R. Frank, Peter Sheridan Dodds, Christopher M. Danforth
TL;DR
The paper addresses whether real-time social-media expressions can reveal geographic, demographic, and health-related variation in population happiness. It combines large-scale geotagged Twitter data with surveyed characteristics of states and urban populations, finding that word use supports happiness estimates and correlates with place characteristics such as education and obesity. The authors conclude that social media may potentially help estimate real-time population-level measures.
Problem
Survey-based well-being measures provide limited real-time, fine-grained evidence about how happiness varies across geographic places and relates to demographic and health characteristics.
Method
The study analyzes geotagged Twitter word frequencies across U.S. states and urban areas, scores words for happiness, and compares word use and happiness with census characteristics.
Results
The analysis estimates happiness for states and cities, groups cities by word-use similarity, and finds correlations between happiness, word use, and characteristics including education, obesity, wealth, and poverty.
Takeaways & Limitations
Social media may potentially be used to estimate real-time levels and changes in population-level measures such as obesity rates.
Abstract
from arXiv · showhide
We conduct a detailed investigation of correlations between real-time expressions of individuals made across the United States and a wide range of emotional, geographic, demographic, and health characteristics. We do so by combining (1) a massive, geo-tagged data set comprising over 80 million words generated over the course of several recent years on the social network service Twitter and (2) annually-surveyed characteristics of all 50 states and close to 400 urban populations. Among many results, we generate taxonomies of states and cities based on their similarities in word use; estimate the happiness levels of states and cities; correlate highly-resolved demographic characteristics with happiness levels; and connect word choice and message length with urban characteristics such as education levels and obesity rates. Our results show how social media may potentially be used to estimate real-time levels and changes in population-level measures such as obesity rates.
I. INTRODUCTION
The paper asks how urban living relates to well-being and develops a data-driven approach using geolocated Twitter expressions alongside demographic and geographic data. It aims to measure happiness across states and cities and explain variation through word use and census characteristics.
- Motivation: Cities are increasingly important settings for studying well-being as urbanization grows and fine-grained data become broadly available.The paper situates its question within the expanding quantitative study of urban phenomena.
- Motivation: Existing city- and state-level well-being measures rely almost exclusively on surveys, while social-network data offer complementary remote-sensing methods.The authors connect the growth of social-network data with data-driven sentiment analysis of large populations.
- Research aims: The paper investigates how geographic place correlates with and potentially influences societal happiness.It first examines states, then studies urban areas in the United States.
- Research aims: Happiness is estimated from word-frequency distributions in geolocated tweets, with individual words independently scored for happiness by Mechanical Turk users.The approach builds on prior word-based happiness measurement methods.
- Research aims: Word shifts and census correlations are used to examine how word usage relates to happiness and social and economic factors.The authors also use word-frequency distributions to group cities by similarities in observed word use.
II. DATA AND METHODOLOGY
The study analyzes over 10 million geotagged tweets from 373 urban areas and estimates happiness from word scores aggregated by frequency. It removes neutral words and compares the resulting measures with census data, while acknowledging that word context is not modeled.
- Data: The corpus contains over 10 million geotagged tweets from 373 urban areas in the contiguous United States collected during 2011.The sample is drawn from Twitter’s 2011 garden-hose feed and includes approximately 1% of tweets that were geotagged.
- Data: Urban areas follow 2010 Census Bureau MAF/TIGER boundaries, which can agglomerate nearby small towns with larger cities.This boundary definition shapes the geographic units analyzed.
- Sentiment measure: Happiness scores use roughly 10,000 LabMT words rated from 1 (sad) to 9 (happy) by Amazon Mechanical Turk users.The word list combines frequently occurring words from Google Books, music lyrics, the New York Times, and Twitter.
- Sentiment measure: A text’s average happiness is computed by weighting each scored word’s happiness value by its normalized frequency.For word wi, normalized frequency is pi = fi/Σ fi, where fi is its frequency.
- Assumption: The method ignores word context and text meaning, although the authors report reliable results for sufficiently large texts.They compare this aggregation to estimating room temperature from many particles rather than a small number.
- Preprocessing: Words with happiness scores between 4 and 6 are removed as neutral stop words before calculating text happiness.The authors describe this filtering as a balance between sensitivity and robustness.
- Validation variables: Happiness results are correlated with 2011 American Community Survey 1-year estimates.These census data provide the demographic and socioeconomic comparison variables.
III. HAPPINESS ACROSS STATES AND URBAN AREAS
The paper estimates happiness from geotagged Twitter word use across U.S. states and cities, then compares these measures with established well-being indicators and geographic patterns. Happiness varies across places, with Hawaii and Napa ranking highest and Louisiana and Beaumont lowest, while city happiness is strongly negatively associated with tweets per capita.
- States: State happiness correlates strongly with most comparison measures, but not with BRFSS, which significantly correlates only with the Gallup well-being index.The BRFSS uses data from 2005–2008, unlike the other measures, which use 2011 data.
- States: Word-use similarities reveal geographically plausible state clusters, including Vermont–New Hampshire and Louisiana–Mississippi, with Nevada as an outlier.The clusters are based on correlations between state word-frequency vectors.
- Urban areas: City-level mapping resolves local happiness differences, including relatively sadder Harlem and Washington Heights compared with Downtown Manhattan.The maps also reveal geographic features such as Manhattan’s outline, Central Park, streets, bridges, and airport terminals.
- Urban areas: 373 cities have a happiness distribution centered at havg = 6.00, with 220 cities above the overall average and 153 below it.The city range extends more than 0.2 around the mean and is skewed toward higher happiness.
- Urban areas: Happiness correlates negatively with tweets per capita, with Spearman correlation coefficient -0.558 and p-value less than 10^-16.This association is stronger than the negative correlation with the total number of tweets gathered per city.
- Urban areas: Napa, California is identified as the happiest contiguous-U.S. city at 6.26, while Beaumont, Texas is the saddest at 5.83.The rankings use average word happiness havg calculated from city tweets.
- Urban areas: Relative city happiness is associated with words such as “lol,” “haha,” “love,” “like,” negations, and profanity, alongside geographically patterned words such as “beach.”Coastal cities including Santa Cruz and Miami show “beach” high on their word lists.
IV. CORRELATING WORD USAGE WITH CENSUS DATA
The paper relates city happiness to demographic, health, and word-use characteristics using census and survey data alongside geotagged Twitter. It finds socioeconomic patterns, obesity associations, food-related word correlations, and city groupings based on word use.
- Census attributes: 432 demographic attributes across 373 cities were correlated with happiness, and only two clusters showed many significant associations at p < 0.01.The clusters broadly represented high and low socioeconomic status; high-status attributes generally correlated positively with happiness, while low-status attributes showed the opposite pattern.
- Word use and demographics: The normalized frequency of “cafe” correlated positively with the percentage of residents holding at least a bachelor’s degree, with Spearman r = 0.481 and p-value 4.90 × 10^-23.Word frequencies were normalized by each city’s total number of collected tweets before correlations were calculated across cities.
- Obesity: Happiness generally decreased as obesity increased across the 190 metropolitan areas included in the Gallup and Healthways 2011 survey.The study used metropolitan statistical areas, which were generally larger than the urban-area boundaries used for the census analysis.
- Obesity: Obesity and happiness had a statistically significant negative Spearman correlation of r = −0.339 with p-value 2.01×10^-6.Boulder had the lowest reported obesity rate at 12.1%, while Beaumont had the fifth highest at 33.8%.
- City word-use groupings: Cities formed clusters based solely on similarities in word usage, with all pairwise correlations exceeding r = 0.8 in the analyzed frequency distributions.Agglomerative hierarchical clustering identified geographically coherent examples but also exceptions to a simple regional pattern; Cleveland and Detroit were most alike at r = 0.995, while Austin and Baton Rouge were most dissimilar at r = 0.813.
V. DISCUSSION
The paper uses word-based happiness measures, word shifts, and census data to characterize urban variation in happiness and its associations with wealth, profanity, and obesity.
- The authors score states and cities for average word happiness and map areas of high and low happiness using a simple mathematical method.Word shift graphs identify words contributing most to differences from the US average, while socioeconomic census data contextualize word usage.
- Profanity was a significant driver of individual-city happiness scores, motivating future study of regional variation in swear-word use.The authors refer to this possible research direction as “geoprofanity.”
- Happiness within the United States showed a large positive correlation with household income and a strong negative correlation with poverty.The authors relate this within-country pattern to the first part of the Easterlin paradox.
- Happiness significantly anticorrelated with obesity, although chronic illnesses accompanying obesity may confound the relationship with psychological well-being.The paper notes that other studies have also reported inverse relationships between weight and depression.
Appendix A: Data set and states
Appendix A documents the census geometry and Twitter preprocessing used to construct city-level word frequencies and state-level geographic summaries.
- 3,592 cities follow an approximate power-law relationship between perimeter and area, with an average fractal dimension of α = 1.29.The smallest city in both measures is Richmond, California, while the largest is New York in the database.
- 15% of a user’s tweets containing selected weather-related words was the threshold for removing likely automated weather-reporting bots.The filtered terms were “humid,” “humidity,” “pressure,” and “earthquake.”
- The normalized word frequency distribution is f_hat(i) = f_i/n, where n is the total number of tweets collected for each city.Its sum represents the average number of LabMT words per tweet, approximately 7.1.
- 9 to almost 12 words per tweet spans the reported average message lengths for qualifying US cities, from Durham to New York.The cities shown had more than 50,000 LabMT words collected during 2011.
- The appendices provide logarithmic choropleths of raw and per-capita geotagged-tweet counts for each US state during 2011.The appendix also lists all state happiness scores and word correlations with demographic attributes.
B,C,D,E,F Online appendices
The online appendices provide complete state happiness scores, obesity-related word correlations, and additional city, demographic, and happiness materials.
- Appendix B contains word-shift graphs for all states, while Appendix C compares city happiness with the Gallup-Healthways wellbeing measure.Appendix C also includes tweet maps and city word-shift graphs.
- Table A1 lists happiness scores for every US state in descending order.
- Appendices D, E, and F provide complete demographic correlations, LabMT words ordered by happiness correlation, and a daily-updating US happiness map.
- Tables A2 and A3 list the 25 words with strongest positive and negative Spearman correlations to obesity in 2011.Neutral stop words with 4 < h_avg < 6 were removed from these lists.
- Table A4 lists 24 food-related words showing the least correlation with obesity, all with p-values greater than 0.9.The words are arranged in decreasing order of p-value.