Source-linked AI summary
Temporal patterns of happiness and information in a global social network: Hedonometrics and Twitter
Peter Sheridan Dodds, Kameron Decker Harris, Isabel M. Kloumann, Catherine A. Bliss, Christopher M. Danforth
TL;DR
The paper addresses how societal happiness can be measured beyond self-report and economic indicators. It uses Twitter text, human ratings of over 10,000 words, and frequency-based analysis to study happiness and information over time. The resulting hedonometer is reported as robust, while its happiness measure is tuned to current experiential rather than long-term reflective evaluations.
Problem
Societal happiness matters scientifically and as a complement to economic measures, but is normally measured through self-report.
Method
The study remotely measures happiness in Twitter expressions by combining word frequencies with independently assessed happiness scores in the labMT 1.0 dataset.
Results
The analysis examines temporal happiness patterns across overall, daily, weekly, keyword-specific, and word-level views, alongside information content that is generally uncorrelated with happiness.
Takeaways & Limitations
Twitter provides a non-invasive way to remotely sense exhibited happiness at very large population scale without asking people how happy they are.
Takeaways & Limitations
The approach measures current experiential happiness rather than individuals’ longer-term reflective evaluations of their lives.
Abstract
from arXiv · showhide
Individual happiness is a fundamental societal metric. Normally measured through self-report, happiness has often been indirectly characterized and overshadowed by more readily quantifiable economic indicators such as gross domestic product. Here, we examine expressions made on the online, global microblog and social networking service Twitter, uncovering and explaining temporal variations in happiness and information levels over timescales ranging from hours to years. Our data set comprises over 46 billion words contained in nearly 4.6 billion expressions posted over a 33 month span by over 63 million unique users. In measuring happiness, we use a real-time, remote-sensing, non-invasive, text-based approach---a kind of hedonometer. In building our metric, made available with this paper, we conducted a survey to obtain happiness evaluations of over 10,000 individual words, representing a tenfold size improvement over similar existing word sets. Rather than being ad hoc, our word list is chosen solely by frequency of usage and we show how a highly robust metric can be constructed and defended.
Introduction
The paper develops a Twitter-based approach to remotely measure societal happiness and information, then examines their temporal variation across multiple timescales. Its method combines human word-happiness evaluations with word-frequency analysis, while acknowledging limits from Twitter’s non-representative user population.
- Societal happiness is treated as a crucial complement to economic measures such as gross domestic product.
- The study remotely senses societal-scale happiness from Twitter’s brief, in-the-moment textual expressions.Twitter’s format is presented as an input signal for a real-time societal hedonometer.
- The hedonometer combines word frequencies with independently assessed happiness scores for over 10,000 words in the labMT 1.0 dataset.The word evaluations were obtained using Amazon’s Mechanical Turk, and the dataset is supplied as supplementary information.
- The analysis covers happiness time series, daily and weekly cycles, keyword-specific expressions, word-level comparisons, and information content.Information is estimated through lexical size or effective vocabulary size derived from generalized entropy measures.
- The authors report that happiness and information are generally uncorrelated quantities.
- Twitter data can reflect current circumstances, with food-related words showing expected daily peaks and cultural terms maximizing around television airing times.
C. Robustness and Refinement of Hedonometer
The hedonometer is refined by excluding words near neutral happiness, producing a tunable family of metrics whose temporal outputs remain highly consistent across a broad parameter range. The selected setting, ∆havg = 1, balances sensitivity and robustness while retaining meaningful corpus coverage, and the enlarged word list improves measurement resolution.
- Metric refinement: Excluding words within ∆havg of neutral happiness creates a tunable family of hedonometer metrics.The excluded band is centered on the neutral score of 5 and has width 2∆havg.
- Robustness: 0.5 ≲∆havg ≲2.5 produces highly correlated happiness time series, demonstrating robustness across parameter choices.Pearson correlations remain impressively high across this central range, while larger values increase sensitivity at the cost of coverage.
- Metric selection: ∆havg = 1 is selected as a compromise between sensitivity and robustness while remaining above the transition near ∆havg ≃0.5.The selected value is supported by the observed correlation structure and coverage analyses.
- Coverage: 3,686 of 10,222 evaluated words remain at ∆havg = 1, covering approximately 23% of the Twitter corpus.The corresponding coverage is substantially greater than the 3.7% reported for the 1,034-word ANEW list.
- Coverage: At ∆havg = 1, coverage reaches 40–50% for words with frequency rank r ≤5,000.Coverage declines as the happiness band expands and neutral words are excluded.
- Instrument refinement: The labMT 1.0 word list reproduces earlier trends while providing greater resolution and fidelity than the previously used ANEW list.The expanded, frequency-selected list also sharpens observations and reveals patterns that were previously hidden.
- Limitations: The method is intended for very large texts because short texts can be too ambiguous for reliable word-frequency-based happiness measurement.The authors identify small-text fallibility as a limitation but focus on large data sets.
- Limitations: The hedonometer measures exhibited happiness as perceived from word frequencies rather than individuals’ internal emotional states.Its simplicity omits text structure but is reported to remain meaningful for sufficiently large texts.
IV. OVERALL TIME DYNAMICS OF HAPPINESS AND INFORMATION
The Twitter happiness time series contains broad trends, recurring weekly cycles, and sharp event-linked deviations. Word-shift analysis attributes these changes to specific increases and decreases in word usage.
- Happiness rose from January to April 2009, then gradually declined, with the decline accelerating during the first half of 2011.
- Weekend happiness generally peaked, while Monday and Tuesday formed the weekly nadir.
- Outlier Dates: Annual celebrations produced positive outliers, whereas disasters, deaths, and societal trauma typically produced negative outliers.
- Outlier Dates: May 2, 2011, following reports of Osama Bin Laden’s killing, was the lowest-happiness day across the entire time frame.
- Word Shift Analysis: Word-shift graphs rank words by absolute contribution to happiness changes and distinguish their emotional valence from their relative prevalence.The analysis uses a 14-day reference window and compares it with the focal date.
- Word Shift Analysis: The Bailout and Bin Laden drops were dominated by more frequent negative words and less frequent positive words, while the Royal Wedding showed the opposite pattern.The first 1,000 words typically account for more than 99% of the complete shift across 3,686 words.
C. Information Content
Information content increased substantially over time, and the paper attributes this change to growing use of non-English languages despite continued growth in English tweets.
- Simpson lexical size NS increased from approximately 300 to 700 beginning around July 2009.
- Monthly estimates of NS were smooth, indicating that the measure was unaffected by missing data and non-uniform sampling rates.Monthly NS was computed from the month’s word distribution rather than by averaging daily NS values.
- The more-than-doubling of NS was attributed to a strong relative increase in non-English languages, especially Spanish.Examples of growing Spanish words included ‘que’, ‘la’, ‘y’, ‘en’, and ‘el’.
- Figure 5 measures weekday happiness using equal-weight averages across Mondays, Tuesdays, and other weekdays from May 21, 2009 to December 31, 2010.
V. WEEKLY CYCLE
Twitter happiness follows a robust weekly cycle: it peaks across Friday–Sunday and reaches its minimum on Tuesday. Word shifts show that happier Saturdays reflect more positive-word use, although negative and socially adverse language remains present.
- Weekly happiness pattern: Saturday has the highest average happiness at h_avg ≃6.06, followed by Friday and Sunday, while Tuesday is the weekly low.Thursday reaches h_avg ≃6.03 after small increases from Tuesday through Wednesday and Thursday.
- Survey comparison: Mechanical Turk ratings broadly preserve the weekday ordering but rate Monday lowest rather than Tuesday and place Sunday above Friday.Isolated-word ratings span 4.30 for Monday to 7.42 for Saturday, a wider range than tweet averages.
- Robustness over time: Friday–Saturday–Sunday forms the peak and Tuesday the minimum in each of four approximately equal time periods, supporting a robust weekly pattern.Only Thursday in one period changes the overall ordering of days.
- Word shift analysis: Saturday’s higher happiness also includes evidence of boredom, fighting, and suffering associated with excessive drinking.The overall positive shift is driven mainly by more frequent positive words and, to a lesser extent, less frequent negative words.
- Word shift analysis: The first 50 words account for approximately 60% of the Saturday–Tuesday happiness shift, with positive changes including ‘love’, ‘haha’, ‘party’, ‘fun’, and ‘happy’.The weekday distributions were averaged over May 21, 2009 to December 31, 2010 with outlier dates removed.
C. Information Content
Information content varies across the week and day on a pattern distinct from happiness in some respects but broadly similar over the daily cycle. Simpson lexical size is highest on Friday and overnight, while word shifts identify the language changes underlying these differences.
- Weekly information cycle: Average Simpson lexical size peaks on Friday, declines through the weekend to a Sunday minimum, and has a smaller workweek low on Tuesday.The pattern remains unchanged under different averaging schemes.
- Weekly information cycle: Friday’s larger Simpson lexical size than Sunday’s is attributed primarily to frequency changes in around 100 words, especially ‘I’, ‘RT’, ‘you’, ‘me’, and ‘my’.These words are typically near the start of a Zipf ranking.
- Daily cycles: The happiest hour is 5–6 am, followed by a decline to the daily low at 10–11 pm; common profanities show a roughly anticorrelated cycle.Profanity use peaks around 1 am and is lowest during the 5–6 am happiness peak.
- Daily word shifts: At 5–6 am, the happiness advantage over 10–11 pm is driven by more positive and less abundant negative words, with the first 50 words explaining approximately 70% of the shift.Examples include ‘morning’, ‘haha’, and ‘happy’, alongside reduced use of ‘no’, ‘don’t’, and ‘shit’.
- Daily information cycle: Simpson lexical size follows a broadly similar daily cycle, rising overnight to NS ≃600 at 5–6 am and reaching NS ≃510 at 10–11 pm.It drops rapidly to a morning local minimum before a smaller early-afternoon crest.
- Daily information cycle: Tweets appear richer and less predictable at night, with an information apex near biological midnight; automated tweets are offered as a possible explanation beyond the study’s scope.Alternate averaging schemes produce remarkably little variation in NS.
VII. HAPPINESS AVERAGES AND DYNAMICS FOR TWEETS CONTAINING KEYWORDS AND PHRASES
The paper examines temporal happiness patterns for tweets containing words, phrases, dates, punctuation, emoticons, and phonemes. These text elements range from long-term topics such as ‘economy’ to contemporary and everyday expressions.
- Scope of text-element analysis: The analysis generates temporal happiness patterns for diverse text elements, including keywords, short phrases, dates, punctuation, emoticons, and phonemes.Examples span ‘economy’, ‘Obama’, ‘today’, and ‘!’.
A. Definition of Ambient Happiness
Ambient happiness measures the average happiness of words surrounding a text element after removing the background happiness of all tweets. This separates contextual associations from the element’s own happiness score.
- Definition: Ambient happiness is computed from co-occurring words while excluding the target text element’s own contribution.The average happiness of all tweets in the same pool is removed to create a differential time series.
- Definition: Normalized happiness includes the text element’s own score, whereas ambient happiness excludes it when evaluating tweets containing that element.Both measures compare the text-element subset with the overall tweet pool.
- Illustrative results: Tweets containing ‘happy’ remain about +0.3 to +0.4 above the overall average, while ‘sad’ stays near −0.2 and ‘:)’ and ‘:(’ average near +0.25 and −0.5.The exclamation point is positive but trends slightly downward toward neutrality.
- Illustrative results: ‘Afghanistan’ has consistently negative ambient happiness, while ‘Tea Party’ shows uneven signals and reaches its lowest score when usage is most frequent.The examples demonstrate variation across contemporary issues and text elements.
B. Overall Ambient Happiness for Specific Tweets
Across selected keywords and text elements, ambient happiness varies meaningfully by topic and social reference, while happiness and lexical information are generally independent. Word-level happiness assessments align strongly with contextual ambient and normalized measures on average.
- Scope and measures: The 100-element list spans political, economic, personal, semantic, and expressive terms, with ambient and normalized happiness rankings reported alongside lexical size.The table includes keyword matches, phrase and punctuation assessments, and lexical-size rankings based on tweets containing each element.
- Happiness and information: Text-element happiness assessments correlate strongly with ambient happiness (rs = 0.794) and normalized happiness (rs = 0.984), although individual sentences need not rigidly follow this structure.These correlations describe average contextual behavior rather than deterministic sentence-level structure.
- Topic patterns: Political terms tend to combine below-average happiness with large lexical sizes, including Obama, Sarah Palin, and George Bush.Their ambient happiness values are −0.173, −0.681, and −0.747, with lexical sizes 326, 275, and 288, respectively.
- Topic patterns: Personal pronouns show a prosocial happiness ordering, while self-reference is associated with richer lexical content than references to others.‘our’ and ‘you’ exceed ‘I’ and ‘me’ in happiness, whereas ‘me’ and ‘we’ have larger lexical sizes than ‘they’ and ‘them’.
- Topic patterns: Happy emoticons generally have higher ambient happiness and information, but semicolon winks rank highest for information despite ranking only third and fourth for happiness.Information levels range from NS=305 to NS=477 across the listed emoticons.
- Happiness and information: Ambient happiness and lexical size show no overall correlation across the 100 elements, with Spearman’s rs = −0.038 (p-value ≃0.71).The authors interpret this as evidence that the two quantities are generally independent, while recommending that both be reported for large-scale text characterization.
C. Analysis of Four Example Ambient Happiness Time Series
Four keyword-specific time series show sharp happiness declines associated with widely covered negative events. Word shifts identify increased negative vocabulary and decreased positive vocabulary as the main contributors to these changes.
- Tiger Woods: Tiger Woods tweets show an abrupt November 2009 happiness decline followed by a rebound toward a slightly below-average steady state.The decline coincided with intense media coverage of publicly reported extramarital affairs, while words such as ‘accident’, ‘crash’, and ‘scandal’ pulled happiness downward.
- Tiger Woods: The Woods word shift also exposes a word-centric limitation: contextually negative uses of relatively happy words such as ‘car’ and ‘sex’ can distort scores.The authors state that such microscopic errors are overcome for sufficiently large texts.
- BP: BP tweets fell by 0.47 in ambient happiness after the Deepwater Horizon explosion, as negative words became more frequent and positive words less frequent.Words including ‘disaster’, ‘damage’, and ‘blame’ increased, while ‘love’, ‘me’, ‘haha’, and ‘lol’ decreased.
- Pope: Pope-related tweets reached a clear happiness minimum in March 2010 while their relative frequency changed little, coinciding with intensified coverage of the Church child-molestation scandal.The corresponding word shift attributes the nadir to more frequent negative words.
- Israel: Israel-related tweets reached their lowest happiness in January during the Gaza War, while tweet volume increased with media reporting of the conflict.The January–February word shift identifies major vocabulary changes associated with the decline.
VIII. CONCLUDING REMARKS
The paper reports robust happiness and information patterns across hours, days, months, and years, while extending a transparent word-based measurement framework. It also identifies methodological improvements and boundaries on generalizing Twitter-derived findings.
- Findings and contribution: Temporal analyses uncover patterns across hours, days, months, and years, with weekly and daily cycles appearing especially robust.The authors describe the seven-day cycle as historically and culturally constructed rather than universal in origin.
- Findings and contribution: The labMT 1.0 word list expands the happiness assessment resource and is presented as useful to other researchers.The paper also highlights word-shift graphs as a comparative tool for explaining differences in word composition and tone.
- Scope and limitations: The study’s scope remains bounded by Twitter data, unresolved generalizability to the broader population, and the possibility that users manipulate online expressions.The authors also note that extracting small-scale patterns for rare topics remains an open question.
- Future work: Future methodological work should incorporate common n-grams, negated sentiments, contextual meanings, and improved information-content handling.The authors specifically mention phrases such as ‘child abuse’ and ‘sex scandal’, negations such as ‘not happy’, and possible future use of Shannon’s entropy.
- Scope and limitations: The paper frames big-data social science as expanding established social science through description and pattern finding before explanation and experimentation.This conclusion emphasizes data abundance as a shift in research practice rather than a replacement of the field’s core.
Supplementary Material
The supplementary material provides supporting figures, the labMT 1.0 Mechanical Turk data, and documentation for using the released word set.
- Supplementary contents: The supplement includes supporting figures and a table for the paper’s analyses.
- labMT 1.0 data: The labMT 1.0 data contain 10,222 words with Mechanical Turk happiness evaluations and associated metadata.The file reports average happiness from 50 user evaluations and the standard deviation of happiness.
- labMT 1.0 data: The word set is ordered by descending average happiness and organized into eight columns.The supplement documents the ordering and column structure for reuse.
- Dataset use: The paper asks researchers to cite the study and use the abbreviation labMT 1.0 when referring to the dataset.
7. New York Times rank,
The supplementary material documents word-shift and time-series analyses of happiness, lexical size, and event-related language patterns. It also presents comparisons across alternative distribution-construction methods and a ranked keyword table.
- Temporal patterns: Supplementary plots examine happiness drops around the tenth anniversary of the 9/11 attacks and lexical-size variation by weekday and local time of day.The figures include simple-average happiness time series and Simpson lexical-size analyses.
- Robustness checks: Alternative distribution-construction approaches produce broadly similar Saturday-versus-Tuesday word-shift patterns, despite words moving within the overall pattern.The approaches differ in sampling-frequency treatment and, in one case, removal of outlier dates.
- Keyword rankings: The supplementary material includes a ranked selection of 100 keywords and text elements ordered by normalized average happiness.The table also reports frequency ranks within the top 5000 words for each specified corpus; “--” marks words absent from that set.