Source-linked AI summary

Human language reveals a universal positivity bias

Peter Sheridan Dodds, Eric M. Clark, Suma Desu, Morgan R. Frank, Andrew J. Reagan, Jake Ryland Williams, Lewis Mitchell, Kameron Decker Harris, Isabel M. Kloumann, James P. Bagrow, Karine Megerdoomian, Matthew T. McMahon, Brian F. Tivnan, Christopher M. Danforth

arXiv:1406.3855v1physics.soc-phcs.CLcs.SI

TL;DR

The paper asks whether emotional patterns in language are universal across languages and usage frequencies. It combines multilingual word evaluations with corpus and word-shift analyses, finding positivity, cross-language agreement, and frequency independence, while noting a limitation of near-equal happiness comparisons.

  • Problem

    The study addresses how to compare emotional word content across corpora and languages when corpora cannot be merged using a single principled weighting scheme.

  • Method

    The authors combine language-specific word-happiness evaluations with corpus ranking, text preprocessing, translation-stable word comparisons, and word-shift analyses.

  • Results

    The analyses report positive shifts associated with increased use of relatively positive words and decreased use of relatively negative words in the Count of Monte Cristo comparison.

  • Takeaways & Limitations

    The word evaluations support adaptable instruments for measuring emotional content in large-scale texts through summaries such as word shifts.

  • Takeaways & Limitations

    Word shifts show approximately the same result when two texts have equal happiness to two decimal places, although their word usage may still differ.

Abstract

from arXiv · show

Using human evaluation of 100,000 words spread across 24 corpora in 10 languages diverse in origin and culture, we present evidence of a deep imprint of human sociality in language, observing that (1) the words of natural human language possess a universal positivity bias; (2) the estimated emotional content of words is consistent between languages under translation; and (3) this positivity bias is strongly independent of frequency of word usage. Alongside these general regularities, we describe inter-language variations in the emotional spectrum of languages which allow us to rank corpora. We also show how our word evaluations can be used to construct physical-like instruments for both real-time and offline measurement of the emotional content of large-scale texts.

Online, interactive visualizations:

The paper provides online resources for measuring and visualizing emotional content in texts, including interactive examples and time series.

  • Online scripts support parsing texts and measuring their average happiness scores.
  • D3 and Matlab scripts generate word shifts for examining changes in word usage and emotional content.
  • Interactive visualizations enable exploration of translation-stable word pairs across languages.
  • Interactive time series are provided for Moby Dick, Crime and Punishment, the Count of Monte Cristo, and other literary works.

Corpora

The study assembled corpora across multiple languages and obtained word-happiness evaluations through language-specific participant surveys using a common rating task.

  • Measuring the Happiness of Words: Participants rated 100 individual words on a 9-point unhappy-happy scale.
  • Measuring the Happiness of Words: The words were selected based on common usage, although some could be offensive, foreign-language, or nonsensical to participants.
  • Corpora: The survey also collected demographic information on gender, age, education, income, origin, current residence, first language, and comments.
  • Corpora: The sizes and sources of all 24 corpora are reported in Table S1.
  • Corpora: Word evaluations covered four English corpora using Mechanical Turk and non-English assessments using Appen Butler Hill.
  • Corpora: Non-English participants were native speakers who grew up where their language is spoken and passed an oral comprehension or proficiency test.

Notes on corpus generation

Corpus generation required a pragmatic cross-corpus ranking and text cleaning procedure, alongside supplementary analyses of translations, emotional distributions, and frequency relationships.

  • Notes on corpus generation: Because no principled weighting exists for merging corpora, the authors chose a method enabling cross-language comparisons and adaptable linguistic instruments.
  • Notes on corpus generation: For each language, a quasi-ranked word list used the smallest rank cutoff whose cross-corpus union contained at least 10,000 words.
  • Notes on corpus generation: Twitter preprocessing removed invalid, invisible, HTML-like, user-address, website, and standard punctuation content before lowercasing Latin letters.
  • Translation analyses: Translation analyses include histograms of happiness changes for translation-stable words and reproduced figures using direct English translation.
  • Frequency analyses: Frequency analyses fit happiness and happiness standard deviation as linear functions of word rank for the most common 5000 words, with stemming noted as a possible influence.
  • Supplementary analyses: Supplementary figures examine happiness distributions and standard deviations by word rank, including English translations across corpora ranked by median happiness.

EXPLANATION OF WORD SHIFTS

Word shifts explain changes in a text’s average happiness by combining word-level emotional scores with changes in normalized word usage relative to a reference text.

  • EXPLANATION OF WORD SHIFTS: Word shifts rank and visualize how individual words contribute to or oppose an overall happiness change between comparison and reference texts.Words are ordered by the absolute value of their contribution and normalized as percentages.
  • EXPLANATION OF WORD SHIFTS: The analysis compares normalized word frequencies after applying a word lens and estimates each text’s happiness from surveyed word scores.The example lens retains words with scores in [1, 3] or, excluding 3 < havg < 7.
  • EXPLANATION OF WORD SHIFTS: Each word’s contribution depends on both its emotional difference from the reference average and its usage difference between the two texts.The framework distinguishes changes in word happiness from changes in normalized word prevalence.
  • EXPLANATION OF WORD SHIFTS: Four categories encode whether words are relatively positive or negative and whether they are used more or less in the comparison text.The categories are +↑, −↓, +↓, and −↑, with yellow indicating relatively positive words and blue indicating negative ones.

Simple Word Shifts

Simple word shifts summarize the largest word-level contributors to an overall happiness change and show how positive and negative word usage combine to produce it.

  • Simple Word Shifts: The simple inset word shifts display the 10 words with the largest absolute contributions to the overall happiness shift.Contributions are ranked by absolute magnitude.
  • Simple Word Shifts: In the Count of Monte Cristo example, increased use of relatively positive words and reduced use of relatively negative words most strongly raise positivity.‘excellence’, ‘mer’, and ‘rêve’ increase, while ‘prison’ and ‘prisonnier’ decrease.
  • Simple Word Shifts: The summary bars total contributions from the four word categories, with relatively negative words used more providing the smallest contribution.Σ+↑ denotes the total shift from relatively positive words that are more prevalent in the comparison text.
  • Simple Word Shifts: The overall summary separates positive and negative word contributions and shows that the example combines more relatively positive words with fewer relatively negative ones.The Count of Monte Cristo example has an overall positive shift.

Full Word Shifts

Full word shifts expand the abbreviated visualization with summaries, distributions, lenses, emotional-distance weighting, and cumulative contribution plots for comparing texts.

  • Full Word Shifts: Each full word shift begins with a summary describing the reference and comparison texts, their average happiness scores, and which text is happier.Average happiness is marked with filled and unfilled diamonds for the reference and comparison texts.
  • Full Word Shifts: When two texts have equal happiness to two decimal places, the shift appears approximately equal even though their word usage may differ.The visualization can remain informative because large-scale texts will most likely use words differently.
  • Full Word Shifts: The figures plot the first 50 words by contribution rank alongside histograms showing how the two texts’ word distributions generate the overall shift.Reference data appear on the left and comparison data on the right.
  • Full Word Shifts: Plot B shows bare frequency distributions, while Plot C applies the lens, renormalizes frequencies, and colors words by relative positivity or negativity.At this stage, relatively positive words dominate in pure counts in the described example.
  • Full Word Shifts: Plot D weights words by emotional distance from the reference average, and Plot E incorporates usage differences to summarize the resulting happiness shift.In the example, the comparison text is generally happier across the negativity-positivity scale.
  • Full Word Shifts: The cumulative plots separately and jointly sum the four contribution categories, with the central black line showing the overall shift.In the described positive example, all contributions sum to +100%.
  • Full Word Shifts: Supplementary figures provide full word shifts corresponding to the simple shifts shown in the main and supplementary figures.The detailed examples include Moby Dick, Crime and Punishment, and the Count of Monte Cristo, including English translations for several comparisons.
Loading 1406.3855v1…