Source-linked AI summary
Engagement, User Satisfaction, and the Amplification of Divisive Content on Social Media
Smitha Milli, Micah Carroll, Yike Wang, Sashrika Pandey, Sebastian Zhao, Anca D. Dragan
TL;DR
The paper asks whether engagement-based ranking reflects users’ deliberate preferences or instead amplifies content driven by automatic biases. A preregistered Twitter audit compares engagement-based, reverse-chronological, and stated-preference timelines, finding more emotional, partisan, and out-group-hostile content under engagement ranking and lower stated value for political recommendations. Stated-preference ranking reduces negativity and hostility but may reinforce in-group bias, while a modified tie-breaking approach reduces out-group hostility without heightened in-group exposure.
Problem
The study addresses whether engagement-based ranking reflects users’ deliberate preferences or exploits automatic biases, a distinction important for understanding and designing ranking systems.
Method
A preregistered Twitter audit compared engagement-based and reverse-chronological timelines and constructed an exploratory timeline by ranking collected tweets according to users’ stated preferences.
Results
Engagement-based ranking amplified emotional, partisan, and out-group-hostile content, users rated its political recommendations lower, and stated-preference ranking reduced negativity and hostility but concentrated content more within users’ political in-groups.
Takeaways & Limitations
A modified stated-preference timeline that uses out-group animosity for tie-breaking reduced out-group-hostile amplification without increasing exposure to in-group content.
Takeaways & Limitations
The study compares timelines at one point in time and therefore does not capture long-term effects on content production and availability.
Abstract
from arXiv · showhide
In a pre-registered algorithmic audit, we found that, relative to a reverse-chronological baseline, Twitter's engagement-based ranking algorithm amplifies emotionally charged, out-group hostile content that users say makes them feel worse about their political out-group. Furthermore, we find that users do \emph{not} prefer the political tweets selected by the algorithm, suggesting that the engagement-based algorithm underperforms in satisfying users' stated preferences. Finally, we explore the implications of an alternative approach that ranks content based on users' stated preferences and find a reduction in angry, partisan, and out-group hostile content, but also a potential reinforcement of pro-attitudinal content. The evidence underscores the necessity for a more nuanced approach to content ranking that balances engagement and users' stated preferences.
1 Introduction
Engagement-based ranking predicts users’ engagement as a proxy for what they want, but may prioritize attention-grabbing content over reflective preferences. This study separates algorithmic selection from following choices by comparing timelines and surveying users’ stated preferences.
- Engagement-based ranking: Engagement-based ranking algorithms predict behaviors such as retweets, replies, video watching, or lingering to personalize users’ feeds.
- Motivation: These algorithms may amplify emotionally charged, in-group, moral, or attention-grabbing content that conflicts with users’ reflective values.
- Research gap: Existing studies often fail to disentangle deliberate choices from automatic biases because users’ following decisions shape the content available to ranking algorithms.
- Study design: The audit compared engagement-based and reverse-chronological timelines, then surveyed users to construct an exploratory timeline ranked by stated preferences.
- Hypotheses: The preregistered hypotheses predicted that engagement-based ranking would increase emotionality, ideological leaning, animosity, and negative perceptions of users’ political out-groups.
2 Results
The audit found that engagement-based ranking amplified emotional, partisan, and out-group-hostile content relative to reverse chronology, while political tweets selected by the algorithm received lower stated-value ratings. Ranking by stated preferences reduced negativity and partisanship but also reduced exposure to users’ political out-groups.
- Study procedure: The survey measured tweet politics, ideological leaning, perceived animosity, author emotions, reader emotions, and users’ stated value for each tweet.
- Political outcomes: 0.24 SD more partisan and 0.24 SD more out-group-hostile content appeared in engagement-based than reverse-chronological timelines.
- User preferences: -0.18 SD lower stated value characterized political tweets selected by the engagement-based algorithm than political tweets in the reverse-chronological timeline.
- Stated-preference ranking: Ranking by stated preferences reduced negativity, partisanship, and out-group animosity relative to engagement ranking, but mainly by reducing out-group content and in-group-directed animosity.
3 Discussion
The discussion finds that engagement-based ranking amplifies emotionally charged, partisan, and out-group hostile content while poorly matching users’ stated political preferences. Ranking by stated preferences may reduce divisive content but could reinforce in-group bias, and the study’s short-term, limited-pool design constrains interpretation.
- User satisfaction: Users were less likely to prefer political tweets selected by the engagement-based algorithm, despite its increasing platform time relative to reverse chronology.This suggests a divergence between engagement-based revealed preferences and users’ stated preferences.
- Content amplification: The engagement-based algorithm selected more emotionally charged, partisan, and out-group hostile tweets than both comparison timelines.The result held beyond users’ following choices and beyond tweet-level stated preferences.
- Polarization: Engagement-based ranking produced more positive perceptions of users’ political in-groups and more negative perceptions of their out-groups than reverse-chronological ranking.The authors interpret this as evidence that engagement-based selection may be more polarizing than following choices alone would predict.
- Limitations: The study cannot establish whether ranking algorithms have lasting effects on attitudes or affective polarization.The audit measured content-specific and short-term responses and did not capture longer-term feedback loops affecting content production and availability.
- Stated-preference ranking: Ranking by stated preferences reduced partisan content and out-group animosity, but these reductions primarily reflected less content from users’ political out-groups.The stated-preference timeline did not significantly differ from the engagement-based timeline in effects on perceptions of political in-groups and out-groups.
- Limitations: The stated-preference results were based on a pool of approximately twenty tweets, limiting conclusions about ranking over the platform’s larger candidate pool.The authors note that respectful political out-group content may have been absent from the limited pool.
Funding.
The study was financially supported by the UC Berkeley Center for Human-Compatible AI and the National Science Foundation Graduate Research Fellowship Program.
- The UC Berkeley Center for Human-Compatible AI provided financial support for the study.
- MC received support from the National Science Foundation Graduate Research Fellowship Program.
- The authors declared no competing interests, and the data and code are publicly available for reproduction.
S1 Materials and methods
The study recruited Twitter users through CloudResearch Connect, collected engagement-based and reverse-chronological timelines using a Chrome extension, and surveyed participants about the displayed tweets and their reactions.
- Participants and recruitment: Participants were U.S. adults who used Twitter at least a few times weekly, followed at least 50 people, and used Google Chrome.Recruitment occurred across four waves from February 11 to February 27, 2023.
- Timeline collection: Participants installed a Chrome extension that collected the first ten engagement-based tweets and ten recent tweets from followed accounts.The reverse-chronological timeline served as the comparison condition.
- Survey procedure: Tweets were shown in randomized order, with overlapping tweets displayed only once and embedded for reference.Replies and quote tweets included both the referenced and main tweet when applicable.
- Survey measures: Participants rated each timeline tweet on political content, ideological leaning, political affect, out-group animosity, and whether they wanted similar tweets.Additional measures assessed author and reader emotions.
- Exclusions: Eighteen percent of consenting participants did not complete Chrome-extension data collection, and failed attention checks or incomplete surveys led to exclusion.Technical problems, conflicting extensions, or Twitter interface experiments could prevent data collection.
- Preregistration deviations: The study stopped near 1,700 timeline pairs rather than the planned 2,000 because of budget constraints.The recruitment platform and repeated participation across waves also differed from the preregistered plan.
- Estimation and inference: Treatment effects were estimated as paired differences in means and tested with two-tailed paired permutation tests.Bootstrap confidence intervals and Benjamini-Krieger-Yekutieli false-discovery-rate adjustments were used across 26 outcomes.
S2 User-level survey questions and demographics
The study surveyed users about their Twitter use and the content shown to them, while comparing participant demographics with a 2020 ANES Twitter population. The study population was younger and more Democratic-leaning than the comparison population.
- Demographics: 53 percent of study participants were aged 18–34, compared with 33 percent in the ANES Twitter population.
- Demographics: 56 percent of study participants identified as Democrats, compared with 43 percent in the ANES Twitter population.
- User-level survey questions: 806 unique users participated 1,730 times across study waves, with demographic information collected at each wave.Users also reported their primary reason for using Twitter and the primary type of content they saw.
- Survey design: Participants answered follow-up questions about ideological leaning or party leaning when they selected moderate, other, independent, or something else.
- Demographics: The study and ANES Twitter populations differed significantly on race and ethnicity, party affiliation, education, and age distributions.The sex/gender distributions did not significantly differ in the reported comparison.
S3 Pre-registered analysis
The pre-registered analysis evaluated Twitter’s engagement-based timeline relative to the reverse-chronological timeline using average treatment effects across 26 outcomes.
- Pre-registered analysis: The analysis reports standardized and unstandardized average treatment effects, p-values, and FDR-adjusted p-values for all 26 pre-registered outcomes.
S4 Exploratory analysis
Exploratory analyses compared the metadata, accounts, emotions, political content, and stated preferences in engagement-based and chronological timelines. The engagement timeline contained more highly liked, retweeted, and angry political content, while political tweets were more often unwanted.
- Timeline metadata: Engagement-based timelines contained tweets with much higher numbers of likes and retweets than chronological timelines.Authors in engagement-based timelines had fewer followers on average, and the average number of links was almost halved.
- Account amplification: News outlets were more dominant in chronological timelines and were among the accounts most de-amplified by the engagement-based algorithm.The authors attribute this pattern partly to news outlets posting more frequently than ordinary accounts.
- Account amplification: Elon Musk received higher amplification than any other account by a wide margin during the study period.
- Emotions: 62 percent of political tweets in engagement timelines expressed anger, compared with 52 percent in chronological timelines.The analysis also examined sad, happy, and anxious emotions for authors and readers across overall and political tweets.
- Stated preference: 22 percent of political tweets in engagement timelines were unwanted, compared with 16 percent in chronological timelines.Across all tweets, both timelines contained 13 percent unwanted tweets.
S4.4 Outcomes by tweet rank
Outcomes did not significantly differ between high- and low-ranked tweets within the examined top ten positions. Robustness analyses varied tweet thresholds, and GPT-4-based labels produced results similar to reader judgments.
- Outcomes by tweet rank: No statistically significant differences appeared between high-ranked and low-ranked tweets for emotions, explicit preferences, or political outcomes.The comparison covered author and reader emotions, stated preference, partisanship, out-group animosity, and in- and out-group perceptions.
- Outcomes by tweet rank: The rank analysis considered only the first 10 tweets, so larger rank differences could appear outside the examined portion of each timeline.
- Robustness checks: Average treatment effects were evaluated across thresholds spanning 5 to 10 tweets, and the findings were consistently robust across these thresholds.
- GPT-4 validation: GPT-4 judgments produced results notably similar to reader judgments, including amplification of emotional, partisan, and out-group-hostile content.GPT-4 judged political tweets selected by the engagement algorithm to be even angrier than human readers did.
- GPT-4 validation: GPT-4 returned validly formatted responses for 24,998 of 28,301 unique tweets, and both human- and GPT-4-based ATEs used that same filtered set.
S4.7 Heterogeneous effects
The analysis examines whether timeline effects vary across demographic groups and users’ survey-reported Twitter use and content categories. It reports conditional average treatment effects across emotional, political, perception, and preference outcomes.
- Subgroup analyses condition effects on political leaning, political party, age, ethnicity, education, household income, and users’ primary reason for using Twitter.
- The analysis also conditions effects on users’ primary category of content seen during the study.
- Reported outcomes include author and reader emotions, with separate overall and political measures for anger, sadness, anxiety, and happiness.
- Figures report conditional average treatment effects with the average treatment effect shown as a blue reference line.
- Additional outcomes cover partisanship, out-group animosity, in-group and out-group perception, and reader preference in overall and political contexts.
S4.8 Effects of stated preference timeline
This section presents the effects of the exploratory stated-preference timeline. The results are reported as average treatment effects in standardized and unstandardized forms, with associated p-values.
- Table 37 reports the stated-preference timeline’s average treatment effects, standardized and unstandardized estimates, and p-values for all outcomes.
- Table 38 reports the corresponding average treatment effects and p-values for the SP-OA timeline.
S4.9 Effects of SP-OA timeline
The SP-OA timeline combines users’ stated preferences with a tie-break rule that down-ranks tweets containing out-group animosity. It reduced angry, partisan, and out-group hostile content while retaining high stated-preference satisfaction.
- Motivation and design: The SP timeline reduced partisan animosity mainly by reducing animosity toward the reader’s in-group, while showing more in-group and less out-group content than comparison timelines.
- Motivation and design: The SP-OA timeline scores tweets by stated preference and subtracts 0.5 points when a tweet contains out-group animosity.
- Results: The SP-OA timeline produced the lowest levels of angry, partisan, and out-group hostile content across the chronological, engagement, and SP timelines.
- Results: 17 percent of political tweets in the SP-OA timeline contained animosity toward users’ out-group, compared with 34 percent in the engagement timeline and 33 percent in the SP timeline.
- Conclusion: The authors conclude that SP-OA maintains high stated-preference satisfaction, mitigates divisive-content amplification, and avoids reinforcing in-group bias.
S5 Survey questionnaires
The survey collected public tweets through a Chrome extension and then asked participants about those tweets, their emotions, political content, and Twitter use. Eligibility and participation procedures were also administered through the online study.
- Eligibility: Eligibility questions covered United States location, age of at least 18, Twitter-use frequency, and the number of people followed.
- Data collection: Participants were instructed to install a Chrome extension that collected public tweets from their Twitter timeline and automatically uninstalled afterward.
- Study administration: The study was conducted online, took about 30 minutes, and recruited participants through CloudResearch Connect.
- Tweet measures: The questionnaire asked how tweets made participants feel, including emotional intensity and feelings toward political groups on the Left and Right.
- Tweet measures: Participants classified tweets by whether they concerned political or social issues and by their political leaning.