Source-linked AI summary
Algorithmic Amplification of Politics on Twitter
Ferenc Huszár, Sofia Ira Ktena, Conor O'Brien, Luca Belli, Andrew Schlaikjer, Moritz Hardt
TL;DR
The paper examines whether Twitter’s personalization algorithms amplify some political groups more than others, addressing a debated question with limited quantitative evidence. Using a large randomized experiment comparing personalized timelines with reverse-chronological control timelines, it finds higher amplification for mainstream right-wing politics and right-leaning U.S. news, but not for extreme ideologies. The authors caution that the study does not establish the precise causal mechanism behind these disparities.
Problem
The paper asks whether Twitter’s personalization algorithms amplify some political groups more than others, a question central to debates about political content consumption.
Method
The study audits Twitter’s recommender system using a long-running randomized experiment comparing personalized timelines with reverse-chronological control timelines across political and news content.
Results
Across seven countries, mainstream right-wing parties benefit at least as much and often more from personalization than left-wing counterparts; right-leaning U.S. media are also amplified, while extreme ideologies are not amplified more than mainstream voices.
Takeaways & Limitations
The findings support evidence-based scrutiny of how personalization algorithms shape the visibility of political content.
Takeaways & Limitations
The study does not provide an explanation for why the observed amplification disparities exist or identify the precise causal mechanism driving them.
Abstract
from arXiv · showhide
Content on Twitter's home timeline is selected and ordered by personalization algorithms. By consistently ranking certain content higher, these algorithms may amplify some messages while reducing the visibility of others. There's been intense public and scholarly debate about the possibility that some political groups benefit more from algorithmic amplification than others. We provide quantitative evidence from a long-running, massive-scale randomized experiment on the Twitter platform that committed a randomized control group including nearly 2M daily active accounts to a reverse-chronological content feed free of algorithmic personalization. We present two sets of findings. First, we studied Tweets by elected legislators from major political parties in 7 countries. Our results reveal a remarkably consistent trend: In 6 out of 7 countries studied, the mainstream political right enjoys higher algorithmic amplification than the mainstream political left. Consistent with this overall trend, our second set of findings studying the U.S. media landscape revealed that algorithmic amplification favours right-leaning news sources. We further looked at whether algorithms amplify far-left and far-right political groups more than moderate ones: contrary to prevailing public belief, we did not find evidence to support this hypothesis. We hope our findings will contribute to an evidence-based debate on the role personalization algorithms play in shaping political content consumption.
Experimental Setup
The experiment compared personalized timelines with a reverse-chronological control condition to measure algorithmic amplification. Its interpretation is constrained by interaction effects, changing treatment conditions, and the inability to estimate unbiased causal effects.
- Experimental design: The study introduced amplification measurement to quantify how different political groups benefit from algorithmic personalization.
- Experimental design: In 2016, Twitter randomly assigned 1% of global users to a nonpersonalized reverse-chronological timeline and sampled 4% of other accounts experiencing personalization as treatment users.Treatment users could still opt out of personalization.
- Limitations: Control users encountered content shared by treatment users, so the experiment did not satisfy the Stable Unit Treatment Value Assumption.This interaction prevents the control group from being fully isolated from treatment effects.
- Limitations: Because of these interactions, simple treatment-control comparisons likely underestimate the true causal effects of personalization.
- Limitations: The personalized treatment changed over time because Twitter used treatment-control differences to improve its ranking experience.
Measuring Amplification
The study defines reach as the number of users who encounter a set of Tweets within a specified audience and time window. Encountering requires sustained visibility of at least half of the Tweet’s interface element for 500 milliseconds.
- Reach is the total number of users in audience U who encounter a Tweet from set T during a specified time window.
- The audience can be restricted to a defined population, such as German Twitter users in the control group.
- A user counts as encountering a Tweet when 50% of its containing interface element remains continuously visible for 500ms.
Measuring Algorithmic Amplification
Algorithmic amplification compares Tweet reach among personalized-treatment users with reach among control users viewing reverse-chronological timelines. The normalized ratio interprets 0% as equal proportional reach and 50% as a 50% greater likelihood of encountering the Tweets under personalization.
- The amplification ratio is the normalized ratio of a Tweet set’s reach in treatment and control audiences.
- 0% amplification means treatment and control users have equal proportional reach for the Tweet set.
- 50% amplification means treatment users are 50% more likely than control users to encounter one of the Tweets.
- Higher amplification ratios indicate that the ranking model assigns greater relevance to Tweets, causing them to appear more often than in reverse-chronological order.
- The study reports both individual-account amplification and aggregate group amplification for sets of political accounts.Individual amplification concerns a single account, while group amplification aggregates Tweets authored by group members.
Results
Across seven countries, algorithmic personalization produced higher aggregate amplification for mainstream right-wing parties than mainstream left-wing parties in all but Germany, while individual amplification varied widely within parties. The analysis also found lower amplification for represented far-left and far-right parties than for moderate parties, and right-leaning patterns in U.S. news amplification, with results depending on the media-bias rating scheme.
- Legislator amplification: Group amplification exceeded 0% for all parties and sometimes exceeded 200%, meaning Tweets reached more than three times their chronological-timeline audience.A 0% value denotes no relative amplification; values above 200% indicate exposure to an audience more than three times as large.
- Legislator amplification: In 6 of 7 countries, pairwise comparisons found statistically significant higher amplification for mainstream right-wing than left-wing parties, except Germany.The strongest differences were Canada, Liberals 43% versus Conservatives 167%, and the U.K., Labour 112% versus Conservatives 176%.
- Individual variation: Individual amplification varied substantially within parties, reaching up to 400% for some politicians and falling below 0% for others.Despite this variation, individual amplification showed no statistically significant association with party affiliation.
- Ideological extremity: In countries with substantially represented extreme parties, far-left and far-right parties were generally amplified less than moderate or centrist parties.The analysis examined examples including VOX, Die Linke, AfD, LFI, and RN.
- News amplification: U.S. news amplification was generally higher for partisan sources than Center sources, with the partisan Right amplified marginally more than the partisan Left under AllSides ratings.Under Ad Fontes ratings, partisan Left amplification was relatively low at 10.5%, while differences among remaining categories were not substantial.
- News amplification: News-outlet amplification varied substantially within bias categories, and conclusions differed between the AllSides and Ad Fontes rating datasets.The rating schemes largely agreed on the political Right but differed most in their assessments of political-left publications.
Discussion
The paper reports a consistent rightward pattern in algorithmic amplification across political parties and U.S. news sources, while finding no support for greater amplification of extreme ideologies. It presents a large-scale audit but leaves the causal mechanism and other forms of platform curation for future work.
- Contribution: The study provides a systematic, large-scale contrast between ranked and chronological Twitter timelines for auditing political amplification.The authors describe it as the first systematic and large-scale study contrasting these timeline types on Twitter.
- Political amplification: Mainstream right-wing parties benefited at least as much, and often substantially more, from algorithmic personalization than left-wing parties across seven countries.This pattern was reported for aggregate party amplification; individual-account comparisons did not show the same association.
- Media amplification: U.S. media outlets with strong right-leaning bias were amplified marginally more than left-leaning sources.The discussion also links strong partisan bias in news reporting with the possibility of higher amplification.
- Ideological extremes: The analysis did not support the hypothesis that personalization amplifies far-left and far-right ideologies more than mainstream political voices.The paper distinguishes partisan bias in reporting from promotion of extreme political ideology.
- Open questions: The precise causal mechanism behind the observed amplification disparities remains unresolved and invites further study.The authors also identify other forms of algorithmic content curation beyond the Home timeline as avenues for similar experiments.
Supplementary Information for
The supplementary methods describe a long-running randomized timeline experiment comparing algorithmically personalized treatment feeds with reverse-chronological control feeds. They define daily, country-specific amplification measures and document additional personalization, ranking, and interference considerations.
- Assignment: Accounts were randomly assigned to treatment or control at experiment onset or account creation, with assignment maintained over the account lifespan.Treatment users could temporarily disable algorithmic recommendations but remained classified as treatment users.
- Experiment scale: The experiment covered 5% of global accounts, with 20% assigned to control and 80% to treatment; about 9.3 million of Twitter’s 186 million monetizable daily active users were included.The included population also contained dormant accounts and bots, only a fraction of which were active during the study period.
- Timeline conditions: Control users saw followed accounts’ Tweets and Retweets in reverse-chronological order, while treatment timelines used algorithmic selection and ranking.Treatment ranking incorporated engagement predictions, content and behavioral signals, heuristics, and other machine-learning outputs.
publications
The supplementary materials document how the political-content study was prioritized, expanded geographically, and supported with multiple media-bias datasets. They also describe internal review, privacy procedures, preregistration, and constrained data access.
- Hypothesis selection: The project shifted toward political-content research because higher-quality third-party data were available, with no other considerations influencing hypothesis selection.The authors initially investigated abusive Tweets and political content in parallel.
- Scope: The analysis expanded beyond the United States and United Kingdom to countries where legislators could be reliably identified and sufficient user data were available.The authors sought to reduce subjectivity in country selection and aimed to include all technically feasible countries.
- Media-bias data: The study used both AllSides and Ad Fontes Media bias datasets after concerns that relying on one source could make findings depend on its validity.The paper presents both result sets without a normative judgment about either underlying source.
- Governance: The internal review involved public-policy, investor-relations, intellectual-property, communications, and technical reviewers, while privacy reviews addressed data protection and retention.The materials state that privacy reviews did not alter research decisions and that communications review could not change interpretation or presentation of findings.
- Data access: Reproduction data were available upon request under the Twitter Developer Agreement, and further sharing was prohibited.The released files support reproducing the main figures, including bootstrap data for political-party amplification and media-bias analyses.