Source-linked AI summary

Quantifying Search Bias: Investigating Sources of Bias for Political Searches in Social Media

Juhi Kulshrestha, Motahhare Eslami, Johnnatan Messias, Muhammad Bilal Zafar, Saptarshi Ghosh, Krishna P. Gummadi, Karrie Karahalios

arXiv:1704.01347v1cs.SIcs.CYcs.HC

TL;DR

Political search results can reflect bias in both the data supplied to a ranking system and the ranking process itself. The paper develops a framework to distinguish and quantify these sources, applies it to political queries on Twitter, and finds that they contribute significantly in different ways.

  • Problem

    Search bias is difficult to characterize because output bias may arise from the input corpus, the ranking system, or both.

  • Method

    The paper develops a framework that quantifies bias in ranked search results and separates contributions from input data and ranking, applying it to Twitter political queries.

  • Results

    Both input data and the ranking system significantly contribute to Twitter search-result bias, with ranking sometimes shifting or altering the input bias’s polarity.

  • Takeaways & Limitations

    Search systems should signal result bias or incorporate bias considerations into ranking and interface design.

  • Takeaways & Limitations

    The study used a limited set of event- or candidate-related queries and simplified users’ political leanings into neutral, pro-democratic, or pro-republican categories.

Abstract

from arXiv · show

Search systems in online social media sites are frequently used to find information about ongoing events and people. For topics with multiple competing perspectives, such as political events or political candidates, bias in the top ranked results significantly shapes public opinion. However, bias does not emerge from an algorithm alone. It is important to distinguish between the bias that arises from the data that serves as the input to the ranking system and the bias that arises from the ranking system itself. In this paper, we propose a framework to quantify these distinct biases and apply this framework to politics-related queries on Twitter. We found that both the input data and the ranking system contribute significantly to produce varying amounts of bias in the search results and in different ways. We discuss the consequences of these biases and possible mechanisms to signal this bias in social media search systems' interfaces.

INTRODUCTION

The paper examines how political search results become biased by separating bias in the input data from bias introduced by ranking. It proposes a framework and applies it to Twitter political queries, finding that both sources materially shape results and user experience.

  • Search systems can influence opinions by ranking one perspective above others, especially for polarizing political topics.
  • The paper asks how search bias can be quantified and how much political search-result bias comes from input data versus ranking.
  • The framework measures output-result bias while distinguishing the contributions of the data set feeding the ranking system and the ranking system itself.
  • Using 25 political Twitter queries collected during a week containing two presidential debates, the study finds that both input data and ranking contribute significantly to output bias.
  • The input stream was democratically biased on average, while ranking shifted or sometimes reversed that polarity, producing substantially different final-result bias.
  • Ranking mitigated opposing bias for the most popular Democratic candidate but enhanced it for the most popular Republican candidate, affecting the user’s search experience.
  • The paper discusses incorporating bias into ranking systems or making result bias transparent through search-interface design.

Bias in Web Search

Prior work studied bias in web search, political content, personalization, and social-media polarization, but generally did not separate input-data bias from ranking-system bias. This paper frames that separation as central to auditing search systems.

  • Bias in Web Search: Earlier web-search studies examined preferences among sites, political leanings, ranking effects, personalization, and geographic differences.
  • Bias in Web Search: Prior Twitter studies often inferred bias from users’ language or social networks rather than from the textual content of individual tweets.
  • Bias in Web Search: The paper extends user-bias inference by incorporating interests, addressing cases where political leaning is not explicit in language or social connections.
  • Bias in Web Search: Research on social-media polarization documented ideologically partitioned retweet networks and selective exposure, while this study focuses on bias in social-media search.
  • Bias in Web Search: The study identifies limited knowledge about how much search bias is inherent in data and how much the search system enhances or mitigates it.
  • Bias in Web Search: Because biased inputs can produce biased outputs alongside algorithm design, separating these sources is important for auditing algorithmic systems.

RQ1: QUANTIFYING SEARCH ENGINE BIAS

The framework separates search bias into input, ranking, and output stages, while using inferred political bias scores for individual Twitter items and users. It evaluates these measures and shows that data composition, ranking, and inference coverage each shape the analysis.

  • Framework: Twitter search retrieves query-relevant tweets as input before ranking them into the displayed result list.For a query such as “Bernie Sanders,” matching tweets form the input to the ranking system.
  • Framework: The framework measures input bias, ranking bias, and output bias across the search process.Input bias concerns retrieved items, ranking bias concerns the ranking system, and output bias concerns the resulting ranked list.
  • Metrics: Output bias weights higher-ranked items more heavily because users attend to and trust top search results more.The proposed output-bias metric is inspired by Average Precision and accumulates bias over the ranked list.
  • Metrics: Time-averaged metrics summarize bias trends across multiple snapshots of search results.The framework defines time-averaged input, ranking, and output bias from measurements collected at different instants.
  • Bias inference: Inference coverage is constrained when accounts are protected, follow no one, or follow fewer than 10 users, although the method supports broad user coverage.For self-identified users, average coverage was 91.12%; inferred scores were usually within [−0.5, +0.5], unlike many AMT scores near the boundaries.

DIA SEARCH

The study collected Twitter search and stream data for political-debate queries to examine how input data and ranking interact in producing search bias.

  • Analysis: The study analyzed collected data to identify sources of Twitter Search bias and the interplay between input data and ranking.
  • Query Selection: The query set was expanded using popular, party-neutral hashtags and filtered to remove queries with apparent partisan leaning.The collection process identified 57 Democratic-debate hashtags and 63 Republican-debate hashtags before selecting the final set.
  • Data Collection: Researchers collected the top 20 non-personalized Twitter search results at 10-minute intervals throughout the study period.Queries were made while logged out to mitigate personalization effects.
  • Data Collection: The search-result dataset comprised 28,800 snapshots, 34,904 distinct tweets, and 17,624 distinct users.
  • Data Collection: The input dataset contained more than 8.2 million tweets posted by 1.88 million distinct users across the selected queries.

RQ2a: Where Does the Bias Come from?

Twitter search bias arose from both the query-filtered input data and the ranking system, with ranking sometimes amplifying or reversing the input data’s political leaning.

  • Input Data: Most selected candidate and debate queries had a democratic-leaning input bias, although the magnitude varied by query.Different queries filter Twitter users with differing political biases, despite the broader corpus having a democratic-leaning bias.
  • Input Data: For Bernie Sanders, the output bias was very democratic at 0.71, with most of that bias originating in the input data rather than Twitter’s ranking system.
  • Ranking System: The ranking system shifted candidate-query bias toward the corresponding party’s leaning on average.
  • Ranking System: For Republican candidate queries, ranking increased Republican bias by 0.18, producing a Republican output bias of −0.11 despite democratic-leaning average input bias.
  • Ranking System: For Chris Christie, Jeb Bush, and Lindsey Graham, ranking changed positive input bias into Republican-leaning output results.
  • Ranking Factors: Twitter’s ranking and rankings based on retweets or favorites showed similar ranking biases, suggesting post popularity explains much of Twitter’s observed ranking bias.Differences for some queries indicate that factors beyond popularity also contribute.

the Ranking System

The ranking system can substantially reshape political bias in Twitter search results, sometimes changing its polarity. Its effects also vary across similarly phrased queries and popular candidates, altering users’ search experience.

  • the Ranking System: More popular candidates’ top results showed greater bias toward the opposing perspective than results for less popular candidates with the same political leaning.For Republican candidates, a negative best-fit slope supported this popularity–opposing-bias relationship.
  • the Ranking System: Tweets from users opposing Hillary Clinton or Donald Trump were included among the top results and criticized or ridiculed the candidates.These results illustrate a potential search-experience consequence for unbiased or undecided voters.
  • the Ranking System: For Hillary Clinton, ranking increased democratic output bias sevenfold relative to input bias, while Donald Trump showed the opposite pattern.The opposing dynamics caused results for searches of the most popular candidates to favor Hillary Clinton over Donald Trump.
  • the Ranking System: Similar queries for the same event can produce noticeably different output biases: republican debate had TOB = 0.53 versus TOB = 0.31 for rep debate.The passages also report differences among queries for the democratic debate.
  • the Ranking System: For republican debate and rep debate, similar input biases were altered in opposing directions by ranking, producing different output biases.Ranking increased republican debate bias by 0.26 toward democratic and decreased rep debate bias by 0.09 toward republican.

Generalizability of the Bias Quantification Framework

The framework can extend beyond Twitter and beyond two political perspectives, but its ability to separate input from ranking bias depends on the search system’s architecture.

  • Generalizability of the Bias Quantification Framework: The framework can quantify search-result bias in web and other social-media search engines without knowing proprietary retrieval or ranking internals.It requires a methodology for measuring the bias of individual data items.
  • Generalizability of the Bias Quantification Framework: In information-retrieval systems that directly rank corpus items without an intermediate relevant set, input and ranking bias are difficult to disentangle.Relative biases of systems operating on similar corpora can still be compared.
  • Generalizability of the Bias Quantification Framework: A multiple-perspective extension would assign each item a bias vector rather than a scalar score and use vector formulations for input, output, and ranking bias.The paper identifies measuring these bias vectors as a challenging aspect for future work.

Signaling Political Bias in Search Results

The paper discusses ranking, interface, and hybrid approaches for addressing political bias in search results, while leaving tool development and user evaluation for future work.

  • Signaling Political Bias in Search Results: Ranking systems could treat bias as a metric and trade it off against relevance, but the appropriate balance may depend on the domain and user.Adding bias to ranking may degrade relevance, popularity, recency, or other metrics.
  • Signaling Political Bias in Search Results: A front-end approach could visualize each result’s bias, making users aware of potential bias without changing ranking efficiency.This approach addresses bias through the search interface rather than the ranking algorithm.
  • Signaling Political Bias in Search Results: A hybrid interface could display separate ranked lists for republican and democratic perspectives while preserving the original ranking within each list.Preserving within-list ranking is intended to avoid degrading relevance, popularity, or recency.
  • Signaling Political Bias in Search Results: Developing political-bias signaling tools and studying how users interact with alternative search interfaces remain future work.The paper states that its proposed solutions have not yet been explored in depth.

Auditing Black Boxes

The paper audits Twitter’s ranking behavior as a black box, showing how its framework can reveal bias without internal algorithmic access and support different stakeholders.

  • Auditing Black Boxes: Proprietary algorithms’ complexity, intellectual-property barriers, and susceptibility to gaming can make their internal operations effectively inaccessible.This motivates auditing algorithmic platforms through their observable behavior.
  • Auditing Black Boxes: The framework characterizes Twitter search-ranking bias without requiring knowledge of the platform’s internal mechanisms.It builds on prior studies that audit algorithmic systems from a black-box perspective.
  • Auditing Black Boxes: Audit results can help users recognize non-neutral search results, designers investigate algorithm-introduced bias, and researchers compare bias across platforms.The stated uses span user awareness, system investigation, and cross-platform measurement.

Distinguishing the Sources of Bias: From Development

The study separates bias originating in input data from bias introduced by ranking, showing that both can substantially shape search results. Ranking can even reverse the input data’s inherent bias, making interface visibility consequential.

  • Input data can contribute a significant part of output bias, so algorithm audits should investigate both input and output data.
  • Ranking algorithms can significantly alter input bias, including changing its polarity.
  • Twitter’s default top results receive more visibility than live results, making ranking-induced bias more prominent to users.Top results are ranked outputs, whereas live results present matching tweets in reverse chronological order.

LIMITATIONS

The study’s conclusions are constrained by its limited query set and simplified political-leaning model. It also leaves proposed bias-signaling mechanisms unevaluated.

  • The study covers a limited set of queries about political events or candidates because of Twitter API submission limits.The authors suggest broader controversial topics, including gun control and abortion, for future analysis.
  • Users are classified as neutral, pro-democratic, or pro-republican, which cannot represent partial affiliation with both political groups.The study retains separate similarity scores that could support more nuanced political-leaning estimates.
  • The proposed mechanisms for signaling political bias in search results were not implemented or evaluated for their effect on users’ search experience.
Loading 1704.01347v1…