Source-linked AI summary

Everyday algorithm auditing: Understanding the power of everyday users in surfacing harmful algorithmic behaviors

Hong Shen, Alicia DeVos, Motahhare Eslami, Kenneth Holstein

arXiv:2105.02980v2cs.HCcs.CY

TL;DR

Formal audits can miss harmful algorithmic behaviors that emerge through everyday use, while users sometimes detect and investigate these issues collaboratively. This paper proposes everyday algorithm auditing, analyzes real-world cases, and argues that such user-driven processes can surface problems that centrally organized audits may miss.

  • Problem

    Existing formal audits often require centralized organization and technical expertise, yet may fail to detect harmful behaviors that emerge in unanticipated contexts during everyday use.

  • Method

    The paper proposes and explores everyday algorithm auditing through theorizing and analyzing several real-world cases of users detecting, interpreting, questioning, and publicizing problematic machine behaviors.

  • Results

    Everyday users collectively surfaced problematic behaviors in cases including Twitter image cropping, Yelp review filtering, and YouTube demonetization.

  • Takeaways & Limitations

    Everyday algorithm auditing can help reveal harmful machine behaviors that are challenging for existing auditing approaches and suggests supporting these user-driven processes.

  • Takeaways & Limitations

    Some harmful algorithmic behaviors may be difficult or infeasible to predict or simulate artificially because they emerge from unanticipated use contexts and changing norms over time.

Abstract

from arXiv · show

A growing body of literature has proposed formal approaches to audit algorithmic systems for biased and harmful behaviors. While formal auditing approaches have been greatly impactful, they often suffer major blindspots, with critical issues surfacing only in the context of everyday use once systems are deployed. Recent years have seen many cases in which everyday users of algorithmic systems detect and raise awareness about harmful behaviors that they encounter in the course of their everyday interactions with these systems. However, to date little academic attention has been granted to these bottom-up, user-driven auditing processes. In this paper, we propose and explore the concept of everyday algorithm auditing, a process in which users detect, understand, and interrogate problematic machine behaviors via their day-to-day interactions with algorithmic systems. We argue that everyday users are powerful in surfacing problematic machine behaviors that may elude detection via more centrally-organized forms of auditing, regardless of users' knowledge about the underlying algorithms. We analyze several real-world cases of everyday algorithm auditing, drawing lessons from these cases for the design of future platforms and tools that facilitate such auditing behaviors. Finally, we discuss work that lies ahead, toward bridging the gaps between formal auditing approaches and the organic auditing behaviors that emerge in everyday use of algorithmic systems.

1 INTRODUCTION

Formal audits can miss harmful behaviors that emerge during everyday use, while everyday users can detect, investigate, and publicize such problems. The paper introduces everyday algorithm auditing and develops a process-oriented account with design implications for supporting it.

  • Motivation: Formal algorithm audits have contributed to detecting harmful behavior but can miss issues that emerge only through everyday use.Users may interact with systems in unanticipated ways or recognize harms that formal auditors did not anticipate.
  • Motivation: Everyday users have exposed harmful algorithmic behaviors, including racial bias in Twitter’s image-cropping algorithm despite prior company testing.Users collectively investigated suspected bias and built evidence through online discussions.
  • Concept: The paper defines everyday algorithm auditing as users detecting, interrogating, and understanding problematic machine behaviors through daily interactions with algorithmic systems.The concept draws on everyday resistance to characterize users’ active questioning of algorithmic systems.
  • Approach: The authors analyze Twitter’s cropping algorithm and online rating algorithms as two types of platforms targeted by everyday audits.The cases are used to examine the nature and dynamics of user-driven auditing.
  • Process model: Comparing cases yields a process model spanning audit initiation, awareness raising, hypothesis formation and testing, and possible remediation.Some audits may end before reaching every phase.
  • Design implications: The paper proposes design interventions involving community, expert, algorithmic, and organizational guidance, plus incentivization.The authors also discuss trade-offs among these intervention categories.

2 ALGORITHM AUDITING

Existing algorithm audits often rely on technical expertise or centralized organization and may fail to detect harms that appear only in situated use or require particular lived experiences. Everyday audits address this understudied gap through more organic, autonomous collective action.

  • Existing approaches: Existing audits include code, user, scraping, sock puppet, and crowdsourced or collaborative approaches conducted by internal teams or external experts.These approaches assess algorithmic systems against legal, ethical, societal, or industry standards.
  • Limitations: Most formal audits require technically knowledgeable people to initiate, conduct, or direct the process, including centralized design in crowdsourced audits.Crowdworkers typically complete subtasks assigned by researchers, who interpret the combined outputs.
  • Limitations: Harmful behaviors may arise only within particular social or cultural contexts, or through unanticipated uses and changing practices that artificial audits cannot predict or simulate well.The Tay chatbot illustrates harmful behavior emerging after users interacted with it in unforeseen ways.
  • Limitations: Auditors may miss harms when they lack the cultural backgrounds or lived experiences needed to recognize them or know where to look.Crowdworkers may not represent an algorithmic system’s user demographics.
  • Research gap: Everyday algorithm auditing remains understudied despite many recent user-driven efforts, motivating the paper’s effort to connect organic audits with academic auditing research.The paper aims to draw lessons for overcoming limitations in existing approaches.
  • Everyday audits: Unlike centrally managed crowdsourced audits, everyday audits are collectively organized and organic, with users retaining autonomy over the audit’s direction.Users’ autonomy and agency are central to the distinction.

3 EVERYDAY ALGORITHM AUDITING

Everyday algorithm auditing is a broad framework for users’ situated detection, interpretation, questioning, and communication of harmful machine behavior. Cases show that audits vary in expertise and collectiveness and can develop into collective resistance and counterpublics.

  • Definition: Everyday algorithm auditing captures users’ situated efforts to detect, understand, interrogate, or publicize socially harmful machine behavior.The framework includes formal or informal testing and investigation arising through routine interactions.
  • Examples: Examples span search, advertising, rating, cropping, translation, recommendation, credit scoring, captioning, and image-recognition systems.Reported cases include Google Search, Yelp, Booking.com, Twitter, Google Translate, YouTube, TikTok, Apple Card, and Google Photos.
  • Research gap: The paper reports many recent everyday audits but notes that these bottom-up, user-driven behaviors have received little academic attention.It aims to connect such practices with formal auditing research and identify lessons for addressing existing limitations.
  • Algorithmic expertise: Everyday audits do not require users to lack technical expertise; both technical experts and ordinary users can act as situated auditors.Examples include Sweeney, Noble, and Buolamwini, whose audits began from everyday encounters with algorithmic outputs.
  • Everyday algorithmic resistance: The authors theorize everyday algorithm auditing as everyday algorithmic resistance, in which users test algorithmic limits through routine interactions.Such resistance may be incidental and incremental or develop into coordinated collective action.
  • Counterpublics: Collective audits can form counterpublics where users share findings, test examples, build consensus, and determine the audit’s direction through discussion.The YouTube demonetization case illustrates collective sensemaking and interdependent probing of possible system mechanisms.

4 METHODS

The study uses exploratory case analysis to characterize everyday algorithm auditing across multiple domains, then examines four cases in depth. It compares image-cropping and rating-platform audits across varying expertise, collectiveness, and organization.

  • Research questions: The authors ask how everyday auditing practices develop and how platforms and tools can support users in detecting, reporting, and theorizing about harmful behavior.The questions address characteristics, dynamics, progression, and support for these practices.
  • Case discovery: An exploratory case study approach began with familiar high-profile cases, iterative discussion, keyword development, and searches of news and social media.The process established an initial scope for the emerging phenomenon.
  • Case set: The search identified 15 cases across image captioning, image search, translation, rating, cropping, credit scoring, advertising, and recommendation systems.The authors state that this set is not comprehensive.
  • Detailed sample: Four primary cases were selected for detailed analysis to balance depth and breadth across domains and the dimensions of expertise, collectiveness, and organization.The selection patterns emerged through iterative discussion among the research team.
  • Comparative design: The detailed comparison covers two image-cropping cases and two rating-platform cases with differing levels of collectiveness and other dimensions.Figure 1 summarizes the four cases and their variation.
  • Data and analysis: For each primary case, the authors reviewed platform discussion threads and supplemented them with relevant media and academic publications.They then traced case progression to derive broader lifecycle dynamics.

5 CASE STUDIES

The four case studies examine everyday audits of Twitter’s image-cropping algorithm, showing how users noticed suspected bias, tested competing explanations, and generated impacts that varied with collective participation.

  • Case-study overview: The cases compare everyday audits across dimensions including collectiveness, algorithmic expertise, and organicness.Figure 1 presents four illustrative cases across these dimensions, with darker shades indicating higher levels.
  • Twitter cropping algorithm: Twitter’s cropping algorithm uses neural networks to identify salient image regions and center them in thumbnails.The system is intended to highlight regions people are likely to look at while cropping out less interesting areas.
  • Case 1: Racial bias: Users suspected racial bias after observing that automated cropping focused on white faces and cropped out Black faces.The concern emerged during ordinary Twitter use and prompted users to investigate the algorithm’s choices.
  • Case 1: Racial bias: Users tested the racial-bias hypothesis by reversing image layouts and varying backgrounds, colors, people, skin tones, contrast, and other image features.Participants used their tests to support or challenge competing explanations of the algorithm’s behavior.
  • Case 1: Racial bias: The racial-bias audit became highly collective, creating a counterpublic space, media attention, and an official Twitter response.Users questioned the platform’s algorithmic authority, developed concrete concerns, and pushed for change.

5.1.2 Case 2. Gender Bias:

The gendered cropping audit attracted little sustained participation and had limited impact, contrasting with the more collective racial-cropping case and the varied dynamics of rating-platform audits.

  • Case 2: Gender bias: A Twitter user noticed that an image of two women and two men cropped out both women’s heads while highlighting the men’s heads and women’s chests.The issue was observed during normal Twitter use.
  • Case 2: Gender bias: Very few users joined the gendered cropping audit, resulting in little additional hypothesis formation or testing.Later discussion and testing remained minimal despite renewed attention in 2020.
  • Case 2: Gender bias: The gendered cropping issue gained little traction and remained mostly individualized despite more than a thousand likes on the resurfaced tweet.Few users engaged in conversation or participated in the audit.
  • Case 2: Gender bias: The audit’s small participant base limited publicity and platform change, unlike the racial-cropping audit.A brief small counterpublic formed, but participants mostly supported the issue rather than directing the investigation.
  • Rating-platform cases: Rating algorithms shape information presented to users, while their opaque mechanisms have generated controversy about algorithmic bias.The section introduces Yelp and Booking.com cases involving suspected review-filtering and rating-calculation problems.
  • Yelp case: Yelp users collectively tested suspected advertising-related review filtering through forum discussion, hypothesis formation, and further analysis.Some explanations were based on users’ folk theories rather than concrete data.
  • Rating-platform cases: The Yelp audit generated lawsuits, public awareness, and platform engagement, while Booking.com users independently tested a suspected rating floor but gained little publicity.Yelp users could communicate on-platform, whereas Booking.com users lacked an on-platform communication channel.

6 THE LIFETIME AND DYNAMICS OF AN EVERYDAY AUDIT

Everyday audits move from detecting problematic behavior through awareness raising, collaborative or individual hypothesis testing, and attempts at remediation. The process is non-linear, may stop early, and can produce platform changes or user-led repair.

  • Process overview: Audits proceed through initiation, awareness raising, hypothesis formation and testing, and remediation, although they may terminate before completing every phase.The phases provide a high-level framework for comparing audit paths and envisioning interventions.
  • Initiation: Users initiate audits individually or collectively after noticing harmful algorithmic behavior during normal use.Some cases begin with one user and spread, while others involve many users noticing potential bias independently.
  • Awareness raising: After detection, users broadcast discoveries to increase visibility and invite participation in discussion, hypothesis formation, and testing.Promotion can instead mainly raise awareness when platform communication channels limit deeper collaboration.
  • Hypothesizing and testing: Users develop folk theories and test algorithmic inputs, with tests sometimes failing and sometimes producing further evidence about system behavior.The paper identifies theory development and testing as important opportunities for users to discover biases.
  • Hypothesizing and testing: Audits combine different forms of collaboration, ranging from individual reporting or systematic testing to collective sensemaking and shared investigation.Testing can be organized by outside stakeholders or arise organically as users assess algorithmic behavior.
  • Remediation: Remediation includes publicity, legal action, platform-level changes, and user-led repair when platforms do not fix harmful behavior.The cases include nearly 700 Yelp lawsuits, Twitter’s cropping commitment, and collective amplification of suppressed TikTok content.

7 DESIGN IMPLICATIONS

The cases suggest that effective everyday auditing can be supported through community, expert, algorithmic, organizational, and incentive-based interventions. These interventions must balance assistance with users’ autonomy and the organic character of audits.

  • Community guidance: Across cases, more collective auditing was associated with potentially greater impact, motivating designs that help users coordinate and build on one another’s findings.Yelp users engaged in more collaborative auditing than Booking.com users.
  • Community guidance: Discussion spaces can help auditors raise awareness, surface severe reports, and review existing hypotheses and findings instead of repeating earlier work.Existing platforms may provide critical mass, but using an audited platform can create censorship and suppression concerns.
  • Expert guidance: Expert guidance can improve audits by supplying technical knowledge, assessing user hypotheses, and suggesting testing strategies.Expert involvement may diminish users’ sense of community and autonomy if handled without care.
  • Expert guidance: Bidirectional feedback systems should help developers guide auditors while also scaffolding users to provide developers with actionable feedback.The paper presents this as a critical direction for future research.
  • Algorithmic guidance: Algorithmic assistance could surface important instances and support comparative audits requiring statistical or causal hypothesis testing.Such tools may help everyday auditors who typically lack statistical or causal-modeling knowledge.
  • Organizational guidance and incentivization: Everyday audits distribute labor fluidly, with users specializing in testing, awareness raising, or hypothesis formation; synthesis mechanisms could help coordinate these roles.The paper also identifies incentives for auditors, developers, and platforms to support audits and remediate discovered issues.

8 DISCUSSION

Everyday algorithm auditing draws power from users’ lived experience, situated knowledge, collective sensemaking, and autonomy. The paper also emphasizes that effective intervention requires unresolved judgments about timing, degree, and trade-offs.

  • The power of everyday auditing: Everyday auditing uses users’ lived experiences to identify sensitive harms that auditors without relevant cultural backgrounds may overlook.The paper distinguishes this from crowdsourced auditing by emphasizing users’ connection to the systems and harms being examined.
  • The power of everyday auditing: Users’ situated knowledge helps reveal harmful machine behaviors that emerge only in real-world contexts with complex social dynamics and changing norms.Day-to-day interaction positions users to notice these context-dependent behaviors.
  • The power of everyday auditing: Counterpublics enable users to build on contributions, test one another’s hypotheses, reach consensus, and support publicity efforts collectively.Collective sensemaking is presented as a distinctive strength of everyday audits.
  • The power of everyday auditing: Users retain control over the audit’s direction rather than completing subtasks determined entirely by outside parties.Their autonomy and agency help steer audits toward issues they find useful and meaningful.
  • Interventions: The paper identifies five intervention categories: community, expert, algorithmic, organizational guidance, and incentivization.Each category involves potential trade-offs that remain open for future research.
  • Open questions: The appropriate timing and degree of intervention remain poorly understood because guidance may support audits while reducing their organic character, community, or autonomy.The authors describe the field as being in its infancy and call for further investigation.

9 CONCLUSIONS

The paper conceptualizes everyday algorithm auditing through real-world cases and proposes a process-oriented account of how users investigate problematic algorithmic behavior. It presents this work as an exploratory first step toward supporting and connecting everyday and formal auditing.

  • Contributions: Real-world case studies and theories of everyday resistance and counterpublics ground a process-oriented view of users auditing problematic machine behaviors.The view covers initiation, awareness raising, hypothesis formation and testing, and remediation.
  • Scope: The research is highly exploratory, and its cases and analytical approaches are illustrative rather than comprehensive or empirically settled.The authors call for follow-up studies of users’ practices and further exploration of the proposed design opportunities.
  • Conclusion: The work is an initial step toward bridging formal algorithm-auditing approaches with auditing behaviors that arise during everyday system use.It aims to inform future research and the design of platforms and tools that empower meaningful auditing.
Loading 2105.02980v2…