Source-linked AI summary

Unpacking the Social Media Bot: A Typology to Guide Research and Policy

Robert Gorwa, Douglas Guilbeault

arXiv:1801.06863v2cs.CY

TL;DR

The paper addresses persistent ambiguity about what social media bots are, how they function, and how they should be distinguished for policy purposes. It develops a history and typology of bots and political automation, finding that clearer concepts and measurements are needed amid difficult detection, data-access, and platform-incentive problems.

  • Problem

    Researchers and policymakers lack sufficiently clear, shared definitions and measurements of bots, complicating efforts to understand political manipulation and formulate effective policy.

  • Method

    The paper provides a history and typology of bots, categorizes their functions and uses, and identifies challenges for researching and governing political automation.

  • Results

    The analysis shows that bot-related policy is complicated by ambiguity in bot concepts, difficulty distinguishing legitimate from illegitimate automation, limited data access, and conflicting platform incentives.

  • Takeaways & Limitations

    Clearer shared concepts and measurements can facilitate cumulative research and provide clearer guidance for bot-policy development.

  • Takeaways & Limitations

    Academic bot research is constrained by limited access to sensitive account information and public APIs, leaving studies concentrated largely on Twitter and dependent on human-coded ground truth.

Abstract

from arXiv · show

Amidst widespread reports of digital influence operations during major elections, policymakers, scholars, and journalists have become increasingly interested in the political impact of social media 'bots.' Most recently, platform companies like Facebook and Twitter have been summoned to testify about bots as part of investigations into digitally-enabled foreign manipulation during the 2016 US Presidential election. Facing mounting pressure from both the public and from legislators, these companies have been instructed to crack down on apparently malicious bot accounts. But as this article demonstrates, since the earliest writings on bots in the 1990s, there has been substantial confusion as to exactly what a 'bot' is and what exactly a bot does. We argue that multiple forms of ambiguity are responsible for much of the complexity underlying contemporary bot-related policy, and that before successful policy interventions can be formulated, a more comprehensive understanding of bots --- especially how they are defined and measured --- will be needed. In this article, we provide a history and typology of different types of bots, provide clear guidelines to better categorize political automation and unpack the impact that it can have on contemporary technology policy, and outline the main challenges and ambiguities that will face both researchers and legislators concerned with bots in the future.

1 Introduction

Bots have become central to concerns about political manipulation, but unclear terminology and definitions hinder research, detection, and policy. The article responds with a historical typology and a framework organized around structural, functional, and usage-based ambiguities.

  • Social media bots have been implicated in misinformation, hyper-partisan messaging, and political conversations surrounding major elections.
  • Stakeholders disagree about what bots are, what they do, and why different academic communities define them differently.
  • Bot-related terminology is used imprecisely across robots, chatbots, social bots, political bots, botnets, sybils, cyborgs, and related concepts.
  • Conceptual ambiguity obstructs research and policy because bots must be defined precisely before their presence or absence can be investigated.
  • The article develops a typology and identifies ambiguity in bots’ structure, function, and uses, alongside challenges involving data access and detection.

2 A Typology of Bots

The history of bots spans diverse automated systems, from crawlers and chatbots to social and political bots, while hybrid human-automated accounts complicate classification and detection. These categories differ in operation, purpose, and relationship to human users.

  • Bots have historically referred to diverse software systems, including web scrapers, crawlers, indexers, chatbots, and autonomous agents.
  • Web crawlers automated large-scale website downloading and indexing, becoming key components of search engines while sometimes burdening servers.
  • Traffic estimates attributing more than half of web traffic to bots can conflate social automation with crawling, indexing, and scraping programs.
  • Chatbots conduct human-computer dialogue through natural language, whereas many social bots primarily communicate through public posts and predetermined or copied messages.
  • Political bots are social bots deployed for political purposes, including smear campaigns, hashtag disruption, and interference with political mobilization.
  • Terms such as troll farm, trolling, bot, and sockpuppet remain culturally and conceptually ambiguous across contexts.
  • Cyborgs combine automation and human curation, but unclear automation thresholds make them difficult to classify and detect accurately.

3 A Framework for Understanding Bots: Three Considerations

The article proposes understanding bots through three considerations: how they are built, what they do, and how they are used. This framework highlights why policy must distinguish technical structure, function, and normative purpose.

  • Framework overview: A typology must account for bots’ changing forms rather than impose a definitive prescriptive classification.
  • Structure: Structure concerns the operating environment, platform, code, and whether automation uses custom scripts, public tools, or content-management software.
  • Structure: Structure also requires distinguishing software bots from humans exhibiting bot-like behavior and hybrid accounts using automation tools.
  • The Bot’s Function: Function concerns whether bots operate accounts, identify or disguise themselves, converse with users, or send one-way mass messages.
  • The Bot’s Function: Function-based distinctions help target policy and prevent confusion between chatbots and social bots with different structures and activities.
  • The Bot’s Use: Use concerns the bot’s purpose, including political, ideological, accountability-oriented, commercial, positive, or harmful objectives.
  • The Bot’s Use: Because identical structures can support positive and negative actors, automation policies necessarily contain normative assumptions about acceptable behavior.
  • The Bot’s Use: Policy must avoid arbitrarily disrupting legitimate citizen-built political bots while addressing manipulative foreign influence operations, with greater transparency.

4 Current Challenges for Bot-Related Policy

Bot-related policy is constrained by difficult measurement, limited data access, unresolved responsibility, and conflicting platform, academic, public, and governmental interests. These challenges make it difficult to distinguish legitimate from malicious automation and to develop accountable interventions.

  • 4.1 Measurement and Data Access: 117,000 malicious applications were suspended and more than 450,000 suspicious logins were detected daily, illustrating the scale and difficulty of bot detection.Twitter reported these figures over a four-month period and per day, respectively.
  • 4.1 Measurement and Data Access: Researchers lack sensitive account information and cannot study Facebook bots through its public API, while human-labeled ground truth remains uncertain.Human coders are not fully reliable at identifying bots, limiting the precision and recall of academic detection methods.
  • 4.1 Measurement and Data Access: Private control of key data prevents the public from reliably determining bot activity’s scale, as competing estimates of election-ad exposure exceeded millions and one hundred million views.The passage presents disagreement over whether some views were generated by illegitimate automated accounts.
  • 4.2 Responsibility and Governance: Responsibility, jurisdiction, and authority remain unresolved as companies’ opaque automation policies acquire political ramifications and attract regulatory scrutiny.The policy debate is connected to governance of political content, hate speech, harassment, and misinformation.
  • 4.3 Contrasting Incentives: Platforms face incentives to preserve automation and user growth while distinguishing benign activity from malicious automation, limiting transparency and data sharing.Twitter’s open API supports productive automated accounts but also complicates efforts to suppress malicious activity; disclosure can threaten platform reputations and revenue.
  • 4.3 Contrasting Incentives: Platform interests often conflict with academic demands for access and public demands for intervention, creating trade-offs with no easy solutions.These tensions connect bot policy to broader debates about platform responsibility, technology governance, and accountability.

5 Conclusion

The article concludes that bot research and policy require clearer concepts and better measurement. Its typology is intended to connect quantitative findings with theoretical foundations while acknowledging persistent detection, data-access, and stakeholder challenges.

  • 5 Conclusion: The debate over bots and political automation is expected to become more central to discussions of social media, polarization, and fake news.The authors describe this debate as being in an embryonic stage.
  • 5 Conclusion: Quantitative studies have improved the identification and measurement of bot influence, but policy-relevant use requires clearer theoretical foundations.The authors frame the typology as a translational effort between quantitative and qualitative research.
  • 5 Conclusion: The article’s typology provides a framework for shared concepts and measurements concerning bots, media manipulation, and political automation.Its stated goal is to provide clearer guidance for developing bot policy.
  • 5 Conclusion: Future work must address imperfect detection, unreliable data, limited access, and overlapping public, corporate, and governmental interests.The authors identify these measurement, access, and interest conflicts as continuing challenges for researchers, policymakers, and journalists.
Loading 1801.06863v2…