Source-linked AI summary
Disinformation and Social Bot Operations in the Run Up to the 2017 French Presidential Election
Emilio Ferrara
TL;DR
Social-media disinformation campaigns can exploit bots to manipulate political discussion, but their audiences and reach require empirical characterization. This study analyzes MacronLeaks using a large French-election Twitter dataset, bot-detection methods, and behavioral comparisons. The campaign peaked near 300 tweets per minute, while its audience was predominantly English-speaking American alt-right users rather than French users.
Problem
Social bots can amplify disinformation and potentially alter public opinion, motivating analysis of such operations in political events.
Method
The study analyzes MacronLeaks using nearly 17 million election-related tweets, machine-learning and cognitive approaches, and comparisons between bots, humans, and broader election discussion.
Results
Nearly 300 tweets per minute marked MacronLeaks’ peak, while its audience was mostly English-speaking American alt-right users rather than French users.
Takeaways & Limitations
The campaign’s audience profile was identified as a reason for its scarce success in affecting the French vote outcome.
Takeaways & Limitations
Public Botometer could not analyze the study’s large user population or accounts that were suspended, quarantined, protected, or deleted.
Abstract
from arXiv · showhide
Recent accounts from researchers, journalists, as well as federal investigators, reached a unanimous conclusion: social media are systematically exploited to manipulate and alter public opinion. Some disinformation campaigns have been coordinated by means of bots, social media accounts controlled by computer scripts that try to disguise themselves as legitimate human users. In this study, we describe one such operation occurred in the run up to the 2017 French presidential election. We collected a massive Twitter dataset of nearly 17 million posts occurred between April 27 and May 7, 2017 (Election Day). We then set to study the MacronLeaks disinformation campaign: By leveraging a mix of machine learning and cognitive behavioral modeling techniques, we separated humans from bots, and then studied the activities of the two groups taken independently, as well as their interplay. We provide a characterization of both the bots and the users who engaged with them and oppose it to those users who didn't. Prior interests of disinformation adopters pinpoint to the reasons of the scarce success of this campaign: the users who engaged with MacronLeaks are mostly foreigners with a preexisting interest in alt-right topics and alternative news media, rather than French users with diverse political views. Concluding, anomalous account usage patterns suggest the possible existence of a black-market for reusable political disinformation bots.
INTRODUCTION
Social media can be exploited through coordinated automation and disinformation to influence political discussion. This paper examines the MacronLeaks operation surrounding France’s 2017 presidential election and defines it through unverified information and coordinated sharing.
- Automated social-media campaigns can generate large volumes of posts supporting or attacking political candidates.
- Social bots used in disinformation campaigns may reach enough users to dominate public discourse and redirect attention toward manufactured information.
- The study investigates MacronLeaks, a potentially disruptive disinformation campaign during the run-up to the 2017 French presidential election.
- MacronLeaks qualifies as disinformation because it combined unverified shared information with a coordinated effort to distribute it.
- The campaign involved 4chan as an incubator for alleged incriminating material and a social-bot operation that amplified its circulation before Election Day.
- The study monitored Twitter from April 27 through May 7, 2017, identified humans and bots, and compared their characteristics and interactions.
METHODS
The study collected election-related Twitter data, isolated MacronLeaks posts, and developed a scalable bot-detection pipeline using account metadata, activity features, and supervised learning. Logistic Regression provided a fast, high-performing basis for classifying more than two million users.
- Data collection: The researchers selected 23 hashtags and keywords covering both candidates and general election discussion to collect the election-related stream.
- Data collection: Researchers collected approximately 17 million unique election-related tweets from 2,068,728 unique users using Twitter’s Search API.
- Data collection: Nearly 350 thousand tweets containing five campaign terms formed the MacronLeaks corpus, about 2% of the overall election dataset.
- Bot detection: Botometer could not support this study because Twitter rate limits prevented large-scale analysis and account restrictions concealed information for suspended or deleted users.
- Bot detection: The custom detector used historical tweets and account metadata, enabling classification of more than two million users without querying recent Twitter data.
- Bot detection: The approach emphasized profile customization, geographic metadata, and activity statistics as signals distinguishing humans from political bots.
- Model evaluation: Random Forests achieved 93% accuracy and 92% AUC-ROC, while Logistic Regression achieved 92% accuracy and 89% AUC-ROC and was substantially faster.
DATA ANALYSIS
The analysis characterizes the MacronLeaks conversation, its bot population, and the campaign’s timing and coordination patterns. The results show a substantial bot presence, distinctive bot behavior, and evidence consistent with reusable political disinformation accounts.
- Timeline and volume: Nearly 300 tweets per minute marked the MacronLeaks peak between May 5 and May 6, briefly approaching the scale of regular election discussion.The campaign began on Twitter on April 30 and remained largely silent beforehand.
- Bot prevalence: 18% of 99,378 MacronLeaks participants—18,324 accounts—were classified as social bots, compared with 81,054 human users.The bot share was consistent with the authors’ prior analysis of the 2016 U.S. presidential election.
- Bot characteristics: Among the top 15 detected bots, 13 were manually verified as correct bots, yielding nearly 87% accuracy versus 92% cross-validation accuracy and 89% AUC-ROC.Four accounts had been deleted, seven suspended, two quarantined, and two remained active or potentially misclassified.
- Bot characteristics: Two suspicious bot families used randomly generated names ending in 2020 or _1337, with additional less-active accounts showing matching behavioral patterns.The _1337 suffix references “leet,” an alternative alphabet associated primarily with hacking communities.
- Bot characteristics: Some bot accounts were created before the 2016 U.S. election, used briefly for an alt-right campaign, and later appeared in MacronLeaks, supporting a possible market for reusable political disinformation bots.The authors present this as evidence supporting a hypothesis, not as definitive proof of such a market.
- Bot characteristics: Bots were less active than humans across tweet, follower, friend, favorite, and list-appearance measures, with statistically very significant distribution differences.Bots averaged 2.86 MacronLeaks-related tweets versus 3.81 for humans, and 1,382 followers versus 2,510.
DISCUSSION AND CONCLUSIONS
The paper analyzes the MacronLeaks disinformation campaign and identifies anomalous bot activity alongside an audience largely composed of English-speaking American alt-right users. These findings support a hypothesis that reusable political disinformation bots may exist, while helping explain the campaign’s scarce effect on the French vote.
- DISCUSSION AND CONCLUSIONS: The study combines machine learning, cognitive heuristics, event reconstruction, and comparison with general election-related discussion to analyze MacronLeaks.The general election conversation serves as a baseline for identifying differences and anomalies.
- DISCUSSION AND CONCLUSIONS: The authors hypothesize that a black market of reusable political disinformation bots may exist.They report bots used during the 2016 U.S. presidential election, later going dark, and returning before the 2017 French election.
- DISCUSSION AND CONCLUSIONS: Most MacronLeaks participants were English-speaking American alt-right users rather than French users, unlike the broader election-related conversation.The baseline involved significantly more French users and showed a trend supporting Emmanuel Macron.
NOTES
The paper notes that the 2016 U.S. election’s event dynamics remained unclear in June 2017, while evidence supported possible foreign-government or vested-interest meddling.
- NOTES: The exact dynamics of the 2016 U.S. election remained unclear and were subject to ongoing federal investigations.The passage reports mounting evidence for possible meddling by foreign governments and organizations with vested interests.
TABLES
The tables define the study’s collection scope, identify prominent accounts and bots, and contrast language and information sources in general election discussion versus MacronLeaks.
- Data collection: 23 keywords were continuously collected from April 27 through May 7, 2017, covering both presidential candidates and general election terms.
- Conversation structure: The top-20 hashtag and mention tables rank topics and users by tweet volume during the election observation window.
- Bots and amplification: The bot tables list accounts detected by the study’s algorithms, ranked by MacronLeaks activity, and report their Twitter enforcement status.Among the top 15 detected bots, four were deleted, seven suspended, two quarantined, and two remained active.
- Bots and amplification: Ten frequently retweeted bots were all suspended, deleted, or quarantined, while six accrued substantial follower counts during the campaign.
- Information sources: General election URLs primarily point to candidates, politicians, and established news media, whereas MacronLeaks URLs include hyper-partisan outlets, leaked data dumps, and fake-news websites.
FIGURES
The figures track tweet volumes, metadata distributions, feature correlations, and the timing of human and bot activity in the MacronLeaks campaign.
- Activity timelines: The timeline compares MacronLeaks tweet volume with generic election-related discussion from April 27 through Election Day on May 7.
- Account features: The human-user and social-bot boxplots display distributions of their respective metadata features.
- Account features: The correlation heat maps show feature relationships separately for human users and social bots.
- Activity timelines: Bot-generated-content spikes often slightly precede human-post spikes, suggesting that bots can trigger disinformation cascades.
- Corpus comparison: The corpus-statistics figure compares MacronLeaks tweets with an equal-sized random sample of French-election tweets across user activity and token distributions.