Source-linked AI summary
Uncovering Coordinated Networks on Social Media: Methods and Case Studies
Diogo Pacheco, Pik-Mai Hui, Christopher Torres-Lugo, Bao Tran Truong, Alessandro Flammini, Filippo Menczer
TL;DR
Social media create vulnerabilities and abuse opportunities that enable coordinated influence and manipulation. This paper introduces an unsupervised network methodology that builds coordination networks from shared behavioral traces, and shows its applicability across diverse campaigns, including cases involving both likely bots and human accounts.
Problem
Social media create new vulnerabilities and abuse opportunities for influence campaigns, while automated-account detection alone does not capture coordinated activity.
Method
The paper builds coordination networks by connecting accounts to extracted behavioral features and identifying unexpectedly similar behaviors, regardless of automation or intent.
Results
The approach detects coordinated campaigns through identities, images, text, retweets, and action timing, including networks mixing likely bot and human accounts.
Takeaways & Limitations
Coordination detection can provide a unified, flexible framework for uncovering different forms of coordinated behavior across information-warfare scenarios.
Takeaways & Limitations
The approach identifies coordination but does not establish participants’ intent, authenticity, or underlying mechanisms.
Abstract
from arXiv · showhide
Coordinated campaigns are used to influence and manipulate social media platforms and their users, a critical challenge to the free exchange of information online. Here we introduce a general, unsupervised network-based methodology to uncover groups of accounts that are likely coordinated. The proposed method constructs coordination networks based on arbitrary behavioral traces shared among accounts. We present five case studies of influence campaigns, four of which in the diverse contexts of U.S. elections, Hong Kong protests, the Syrian civil war, and cryptocurrency manipulation. In each of these cases, we detect networks of coordinated Twitter accounts by examining their identities, images, hashtag sequences, retweets, or temporal patterns. The proposed approach proves to be broadly applicable to uncover different kinds of coordination across information warfare scenarios.
Introduction
Social media broaden participation but also create opportunities for coordinated manipulation that individual-account bot detectors may miss. The paper proposes an unsupervised network approach and demonstrates it across five Twitter coordination traces.
- Malicious actors exploit social media to spread disinformation, manipulate users, and conceal control of social bots.
- Individual-account detection can miss coordination tactics that appear innocuous separately but suspicious when accounts interact.A single handle change may be normal, whereas accounts rotating handles are unlikely to be coincidental.
- The paper builds coordination networks by linking accounts with unexpectedly similar behavioral traces, without requiring labels.The approach can operate regardless of whether accounts are automated or organic and whether intent is malicious or benign.
- Five Twitter case studies examine coordination through handle changes, image sharing, hashtag sequences, co-retweets, and synchronization.
- The case studies detect campaigns through identities, images, text, retweets, and timing, illustrating the approach’s generality.
- Coordinated malicious campaigns can combine likely bot and human accounts rather than consisting only of automated accounts.
Related Work
Prior work largely detects individual bots or coordination along particular dimensions. This paper addresses that fragmentation with a broader unsupervised framework that can also capture human-controlled coordination.
- Most abuse-detection tools target social bots using features from individual accounts or tweets.
- Single-account supervised methods are less effective for coordinated accounts, motivating a shift toward unsupervised learning.
- Existing approaches often consider only one coordination dimension and emphasize patterns associated with automated accounts.
- The proposed methodology unifies multiple similarity criteria and extends coordination detection to human-controlled accounts.
Methods
The methodology converts behavioral traces into an account coordination network and extracts suspicious clusters through unsupervised network analysis. Its flexible feature design supports content, activity, identity, and combined traces.
- The approach has four phases: behavioral-trace extraction, bipartite network construction, projection onto an account network, and cluster analysis.
- Behavioral traces encode suspicious content, activity, identity, or combinations of dimensions, with combined traces potentially reducing false positives.
- Low-support accounts and low-weight edges can be filtered before community detection to prioritize precision over recall.
- Accounts connect to extracted features in a weighted bipartite network, with weights reflecting association strength and optional normalization for popular features.
- The bipartite network is projected into an undirected account network whose edge weights quantify feature similarity.
- Eight actionable steps operationalize the method, from formulating suspicious behavior through extracting coordinated groups.
- Five case studies instantiate the framework using shared identities, images, hashtag sequences, co-retweets, and activity patterns.
Case Study 1: Account Handle Sharing
The first case study detects coordinated Twitter accounts through reused handles. The resulting network reveals star-like hijacking or squatting patterns, a large multi-campaign component, and concrete suspicious clusters.
- Handle reuse exposes users to username squatting and impersonation because handles are changeable and generally reusable.
- The study analyzes 54 million Botometer records containing 1.9 million handles, retaining users with at least ten queries to observe repeated changes.
- Suspicious handles are shared by at least two accounts, and projected account components represent coordinated clusters based on handle co-occurrence.
- The resulting weighted undirected network contains 7,879 Twitter-account nodes.
- Reciprocal handle switches occur 12 times more often in star-like components than in other components, consistent with hub-mediated squatting or hijacking.
- The giant component contains 722 accounts sharing 181 names and divides into 13 Louvain subgroups suspected to represent campaigns by one group.
- One documented cluster saw 23 accounts adopt @GullyMN49 within five days, while 21 of those accounts remained active after the handle was banned.
- Another six-account cluster shared seven politically conflicting handles and was linked to apparent fundraising attempts targeting opposing groups.
Case Study 2: Image Coordination
This case study detects suspicious image-sharing coordination among Twitter accounts discussing the 2019 Hong Kong protests. The resulting clusters include both pro- and anti-protest groups and account for image variants as well as exact matches.
- Data and preprocessing: 31,772 protest-related tweets containing images were collected across six languages for the Hong Kong case study.The collection used a couple dozen protest-related hashtags.
- Coordination detection: Image coordination is inferred from similar 384-dimensional RGB color-histogram vectors rather than image URLs.Each color channel is binned into 128 intervals, allowing identical images and slight variants to match.
- Data and preprocessing: Accounts tweeting fewer than five images are excluded to reduce noise from insufficient evidence.The threshold can be adjusted to trade precision and recall and was selected to prioritize precision while retaining reasonable recall.
- Coordination detection: The projected account network retains the largest 1% of Jaccard-weighted edges, then ranks connected components after removing singletons.Accounts are linked through shared image features in a bipartite network before projection.
- Analysis: Three suspicious clusters involving 315 accounts shared pro- or anti-protest images, with Chinese text in anti-protest content and English text in pro-protest content.The network visualization distinguishes likely coordinated accounts and colors the three largest components by image content.
- Analysis: Shared image features included exact matches and brightness or cropping variants, represented by 59 pro-protest and 61 anti-protest image URLs.The same feature can therefore correspond to multiple URLs for visually similar images.
Case Study 3: Hashtag Sequences
This case study identifies coordinated accounts through repeated ordered hashtag sequences, motivated by the expectation that paraphrased campaign messages may retain common targeting hashtags. Applied to 2018 U.S. midterm-election tweets, the method found 617 daily coordination instances across 1,809 accounts.
- Motivation: The approach targets coordinated campaigns because paraphrased messages may still preserve identical hashtags associated with campaign targets.This motivation is presented as a conjecture rather than an established guarantee.
- Data: The dataset contains original tweets collected around the 2018 U.S. midterm election and is split into daily intervals.Daily intervals are used to detect when account pairs become coordinated.
- Method: Accounts must produce at least five tweets and five unique hashtags within 24 hours to provide sufficient coordination support.More stringent filtering could reduce the chance of coincidental sequence matches.
- Method: Ordered hashtag sequences combine content and activity traces to connect accounts posting identical sequences.Accounts and hashtag sequences form a bipartite network whose projection produces account edges.
- Analysis: Large connected components are treated as more suspicious because many accounts sharing hashtag sequences are less likely to arise by chance.Singleton nodes are removed before connected components are extracted.
- Results: 617 daily coordination instances involved 1,809 unique accounts, producing 32 suspicious groups on one illustrated day.The largest component had 404 nodes using the “Backfire Trump” application, while the smallest groups contained pairs of accounts.
Case Study 4: Co-Retweets
This case study applies co-retweet analysis to Twitter narratives about the White Helmets during the Syrian civil war. The resulting network highlights groups amplifying pro- and anti-White Helmets messages.
- Motivation: Retweeting the same tweets or accounts is used as a behavioral trace because shared amplification may signal coordination.The case study focuses on retweets as a common form of information-source amplification.
- Data: The dataset covers White Helmets-related Twitter activity collected with English and Arabic keywords.The organization had been targeted by disinformation campaigns during the Syrian civil war.
- Method: Accounts and retweeted messages form a bipartite network, weighted with TF-IDF and projected using cosine similarity between account vectors.Self-retweets and accounts with fewer than ten retweets are excluded.
- Results: The co-retweet network highlights two coordinated groups: orange accounts retweet pro-White Helmets messages, while purple accounts retweet anti-White Helmets messages.The figure includes exemplar retweets for both groups.
Case Study 5: Synchronized Actions
This case study uses synchronized posting times to investigate coordinated cryptocurrency pump-and-dump campaigns on Twitter. The network flags suspicious clusters across multiple cryptocurrencies, while market volatility limits quantitative validation.
- Data: Cryptocurrency-related original tweets and retweets were collected for 25 vulnerable coins using keywords and cashtags such as $BTC.Both tweet types were included because each contributes to the information stream considered by potential buyers.
- Method: Accounts are connected through tweet-time bins weighted with TF-IDF and projected using cosine similarity between account vectors.The approach uses temporal proximity as the behavioral trace and manual inspection found much unrelated content in the network.
- Method: At least eight messages are required per account to reduce false-positive matches from coincidental temporal overlap.Shorter intervals reduce coincidences but produce fewer matches and require more computation.
- Results: Purple network subgraphs flagged coordinated accounts associated with suspected pump-and-dump schemes involving many cryptocurrencies.Example tweets promoted the Indorse Token and Bitcoin by alleging business intelligence and hinting at future price increases.
- Limitations: Market volatility and difficulty attributing price changes to Twitter activity make quantitative validation difficult.The authors nevertheless observed synchronized-tweet surges preceding later record prices for Verge, Enjin, and DigiByte.
- Results: Dense gray clusters represented spam or coordinated advertising rather than pump-and-dump schemes, but still reflected coordinated manipulation.The figure omits singleton accounts and identifies suspicious clusters through connected components.
Discussion
The framework offers a flexible, unified approach to detecting coordination across platforms and behavioral traces, while its interpretations remain bounded by methodological and data limitations. Results also show that coordinated accounts are not reliably identified through bot detection alone.
- The methodology can extend beyond Twitter, including image coordination on Instagram and content-based coordination among Facebook pages.
- The framework unifies diverse unsupervised coordination methods by representing them as special cases based on different similarity or temporal schemes.
- The approach identifies coordination but does not determine participants’ intent, authenticity, or underlying mechanisms.
- Coordinated accounts are not necessarily bots: many have low, human-like bot scores, with the majority in two of three analyzed cases.
- Case studies 1 and 3 resemble activity-biased random tweet samples, whereas Case Study 2 has higher bot scores, possibly reflecting high-volume image posting by bots.
- The method prioritizes minimizing false positives through design and parameter choices, but more rigorous null models are needed to exclude chance links formally.
- The implementations mainly examine single behaviors, although combining dimensions can reveal either larger groups or separate, independent coordination campaigns.
Conclusion
The paper concludes that a network-based approach can identify coordinated accounts across multiple coordination types on Twitter. It is intended to complement individual-account methods and provide a unified framework for studying coordinated campaigns.
- The network approach identifies coordinated accounts on social media and detects multiple coordination types on Twitter.
- The approach complements rather than replaces individual-level bot or troll detection by targeting coordinated behavior at the group level.
- The unified framework may help researchers compare similarities and differences among approaches to coordinated-campaign detection.
- The framework is planned for incorporation into BotSlayer to broaden participation in countering social-media disinformation.