Source-linked AI summary
Serf and Turf: Crowdturfing for Fun and Profit
Gang Wang, Christo Wilson, Xiaohan Zhao, Yibo Zhu, Manish Mohanlal, Haitao Zheng, Ben Y. Zhao
TL;DR
Crowdturfing uses organized human labor to perform malicious online tasks, challenging defenses built around automated attackers. The paper measures these systems, analyzes their campaigns and workforce, and tests campaign effectiveness. It finds effective campaigns, international worker mobility, and continued growth that threatens online communities.
Problem
Existing security mechanisms often assume malicious activity is automated, leaving them challenged by crowd-sourcing systems that organize real humans to perform malicious tasks.
Method
The paper crawls large crowdturfing systems, analyzes campaigns and worker sources, measures Weibo cascades, and runs benign active campaigns to evaluate end-to-end effectiveness.
Results
Crowdturfing workers generate large information cascades, avoid systems designed to catch automated spam, and drive hundreds of clicks from normal users.
Takeaways & Limitations
Crowdturfing’s continuing growth and global workforce pose a concrete threat to U.S.-based online communities and other social networks.
Takeaways & Limitations
The empirical study focuses on two large crowdturfing systems hosted in and targeting users in China, with incomplete campaigns comprising 8% on ZBJ and 12% on SDH.
Abstract
from arXiv · showhide
Popular Internet services in recent years have shown that remarkable things can be achieved by harnessing the power of the masses using crowd-sourcing systems. However, crowd-sourcing systems can also pose a real challenge to existing security mechanisms deployed to protect Internet services. Many of these techniques make the assumption that malicious activity is generated automatically by machines, and perform poorly or fail if users can be organized to perform malicious tasks using crowd-sourcing systems. Through measurements, we have found surprising evidence showing that not only do malicious crowd-sourcing systems exist, but they are rapidly growing in both user base and total revenue. In this paper, we describe a significant effort to study and understand these "crowdturfing" systems in today's Internet. We use detailed crawls to extract data about the size and operational structure of these crowdturfing systems. We analyze details of campaigns offered and performed in these sites, and evaluate their end-to-end effectiveness by running active, non-malicious campaigns of our own. Finally, we study and compare the source of workers on crowdturfing sites in different countries. Our results suggest that campaigns on these systems are highly effective at reaching users, and their continuing growth poses a concrete threat to online communities such as social networks, both in the US and elsewhere.
1. INTRODUCTION
Crowdturfing organizes real people to perform malicious tasks, undermining defenses designed around automated attackers. The paper measures these systems’ growth, operations, campaign effectiveness, and international workforce.
- 1. INTRODUCTION: Crowdturfing breaks security mechanisms that assume malicious tasks cannot be performed by real humans en masse.Such mechanisms include CAPTCHAs and machine-learning detectors of abnormal behavior.
- 1. INTRODUCTION: Measurements show malicious crowd-sourcing systems are rapidly growing in user base and revenue.The paper names these systems crowdturfing because they combine crowd-sourcing with astroturfing-like behavior.
- 1. INTRODUCTION: The study examines large-scale crowdturfing systems, focusing on two major systems hosted in and targeting users in China.Evidence of such systems also exists in countries including the United States and India.
- 1. INTRODUCTION: Detailed crawls quantify system size, operational structure, tasks, and revenue, with both tasks and revenue growing exponentially.The analyzed tasks include mass account creation, social-network posting, and attempts to start microblog information cascades.
- 1. INTRODUCTION: Benign active campaigns show crowdturfing can be cost-effective at soliciting real user responses, while workers cross national borders.Workers in less-developed countries may receive payment through global services for tasks affecting U.S.-based targets.
- 1. INTRODUCTION: The study is among the first to examine both the organization and effectiveness of large-scale crowdturfing systems.Its motivation is to support development of defenses for online social networks and communities.
2. CROWDTURFING OVERVIEW
Crowdturfing combines sponsored, concealed information campaigns with paid human labor organized through customer–agent–worker systems. Distributed structures resist scrutiny but lack accountability, whereas centralized sites simplify coordination while exposing themselves to third-party investigation.
- 2.1 Introduction to Crowdturfing: Crowdturfing combines crowd-sourcing with astroturfing to mobilize paid users for sponsored campaigns that appear decentralized.The campaigns may disseminate defamatory rumors, false advertising, or suspect political messages.
- 2.1 Introduction to Crowdturfing: Customers initiate campaigns, agents organize and fund workers, and workers complete specific tasks for fees.These three actors form the basic organizational structure of a crowdturfing campaign.
- 2.1 Introduction to Crowdturfing: Campaigns decompose into tasks whose submissions provide evidence for customer or agent verification.Submissions may be rejected or absent, so task and submission counts need not match.
- 2.2 Crowdturfing Systems: Distributed systems use small private groups led by intermediaries, while centralized systems directly connect customers and workers through websites.Centralized sites automate campaign management, payments, and reputation scoring.
- 2.2 Crowdturfing Systems: Distributed systems resist external threats but limit participation through weak accountability and fragmented access.Customers receive little assurance of satisfactory work, and workers have no guarantee of payment.
- 2.2 Crowdturfing Systems: Centralized systems reduce uncertainty through accessible websites, automated payments, and reputation mechanisms, but are susceptible to third-party scrutiny.Their public accessibility also enabled the paper’s crawling and analysis of large sites.
3. CAMPAIGNS, TASKS, AND REVENUE
The study measures two large Chinese crowdturfing systems and finds substantial, rapidly expanding activity across campaigns, tasks, workers, submissions, and revenue. Campaigns span account creation, social-media spamming, and other services, with highly skewed worker participation and rapid responses.
- Scale and revenue: ZBJ attracts more campaigns, workers, and money than younger SDH, while both systems process many tasks per campaign.The study focuses on these two large Chinese systems because readily available data enabled detailed measurement.
- Tasks and submissions: 36% of ZBJ tasks receive submissions versus 130% of SDH tasks, and roughly 50% of all submissions are accepted.The 130% figure indicates competition among workers for the same SDH tasks.
- Scale and revenue: More than $4 million has been spent on ZBJ and SDH over five years, while campaigns and total spending grow exponentially.Both sites charge a 20% fee on campaign dollars.
- Tasks and submissions: ZBJ campaigns tend to contain an order of magnitude fewer tasks than SDH campaigns, although both accept approximately 50% of submissions.SDH’s excess submissions make accepted submissions closely track requested tasks, especially above 100 tasks.
- Campaign types: Account-registration campaigns make target websites or games appear popular, while other campaigns pay workers to post content in QQ groups, forums, blogs, and microblogs.These five major campaign categories account for 88% of ZBJ campaigns and 91% of SDH campaigns; microblog campaigns grow fastest in both systems.
- Workers: Worker activity is highly skewed: career crowdturfers generate approximately 75% of submissions, while average workers complete about 5–7 tasks.About 40% of SDH workers complete one task, compared with 20% on ZBJ.
- Workers: 75% of ZBJ campaigns and 50% of SDH campaigns receive a first submission within 24 hours, but ramp-up can take 15 days on ZBJ and 30 days on SDH.The larger ZBJ worker population corresponds to faster activation in this comparison.
- Workers and compensation: Most submissions occur during workdays and evenings, indicating that workers rather than automated bots generate them.Workers generally earn $0.11 per accepted submission, and close to 70% earn less than $1 overall.
4. CROWDTURFING ON MICROBLOGS
The study measures how crowdturfing campaigns spread through Sina Weibo, from worker-generated cascades to exposure among normal users. Campaign reach is substantial, but virality depends more on campaign cost, format, and customer behavior than on individual workers.
- Data collection: Campaigns are modeled as forests of information-cascade trees, with customer or worker origin posts and retweets attributed to workers or normal users.The number of messages includes posts from customers, workers, and normal users who retweet campaign content.
- Campaign scale: 2,869 campaigns involving 1,280 customers reached more than 463,000 non-worker users through submissions from over 12,000 Weibo accounts.Two percent of worker accounts were inaccessible, compared with 0.08% of non-worker accounts; customer accounts remained active.
- Account characteristics: Worker accounts were designed to blend in: they tweeted as frequently as normal accounts and generally had follow rates ≈1.Customers tended to have follow rates >1, consistent with their role as commercial information disseminators.
- Campaign reach: 50% of campaigns generated ≤146K messages, while 8% exceeded 1M messages; workers produced the vast majority because normal-user retweets were rare.Total messages provide an upper bound on audience size because the social graph was incomplete and duplicate exposures could not be quantified.
- Campaign reach: Pay-per-retweet campaigns engaged normal users more successfully than pay-per-tweet campaigns, with 50% reaching cascade depths greater than 2.Pay-per-tweet campaigns were very shallow, whereas pay-per-retweet campaigns typically included at least one normal-user retweet.
- Factors impacting success: Campaign size increased linearly with spending, but individual workers did not explain viral success; a small group of customers showed consistent virality.The highlighted group contained 20 customers, or 1.5% of customers, who initiated several campaigns and achieved significant numbers of viral campaigns.
5. ACTIVE EXPERIMENTS
The authors ran benign advertising campaigns through ZBJ to measure crowdturfing’s end-to-end delivery and user response. Workers responded quickly, campaigns generated substantial traffic, and effectiveness varied by target network.
- Experimental setup: Nine campaigns targeted Weibo, QQ groups, and discussion forums using a measurement server that logged worker and recipient clicks before redirecting users.Workers were directed to post advertisements for legitimate online stores, while search-engine and bot clicks were filtered from the logs.
- Campaign execution: Seven short campaigns received sufficient submissions, and six completed within a few hours; together they generated 894 submissions from 224 workers.Workers continued submitting after campaigns filled, hoping earlier submissions would be rejected and rewards would become available.
- Campaign execution: More than 80% of submissions arrived within one hour for Weibo and forum campaigns, and within six hours for QQ campaigns.Response times were aggregated across campaign types because completion depended on the number of accounts workers controlled on each network.
- Results and analysis: QQ campaigns generated more clicks than Weibo campaigns despite producing only 1/5 as many messages.The authors suggest QQ pop-up messages may receive more views and clicks than Weibo tweets.
- Results and analysis: The iPhone4S and Maldives campaigns generated 491 and 218 click-backs, respectively, at a cost of $45 each.Their costs per click were $0.21 and $0.09, respectively, although the authors note that better targeting could reduce these costs.
- Results and analysis: The Maldives campaign coincided with sales increasing from 4 trips in the prior month to 11 trips on launch day, followed by no additional sales the next month.The authors state they cannot be sure of causation, but consider the 218 campaign clicks likely responsible for the sales.
6. CROWDTURFING GOES GLOBAL
Crowdturfing extends beyond China through active U.S. sites and an Indian service designed for low-bandwidth users. The market’s prevalence varies sharply across platforms and is supported by international labor access.
- United States: Crowdturfing systems in the United States are active and supported by an international workforce.The authors combine additional crawls with prior research to characterize the U.S. market.
- United States: Crowdturfing comprised 12% of Mechanical Turk campaigns in October 2011, down from the 41% reported for 2010.The authors measured Mechanical Turk hourly for one month and classified tasks using keyword analysis.
- United States: Between 70-95% of campaigns on four other U.S.-based sites were classified as crowdturfing.MinuteWorkers, MyEasyTask, and Microworkers were crawled daily for October 2011, while ShortTask provided one year of historical data.
- United States: Alternative sites support the underground market by tolerating crowdturfing and offering payment methods accessible to an international workforce.This contrasts with Mechanical Turk’s enforcement against spammy jobs and restrictions associated with payment access.
- India: Paisalive uses email for task requests and submissions, targeting rural workers with low-bandwidth or intermittent Internet connectivity.The Indian site is small and pays lower wages than the other measured services.
7. RELATED WORK
Related research has studied crowdsourcing labor and spam manifestations, but this paper focuses on the underlying systems that organize crowdturfing activity. It connects crowdturfing to broader work on opinion manipulation and deceptive content.
- Crowd-sourcing research: Prior crowdsourcing research examined Mechanical Turk worker demographics, task pricing, user-study methodology, and Micro Workers’ characteristics.These studies provide context for the labor platforms on which crowdturfing can operate.
- OSN spam and detection: OSN-spam research identified fake accounts and campaigns on Facebook, Twitter, and Renren and developed machine-learning defenses.Those studies primarily analyze and defend against the outward manifestations of spam, whereas this paper identifies underlying organizing systems.
- Opinion spam: Opinion-spam research has addressed fake reviews, fake news-site comments, political astroturfing, and deceptive reviews generated by Mechanical Turk workers.The authors position these findings as consistent with crowdturfing’s growth as a global threat.
8. CONCLUSION
Crowdturfing has already produced substantial spending and is expanding rapidly, while career workers generate effective spam and evade automated defenses. Its international workforce and payment infrastructure extend the threat beyond China to U.S. websites.
- $4 million has already been spent on ZBJ and SDH, while their campaigns and spending grow exponentially.
- Career crowdturfers control thousands of OSN accounts manually, generate large information cascades, and avoid automated-spam defenses.
- International payment systems and a global workforce enable crowdturfing sites to target U.S. websites.