Source-linked AI summary

Social Turing Tests: Crowdsourcing Sybil Detection

Gang Wang, Manish Mohanlal, Christo Wilson, Xiao Wang, Miriam Metzger, Haitao Zheng, Ben Y. Zhao

arXiv:1205.3856v2cs.SIphysics.soc-ph

TL;DR

Sybils increasingly threaten online social networks while sophisticated accounts evade automated detectors built on behavior and graph assumptions. The paper evaluates crowdsourced human detection through a multi-network user study and uses the findings to design a scalable system. Experts and students achieve exceptionally good detection with near-zero false positives, while turkers vary more and make additional mistakes with fatigue; simulations show a hierarchical two-tier system can be accurate and scalable at reasonable cost.

  • Problem

    Sophisticated Sybils can befriend legitimate users and evade automated detectors that rely on specific behavior and graph assumptions.

  • Method

    The paper conducts a user study using ground-truth Sybil accounts from Renren, Facebook-US, and Facebook-India, then designs a crowdsourced detection system from the results.

  • Results

    Experts and undergraduate students produced exceptionally good detection with near-zero false positives, while turkers missed more Sybils and made more mistakes as profiles accumulated.

  • Takeaways & Limitations

    A hierarchical two-tier crowdsourced system can provide accurate, scalable Sybil detection with reasonable total costs.

  • Takeaways & Limitations

    Evaluation is constrained by the ground-truth Sybils included, so reported detection accuracy is a lower bound.

Abstract

from arXiv · show

As popular tools for spreading spam and malware, Sybils (or fake accounts) pose a serious threat to online communities such as Online Social Networks (OSNs). Today, sophisticated attackers are creating realistic Sybils that effectively befriend legitimate users, rendering most automated Sybil detection techniques ineffective. In this paper, we explore the feasibility of a crowdsourced Sybil detection system for OSNs. We conduct a large user study on the ability of humans to detect today's Sybil accounts, using a large corpus of ground-truth Sybil accounts from the Facebook and Renren networks. We analyze detection accuracy by both "experts" and "turkers" under a variety of conditions, and find that while turkers vary significantly in their effectiveness, experts consistently produce near-optimal results. We use these results to drive the design of a multi-tier crowdsourcing Sybil detection system. Using our user study data, we show that this system is scalable, and can be highly effective either as a standalone system or as a complementary technique to current tools.

1 Introduction

Sybils threaten OSN security while increasingly sophisticated accounts evade automated detectors that assume limited interaction with legitimate users. This paper studies whether crowdsourced human judgments can provide accurate, scalable detection.

  • Threat and detection gap: Sybils are rapidly growing fake identities used in coordinated spam and malware campaigns across online social networks.Measurement studies have detected hundreds of thousands of Sybil accounts, and Facebook reported up to 83 million potentially fake users.
  • Threat and detection gap: Most automated detectors assume Sybils struggle to befriend legitimate users and cluster together, but sophisticated attackers increasingly violate these assumptions.Renren Sybils rarely linked to one another and instead attempted to infiltrate legitimate users’ networks.
  • Research questions: The study asks how human accuracy, cost, and scalability vary with experience, motivation, fatigue, language, and cultural barriers.These questions guide the feasibility assessment of crowdsourced authenticity checks for suspicious profiles.
  • Study design: The authors collect ground-truth Sybil accounts from Renren, Facebook-US, and Facebook-India for a large user study.Renren provided Sybil data, while Facebook accounts were obtained by crawling highly suspicious profiles before banning.
  • Findings and system: Experts and undergraduate students achieved exceptionally good detection with near-zero false positives, while turkers missed more Sybils but also produced near-zero false positives.Experts maintained accuracy as they examined more profiles, whereas crowdworkers made more mistakes over time.
  • Findings and system: The proposed crowdsourced system achieves accuracy and scalability with reasonable costs in trace-driven evaluation.The system is presented as a standalone detector or complement to existing tools.

2 Background and Motivation

Crowdsourcing offers flexible human labor for tasks that automated systems handle poorly, but OSN defenders have not generally crowdsourced fake-account identification. Existing Sybil detectors remain vulnerable because their behavior and graph assumptions do not generalize across networks and attack strategies.

  • Crowdsourcing: Crowdsourcing decomposes work into short tasks assigned to a flexible, on-demand group of workers for small fees.Its benefits include distributing effort, dynamically changing the workforce, and enabling elasticity.
  • Crowdsourcing: MTurk illustrates this model through Human Intelligence Tasks completed by turkers, with around 100,000 HITs available at any time.The passage reports that 90% of these tasks pay ≤$0.10 each.
  • Crowdsourcing in OSNs: OSNs use crowdsourcing for content moderation, but Facebook and Tuenti retained dedicated in-house staff for fake-account identification.The paper presents crowdsourced Sybil detection as an unestablished OSN application at the time of study.
  • Adversarial crowdsourcing: Attackers also crowdsource Sybil creation and spam, producing accounts managed by real people that appear more authentic and can bypass CAPTCHAs.Human-managed fake profiles are particularly dangerous because they are less visibly automated.
  • Limits of automated detection: Existing Sybil detectors depend on network-specific assumptions, so none is general enough to perform well across all OSNs and attack strategies.Common assumptions include difficulty friending legitimate users and forming many links among Sybils.
  • Limits of automated detection: Feature-based detectors can work in settings such as Renren but fail to generalize when platforms like Twitter do not require social connections for spam.This illustrates how detector effectiveness depends on platform behavior.
  • Proposed direction: The paper proposes crowdsourced Sybil detection because humans can assess complex profile cues, social-Turing tests avoid fixed features, and crowdsourcing costs less than full-time moderation.The study then tests accuracy, demographics, fatigue, and system effectiveness.

3 Experimental Methodology

The study evaluates crowdsourced Sybil detection using controlled tests on ground-truth profiles from Facebook and Renren. It constructs regional datasets, verifies classifications, and measures how testers distinguish fake from legitimate accounts.

  • Datasets: The experiments use Renren, Facebook-US, and Facebook-India populations, with testers evaluating profiles classified as Sybil, legitimate, or suspicious.Renren data came from the security team; Facebook data were collected through crawling and regional sampling.
  • Facebook collection: Legitimate Facebook profiles were sampled from 86K profiles reached through friends-of-friends crawling from eight trusted laboratory seeds.The sampling rationale relied on the reported transitivity of trust in social networks.
  • Facebook collection: 8779 suspicious Facebook profiles were located by snowball crawling from known Sybils and searching profile images with Google Search by Image.Profiles were considered suspicious when they had at least two profile images and at least 90% were available on the web.
  • Ground truth: Ground truth was established from banned or security-verified accounts, while suspicious unbanned Facebook profiles were excluded from the final labeled statistics.Renren supplied confirmed Sybil and legitimate profiles directly; Facebook Sybils were monitored until accounts became inaccessible.
  • Limitations: The Facebook dataset covers only public profiles, so its characteristics may not represent private legitimate or Sybil accounts.The paper identifies this as a limitation shared by studies based on crawled OSN data.

4 User Study Results

The study finds that experts generally identify Sybil profiles more accurately than turkers, while aggregated votes reduce classification errors. Testers also use inconsistent reasons, suggesting that detecting Sybils requires diverse profile information rather than a single reliable feature.

  • Demographics: Most testers report extensive OSN experience, although Indian and Chinese turkers include larger fractions with less than two years’ experience.US experts, Chinese experts, and social science undergraduates almost uniformly report at least two years of experience.
  • Individual Accuracy: 50% of Chinese experts achieve at least 90% accuracy, whereas 50% of Chinese and Indian turkers achieve at most 65%.US and Indian experts also perform highly; US turkers and social science students fall between the strongest and weakest groups.
  • Individual Accuracy: False positives remain below 20% for 90% of testers, while false negatives exceed false positives across all test groups.Turkers generally have higher error rates than experts.
  • Accuracy of the Crowd: Aggregated votes produce uniformly low false positive rates; in the worst case, US turkers and social science students misclassify 1 of 50 legitimate profiles.False negative rates vary widely, with experts and social science students below 10% but Chinese and Indian turkers at least 50%.
  • Accuracy of the Crowd: Aggregation significantly lowers both false positive and false negative rates, but turkers remain less accurate than experts.The low aggregate false positive rates indicate that crowdsourcing does not harm legitimate social-network users.
  • Reasons for Suspicion: Chinese turkers show the greatest disagreement, with average Jaccard coefficients at most 0.4 for 50% of Sybils.Chinese experts and all three US groups have coefficients at most 0.5 for 50% of Sybils, while near-total agreement or disagreement is rare.
  • Reasons for Suspicion: Testers’ reasons vary across Sybils, with no single profile feature consistently indicating Sybil activity.The authors conclude that testers benefit from a large, diverse set of information when classifying profiles.
  • Answer Revisions: Only 28 answer revisions occurred: 16 changed incorrect answers to correct and 12 changed correct answers to incorrect.The authors attribute the low revision count partly to turkers’ incentive to complete paid tasks quickly.

5 Turker Accuracy Analysis

The study examines how demographic factors, evaluation time, fatigue, worker selection, and vote aggregation affect turker accuracy in Sybil detection. Experts are consistently strong, while turker performance varies and can improve through selection, although difficult Sybils remain.

  • 5.1 Demographic Factors: Higher education correlates with better turker Sybil detection, while OSN experience helps Chinese and US workers but not Indian workers.Gender has no meaningful effect; filtering workers with low education or limited OSN experience could improve accuracy.
  • 5.2 Temporal Factors and Survey Fatigue: Absolute profile evaluation time does not reliably indicate accuracy: Chinese experts are faster yet more accurate than Chinese turkers, whereas Facebook experts are slower and more accurate.Chinese experts average one profile every 23 seconds.
  • 5.2 Temporal Factors and Survey Fatigue: Turkers speed up during surveys, and their accuracy decreases over time, indicating survey fatigue; experts speed up without losing accuracy.The reported increase around Chinese turker profile 10 is described as a nonsignificant statistical anomaly.
  • 5.3 Aggregating Turker Votes: After 4 votes, additional turker votes yield diminishing false-positive improvements, while false negatives stay flat for US turkers and worsen slightly for Chinese and Indian turkers.The result comes from simulations that progressively aggregate randomized classifications for each profile.
  • 5.3 Filtering Inaccurate Turkers: At a 70% accuracy threshold, all three turker groups achieve false negative rates ≤10%, matching experts, but higher thresholds leave too few workers for coverage.Selection simulates pre-screening workers before they classify unknown suspicious accounts.
  • 5.4 Profile Difficulty: Experts classify most Sybils that turkers miss, although a few rare “stealth” Sybils evade both groups.High-accuracy turkers still saw 97% of the Renren and Facebook US Sybils in the study datasets.

6 A Practical System

The proposed practical system combines automated filtering with crowdsourced validation to scale Sybil detection. Simulations indicate that a two-layer design can achieve very low error rates with limited votes, while deployment must address calibration and privacy.

  • 6.1 System Design and Scalability: The hierarchical system uses a filtering layer to find suspicious profiles and a crowdsourcing layer to obtain human legitimacy classifications.The design targets social networks with hundreds of millions of users by focusing human effort on suspicious accounts.
  • 6.1 System Design and Scalability: Majority voting and continuous turker selection improve group accuracy by aggregating workers and removing workers whose ground-truth performance declines.Ground-truth profiles are mixed into ongoing work to detect quality changes and undercover malicious testers.
  • 6.2 System Simulations and Accuracy: The simulations evaluate 2,000 suspicious profiles using random turkers whose correctness probabilities come from the user study, excluding turkers below 60% accuracy.The simulated set contains 1,000 Sybil and 1,000 legitimate profiles.
  • 6.2 System Simulations and Accuracy: The two-layer configuration with R = [0.2, 0.5] and T = 90% reaches a 0.7% false negative rate with 6 average votes per profile.These parameters were selected for the remainder of the analysis.
  • 6.2 System Simulations and Accuracy: With an average of 6 votes per profile, the system achieves false positive and false negative rates both below 1%.The one-layer design is cheaper but incurs more false negatives; the two-layer design provides superior false-negative results.
  • 6.3 Cost and Accuracy Tradeoffs: Reducing false positives below 0.1% requires two additional votes per turker, increasing costs by 33% while reducing false positives by an order of magnitude.The tradeoff is evaluated by varying the target false-positive rate in simulation.
  • 6.4 Privacy and Calibration: System parameters may not suit every user population, and real deployment must handle privacy restrictions for users with strict settings.The authors note that periodic recalibration may be needed, but lack sufficient data to estimate when it should occur.
  • 6.3 Cost and Accuracy Tradeoffs: The estimated workload is 50 full-time turkers for Tuenti’s reports and 500 for a tenfold larger OSN workload.Tuenti averages 12,000 user reports per day; the reported moderation cost is $2,240 per day.

7 Related Work

Prior crowdsourcing research studies worker populations, pricing, and methods for improving unreliable-worker accuracy. Related work also documents malicious uses of crowdsourcing and existing social-network moderation practices.

  • Studies of Amazon Mechanical Turk examine worker demographics, task pricing, and the advantages and disadvantages of using MTurk for user studies.
  • Majority voting is common for improving crowdsourced accuracy, while pre-screening filters unreliable workers and tournament algorithms address difficult tasks.Majority voting can be vulnerable to collusion attacks by malicious turkers.
  • Crowdsourcing has been abused through malicious tasks involving social spam, search-engine optimization, fake reviews, and malware installation.

8 Conclusion and Open Questions

The paper proposes crowdsourced Sybil detection as a scalable and accurate response to increasingly realistic fake accounts, while identifying ground-truth, deployment, and attacker-countermeasure challenges.

  • Crowdsourced Sybil detection addresses increasingly realistic fake accounts that challenge online social-network security and existing ad hoc defenses.Social networks currently rely on measures such as employee manual inspection.
  • A hierarchical two-tiered system can be accurate and scalable in total costs when experts calibrate ground-truth filters and separate accurate turkers from low-accuracy workers.The calibration process can eliminate low-accuracy turkers and identify the most accurate workers.
  • Open Questions: Evaluation accuracy is a lower bound because the ground-truth data may omit additional Sybils that existing Facebook or Renren mechanisms failed to catch.Those undetected Sybils could potentially be caught by the crowdsourced system.
  • Open Questions: Effective deployment remains open, with the system envisioned as a complement to content-filtering and statistical models rather than a fully resolved deployment solution.The paper also discusses using accurate turker outputs to teach automated tools and using social-network users as crowdworkers to reduce costs.
  • Open Questions: Attackers may poison worker results, refresh profiles to evade detection, or infiltrate the system, and handling such undercover attackers remains an open question.The proposed countermeasures include randomly mixing ground-truth profiles with test profiles and periodically refreshing ground truth.
Loading 1205.3856v2…