Source-linked AI summary

Uncovering Social Network Sybils in the Wild

Zhi Yang, Christo Wilson, Xiao Wang, Tingting Gao, Ben Y. Zhao, Yafei Dai

arXiv:1106.5321v1cs.SIphysics.soc-ph

TL;DR

Sybil attacks create fake identities to increase a malicious user’s power, yet large-scale measurements of Sybil behavior in online social networks have been limited. The paper uses Renren ground truth to build and deploy measurement-based detectors, then analyzes 660,000 Sybil accounts and their link creation behavior. It finds that Sybils generally do not form tight-knit communities, so existing defenses are unlikely to succeed on today’s online social networks.

  • Problem

    Large-scale studies had not characterized Sybil behavior in online social networks, leaving the assumptions of community-based defenses untested.

  • Method

    The paper uses Renren ground truth to build a measurement-based, real-time detector and analyzes Sybil graph topology using edge creation information.

  • Results

    Sybil accounts generally do not form tight-knit communities: >70% have no social edges to other Sybils, while 69% of the connected remainder form one component accidentally.

  • Takeaways & Limitations

    Existing Sybil defenses are unlikely to succeed on today’s online social networks, motivating new techniques to detect and defend against Sybil attacks.

  • Takeaways & Limitations

    Existing defenses had demonstrated efficacy on synthetic graphs, but their efficacy at detecting Sybils in the wild had not been demonstrated.

Abstract

from arXiv · show

Sybil accounts are fake identities created to unfairly increase the power or resources of a single malicious user. Researchers have long known about the existence of Sybil accounts in online communities such as file-sharing systems, but have not been able to perform large scale measurements to detect them or measure their activities. In this paper, we describe our efforts to detect, characterize and understand Sybil account activity in the Renren online social network (OSN). We use ground truth provided by Renren Inc. to build measurement based Sybil account detectors, and deploy them on Renren to detect over 100,000 Sybil accounts. We study these Sybil accounts, as well as an additional 560,000 Sybil accounts caught by Renren, and analyze their link creation behavior. Most interestingly, we find that contrary to prior conjecture, Sybil accounts in OSNs do not form tight-knit communities. Instead, they integrate into the social graph just like normal users. Using link creation timestamps, we verify that the large majority of links between Sybil accounts are created accidentally, unbeknownst to the attacker. Overall, only a very small portion of Sybil accounts are connected to other Sybils with social links. Our study shows that existing Sybil defenses are unlikely to succeed in today's OSNs, and we must design new techniques to effectively detect and defend against Sybil attacks.

1. INTRODUCTION

Sybil attacks create fake identities to increase a malicious user’s power, but large-scale evidence about their behavior in online social networks has been lacking. Using Renren ground truth and operational data, this study detects and characterizes Sybils, finding that they generally integrate into the social graph rather than forming tightly connected communities.

  • Sybil attacks let malicious users create fake identities to unfairly increase their power and influence in a target community.
  • Existing decentralized defenses look for tightly connected Sybil communities, but their assumptions had not been tested through large-scale studies of Sybils in the wild.
  • Over 100,000 Sybil accounts were identified and banned by a real-time detector deployed on Renren between August 2010 and February 2011.
  • >70% of Sybils had no social edges to other Sybils, while attackers used snowball sampling to friend popular users and integrate into the social graph.
  • Among the remaining 30% connected to other Sybils, 69%—65,000 accounts—formed one component that arose accidentally rather than through coordinated attacker efforts.
  • The analysis concludes that existing Sybil defenses are unlikely to succeed on today’s online social networks, motivating new detection and defense techniques.

2. DETECTING SYBILS

The paper develops a near-real-time Sybil detector for Renren using verified account data and behavioral features, then deploys it operationally. The approach identifies distinctive invitation, acceptance, and clustering patterns and detects approximately 100,000 Sybil accounts.

  • Ground truth and features: Renren provided 1,000 verified Sybil and 1,000 non-Sybil accounts as ground truth for identifying distinguishing behavioral attributes.Researchers searched this dataset for features that could support Sybil detection.
  • Behavioral indicators: More than 20 invitations per time interval separates Sybils from normal users at both long and short time scales.A threshold of 40 requests/hour identifies approximately 70% of Sybils with no false positives.
  • Behavioral indicators: Only 26% of outgoing friend requests from Sybils are accepted, compared with 79% for non-Sybil users.Sybils target strangers, whereas normal users typically invite people with prior relationships.
  • Behavioral indicators: Approximately 80% of Sybils accepted all incoming friend requests, although this feature can delay detection because Sybils receive few requests.Some lower acceptance rates resulted from Renren banning accounts before they responded to all outstanding requests.
  • Behavioral indicators: Average clustering coefficients were 0.0386 for non-Sybil users and 0.0006 for Sybils, using each user’s first 50 friends.Because clustering can be computed from invitations alone, it can support real-time detection without user responses.
  • Deployment and evaluation: The detector combines invitation frequency, outgoing acceptance rates, and clustering coefficient in an adaptive threshold-based scheme.Renren operated the detector continuously from August 2010 and used it to detect and ban approximately 100,000 Sybil accounts through February 2011.

3. SYBIL TOPOLOGY

The paper tests whether Sybils in Renren form detectable communities and finds that most integrate into the social graph, while the connected minority forms loose, largely accidental structures.

  • Sybil Community Detectors: Community-based Sybil detectors assume Sybils form tightly connected clusters with few links to normal users.These detectors use random walks or max-flow to locate the small edge cuts separating Sybil regions from the social graph.
  • Sybil Edges: 20% of Sybils are friends with at least one other Sybil, so most Sybils form only attack edges and integrate into the normal social graph.The degree distribution covers 667,723 Sybil accounts.
  • Sybil Communities: 7,094 Sybil components are highly fragmented, with 98% containing fewer than 10 members, although most Sybil accounts belong to one large component.The component-size distribution is heavy tailed.
  • Sybil Communities: Every Sybil component has more attack edges than Sybil edges, so none meets the requirements of existing community-based detection algorithms.Figure 7 places all components above the 45° line.
  • Sybil Edge Formation: Sybil edges in the largest component appear mostly accidental: their creation order is almost uniformly random, with intentional connections observed for only a handful of accounts.The analysis samples 1,000 Sybils from a component containing 63,541 Sybils.
  • Sybil Edge Formation: 34.5% of Sybils connect to one other Sybil and 93.7% connect to ≤10, a loose structure unlikely to justify deliberate linking.The authors attribute accidental edges to snowball-sampling tools that sometimes target Sybil nodes, which generally accept incoming requests.

4. RELATED WORK

Prior OSN research has focused on detecting spam content and aberrant behavior, whereas this study emphasizes the graph topology and invitation information of malicious accounts.

  • OSN Spam: Studies of Facebook and Twitter spam use offline heuristics to locate spam messages and aberrant behavior in coordinated campaigns.Those studies report millions of spam messages on each OSN.
  • OSN Spam Detection: Social honeypots can trap spammers, but the study suggests they are unlikely to attract spammers unless engineered to appear popular.This implication follows from the observed targeting of popular users.
  • OSN Spam Detection: Bayesian filters and SVMs work well on Twitter using public following data, but detection is less successful on Facebook and Renren because public indicators are laggy.The paper’s detector instead leverages friend-invitation information that is not publicly available.

5. CONCLUSION

The paper contributes a real-time Sybil detector and a large-scale characterization of Sybil topology on Renren. Its findings challenge community-based defenses because most Sybils do not connect to one another and observed Sybil edges are usually accidental.

  • Contributions: A computationally efficient threshold classifier catches 99% of Sybils with low false positive and negative rates.Deployed on Renren, it identified and banned over 100,000 Sybil accounts.
  • Contributions: Analysis of 660,000 Sybil accounts shows that 80% do not connect to other Sybils, while connected clusters are loose rather than tightly knit.Temporal analysis indicates that Sybil edges form accidentally rather than intentionally.
  • Contributions: The topology findings indicate that new approaches are needed beyond existing decentralized community-based Sybil detectors.The paper presents this as a consequence of Sybil behavior observed in the wild.
Loading 1106.5321v1…