Source-linked AI summary
Social Information Processing in Social News Aggregation
Kristina Lerman
TL;DR
The paper asks how social media can support document recommendation and rating through user-created networks and collective evaluation. Using Digg data and mathematical models, it studies social recommendation, collaborative voting, and user-rank evolution. The models qualitatively agree with observed story-vote and user-rank behavior, while the modeling scope excludes some browsing modalities and user-behavior variation.
Problem
The paper examines how Digg can address document recommendation and rating through social networks and independent user evaluations.
Method
The paper tracks Digg stories and users and develops mathematical models of collaborative voting, story promotion, and user-rank evolution.
Results
The model solutions qualitatively agree with the evolution of votes received by actual Digg stories and with observed user-rank behavior.
Takeaways & Limitations
Personal social networks form the basis for an effective social recommendation system, while mathematical modeling can help explore collaborative-system interface designs.
Takeaways & Limitations
The models consider only default “Newly popular” browsing and omit other browsing modalities and variance in user behavior needed for quantitative reproduction or substantial disagreement analysis.
Abstract
from arXiv · showhide
The rise of the social media sites, such as blogs, wikis, Digg and Flickr among others, underscores the transformation of the Web to a participatory medium in which users are collaboratively creating, evaluating and distributing information. The innovations introduced by social media has lead to a new paradigm for interacting with information, what we call 'social information processing'. In this paper, we study how social news aggregator Digg exploits social information processing to solve the problems of document recommendation and rating. First, we show, by tracking stories over time, that social networks play an important role in document recommendation. The second contribution of this paper consists of two mathematical models. The first model describes how collaborative rating and promotion of stories emerges from the independent decisions made by many users. The second model describes how a user's influence, the number of promoted stories and the user's social network, changes in time. We find qualitative agreement between predictions of the model and user data gathered from Digg.
1 Introduction
Social media turns users into participants who create, evaluate, distribute, and socially organize information. The paper studies Digg's use of this social information processing for recommendation, rating, and user-rank dynamics.
- Social information processing: Social media users create content, annotate it, evaluate it, and form networks with users sharing similar interests.These activities add social networks, annotations, and ratings as metadata for collaborative problem solving.
- Digg as a case study: Digg lets users submit stories, vote on them, form friendships, and track friends' activities; it promotes selected stories to front pages.The paper examines these functions as mechanisms for recommendation and rating.
- Social recommendation: Social networks provide social filtering that helps users discover stories their contacts found interesting, avoiding reliance on explicit product ratings.The paper contrasts this approach with collaborative filtering, whose users may resist rating products.
- Collaborative evaluation: The paper frames collaborative evaluation as a second information-processing problem, analogous to using independently created Web links to assess page importance.Digg and Reddit use independent user opinions to evaluate news stories.
- Modeling contributions: Two mathematical models describe collaborative rating and user-rank evolution, and their solutions predict observed Digg voting and ranking behavior.The models address story promotion dynamics and changes in users' influence, promoted stories, and social networks.
2 Anatomy of Digg
Digg combines user-submitted stories, voting, social activity tracking, and front-page promotion to organize news. Its interface exposes both story activity and friends' recommendations while ranking users by platform activity and promotion history.
- Story flow: Submitted stories enter an upcoming queue, while stories receiving enough votes are promoted to the front page, which most daily visitors read.Front-page promotion therefore greatly increases a story's visibility.
- User moderation: Users can digg stories to vote and save them in their history, or bury stories identified as spam, duplicates, or inappropriate material.Buried stories are removed only after enough users bury them; burying does not reduce their rating.
- Emergent selection: Digg's front page showcases selected content through an emergent outcome of many users' evaluations, similar to featured content on Flickr and Delicious.Flickr uses views, comments, and favorites for Explore selection, while Delicious showcases recently tagged popular pages.
- Social filtering: The Friends interface summarizes friends' recent submissions, comments, and likes, marking those stories for easy discovery.This supports social filtering instead of requiring users to search actively for new content.
- User ranking: Before February 2007, Digg ranked users mainly by the number of their stories promoted to the front page, with activity breaking ties.The ranked list was accessible through the Top Users link and may have encouraged competition.
3 Dynamics of collaborative rating
The study tracks Digg stories, voters, and users to examine how stories accumulate votes and reach the front page. It finds slow queue-stage voting, faster front-page accumulation, and a relationship between submitter rank and story outcomes.
- Data collection: The study scraped front-page, upcoming-story, and top-user data, including story metadata, voter lists, activity measures, ranks, and social-network ties.Front-page and upcoming-story wrappers ran hourly for a week, while the top-users wrapper gathered weekly snapshots.
- Observed dynamics: Of 2,858 stories submitted by 1,570 users, 98 stories by 60 users reached the front page.Stories were followed from approximately one day of submissions across several subsequent days.
- Observed dynamics: Stories accumulated votes slowly in the upcoming queue and much faster after front-page promotion.Figure 2(a) tracks selected stories over four days and marks their transitions to the front page.
- Submitter rank: Top-ranked users did not submit stories receiving the most votes, although the top 3% of users produced 35% of more than 15,000 front-page stories.Their submitted stories had average interestingness 600, nearly half the average for stories from low-rated users.
3.1 Social networks and social filtering
Digg’s social networks help users discover and promote stories through friends’ submissions and voting activity. Evidence shows network size correlates with submitter success, while social filtering increases visibility for otherwise poorly connected stories.
- Users employ the Friends interface to filter Digg’s many submissions and discover new interesting stories.
- Users with larger social networks, especially more reverse friends, have higher success rates getting stories promoted.Success rate is the fraction of submitted stories promoted to the front page; analysis included users who submitted at least 50 stories.
- Users digg stories their friends submit: 99 of 195 front-page stories were submitted by users with more than 20 reverse friends, and all but two were dugg by the submitter’s reverse friends.
- Users digg stories their friends submit: The chance of observing reverse-friend votes randomly was P = 0.005, falling to P = 0.003 when considering the first 25 voters.Including two stories with no reverse-friend diggs raises the average probability to P = 0.023, which remains significant.
- Users digg stories their friends digg: For 96 stories from unknown users, visibility through the Friends interface rose from 26 at submission to 75 after five additional voters and all 96 after 25 voters.Almost half of the stories visible after 25 votes were dugg by friends; chance probabilities exceeded the 0.05 significance level for m = 26 through m = 46.
- Social networks can create a tyranny of the minority, yet active users also help surface stories that might otherwise be buried among new submissions.
3.2 Mathematical model of collaborative rating
The model explains collaborative rating by combining story interestingness with visibility through Digg’s queues, front page, and Friends interface. Its predictions qualitatively match observed voting patterns, while highlighting approximations that affect individual stories.
- Model formulation: The model predicts vote accumulation from a story’s interestingness coefficient r and its visibility across Digg’s front page, upcoming queue, and Friends interface.Interestingness determines the probability of a positive vote once a story is seen.
- Visibility through social networks: A submitter’s reverse-friend network exposes the story at rate a = S/24, with visibility through Friends limited to the first 48 hours.The model treats reverse friends as users watching the submitter’s activities.
- Visibility through social networks: As users vote, the story reaches additional audiences through voters’ friends; the combined network grows on average as S_m = 112.0*log(m) + 47.0.The model limits this Friends-interface visibility to 48 hours.
- Model solutions: Without social-network visibility, the model gives a maximum of 43 votes on the upcoming pages, approximately the promotion threshold, so additional effects are needed for front-page promotion.This result follows from the asymptotic solution m(T →∞) ≈ 42r + 1.
- Model solutions: For S = 80, sufficiently interesting stories are promoted, while larger submitter networks lower the interestingness required; model predictions qualitatively agree with six observed stories.A story with r = 0.1 and S = 400 is also predicted to reach the front page.
- Model limitations: The comparison is limited because the model assumes identical growth rates for voters’ combined networks and a common interestingness parameter across audiences.These approximations help explain discrepancies between observed promotion order and model predictions.
4 Dynamics of user rank
Digg’s user rank dynamics are modeled through front-page stories and social-network growth, with model solutions qualitatively reproducing observed user trajectories.
- Digg ranked users primarily by the number of their stories promoted to the front page, with activity breaking ties.Rank 1 was the highest standing, and the Top Users list provided public prestige.
- The model uses front-page story count F as a proxy for rank because Digg’s exact ranking formula is unknown.Observed rank follows a power law with exponent -1: rank ∝1/F.
- A user’s promoted-story count depends on submission rate M and success rate, which is linearly correlated with social-network size S.The success rate is the fraction of newly submitted stories promoted to the front page.
- The fitted parameters are c = 0.002 for success-rate dependence on S, a = 0.03 for g(F) = aF, and b = 1.0 for promoted-story effects.These values were estimated from trends and linear fits in the observed data.
- The rank-dynamics model represents social-network growth as contributions from user rank and newly promoted stories.The model estimates g(F) for rank-dependent growth and uses b for growth associated with newly promoted stories.
- Model solutions qualitatively reproduce rank and social-network evolution: active users grow in both, while inactive users’ rank stagnates as their networks grow more slowly.The model links rank change to both new submissions and social-network size.
5 Limitations of modeling
The models remain tractable through simplified browsing, behavioral, and story assumptions, but these abstractions limit quantitative fidelity and leave possible factors unmodeled.
- The model omits several Digg browsing modalities, including some Friends-interface activities and other ways users may encounter stories.The authors note that omitted terms could be added if data show those browsing options are popular.
- The authors have not proved their conjecture that users can influence social-network growth by reciprocating friend requests.The conjecture concerns implicit social etiquette around friend requests.
- The collective-voting model describes average vote dynamics across many similar stories rather than the changing rating of a specific story.This abstraction simplifies story-level variation.
- User behavior is represented with single-valued parameters, assuming a constant visit rate and identical story interestingness across users.The rank model likewise uses characteristic mean parameter values rather than distributions.
- The abstractions may remove important factors, so quantitative reproduction or substantial model-data disagreement would require richer browsing and behavioral-variance terms.The authors claim the simple models capture the most salient features while identifying these extensions as future work.
6 Previous research
Previous research covers social-media analysis, collaborative filtering, social navigation, and mathematical models of collective behavior, providing context for Digg’s analysis.
- Research on social media has examined blogs for public-opinion trends and tagging for organizing information.The passage identifies the blogosphere as the most mature area among these research directions.
- Collaborative filtering recommends documents or products by comparing ratings from users with similar interests.Amazon and Netflix are cited as examples of services using this technology.
- Social navigation guides users through information traces left by previous users and is closely linked to collaborative filtering.The passage contrasts this related concept with the paper’s research.
- The paper adapts mathematical analysis of multi-agent collective behavior, previously applied to groups of robots, to model Digg users.Comparison with real Digg data is presented as evidence that online collective behavior can be mathematically analyzed.
7 Conclusion
The paper argues that social information processing on Digg supports recommendation and collaborative evaluation, and that mathematical models can reproduce important user and voting dynamics qualitatively.
- Social information processing combines user participation, personal social networks, and collaborative evaluation to address document recommendation and quality assessment.Users create, evaluate, and disseminate information rather than merely consuming it.
- Personal social networks form the basis of an effective social recommendation system by suggesting stories that friends found interesting.The paper also studies how distributed opinions produce Digg’s front page by consensus.
- Mathematical models qualitatively agree with both the evolution of votes on actual stories and changes in users’ rank over time.The rank analysis incorporates users’ new submissions and social-network growth.
- Mathematical modeling can help explore global consequences of alternative story-promotion algorithms before implementation.The paper frames this as a way to investigate interface design choices such as timeliness or submitter popularity.
- Digg illustrates that activities of other users can be exploited to solve difficult information-processing problems.The authors expect related progress in personalization, search, and discovery.