Source-linked AI summary

Tracking User Attention in Collaborative Tagging Communities

Elizeu Santos-Neto, Matei Ripeanu, Adriana Iamnitchi

arXiv:0705.1013v4cs.DLcs.CY

TL;DR

As tagging communities grow, user-generated metadata becomes harder to use for retrieval, motivating methods that improve navigability. The paper characterizes CiteULike and Bibsonomy through activity and interest-sharing analyses, finding preliminary evidence that neighbors’ item sets can predict future user activity and support recommendation systems.

  • Problem

    As tagging communities grow, retrieval from user-generated tags becomes less efficient, creating a need to understand whether usage patterns can improve navigability.

  • Method

    The paper characterizes CiteULike and Bibsonomy, analyzes tagging activity and user-interest structure, and evaluates neighbor item sets for predicting future attention.

  • Results

    Preliminary evidence suggests that a user’s activity can be predicted from the union of neighboring users’ item sets in an interest-sharing graph.

  • Takeaways & Limitations

    Interest-sharing structure may support efficient online recommendation systems and reduce the diversity of items exposed to users.

  • Takeaways & Limitations

    The activity-distribution fit does not establish that user-library or vocabulary diversity is a biodiversity-like phenomenon, and the interest-graph analysis remains preliminary across metrics and thresholds.

Abstract

from arXiv · show

Collaborative tagging has recently attracted the attention of both industry and academia due to the popularity of content-sharing systems such as CiteULike, del.icio.us, and Flickr. These systems give users the opportunity to add data items and to attach their own metadata (or tags) to stored data. The result is an effective content management tool for individual users. Recent studies, however, suggest that, as tagging communities grow, the added content and the metadata become harder to manage due to an ease in content diversity. Thus, mechanisms that cope with increase of diversity are fundamental to improve the scalability and usability of collaborative tagging systems. This paper analyzes whether usage patterns can be harnessed to improve navigability in a growing knowledge space. To this end, it presents a characterization of two collaborative tagging communities that target scientific literature: CiteULike and Bibsonomy. We explore three main directions: First, we analyze the tagging activity distribution across the user population. Second, we define new metrics for similarity in user interest and use these metrics to uncover the structure of the tagging communities we study. The structure we uncover suggests a clear segmentation of interests into a large number of individuals with unique preferences and a core set of users with interspersed interests. Finally, we offer preliminary results that demonstrate that the interest-based structure of the tagging community can be used to facilitate content usage as communities scale.

1. INTRODUCTION

As collaborative tagging communities grow, user-generated tags become less effective for information retrieval, motivating usage-pattern analysis to improve navigability. The paper characterizes CiteULike and Bibsonomy as preliminary steps toward modeling user interests and contextualized attention.

  • Growing tagging communities make information retrieval based on user-generated tags less efficient, increasing the need to track user attention.
  • The study investigates whether usage patterns can present relevant, contextualized information and reduce navigability problems caused by informational overload.
  • The paper characterizes CiteULike and Bibsonomy to support a model of user interests based on tagging activity.
  • The planned analysis covers tagging-activity distributions, shared-interest structure, and preliminary use of that structure to improve navigability.

2. RELATED WORK

Prior work studies explicit and implicit preference elicitation, usage tracking, tagging behavior, and recommendation, but this paper focuses on dynamic metrics for shared user interests in tagging communities.

  • Implicit preference techniques infer user interests from activity, while explicit techniques rely on users’ direct reports of preferences and interests.
  • In tagging communities, tags provide explicit metadata, while item volume, tagging frequency, tagged-item counts, and vocabulary size provide implicit activity information.
  • Prior studies report low correlation between users’ bookmark-list sizes and tag counts and use an urn model to describe tag-frequency evolution.
  • The paper differs from structural tagging studies by using dynamic metrics to define shared user interests and from preference-based recommendation by using implicit profiles and entropy.
  • Social-tagging efficiency decreases as communities grow because tags become less descriptive, making items and useful retrieval tags harder to find.

3. BACKGROUND

A collaborative tagging community lets users assign uncontrolled-vocabulary tags to shared items, creating user, item, tag, and assignment relationships. In CiteULike and Bibsonomy, shared libraries and tags support collaborative discovery of scientific publications.

  • Users interact by searching for items, adding items, or tagging existing items; each tagging action is a tag assignment.
  • In CiteULike and Bibsonomy, users share libraries of scientific publications and books, and can inspect or reuse other users’ tags.
  • Users can add publications by browsing scientific-literature portals or searching items already present in other users’ libraries.
  • Although tagging may be self-centered, shared tags can help other users find content of interest.
  • A collaborative tagging community consists of users, items, a tagging vocabulary, and tag assignments represented as C=(U,I,T,A).
  • Each user is associated with previously tagged items and used tags, while items and tags are associated with their participating users and linked entities.

4. DATA SETS AND DATA CLEANING

The study analyzes administrator-provided snapshots of CiteULike and Bibsonomy scientific-literature data after cleaning anomalous or non-interactive activity. Cleaning removed automated behavior and users relying only on automatically assigned tags.

  • Both datasets come from CiteULike and Bibsonomy, systems designed to organize research publications and support citation-record import and export.
  • The data are global system snapshots bounded by trace timestamps, with Bibsonomy’s analysis restricted to its scientific-literature dataset.
  • CiteULike’s most popular tags, “bibtex-import” and “no-tag,” indicate substantial citation conversion use and limited tagging when items are posted.
  • A CiteULike user posted and tagged more than 3,000 items in approximately 5 minutes, behavior attributed to an automatic mechanism.
  • The analysis targets users who interactively bookmark and share articles rather than automated or citation-conversion activity.

5. TAGGING ACTIVITY

Tagging activity is highly heterogeneous across CiteULike and Bibsonomy: activity spans multiple orders of magnitude, follows strong cross-metric correlations, and is better modeled by a Hoerl than a Zipf-like distribution.

  • Users are heterogeneous in library activity, with library sizes ranked separately for CiteULike and Bibsonomy.Figure 2 plots users in decreasing order of library size and confirms heterogeneous activity intensity.
  • Library size and vocabulary size are strongly correlated in CiteULike (R2 = 0.98, n = 5954) and positively correlated in Bibsonomy (R2 = 0.80, n = 654).This differs from reported del.icio.us behavior, where tag suggestions may limit vocabulary growth; the effect of tagging recommendations requires further investigation.
  • Tagging assignments and library size are strongly correlated in both communities (R2 above 0.97), while assignment–vocabulary correlation is stronger in CiteULike (R2 = 0.99) than Bibsonomy (R2 = 0.67).
  • Tagging activity distributions are better fit by a Hoerl model than by a Zipf-like distribution.The Hoerl parameters are determined by curve fitting for each observed ranking distribution.
  • Although Hoerl fits activity distributions, it does not establish that user-library or vocabulary diversity is analogous to geographic biodiversity.The model remains useful for studying user diversity in collaborative tagging systems.
  • Activity intensity spans multiple orders of magnitude across users in the studied communities.The section evaluates item counts, tagging assignments, and vocabulary size as activity metrics.

6. EVALUATING USER SIMILARITY

The study evaluates interest-sharing graphs built from shared items or tags to determine whether tagging communities contain distinct user-interest groups. Across similarity definitions, the graphs reveal many isolated users alongside larger components, with threshold choice and graph direction shaping the resulting segmentation.

  • Interest-sharing structure: The interest-sharing graph connects users who share items or tags, using thresholded similarity based on item libraries, tag vocabularies, or directed item-library overlap.The directed definition introduces edge direction to examine the role of users with large libraries.
  • Interest-sharing structure: 2,672 CiteULike users (44.87%) remain isolated even when users are connected by sharing one item, indicating many users have individual preferences.The analysis uses the lowest sharing threshold, t = one item, as evidence for this pattern.
  • Threshold effects: As the similarity threshold rises, connected components initially increase as large groups split into similarity-based islands, then decrease when stricter definitions create more isolated nodes.The threshold at which the trend reverses differs across similarity definitions.
  • Threshold effects: The User-Item graph loses components faster than the Directed User-Item graph, showing that asymmetric shared-interest edges alter component structure.The comparison concerns how quickly the number of components decreases as the threshold increases.
  • Interest-sharing structure: User-Item and Directed User-Item graphs contain more isolated nodes than User-Tag graphs, suggesting vocabulary overlap exceeds library overlap.The paper proposes that users may browse and consume items from others without adding them to their own libraries.
  • Interest-sharing structure: Across the metrics, the graphs generally contain one giant component, several tiny components, and many isolated nodes, supporting user segmentation by manifested interest.The study uses component sizes and largest-component sizes across thresholds to characterize this structure and motivate recommendation mechanisms.

7. IMPROVING NAVIGABILITY

As CiteULike grows, item diversity and navigation difficulty increase; the paper evaluates whether interest-sharing neighborhoods can reduce this entropy and predict future attention.

  • CiteULike item diversity grows over time, increasing the filtering burden users face when searching for relevant items.Entropy rises with item-set size or a more uniform popularity distribution, and higher entropy makes relevant items harder to find.
  • Using users’ interest-sharing neighborhoods, the system can present lower-entropy item sets that may improve navigability.The neighborhood is formed from users connected through shared item or tag interests.
  • Neighborhood entropy is one half to one tenth of total dataset entropy, and 24% to 400% lower than comparable random constructions except at the 1% threshold.The comparison uses interest-sharing graphs across analyzed thresholds; the 1% threshold lacks sufficient discrimination.
  • Neighbor item sets predict future attention with hit rates from 20% at one-hour granularity to 5% at one-month granularity.The predictive effectiveness decreases as the forecasting interval becomes longer.
  • The findings support using interest-sharing graphs as a basis for recommendation systems, but the analysis remains preliminary across similarity metrics and thresholds.The authors identify further exploration of diverse graphs as incomplete.

8. CONCLUSIONS & FUTURE WORK

The study characterizes CiteULike and Bibsonomy, finding heterogeneous activity and segmented interest structures that support preliminary approaches to reducing content diversity and predicting user activity.

  • A few active users contribute many tag assignments and maintain large item libraries and tag vocabularies, while most users remain modestly active.Users with large libraries tend to have large tag vocabularies, unlike the uncorrelated pattern previously observed in del.icio.us.
  • Interest-sharing graphs contain many isolated users with unique preferences, while directed connectivity roughly doubles the size of the main connected component.The graphs also reveal numerous small interest sub-communities that are completely separated from one another.
  • Neighborhood item sets have lower entropy than broader community and comparison item sets, indicating that interest structure can reduce the diversity of exposed content.The comparison includes the full item set, the main connected component, and random item sets of similar sizes.
  • Preliminary evidence suggests that a user’s activity can be predicted from the union of neighboring users’ item sets.The authors conjecture that this property could support efficient online recommendation systems for tagging communities.
  • The work opens further questions about information propagation, evolving user attention, malicious behavior, and the formation of user-similarity graphs.These directions include detecting tagging misbehavior and modeling how user-interest structures evolve over time.
Loading 0705.1013v4…