Source-linked AI summary

An Empirical Investigation of Personalization Factors on TikTok

Maximilian Boeker, Aleksandra Urman

arXiv:2201.12271v1cs.HCcs.CYcs.SI

TL;DR

TikTok’s recommendation algorithm is central to content distribution, yet empirical evidence about how user actions and characteristics shape recommendations remains limited. Using sock-puppet auditing with a custom web-based bot, the study tests language, location, following, liking, and video-watching effects, finding that all tested factors influence recommendations, with following strongest among the reported factors.

  • Problem

    Empirical evidence on how user actions and characteristics influence TikTok’s important recommendation system remains limited.

  • Method

    The study uses a controlled sock-puppet audit with a custom web-based bot to test language, location, following, liking, and video-watching effects on TikTok recommendations.

  • Results

    All tested factors affect recommended content, with following having the strongest influence among the examined factors.

  • Takeaways & Limitations

    The findings indicate that TikTok users have some control over their feeds while algorithmic personalization may contribute to filter bubbles and problematic-content risks.

  • Takeaways & Limitations

    The analysis is not exhaustive, and technical failures led the authors to exclude some runs from analysis.

Abstract

from arXiv · show

TikTok currently is the fastest growing social media platform with over 1 billion active monthly users of which the majority is from generation Z. Arguably, its most important success driver is its recommendation system. Despite the importance of TikTok's algorithm to the platform's success and content distribution, little work has been done on the empirical analysis of the algorithm. Our work lays the foundation to fill this research gap. Using a sock-puppet audit methodology with a custom algorithm developed by us, we tested and analysed the effect of the language and location used to access TikTok, follow- and like-feature, as well as how the recommended content changes as a user watches certain posts longer than others. We provide evidence that all the tested factors influence the content recommended to TikTok users. Further, we identified that the follow-feature has the strongest influence, followed by the like-feature and video view rate. We also discuss the implications of our findings in the context of the formation of filter bubbles on TikTok and the proliferation of problematic content.

1 INTRODUCTION

TikTok’s algorithmic “For You” feed is central to the platform’s rapid growth, yet its user-facing recommendation logic remains poorly understood. This study addresses that gap by auditing how selected user characteristics and actions shape recommendations.

  • 1 INTRODUCTION: TikTok’s algorithmically driven “For You” feed is a major success factor, but its recommendation system remains a black box requiring empirical audit.The platform’s popularity and concentration of younger users heighten interest in how its recommendation system distributes content.
  • 1 INTRODUCTION: The study develops a user-centric audit of how user actions and characteristics affect TikTok recommendations.It focuses on language, location, liking, following, and video-watching behavior.
  • 1 INTRODUCTION: The methodology can be applied repeatedly to trace recommendation-system changes and to examine similar feeds such as YouTube Shorts or Instagram Reels.This extends the approach beyond a single observation period or platform.

2 RELATED WORK

Prior work has examined recommendation-system auditing and TikTok’s technical or supply-side dynamics, but scientific evidence about how users’ actions and characteristics shape their personalized feeds remains limited. The paper positions its hypotheses and sock-puppet audit within this demand-side gap.

  • 2.1 Auditing Recommendation Systems: Algorithm audits can uncover personalization problems including bias, filter bubbles, and unequal access to critical information.The literature distinguishes code, noninvasive user, scraping, sock-puppet, and collaborative audits.
  • 2.2 TikTok-focused research: Existing TikTok research has studied users, platform characteristics, recommendation mechanics, and factors affecting whether videos reach larger audiences.Technical accounts describe tagging, candidate retrieval and ranking, and staged distribution through increasingly broad user buckets.
  • 2.2 TikTok-focused research: This study addresses the underexamined demand side by testing how users’ actions and characteristics affect their “For You” feeds.Earlier demand-side analysis was largely journalistic rather than scientific.
  • 2.2 TikTok-focused research: The hypotheses predict feed divergence over time when paired users receive different interactions, stronger effects for some personalization factors, and language- and location-specific recommendations.The tested factors include language, location, likes, follows, and longer viewing.
  • 2.2 TikTok-focused research: The study also considers whether increasingly personalized feeds become more niche and less popular, while noting that users here means content consumers rather than creators.This expectation is framed as a hypothesis rather than an established result.

3 METHODOLOGY

The study uses controlled sock-puppet experiments to isolate personalization factors, compare active and control users, and analyze resulting feed differences. Its pipeline combines scripted browsing and interaction with several content-similarity and popularity analyses under stated ethical constraints.

  • 3 METHODOLOGY: The methodology uses a custom bot and controlled environments to isolate tested personalization factors while mimicking realistic TikTok user behavior.Incognito ChromeDriver sessions reduce noise from cookies and browsing history.
  • 3.1 Data Collection: The bot logs in, scrolls through predefined feed batches, performs scripted actions, extracts post metadata and request data, and stores the results.TikTok automatically preloaded about 30 posts per access, which the study treated as one batch.
  • 3.1 Data Collection: Dedicated proxies, manually created accounts, incognito mode, and paired execution controlled location, identity, and machine-related differences.Each test user received a specific IP address, and paired users ran on the same local machine.
  • 3.2 Experiment Overview: Each scenario paired an active user performing one tested action with a control user that only scrolled, allowing feed differences to be attributed to the manipulated factor beyond expected noise.Scenarios covered following, liking, longer viewing, language, and location, with about 20 runs per scenario.
  • 3.3 Data Analysis: The analysis measured feed overlap and divergence with Jaccard indices and regression, post popularity, recurring hashtags, sounds and creators, and hashtag semantics using a Skip-Gram model.Common hashtags such as “#fyp” were removed to avoid obscuring similarity patterns.
  • 3.4 Ethical considerations: The audit was conducted for academic, noncommercial purposes with few agents and was classified as exempt from the university’s human-subjects ethical review.The authors characterize the interactions as marginal, non-intrusive, and harmless.

4 EXPERIMENTS

Across 39 completed scenarios, the audit controlled for feed divergence caused by noise and tested how location, language, liking, following, and viewing duration affect TikTok recommendations. All tested factors influenced recommendations, with following strongest, followed by video view rate and liking.

  • 4 EXPERIMENTS: The experiments comprised 39 successfully completed scenarios covering 30’436 posts, 34’905 hashtags, 21’278 content creators, and 20’302 sounds.Experiments ran between late June and mid-August 2021.
  • 4.1 Controlling Against Noise: Control scenarios measured default feed divergence, while matched experimental controls and adjustments for recurring drops helped distinguish personalization from noise.Control feed differences fluctuated substantially, including presumed software-update drops.
  • 4.2 Language and Location: Location strongly affected recommended posts, whereas language had a weaker influence; US users generally saw more English content regardless of language settings, except French feeds were more similar to one another.The location effect persisted when users switched locations during testing.
  • 4.3 Like-Feature: Liking posts increased feed divergence relative to controls, with stronger effects for persona-, creator-, or sound-based selections than for arbitrary likes.Liked creators and sounds reappeared more often, while common broad hashtags produced limited divergence.
  • 4.4 Follow-Feature: Following a content creator made active-user feeds more similar over time and increased exposure to followed or similar creators.In scenario 28, hashtag similarity increased faster for the active user than the control user, 21% > 18%.
  • 4.5 Video View Rate: Video view rate influenced recommendations, especially when longer viewing was assigned to niche persona-based interests rather than randomly selected posts.The study also found stronger feed divergence for some shorter-viewing conditions, while overall longer viewing still exerted stronger influence.
  • 4.6 Concluding Results: Following specific content creators produced the strongest recommendation effect, ahead of video view rate and liking among the tested factors.The influence of video view rate was only marginally higher than that of liking.
  • 4.6 Concluding Results: Active-user feeds generally became more different from controls and more internally similar, but post popularity did not consistently decline faster; language and location also affected recommendations.The authors reject hypothesis 5 while considering hypothesis 4 supported.

5 DISCUSSION

The discussion interprets how user actions and characteristics shape TikTok recommendations, emphasizing user control, filter bubbles, and risks from problematic content.

  • 5 DISCUSSION: Following has the largest influence among examined factors, while video view rate is similarly important to liking action.Following is conscious and reversible, whereas users cannot “unwatch” videos, limiting control over recommendation effects.
  • 5 DISCUSSION: Video view rate may drive users toward harmful filter bubbles because lingering on problematic videos can influence recommendations without explicit approval.The authors connect this concern to extremist content, insufficient platform safeguards, and randomization in served videos.
  • 5 DISCUSSION: The authors recommend filtering problematic content and giving users more control over inferred interests and feed personalization.They specifically suggest detailed, continuously updated interest lists that users can inspect and adjust.
  • 5 DISCUSSION: Location and language analyses, together with hashtag similarity, imply filter bubbles based on individual interests and geographic context.The authors propose more serendipitous recommendations to increase diversity while preserving perceived preference fit and enjoyment.

6 CONCLUSION

The study uses sock-puppet auditing to test how TikTok user characteristics and actions affect recommendations. All tested factors influenced recommendations, with following strongest and location stronger than language.

  • 6 CONCLUSION: All tested factors affected TikTok recommendations, while following was strongest, followed by video view rate and liking; location outweighed language.The tested factors were language, location, following, liking, and watching posts longer.
  • 6 CONCLUSION: The sock-puppet audit provides a foundation for studying additional recommendation influences, filter-bubble formation, and problematic-content distribution.The analysis does not exhaustively cover other possible factors such as commenting or sharing.

A EXPERIMENTAL SCENARIO DETAILS

Table 1 organizes experimental scenarios across noise controls, language and location, liking, following, and video view rate conditions.

  • A EXPERIMENTAL SCENARIO DETAILS: Table 1 covers experimental groups for noise control, language and location, liking, following, and video view rate.Yellow-highlighted users are active users, while red-highlighted scenarios are failed runs.

B DIFFERENCE ANALYSIS RESULTS

Tables 2–4 summarize average analysis metrics for control scenarios compared with liking, following, and video-view-rate tests.

  • B DIFFERENCE ANALYSIS RESULTS: Table 2 compares average analysis metrics between control and like test scenarios.The supplied caption identifies the comparison scope but does not report metric values or a winner.
  • B DIFFERENCE ANALYSIS RESULTS: Table 3 compares average analysis metrics between control and follow test scenarios.The supplied caption identifies the comparison scope but does not report metric values or a winner.
  • B DIFFERENCE ANALYSIS RESULTS: Table 4 compares average analysis metrics between control and video-view-rate test scenarios.The supplied caption identifies the comparison scope but does not report metric values or a winner.

C ADDITIONAL FIGURES

This section presents additional figures for two test scenarios: post-metric changes in scenario 7 and hashtag similarity within users’ feeds in scenario 28.

  • Figure 7 presents changes in likes, shares, comments, and views for test scenario 7.
  • Together, the figures cover post metrics in scenario 7 and feed hashtag similarity in scenario 28.
  • Figure 8 presents hashtag similarity within each user’s feed for each test run in scenario 28.
Loading 2201.12271v1…