Source-linked AI summary

Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy

Liwei Deng, Jing Jiang, Zhiwei Li, Yang Wang, Guodong Long

arXiv:2608.13990v1cs.AIcs.IR

TL;DR

Short-video recommenders favor rapidly attention-grabbing content, but content depth lacks a measure for evaluation alongside engagement. This paper introduces CDS and SCOPE-Bench, finding that all 13 evaluated systems’ content-depth performance remains close to random recommendation.

  • Problem

    Short-video recommenders favor rapidly attention-grabbing content, but the field lacks a measure for evaluating content depth alongside engagement.

  • Method

    The paper introduces CDS, a seven-level content-depth metric, and SCOPE-Bench, a benchmark with human-aligned annotations for 150K videos.

  • Results

    All 13 representative recommender systems’ content-depth performance remains close to random recommendation, so stronger engagement does not translate into greater recommended depth.

  • Takeaways & Limitations

    Content depth becomes an explicit dimension for assessing recommendation lists beyond conventional engagement-oriented performance.

  • Takeaways & Limitations

    SCOPE-Bench has incomplete modality coverage because captions, visual features, and ASR transcripts are not uniformly available for all videos.

Abstract

from arXiv · show

Driven by the attention economy, short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds. These systems inherently favor shallow-content videos that are effective at attracting immediate attention. However, growing evidence suggests that prolonged exposure to such content may negatively affect users' cognitive engagement and mental well-being, raising concerns about the long-term societal impact of the short-video platform. To tackle this challenge, this paper introduces a new metric, the \textbf{Content Depth Score (CDS)}, to quantify the content depth of short videos. CDS measures the extent to which a video is expected to stimulate higher-order cognitive processes, using a seven-level scale grounded in established theories of cognitive psychology and learning. As an initial step toward this vision, we present \textbf{SCOPE-Bench}, the first benchmark for content-depth evaluation in short-video recommendation. Built upon a large-scale open-source short-video dataset, SCOPE-Bench provides CDS annotations for 150K videos, enabling systematic evaluation of RSs from a cognitive-content perspective. Leveraging SCOPE-Bench, we evaluate 13 representative RSs and reveal a consistent preference for shallow-content videos. Moreover, we find that these algorithms recommending cognitively deep content are only marginally better than random selection, highlighting a previously overlooked limitation of existing recommendation objectives. Our code and datasets are available at https://liweidengdavid.github.io/SCOPE-Bench/.

1 Introduction

Short-video recommender systems optimize engagement, favoring attention-grabbing content while giving limited attention to content depth. This paper introduces the seven-level Content Depth Score (CDS) and uses it to expose existing systems’ near-random content-depth performance.

  • Motivation: Engagement-optimized short-video RSs turn attention into revenue through signals such as watch time and clicks.Adult users spend more than one hour per day on these platforms.
  • Problem: Content depth concerns how deeply videos develop information, creating opportunities for understanding, reasoning, and reflection.The introduction frames content depth as a dimension that may be overlooked when systems prioritize attention retention.
  • Findings: 13 representative RSs perform competitively on engagement but remain low and near the random-selection baseline on content depth.Figure 1 evaluates engagement on the X axis and content depth on the Y axis, with low, medium, and high depth regions.
  • Motivation: Existing interventions mainly restrict access or usage, whereas the paper addresses the absence of a measure for short-video content depth.The introduction cites under-16 restrictions in Australia and the UK and adolescent daily-usage caps by platforms.
  • Contribution: The Content Depth Score (CDS) measures how deeply a short video presents and delivers information using a seven-level rubric grounded in cognition and learning theories.The paper proposes CDS as an explicit objective for recommender-system evaluation and optimization.

2 Related Works

Prior work assesses content through metric-based and judgment-based paradigms, while recommender systems are evaluated from system-centric and user-centric perspectives. Existing judgment-based methods offer limited conceptualizations of content depth, motivating CDS as a quantitative metric grounded in cognitive processes.

  • Content Quality Assessment: Content quality assessment distinguishes metric-based evaluation of quantifiable properties from judgment-based evaluation of open-ended or interpretive properties.Metric-based assessment includes perceptual fidelity, temporal consistency, generative realism, and cross-modal alignment.
  • Content Quality Assessment: CDS measures opportunities that video content provides for cognitive processes, defined as mental operations for acquiring, processing, storing, and using information.This definition is attributed to Sternberg, Sternberg, and Mio (2006).
  • Content Quality Assessment: LLM-as-a-Judge methods span scalar scoring, pairwise comparison, rubric-based evaluation, and fine-grained assessment, but provide only a limited conceptualization of content depth.The paper responds by systematically defining content depth and proposing CDS as the first quantitative metric.
  • Recommender-System Evaluation: Recommender-system evaluation divides into system-centric and user-centric evaluation, with system-centric assessment covering accuracy-oriented and beyond-accuracy evaluation.Accuracy-oriented evaluation uses Precision, Recall, and NDCG to assess whether relevant items are retrieved and ranked.

3 Content Depth Score Metric

The Content Depth Score (CDS) quantifies how deeply short videos develop topics into structured understanding and higher-order reasoning. It uses a seven-level rubric grounded in complementary cognitive and learning theories.

  • Content depth measures how a video develops isolated information into structured understanding through explanation, procedural demonstration, mechanism analysis, evaluation, and generalizable insights.
  • CDS assigns higher scores to videos with richer semantic information, clearer explanations, and more advanced reasoning structures that offer opportunities for higher-order cognitive processes.
  • The seven-level CDS rubric combines Dual Process Theory, Bloom’s Taxonomy, and the SOLO Taxonomy to represent progressively complex cognitive processing and understanding.System 1 defines the lowest level, while Bloom’s six cognitive levels define the remaining levels.
  • Levels 0–6 are grouped into low, medium, and high CDS categories, progressing from affective or entertaining responses to structured knowledge, application, and higher-order reasoning.The medium group covers Levels 1–3, while the high group covers Levels 4–6.

4 Benchmark for CDS Evaluation

SCOPE-Bench evaluates short-video content depth at both the individual-item and recommendation-list levels, extending assessment beyond engagement-oriented performance. It combines multimodal dataset resources with text-based CDS annotation, human-aligned scalable labeling, and the LCDS metric for top-K lists.

  • Benchmark scope: SCOPE-Bench assesses individual-video depth from the user perspective and recommendation-list depth from the platform perspective.This two-level design enables evaluation of recommender systems beyond conventional engagement-oriented performance.
  • Dataset: ShortVideo3 contributes approximately 1M chronologically ordered interactions from 10K real-world users over one week, with Full and Sampled versions.The dataset includes user-item interactions and multimodal information, but the sampling strategy for the Sampled version is not fully specified in the passage.
  • CDS annotation: The CDS protocol uses captions, category labels, and ASR transcripts as complementary textual inputs rather than raw visual and audio streams.The modality coverage is incomplete, and videos relying mainly on visual or auditory content or lacking usable transcripts receive an insufficient-information score ∅.
  • Annotation workflow: Five human evaluators establish a gold-standard subset, after which candidate LLMs are compared against human annotations and the strongest-aligned evaluator labels the larger dataset.The human annotation follows a Delphi-method process, while the selected LLM produces structured CDS outputs for scalable annotation.
  • List-level evaluation: LCDS aggregates CDS values across top-K recommendation lists using marginal-level gains and position-dependent exposure weights.Its parameter β controls peak orientation: larger β emphasizes the highest-depth positions, whereas smaller β increases sensitivity to low-CDS positions and sustained depth.

B. Content-depth-aware Recommender System Evaluation

This section presents a content-depth-aware evaluation framework that jointly assesses engagement and recommendation-list content depth, supported by reliable human CDS annotations. It defines exposure-aware depth metrics whose values range from 0 to 1.

  • The evaluation framework jointly assesses user engagement and the content depth of recommendation lists.
  • Human CDS annotations show high agreement, evaluated through score-difference tolerance and bootstrap distributions of Ordinal Krippendorff’s α.
  • E-LCDS@K measures recommendation-list content depth with rank weights wi = 1/ log2(i + 1), followed by NDCG to account for greater exposure at higher ranks.
  • Both content-depth metrics lie in [0, 1], with larger values indicating greater content depth in top-K recommendation lists.

5 Experiment

Experiments show that CDS can be applied reliably by humans and that short-video content is predominantly shallow, with depth associated more with meaningful informational structure than superficial cues. Counterfactual tests further indicate robustness to category and verbosity biases, with transcript shortening affecting higher-depth videos most but never shifting scores by more than two levels.

  • Annotation Reliability: Human evaluators exhibit high agreement when independently applying the CDS rubric before conflict resolution.This supports the reliable application of the proposed scoring rubric across evaluators.
  • CDS Distribution: Most videos fall into the low-CDS or NaN groups, indicating that short-video platforms contain a large proportion of limited-depth content.NaN cases mainly arise from missing raw videos or ASR transcripts that are too short or noisy for reliable assessment.
  • CDS Distribution: Medium- and high-CDS videos are more associated with knowledge-intensive categories, whereas low-CDS videos are more associated with entertainment-oriented categories.Examples include Finance, History, and Law for higher-CDS content, versus Dance, Comedy, and Beauty for low-CDS content.
  • CDS Correlates: Entertainment reaction and plot/dialogue signals decrease as CDS rises, while most other theory-oriented lexical categories increase.The enrichment patterns align low CDS with entertainment reactions and plot-level descriptions, and high CDS with other cognitive-content signals.
  • CDS Correlates: ASR length and category exhibit stronger associations with valid CDS than caption length, consistent with higher-depth content requiring textual space for explanations, procedures, and reasoning.Category differences similarly reflect the distinction between entertainment-oriented and knowledge-oriented content.
  • Robustness: No CDS shifts exceed two levels under transcript shortening or category perturbations, while length expansion has minimal impact and low-CDS videos remain unchanged when mapped to higher-CDS categories.Overall, the protocol primarily responds to meaningful informational and reasoning structures rather than category labels or transcript length alone.

6 SCOPE-Bench Leaderboard

SCOPE-Bench evaluates 13 representative recommendation baselines and finds that engagement strength does not translate into recommending cognitively deeper content. Across two datasets, methods remain in the Low-LCDS range and generally perform close to random recommendation.

  • Baselines: 13 baselines are evaluated, comprising three ID-based methods and ten multimodal methods.The ID-based methods are BPR, NCF, and LightGCN; the multimodal methods include VBPR, GRCN, LATTICE, BM3, FREEDOM, MGCN, LGMRec, DiffMM, REARM, and FITMM.
  • Performance: Baselines show competitive engagement performance but remain in the Low-LCDS range on content-depth metrics.Low-LCDS is defined as [0, 1/6), and Table 8 reports metrics scaled by ×100.
  • Performance: Most methods perform close to the Random baseline, indicating that stronger engagement performance does not necessarily yield deeper-content recommendations.This result is reported for both A-LCDS and E-LCDS content-depth metrics.

7 Conclusion

The paper introduces CDS as a principled metric for measuring short-video content depth and constructs SCOPE-Bench for item- and list-level evaluation. Empirical results support the framework’s interpretability, practical utility, and robustness.

  • CDS measures the content depth of short videos and provides a basis for evaluating and optimizing this dimension in existing recommender systems.
  • SCOPE-Bench is the first benchmark supporting both item-level and list-level content-depth evaluation.
  • Empirical results demonstrate the proposed evaluation framework’s interpretability, practical utility, and robustness.
Loading 2608.13990v1…