Source-linked AI summary
Expecting to be HIP: Hawkes Intensity Processes for Social Media Popularity
Marian-Andrei Rizoiu, Lexing Xie, Scott Sanner, Manuel Cebrian, Honglin Yu, Pascal Van Hentenryck
TL;DR
The paper asks how external promotions relate to online popularity, a relationship that remains insufficiently quantified. It develops HIP and the endo-exo map to model and characterize that relationship, and reports 28.6% lower average percentile error than popularity-history methods when forecasting future popularity.
Problem
The paper addresses the unresolved problem of quantifying how external promotions relate to online-item popularity and popularity dynamics.
Method
HIP models expected popularity volumes by linking exogenous promotion to endogenous responses in online social networks.
Results
28.6% lower average percentile error was obtained than with state-of-the-art methods based on popularity history.
Takeaways & Limitations
The endo-exo map identifies videos with high potential to become viral and videos for which promotion is unlikely to have an effect.
Takeaways & Limitations
HIP does not capture seasonality, individual influence, or promotion sources that are unobserved or difficult to obtain.
Abstract
from arXiv · showhide
Modeling and predicting the popularity of online content is a significant problem for the practice of information dissemination, advertising, and consumption. Recent work analyzing massive datasets advances our understanding of popularity, but one major gap remains: To precisely quantify the relationship between the popularity of an online item and the external promotions it receives. This work supplies the missing link between exogenous inputs from public social media platforms, such as Twitter, and endogenous responses within the content platform, such as YouTube. We develop a novel mathematical model, the Hawkes intensity process, which can explain the complex popularity history of each video according to its type of content, network of diffusion, and sensitivity to promotion. Our model supplies a prototypical description of videos, called an endo-exo map. This map explains popularity as the result of an extrinsic factor - the amount of promotions from the outside world that the video receives, acting upon two intrinsic factors - sensitivity to promotion, and inherent virality. We use this model to forecast future popularity given promotions on a large 5-months feed of the most-tweeted videos, and found it to lower the average error by 28.6% from approaches based on popularity history. Finally, we can identify videos that have a high potential to become viral, as well as those for which promotions will have hardly any effect.
1. INTRODUCTION
The paper addresses how online popularity evolves under continuous external promotion and develops HIP to connect exogenous discussions with endogenous reactions. It uses this link to explain popularity, identify promotion-responsive videos, and forecast future popularity.
- Motivation: Popularity dynamics describe how the attention received by an online cultural item evolves over time.Understanding popularity supports information dissemination for producers and information-overload management for consumers.
- Motivation: Continuous external influence remains difficult to model using existing popularity prototypes based on shocks, relaxation, decay, or recurrence patterns.The paper specifically asks how popularity would evolve under ongoing external influence.
- Method: HIP links exogenous promotion volumes to endogenous social-network responses and produces a total popularity series by averaging over Hawkes-process event histories.It extends the Hawkes point process to describe expected event volumes rather than individual event times.
- Results: The HIP-modeled popularity series matches observed view counts closely, including videos with complex popularity lifecycles.This result supports using the model to represent varied popularity histories.
- Results: The endo-exo map combines endogenous response and exogenous sensitivity to identify videos with high viral potential and videos unlikely to benefit from promotion.High values on both dimensions indicate videos expected to go viral if promoted, while very low values on either dimension indicate limited promotional effect.
- Results: 28.6% lower average percentile error was achieved by HIP forecasts than by state-of-the-art methods based on popularity history.Parameters were estimated from each video’s first 90 days, with forecasts made for the following 30 days on more than 13K actively discussed YouTube videos.
2. THE MODEL
The Hawkes Intensity Process models online attention as the interaction of continuous external promotion and endogenous, self-exciting responses. It converts event-level Hawkes dynamics into an intensity model that can be estimated from aggregated daily popularity data and explain complex histories.
- 2.1 Problem setting: HIP models each viewing event as triggered either by a previous event or by external influence.Its arrival rate combines exogenous promotion with endogenous responses generated by earlier events.
- 2.2 Hawkes process for social events: The model uses a power-law triggering kernel to represent memory decay, user influence, and video-dependent scaling.The kernel includes parameters for video quality, user influence, influence nonlinearity, cutoff behavior, and social-memory decay.
- 2.3 From Hawkes to HIP: HIP defines Hawkes intensity as the expected event rate over stochastic event histories, enabling popularity estimation from aggregated data.This addresses the practical setting in which daily views are observed but individual viewing-event times are unavailable.
- 2.3 From Hawkes to HIP: The resulting intensity is driven by external stimulus and its own history convolved with a power-law memory kernel.The model estimates video-dependent parameters from popularity histories and can incorporate an initial shock and constant background excitation for unobserved influence.
- 2.4 Model interpretation: HIP captures multiple rises and falls in popularity and links each external stimulus to an unfolding endogenous response.Example fits show fast memory decay tracking stimuli closely for a music video and slow decay producing a rising trend for a news video.
- 2.5 Virality metrics: The area under the impulse response, Âξ, measures a video’s endogenous response and is used with exogenous sensitivity to assess virality.These quantities provide two key metrics for comparing individual videos and collections.
3. THE TWEETED VIDEOS DATASET
The dataset links YouTube popularity and promotion signals to Twitter activity at large scale. Its popularity scale uses percentile rankings to represent long-tailed views and shares and reveals that some videos substantially change rank over time.
- 3.1 Dataset construction: The study links 81.9 million YouTube videos to 1.06 billion tweets over a six-month period.The dataset connects popularity on YouTube with external discussion on Twitter.
- 3.1 Dataset construction: The Active dataset retains videos online for at least 120 days with at least 100 tweets and 100 shares by day 120.These restrictions select videos with sufficiently long popularity and promotion histories for estimating external influence.
- 3.2 The popularity scale: Popularity is represented on a percentile scale from 0.0% for least popular to 100% for most popular.The scale is applied to views, shares, and tweets to handle their long-tailed distributions.
- 3.2 The popularity scale: The 2.5% most popular videos span more than one order of magnitude in both views and shares during their first 60 days.Figure 3 divides videos into 40 equally sized percentile bins and displays views and shares with boxplots.
- 3.2 The popularity scale: Videos from any popularity bin at 30 days can jump into the top popularity bins by 60 days.Most videos retain similar ranks or decline slightly, but upper-right outliers show substantial late popularity gains.
4. THE ENDO-EXO MAP
The endo-exo map characterizes videos by endogenous response and sensitivity to external stimuli, while popularity and promotion are shown separately. It supports comparisons of viral potential, explains differences among similarly promoted videos, and identifies videos that are difficult to promote.
- 4.1 Explaining popularity: The endo-exo map places videos by endogenous response Âξ on the x-axis and exogenous sensitivity µ on the y-axis.Circle radius represents popularity percentile, while color represents the percentile of total shares received.
- 4.1 Explaining popularity: Each unit of exogenous excitation generates µÂξ expected events, so large values indicate greater potential to become viral.Videos near one another on the map have similar intrinsic potential, making differences in external attention relevant to their popularity.
- 4.1 Explaining popularity: 3.22 times more promotions explain why v1 has 4.61x more views than the similarly positioned v2.The two videos have similar endogenous response and exogenous sensitivity but differ in external promotion.
- 4.2 What describes the most popular videos?: Popular Film & Animation videos tend to have higher exogenous sensitivity, whereas popular Gaming videos mainly have higher endogenous response.The comparison uses density plots for all videos versus the most popular 5% within each category.
- 4.3 Identifying unpromotable videos: Videos with µÂξ < 10^-3 are classified as unpromotable because one unit of external promotion generates very few expected views.The Active dataset contains 549 unpromotable videos, approximately 3.9% of the set.
- 4.3 Identifying unpromotable videos: An example unpromotable video has µ = 2.88 × 10^-15 and Âξ = 1, while v1 is expected to generate 598 views per promotion.These contrasting values illustrate how the map distinguishes low promotion responsiveness from high response potential.
5. FORECASTING POPULARITY GROWTH
HIP forecasts popularity over a temporal holdout using observed or planned external promotions, while also identifying videos whose popularity can surge after promotion. On the Active set, HIP reduced average percentile error relative to popularity-history baselines, with larger gains for videos exposed to large future shocks.
- A video that will go viral: A video explaining a brain disorder rose from the 5.85% to 94.9% popularity percentile after receiving 229 shares and gaining 2.42 million views between days 91 and 120.During its first 90 days, it had received 36 shares and 15,687 views, while HIP estimated high exogenous sensitivity and high endogenous response.
- Forecasting setup: HIP estimates parameters from each video’s first 90 days and forecasts popularity for the following 30 days using known or planned promotions.The model takes exogenous promotion s[t] as input and produces estimated viewcounts ξ[t].
- Forecasting results: 28.6% lower average percentile error was achieved by HIP (#shares) than MLR (#shares), with mean errors of 4.96% and 6.94%, respectively.The differences were statistically significant with paired t-test p < 0.001; HIP using shares slightly outperformed HIP using tweets, without a statistically significant difference at p = 0.001.
- Forecasting results: HIP (#shares) achieved a 5.11% mean percentile error versus 9.24% for MLR (#shares) on 4006 videos with large exogenous shocks during the forecasting period.The corresponding medians were 3.25% and 6.5%.
- Limitations of forecasting: HIP can miss forecasts when little external influence is recorded and popularity was likely driven by unseen exogenous sources.The forecasting protocol uses historical temporal holdouts rather than generating realistic promotions and responses in a large-scale social network.
- Causality: HIP is causal only in the linear-systems sense and does not establish that promotions change views under confounding factors.Granger-causality tests showed inconsistent results for shares influencing views or views influencing shares.
6. RELATED WORK
The related-work discussion positions HIP among popularity models, volume-based temporal descriptions, and graph-centric influence methods. Its distinctive focus is forecasting total popularity from linked external activity while recovering model parameters and non-stationary effects.
- Popularity modeling and prediction: Earlier popularity models used stylistic curves, point processes, reinforced Poisson processes, and double stochastic processes to represent bursts, decay, infectiousness, or social contagion.Examples include Hawkes point processes, SEISMIC, TiDeH, and models incorporating rich-get-richer effects.
- Modeling volumes of popularity: HIP differs by modeling popularity volumes under continuous external influence rather than only event times or final cascade size.It extends Hawkes processes by taking expectations over stochastic event histories to describe expected event volumes.
- Modeling volumes of popularity: HIP simultaneously forecasts total popularity, recovers all parameters from data, and explains non-stationary variation from linked external activity sources.These properties are presented as shortcomings addressed relative to earlier volume-based approaches.
- Influence estimation and maximization: Influence estimation and maximization learn user-level influence probabilities or choose promoters, whereas HIP measures promotion volume to forecast popularity.HIP therefore takes a volume-based rather than graph-centric view of external promotion.
7. SUMMARY AND DISCUSSION
The paper links external promotion to endogenous popularity dynamics, quantifies video-level sensitivity and virality, and applies these quantities to forecasting and identifying promotion-responsive videos. Its scope is limited by unobserved influences, omitted seasonality, and aggregation over users.
- Summary and contribution: HIP systematically links endogenous responses to exogenous stimuli and explains complex, multi-phased popularity dynamics over time.The model is validated on popularity and promotion histories of a large set of YouTube videos.
- Summary and contribution: HIP quantifies endogenous virality and exogenous sensitivity for each video, identifying videos likely to respond well or poorly to promotion.The analysis uses aggregated attention and promotion data from YouTube and public sources such as Twitter.
- Scope: The model can operate with aggregated popularity and promotion series generated by human behavior without platform-dependent assumptions.The paper suggests analogous attention dynamics for webpages, podcasts, and blogs.
- Limitations: HIP omits seasonality, models expected influence across users rather than individual influence, and cannot fully observe diverse online or offline promotion sources.The paper identifies seasonality components, individual influences, and unknown exogenous sources as directions for further investigation.
1 Details of HIP
HIP models popularity as an exogenously driven self-exciting process, connecting external promotions with endogenous responses. It averages marked Hawkes event rates over event histories to describe expected popularity volumes and supports interpretable response metrics.
- Model definition: External events and endogenous events jointly determine online attention, with endogenous events spawned by previous events.Exogenous events originate outside the system, while endogenous events respond to earlier exogenous or endogenous events.
- Marked Hawkes process: Marked Hawkes processes use event times and magnitudes to represent how strongly each event can amplify future events.Event magnitude can capture user influence, such as follower or friend counts, while the trigger kernel decays over time.
- Expected event rate: The expected event rate is derived from the underlying counting process by averaging the conditional event rate over possible event histories.The derivation converts event-history expectations into counting-process increments and then extends the unmarked result to marked processes.
- Model definition: HIP extends the Hawkes point process by modeling expected event volumes rather than individual event times.It takes expectations over stochastic event histories to obtain the Hawkes intensity ξ(t).
- Expected event rate: The resulting intensity function is presented as a new analytical form for event intensity, distinct from Hawkes’s earlier covariance-density equation.The paper states that this definition and derivation are new to the best of its knowledge.
- Model properties: HIP supports additive, scalable, and time-shifted responses to external stimulation, enabling separate source attribution and promotion scheduling.Scaling promotions scales the endogenous reaction, while shifting promotions shifts the corresponding views.
2 Details about fitting HIP
The fitting procedure estimates HIP parameters and unobserved external influence by minimizing regularized squared error between observed and modeled popularity. It uses iterative gradients, video-level random restarts, and temporally held-out tuning and forecasting.
- Objective: HIP fitting minimizes squared error between observed and modeled popularity while estimating model parameters and unobserved external influence.The fitted quantities include {µ, θ, C, c} and external-stimulus parameters γ and η.
- Objective: The recursive model uses previously estimated popularity values rather than observations when reproducing the full observed time series.This recursion is computed iteratively, and an L2 regularizer is added to improve fitting stability.
- Optimization: L-BFGS with supplied gradients optimizes the non-convex objective from multiple random initializations, retaining the solution with the lowest error.The implementation fits each video in parallel and uses eight random starting points per video.
- Regularization: L2 regularization is applied to selected linear coefficients, with reference values normalizing parameter weights; c and θ are excluded.The regularization coefficient is tuned per video by line search on a temporally held-out sequence.
- Validation and forecasting: The temporal split uses the first 75 days for estimation, the next 15 days for tuning, and days 91–120 for forecasting.The regularizer parameter ω is selected using the tuning sequence.
- Optimization properties: Given fixed nonlinear parameters θ and c, the loss is convex in the remaining linear parameters, but the overall procedure converges only to a local minimum.The method assumes stationarity over time and across different parts of the activated online social network, with no known convergence-rate results for sequence length.
3 Data
The study constructs large YouTube–Twitter datasets and two cleaned subsets for analyzing popularity and forecasting. It represents popularity on percentile-based scales to handle YouTube’s long-tailed view distribution.
- Dataset construction: 1,061,661,379 tweets produced 81,915,174 distinct YouTube video IDs in the raw Tweeted Videos dataset.The data covered 2014-05-29 to 2014-12-26 and used Twitter and YouTube APIs.
- Dataset construction: The 5Mo subset contains 16,417,622 videos with at least 60 days of popularity history for forecasting.More than half of the videos lacked complete popularity histories because of data loss, including videos no longer online.
- Dataset construction: The Active subset requires at least 100 tweets and 100 shares during recorded lifetime, with sufficient history for model estimation and forecasting.Its largest four categories cover more than 70% of videos, and Music accounts for more than 25%.
- Popularity measurement: Popularity bins use view-count percentiles, with 0.0% representing the least popular videos and 100% the most popular.This percentage scale makes comparisons possible despite the long-tailed distribution of view counts.
- Popularity measurement: The first 15% of videos each receive fewer than 10 views, whereas the top 5% span almost two orders of magnitude in view counts.The popularity scale remains similar at 30 and 60 days, with a slight increase in the dynamic range.
- Popularity measurement: Active videos occupy the top 5% popularity percentiles of 5Mo at 30 days, making Active a subset of the most popular videos.Figure 7 positions Active relative to the 5Mo popularity distribution.
4 Understanding popularity dynamics
The HIP analysis links popularity to exogenous sensitivity, endogenous response, memory, and category or channel characteristics. The endo-exo map exposes systematic differences in how videos respond to external influence and retain attention.
- Parameter interpretation: The endo-exo map relates fitted HIP parameters to video groups, including channels, content categories, and popularity.The supplementary analyses explicitly connect endogenous and exogenous components to model parameters.
- Channels and categories: Game-recording videos are generally more popular than the reporter’s news videos, explained by higher exogenous sensitivity µ.Figure 9 uses circle radii for popularity percentile and colors for the two channels.
- Channels and categories: Gaming, Comedy, and Entertainment show higher initial impulse γ and constant exogenous stimuli η, consistent with activity outside observed Twitter or YouTube sharing.γ and η represent unobserved exogenous influence not captured in observed activity s(t).
- Channels and categories: Comedy, Gaming, and Sport appear particularly sensitive to external influence, while Comedy, Entertainment, and Gaming show higher-than-median endogenous response.These category-level patterns provide an alternative view of the endo-exo map.
- Memory across categories: Music has mean θ=14.95 versus 15.94 overall, while Nonprofits & Activism and News & Politics have means of 17.56 and 19.45.All groups have long-tailed θ distributions peaking near θ≃3.35; smaller θ implies slower decay and larger endogenous response.
5 Popularity forecasting and comparison to baseline
The study evaluates HIP against multilinear regression using observed views and either shares or tweets as external stimuli. HIP forecasts popularity more accurately, including for videos with large exogenous shocks.
- Evaluation setup: The forecasting pipeline fits HIP parameters on the first 75 days, tunes ω on the next 15 days, and predicts views for days 91–120.Either shares or tweets can provide the known exogenous stimulus series, and MLR is trained on the same data.
- Forecasting performance: HIP forecasts 87% of videos with at most 10% error, compared with 78% for MLR at the same threshold.Adding shares to MLR yields only marginal improvement over MLR without exogenous information.
- Forecasting performance: HIP (#shares) has a median absolute forecasting error of 3%, versus 3.75% for MLR (#shares).The error distribution for HIP concentrates more strongly at lower errors.
- Statistical comparison: No significant forecasting difference is found between shares and tweets as exogenous stimulus sources for either HIP or MLR.The two series are highly correlated for most videos, supporting similar forecasting performance from either source.
- Statistical comparison: The HIP–MLR forecasting differences are statistically significant, with p-values of approximately 10^-151 for shares and 10^-95 for tweets.Both comparisons show small but non-negligible Cohen’s d effect sizes.
- High exogenous shocks: Among 4,006 videos with high testing-period exogenous shocks, HIP (#shares) has a 3.25% median error versus 6.5% for MLR (#shares).HIP maintains performance similar to the full Active dataset, while MLR largely misses these difficult predictions.
Video 8QWg0YrcFhI: example from the exogenous set
The example concerns a relatively obscure Brazilian-politics video that received a delayed burst of attention before returning to obscurity.
- Example video: The video experienced a considerable attention increase more than three months after upload, followed by declining attention after December 2015.The example illustrates a high exogenous shock during the testing period.
- Example video: The example is used to compare Hawkes intensity and MLR forecasting performance on a difficult high-shock video.Figure 15 combines the video trajectory with aggregated forecasting-error comparisons.
Autos & Vehicles: all
The figure set compares category-specific video distributions on the endo-exo map, including Autos & Vehicles and Howto & Style top-5% views. Figure 17 continues the remaining categories.
- Autos & Vehicles: all: Autos & Vehicles: top 5% videos are shown as a category-specific endo-exo-map distribution.
- Autos & Vehicles: all: Howto & Style: top 5% videos are also represented in the category comparisons.
- Autos & Vehicles: all: Each category uses paired 2-dimensional density heatmaps: all videos on the left and the most popular 5% on the right.
- Autos & Vehicles: all: Figure 17 continues Figure 16 with six additional categories.