Source-linked AI summary

Inferring Temporal Dependencies from Social Time Series with the Cross-Correlogram

Bridget Smart, Renaud Lambiotte, Takaaki Aoki, Ryota Kobayashi

arXiv:2609.16633v1cs.SIphysics.soc-ph

TL;DR

Social event streams are bursty and non-stationary, limiting temporal-dependence methods that rely on stationarity or separated timescales. The paper adapts cross-correlograms with smooth intensity null models, recovering interpretable lag profiles and revealing delayed online relationships that co-occurrence measures miss.

  • Problem

    Burstiness, non-stationarity, and overlapping temporal scales make it difficult to distinguish genuine social dependencies from correlations induced by shared rhythms.

  • Method

    The paper models temporal structure in smooth functional or empirically estimated intensity nulls while retaining observed event times for cross-correlogram analysis.

  • Results

    Applied to 3.1 million social-media event times, the smooth-null cross-correlogram revealed delayed hashtag relationships whose television-program lags matched published broadcast times.

  • Takeaways & Limitations

    Embedding behavioral rhythms in statistical null models improves the robustness and interpretability of temporal inference in complex event-time systems.

  • Takeaways & Limitations

    The formulation assumes Poisson event generation, may underrepresent self-exciting dynamics, and does not establish causality.

Abstract

from arXiv · show

Characterizing temporal interactions in social systems is challenging because social behavior can be bursty and non-stationary, violating the stationarity assumptions of many methods used to measure temporal dependence. The cross-correlogram, an existing technique used to profile neural excitations and inhibitions, offers an interpretable alternative to methods such as Granger causality or co-occurrence, as it produces a full profile of lagged dependence directly from event times rather than a single summary statistic. We adapt the cross-correlogram by integrating functional models of behavior with data-driven temporal response profiling. By characterizing how periodic structure biases traditional cross-correlograms, we propose a correction based on smooth intensity functions, specified from a known functional form or estimated empirically. This approach provides a robust, interpretable estimator of temporal dependency profiles even when collective rhythms operate on timescales that overlap those of the interactions of interest. We demonstrate theoretically and through simulation that the proposed method recovers temporal dependencies in periodic regimes, outperforming interval-jitter and Granger causality methods. Finally, we apply the method to 3.1 million event times from X (formerly Twitter) collected between 2019 and 2020, demonstrating how cross-correlograms reveal delayed temporal relationships in collective online behavior that are missed by co-occurrence measures. For a subset of television-related hashtags, recovered delays align with known broadcast schedules, providing evidence that the proposed method captures genuine temporal structure rather than artifacts of shared attention cycles.

1. Introduction

Social event streams are bursty and non-stationary, making standard temporal-dependence measures vulnerable to spurious associations. The paper adapts cross-correlograms with temporal null models to recover interpretable lagged dependencies in social systems.

  • Social rhythms and burstiness can create correlations from shared exogenous variation, confounding association with direct interaction.Regular behavioral patterns are often omitted from system-level analyses, despite their importance for social event timing.
  • Standard cross-correlograms may fail in social settings because homogeneous baselines, stationarity, and timescale separation assumptions are often violated.These violations can distort baseline structure and introduce spurious dependencies.
  • The proposed framework combines data-driven estimation with behavioral mechanisms to produce full temporal profiles of dependence.The approach first characterizes periodic and bursty bias, then develops corrections adaptable to specified temporal patterns.
  • Interval jitter can fail when unwanted temporal structure overlaps the timescale of true dependencies because it assumes linear, time-invariant processes.The paper derives jitter-induced error in periodic and bursty settings.
  • The smooth-intensity null models event rates as functions in an inhomogeneous Poisson framework while leaving observed event times unchanged.Expected coincidence counts account for interactions without removing lagged dependencies from the data.
  • Unlike co-occurrence, Granger causality, and transfer entropy, the cross-correlogram yields an interpretable lag distribution rather than a single summary statistic.It is symmetric and interpolation-free, complementing windowed and directed temporal measures.

2. Related work

Related work uses event-time data to study dependencies, coordination, and prediction, while highlighting challenges from irregularity, non-stationarity, periodicity, and burstiness. Existing corrections and null models motivate incorporating temporal structure directly when interpreting cross-correlograms.

  • Event-time correlations support studying physical dependencies, shared external responses, social-media coordination, unrest forecasting, and prediction.
  • Classical and event-based methods estimate directional dependence but often require temporal binning, continuous approximations, or large samples.
  • Cross-correlograms originated for neural spike trains, where non-stationarity can create lagged structures that are incorrectly interpreted as causal.
  • Jittering removes temporal structure below a chosen timescale, but its null can absorb structure at the interaction timescale, especially under periodic or bursty modulation.The paper quantifies this effect for periodic and bursty rate modulation.
  • Simulations illustrate how a temporal-structure-aware null model separates genuine source-target interaction from correlation induced by shared periodic variation.
  • Social activity is bursty and shaped by circadian, weekly, and seasonal rhythms, motivating models that combine excitation with explicit temporal modulation.Related work also applies generalized linear models, convolutional networks, community detection, and Hawkes processes to temporal network and hashtag data.

3. Measuring Temporal Dependencies in Social Systems

The cross-correlogram estimates lagged dependence from event times, but periodicity, burstiness, and non-stationarity can create misleading peaks. The proposed null-based corrections model temporal structure while preserving observed event times and can isolate interaction-related deviations.

  • 3.1. The cross-correlogram: A cross-correlogram is a histogram of inter-event lags whose positive or negative peaks indicate temporal ordering between source and target events.Its lag window is chosen to match the timescale of the behavior under study.
  • 3.2. Non-Stationarity: Homogeneous Poisson nulls become unreliable when rates are non-stationary or bursty, especially when unwanted rhythms overlap the timescales of genuine interactions.These conditions can make apparent correlogram peaks reflect shared external modulation rather than influence.
  • 3.2. Non-Stationarity: Shared periodic intensities can produce directional-looking peaks without interaction, because the cross-correlogram reflects phase offsets between independent processes.For a 24-hour periodic intensity, the observed peak occurs at the phase offset even when the processes are independent.
  • 3.3. Interval Jitter: Interval jitter can introduce error when temporal structure and target interactions occur on overlapping scales, with error increasing alongside burstiness and periodic modulation.The paper derives and validates this error theoretically and on synthetic data.
  • 3.4. Smooth intensity function: The proposed correction incorporates temporal structure into the null’s expected bin heights instead of modifying event times, preserving all observed lagged dependencies.Functional and empirical intensity models account for predictable or estimated rate variation, while remaining deviations can represent interaction.

4. Evaluation and Comparison

The evaluation tests cross-correlogram corrections under homogeneous, unimodal periodic, and bimodal periodic simulations, including misspecified nulls and baseline methods. Performance depends on matching the null model to temporal structure, while smooth intensity nulls remain useful when structure is difficult to parameterize or interaction and nuisance timescales overlap.

  • Simulation design: Detection probability rises with interaction strength when the null model matches the data-generating temporal structure.Correctly specified homogeneous and periodic cross-correlograms recover increasing detection sensitivity as ρ increases.
  • Periodic settings: A misspecified periodic K = 1 null falsely detects a response at ρ = 0 in 86% of bimodal simulations.The method is invalid in this setting because the null fails to represent the bimodal temporal structure.
  • Periodic settings: The smooth null performs well for both unimodal and bimodal periodic data, but detects unimodal spikes only when ρ ≥0.3.It outperforms the correctly specified null in the bimodal setting but is less sensitive than the correctly specified K = 1 null in the unimodal setting.
  • Baseline comparison: Mismatched nulls produce spurious detections or reduced sensitivity, while Granger causality remains near-uninformative for periodic data.In homogeneous data, Granger causality outperforms cross-correlogram methods, whereas homogeneous nulls applied to periodic data falsely detect weak or absent spikes.
  • Method choice: Interval jitter is appropriate when unwanted and target timescales are separated, but smooth nulls are needed when those scales overlap.The smooth-null approach preserves observed event times and therefore retains lagged dependencies that jittering could remove.

5. Application: Temporal Interaction Networks in Social Media

The application represents hashtags as event-time processes and compares smooth-null cross-correlogram networks with exact-time co-occurrence networks. The two measures identify substantially different relationships, including delayed television-related responses aligned with broadcast schedules.

  • Results and Network Analysis: 50% ranking agreement is reached only at R = 6,252, covering 31% of all 19,900 unordered hashtag pairs.Agreement remains low for small and moderate rank thresholds, showing that the measures prioritize different relationships.
  • Results and Network Analysis: The top-100 cross-correlogram-only network forms thematic clusters around politics, news, television, sport, and music fandom, with #auspol as a hub.These relationships include ongoing event-driven patterns that co-occurrence is less likely to detect.
  • Television response profiles: Television-program cross-correlograms peak near known broadcast delays, including relationships among Australian entertainment and current-affairs programs.The analysis uses 3,133,247 hashtag-time pairs from posts collected between January 2019 and September 2020.
  • Results and Network Analysis: At R = 100, the cross-correlogram network connects 55 of 57 active nodes, whereas the co-occurrence network connects only 9 of 96.The cross-correlogram network therefore exposes a much broader set of delayed relationships at this threshold.
  • Results and Network Analysis: The cross-correlogram identifies delayed relationships missed by co-occurrence, with weak overall rank correlation of 0.379 across 19,900 hashtag pairs.This complements exact-time co-occurrence by detecting structured temporal responses rather than only simultaneous activity.

6. Discussion and Conclusion

The paper concludes that smooth, inhomogeneous null models make cross-correlograms interpretable for bursty, non-stationary social event data. Simulations and a 3.1-million-event application support recovery of structured delayed relationships, while the method remains limited by Poisson assumptions, quadratic scaling, and lack of causal identification.

  • Conclusions: Functional and empirical smooth intensity models account for behavioral rhythms without perturbing event times, recovering dependency profiles under known periodic structure.The framework analytically characterizes bursty and periodic bias, derives jitter error, and validates the proposed nulls through simulation.
  • Conclusions: Smooth-null cross-correlograms recover delayed, asymmetric hashtag relationships and television lags matching published broadcast times.These matches provide external validation that recovered lags reflect structured behavior rather than shared attention cycles.
  • Limitations: The method assumes Poisson event generation and may underrepresent higher-order dependencies or self-exciting social dynamics.The authors suggest extensions to Hawkes or renewal processes to model such feedback explicitly.
  • Limitations: All pairwise cross-correlograms scale quadratically with the number of processes, motivating approximate or sparse representations for large systems.The framework also infers directional timing relationships without establishing causality.
  • Broader implications: Interpretable lag profiles and temporal-network visualizations support studying collective attention, information diffusion, and coordination across complex systems.The conclusion presents this as a broader potential of combining temporal-correlation analysis with functional modeling.

8. Funding

The work received support from the listed NSF, EPSRC, JSPS KAKENHI, JST FOREST, and AMED grants.

  • Funding: The authors acknowledge funding from NSF, EPSRC, JSPS KAKENHI, JST FOREST, and AMED grants.The passage lists grant numbers and individual researcher support.

A. Comparison of methods on periodic data

When the null model matches the periodic data-generating process, detection power increases reliably with interaction strength, while mismatched nulls produce false positives or reduced sensitivity.

  • A. Comparison of methods on periodic data: Matching the null model to the periodic data-generating process makes detection power increase reliably with interaction strength.The functional model fit is used for the periodic setting, whereas mismatched nulls can distort inference.
  • A. Comparison of methods on periodic data: Mismatched null models can generate false positives or reduce sensitivity to temporal interactions.
  • A. Comparison of methods on periodic data: Visual inspection of jitter-based cross-correlograms often cannot establish significance or separate null-model artifacts from structure at particular timescales.With sufficient data, smooth functional nulls provide clearer identification of meaningful deviations from theoretical expectations.

B. Bursty dynamics

The bursty-process evaluation tests detection across interaction strengths using recursive burst generation and repeated simulations, showing that interval-based corrections improve detection but remain sensitive to interval choice.

  • B. Bursty dynamics: Bursty simulations generate recursive short-delay event sequences, with burst probability BR and delay U(l, u), optionally switching states through an HMM.In non-bursty states, events follow a standard Poisson process.
  • B. Bursty dynamics: Detection sensitivity is evaluated across spike strengths ρ = 0, 0.05, . . . , 1, with each configuration repeated over 100 runs.Artificial delayed copies are inserted into the target process for a proportion ρ of source events.
  • B. Bursty dynamics: Smooth null and jittered cross-correlograms have true positive rates that increase with interaction strength, although both show elevated false positive rates.Granger causality does not perform well in this bursty setting.
  • B. Bursty dynamics: In bursty data, both interval jitter and smooth intensity approaches depend on interval size, while smaller intervals better approximate stochastic short burst periods.

C. Distribution of cross-correlogram bin heights in the non-stationary case

For independent inhomogeneous Poisson processes, cross-correlogram bin heights have expectations obtained by integrating target intensity over source events, while non-stationarity adds variance beyond the conditional mean.

  • C. Distribution of cross-correlogram bin heights in the non-stationary case: A source event contributes the number of target events falling within a lagged bin interval of width ∆.
  • C. Distribution of cross-correlogram bin heights in the non-stationary case: The expected bin height is obtained by integrating the target intensity over all possible source events and averaging across the bin interval.
  • C. Distribution of cross-correlogram bin heights in the non-stationary case: The bursty and periodic simulations compare these theoretical distributions with null-model detection performance under non-stationary event rates.
  • C. Distribution of cross-correlogram bin heights in the non-stationary case: Conditioned on source event times, bin contributions are independent Poisson variables, so conditional variance equals conditional mean.
  • C. Distribution of cross-correlogram bin heights in the non-stationary case: Under inhomogeneous intensities, the unconditional variance of a cross-correlogram bin includes a generally nonzero extra term and therefore need not equal its mean.That term vanishes for homogeneous processes, yielding Var(h_i) = E[h_i].

D.1. Cross-correlogram error

The analysis quantifies relative cross-correlogram errors caused by interval jitter, using theoretical calculations for periodic effects and simulations for bursty effects.

  • D.1. Cross-correlogram error: The analysis evaluates jittering error by comparing theoretical bin heights from original and jittered time series under periodic or bursty effects.
  • D.1. Cross-correlogram error: Jitter error is quantified by comparing original bin heights h_j with jittered heights h̃_j through relative error, median error, and a 5th–95th percentile error band.
  • D.1. Cross-correlogram error: Even a 1% single-bin difference between original and jittered cross-correlograms can greatly reduce statistical-test power.The analysis therefore considers how jitter-induced errors propagate through the cross-correlogram procedure.
  • D.1. Cross-correlogram error: Periodic non-stationary effects permit exact per-bin error derivation, whereas bursty effects require simulations to approximate errors.

D.2. Periodic non-stationary effects

The simulations examine how periodic and bursty non-stationarity affect cross-correlogram errors and show that jittering bias increases with burstiness and periodic amplitude.

  • D.2. Periodic non-stationary effects: The simulations use sinusoidal target intensities with amplitude aT and period bT, alongside a constant background intensity and specified cross-correlogram and jitter parameters.Default settings include aT = 1, bT = 24, cross-correlogram width w = 6, bin width Δ = 0.25, and jitter interval k = 24.
  • D.2. Periodic non-stationary effects: Jitter-window size has little relationship with typical error, although maximum individual error increases slightly as the window grows.
  • D.2. Periodic non-stationary effects: Increasing the amplitude of periodic non-stationarity raises both the median and spread of individual relative errors.
  • D.2. Periodic non-stationary effects: Relative error varies with the period of periodic non-stationarity, with stronger effects at some periods such as 12 and 20 units.
  • D.2. Periodic non-stationary effects: As burstiness increases, jittering introduces greater bias in the cross-correlogram under both burst-ratio and HMM models.The difference between original and jittered bin heights grows with burst ratio and with time spent in the HMM's bursty state.

E. Topic classifications

The study classifies Twitter hashtags into broad topics and specific subtopics, then manually checks and merges labels for consistent thematic analysis.

  • E. Topic classifications: ChatGPT was used to assign general topics and specific subtopics, after which classifications were manually checked and some topics were merged.
  • E. Topic classifications: The resulting labels support analysis of hashtag dynamics as thematic groups.
  • E. Topic classifications: The final classifications include music with 29 unique hashtags and politics with 23 unique hashtags.
  • E. Topic classifications: Television and news/current-affairs hashtags form distinct classified groups used in the temporal analysis.The television group contains 10 unique hashtags, while the news group contains 11.
  • E. Topic classifications: Other classified topics include general, literature, gaming, health, music awards, adult, sports, travel, environment, finance, and other.
Loading 2609.16633v1…