Source-linked AI summary

A Poissonian explanation for heavy-tails in e-mail communication

R. Dean Malmgren, Daniel B. Stouffer, Adilson E. Motter, Luis A. N. Amaral

arXiv:0901.0585v1physics.soc-phcs.CYphysics.data-an

TL;DR

The paper examines whether matching heavy-tailed inter-event-time scaling is sufficient to explain e-mail activity mechanistically. It proposes a cascading non-homogeneous Poisson process incorporating periodic activity and tests it against empirical data, finding agreement with the full distribution.

  • Problem

    Matching asymptotic power-law scaling alone may not establish an accurate mechanistic description of the underlying process.

  • Method

    The model combines time-varying daily and weekly activity with cascades of additional events during active intervals, then fits the active-interval configuration to observed inter-event times.

  • Results

    Only 1 user rejected the cascading model at the 5% significance level, compared with 344 users for the truncated power-law null model.

  • Takeaways & Limitations

    Circadian and weekly cycles coupled with cascading activity can accurately describe heavy tails in e-mail communication without requiring rational decision making.

  • Takeaways & Limitations

    The model describes when e-mails are sent but does not determine their probable recipients or explain how e-mail social-network structure evolves.

Abstract

from arXiv · show

Patterns of deliberate human activity and behavior are of utmost importance in areas as diverse as disease spread, resource allocation, and emergency response. Because of its widespread availability and use, e-mail correspondence provides an attractive proxy for studying human activity. Recently, it was reported that the probability density for the inter-event time $τ$ between consecutively sent e-mails decays asymptotically as $τ^{-α}$, with $α\approx 1$. The slower than exponential decay of the inter-event time distribution suggests that deliberate human activity is inherently non-Poissonian. Here, we demonstrate that the approximate power-law scaling of the inter-event time distribution is a consequence of circadian and weekly cycles of human activity. We propose a cascading non-homogeneous Poisson process which explicitly integrates these periodic patterns in activity with an individual's tendency to continue participating in an activity. Using standard statistical techniques, we show that our model is consistent with the empirical data. Our findings may also provide insight into the origins of heavy-tailed distributions in other complex systems.

Empirical patterns

The study analyzes e-mail records to characterize deliberate human activity and finds that circadian and weekly periodicity produces systematic deviations from a truncated power-law null model.

  • Empirical dataset: 3,188 e-mail accounts were observed over 83 days, yielding 394 accounts suitable for quantifying activity after preprocessing.Records included sender, recipient, message size, and one-second timestamps; likely spammers and listservs were excluded.
  • Activity patterns: A fictitious user illustrates sporadic daily activity interspersed with active intervals during which e-mails are sent in rapid succession.The activity rate changes periodically with sleep and work patterns, while active intervals vary in length.
  • Periodic deviations: The data contain significantly more inter-event times between 16 and 32 hours than predicted by the truncated power-law null model.The excess is expected from e-mails sent during similar eight-hour periods on consecutive workdays; normalization causes overestimation elsewhere.

Model

The paper models e-mail activity with a periodic non-homogeneous Poisson process and nested active-interval cascades. The primary rate follows daily and weekly activity patterns, while secondary processes generate additional events within active intervals.

  • Primary process: The primary process is a non-homogeneous Poisson process whose rate ρ(t) varies periodically with period W.The model uses ρ(t) = ρ(t + W) to represent recurring activity patterns.
  • Primary process: Daily and weekly distributions of active-interval initiation determine the periodic primary rate.The model relates ρ(t) to pd(t) and pw(t), the distributions of starting an active interval by time of day and week.
  • Primary process: The period W is one week, and Nw denotes the average number of active intervals per week.These quantities parameterize the recurring initiation of activity intervals.
  • Cascading process: Each primary-process event initiates a secondary homogeneous Poisson process producing Na additional events at rate ρa before control returns to the primary process.These cascades represent bursts of e-mail activity within active intervals.

Results

Because active-interval configurations are unobserved, the authors infer model distributions nonparametrically and evaluate agreement with e-mail data using Monte Carlo hypothesis testing. The cascading model is consistent with the observed inter-event-time distribution for nearly all users, unlike the truncated power-law null model.

  • Parameter inference: The data do not specify which events belong to the same active interval, leaving the distributional form of Na uncertain.For example, it is unclear whether p(Na) should be modeled as normal or exponential.
  • Parameter inference: A new method nonparametrically infers pd(t), pw(t), and p(Na) without assuming their functional forms.This addresses the absence of a priori knowledge about the cascading activity pattern.
  • Parameter inference: The best-estimate active-interval configuration is selected by manipulating configurations to match the observed inter-event-time distribution.The configuration determines Nw, initiation distributions, ρa, and p(Na).
  • Assumptions: The analysis assumes that the fraction of time spent in active intervals is very small.The authors state that this condition was verified for all users considered.
  • Model evaluation: Monte Carlo hypothesis testing found p-values clearly above the 5% rejection threshold for the model’s agreement with empirical data.The test accounts for the fact that model parameters were estimated from the same data.
  • Model evaluation: At the 5% significance level, the cascading model was rejected for one user, compared with 344 users for the truncated power-law null model.The cascading model also lacked the null model’s systematic deviations from the data.

Discussion

The model explains heavy-tailed e-mail timing through periodic activity cycles coupled with cascading activity, while hypothesis testing distinguishes mechanistic agreement from asymptotic scaling alone. Its scope extends beyond e-mail but leaves recipient choice unresolved.

  • Discussion: Circadian and weekly cycles coupled with cascading activity accurately describe heavy tails in e-mail communication patterns.The authors argue that rational decision making is not necessary given this simpler explanation.
  • Discussion: The model may apply to other conscious activities, including telephone calls and errands, where repeated actions cluster within routines.These examples illustrate the model’s periodic and cascading mechanisms for optimizing time and effort.
  • Discussion: The model’s periodic and cascading features must be adapted to the activity, and its parameters can be generalized beyond stationary settings.Examples include menstrual cycles, airline seasonality, and changing annual letter volume.
  • Discussion: The single-activity model can be extended to multiple activities by treating it as a non-stationary hidden Markov point process.Individuals switch between activities according to time-dependent transition probabilities.
  • Discussion: Additional records of computer or e-mail-client use could directly measure active intervals; without them, simulated annealing infers the hidden structure.The inferred structure supports comparison with other cascading point processes.
  • Discussion: The model describes when e-mails are sent but does not determine probable recipients, leaving social-network evolution incompletely modeled.Recipient choice may depend on random contact, shared interests, task priority, or previous correspondence.
  • Discussion: Both models reproduce asymptotic scaling, but only the proposed model is consistent with the entire inter-event time distribution.This illustrates why hypothesis testing is needed to assess model validity.
  • Discussion: Matching asymptotic power-law scaling alone does not establish an accurate mechanistic description of the underlying process.Unrecorded active intervals can conceal multiple activity scales whose mixture produces scale-free patterns.

Area test statistic.

The area test statistic measures disagreement between empirical and model cumulative distributions after transforming inter-event times logarithmically. Simulated annealing searches for active-interval configurations that minimize this disagreement.

  • Area test statistic.: The area A measures the difference between empirical and model cumulative distribution functions for inter-event times.It quantifies agreement between model M(θ) and data set D.
  • Area test statistic.: Using u = ln τ improves simulated-annealing efficiency because the transformed variable is roughly uniformly distributed.The area statistic is easy to interpret and retains more distributional information than many alternatives.
  • Area test statistic.: The procedure seeks an active-interval configuration that minimizes area A between empirical data and cascading-process predictions.Starting from a random configuration, it estimates parameters, computes the model cumulative distribution, and evaluates A before modifying the configuration.
Loading 0901.0585v1…