Source-linked AI summary

A Review of Self-Exciting Spatio-Temporal Point Processes and Their Applications

Alex Reinhart

arXiv:1708.02647v2stat.ME

TL;DR

The paper addresses how to model and understand clustered events whose rates depend on prior history, especially where existing application literatures have developed separately. It reviews the theory, estimation, inference, declustering, diagnostics, and applications of self-exciting spatio-temporal point processes. The review concludes that these models are fast, flexible, and interpretable tools when the underlying generative process is clustered.

  • Problem

    Self-exciting models capture history-dependent triggering that spatial regression and log-Gaussian Cox processes do not, while theory, estimation, and inference have largely remained isolated across applications.

  • Method

    The paper reviews basic theory, estimation and inference techniques, declustering, diagnostics, and major applications of self-exciting spatio-temporal point processes.

  • Results

    The review identifies advances supporting maximum likelihood and Bayesian estimation, event declustering, and model diagnostics across earthquake, crime, and infectious-disease applications.

  • Takeaways & Limitations

    When the underlying generative process is clustered, self-exciting spatio-temporal point processes provide fast, flexible, and interpretable tools for scientific applications.

  • Takeaways & Limitations

    Parameter estimates can be biased by boundary effects because unobserved events outside the observation region or before the study period may generate observed offspring.

Abstract

from arXiv · show

Self-exciting spatio-temporal point process models predict the rate of events as a function of space, time, and the previous history of events. These models naturally capture triggering and clustering behavior, and have been widely used in fields where spatio-temporal clustering of events is observed, such as earthquake modeling, infectious disease, and crime. In the past several decades, advances have been made in estimation, inference, simulation, and diagnostic tools for self-exciting point process models. In this review, I describe the basic theory, survey related estimation and inference techniques from each field, highlight several key applications, and suggest directions for future research.

1. INTRODUCTION

Self-exciting spatio-temporal point processes model event rates from spatial location, time, and prior event history, making them useful for studying triggering and clustering. This review motivates these models over aggregated regression and surveys their theory, inference, and applications.

  • Motivation: Aggregated regression can produce coefficients and variances that depend strongly on arbitrary spatial boundaries or grids.This is a consequence of the Modifiable Areal Unit Problem.
  • Motivation: Point-process models estimate event intensity directly at any spatial location and time, avoiding aggregation into cells and discrete intervals.The intensity function predicts the rate of events at location s and time t.
  • Self-excitation: Self-exciting models condition event rates on prior history, allowing events to trigger new events and representing the process as a cluster process.This dependence distinguishes them from spatial regression and log-Gaussian Cox processes.
  • Scope: The review covers maximum likelihood estimation, stochastic declustering, Bayesian approaches, model selection, and diagnostic techniques.Declustering attributes events to prior triggering events or the underlying background process.
  • Applications: Applications include earthquake forecasting, crime dynamics, infectious disease, and extensions to events occurring on networks.These applications illustrate the utility of self-exciting models and the reviewed techniques.

2. SELF-EXCITING SPATIO-TEMPORAL POINT PROCESSES

Self-exciting spatio-temporal point processes extend Hawkes models by making event rates depend on prior events across space and time. Their conditional-intensity and cluster-process representations capture triggering, offspring generations, and event marks.

  • 2.1 Hawkes Processes: Self-exciting processes model current event intensity from the past history, with triggering functions controlling how long prior events influence rates.Rapidly decaying triggering functions produce recent-history dependence, whereas slower decay permits longer-term effects.
  • 2.1 Hawkes Processes: Hawkes processes can be represented as Poisson cluster processes containing background events and recursively triggered offspring events.When m < 1, expected total cluster size is 1/(1− m), including the initial background event.
  • 2.2 Spatio-Temporal Form: Spatio-temporal models extend conditional intensity to event locations and times, with self-excitation commonly represented by a nonnegative spatial-temporal triggering function.The triggering function is often separable as g(s−si, t−ti) = f(s−si)h(t−ti).
  • 2.2 Spatio-Temporal Form: The triggering function serves as the offspring-process intensity, making the cluster representation useful for estimating and simulating self-exciting processes.Proper normalization turns it into a probability distribution for offspring locations and times.
  • 2.3 Marks: Marked point processes augment event locations and times with features such as earthquake magnitude or crime type, while the ground process omits those marks.The ground-process intensity can be combined with a conditional mark density; the ground process may depend on prior marks.
  • 2.4 Log-Likelihood: Janossy densities provide a route to likelihood construction when self-excitation makes event-count and spatial-distribution probabilities difficult to obtain directly.For n events, the Janossy density describes the probability of exactly n events occurring in specified infinitesimal intervals.

3. ESTIMATION AND INFERENCE

The review surveys approaches for estimating parameters, performing inference, and simulating self-exciting spatio-temporal point-process data. It focuses largely on maximum likelihood while distinguishing descriptive-statistics approaches as a related literature.

  • 3. ESTIMATION AND INFERENCE: The estimation and inference literature commonly begins with an observed realization and a conditional-intensity model whose parameters must be estimated.The same framework can support inference and simulation of new data.
  • 3. ESTIMATION AND INFERENCE: Maximum likelihood is the review’s primary focus, while descriptive-statistics methods instead emphasize first- and second-order moments of the process.The review does not delve into the descriptive-statistics literature, citing other sources for broader treatments.

3.1 Maximum Likelihood

Maximum likelihood is the standard fitting approach for self-exciting point-process models, but direct optimization is computationally and numerically difficult. EM addresses this by introducing latent branching assignments, simplifying maximization and supporting declustering.

  • Challenges: Maximum-likelihood fitting is often analytically intractable because conditional intensities contain sums over previous points, making numerical evaluation O(n^2).The likelihood can be nearly flat, convergence can be very slow, and optimization may fail for some examples.
  • EM algorithm: The latent variable u_i records whether event i is a background event or was triggered by a previous event.This branching structure follows the model’s cluster-process representation.
  • EM algorithm: Conditioning on the branching structure simplifies the complete-data log-likelihood because each event’s intensity comes only from its source.The source is either the background process or a previous event.
  • EM algorithm: The EM E step estimates triggering probabilities, while the M step maximizes the expected complete-data log-likelihood and repeats until convergence.Termination occurs when the log-likelihood converges or parameter changes fall below a specified tolerance.
  • Benefits: EM makes each maximization easier than other numerical methods and reuses triggering probabilities for stochastic declustering.The shared probabilities connect parameter estimation with subsequent event attribution.
  • Caveat: Boundary effects can bias estimates when events outside the observed region or before the observation period influence observed offspring.The estimated mean number of offspring may be biased downward, while intensity bias is not well characterized for common models.

3.2 Stochastic Declustering

Stochastic declustering separates background events from triggered events using estimated triggering probabilities. The reviewed methods include iterative model-based procedures, predictive semiparametric estimation, and model-independent approaches with uncertainty quantification.

  • Model-Based Stochastic Declustering: Stochastic declustering estimates which events belong to the background process rather than being triggered, requiring iterative updates because background estimation and triggering-function estimation depend on each other.The model-based procedure uses a parametric triggering function and a nonparametric background estimate.
  • Model-Based Stochastic Declustering: The thinning procedure keeps each event with probability 1 − Pr(u_i ≠ 0), treating retained events as background and deleting the rest as triggered.The probabilities are recalculated using the final estimated background intensity.
  • Model-Based Stochastic Declustering: Family-tree construction samples background or ancestor assignments for each event using the probabilities Pr(u_i = 0) and Pr(u_i = j).Repeating the stochastic procedure reveals uncertainty by showing whether features persist across declustered realizations.
  • Forward Likelihood-based Predictive approach: The Forward Likelihood-based Predictive approach estimates parameters from sequential log-likelihood increments, using earlier observations to predict each subsequent event.Its semiparametric extension alternates estimation of smoothing parameters and triggering-function parameters until convergence.
  • Forward Likelihood-based Predictive approach: Improved performance over fixing smoothing parameters with Silverman’s rule was reported for a large earthquake catalog in Italy.The comparison concerns the semiparametric method applied to earthquake models.
  • Model-Independent Stochastic Declustering: MISD removes the need for a parametric triggering function by estimating its shape from the data, while later extensions allow spatially varying backgrounds and bootstrap uncertainty intervals.These extensions support inference about the estimated triggering function’s shape.

3.3 Simulation

Simulation methods generate self-exciting spatio-temporal event catalogs either by thinning candidate events according to conditional intensity or by exploiting the process’s cluster structure. Both approaches require care with boundaries, while cluster-based simulation can avoid repeated thinning computations.

  • Thinning methods: Thinning-based simulation generates candidate events and retains them according to the conditional intensity so history dependence is incorporated.The approach was originally developed for nonhomogeneous Poisson processes and adapted to spatio-temporal settings.
  • Thinning methods: Ogata’s two-stage algorithm simulates event times using an intensity integrated over space, then samples locations from the induced spatial distribution.Proposed events are thinned in proportion to their intensities.
  • Cluster-based simulation: Cluster-based simulation eliminates thinning and repeated evaluation of the conditional intensity, addressing a computational cost of Ogata’s method.Direct intensity evaluation can involve a large sum, and thinning may require many candidate times.
  • Cluster-based simulation: Cluster-based simulation first generates background events, then recursively generates Poisson-distributed offspring with locations and times drawn from the triggering function.The procedure continues until no new offspring remain, after which all generations are combined.
  • Simulation boundaries: Simulating only within the observation window creates edge effects because unobserved events outside the spatial or temporal boundary can produce observed offspring.A larger simulation window followed by restriction to the target region avoids these effects; a perfect spatio-temporal extension remains undeveloped.

3.4 Asymptotic Normality and Inference

Inference for self-exciting spatio-temporal point processes combines asymptotic likelihood theory, covariance estimation, parametric bootstrap, and partial likelihood. Finite-sample performance can diverge from asymptotic guarantees, especially for short observation periods or small samples.

  • The Hessian-based covariance estimator can produce heavily biased standard errors for small to moderate observation periods.Repeated-simulation comparisons indicate poor finite-sample behavior.
  • Maximum likelihood estimates are consistent and asymptotically normal as observation time T →∞ under regularity conditions on λ(s, t).
  • Wald tests derived from the estimated covariance matrix provide parameter tests and confidence intervals through test inversion.
  • Parametric bootstrap repeatedly simulates and refits the model, then uses estimated-parameter quantiles to construct 95% confidence intervals.The procedure may use a fixed number of simulations, such as B = 1000, or stop adaptively after convergence.
  • Bootstrap inference relies on minimal assumptions but has no performance guarantees on small samples, so its behavior should be tested by simulation.
  • Partial likelihood estimates can equal complete-likelihood estimates under separability, remain consistent in broader additive models, and retain asymptotic normality.The latter result assumes omitted parameters have relatively small effects on intensity.

3.5 Bayesian Approaches

Bayesian approaches use either direct likelihood-based MCMC or latent cluster-process representations. The latter partitions events into Poisson-process components, reducing parameter dependence and improving sampling performance.

  • Direct MCMC uses Metropolis updates within a Gibbs sampler, but repeated likelihood evaluation costs O(n^2) and correlated parameters can hinder convergence.
  • A latent-variable approach partitions events into N + 1 sets that can each be treated as an inhomogeneous Poisson process.The background set uses µ(s), while other sets use intensities proportional to the triggering function g.
  • Partitioning the log-likelihood reduces dependence between parameters and dramatically improves sampling performance.The algorithm alternates sampling latent assignments and model parameters.

3.6 Model Selection and Diagnostics

Model selection uses information criteria, while diagnostics assess whether fitted conditional intensities adequately represent spatial structure and residual patterns. Thinning, super-thinning, and Voronoi residual maps provide complementary checks, with boundary effects requiring caution.

  • AIC is more effective than related criteria in small samples but less effective in larger samples when selecting the correct model.
  • Thinning events with probability p_i can yield a homogeneous Poisson process of rate b under the estimated intensity, enabling spatial homogeneity checks.The K-function can detect residual clustering not accounted for by the model.
  • Super-thinning adds a simulated inhomogeneous Poisson process so the resulting process is homogeneous with rate k when the estimated model is correct.The chosen rate satisfies b ≤ k ≤ sup_s,t λ(s, t).
  • Voronoi residual maps compare predicted and actual event counts over event-centered convex polygons instead of fixed grid cells.Voronoi tessellation avoids low-count skew in small cells and cancellation of over- and under-prediction in large cells.
  • A misspecified constant-background model produces positive residuals where the background rate is higher and negative residuals elsewhere.

4. APPLICATIONS

Self-exciting spatio-temporal point processes have been applied to earthquakes, crime, and epidemic forecasting because they model clustering, triggering, spatial variation, and temporal change. Applications include flexible ETAS earthquake models, crime intensities combining background risk with near-repeats, and epidemic models using covariates, marks, and intensity-dependent productivity.

  • Applications: The review covers earthquake models, crime forecasting, epidemic infection forecasting, and events on networks, while noting additional applications such as wildfires and civilian deaths.These applications illustrate features that make self-exciting point processes valuable.
  • Earthquake Aftershock Sequence Models: ETAS models capture earthquake clustering, spatial dependence, temporal trends, and spatial inhomogeneity through background and triggering components.Spatio-temporal extensions use varied triggering functions and can estimate an inhomogeneous background from spatial structure or stochastic declustering.
  • Earthquake Aftershock Sequence Models: AIC comparisons found spatial power-law triggering more effective than normal kernels, suggesting long-range aftershock triggering and magnitude-dependent triggering rates.Poor background fits can instead classify background events as triggered, overestimating the triggering rate and expected offspring number m.
  • Earthquake Aftershock Sequence Models: Nonstationary earthquake models use change points or continuously varying parameters, with penalized likelihood enforcing temporal smoothness; analyses found evidence of nonstationarity in a Japanese earthquake swarm.AIC was used to compare fits to earthquake series.
  • Crime Forecasting: Crime models divide conditional intensity into a spatially varying chronic background and a self-exciting component for near-repeats and retaliatory shootings.The background rate reflects spatial variation in offenders, targets, guardians, socioeconomic conditions, development, and policing; temporal variation can reflect weather and seasonality.
  • Epidemic Forecasting: Epidemic models can use population density, spatio-temporal covariates, event marks, Gaussian spatial triggering, and constant temporal triggering when direct transmissions are sparse.The infection model restricts triggering to previous infections within a fixed distance and time window; marks include patient age and bacterial strain.
  • Epidemic Forecasting: IMD applications used strain and age marks to compare epidemic potential across finetypes and spread behavior across age groups.The reported results show that self-exciting models can estimate epidemic behavior of IMD.
  • Epidemic Forecasting: Intensity-dependent productivity models represent epidemic spread with offspring number varying by conditional intensity rather than remaining constant.The recursive formulation is intended to reflect high infection rates among less-exposed populations and slower spread as prevalence and prevention increase.

5. CONCLUSIONS

The review presents self-exciting models as fast, flexible, and interpretable tools for clustered spatio-temporal processes, while emphasizing careful interpretation and open methodological problems. In infectious disease settings, unobserved infections can distort the background component and underestimate estimated transmission.

  • Conclusions: Self-exciting models are powerful when events form clusters triggered by common causes, and the review surveys estimation, declustering, and diagnostic tools across applications.The review highlights fast maximum likelihood and Bayesian estimation alongside model diagnostics.
  • Conclusions: Graphical diagnostics are not yet widely adopted, Bayesian hierarchical models remain an open direction, and network applications are still in their infancy.The proposed hierarchical setting includes multiple realizations such as crime data from different cities.
  • Conclusions: For infectious disease without known carriers, separating background and triggered cases can be conceptually misleading because the background component may represent unobserved infections.Improved case reporting would decrease the apparent importance of the background process.
  • Conclusions: Unobserved infections cause the estimated number m of infections triggered by each case to be underestimated because some triggered infections are absent from the data.This limitation follows from incomplete observation of the transmission process.
Loading 1708.02647v2…