Source-linked AI summary

A toolbox for fitting complex spatial point process models using integrated nested Laplace approximation (INLA)

Janine B. Illian, Sigrunn H. Sørbye, Håvard Rue

arXiv:1301.1817v1stat.AP

TL;DR

Complex spatial point-pattern data often require joint models for locations, marks, and covariates, but routine Cox-process fitting remains limited. The paper develops log-Gaussian Cox-process models with constructed covariates and fits them using INLA, yielding a toolbox for efficient fitting, comparison, and assessment. Simulations and realistic examples demonstrate the approach's applicability, while its use requires care with model complexity and spatial scales.

  • Problem

    Routine fitting of complex spatial point process models and joint models involving marks or covariates remains limited in the literature.

  • Method

    The paper uses log-Gaussian Cox-process latent Gaussian models with constructed covariates for local interaction and INLA for parameter estimation and model fitting.

  • Results

    The approach provides a toolbox for fitting, comparing, and assessing realistically complex spatial point-pattern models, including models with large-scale covariates, local clustering, and dependent marks.

  • Takeaways & Limitations

    Complex spatial point-pattern models can be made accessible to scientists outside statistics through the methodology and the R-INLA library.

  • Takeaways & Limitations

    The approach requires careful choice of spatial scales and constructed covariates because scale mismatch and data-derived covariates can lead to overfitting.

Abstract

from arXiv · show

This paper develops methodology that provides a toolbox for routinely fitting complex models to realistic spatial point pattern data. We consider models that are based on log-Gaussian Cox processes and include local interaction in these by considering constructed covariates. This enables us to use integrated nested Laplace approximation and to considerably speed up the inferential task. In addition, methods for model comparison and model assessment facilitate the modelling process. The performance of the approach is assessed in a simulation study. To demonstrate the versatility of the approach, models are fitted to two rather different examples, a large rainforest data set with covariates and a point pattern with multiple marks.

1. Introduction.

Routine fitting of complex spatial point process models remains limited despite accessible statistical software and increasingly rich spatial data. The paper addresses this gap by developing complex joint models and methods for fitting, comparing, and assessing them.

  • 1.1. Complex point process models.: Routine fitting of complex spatial point process models remains in its infancy despite rapidly improving technology.
  • 1.1. Complex point process models.: Spatially explicit data are increasingly available, but analyses often fail to use the full spatial information they contain.
  • 1.1. Complex point process models.: Real data commonly combine point locations with dependent qualitative or quantitative marks and spatial covariates, motivating complex joint models.
  • 1.1. Complex point process models.: Existing Gibbs-process approaches generally handle simpler models, while approximation reliability decreases as interaction strength increases.
  • 1.1. Complex point process models.: Cox-process fitting, model comparison, and model assessment have received little attention, especially for joint models of patterns with marks or covariates.
  • 1.1. Complex point process models.: The paper models local and larger-scale spatial behavior together using constructed covariates and spatial effects, with simulation and realistic data examples.

2. Methods.

The methods represent complex spatial patterns, marks, and covariates with log-Gaussian Cox-process latent Gaussian models. Constructed covariates capture local interaction, while INLA, model comparison, and spatial effects support efficient fitting and assessment.

  • 2. Methods.: Log-Gaussian Cox processes model spatial patterns through a latent Gaussian random field governing a stochastic intensity.
  • 2. Methods.: The proposed models jointly represent spatial patterns with marks and covariates, while incorporating both small- and large-scale spatial behavior.
  • 2. Methods.: Because the models are latent Gaussian models, INLA provides a route to parameter estimation and model fitting.
  • 2. Methods.: INLA approximates marginal posteriors by nesting approximations over latent variables and low-dimensional hyperparameters, then integrating numerically.
  • 2. Methods.: The Gaussian approximation uses the conditional mode and curvature, substantially reducing fitting time when applied to latent Gaussian models.
  • 2. Methods.: A constructed covariate summarizes local inter-individual behavior, and its smooth effect can represent nonlinear spatial dependence.
  • 2. Methods.: The nearest-point-distance covariate can model either local repulsion or local clustering without Gibbs-process intractable normalizing constants.
  • 2. Methods.: DIC supports comparisons among models, while structured and unstructured spatial effects help assess unexplained structure and extreme observations.

3. Using a constructed covariate to account for local spatial structure— a simulation study.

The simulation study evaluates constructed covariates across repulsive, clustered, and inhomogeneous clustered patterns. They identify local structure, but fit repulsive patterns more accurately than clustered ones, where clustering is underestimated.

  • Simulation design: The study evaluates repulsion, clustering, and small-scale clustering alongside larger-scale inhomogeneity using simulated point patterns.Model evaluation compares patterns generated from fitted models with example patterns using simulated-pattern characteristics and L-functions.
  • Repulsion: For repulsive patterns, the functional relationship identifies the interaction distance as 0.05 and levels out beyond that distance.Higher constructed-covariate distances correspond to higher intensities, reflecting repulsion, while larger distances show no relationship with intensity.
  • Repulsion: The fitted repulsion model produces patterns and L-functions similar to the original, with the original L-function remaining within simulation envelopes.The envelopes use 50 simulated patterns and 100,000 Metropolis iterations for each pattern.
  • Clustering: For clustered patterns, the constructed covariate introduces clustering, but fitted patterns have fewer and less distinct clusters and weaker local clustering than the originals.The simulation envelopes do not include the true L-function, indicating inadequate recovery of the clustering strength.
  • Inhomogeneous clustering: With large-scale inhomogeneity, the model captures small-scale clustering and larger-scale spatial behavior, but simulated patterns show slightly weaker local clustering than the originals.The fitted model’s mean L-function lies near or partly outside the simulation-envelope upper edge, indicating insufficient clustering strength.
  • Overall findings: Across scenarios, local spatial structure is identifiable, repeated simulations give essentially the same results, and homogeneous Poisson data produce no spurious clustering or regularity.The approach therefore handles regularity and clustering, although it performs clearly better for repulsive than clustered patterns.

4. Joint model of a point pattern and environmental covariates.

The paper jointly models a large rainforest point pattern and sparsely observed environmental covariates, combining broad-scale covariate effects with constructed-covariate local clustering. Applied to 7,416 Aporusa microstachya trees, the model separates spatial behavior across scales and supports assessment of remaining structure.

  • Model formulation: The joint model uses original, noninterpolated covariate data, accounts for measurement error, and links covariates with the point pattern through joint spatial fields.An additional spatially structured effect captures pattern structure unexplained by the covariate fields.
  • Model formulation: The model combines spatial effects for phosphorus and magnesium with a constructed covariate for local clustering and additional unstructured fields for measurement or sampling error.The constructed covariate targets clustering at a smaller scale than environmental covariate associations.
  • Rainforest data: The Pasoh Forest Reserve application analyzes 7,416 Aporusa microstachya individuals with environmental covariates measured at 83 locations distinct from tree locations.The analysis uses a 50 ha forest dynamics plot in Peninsular Malaysia.
  • Results: DIC increases from 15,379 to 15,440 when the two empirical covariate terms are omitted, indicating that those terms explain some spatial structure.The estimated phosphorus and magnesium effects are very smooth, although covariate information is available in only 83 grid cells.
  • Results: The constructed covariate accounts for clustering up to 15 metres, and including it removes local structure from the estimated spatial effect, making that effect smoother.Without the constructed-covariate term, the estimated spatial effect shows clear local structure and may overfit the observed pattern.
  • Model assessment: Model assessment reveals remaining spatial structure that the current model cannot explain, suggesting that further covariates could improve the model.The approach also provides model comparison and assessment methods for practical use.

5. Modeling marks and pattern in a marked point pattern with multiple marks.

The section develops a joint marked Cox process model for a spatial pattern and two dependent marks, then applies it to koala tree data. The results show how shared spatial effects and mark dependencies support interpretation of aggregation, palatability, and visitation frequency.

  • Model: The model jointly represents tree locations and two nonindependent marks, with the first mark depending on pattern intensity and the second depending on both intensity and the first mark.The framework uses exponential-family distributions and can be generalized beyond two marks.
  • Koala data: The koala data cover 915 eucalyptus trees with leaf chemistry and monthly frequency-of-visit information collected between 1993 and March 2004.Leaf chemistry is summarized by palatability, while frequency marks describe diurnal tree use.
  • Model specification: The fitted model uses a normal distribution for leaf marks and a Poisson distribution for frequency marks, with both linked to the spatial pattern.The analysis uses 1571 grid cells and applies a simple edge correction for the constructed covariate.
  • Results: DIC decreases to 6943 after adding the spatial effect, whereas the constructed covariate does not improve fit for this data set.The estimated constructed-covariate function is not significantly different from 0, consistent with weak local clustering.
  • Results: The common spatial effect captures shared spatial autocorrelation, while the posterior mean for β4 is 1.38, indicating a significant positive influence of palatability on visit frequency.The negative β2 estimate indicates lower palatability where trees are aggregated, and positive β3 reflects higher koala presence where intensity is higher.
  • Implications: Marked Cox process models can treat marks as the primary scientific focus while accounting for spatial dependence through a joint model.The approach is presented as applicable to related marked data sets, including ecological data with locations and properties.

6. Discussion.

The discussion positions the methodology as a flexible, accessible framework for fitting and assessing realistic spatial point-process models. It also identifies important scope boundaries, including limited handling of inter-individual interactions and reliance on a discretized spatial field.

  • Model assessment: Structured spatial effects can reveal unexplained spatial correlations and help identify suitable covariates, while unstructured effects can identify extreme observations and measurement-error issues.These effects support assessment of model fit beyond parameter estimation.
  • Context: Exploratory summaries such as Ripley’s K-function and the pair-correlation function characterize spatial behavior across scales, but increasing data complexity makes suitable summaries less obvious.Likelihood-based modeling remains useful when data include additional marks or covariates.
  • Scope boundary: Unlike Gibbs processes, log-Gaussian Cox processes do not allow second-order inter-individual interactions, making them unsuitable when those interactions are the primary interest.The approach instead represents spatial trend as a latent random field in a hierarchical, doubly stochastic structure.
  • Computational representation: The continuous Gaussian random field is approximated by a discrete Gauss Markov random field to make model fitting feasible.The discussion argues this approximation is justified through an explicit link between covariance functions and Gauss Markov random fields.
  • Practical contribution: The methodology and R-INLA make complex spatial point-process models accessible for routine fitting and fit assessment with little computational effort.The framework accounts for both local and global spatial behavior and permits joint models containing marks or covariates.
Loading 1301.1817v1…