Source-linked AI summary
Adaptive estimation for Hawkes processes; application to genome analysis
Patricia Reynaud-Bouret, Sophie Schbath
TL;DR
The paper addresses detection of favored or avoided genomic distances by modeling events with Hawkes processes, where existing model-selection approaches are limited. It develops non-asymptotic penalized projection selection with oracle guarantees and studies adaptive estimation, practical penalty calibration, and genomic applications. The resulting procedure includes a sparse Islands strategy and yields estimates coherent with biological knowledge on real genomic data.
Problem
Existing genomic Hawkes-process analysis relies on approaches with limitations, while non-asymptotic model selection and minimax estimation had not been developed for this setting.
Method
The paper uses penalized piecewise-constant projection estimators over several model families, including the sparse Islands strategy, with a theoretically derived penalty.
Results
The estimator satisfies an oracle inequality, is minimax on the stated model class up to a logarithmic factor, and performs well in simulations and real genomic data.
Takeaways & Limitations
Sparse Islands models provide a practical way to identify localized favored or avoided distances while producing genomic results coherent with biological knowledge.
Takeaways & Limitations
The method uses piecewise constant estimators and theoretical controls exclude regimes where ν tends to zero or infinity or the excitation parameter approaches one.
Abstract
from arXiv · showhide
The aim of this paper is to provide a new method for the detection of either favored or avoided distances between genomic events along DNA sequences. These events are modeled by a Hawkes process. The biological problem is actually complex enough to need a non asymptotic penalized model selection approach. We provide a theoretical penalty that satisfies an oracle inequality even for quite complex families of models. The consecutive theoretical estimator is shown to be adaptive minimax for holderian functions with regularity in (1/2, 1]: those aspects have not yet been studied for the Hawkes process. Moreover we introduce an efficient strategy, named Islands, which is not classically used in model selection, but that happens to be particularly relevant to the biological question we want to answer. Since a multiplicative constant in the theoretical penalty is not computable in practice, we provide extensive simulations to find a data-driven calibration of this constant. The results obtained on real genomic data are coherent with biological knowledge and eventually refine them.
1. Introduction.
The paper models genomic event occurrences with Hawkes processes to detect favored or avoided distances, addressing limitations of existing likelihood-based model selection. It develops a non-asymptotic penalized approach with theoretical guarantees and practical estimators.
- Motivation: Genomic event occurrences are modeled as points on the real line so favored or avoided distances can be identified through the interaction function h.Positive values of h at a distance indicate self-excitation, while negative values indicate self-inhibition.
- Limitations of existing methods: FADO uses maximum likelihood estimation with splines and AIC, but its equally spaced knots and support specification limit model flexibility.Its theoretical guarantees require a fixed candidate family and an existing true knot set as sequence length grows.
- Contribution: The paper develops non-asymptotic penalized model selection for Hawkes processes because Gaussian-model heuristics do not determine an appropriate penalty here.The method targets oracle inequalities and minimax adaptivity, which had not previously been studied for Hawkes models according to the authors.
- Method: The theoretical procedure uses penalized projection estimators for nonnegative h and piecewise constant functions, while practical calibration addresses unknown penalty constants.These restrictions reduce the gap between the theoretical and practical procedures but limit the estimator family.
- Scope: The framework is kept general enough for Hawkes applications beyond genome analysis, including earthquake, financial, and economic event data.The biological application motivates the specific setting, but the formalism is not restricted to genomic events.
2. Framework.
The framework estimates a Hawkes intensity with projection estimators over selectable interval models, then penalizes model complexity to obtain an oracle-style choice. It includes nested, regular, irregular, and sparse Islands strategies, while imposing explicit stability and parameter-scope conditions.
- Framework: The estimator observes a stationary Hawkes process on [-A,T] and estimates an intensity parameterized by a baseline and a function supported on (0,A].The support bound A is treated as known and represents the maximal distance considered relevant for linear interaction.
- 2.1. Projection estimator.: Projection estimators approximate the interaction function with piecewise constant functions on interval models, balancing approximation bias against stochastic error.The bias is associated with projection error, while the variance term scales with model dimension divided by T, up to a slowly varying factor.
- 2.2. Penalized projection estimator.: Penalized projection selection chooses among candidate models so the selected estimator is nearly as accurate as the best projection estimator in the family.The target guarantee is an oracle inequality, obtained with high probability or in expectation up to residual and multiplicative terms.
- 2.3. Strategies.: The strategies range from nested and regular partitions to irregular partitions with unknown cuts and Islands models designed for sparse, localized support.The irregular family has approximately 2^N models, while Islands represents nonzero interaction intervals as separated islands within a larger support.
- 2.4. Self-inhibition.: The theoretical treatment assumes a self-exciting model with nonnegative h, although the estimators can also be computed when h permits self-inhibition.For self-inhibition, stationarity requires |h| < 1, but the paper does not provide a full theoretical analysis because of technical difficulties.
- 2.4. Self-inhibition.: The theoretical controls exclude regimes where ν approaches zero or infinity or the excitation parameter approaches one, because point counts then vanish or explode.These restrictions define the parameter subset on which projection-estimator behavior can be controlled.
3. Main results.
The paper develops non-asymptotic risk bounds and penalized projection model selection for Hawkes-process intensity estimation. The resulting procedures provide oracle guarantees, near-sharp risk control, and minimax adaptivity, while practical use requires calibrating an unknown penalty constant.
- Projection estimation: For a fixed model containing s, the estimator is consistent non-asymptotically, with convergence rate smaller than log(T)/T, and the model may depend on T under the stated condition.The result controls risk without requiring the model to remain fixed as the observation length grows.
- Projection estimation: The clipped projection estimator has a risk upper bound combining approximation bias with a stochastic term, and the variance contribution is nearly sharp.When the target belongs to the selected model, the bias disappears and the variance term is bounded by a constant times |m| log(T)/T.
- Risk bounds: The clipped estimator’s risk is bounded below by a constant times |m|/T and above by |m| log(T)/T, leaving only a logarithmic gap caused by Hawkes-process intensity unboundedness.The upper-bound logarithm arises from controlling the process intensity over the observation interval.
- Model selection: A penalized projection estimator satisfies an oracle inequality, selecting a model whose performance is close to the best projection estimator in the candidate family.The theoretical penalty has order |m| log(T)^2/T, up to a multiplicative constant depending on model and process parameters.
- Practical implementation: The theoretical penalty’s multiplicative constant depends on unknown quantities such as H, η, and P, so it is not directly computable in practice.The paper therefore motivates an implementable calibration strategy based on a dimension-scaled penalty.
4. Practical data-driven calibration via simulations.
The simulations calibrate practical penalties for Regular, Irregular, and Islands strategies, comparing them with hold-out procedures under computationally constrained model families. Methods 1, 2, and 4 generally achieve the best risks, near-oracle performance, and dimension recovery, with Islands especially effective for localized structure.
- Calibration: The practical study restricts simulations to models with at most 15 intervals and excludes Nested because its family would contain only five models.The theoretical penalty's multiplicative constant is not directly computable, so simulations are used to calibrate practical constants.
- Results: Methods 1, 2, and 4 have the lowest risks across increasing T and are the most precise methods in the rescaled-risk experiments.Risk decreases as T increases, and rescaled risk decreases as either T or c increases.
- Results: The oracle ratio is close to 1 for Methods 1, 2, and 4 at T = 500000, indicating performance near the best estimator in their respective families.For the rescaled-risk setting, the oracle ratio reaches 1 when T = 500000 and c = 0.8 for the three favorite methods.
- Calibration: The angle penalty automatically locates the contrast curve's dimension-changing angle for the Irregular and Islands strategies.This provides a practical selection rule without manually inspecting the contrast curve.
- Results: For negative jumps and smooth functions, Islands more precisely detects localized fluctuations, although piecewise-constant estimators cannot closely match smooth functions.Method 4 has a slight risk advantage for the negative function and identifies spike positions more clearly than Method 1; Method 2 misses some smaller features.
5. Applications on real data.
The method detects favored and avoided genomic distances in two E. coli datasets, producing patterns coherent with biological knowledge and a more compact interaction range than FADO.
- Gene occurrences are uncorrelated below 2600 basepairs, avoided at 0–500 bps, and favored at approximately 700–2000 bps.
- Shortening the support reveals a negative effect below 250 bps and a positive effect around 1000 bps, consistent with nonoverlapping genes and compact bacterial genomes.
- Motif tataat shows favored distances below 600 bps, around 0–1500 bps, and near 3000 bps, alongside no correlation down to 4000 basepairs.
- For tataat, shortening support identifies favored distances below 300 bps, around 1000 bps, and around 3500 bps, consistent with motif self-overlap and promoter biology.
- For the gene dataset, Islands selects a four-part model versus FADO’s 15-part model and indicates no significant interaction after 3000 bps.
- The theoretical Nested strategy is adaptive minimax, although it was not implemented for technical reasons.
6. Minimax properties.
The paper establishes minimax behavior for clipped projection estimators and shows that penalized Nested selection adapts to unknown Hölder regularity over the range a ∈ (1/2, 1].
- Clipped projection estimators are minimax over the stated Hölder classes for regularity a ∈ (1/2, 1], up to a logarithmic factor.
- The clipped penalized estimator with the Nested strategy is adaptive minimax over these classes up to a logarithmic term.
- Irregular and Islands strategies provide separate model families for sparsity in partitions and sparsity in the support of h.
- The estimator’s rate matches 1/T up to logarithmic factors for the considered sparse-function settings.
7. Technical results.
The technical analysis derives oracle inequalities for practical penalized projection estimators by controlling Hawkes-process fluctuations on a suitable observable event.
- With probability at least 1 − 3#{M_T}e^−x, the theorem supplies simultaneous inequalities for every model in the candidate family.
- On the event B, local point counts are controlled and the random empirical norm is equivalent to the deterministic norm, enabling dimension-based penalties.
- Controlling B requires carefully choosing N, R, r, and the model family, preventing treatment of very high-complexity families.
- The technical assumptions imply the main oracle result once T exceeds a threshold depending on the model and process parameters.
8. Proofs of the technical and minimax results.
The proofs connect Hawkes-process concentration and norm equivalence to the oracle and minimax results, while the final construction combines model selection with the Islands strategy for genomic interactions.
- The proof uses norm equivalence, contrast minimality, and concentration inequalities for Hawkes processes to establish the projection-estimation bounds.
- The technical proof controls empirical deviations by stopping the process on B and applying concentration bounds to finite-dimensional model spaces.
- The resulting model-selection method is adaptive minimax for certain function classes and supports data-driven penalty calibration in practice.
- The Islands strategy targets sparsity in the support of h, matching the biological goal of identifying interaction ranges.
9. Conclusion.
The conclusion identifies two extensions needed for the Islands strategy and calls for testing whether the interaction function h is nonzero.
- The Islands strategy should be extended to interactions between another pair of event types, such as promoters and genes.
- A test procedure is needed to determine whether h is nonzero, which is equivalent to testing for the existence of an interaction.
- The paper acknowledges contributions from collaborators and referees.