Source-linked AI summary
Forecasting Global Volatility Across Asynchronous Markets: Incremental Accuracy from Constrained Cross-Market Attention
Xinlin Zhao, Haotian Qiao, Ziyao Lin
TL;DR
Global volatility forecasting must respect asynchronous exchange closures while determining whether cross-market information improves established benchmarks. The paper develops PGA-Trans-HAR, combining origin-admissible connectedness priors with constrained attention and HAR-anchored corrections. It reports broad but qualified gains, strongest at weekly and monthly horizons, without claiming universal dominance or causal identification.
Problem
Asynchronous exchange closures complicate the information set for multivariate international volatility forecasts, while unconstrained models and full-sample graphs risk overfitting or look-ahead bias.
Method
PGA-Trans-HAR combines an origin-admissible rolling ridge-VAR/GFEVD prior with constrained spatial attention, a fixed market gate, asymmetric masking, and a HAR forecast anchor.
Results
PGA-Trans-HAR lowers MSE and MAE versus HAR for all markets at h = 1 and h = 5 and seven of eight at h = 22, with best cross-market averages at longer horizons.
Takeaways & Limitations
Origin-aligned, structurally constrained cross-market information can improve volatility forecasts, particularly at medium and long horizons.
Takeaways & Limitations
The origin-admissibility design prevents look-ahead but does not establish structural causal identification.
Abstract
from arXiv · showhide
Multivariate volatility forecasting across international equity markets presents a fundamental information-set problem: asynchronous exchange closures dictate which market observations belong to the information filtration at any forecast origin. We investigate whether regularized, origin-admissible cross-market information yields incremental accuracy beyond established benchmarks. We develop PGA-Trans-HAR, combining an origin-admissible ridge-VAR/GFEVD connectedness prior with spatial self-attention. A time-invariant market gate governs their allocation, asymmetric attention masking prevents closed exchanges from transmitting spurious signals, and a direct-horizon HAR baseline anchors residual corrections. Using high-frequency data from eight major indices (2006--2022), we evaluate direct forecasts at 1-, 5-, and 22-day horizons across all-days and common-days panels, five-seed ensembles, structural ablations, HAC-adjusted Diebold--Mariano tests, and Model Confidence Sets. Relative to univariate HAR, the framework reduces MSE and MAE across all markets at daily and weekly horizons, and seven of eight monthly. Among linear and deep learning benchmarks, it achieves the lowest daily average MAE and the lowest weekly/monthly average MSE and MAE. Structural ablations show that spatial restrictions are essential: learned market gates improve accuracy over uniform weighting at medium-to-long horizons, while daily forecasts favor stronger scalar shrinkage. Disciplined, origin-aligned cross-market information yields genuine predictive gains, especially at medium and long horizons where structural spillovers persist.
1 Introduction
The paper addresses asynchronous-market information sets by combining origin-admissible economic priors with constrained attention around a frozen HAR benchmark. It finds qualified accuracy gains, especially at longer horizons, without establishing uniform model dominance.
- Motivation: Asynchronous closures misalign global volatility panels, forcing a choice between discarding open-market information and misreading absent observations as low volatility.The union calendar retains active-market dates while masks distinguish closures from observations.
- Motivation: Highly parameterized models risk overfitting low-signal volatility data, while full-sample relational graphs can violate out-of-sample information boundaries.These concerns motivate parsimonious, origin-aligned relational modeling.
- Research question: The central question is whether constrained, origin-admissible cross-market information adds statistically significant accuracy beyond an established HAR benchmark.HAR remains a frozen anchor rather than being replaced by an unconstrained black box.
- Approach: The framework combines forecast-origin alignment, a rolling ridge-VAR/GFEVD prior, data-driven spatial attention, and a time-invariant market-specific gate.The prior is refreshed by forecast origin, and the gate is fixed across observations, dates, source markets, and spatio-temporal blocks.
- Findings: PGA-Trans-HAR improves both MSE and MAE over HAR for every market at h = 1 and h = 5, and for seven of eight markets at h = 22.It also has the lowest average daily MAE and the lowest average MSE and MAE at weekly and monthly horizons.
- Findings: The evidence supports constrained cross-market information as useful particularly at longer horizons, but does not establish uniform model dominance or structural volatility transmission.The reported evidence is horizon- and loss-dependent.
2 Related Literature
The related literature spans HAR persistence models, connectedness-based relational priors, graph and attention architectures, and the asynchronous-calendar problem. This paper positions constrained, origin-aligned flexibility as a response to look-ahead risk and overfitting concerns.
- Volatility forecasting: HAR models approximate long-memory realized volatility with transparent daily, weekly, and monthly components and provide the study’s foundational benchmark.The benchmark is intended to test whether foreign-market information delivers demonstrable improvement.
- Volatility forecasting: VHAR and HAR-KS incorporate foreign-market predictors through linear regressors, helping distinguish cross-market information from nonlinear architectural effects.These alternatives support comparison against the univariate HAR baseline.
- Connectedness priors: GFEVD produces order-invariant variance shares that summarize directional predictive connectedness across source and destination markets.This compresses high-dimensional multivariate histories into a parsimonious spillover representation.
- Graph forecasting: Graph-based models propagate volatility through relational structures, but their out-of-sample validity depends on how those structures are estimated and temporally aligned.Full-sample or fixed-pre-sample graphs can conflict with forecast-origin information boundaries.
- Attention architectures: Self-attention captures complex dependencies flexibly, yet unconstrained neural architectures are hazardous in low-signal financial forecasting because they can overfit transient noise.The paper responds with restrictions on where and how flexibility enters the system.
- Asynchronous calendars: Asynchronous holidays create dates with partial market coverage, so common-day deletion discards active-market information while unified panels require explicit closure handling.Estimated relational objects must also use windows ending no later than the forecast origin.
- Evaluation: HAC-adjusted Diebold–Mariano tests and Model Confidence Sets account for overlapping multi-step losses and uncertainty in multi-model comparisons.Neither procedure treats a marginally lower average loss as universal predictive dominance.
3 The Forecasting System
The forecasting system restricts information by forecast origin, combines masked temporal and spatial channels with a rolling connectedness prior, and anchors neural corrections to a positive HAR forecast.
- System overview: The calendar mask, rolling connectedness matrix, fixed market gate, and frozen HAR forecast impose successive restrictions on information use.The architecture is designed to test whether disciplined cross-market information improves forecasting without relying on an unconstrained black box.
- Information set: For target t and horizon h, the forecast origin is o = t −h, and neural and HAR inputs use observations no later than o.The target remains a one-day model-scale volatility observation rather than an h-day aggregate.
- Connectedness prior: The rolling graph uses a half-open historical window ending no later than the origin, is reused for at most 20 origin endpoints, and follows the same schedule for h = 1, 5, and 22.This origin-indexed refresh rule prevents future-target information from entering the estimated prior.
- Constrained attention: Masked temporal and spatial attention prevents inactive markets from transmitting information while retaining their destination-query states.Masked-softmax rows are set to zero when no active source exists, with a corresponding defensive fallback for empty spatial sources.
- Market-specific allocation: The spatial mixture uses one learned gate per destination market, shared across dates, positions, source markets, and spatio-temporal blocks within each fitted horizon–seed model.Separate horizon–seed fits may estimate different gate vectors, but each vector is time-invariant after training.
- HAR-anchored fusion: The frozen HAR path anchors the forecast, while a bounded neural residual modifies it through an input-dependent gate and produces a strictly positive output.The inverse-softplus fusion initializes the forecast at the HAR anchor and limits the neural correction amplitude.
M HAR n
The node-balanced objective combines average HAR-relative performance with a smooth penalty for severe deterioration in any market.
- Node-balanced optimization: The log-mean-exp term smoothly approximates the largest HAR-relative deterioration while retaining a differentiable panel objective.With λ = 0.10 and τ = 0.20, the objective discourages severe degradation for one market.
- Node-balanced optimization: The optimization prioritizes panel-wide accuracy while limiting the risk that one market experiences a large HAR-relative deterioration.This is a training objective rather than an out-of-sample performance guarantee.
4 Experimental Design
The empirical design tests incremental accuracy, structural sources of gains, and calendar robustness using origin-safe splits, multiple benchmarks, ablations, and HAC-based inference.
- Research questions: The study asks whether PGA-Trans-HAR improves on established volatility models, whether gains reflect disciplined cross-market information, and whether rankings survive common-day evaluation.Model selection procedures avoid using test-period outcomes.
- Data and calendar panels: The dataset contains daily realized volatility for eight major equity indices from October 24, 2006, through June 28, 2022, across 4,079 union-calendar dates.Union-calendar coverage retains dates when at least one international exchange is active.
- Estimation protocol: Models are selected on a 60–10 development split, refit from fresh initialization on the first 70%, and evaluated once on the final 30%.The reported ensemble forecasts average predictions across five jointly seeded runs.
- Benchmarks and ablations: The benchmark hierarchy compares univariate HAR with temporal attention, dynamic prior graphs, fixed scalar fusion, and the complete trainable market-specific gate.The proposed model replaces g = 0.5 with gn = σ(γn).
- Benchmarks and ablations: Incremental gains are measured as percentage out-of-sample error reductions for nested MSE or MAE comparisons, with positive values favoring the augmented model.The transitions isolate temporal attention, prior graphs, scalar fusion, and learned gating.
- Data and calendar panels: Evaluation uses both an all-days panel preserving each market’s active dates and a common-days panel restricted to dates when all eight exchanges were open.This design separates broad union-calendar coverage from strict simultaneous-trading robustness.
- Inference: Inference uses HAC standard errors with a Newey–West Bartlett kernel to accommodate serial correlation from overlapping multi-step forecast errors.The bandwidth is adapted to the effective sample size of the loss-difference series.
5 Empirical Results
PGA-Trans-HAR delivers broad but horizon- and loss-dependent incremental accuracy over HAR and competing benchmarks, with strongest average performance at weekly and monthly horizons. Statistical tests and ablations support constrained cross-market information while showing that market-level dominance is not universal.
- Market-level comparison: PGA-Trans-HAR lowers both MSE and MAE for all eight markets at h = 1 and h = 5, and for seven markets at h = 22, relative to HAR.HSI deteriorates slightly at h = 22, by 0.9% under MSE and 0.3% under MAE.
- Cross-market averages: 0.180997 is the lowest cross-market average daily MAE for PGA-Trans-HAR, while HAR-KS has the lowest average daily MSE at 0.089199.PGA-Trans-HAR's daily MSE is 0.090442, above VHAR's 0.089624 and HAR-KS's 0.089199.
- Cross-market averages: 0.150782 MSE and 0.230454 MAE are PGA-Trans-HAR's lowest cross-market weekly averages, with the model ranking first on both metrics.It achieves the lowest MSE for four markets and the lowest MAE for six, but several benchmarks lead individual markets.
- Cross-market averages: 0.252361 MSE and 0.289352 MAE are PGA-Trans-HAR's lowest monthly cross-market averages, although its margins over Pure-ST-Transformer are approximately 0.14% and 0.18%.The monthly advantage is distributed across several markets rather than reflecting universal index-level dominance.
- Statistical evidence: At h = 5, PGA-Trans-HAR records lower MAE in 46 of 48 comparisons, including 27 favorable 5% rejections; five-day MSE favors it in 41 comparisons, with seven significant.At h = 1, 42 of 48 MAE and 35 of 48 MSE comparisons favor the model, while monthly horizons show 40 of 48 MAE and 39 of 48 MSE differentials favoring it.
- Statistical evidence: PGA-Trans-HAR is retained in the 90% MCS superior set across all markets for weekly MSE and both monthly metrics, but daily and weekly squared-loss separation remains limited.At h = 5 for MSE, all 11 models are retained for every market; at h = 1, retention is five markets under MSE and one under MAE.
6 Discussion
The discussion interprets the gains as arising from disciplined, regularized cross-market information rather than model complexity alone. Benefits are strongest at longer horizons, while fixed gates, point-forecast scope, and the limited panel constrain interpretation and generalization.
- PGA-Trans-HAR’s most consistent gains occur at weekly and monthly horizons, while daily rankings remain heterogeneous across losses and markets.
- The Value of Parsimony with Bounded Corrections: 0.357–0.461: estimated gating parameters remain below 0.5 across five seeds, preserving dominant reliance on the econometric prior.Mean attention weights are 0.389, 0.397, and 0.414 for h = 1, 5, and 22.
- The Value of Parsimony with Bounded Corrections: Learned market-specific gates improve uniform MSE and MAE at h = 5 and h = 22, whereas ungated combinations deteriorate accuracy at longer horizons.
- The Value of Parsimony with Bounded Corrections: The Pure-ST-Transformer’s higher daily and weekly losses but monthly competitiveness support a bias–variance interpretation of bounded corrections.
- Limits of the Information Set: The learned gate is a time-invariant regularization parameter, not a crisis indicator or causal explanation of market transmission.
- Limits of the Information Set: The asynchronous mask enforces information admissibility but cannot represent intraday closing sequences or identify structural causal integration.
- Limits of the Information Set: The study uses conditional point forecasts for eight equity indices during 2006–2022, limiting tail-risk assessment and post-2022 or cross-asset generalization.The framework also excludes several potentially informative conditioning variables and uses fixed prior settings across horizons.
7 Conclusion
PGA-Trans-HAR tests whether origin-admissible foreign-market information can add forecast accuracy beyond HAR while preserving strict information-set discipline. Its constrained architecture produces broad but horizon- and specification-dependent gains, supporting transparent rather than unconstrained use of neural flexibility.
- Conclusion: PGA-Trans-HAR combines four restrictions: an origin-admissible rolling ridge-VAR/GFEVD prior, learned market-specific gating, asymmetric calendar masking, and a frozen horizon-specific HAR residual anchor.The restrictions regulate both cross-market information flow and the neural correction while preserving a classical benchmark.
- Conclusion: The framework is designed as a controlled test that isolates the marginal predictive contribution of cross-market information rather than replacing econometric models with a black box.This framing makes the source of any predictive increment more transparent.
- Conclusion: Relative to univariate HAR, it reduces both MSE and MAE for all eight markets at daily and weekly horizons and for seven of eight markets monthly.The result is qualified by forecast horizon: monthly gains do not cover every market.
- Conclusion: Across linear and deep-learning benchmarks, it has the lowest average daily MAE and the lowest average MSE and MAE at weekly and monthly horizons.Pairwise HAC-adjusted Diebold–Mariano tests find significant improvements in numerous comparisons, especially under daily and weekly absolute loss.
- Conclusion: Ablations show that temporal attention and the rolling prior graph are not unconditional complements, while learned gating improves on uniform scalar shrinkage weekly and monthly but not daily.The preferred allocation between structural and data-driven components therefore depends on forecast horizon.
- Conclusion: The methodological implication is that economically interpretable information filtration and adjustment restrictions can govern flexible architectures in low-signal-to-noise forecasting environments.The rolling prior, asymmetric mask, market-level allocation, and HAR anchor make the predictive increment more transparent.