Source-linked AI summary
Weighted Empirical Risk Minimization for Machine Learning under Long-Range Dependence: Exact Pathwise Rates and Learning-Error Geometry
Elina Moldavskaya
TL;DR
The paper asks how regularly varying sample weights affect exact almost-sure learning fluctuations for smooth ERM under long-range dependence. It reduces finite-lag scores to weighted Hermite chaos and propagates the resulting LIL through risk curvature. Chaos rank and memory determine polynomial rates, while weights change sharp pathwise constants and cluster geometry; rank-one optimization favors strictly positive exponents.
Problem
The paper studies exact almost-sure learning behavior for regularly weighted ERM trained on finite-window observations from long-range dependent Gaussian data.
Method
It combines finite-lag score reduction with weighted LIL theory and smooth-risk expansions to obtain pathwise learning-error and trajectory results.
Results
The chaos rank m determines parameter-error exponent −αm/2 and excess-risk exponent −αm, while weighting changes the exact LIL constant and nonlinear cluster geometry.
Takeaways & Limitations
Long memory and chaos rank set polynomial learning rates, whereas regularly varying weights provide a way to compare sharp persistent pathwise fluctuations.
Takeaways & Limitations
The results do not provide finite-sample anytime guarantees and do not automatically extend to stochastic gradient descent.
Abstract
from arXiv · showhide
We develop an exact almost-sure learning theory for smooth parametric models trained by regularly weighted empirical risk minimization on long-range dependent data. The training observations are generated from a fixed finite window of a stationary Gaussian sequence, and the sample weights are regularly varying. If the loss gradient at the population minimizer has Wiener-chaos rank $m$ and a nonzero low-frequency coefficient, then, in the long-memory interior regime, the finite-lag score reduces on the iterated-logarithm scale to a single weighted Hermite chaos. This yields an almost-sure Bahadur representation, an exact limsup law for the learned parameter, and, for $m\ge2$, the functional cluster set of the complete learning trajectory. The polynomial learning exponent is determined by the memory parameter and the chaos rank and is invariant under the admissible power weighting, whereas the sharp pathwise constant and cluster geometry depend on the weights. In the rank-one case, global optimization over the admissible power exponents shows that every optimizer is positive. Time-series prediction and classification examples illustrate the results.
1. Introduction
The paper develops exact almost-sure learning theory for regularly weighted ERM on long-range dependent Gaussian data, where chaos rank controls leading fluctuations. It connects finite-lag score reduction and weighted LIL behavior to parameter learning, weighting effects, and illustrative machine-learning models.
- Learning setting: The model trains a smooth parametric predictor on finite-window observations from a stationary Gaussian sequence with regularly varying sample weights.Weights are a_t = t^-κ L_a(t), with ordinary ERM recovered when a_t ≡ 1.
- Learning setting: The loss gradient’s first nonzero Wiener chaos, with rank m and low-frequency coefficient J, determines the leading long-memory fluctuations.The finite-lag gradient is formed from the Gaussian window and its first nonzero chaos governs dependence.
- Motivation and related work: The paper targets exact almost-sure normalizations, limsup constants, and functional cluster geometry beyond existing weak-dependence and distributional results.Its probabilistic input is a weighted LIL for nonlinear functionals of long-memory Gaussian sequences.
- Main contributions: A finite-lag reduction connects the scalar weighted LIL to learning by replacing the leading score with a coefficient J times a weighted Hermite chaos.This bridge supports the subsequent pathwise Bahadur representation and learning-error results.
- Main contributions: The learning error has a Bahadur-type pathwise representation, with leading direction V^-1J and functional cluster behavior for m ≥ 2.Here V is the population-risk Hessian at the minimizer.
- Examples and experiments: Examples cover one-step prediction, lagged threshold classification, and a quadratic-target model with a symmetry-induced transition between chaos ranks.Numerical experiments illustrate weighting, one-dimensional learning-error geometry, and rank transitions.
3. Finite-lag reduction for the weighted learning score
The finite-lag reduction shows that a fixed Gaussian lag window changes the leading weighted learning score only through the deterministic coefficient J. Lag effects and higher chaoses are negligible on the iterated-logarithm scale, enabling exact score LILs and one-dimensional leading geometry.
- Reduction principle: A fixed lag window changes the leading score only through the deterministic low-frequency coefficient J.All remaining lag effects and higher Wiener chaoses are negligible on the iterated-logarithm scale.
- Scaling: The weighted reference variance is regularly varying with index H = 1 − κ − αm/2 in the long-memory interior regime.This variance scaling supplies the normalization used in the reduction and subsequent LIL results.
- Reduction principle: The score decomposes into J times the weighted Hermite chaos plus lag and higher-chaos remainders.The lag component remains in the mth chaos, while the higher-chaos remainder contains orders at least m + 1.
- Uniform control: The remainders vanish almost surely after normalization, including uniformly over interpolated score trajectories on every fixed time interval.The uniform reduction transfers relative compactness and cluster-set results from the scalar reference statistic.
- Geometry: Even for multidimensional parameters, the leading score fluctuations are confined to the deterministic direction J.This rank-one pathwise geometry follows from scalar Gaussian input and finite-lag subordination, not from a one-dimensional parameter restriction.
- LIL consequences: The normalized learning score has an exact weighted LIL and, for m ≥ 2, a functional cluster set inherited from the scalar weighted Hermite process.For m = 1, the corresponding result uses the sharp linear weighted LIL.
4. Exact LIL for the learning error
The paper transfers weighted score asymptotics to the empirical-risk minimizer, obtaining an LIL-scale Bahadur representation and exact learning-error limits. The leading error is one-dimensional, with parity-dependent endpoint cluster geometry.
- 4. Exact LIL for the learning error: The LIL-scale Bahadur representation transfers the weighted score asymptotics to the learned parameter.The result relies on weighted empirical-risk consistency, local curvature, and Taylor expansion.
- 4. Exact LIL for the learning error: For m≥2, the exact LIL also identifies the full endpoint cluster set of the normalized learning error.This follows from the sharp weighted LIL and the endpoint cluster theorem for the reference statistic.
- 4. Exact LIL for the learning error: The leading multidimensional learning error lies in the deterministic direction V−1J, while orthogonal directions are negligible on the LIL scale.Thus the pathwise learning geometry is effectively one-dimensional even for vector-valued parameters.
- 4. Exact LIL for the learning error: The endpoint cluster set is one-sided for even m and symmetric for odd m.The parity distinction follows from the sign structure of the underlying Hermite-chaos kernel.
5. Functional LIL for the learning trajectory
The paper extends the learning-error LIL from endpoints to complete interpolated trajectories. The normalized paths are relatively compact, and terminal normalization produces a weight-dependent profile while preserving the underlying polynomial learning profile.
- 5. Functional LIL for the learning trajectory: The normalized weighted learning trajectories are almost surely relatively compact in C([0,T];R^p) for every fixed T>0.Their cluster set is inherited from the functional score LIL through the map f ↦ −V−1f.
- 5.1. The normalized learning curve: The normalized learning curve is relatively compact on [ε,T], with terminal normalization multiplying the trajectory by t^(κ−1).The restriction excludes zero because t^(κ−1) is singular there under κ<1.
- 5.2. Pointwise learning envelope: Weighting changes the sharp constant and cluster geometry but not the polynomial learning profile t^(−αm/2).The power profile is independent of κ after combining the functional normalization with the pointwise scaling.
- 5. Functional LIL for the learning trajectory: The same functional LIL transfers through differentiable functionals, including prediction-error cluster sets for fixed feature vectors.The effective direction becomes DF(θ0)V−1J.
6. Learning rates, excess risk, and optimal weighting
The paper makes the learning and excess-risk scales explicit and studies how regularly varying weights affect their exact pathwise constants. Memory and chaos rank determine the polynomial exponent, while rank-one optimization favors strictly positive power weighting.
- 6.1. Exact learning and excess-risk rates: The exact learning rate has polynomial exponent −αm/2 for every admissible regularly varying weight.The sharp constant additionally depends on the weight through W_m,α(κ), while the direction is V−1J.
- 6.1. Exact learning and excess-risk rates: The excess population risk satisfies an exact LIL on the corresponding squared learning-error scale.For m≥2, its cluster sets follow from the parameter cluster sets through a continuous quadratic map.
- 6.3. Global optimization for m=1: For m=1, every globally optimal admissible power exponent is strictly positive.Negative exponents and the unweighted benchmark κ=0 do not minimize the exact rank-one LIL constant.
- 6.3. Global optimization for m=1: The optimal rank-one exponent lies in the interior of the admissible interval, and uniqueness is not asserted.The admissible set is κ<(1−α)/2, while the theorem excludes both κ=0 and the right endpoint as minimizers.
7. Machine-learning examples and verification of the assumptions
Three smooth learning examples verify how loss-gradient chaos rank determines pathwise learning behavior. One-step prediction has rank two, threshold classification rank one, and the quadratic target switches rank at exact symmetry.
- One-step prediction: One-step prediction has chaos rank two because its score is quadratic in the Gaussian data despite linear parameterization.Its exact learning rate therefore follows the m=2 weighted LIL.
- Threshold classification: Threshold classification has a nonzero first-chaos projection, yielding rank-one parameter and excess-risk learning laws.The model uses V=I2 and a nonzero low-frequency coefficient Jcls.
- Weighting: For the rank-one classifier, a strictly positive admissible exponent κ★ makes the exact constants smaller than at the unweighted choice.The result applies for every α∈(0,1).
- Quadratic target: For the quadratic target, s≠0 gives rank one with Js=2s, whereas exact symmetry s=0 gives rank two with J0=−1.The corresponding learning exponents change from −α/2 to −α at symmetry.
- Examples: The examples cover one-step prediction, threshold classification, and a quadratic target, with their ranks, coefficients, and learning scales summarized in Table 1.The corresponding sharp constants are given in the cited corollaries.
8. Numerical experiments
Numerical experiments evaluate weighting objectives and finite-sample learning geometry. Positive weighting produces modest sharp-constant improvements, while negative weighting is more costly and simulations display the predicted rank-dependent geometry.
- Weighting objective: At α=0.2, the best positive weights reduce the sharp constant by 0.100%, 0.205%, and 0.315% for m=1,2,3, respectively.The corresponding minimizers are κ★=0.1361, 0.1310, and 0.1250.
- Weighting objective: At κ=−0.6, the sharp constant increases by 1.44%, 3.10%, and 5.04% for m=1,2,3, respectively, relative to κ=0.Negative weighting is therefore substantially more costly than the modest positive-weighting gains.
- Learning geometry: In threshold classification, RMSE⊥/RMSE∥ decreases from about 0.46 at n=2^8 to about 0.17 at n=2^16 across all four weights.The fitted slopes range from −0.163 to −0.175, providing finite-sample evidence for rank-one geometry.
- Rank transition: For the quadratic target, fitted slopes move from −0.370 at s=0 toward −0.231 at s=0.6, showing crossover from rank-two to rank-one behavior.The asymptotic exponents are −α=−0.4 at symmetry and −α/2=−0.2 for fixed s≠0.
- Interpretation and scope: The theory separates polynomial learning exponents from weighting-dependent sharp constants and cluster geometry, while the LIL itself remains asymptotic rather than finite-sample anytime-valid.The results concern ERM and do not automatically extend to SGD.
2. Pathwise delta method for smooth functionals
The pathwise delta method transfers normalized learning-curve limits through smooth functionals. It maps parameter cluster sets linearly and yields corresponding prediction-error and parity results.
- Pathwise delta method: A continuously differentiable functional F maps the normalized parameter learning curve through its derivative at θ0 on the same pathwise scale.The resulting normalized functional curve has the image cluster set under DF(θ0).
- Pathwise delta method: The effective leading direction is obtained by applying the functional derivative to the parameter trajectory's leading direction.The construction relies on uniform control of Taylor remainders over compact time intervals.
- Prediction functionals: For a smooth prediction functional, the same mapping gives a prediction-error cluster set determined by the corresponding derivative and parameter cluster set.The result applies to fixed feature vectors for which θ↦fθ(u) is continuously differentiable.
- Parity: For even chaos rank, all nonzero pointwise normalized prediction-error cluster values have the same sign, while opposite-direction excursions are negligible on the LIL scale.This is a pathwise cluster-geometry statement, not a claim about finite-sample bias.
3. Functional LIL for the excess population risk
The excess population-risk trajectory is obtained by applying the quadratic curvature map to the normalized parameter trajectory. Its functional cluster set and pointwise envelope follow from this quadratic transformation.
- Functional excess risk: The excess-risk scaling is the squared parameter scale, and its cluster behavior follows from the local curvature V at θ0.Continuity of the Hessian controls the Taylor remainder.
- Functional excess risk: The normalized excess-risk trajectory has a functional cluster set equal to the quadratic image of the parameter learning-curve cluster set.The map is QV(h)=1/2 h^T Vh, with interpolation errors controlled uniformly.
- Pointwise envelope: At each fixed t>0, the pointwise excess-risk cluster set is an interval whose upper endpoint equals the parameter LIL limsup constant after the quadratic map.Squaring removes the parity distinction between one-sided even-rank and symmetric odd-rank parameter cluster sets.
4. Continuity and bounds for the weighting objective
For fixed chaos rank and memory parameter, the weighting objective is continuous, finite, and strictly positive throughout its admissible power-exponent domain. The available bounds establish regularity but do not locate the minimizer.
- Continuity: W_m,α(κ) is continuous on D_m,α when m≥2 and αm<1.The proof uses positivity and continuity of the beta-function representation together with dominated convergence on compact subsets.
- Bounds: W_m,α(κ) is finite and strictly positive for every fixed admissible κ.The kernel representation and two-sided estimates provide these bounds throughout D_m,α.
- Interpretation: The bounds establish regularity but are not intended to identify the minimizer, which is computed directly in the numerical section.Optimization therefore requires evaluating the weighting objective rather than relying on the continuity and bound results alone.
5. Digamma estimates for rank-one weighting
The rank-one weighting analysis derives sign properties for a digamma-based auxiliary function and uses them to support the global positivity result for optimal power weighting. The accompanying checks verify the least-squares assumptions and identify chaos ranks in the examples.
- Digamma sign estimates: D_α(κ)<0 for κ≤0, while lim as κ approaches (1−α)/2 from below D_α(κ)>0.These endpoint sign changes supply the key sign information used in the rank-one weighting argument.
- Digamma sign estimates: The sign proof reduces to sinh(βu)<β sinh(u), obtained from strict convexity of sinh with sinh(0)=0.This transformation establishes the inequality needed for the nonpositive-exponent region.
- Model verification: The common linear least-squares verification establishes the model conditions under finite fourth moments and positive definiteness of A.It also gives existence of a measurable constrained minimizer through continuity and measurability of the weighted empirical risk.
- Example verification: The prediction example has score chaos rank m=2 with nonzero coefficient J_pred=ρ_1−1.The result follows because the score lies entirely in the second Wiener chaos and |ρ_1|<1 under the stated dependence condition.
- Example verification: The quadratic-target example uses ζ=1, Y=(X_{t−s})^2, with A=1 and b_ζ=1+s^2.The target has moments of every order, supporting the same least-squares condition checks.
7. Numerical methods and additional results
The numerical study evaluates the weighting objective with grid-based deterministic procedures and examines finite-sample learning diagnostics for classification and quadratic targets. Results are accompanied by explicit warnings that refinement and transition scans do not certify global optimality or almost-sure LIL constants.
- Grid refinement and weighting objective: 0.205184% and 0.314589% are the refined reductions relative to uniform weighting for ranks two and three at α=0.2.Refining J from 1500 to 3000 changes the corresponding minimizers to 0.131032 and 0.125006, while the rank-one value remains 1.000000.
- Grid refinement and weighting objective: At (α,κ)=(0.2,0.1), rank-two values converge to 0.680863 and rank-three values round to 0.354299 across J=800,1500,3000.The rank-one values round to 1.000000 on the finer grids, consistent with the stated normalization check.
- Boundary behavior: The reported optimizer transitions near the admissible boundary are algorithm- and grid-dependent observations, not certified critical intervals or proofs of a boundary infimum.The transition scans use J=3000 and a fine α grid, but the imposed upper search limit can produce gaps near the search cutoff.
- Monte Carlo design: The Monte Carlo design uses fractional Gaussian noise, pure power weights, 120 independent paths, N=216, and dyadic sample sizes n=2^j.Classification uses α=0.2 and κ∈{−0.15,0,0.10,0.20}; the computations use shared paths across weights or target shifts.
- Learning diagnostics: The total slopes remain more negative than the reference exponent −α/2=−0.1, so the comparison is pre-asymptotic.The diagnostics are finite-sample second-moment summaries and do not estimate an almost-sure limsup constant.