Source-linked AI summary

Variation Brownian Kernel Ladders

Mahdi Mohammadigohari

arXiv:2608.13882v1cs.LG

TL;DR

Claims about depth depend on how representation complexity is defined, motivating a framework that separates recursive feature construction from linear superposition. The paper introduces VBKL and establishes depth-dependent function-space geometry, architecture-dependent statistical guarantees, and separated-resource approximation bounds.

  • Problem

    The paper asks how placing linear superposition within or after recursive feature construction changes the resulting function-space theory.

  • Method

    VBKL recursively composes normalized Brownian RKHS profiles into path dictionaries, then forms their signed-measure variation hulls.

  • Results

    The framework establishes compactness, attained variation representations, Hölder regularity, strict depth growth, finite-architecture statistical bounds, and M^-1/2 + m^-1/2 approximation error.

  • Takeaways & Limitations

    VBKL provides a function-space account of depth that separates recursive geometry, variation complexity, statistical control, and approximation resources.

  • Takeaways & Limitations

    The statistical guarantees apply to explicitly parameterized finite lower-support architectures, not unrestricted infinite-dimensional variation balls.

Abstract

from arXiv · show

Claims about the benefit of depth depend on the complexity assigned to a representation. We introduce the \emph{Variation Brownian Kernel Ladder} (VBKL), a path-atomic function-space framework that separates nonlinear recursive dictionary construction from linear variation superposition. Starting from linear projections, each atom recursively composes unit-ball profiles from the Brownian reproducing kernel Hilbert space; the full VBKL space is then the signed-measure variation hull of the completed dictionary. We identify each recursive dictionary as a union of Brownian pullback RKHS balls and establish variation-controlled Hölder regularity, compactness and attainment, and strict growth with depth under a local non-degeneracy condition whose trace lies in the support of the input measure. For associated finite lower-support architectures, we derive Rademacher and generalization bounds through Brownian quadratic chaos, signed threshold traces, and VC entropy. We also construct two-stage approximants by discretizing the outer measure and the selected outer Brownian profiles, obtaining an $M^{-1/2}+m^{-1/2}$ error bound, a sharp interpolation constant $\sqrt{A/2}$, and at most $2M$ active outer-profile basis contributions per evaluation. Controlled experiments illustrate the approximation mechanisms and indicate a favorable limited-data accuracy--complexity trade-off.

1 Introduction

The VBKL separates recursive nonlinear path-dictionary construction from outer signed-measure superposition, enabling depth-sensitive function-space analysis. The paper establishes geometric, statistical, approximation, and experimental results for this framework while distinguishing full-space theory from finite architectures.

  • 1 Introduction: VBKL recursively composes lower-level atoms with unit-ball Brownian RKHS profiles, then applies the signed-measure variation hull only after reaching the target depth.Each depth-L atom contains one support path and exactly L −1 normalized Brownian profiles; intermediate levels are not convexified.
  • 1 Introduction: The full-space theory proves compactness, attained minimum variation representations, variation-controlled 2−(L−1)-Hölder regularity, and strict depth growth under local non-degeneracy supported by the input measure.Under full input support, representatives are unique and define a continuous Hölder embedding.
  • 1 Introduction: Finite lower-support architectures receive Rademacher and high-probability generalization bounds via Brownian quadratic chaos, signed threshold traces, and VC entropy.These statistical guarantees concern explicitly parameterized finite Brownian variation classes, distinct from the full infinite-dimensional path-atomic space.
  • 1 Introduction: Approximation separates resources: recursive-atom discretization contributes M−1/2 error, outer-profile interpolation contributes m−1/2 error, and the interpolation constant √(A/2) is optimal.The resulting evaluation uses at most 2M active outer-profile basis contributions, independently of m.
  • 1 Introduction: Controlled experiments isolate discretization mechanisms, exactly attain the worst-case normalized-tent interpolation bound, and indicate strongest VBKL behavior in limited-data regimes with a favorable accuracy–parameter trade-off.The experiments do not claim universal predictive dominance.

2 Notation

The section defines the Brownian-kernel notation and constructs VBKL by recursively composing normalized Brownian profiles after linear projections, followed by an outer signed-measure variation hull. It also introduces a separate finite lower-support architecture for statistical analysis, which is not generally identical to the full infinite-dimensional VBKL ball.

  • Recursive Brownian dictionary: VBKL depth-L atoms begin with a linear projection and compose exactly L −1 normalized Brownian profiles, forming the recursive dictionary U_L.Each recursive step propagates one support rather than averaging over a support family; signed-measure superposition is added only after the target depth is reached.
  • Variation-space construction: The depth-L VBKL space is the signed-measure variation hull of U_L, with intrinsic variation complexity given by the infimal total variation of representing measures.Finite-support measures provide discrete atomic expansions and approximants, while compactness and weak-* convergence preserve barycentric representations under the stated conditions.
  • Relation to BKL: Each U_l is a union of Brownian pullback RKHS unit balls, retaining recursive Brownian geometry while replacing layerwise kernel averaging with an outer variation hull over compositional paths.Scaling a support preserves its RKHS as a set but not its unit ball, so intermediate supports are not implicitly normalized.
  • Finite lower-support architecture: The finite lower-support architecture replaces infinite-dimensional Brownian profiles with grid-based piecewise-linear profiles and constrained mixing and readout coefficients.The row-sum and readout constraints ensure uniformly bounded intermediate supports, and the parameter count is PL−1,m,G = md + (L −2)m(G + 1) + (L −3)m^2 + m.
  • Relation to the infinite-dimensional dictionary: The finite architecture is an associated class rather than the full infinite-dimensional ball V(L), although its path-only specialization yields recursively composed Brownian atoms.The distinction arises because general mixing and linear readout need not select a single compositional path.

3 Analytical and Statistical Properties of the VBKL spaces

The section establishes analytical regularity, pointwise control, and depth growth for full VBKL spaces, then derives architecture-dependent statistical guarantees for finite Brownian variation classes.

  • Analytical properties: Variation complexity yields Hölder representatives, pointwise control, and continuous embedding into a Hölder space; full support of ν ensures representative uniqueness.These results apply to the full infinite-dimensional VBKL spaces.
  • Analytical properties: Depth inclusions are monotone and become strict under local non-degeneracy along a Lipschitz trace whose support lies in supp(ν).The condition uses a first-layer generator with linear growth along the trace.
  • Statistical complexity: The finite-architecture Rademacher bound separates outer variation radius, recursive Brownian geometry, and lower-support parameter complexity.Its proof combines Brownian quadratic chaos with signed threshold entropy for the finite lower-support architecture.
  • Statistical complexity: The statistical bound applies to finite-architecture classes rather than the full infinite-dimensional variation ball, while covering the restricted path-only VBKL subclass.The distinction reflects intermediate linear mixing and final linear readout in the general finite architecture.
  • Statistical complexity: Under the theorem’s architecture and Lipschitz-loss conditions, the Rademacher result yields a high-probability generalization guarantee for every finite-architecture predictor.The guarantee holds with probability at least 1 − δ over the sample draw.

4 Constructive Approximation of VBKL spaces

This section constructs finite approximants to the full infinite-dimensional VBKL space by separately discretizing the outer signed measure and the selected atoms’ outer Brownian profiles. The resulting error has M^-1/2 and m^-1/2 contributions, while lower-level supports remain unchanged, and the interpolation rate m^-1/2 with constant sqrt(A/2) is sharp.

  • Finite-atomic approximation: Finite-atomic approximation replaces each full-space element by a finite-support VBKL element under compact-domain and probability-measure assumptions.The zero-variation case is represented by the zero function, while positive-variation elements admit finitely many selected atoms and signs.
  • Computational structure: Evaluating the two-stage approximant requires at most 2M active outer-profile basis contributions, independently of interpolation resolution m.The zero realization requires no active outer-profile contribution.
  • Two-stage approximation: The two-stage approximant combines finite outer atomic support with piecewise-linear outer-profile interpolation while preserving selected lower-level supports exactly.The outer measure and outermost Brownian profiles are discretized separately; lower-level supports remain elements of U_L−1.
  • Two-stage approximation: The approximation error separates into an M^-1/2 term from outer-measure discretization and an m^-1/2 term from outer-profile interpolation.Theorem 4’s construction does not discretize the lower-level supports.
  • Outer-profile interpolation: Piecewise-linear interpolation of Brownian RKHS profiles achieves the m^-1/2 rate with sharp constant sqrt(A/2), so neither quantity can be improved.The interpolation estimate is established on a symmetric interval [−A,A] using a uniform grid.

5 Experiments

Experiments test VBKL’s approximation behavior, supervised accuracy and representation efficiency, and numerical optimization. Results show sharp worst-case interpolation, competitive limited-data performance with fewer parameters, and stable directional estimation under tested conditions.

  • Approximation: The two-stage approximants follow the predicted M^-1/2 and m^-1/2 behaviors, while fixed Brownian profiles converge faster and normalized tent profiles attain the theoretical interpolation bound.Figure 2 separates signed-measure and profile discretization, and Figure 1 demonstrates sharpness of the worst-case interpolation estimate.
  • Representation efficiency: At n = 100, VBKL wins four of five matched seed-wise comparisons while using approximately 4.6× fewer parameters than DNVS; the ratios reach 13.7× and 18.3× at n = 250 and n = 500.VBKL uses fewer parameters at every Energy Efficiency training size, even when DNVS wins more accuracy comparisons.
  • Supervised learning: At n = 100, VBKL has lower mean error than DNVS on Energy Efficiency, but at n = 250 and n = 500, DNVS, KRR, and RBF perform better.The evidence supports a limited-data advantage rather than universal predictive dominance.
  • Supervised learning: On the recursive-teacher benchmark, validation-selected VBKL performs particularly well with limited data, while DNVS attains the lowest mean error at the largest sample sizes.Performance differences narrow as sample size grows, consistent with validation-adapted representation complexity.
  • Optimization validation: Directional-estimator relative error improves from interaction scales 10^-1 to 10^-3, rises modestly at 10^-4, and variance decreases when averaging more sampled directions.These tests validate numerical stability under tested conditions but do not establish a global optimization-convergence theorem.

6 Conclusion

The conclusion presents VBKL as a recursive Brownian dictionary framework followed by outer variation superposition, with regularity, statistical, and approximation guarantees. Experiments suggest an accuracy–complexity trade-off in limited-data regimes, while future work targets broader statistical, discretization, and optimization theory.

  • Core framework: VBKL combines recursive Brownian dictionary construction with outer variation superposition, yielding variation-controlled Hölder representatives and continuous Hölder embeddings under full support.The recursive dictionaries are unions of Brownian pullback RKHS unit balls, with embedding results requiring the stated local non-degeneracy and trace-support conditions.
  • Approximation and statistical guarantees: The approximation error splits into M^-1/2 for finite outer-measure support and m^-1/2 for interpolation of selected outer Brownian profiles, while lower-level supports remain in U_L−1.For finite lower-support architectures, the statistical bounds separate the outer variation radius from the architecture parameter count.
  • Empirical scope and future work: Experiments indicate a favorable accuracy–complexity trade-off in limited-data regimes without claiming universal predictive dominance or global optimization convergence.Future work includes intrinsic full-space statistical bounds, recursive discretization of lower-level supports, and optimization theory for explicit finite architectures.

7 Proofs

This section proves the paper’s main analytical, statistical, and approximation results, including regularity and attainment, strict depth growth, finite-architecture complexity bounds, and sparse two-stage evaluation.

  • Analytical results: The proofs establish variation-controlled Hölder representatives, compactness, attainment, and uniqueness under full support, with the induced map into C_0,αL(X) linear, injective, and continuous.Arzelà–Ascoli supplies uniform convergence of near-minimizing representatives, while full support identifies continuous representatives of the same L2(ν) element.
  • Analytical results: V(L) is strictly contained in V(L+1) under the local non-degeneracy and trace-support conditions.The contradiction uses a fractional-power Brownian RKHS profile whose exponent a satisfies aL < αL, violating the Hölder behavior required at depth L.
  • Statistical results: The finite lower-support architecture proofs derive the stated Rademacher and generalization bounds by combining variation-hull estimates, Brownian kernel identities, envelope control, and VC entropy.The argument proceeds through five steps and concludes with the expected Rademacher complexity equation (39).
  • Approximation results: At most 2M active outer-profile basis contributions are needed to evaluate F_M,m(x), independently of interpolation resolution m.Each of the M summands contributes at most two active interpolation basis functions.

Appendix A. Experimental Protocols and Extended Results · A.1 Constructive-approximation protocol

The appendix supplies protocols, numerical details, controlled constructions, validation-selected configurations, exact comparisons, and computational measurements supporting Section 5. Its constructive studies use a known deterministic teacher representation, enabling direct approximation-error evaluation without training or statistical-estimation error.

  • Appendix A. Experimental Protocols and Extended Results: The appendix records protocols and numerical details supporting the principal empirical findings presented in Section 5.The main body retains the principal findings and all figures needed to assess them.
  • Appendix A. Experimental Protocols and Extended Results: The appendix documents controlled constructions used to support the reported experiments.
  • Appendix A. Experimental Protocols and Extended Results: It reports validation-selected configurations for the experimental evaluations.
  • Appendix A. Experimental Protocols and Extended Results: It provides exact numerical comparisons for the evaluated constructions.
  • Appendix A. Experimental Protocols and Extended Results: The appendix also records computational measurements associated with the experiments.
  • A.1 Constructive-approximation protocol: The constructive studies use a deterministic teacher F ∈V(L) represented by an explicit finite signed measure on recursive Brownian atoms.Each selected atom has exact recursive support and a deterministic unit-energy Brownian profile.
  • A.1 Constructive-approximation protocol: Because the reference representation is known, approximation errors are computed directly without training or statistical-estimation error.Figure 2 isolates the components of the construction across three panels.

A.2 Statistical-learning protocol and extended results · A.2.1 Recursive-teacher configurations

The supervised studies use seed-specific, validation-selected VBKL configurations retrained and tested under matched split protocols. The recursive-teacher benchmark uses a controlled hierarchical Brownian target and records adaptive configuration choices across sample sizes and seeds.

  • A.2 Statistical-learning protocol and extended results: VBKL architecture selection is performed separately for every training-set size and random seed using the validation set.
  • A.2 Statistical-learning protocol and extended results: Selected configurations are retrained and evaluated on held-out test sets, with results aggregated over five independent seeds.
  • A.2 Statistical-learning protocol and extended results: DNVS, KRR, and RBF follow the same split protocol as VBKL within each benchmark.
  • A.2 Statistical-learning protocol and extended results: The supervised studies include the recursive-teacher benchmark and the Energy Efficiency benchmark.
  • A.2.1 Recursive-teacher configurations: The recursive-teacher target uses the proposed model’s hierarchical Brownian mechanism, creating a controlled matched setting.
  • A.2.1 Recursive-teacher configurations: Validation jointly selects the number of recursive Brownian paths and the optimization horizon.
  • A.2.1 Recursive-teacher configurations: Table 3 reports selected recursive-teacher configurations in seed order.
  • A.2.1 Recursive-teacher configurations: Configuration variation across sample sizes and seeds intentionally records the adaptive model-selection procedure underlying Figure 3a’s learning curve.

A.2.2 Energy Efficiency numerical results · A.3 Optimization protocol

The Energy Efficiency results compare VBKL with other methods under a common benchmark protocol, while validation selects VBKL configurations separately by training size and seed. The optimization-protocol experiments examine finite-difference scale effects and Monte Carlo estimator stability against automatic differentiation.

  • A.2.2 Energy Efficiency numerical results: At n = 100, VBKL’s mean Energy Efficiency error is close to the best reported value and lower than the DNVS mean.Table 4 reports means and standard deviations over five random seeds; the passage also notes results at two larger training sizes without providing their values.
  • A.2.2 Energy Efficiency numerical results: The Energy Efficiency benchmark uses the same train/validation/test protocol for every compared method.The benchmark is publicly available, and the comparison is conducted under a common evaluation procedure.
  • A.2.2 Energy Efficiency numerical results: VBKL validation independently selects path count, Brownian profile resolution, and optimization horizon for each training size and random seed.Table 3 lists validation-selected path counts and training epochs in seed order.
  • A.2.2 Energy Efficiency numerical results: Table 4 reports Energy Efficiency test MSE as means with standard deviations computed over five random seeds.These statistics provide the exact numerical values underlying Figure 3b, although the supplied passage does not include the full table entries.
  • A.3 Optimization protocol: The consistency experiment uses automatic differentiation as the reference for directional derivatives and evaluates finite-difference averaged directional estimates.Interaction scales are h ∈ {10^-1, 10^-2, 10^-3, 10^-4}.
  • A.3 Optimization protocol: The Monte Carlo stability study varies sampled directions from 20 to 27 and records estimator standard deviation, revealing finite-difference scale trade-offs and stability behavior.Figure 4 presents both the interaction-scale experiment and the direction-sampling study.

A.4 Computational characteristics · Appendix B. Additional notation

The appendix reports computational characteristics showing finite VBKL models are tractable but not uniformly fastest, while defining notation for norms, function spaces, integrals, and indicators used in the proofs.

  • A.4 Computational characteristics: Finite VBKL models remain computationally tractable at benchmark scale, although they are not uniformly the fastest method.The benchmark records training time, complete-test-set inference time, and fitted or trainable parameter counts over training sizes and seeds.
  • A.4 Computational characteristics: DNVS has lower inference time throughout the comparison, while feature and kernel baselines have lower training time for tested sample sizes.
  • A.4 Computational characteristics: VBKL’s principal computational distinction is an explicit recursive representation that is substantially more compact than DNVS in the limited-data Energy Efficiency comparison.
  • A.4 Computational characteristics: Table 5 summarizes computational quantities across all training sizes and seeds, with means and standard deviations over five seeds and median parameter counts.Tables 6–8 provide training-size-specific results for training time, inference time, and trainable or fitted parameter counts.
  • A.4 Computational characteristics: The computational tables identify best and second-best results by boldface and underlining, respectively.Reported values use means with standard deviations over five random seeds.
  • Appendix B. Additional notation: Appendix B defines norm balls and spheres, including Bd and Sd−1, together with Euclidean inner-product notation in Rd.
  • Appendix B. Additional notation: The notation appendix specifies C(X), Hölder spaces C0,α(Ω), bounded-function spaces L∞(A), and square-integrable spaces L2(ν) with their norms or inner products.Vector-valued integrals use the Bochner sense, and 1E denotes the indicator of a statement E.

Appendix C. Sharpness of the Brownian profile interpolation estimate

Appendix C proves that the Brownian profile interpolation estimate is sharp: neither its m^-1/2 convergence rate nor its constant sqrt(A/2) can be improved. The result applies to continuous piecewise-linear interpolation on a uniform grid over [-A,A].

  • Sharpness result: The interpolation estimate is optimal in both its m^-1/2 rate and its sharp constant sqrt(A/2).Thus, neither the exponent 1/2 nor the constant sqrt(A/2) admits improvement.
  • Sharpness construction: For every m ≥ 1, a Brownian RKHS profile gm attains the reverse interpolation bound on the uniform grid.The proof constructs gm through its weak derivative and evaluates the interpolation error at interval midpoints.
  • Proof mechanism: The extremizing profile has equal endpoint values on each interpolation interval, making its piecewise-linear interpolant constant there and exposing the midpoint error.Absolute continuity and the weak-derivative construction establish the endpoint equality used in the argument.

Appendix D. Internal Lemmas

Appendix D establishes internal lemmas ensuring the measure-valued VBKL space is well defined, its variation complexity is non-degenerate, and finite-support architectures and recursive pullback RKHSs admit precise structural characterizations.

  • Lemma B1: Finite signed-measure Bochner representations define elements of L2(ν), and the associated variation complexity is finite and well defined.Strong measurability and a uniform atomic bound ensure Bochner integrability with respect to the total-variation measure.
  • Lemma B2: Variation complexity is non-degenerate: zero variation implies F = 0 in L2(ν).The lemma first bounds the L2(ν) norm by the uniform atomic bound times variation complexity, then uses positive definiteness of the L2 norm.
  • Lemma B3: Finite lower-support profiles are exactly continuous functions affine on each grid interval and constant outside [−AX, AX], with an exact Brownian RKHS norm formula.The same lemma gives the finite architecture parameter count PL−1,m,G = md + (L −2)m(G + 1) + (L −3)m2 + m.
  • Lemma B4: For each recursive atom, the Brownian pullback operator has a closed nullspace in Hk(B), yielding the recursive RKHS structure of the Variation BKL dictionary.The operator is defined by composing a Brownian RKHS profile with u, and its nullspace consists of profiles vanishing on the trace u(X).

This representative satisfies

The depth-l variation space is exactly the variation hull generated by recursively constructed Brownian pullback RKHS unit balls. For finite symmetric dictionaries, empirical Rademacher complexity reduces exactly to the underlying dictionary’s complexity, with an envelope bound.

  • Recursive dictionary: V(l) is the variation hull generated by the recursively constructed union of Brownian pullback RKHS unit balls.This identifies the depth-l space through its recursive atomic dictionary and outer signed-measure superposition.
  • Recursive dictionary: Each Brownian pullback RKHS unit ball is exactly represented by minimum-norm representatives in the ambient RKHS, preserving the recursive dictionary construction.The pullback operator is linear and continuous, and the minimum-norm representative attains the pullback norm bound.
  • Finite variation hulls: For any symmetric dictionary, the empirical Rademacher complexity of its finite signed-measure variation hull equals the corresponding variation-scaled dictionary complexity.The result is obtained through a fixed-Rademacher supremum identity followed by expectation over the Rademacher variables.
  • Finite variation hulls: The variation-hull Rademacher complexity is bounded by the variation radius times the dictionary envelope, and the finite outer Brownian dictionary satisfies the required symmetry and envelope conditions.The finite outer Brownian dictionary is symmetric because each constituent RKHS unit ball is symmetric, enabling the specialization of the general reduction.

Lemma B6 (Rademacher bound for unions of Brownian pullback RKHS balls)

Lemma B6 bounds the empirical Rademacher complexity of a union of Brownian pullback RKHS balls by reducing each fixed-support supremum to an exact RKHS norm and then controlling the resulting Brownian quadratic chaos. The proof further decomposes diagonal and off-diagonal kernel terms and specializes the bound to finite lower-support architectures.

  • Symmetry: The union class D_r(A) is symmetric, allowing its empirical Rademacher complexity to be written using absolute values without changing the quantity.Symmetry follows from the absolute homogeneity of each Hilbert-space norm and the stated kernel normalization.
  • Fixed-support reduction: For each fixed support a, the supremum over the Brownian pullback RKHS ball reduces exactly to r times the RKHS norm of the Rademacher kernel-section aggregate.The supremum is attained in the direction of the aggregate whenever it is nonzero, and both sides vanish otherwise.
  • Quadratic-chaos reduction: Taking the union over supports converts the complexity estimate into a Brownian quadratic-chaos expression, followed by a Jensen bound on its square-root expectation.The reduction applies the fixed-support identity to every a in A and uses nonnegativity and concavity of the square root.
  • Chaos decomposition: The Brownian quadratic chaos is bounded by separating diagonal and off-diagonal kernel contributions, using the kernel diagonal normalization for the former and absolute-value control for the latter.This decomposition yields the stated Rademacher estimate after taking expectations and simplifying the n-dependent factors.
  • Finite lower-support architecture: For finite lower-support architectures, the general union bound specializes through the architecture’s defining class and its support-function kernel bounds, yielding the stated finite-architecture estimate.The specialization uses the definition of the finite class with r = 1, the Brownian diagonal identity, and the resulting maximum over sample points and supports.
Loading 2608.13882v1…