Source-linked AI summary

Neural operators approximate strongly continuous convex monotone semigroups

Jonas Blessing, Philipp Schmocker, Alessandro Sgarabottolo

arXiv:2609.02727v1math.NAcs.LGmath.APmath.PRstat.ML

TL;DR

The paper addresses approximation of strongly continuous convex monotone semigroups by learning their Chernoff-type one-step operators. It proves universal approximation for general Chernoff-neural operators, derives quantitative rates for envelope-neural operators, and demonstrates the approach in several stochastic and nonlinear-PDE examples.

  • Problem

    The paper studies how to approximate strongly continuous convex monotone semigroups without directly approximating the full solution operator from supervised PDE solution data.

  • Method

    It learns Chernoff-type one-step operators with neural operators, propagates their approximation errors through iteration, and uses finite parameter maxima with ReLU networks for envelope operators.

  • Results

    The paper proves universal approximation of the corresponding semigroups and derives explicit approximation rates for envelope-neural operators under suitable regularity assumptions.

  • Takeaways & Limitations

    The approach applies to nonlinear PDEs, stochastic optimal control, HJB equations, and stochastic processes under model uncertainty, with envelope-neural operators offering efficient training and small architectures.

  • Takeaways & Limitations

    The results rely on assumptions about operator extensions, weight functions, and regularity or parameter dependence for the respective neural-operator classes.

Abstract

from arXiv · show

We approximate strongly continuous convex monotone semigroups by learning their Chernoff-type one-step operators with neural operators. First, we introduce the general class of so-called Chernoff-neural operators and show in a universal approximation theorem that they can approximate the Chernoff one-step operators arbitrarily well. By using stability estimates between weighted Hölder spaces, the one-step approximation error can be propagated through the iterations which yields universal approximation of the corresponding semigroup. Second, we introduce the more specialized class of envelope-neural operators for envelope semigroups which allows us to derive quantitative approximation rates. Finally, we illustrate the effectiveness of these neural operators in several numerical examples arising from non-linear partial differential equations, stochastic optimal control and stochastic processes under model uncertainty.

1. Introduction

The paper approximates strongly continuous convex monotone semigroups by learning Chernoff-type one-step operators and iterating the learned operators. It develops general and envelope-specialized architectures, with universal approximation, quantitative rates, and applications across nonlinear PDEs and stochastic models.

  • Chernoff-neural operators: Weighted Hölder-space approximation and stability estimates propagate one-step errors through iterations to approximate the semigroup.The universal approximation theorem covers the one-step operators and the resulting semigroup.
  • Envelope-neural operators: Envelope-neural operators replace the supremum over parameters with a maximum over finitely many trainable parameters and use ReLU networks to represent finite maxima.The construction preserves the envelope structure of the one-step operators.
  • Applications and positioning: The framework covers nonlinear PDEs, stochastic optimal control, HJB equations, and stochastic processes under model uncertainty, with three numerical examples.The paper contrasts its approach with direct PDE-solution and full solution-operator approximation methods.
  • Envelope-neural operators: The envelope architecture yields explicit approximation rates in terms of neural-operator size under suitable regularity assumptions.These rates are combined with one-step convergence rates to obtain quantitative semigroup error bounds.
  • Overview: The method learns Chernoff-type one-step operators with neural operators and recovers the semigroup by iteration.Training requires evaluations of the one-step operators rather than supervised PDE solution data.

2. Setup and notation

The paper formulates Chernoff approximations and strongly continuous convex monotone semigroups on weighted function spaces. It establishes assumptions, stability properties, uniqueness, and convergence mechanisms used by the neural-operator constructions.

  • Function spaces: Weighted spaces include continuous functions controlled by weights, with polynomial weights κ(x) := (1 + |x|^q)−1 among the motivating examples.The setup also introduces weighted Hölder spaces and related embeddings between different regularity and weight scales.
  • Topology: The analysis uses the mixed topology on Cκ, combining the weighted norm with uniform convergence on compact subsets.This topology is stated to be crucial for analyzing strongly continuous convex monotone semigroups.
  • Semigroup properties: The resulting semigroup satisfies the semigroup law, preserves convexity and monotonicity, and is uniquely identified under the stated generator conditions.The uniqueness statement applies when the comparison semigroup satisfies the corresponding regularity conditions.
  • Chernoff approximation: Chernoff approximations iterate one-step operators In a number of times determined by n := max{k ∈N0 : khn ≤t}, with hn →0.Under the stated conditions, the iterates converge to a semigroup uniquely determined by the infinitesimal behavior of the one-step operators.
  • Semigroup construction: The assumptions require convexity, monotonicity, normalization, stability bounds, regularity, Lipschitz preservation, and existence of the infinitesimal limit.These conditions yield a strongly continuous convex monotone semigroup with generator A and stability bound ∥S(t)f −S(t)g∥κ ≤eωt∥f −g∥κ.

3. Chernoff-neural operators

This section develops a general neural-operator architecture for approximating Chernoff one-step operators on weighted Hölder spaces, then propagates the approximation through iteration to approximate the associated semigroup.

  • Chernoff-neural operators: The architecture learns Chernoff one-step operators uniformly over functions so their iterations approximate the corresponding strongly continuous convex monotone semigroup.The construction uses weighted spaces and stability estimates across Hölder regularity and weight scales.
  • Qualitative approximation of the semigroup: Stability estimates between weighted Hölder spaces allow the one-step error to be propagated through repeated iterations over finite time horizons.The iteration argument uses intermediate Hölder exponents and weights together with operator growth estimates.
  • Universal approximation: The universal approximation proof adapts infinite-dimensional neural-network results using trigonometric feature algebras, A-submodules, and density of the linear readout space.Point separation and nowhere-vanishing properties support the density argument.
  • Universal approximation: Under structural assumptions on the operator sequence and neural features, Chernoff-neural operators achieve arbitrarily small one-step approximation error.The assumptions require suitable extensions, growth bounds, and separating feature mappings on weighted Hölder spaces.
  • Qualitative approximation of the semigroup: The resulting iterated Chernoff-neural operators universally approximate the strongly continuous convex monotone semigroup on the stated weighted Hölder-space domains.The theorem provides approximation for arbitrary finite horizon T and tolerance ε under the listed assumptions.

4. Envelope-neural operators

Envelope-neural operators approximate one-step operators defined as suprema over parameterized linear operator families by replacing the supremum with a finite ReLU-implemented maximum. Under regularity and compactness assumptions, their iterations approximate the associated semigroup, with explicit rates tied to parameter-space coverage.

  • Architecture and setting: Envelope-neural operators target one-step operators expressed as suprema over parameterized linear operators, covering Nisio semigroups and related uncertain stochastic processes.The construction uses operators I_n,λ, penalization functions, and step sizes h_n tending to zero.
  • Architecture and setting: The semigroup construction requires monotonicity, stability, approximation, regularity, and generator conditions collected in Assumption 4.1.These include linear monotone one-step operators and norm growth bounded by e^(ωh_n).
  • Finite-envelope construction: The supremum over parameters is replaced by a maximum over finitely many trainable parameters and implemented by a ReLU network.The resulting operator composes finitely many hidden-layer maps I_n,λ with a ReLU network that computes the penalized maximum.
  • Universal approximation: Iterations of these envelope-neural operators approximate the strongly continuous convex monotone semigroup on bounded time intervals, and suitable sequences of neuron counts generate the same semigroup.The result propagates one-step errors through the iterations and then chooses n and network parameters to meet any prescribed ε.
  • Quantitative approximation: The quantitative theory bounds approximation error using the fill-in distance δ_r(M), while the output network has depth O(log2(M)) and width O(M).For Euclidean-equivalent parameter metrics in R^m, δ_r(M) ≍ M^-1/m; neuron counts can then be scaled with h_n to preserve semigroup convergence.

5. Numerical experiments

The experiments apply Chernoff-neural and envelope-neural operators to semilinear PDEs, stochastic control, and Wasserstein-perturbation semigroups. They evaluate learned iterations against analytical or finite-difference reference solutions and examine step-number effects.

  • Overview: Three numerical examples test neural approximations of Chernoff one-step operators and their corresponding strongly continuous convex monotone semigroups.The examples increase in abstraction and cover nonlinear PDEs, stochastic optimal control, and stochastic processes under model uncertainty.
  • Semilinear PDE: The semilinear PDE experiment splits diffusion and nonlinear transport, then learns the Chernoff one-step operator with a neural operator.The experiment uses n = 30 and compares out-of-sample evaluations with reference solutions based on explicit representations and Gauss–Hermite quadrature.
  • Stochastic optimal control: The stochastic-control experiment approximates a value function using Chernoff and envelope neural operators trained on sampled strike-price functions.Reference solutions come from a finite-difference scheme for the corresponding fully nonlinear HJB equation, with out-of-sample functions excluded from training.

Appendix A. Weighted Hölder spaces

Appendix A develops foundational properties of weighted Hölder spaces, including completeness, continuous embeddings under weaker regularity or weights, and approximation by compactly supported smooth functions.

  • Embeddings: Weighted Hölder spaces are treated through comparisons between regularity exponents and weight functions.The appendix establishes embeddings when α′ ≤ α and κ′ is controlled by κ, with stronger compactness-related conditions when α′ < α.
  • Completeness: The weighted Hölder space Cα_κ is complete under its weighted Hölder norm.The proof combines convergence in the weighted continuous-function space with control of Hölder seminorms.
  • Compactness: Theorem A.3 provides a compactness statement for bounded sequences in weighted Hölder spaces under lower regularity and weaker-weight conditions.The proof uses Arzelà–Ascoli and local uniform convergence before controlling the weighted tail.
  • Smooth approximation: For α < 1, weighted Hölder functions can be approximated by compactly supported smooth functions.A cutoff argument first removes the tail, and mollification then gives smooth approximation on compact support.

Appendix B. Weighted Stone-Weierstrass theorems for Banach scales

Appendix B establishes scalar and vector-valued weighted Stone–Weierstrass results on Banach scales. The proofs combine compactness, local approximation, partitions of unity, and algebraic closure.

  • Banach-scale framework: The Banach-scale setting orders spaces by continuous inclusions and requires bounded balls at stronger levels to be compact at weaker levels.This structure supports the weighted approximation arguments used later for neural-operator universality.
  • Application to weighted spaces: The appendix specializes the ordered index set to pairs of Hölder exponents and weights indexed by a totally ordered set.The relations compare both regularity and weight strength across the Banach scale.
  • Scalar approximation: The scalar theorem assumes a point-separating, nowhere-vanishing subalgebra and concludes density in the target weighted mapping space.The proof applies classical Stone–Weierstrass on compact subsets, then uses truncation and polynomial composition.
  • Vector-valued approximation: The vector-valued theorem extends density to an A-submodule whose pointwise values are dense in the target Banach spaces.A partition of unity combines locally selected approximants while preserving the required weighted mapping properties.

Appendix C. Proof of Lemmas 5.1–5.4

Appendix C proves a polynomial moment bound for a Brownian-driven process by applying Itô’s formula to a Lyapunov function and then Gronwall’s lemma.

  • Moment estimate: For X_t = μt + σW_t, the appendix bounds the q-th polynomial growth of x + X_t by e^(ω_q t)(1 + |x|^q).The bound holds for all t ≥ 0 and x ∈ R^d.
  • Proof strategy: The proof uses V_q(x) = 1 + |x|^q and computes its gradient and Hessian to control the generator.Itô’s formula yields an integral inequality for E[V_q(x + X_t)], which Gronwall’s lemma closes.

C.1. Proof of Lemma 5.1.

The proof verifies the assumptions needed for the parameterized operators by combining an external theorem, Lemma C.1, and properties of the operators I_n,λ.

  • Assumptions 2.2 and 4.1, together with the extension I_n: C_κ0 → C_κ0, are referred to [12, Theorem 3.6].
  • Assumption 3.1(i) is stated to be clearly satisfied.
  • Lemma C.1 supplies a constant ω ≥ 0 for the weighted Hölder estimates over n, λ, α′, κ′, and f.
  • The operators use Gaussian increments Z_n,λ := λh_n + W_hn and act as (I_n,λf)(x) := E[f(x + Z_n,λ)].
  • Linearity of I_n,λ and taking the supremum over λ establish Assumptions 3.1(ii)–(iii) and 4.5(i).
  • Assumption 4.8 holds with L_r := (c + r) sup_n∈N h_n, d_r := | · |, and β_r := 1.

C.2. Proof of Lemma 5.2.

This proof verifies the assumptions for operators indexed by drift and volatility parameters, using Gaussian increments, linearity, and supremum estimates.

  • Assumptions 2.2 and 4.1, together with the extension I_n: C_κ0 → C_κ0, are referred to [8, Theorem 6.2].
  • Assumption 3.1(i) is stated to be clearly satisfied, while Lemma C.1 provides a constant ω ≥ 0 for the required estimates.
  • The increments are Z_n,(μ,σ) := μh_n + σW_hn, and the operators satisfy (I_n,(μ,σ)f)(x) := E[f(x + Z_n,(μ,σ))].
  • Linearity of I_n,(μ,σ) and taking the supremum over Ξ × Σ establish Assumptions 3.1(ii)–(iii) and 4.5(i).
  • The proof invokes Lemma C.1, Hölder’s inequality, and Jensen’s inequality in deriving the estimate involving h_n^2 + h_n.
  • Assumption 4.8 holds with L_r := 2^(q−2)C_r, d_r := d, and β_r := α.

C.3. Proof of Lemma 5.3.

The proof establishes the assumptions for an envelope indexed by probability measures and nonnegative parameters, using bounded Lipschitz estimates and compact parameter sets.

  • Assumptions 4.1(i)–(vii) are referred to, and monotonicity of η is used in the subsequent estimate.
  • The parameter set is Λ := P_p(R^d) × R_+, and the proof works with f ∈ C_b and later with f ∈ Lip_b.
  • Using α = 1 and κ ≡ 1 establishes Assumption 4.5(i).
  • For every r ≥ 0, the proof introduces R_r and the restricted set Λ_r := {ν ∈ P_p(R^d): W_p(ν, δ_0) ≤ C_r} × [0, R_r].
  • Equation (5.11) gives C_r < ∞, and p > 1 implies that Λ_r is compact in the relevant metric.
  • The Kantorovich–Rubinstein inequality, Lipschitz continuity of η, and bounded sup_n h_n yield Assumption 4.8 with L_r := max{r, h̄c_r}, d_r := d, and β_r := 1.

C.4. Proof of Lemma 5.4.

The proof verifies Assumption 4.8 for a parameter set in R^d by combining prior assumption checks, total boundedness, and Lipschitz continuity of η(| · |).

  • Assumption 4.1 is referred to, while Assumption 4.5(i) is verified as in the proof of Lemma 5.3.
  • The identity η(0) = 0 is used in an estimate valid for n, r, f ∈ Lip_b(r), and λ, x ∈ R^d.
  • For p > 1, the proof introduces a totally bounded set Λ_r ⊂ R^d satisfying the stated estimates for n, r, and f.
  • Lipschitz continuity of η(| · |) on Λ_r, with constant c_r, and h̄ := sup_n∈N h_n < ∞ support the parameter estimate.
  • Assumption 4.8 holds with L_r := h̄(r + c_r), d_r := | · |, and β_r := 1.
Loading 2609.02727v1…