Source-linked AI summary

Sequential operator learning under dependent data

Rafael Oliveira

arXiv:2608.24426v1stat.MLcs.LG

TL;DR

Sequential operator learning lacks broad guarantees when observations and sensing choices depend on preceding data, especially in infinite-dimensional settings. The paper derives time-uniform self-normalized concentration for Hilbert-valued noise and uses it to establish linear and nonlinear regression-error guarantees without independence or mixing assumptions. Its linear analysis also accommodates non-Hilbert-Schmidt targets and trace-class induced noise proxies, while the nonlinear analysis remains more preliminary for pointwise confidence.

  • Problem

    Adaptive operator learning involves dependent observations, partial outputs, and data-dependent inputs or sensing operators, but available theoretical guarantees remain limited for such settings.

  • Method

    The paper derives time-uniform self-normalized concentration for Hilbert-valued stochastic processes and applies it to linear and nonlinear parametric operator regression.

  • Results

    The framework provides regression-error guarantees under dependent data without independence or mixing assumptions, including linear targets beyond the Hilbert estimation space and nonlinear models trained with losses and regularizers.

  • Takeaways & Limitations

    Operator-valued covariance proxies can yield confidence bounds when induced proxies are trace class, including conditional mean embedding regression with infinite-dimensional outputs.

  • Takeaways & Limitations

    The nonlinear result is more preliminary because pointwise confidence bounds additionally require controlling a variance-like evaluation term.

Abstract

from arXiv · show

Learning operators from sequentially collected data arises in adaptive experimental design, Bayesian optimization, and dynamical-system modelling, where observations may be dependent, and future inputs or sensing operators may depend on preceding data. We derive time-uniform self-normalized concentration bounds for stochastic processes in Hilbert spaces with vector-valued noise. We use these bounds to obtain regression-error guarantees for linear operators, including targets outside the Hilbert estimation space, and for nonlinear parametric operators trained with strongly convex losses and regularizers. Our results allow possibly infinite-dimensional inputs and outputs without independence or mixing assumptions, providing a major step towards convergence guarantees for adaptive operator learning and learning from stochastic dynamical data.

1 Introduction

Operator learning theory has largely focused on approximation or independently sampled data, while adaptive operator learning involves dependent observations and data collection. This work develops guarantees for sequentially dependent trajectories in function spaces.

  • Operator learning studies mappings between function spaces for scientific machine learning and dynamical-system modelling.
  • Adaptive operator-learning settings include optimal experimental design, active learning, and Bayesian optimization, often with partial observations and structured noise.
  • The paper derives time-uniform concentration and regression-error guarantees for linear and nonlinear operators trained on sequentially dependent data.The linear result allows targets beyond the Hilbert estimation space, while the nonlinear result covers parametric models with general losses and regularizers.
  • Existing nonlinear statistical analyses typically assume independently sampled training data, limiting guarantees for sequentially collected observations.
  • The framework provides infinite-dimensional regression-error bounds under general predictable data collection without independence or mixing assumptions.It covers both dependent trajectories and observations selected adaptively from preceding data.

2 Background and problem formulation

The problem formulation models an unknown operator learned from sequential observations in Hilbert spaces. Inputs and observation maps may depend on prior data, while noise is conditionally sub-Gaussian.

  • The unknown target operator maps an input Hilbert space U to an output Hilbert space V and may be nonlinear.
  • Observation maps can represent partial observation or discretization from V into an iteration-dependent Hilbert space Y_t.
  • Inputs and observation operators are predictable with respect to a filtration containing observed data and revealed algorithmic randomness.
  • The noise is conditionally sub-Gaussian given the preceding filtration, while inputs and sensing operators may depend arbitrarily on preceding data.
  • The formulation includes stochastic dynamics and deliberate adaptivity, without requiring the observed process to be Markovian or fully observed.

3 Main result

The main technical result extends self-normalized concentration to Hilbert-valued stochastic processes. Its proof supports time-uniform analysis under predictable, operator-valued noise covariance assumptions.

  • The main technical result extends self-normalized concentration to general Hilbert-valued increments.
  • Theorem 3.1 assumes an adapted Hilbert-valued sequence with conditionally Σ_t-sub-Gaussian noise and predictable positive-semidefinite trace-class covariance proxies.
  • The theorem’s proof uses finite-rank approximations and continuity of quadratic forms and Fredholm determinants to extend finite-dimensional self-normalization.
  • The resulting concentration framework is applied to regression analysis in subsequent sections.

4 Applications to operator learning

The paper applies time-uniform concentration bounds to sequentially dependent operator regression, covering linear targets outside the estimation space and nonlinear parametric models. It provides high-probability error guarantees under adaptive observations, with convergence rates in persistently excited finite-dimensional settings.

  • Applications to operator learning: The framework applies to linear and nonlinear operator regression under sequentially dependent observations, including targets outside the Hilbert estimation space.The nonlinear analysis assumes strongly convex losses and regularizers for parametric operator models.
  • Linear regression: Linear observations use predictable bounded maps and regularized least squares, with conditional operator-valued sub-Gaussian noise entering the regression guarantee.The estimator minimizes accumulated observation errors plus a regularization term λ∥F∥².
  • Linear regression: The linear theorem gives probability at least 1 −δ guarantees for regression error under the stated trace-class proxy and predictability assumptions.The result is formulated through the covariance-normalized operator Ct and holds on a high-probability event.
  • Linear regression: Under finite-dimensional persistent excitation, the linear bound yields an explicit convergence rate for observations confined to a fixed d-dimensional subspace.The setting assumes bounded observation operators and excitation of the relevant subspace; for fixed d and δ, the error is O(...).
  • Linear regression: More general rates require controlling variance-like norms and Fredholm log-determinant terms, potentially through operator-valued information-gain analogues.The paper identifies this control as necessary when finite-dimensional persistent excitation is unavailable.
  • Nonlinear regression: The nonlinear parametric theorem assumes twice Fréchet-differentiable, α-strongly convex losses, λt-strongly convex regularizers, realizability, and a globally optimized estimator.Its proof combines global optimality with self-normalized control of cumulative gradient noise; the factor of two reflects optimization over the parametric class rather than all of H.

5 Discussion

The discussion emphasizes a general concentration framework for sequential operator learning with dependent, adaptive, and potentially infinite-dimensional observations. Operator-valued covariance proxies broaden the linear-regression setting, while nonlinear pointwise confidence remains more preliminary.

  • Operator-valued sub-Gaussian covariance proxies allow spectral structure to enter confidence bounds for Hilbert-valued noise.
  • Trace-class induced proxies suffice in linear regression, rather than requiring Hilbert-Schmidt observation operators.
  • This includes conditional mean embedding regression with K(x, x′) = k(x, x′)I, whose Gram operator is not trace class for infinite-dimensional outputs.
  • Self-normalized bounds support confidence sets, regret bounds, and convergence guarantees throughout sequential learning.
  • The nonlinear result remains more preliminary because pointwise confidence bounds require control of a variance-like evaluation term.

A.2 Existing results

This appendix records operator identities and finite-dimensional concentration ingredients used by the proofs. It then supports extension from matrix subspaces to infinite-dimensional Hilbert spaces.

  • The proofs use operator versions of Woodbury’s identity and the push-through identity under bounded-invertibility assumptions.
  • The push-through identity rewrites (A + BB∗)−1B as A−1B(I + B∗A−1B)−1.

B Auxiliary results

The auxiliary section supplies the Hilbert-space concentration and convex-loss tools underlying the paper’s main regression guarantees. Its proofs lift finite-dimensional arguments through projections, operator limits, and determinant continuity.

  • Theorem B.1 extends self-normalized concentration to Hilbert-valued noise and is identified as potentially independently useful.
  • The proof reduces the result to finite-rank projections and then lifts it to the full Hilbert space using strong-operator limits and Fatou’s lemma.
  • Fredholm determinant continuity under trace-norm convergence supports the infinite-dimensional limit arguments.
  • Corollary B.2 extends Theorem B.1 to any boundedly invertible positive-definite regularizer.
  • Lemma B.3 provides loss-difference bounds for operator-valued observations when losses and regularizers are strongly convex.

C.1 Proof of Theorem 3.1

The proof establishes a time-uniform concentration statement by controlling the first time a bound is crossed. A stopping-time argument yields simultaneous validity over all times with probability at least 1 −δ.

  • The proof uses the standard stopping-time argument for self-normalized concentration.
  • The first crossing time is stopped at τt := min{τ, t} to analyze violations up to each fixed time.
  • The probability of a crossing by time t is bounded by δ for every t.
  • The resulting inequality holds simultaneously for every t ∈N with probability at least 1 −δ.

C.2 Proof of Theorem 4.1

The proof derives a closed-form least-squares estimator and analyzes its approximation error using operator identities and gradient calculations. A high-probability bound then holds simultaneously for every Φ in the function space.

  • The proof begins by deriving a closed-form least-squares estimator before analyzing its approximation error.
  • Differentiating the least-squares objective over H yields the gradient relation ∇L_t(F) = 2C_tF − 2P_t.
  • The proof defines a cumulative operator M_1:t mapping sequential observations into H and uses Woodbury’s identity.
  • Operator identities show that the relevant terms map F into F, with C_t defined as λI + P_t.
  • Trace duality and Cauchy–Schwarz complete the bound, which holds with probability at least 1 − δ simultaneously for every Φ ∈ F.

C.3 Proof of Corollary 4.2

The corollary specializes the theorem to a fixed finite-dimensional subspace S containing the ranges of the sensing operators. Under boundedness and rank conditions, the resulting error bound has a regularization term of order O(t^-1) and a stochastic term of order O(log t/t).

  • The corollary introduces G_t := P_t under the assumptions of Theorem 4.1.
  • The result assumes a fixed d-dimensional subspace S ⊂ F containing Ran(M_t), with uniformly bounded operator norms ∥M_t∥op ≤ L.
  • With probability at least 1 − δ, the bound holds simultaneously for every t ≥ t0 and Φ ∈ S.
  • Because the ranges lie in S, that subspace is invariant under G_t and C_t = λI + G_t, enabling finite-dimensional norm comparisons and determinant bounds.
  • For fixed δ, Φ, and problem constants, the regularization term is O(t^-1), while the stochastic term is O(log t/t).

C.3.1 Proof of Theorem 4.3

The proof combines conditional sub-Gaussian noise properties with operator inequalities, determinant simplifications, and gradient decompositions. It then applies these results to the parametric estimator and target operator to establish the stated bound.

  • Letting F_t be the whole-space least-squares minimizer, the proof compares F(·, bθ_t) and F⋆ using their global-minimizer relationship.
  • Applying Lemma B.3 to both the estimated parametric operator and F⋆, then using the triangle inequality, yields the comparison bound.
  • Each M_iξ_i is conditionally M_iΣ_iM_i*-sub-Gaussian, with a trace-class covariance proxy because M_i is Hilbert–Schmidt and Σ_i is bounded.
  • The proof inverts an operator inequality to obtain a time-uniform high-probability statement.
  • The determinant is simplified and the resulting expression is substituted into the preceding bound to prove the stated result.
Loading 2608.24426v1…