Source-linked AI summary
Operator-Theoretic Generalization Bounds for Multitask Deep Learning
Mahdi Mohammadigohari, Thomas Borsani, Giuseppe Di Fatta
TL;DR
Deep-network generalization analysis often captures transformation geometry only coarsely. This paper uses Koopman composition operators to derive multitask RKHS bounds across Sobolev and Brownian regimes, alongside shared-operator learning results, with structurally distinct guarantees rather than a uniform ordering.
Problem
Existing deep-network capacity bounds often retain only coarse information about the geometry of layer transformations.
Method
The paper represents layers as Koopman operators on vector-valued RKHSs, analyzes Sobolev and Brownian regimes, and develops shared-operator learning theory.
Results
The analysis derives multitask Rademacher bounds, exact Brownian layer estimates, a finite-rank representer theorem, and a source-conditioned target-transfer bound.
Takeaways & Limitations
The Sobolev and Brownian guarantees apply to different hypothesis spaces and should be compared structurally rather than treated as uniformly ordered.
Takeaways & Limitations
The theorems exclude rank-deficient layers, and the Brownian result is one-dimensional with domain-preserving, anchored, bias-free architecture assumptions.
Abstract
from arXiv · showhide
We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces. In vector-valued Sobolev RKHSs, we derive Rademacher complexity bounds for invertible and width-expanding injective architectures. The estimates separate the output-coupling contribution, represented by the trace of the task matrix, from the layerwise operator norms, Sobolev symbol ratios, determinant factors, and restriction constants generated by the linear maps. We then analyze a distinct one-dimensional Brownian/Cameron--Martin regime. Using the exact anchored derivative-norm characterization of the vector-valued Brownian RKHS, we obtain layerwise bounds for domain-preserving scalar linear maps and anchored diffeomorphic activations; the corresponding factors scale as $|W_l|^{1/2}$ and $\|σ_l'\|_\infty^{1/2}$, respectively, and do not involve Sobolev smoothness exponents. Because the Sobolev and Brownian results concern different hypothesis spaces, neither is asserted to dominate the other uniformly. We additionally formulate shared operator learning across tasks, prove a finite-rank representer theorem, derive the exact finite-dimensional problem for squared loss, and establish a target-transfer bound when the learned operator is obtained independently of the target sample. Synthetic and MNIST studies examine stabilized Sobolev-inspired and Brownian-inspired complexity proxies; these empirical proxies are not evaluations of the proved bounds for rank-deficient architectures.
1 Introduction
The paper studies deep multitask generalization through induced composition operators in vector-valued Sobolev and one-dimensional Brownian/Cameron–Martin RKHSs. It separates task coupling and layer geometry while emphasizing that the resulting bounds apply to different, non-uniformly comparable hypothesis classes.
- Motivation: Classical generalization analyses control capacity through parameter counts, margins, compression, or weight-matrix norms, while norm-based estimates can avoid explicit width dependence.The introduction motivates a function-space operator viewpoint as an alternative to these capacity controls.
- Brownian/Cameron–Martin viewpoint: In the one-dimensional Brownian/Cameron–Martin regime, the exact RKHS norm is an anchored first-derivative energy, enabling direct change-of-variables analysis without Fourier extensions.Domain-preserving conditions ensure that compositions remain inside the RKHS.
- Brownian/Cameron–Martin viewpoint: The Brownian and Sobolev bounds concern different function-space geometries and are not presented as uniformly comparable numerical estimates.The distinction follows because the bounds apply to different hypothesis classes.
- Vector-valued Sobolev Koopman bounds: The paper extends Koopman factorization to separable vector-valued Sobolev RKHSs, deriving multi-output Rademacher bounds with output coupling through Tr (M) and layerwise geometric factors.Linear layers contribute operator-norm, Sobolev-symbol, and determinant factors; the terminal norm is M−1-weighted.
- Injective width expansion: For injective non-square layers, the analysis includes explicit Sobolev trace conditions, inverse transposes on layer ranges, and restriction costs from lower-dimensional ranges.These restriction constants arise from the geometry of width-expanding injective maps.
2 Related Work
Prior work develops norm-based capacity bounds, Koopman formulations, vector-valued RKHS multitask theory, and finite-rank operator-learning reductions. This paper extends these directions to multi-output Sobolev and Brownian settings and vector-valued multitask operator learning.
- Norm-based generalization bounds: Norm-based analyses control deep-network capacity using spectral, Frobenius, path, or reference-matrix norms, while compression methods exploit approximate low-rank structure.These approaches avoid direct parameter counting but generally summarize layers through a limited collection of matrix norms.
- Koopman formulations: Koopman learning represents nonlinear dynamics through linear composition operators, including determinant-sensitive Sobolev-RKHS bounds for scalar networks with full-rank weights.The present work extends this construction to multi-output vector-valued RKHSs and injective width-expanding maps, while using an anchored one-dimensional regime for its Brownian result.
- Vector-valued RKHSs and multitask learning: Vector-valued RKHSs provide a standard framework for coupled outputs and multitask learning, with prior guarantees covering linear classes, trace-norm regularization, and representation learning.The cited literature includes Rademacher and excess-risk guarantees for multitask settings.
- Operator learning and transfer: Prior Koopman operator-learning work uses representer theorems to reduce Hilbert-space operator problems to finite rank; this paper applies the projection principle in a vector-valued multitask setting.The transferable object is a bounded operator between Hilbert function spaces rather than only an output kernel or finite-dimensional latent vector.
3 Preliminaries
This section establishes the notation and vector-valued RKHS framework used throughout, including Sobolev and Brownian spaces, Rademacher complexity, and Koopman composition operators. It also specifies the deep-network architecture and explains how layerwise operator norms, trace terms, determinant factors, and Sobolev symbols enter later bounds.
- Notation and RKHS conventions: The preliminaries define operator notation, range-volume determinants, positive-semidefinite matrix cones, scalar and matrix-valued kernels, and reproducing properties for vector-valued RKHSs.These conventions include ran(A), ker(A), operator norms, S_m^+, and the separable-kernel construction K(x,x′)=k(x,x′)M.
- Sobolev and matrix assumptions: The main Sobolev and Brownian theorems assume M ≻ 0, while target transfer later permits M ⪰ 0 using only reproducing and diagonal trace bounds.Sobolev spaces use Bessel-potential kernels with order s>d/2 and a specified Fourier norm; multi-index weak derivatives are also fixed.
- Koopman representation: Koopman operators represent composition by measurable maps, including translations, and provide the operator-theoretic language for network-layer analysis.For a map ϕ, K_ϕg=g∘ϕ, while translations are induced by b(x)=x+a.
- Brownian RKHS: The one-dimensional Brownian setting uses a separable signed kernel on [−R,R] and the anchored Cameron–Martin geometry of the vector-valued Brownian RKHS.The exact two-sided signed-kernel characterization is supplied later, and ∥σ′∥∞ denotes the supremum of |σ′| on the stated domain.
- Deep-network setup: The network class composes affine maps, coordinatewise activations, and a terminal map, with Sobolev spaces H^s_l induced by K^s_l=k^s_lM and bounded activation Koopman operators.The layer composition order is right to left, and admissibility requires activation boundedness and a terminal map in H^s_L.
- Layerwise complexity factors: Koopman factorization separates translations, activations, and linear maps, combining their norms submultiplicatively while assigning trace, determinant, and Sobolev-symbol factors to distinct mechanisms.The trace of M arises from vector-valued RKHS evaluation, whereas determinant and Sobolev-symbol factors arise from linear composition operators.
4 Vector-Valued Sobolev Koopman Bounds
This section derives vector-valued Sobolev Rademacher bounds for invertible and width-expanding injective Koopman architectures. The bounds separate task coupling, layerwise operator effects, Sobolev-symbol and determinant factors, and restriction costs, while excluding rank-deficient extensions.
- Core Sobolev decomposition: The task matrix enters through Tr(M) and the M^-1-weighted terminal norm, while layers contribute Koopman norms, Sobolev-symbol ratios, determinant terms, and nonlinear composition norms.These components jointly determine the Sobolev estimates for vector-valued RKHS network classes.
- Invertible architectures: Theorem 1 establishes the invertible Sobolev bound under the stated assumptions and monotone Sobolev-order condition.The monotonicity ensures finiteness of every Sobolev-symbol supremum for invertible linear maps.
- Invertible architectures: Translations act isometrically on the weighted Sobolev spaces, so the bound contains no bias-dependent factor.This removes translation contributions from the layerwise complexity estimate.
- Injective architectures: Theorem 2 extends the Sobolev bound to width-expanding injective maps under Sobolev trace conditions, with finite uniform restriction constants supplied by the trace theorem.The restriction constants arise from bounded traces onto lower-dimensional linear subspaces and rotational invariance.
- Injective architectures: The realized restriction ratio yields a function-dependent estimate for a fixed network but cannot replace the uniform restriction constant in the class-level Rademacher theorem.The ratio measures the restriction cost of the actual downstream function in the Koopman chain.
- Scope and limitations: Rank-deficient weight matrices fall outside the theorem because injectivity supports bounded pullbacks and nonzero range-volume determinants; stabilized factors are empirical surrogates, not proofs of boundedness.A rigorous non-injective extension requires a different function-space or quotient-space analysis.
5 Vector-Valued Brownian/Cameron–Martin Koopman Bound
The Brownian/Cameron–Martin regime derives a one-dimensional Koopman generalization bound from the anchored derivative-norm characterization of a vector-valued Brownian RKHS. Its assumptions and conclusions are separate from the Sobolev analysis rather than a uniform improvement over it.
- Architecture assumptions: The theorem analyzes domain-preserving scalar linear maps and anchored C1 diffeomorphic activations on X = [−R, R].The conditions C ≤ 1, σ_l(0) = 0, and σ_l : X → X exclude bias translations.
- Hypothesis space: The Brownian RKHS consists of anchored, absolutely continuous vector-valued functions with square-integrable derivatives.Every function satisfies f(0) = 0, with f′ ∈ L2(X, Rm).
- Scope and comparison: The Brownian theorem is one-dimensional, requires every layer to preserve the interval and anchor, and applies to a different RKHS from the Sobolev results.It should therefore be interpreted as a separate operator-theoretic regime, not a uniform improvement over multidimensional Sobolev bounds.
6 Shared Operator Learning and Transfer
This section develops shared operator learning across source tasks, proving a finite-rank representer theorem and an exact squared-loss coefficient formulation. It also derives a conditional target-transfer bound that separates approximation from estimation without asserting universal transfer improvement.
- Shared operator learning: Shared operator learning has a unique minimizer in the Hilbert–Schmidt operator space under the stated convex, continuous, finite-valued loss assumptions.The learned operator maps terminal representations from the source-task Hilbert space into the vector-valued RKHS.
- Shared operator learning: The unique minimizer admits a finite-rank representer expansion built from kernel sections associated with the source training data.This is the section’s central structural result for shared operator learning.
- Squared-loss reduction: Under squared loss, representer coefficients solve an exact finite-dimensional optimization problem with the same optimal value as the original operator problem.Conversely, every coefficient tensor defines a finite-rank operator of the representer form.
- Target transfer: Conditionally on a source-learned operator independent of the target sample, the target bound controls approximation and estimation through the induced class’s complexity.The result assumes bounded loss, Lipschitz dependence on predictions, and measurability of the relevant suprema.
- Target transfer: The transfer result does not claim that sharing always improves target performance; it requires a useful source-learned operator with controlled norm.If the operator is a Koopman chain, its norm can be bounded using the corresponding layerwise estimates.
7 Empirical Proxy Study
The study evaluates stabilized numerical proxies motivated by the theory rather than the proved bounds, using architectures that include rectangular and rank-deficient layers. Across five initializations, BP is numerically smaller and less variable than SP, but this descriptive comparison does not establish uniform theorem dominance.
- Experimental scope: The experiments use stabilized factors motivated by layer geometry, not direct evaluations of Theorems 1 to 3.Rectangular and rank-deficient architectures are included, and determinants are stabilized by adding the identity.
- Experimental setup: The setup uses W1 ∈ R3×3, W2 ∈ R6×3, b1 ∈ R3, and b2 ∈ R6.The activation is a smooth Leaky ReLU, with orthogonal weight initialization and uniformly initialized biases.
- Experimental setup: Training runs for 1600 epochs with learning rate 3 × 10^-3 and L2 penalty 10^-4.These implementation choices accompany the stated initialization and activation scheme.
- Proxy comparison: Across five initializations, BP is numerically smaller and less variable than SP.The comparison concerns the selected stabilized surrogates and is reported descriptively.
- Interpretation and limitation: Identity stabilization keeps both proxies finite for rectangular or rank-deficient matrices, so they cannot be read as determinant factors from the proved injective theorem.The observed BP–SP ordering does not show that the Brownian theorem uniformly dominates the Sobolev theorem.
8 Limitations
The results have distinct scope limitations: the Sobolev theorems require specific operator and map conditions, while the Brownian theorem is one-dimensional. Rank-deficient architectures are not covered by stabilized determinants; that replacement is used only as an experimental proxy.
- Sobolev limitations: Sobolev theorems require invertible or injective linear maps and bounded activation composition operators on selected weighted Sobolev spaces.The injective theorem also requires stated trace conditions, and its restriction constants may be large.
- Empirical proxy: Rank-deficient layers are not covered by replacing a singular determinant with a stabilized one; experiments use that replacement only as a proxy.Thus, stabilized-determinant experiments are not evaluations of the proved bounds for rank-deficient architectures.
- Brownian limitations: The Brownian theorem is one-dimensional and requires |W_l| ≤ 1.Its scope differs from the Sobolev results, which concern different hypothesis spaces.
9 Conclusion … A.2 Translation and invertible linear composition
The paper develops operator-theoretic Rademacher bounds that separate task coupling from layerwise Sobolev composition costs, while also establishing distinct Brownian/Cameron–Martin estimates. The appendices derive a common vector-valued RKHS estimate and prove translation isometry and invertible linear pullback bounds.
- 9 Conclusion: The Sobolev analysis separates task coupling, terminal-map size, activation norms, determinant-based volume distortion, and restriction costs for injective width expansion.The conclusion identifies these as distinct contributions in the multi-output composition bounds.
- 9 Conclusion: The Brownian/Cameron–Martin regime uses an exact anchored derivative norm to obtain direct linear and activation estimates without Sobolev exponents or Fourier extensions.This regime is presented as distinct from the Sobolev hypothesis-space analysis.
- A Proofs of the Sobolev Bounds: The Sobolev proofs use the M−1-weighted Sobolev norm, first establishing a common vector-valued RKHS Rademacher estimate before deriving layerwise composition bounds.The appendix states that every calculation is performed in the weighted norm used by the theorem statements.
- A.1 A common vvRKHS Rademacher estimate: For operator-image classes, the common estimate bounds Rademacher complexity through the family’s operator norms and the vector-valued RKHS evaluation term.The estimate is formulated for bounded linear operators from a Hilbert space into a matrix-valued-kernel vvRKHS.
- A.1 A common vvRKHS Rademacher estimate: When K(x,x′)=k(x,x′)M and k(xi,xi)≤κ, the common estimate further reduces using the task matrix and the diagonal kernel bound.The appendix derives this specialization through the reproducing property and Rademacher-coordinate calculations.
- A.2 Translation and invertible linear composition: Translations preserve the weighted Sobolev norm exactly, so the associated Koopman operator is an isometry.The Fourier transform contributes only a scalar phase of modulus one.
- A.2 Translation and invertible linear composition: For an invertible linear pullback, the weighted Sobolev operator norm is controlled by a supremum of the transformed Sobolev-weight quotient together with the Jacobian factor |det(W)|.The proof applies Fourier scaling and the change of variables ξ=W^Tω before taking square roots.
A.3 Proof of the invertible theorem and its corollary … B.1 Exact vector-valued Brownian RKHS
The appendices prove the invertible and injective Sobolev composition bounds through Koopman-operator factorization, restriction, and pullback estimates. They also identify the vector-valued Brownian RKHS exactly as an anchored absolutely continuous Cameron–Martin space with weighted derivative norm.
- A.3 Proof of the invertible theorem and its corollary: The invertible theorem follows by representing each network as a complete Koopman chain and applying operator-norm submultiplicativity.The resulting estimate is obtained after taking the supremum over admissible weight tuples and applying the terminal Hilbert-space Rademacher bound.
- A.3 Proof of the invertible theorem and its corollary: The invertible corollary specializes all Sobolev smoothness indices to a common value and uses the determinant constraint to uniformly bound each linear-layer contribution.The determinant condition yields |det(W)|^1/2 ≤ 1/D^1/2, after which substitution into Theorem 1 proves the corollary.
- A.4 Injective linear composition and proof of the injective theorem: For injective maps, weighted Sobolev restriction to the range is bounded with a constant independent of the subspace orientation and task matrix M.The restriction threshold is s_out > (d_out − d_in)/2, and the vector-valued estimate is obtained componentwise after output whitening.
- A.4 Injective linear composition and proof of the injective theorem: The injective pullback factors through the range restriction and an invertible square pullback in orthonormal range coordinates.Writing W = QA and using injectivity to make A invertible transfers the square-coordinate estimate into coordinate-free form.
- A.4 Injective linear composition and proof of the injective theorem: The injective theorem combines layerwise restriction and range-pullback estimates with Koopman-chain submultiplicativity before applying the terminal RKHS bound.Every function has the form Tg with terminal norm at most B_g, while the input kernel is K_s0 = k_s0 M and has scalar diagonal bounded by κ.
- B.1 Exact vector-valued Brownian RKHS: The Brownian feature map uses signed interval indicators, whose L2 inner products reproduce the scalar Brownian kernel, including the zero contribution for opposite signs.The construction is verified by separating nonnegative, nonpositive, and opposite-sign cases.
- B.1 Exact vector-valued Brownian RKHS: The vector-valued Brownian RKHS is exactly the space of anchored absolutely continuous functions with square-integrable derivatives, equipped with the M-weighted derivative inner product.The identification is isometric and follows by constructing a bijection from weighted L2 derivative space, then verifying the reproducing property.
B.2 Brownian composition operators … C.6 A two-sided expected-Rademacher deviation lemma
The paper proves Brownian-RKHS composition bounds for domain-preserving scalar maps and anchored diffeomorphisms, then extends these tools to shared-operator learning and transfer analyses. The shared-operator problem has a unique finite-rank representer, an exact finite-dimensional squared-loss reduction, fixed-image complexity bounds, and two-sided expected-Rademacher deviations.
- B.2 Brownian composition operators: |W| ≤1 ensures scalar linear composition preserves the interval and yields the Brownian-RKHS operator-norm bound.The proof uses the chain rule, a change of variables, and nonnegativity of the quadratic form.
- B.2 Brownian composition operators: Anchored C1 diffeomorphisms preserve the Brownian-RKHS structure, with the inverse derivative canceling exactly rather than contributing a multiplicative norm factor.The anchor and domain conditions justify the composition, while the diffeomorphic change of variables supplies the cancellation.
- B.3 Proof of the Brownian theorem: The Brownian theorem follows by combining submultiplicativity with the scalar-map and anchored-diffeomorphism lemmas, then applying the RKHS complexity lemma to complete Koopman chains.The proof takes the supremum over admissible layer weights and applies the input-kernel diagonal bound.
- C.1–C.2 Shared operator foundations: Hilbert–Schmidt rank-one identities provide the operator norms and inner products used throughout the shared-operator analysis, while coercivity and strict convexity establish a unique minimizer.The shared prediction map is bounded on a Hilbert space, and the regularized finite-dimensional loss attains its minimum.
- C.3 Proof of the shared-operator representer theorem: Every shared-operator minimizer lies in the finite span of training-derived rank-one operators, yielding a finite-rank representer with coefficients ctia.Orthogonal components do not affect training predictions but strictly increase regularization when λ > 0.
- C.4 Proof of the finite-dimensional reduction: Under squared loss, substituting the representer produces an exact coefficient objective whose value matches the original infinite-dimensional optimization problem.The reduction follows from rank-one inner-product identities and the vector-valued RKHS reproducing property.
- C.5 Rademacher complexity of a fixed operator image: The fixed-operator target class admits conditional Rademacher bounds in both operator norm and Hilbert–Schmidt norm forms.The second estimate follows because every Hilbert–Schmidt operator is bounded in the relevant norm comparison.
- C.6 A two-sided expected-Rademacher deviation lemma: For functions bounded in [0, C], the deviation lemma gives simultaneous upper and lower empirical-to-population inequalities with probability at least 1 −δ.The proof combines ghost-sample symmetrization, bounded differences, McDiarmid’s inequality, and a union bound.
C.7 Proof of the conditional target-transfer bound … E.1 Proxy regularization on MNIST
The paper proves a conditional target-transfer bound under source–target independence, establishes bounded Sobolev composition for sufficiently smooth bi-Lipschitz activations, and evaluates theory-motivated MNIST penalties only as stabilized proxies. The experiments compare Brownian-inspired and Sobolev-inspired regularization without claiming direct evaluation of the proved bounds.
- C.7 Proof of the conditional target-transfer bound: Conditioning on source data fixes the learned target operator while preserving i.i.d. target observations, enabling a conditional generalization argument.The proof explicitly relies on independence between source data and target observations.
- C.7 Proof of the conditional target-transfer bound: The conditional proof combines vector contraction, empirical Rademacher control, and conditional deviation inequalities to compare the empirical minimizer with the best admissible predictor.The resulting comparison introduces the conditional complexity radius and approximation slack before taking the infimum.
- C.7 Proof of the conditional target-transfer bound: The asserted bound remains valid when the operator norm is replaced by the Hilbert–Schmidt norm, because every Hilbert–Schmidt operator has operator norm no larger than its Hilbert–Schmidt norm.This replacement is made after establishing the conditional bound.
- D A Sufficient Sobolev Activation Condition: A C^s bi-Lipschitz diffeomorphism with bounded derivatives through order s induces bounded scalar composition on integer-order Sobolev spaces.The proof uses weak-derivative norm equivalence, change of variables, and the multivariate chain rule.
- D A Sufficient Sobolev Activation Condition: The same activation composition operator is bounded on the M−1-weighted vector-valued Sobolev space by applying the scalar estimate componentwise after the constant output transformation.Taking square roots yields boundedness in the weighted vector norm.
- E.1 Proxy regularization on MNIST: The MNIST study trains a fully connected network with hidden widths 1024, 2048, and 2048, followed by a ten-dimensional output layer, for 1800 epochs on 1000 examples.Adam optimization uses the stated initialization schemes and smooth Leaky ReLU activation.
- E.1 Proxy regularization on MNIST: γ = 0.01 weights each proxy penalty, whose identity-shifted determinant stabilizes rectangular or rank-deficient weights but differs from determinant terms in the proved injective bound.The proxies are optimization penalties motivated by theory rather than evaluations of Theorems 1 to 3 for the experimental architecture.
- E.1 Proxy regularization on MNIST: The Brownian-inspired penalty produces a higher mean test-accuracy trajectory than the unregularized baseline, while the experiment reports trajectories across five independent initializations.The passage contrasts this result with the Sobolev-inspired penalty but does not provide its completed outcome here.