Source-linked AI summary
Neural Network with Unbounded Activation Functions is Universal Approximator
Sho Sonoda, Noboru Murata
TL;DR
The paper asks whether neural networks with unbounded activations such as ReLU retain universal approximation, addressing transform settings that previously did not directly admit these activations. It constructs ridgelet transforms for Lizorkin distributions and derives reconstruction formulas, showing universal approximation and a constructive interpretation of what the network learns. Numerical examples support the admissibility analysis while indicating low-pass filtering for some non-admissible activation–ridgelet combinations.
Problem
Existing ridgelet-transform settings did not directly admit fundamental unbounded activations such as the sigmoidal function and ReLU, while their approximation property required analysis.
Method
The paper constructs ridgelet transforms for Lizorkin distributions and develops reconstruction formulas through Fourier, Radon, and related transform representations.
Results
Neural networks with unbounded non-polynomial activation functions have the universal approximation property, with numerical examples consistent with theoretical admissibility diagnoses.
Takeaways & Limitations
Under a constructive admissibility condition, the trained network can be obtained by discretizing the ridgelet transform, while some non-admissible combinations act as low-pass filters.
Takeaways & Limitations
The admissibility framework excludes polynomial activations, and the paper identifies extending the analysis to deep multi-input multi-output cascades as future work.
Abstract
from arXiv · showhide
This paper presents an investigation of the approximation property of neural networks with unbounded activation functions, such as the rectified linear unit (ReLU), which is the new de-facto standard of deep learning. The ReLU network can be analyzed by the ridgelet transform with respect to Lizorkin distributions. By showing three reconstruction formulas by using the Fourier slice theorem, the Radon transform, and Parseval's relation, it is shown that a neural network with unbounded activation functions still satisfies the universal approximation property. As an additional consequence, the ridgelet transform, or the backprojection filter in the Radon domain, is what the network learns after backpropagation. Subject to a constructive admissibility condition, the trained network can be obtained by simply discretizing the ridgelet transform, without backpropagation. Numerical examples not only support the consistency of the admissibility condition but also imply that some non-admissible cases result in low-pass filtering.
1 Introduction
The paper studies universal approximation with unbounded, non-polynomial activations by connecting neural-network integral representations to ridgelet transforms and harmonic analysis. It develops constructive admissibility and reconstruction ideas, including applications to ReLU and related activation functions.
- Integral representation: A neural network gJ is obtained by discretizing an integral representation whose hidden parameters are (a_j,b_j) and output parameters are c_j.The continuous output parameter T leads to a dual ridgelet transform representation.
- Approximation framework: Under admissibility of (ψ,η) and regularity of f, discretizing the reconstruction formula verifies the approximation property of neural networks with activation η.The paper presents this as a constructive route rather than relying only on backpropagation.
- Motivation and scope: The study constructs ridgelet transforms with respect to Lizorkin distributions to analyze neural networks whose activation function η is unbounded.The Lizorkin distribution space contains ReLU and truncated power functions.
- Activation functions: Unbounded non-polynomial activations, including ReLU and truncated power functions, are treated alongside traditional bounded activations such as sigmoidal and radial basis functions.The paper frames this analysis as addressing the limited analytical evaluation of ReLU-related hypotheses.
- Harmonic-analysis connection: The paper relates neural networks to harmonic analysis and tomography because the dual ridgelet transform represents the network and the ridgelet transform combines Radon and wavelet analysis.This connection builds on earlier constructive uses of Fourier and Radon transforms in approximation theory.
- Research gap: Existing ridgelet-transform settings did not directly admit sigmoidal functions and ReLU, motivating a formulation with W = S1_0 that admits them.This extends the choice of dual ridgelets while retaining the Lizorkin-distribution framework.
2 Preliminaries
The preliminaries define the function and distribution spaces, transforms, and convolution operations used to formulate ridgelet analysis. They also introduce Radon, Fourier, Hilbert, and backprojection operators as analytical foundations.
- Notation: The paper uses Y^{m+1} = R^m × R for hidden parameters (a,b), while X(R^m), Y(Y^{m+1}), Z(R), and W(R) denote transform domain, range, ridgelets, and dual ridgelets.These spaces organize the neural-network and ridgelet-transform notation.
- Function and distribution spaces: Lizorkin functions S0(R^k) have all moments vanishing, and Lizorkin distributions S1_0(R^k) are represented as S1(R^k) modulo polynomials.Within this work, polynomials are identified with zero in the Lizorkin distribution space.
- Convolution: The preliminaries characterize convolution classes and note that general distribution convolutions need not commute or associate, whereas specified Schwartz-class combinations are commutative and associative.These convergence and algebraic properties support later distributional constructions.
- Fourier analysis: The Fourier transform, inverse Fourier transform, and Hilbert transform are defined before the paper uses them in its transform and backprojection formulas.The Hilbert transform appears in the definition of the backprojection filter for odd dimensions.
- Radon analysis: The Radon transform integrates f over hyperplanes orthogonal to u, with (Ru)^⊥ denoting the orthogonal complement and du the sphere’s surface measure.The dual Radon transform and Radon inversion formula are introduced as core tools.
- Fourier and backprojection analysis: The Fourier slice theorem identifies the m-dimensional Fourier transform with the one-dimensional Fourier transform of the Radon transform along corresponding directions.The backprojection filter Λ_m is defined as a one-dimensional Fourier multiplier in the Radon variable.
3 Classical Ridgelet Transform
The classical ridgelet transform is presented through direct, polar, convolution, Fourier, and Radon-domain forms. Its reconstruction framework connects ridgelet analysis to wavelets and tomography under convergence and admissibility conditions.
- Definition: The ridgelet transform R_ψf integrates f against scaled and translated ridgelet functions ψ(a·x − b), with a technical weight |a|^s.The paper later sets s = 1 for simpler notation, while other formulations use different exponents.
- Well-definedness: For f ∈ L1(R^m) and ψ ∈ L∞(R), the ridgelet transform is absolutely convergent and defines a bounded bilinear operator into L∞(Y^{m+1}) when s = 0.The convergence estimate follows from Hölder’s inequality.
- Reconstruction: When ψ and η are admissible and the relevant regularity conditions hold, the reconstruction formula recovers f from the ridgelet transform and its dual.The admissibility constant must be finite and nonzero.
- Polar representation: In polar coordinates, ridgelet analysis decomposes into directional, scale, and translation variables, with the radius α defined reciprocally to emphasize its wavelet connection.The same parameter space is used under both Cartesian and polar parametrizations.
- Radon-domain interpretation: The ridgelet transform can be written as wavelet analysis of the Radon transform, using orthogonal decomposition and Fubini’s theorem to obtain equivalent expressions.The Fourier form follows from the convolution identity and the Fourier multiplier relation.
- Dual transform: The dual ridgelet transform has a Mellin-convolution representation and equals the composition of the dual Radon transform with the dual wavelet transform.This composition is stated explicitly after rewriting the scale and translation variables.
4 Ridgelet Transform with respect to Distributions
The section extends ridgelet transforms to distributions and establishes when they are well defined, bounded, injective, and compatible with dual operators. It also identifies trade-offs between function and ridgelet spaces and shows that non-admissible polynomial ridgelets can destroy injectivity.
- The admissible distributional framework requires weak definitions that agree with ordinary and convolution forms when the distributions are sufficiently regular.The convolution, dilation, reflection, and conjugation operations are interpreted for Schwartz distributions.
- The ridgelet transform is defined as a bilinear map on distributional function and ridgelet spaces selected from Table 4.The table specifies compatible domains, Radon-transform ranges, and ridgelet-transform ranges.
- The allowable function and ridgelet classes trade off against each other because convolution convergence limits the largest compatible ridgelet class for each function class.As the function space X increases, the compatible ridgelet space Z decreases, and conversely.
- Direct ridgelet-transform extension to non-integrable functions may require more sophisticated methods because the Radon transform can diverge.The section notes bounded-extension techniques for L2 as one later remedy.
- For ψ in the Schwartz space, the ridgelet transform is a bounded operator from L1(R^m) to L∞(Y^{m+1}).The proof uses the convolution form, Young’s inequality, the Radon-transform bound, and rapid decay of ψ.
- The ridgelet transform is injective for admissible ψ, but a polynomial ψ can make Rψf vanish for nonzero Laplacian inputs.The example uses f=Δg and a polynomial ψ satisfying ψ^(2)=0.
- When it exists, the dual ridgelet transform coincides with the dual operator of the ridgelet transform.This identification follows from injectivity and uniqueness of the dual operator.
5 Reconstruction Formula for Weak Ridgelet Transform
The section develops admissibility conditions and reconstruction formulas for weak ridgelet transforms in Fourier, real, Radon, and L2 settings. These results provide constructive conditions for admissible pairs while exposing algebraic and distributional scope limitations.
- 5 Reconstruction Formula for Weak Ridgelet Transform: The section derives reconstruction formulas in the Fourier and Radon domains and frames both as constructive routes beyond Fourier-only treatments.It also extends the ridgelet transform to L2 in a later subsection.
- 5.1 Admissibility Condition: The admissibility integral excludes the origin because products of singular distributions there are generally indeterminate, while its convergence is independent of the chosen neighborhood.Polynomial distributions are excluded because distributions supported at the origin cannot be admissible.
- 5.1 Admissibility Condition: An admissible pair (ψ,η) is characterized through local behavior near zero, integrability, and a backprojection-filter condition involving Λ.The structure theorem requires bounded one-sided limits for the auxiliary Fourier-domain quantity u.
- 5.1 Admissibility Condition: Corollary 5.5 gives a constructive procedure for building admissible pairs from a distribution η by selecting a suitable Schwartz function ψ.The construction relies on the sufficient conditions of Theorem 5.4 and a convergent nonzero normalization quantity.
- 5.2 Reconstruction Formula: For sufficiently regular f and an admissible pair, the reconstruction formula recovers f almost everywhere and at every point where f is continuous.The theorem imposes a stronger condition on u than the preceding admissibility theorem.
- 5.2 Reconstruction Formula: Extensions between broader distribution spaces are constrained because multiplication may be noncommutative or nonassociative and Fourier transforms may be undefined on some spaces.These issues affect extensions involving S′_0 and D′.
- 5.2 Reconstruction Formula: The Radon-domain reconstruction interprets wavelet analysis as a backprojection filter, with admissibility determining the filter Λ^m.The result assumes f is sufficiently smooth and u is real-valued, smooth, and integrable.
- 5.3 Extension to L2: For L2, admissibly decomposable pairs with Kψ,η=1 yield reconstruction, while self-admissible ψ permit a unique bounded extension preserving the L2 norm.The extension is obtained through approximation from L1∩L2 and Plancherel’s identity.
6 Neural Network with Unbounded Activation Functions
This section determines universal approximation by checking admissibility for activation functions represented as dual ridgelet functions. It verifies admissibility for several Lizorkin-distribution activations, including truncated powers, Dirac derivatives, Gaussian derivatives, sigmoidal derivatives, and softplus.
- Admissibility: Admissibility of η determines whether the corresponding neural network is a universal approximator through the reconstruction formulas.The criterion is based on Theorems 5.6, 5.7, and 5.11.
- Activation classes: Truncated power functions, including ReLU and the unit step function, are included among the candidate Lizorkin-distribution activation functions.The truncated powers z^k contain ReLU and the step function.
- Activation classes: RBFs and their derivatives belong to S(R), while Dirac’s delta and its derivatives belong to S1(R).These class memberships support their treatment as candidate activations in the paper’s framework.
- Admissibility results: Gaussian derivatives, Dirac derivatives, and sigmoidal derivatives satisfy parity-dependent admissibility conditions with Gaussian-based ridgelet functions.For the listed examples, odd or even derivative orders can make Kψ,η vanish and therefore fail admissibility.
- Admissibility results: Dirac’s delta can be admissible, unlike polynomial functions, while Gaussian derivatives and softplus have separately specified admissibility cases.The paper gives Dirac’s delta as a non-polynomial admissible activation and identifies σ^(-1) as admissible with ψ = Λ^mG^2.
7 Numerical Examples of Reconstruction
The numerical experiments test theoretical admissibility diagnoses by reconstructing a sinusoidal signal and a Shepp–Logan phantom. The results are generally consistent with theory, while several non-admissible or difficult-to-implement cases exhibit failure or low-pass filtering.
- Experimental setup: Admissibility diagnoses classify pairs as admissible, vanishing, or divergent according to whether Kψ,η converges to a non-zero constant, zero, or diverges.A non-zero limiting constant implies universal approximation by Theorem 5.6.
- Sinusoidal curve: The one-dimensional experiment reconstructs sin 2πx on [-1,1] using Gaussian-derived ridgelets and activations including softplus, sigmoidal functions, ReLU, the unit step, and Dirac’s delta.The signal uses Δx = 1/100, while the reconstruction parameters are discretized over a = b = [-30,30] with Δa = Δb = 1/10.
- Sinusoidal curve: Increasing the Gaussian-derivative order localizes the ridgelet transform, and the resulting transforms can be reconstructed with admissible activations.The case ψ = ΛG^2 can be reconstructed with two different activation functions.
- Sinusoidal curve: Softplus reconstruction can be incomplete because convolution with ΛG acts as an integrator and low-pass filter.The paper attributes this behavior to the pole ζ^-2 in the softplus Fourier transform.
- Truncated power functions: Dirac reconstructions fail in the experiment because discretized ridge parameters rarely make the activation argument exactly zero.This is presented as an implementation difficulty rather than a contradiction of the theory.
- Linear activation: Linear-function reconstructions fail consistently with the theory because polynomials cannot be admissible when their Fourier transforms are supported at the origin.The same theoretical and experimental consistency is reported across the reconstruction figures.
8 Concluding Remarks
The paper establishes universal approximation for neural networks with unbounded non-polynomial activations by constructing distributional ridgelet transforms and reconstruction formulas. It also identifies the learned representation and records extensions toward deep networks and admissibility choices.
- Contributions: Neural networks with unbounded non-polynomial activation functions have the universal approximation property.The result covers traditional radial basis, sigmoidal, unit-step, and truncated-power activations, including ReLU-related cases.
- Interpretation: Backpropagation indirectly searches for an admissible ridgelet function by constructing a backprojection filter for the target function.Under admissibility, the trained network can therefore be related to a discretized ridgelet transform rather than only to an optimization procedure.
- Ridgelet construction: The ridgelet transform is constructed for distributions, and the neural network’s integral representation coincides with its dual ridgelet transform.The construction uses weak-form expressions and establishes existence and operator properties for the distributional transform.
- Admissibility: Admissibility requires an unbounded activation to be non-polynomial and associated with a backprojection filter.A constructive sufficient condition is obtained, while admissible ridgelet choices are not unique for a fixed activation.
- Reconstruction: Reconstruction formulas are established through the Fourier slice theorem, Radon-transform inversion, and Parseval’s relation on L1 and L2 settings.The L1 results are extended to L2 through a weak reconstruction formula and bounded extension of the ridgelet transform.
- Numerical findings: Numerical examples agree with the admissibility analysis, while some non-admissible activation–ridgelet pairs behave as low-pass filters.Examples include pairs using Λ_m applied to Gaussians with ReLU or softplus.
A Proof of Theorem 4.2
This proof establishes that the ridgelet transform is well defined across several function and distribution spaces by analyzing Radon transforms, convolutions, growth, and parameter regularity.
- Transform definition: The ridgelet transform is defined as a convolution of the Radon transform with a dilated ridgelet distribution.The proof tracks the resulting function or distribution class under different choices of input and ridgelet spaces.
- Radon-transform mapping: The Radon transform continuously injects admissible function spaces on R^m into corresponding spaces on S^{m−1} × R.This embedding determines the domain classes considered in the proof.
- Distributional cases: For compactly supported inputs and distributional ridgelets, the convolution is shown to define a smooth function on the parameter space.In particular, the proof obtains W[ψ;g] ∈ E when g ∈ D and ψ ∈ D1.
- Growth and integrability: For Schwartz inputs and Lizorkin distributions, the convolution is shown to belong to OM, while related growth estimates establish local integrability and polynomial growth in L1-based cases.The proof also derives weighted Lp integrability under a condition relating the growth exponents and dimension.
- Tempered-distribution cases: The remaining cases extend the same estimates to tempered-distribution inputs and establish membership in S1 through continuity, convolution, and distributional bounds.The argument reduces some Lp cases to the previously treated L1 case and uses standard inequalities to control parameter growth.
- Admissibility products: The admissibility analysis defines products involving the Fourier transforms of the ridgelet and activation distributions as distributions, including when the activation transform is singular at zero.The proof separately establishes sufficiency and necessity under a local continuity assumption away from zero.
C Proof of Theorem 5.6
The proof of Theorem 5.6 derives a reconstruction formula for L1 functions by applying the Fourier slice theorem in the sense of distributions and using the admissibility normalization.
- Assumptions: For f ∈ L1(R^m) with Fourier transform in L1 and an admissible pair (ψ,η), the reconstruction formula is normalized by setting Kψ,η = 1.The theorem assumes the relevant ridgelet and activation distributions satisfy the admissibility requirements.
- Fourier-slice step: The Fourier slice theorem converts the m-dimensional Fourier transform into the one-dimensional Fourier transform of the Radon transform.The proof then changes variables between radial frequency, scale, and direction parameters.
- Reconstruction: Substituting β = u · x into the ridgelet representation yields the target function through Fourier inversion.The equality holds almost everywhere and at every point where f is continuous.
D Proof of Theorem 5.7
The proof of Theorem 5.7 reconstructs functions through approximation to the identity, controlling limiting convolutions and applying convergence arguments to the Radon-domain representation.
- Approximation kernel: A normalized kernel k with integral one is used as an approximation of the identity.The proof establishes k ∈ L1 ∩ L∞ and verifies its decay near zero and at infinity.
- Small-scale limit: As ε tends to zero, convolution with kε recovers J(u,p) almost everywhere.This provides the local limiting step in the reconstruction argument.
- Large-scale limit: As δ tends to infinity, convolution with kδ converges to zero almost everywhere.The proof uses this contrasting limit to isolate the reconstruction contribution.
- Limit interchange: Uniform integrability and the Vitali convergence theorem justify passing the limit through the directional integral over S^{m−1}.The resulting identity reconstructs the target function almost everywhere in R^m.
- L2 extension: The L2 argument restricts the ridgelet-domain integral to truncated regions and controls the remainder as those regions expand.This supplies the convergence mechanism used for the L2 reconstruction formula.
F Proofs of Example 6.4 and Example 6.10
The proofs establish growth and smoothness properties for σ, tanh, and derivatives of σ, then use parity and constructed functions to determine admissibility cases.
- Example 6.4: σ and tanh belong to OM(R), with bounded derivatives of σ established through polynomial derivative representations.The proof uses boundedness of σ and induction on derivative order; tanh follows similarly.
- Example 6.4: σ1 decays faster than any polynomial, so σ belongs to S(R) after establishing the corresponding bounds for all derivatives.The proof concludes σ1 ∈ S(R), and therefore σ^(k) ∈ S(R) for every k ∈ N.
- Example 6.10: σ^(-1) has at most polynomial growth and consequently belongs to OM(R).The proof analyzes the maximum of ρ at zero to establish the required growth control.
- Example 6.10: For ψ = Λ^mG, σ^(k) is admissible when k is positive and odd because σ^(k) is odd and its pairing with ψ0 vanishes.For even k, σ^(k) is even and the pairing does not vanish, so the stated admissibility conclusion does not apply.
- Example 6.10: σ and σ^(-1) are not admissible with ψ = Λ^mG, but they are admissible with ψ = Λ^mG1 and ψ = Λ^mG2, respectively.The constructed convolutions u0 and u-1 belong to S(R), satisfying the sufficient condition in Theorem 5.4.