Source-linked AI summary
What Neural Network Field Theory Can and Cannot Realise on a Computer
Thomas R. Harvey
TL;DR
The paper asks when neural-network ensembles can serve as computational QFTs or EFTs rather than merely provide formal representations. It proves an obstruction for regular pointwise ensembles, analyzes the four finite- and infinite-width interpretations, and finds that only the infinite-width EFT is fully computable, while the QFT limit is computable only for smeared correlators.
Problem
The central question is whether neural-network field-theory representations can be made amenable to finite computation with controlled errors for observables, rather than existing only as infinite representations.
Method
The paper proves a no-go theorem for pointwise, Euclidean-invariant, mean-square-continuous ensembles with reflection positivity, then applies it across four finite- and infinite-width QFT/EFT interpretations.
Results
In d ≥2, finite-width ensembles with finite pointwise variance cannot be exact reflection-positive QFTs; the infinite-width EFT is fully computable, while the QFT limit is only partly computable through smeared correlators.
Takeaways & Limitations
At the level of controlled numerical computation, the QFT and EFT limit versions cannot be distinguished when only smeared correlators are accessible.
Takeaways & Limitations
The conclusions assume finite pointwise variance, exact Euclidean invariance, mean-square continuity, and reflection positivity, with possible escapes requiring divergent pointwise variance or broken rotation invariance.
Abstract
from arXiv · showhide
One aim of neural network field theory is to put a quantum or effective field theory on a computer, with the network ensemble itself as the theory. We ask how far that aim can be pushed for a function class regular enough to be computed with. Our main result is a no-go theorem with assumptions that hold for standard network architectures. We use it to separate four versions of neural network field theory, according to whether the defining object is the finite width ensemble or its infinite width limit, and whether the target we want to compute is a quantum or an effective field theory. Neither finite width interpretation is straightforwardly consistent. For finite width ensembles with finite variance at each point, the QFT interpretation fails reflection positivity, while the EFT interpretation establishes no scale separation by which the positivity violation can be placed outside its domain of validity. Of the two limit versions, one can be simulated in full and the other only in part, as only its smeared correlators are computable with a controlled error. As such, at the level of a controlled numerical computation, the QFT and EFT versions cannot be distinguished. One dimension escapes the obstruction, yet reflection positivity is shown to still fail there at every finite width for the cosine network. Two escapes from the theorem remain, giving up either finite variance at a point or exact rotation invariance, and we discuss both of these possibilities.
1 Introduction
NN-FT seeks to use neural-network ensembles as field theories that can be computationally evaluated, while distinguishing this goal from representation and sampling applications. The paper frames interactions, finite computation, and four QFT/EFT interpretations as central challenges.
- Motivation: Infinite-width neural-network ensembles behave as Gaussian processes and therefore describe free, generally non-local Euclidean field theories.Their two-point function is determined by the neural network Gaussian-process kernel; finite-width higher-point connected functions are suppressed by 1/N.
- Motivation: Interactions can arise from 1/N effects at finite width or from parameter correlations that survive the infinite-width limit.Matching only the perturbative 1/N expansion is insufficient because the full finite-width theory also contains non-perturbative contributions.
- Computational aim: A computational representation must make observables computable from finite samples with errors quantified from the ensemble, rather than merely match correlators in principle.Monte Carlo averaging over parameter draws introduces statistical error, while finite computation requires controlled errors without a previously known answer.
- Computational aim: Finite-width network measures need not coincide with perturbative effective actions, and their finite pointwise variance conflicts with typical continuum propagator singularities.Edgeworth expansions can match correlators order by order in 1/N without defining the exact finite-width probability measure.
- Four interpretations: The paper separates four NN-FT versions by whether the defining object is finite width or infinite width and whether the target is a QFT or EFT.At finite width, N is a theory parameter and 1/N terms are interactions; in the limit, finite N is a truncation and interactions must survive N →∞.
2 The Obstruction
The paper proves that, in d ≥2, pointwise ensembles with finite variance, Euclidean-invariant continuous two-point functions, and reflection positivity must have constant two-point functions. This rules out nontrivial propagating theories under these assumptions.
- Reflection positivity: Reflection positivity is the Osterwalder–Schrader condition that reconstructs a physical Hilbert space from Euclidean correlators.Its failure means the correlators do not define a reflection-positive measure or a corresponding quantum theory.
- Assumptions: Finite variance at a point means each field value is an L2 random variable, whereas continuum QFT fields are generally distribution-valued and only smeared observables have finite variance.Typical network ensembles satisfy the stronger pointwise condition by construction.
- Theorem 2.1: In d ≥2, a pointwise ensemble satisfying Euclidean invariance, mean-square continuity, and reflection positivity has a constant two-point function and is therefore trivial.The resulting field does not fluctuate from point to point and carries no propagating degrees of freedom.
- Theorem 2.1: For every non-constant ensemble satisfying the other assumptions, the two-point Osterwalder–Schrader form has a finite configuration with a negative eigenvalue.The obstruction uses only the field-evaluation restriction of the full positivity form, so higher correlators are unnecessary for the theorem.
- Finite width: At finite width, the minimum eigenvalue of the two-point restriction is negative whenever the finite coincident variance condition holds.For usual iid constructions, this two-point matrix can be independent of N even though higher-point functions receive finite-width corrections.
3 Consequences of the Obstruction in d ≥2
The obstruction rules out the finite-width QFT interpretation under the theorem’s assumptions and makes the finite-width EFT interpretation difficult to control. The infinite-width limit can represent an EFT computationally, while a QFT limit is only partly computable if its pointwise variance diverges.
- (QFT | finite N): A finite-width Euclidean-invariant continuous ensemble with non-constant two-point function cannot define an exact reflection-positive QFT in d ≥2.The theorem applies at every N, so no Hilbert space or Wick-rotated quantum theory exists for this interpretation.
- (QFT | the limit): If the infinite-width limit remains pointwise with finite coincident variance, the same theorem forces its two-point function to be constant.Escaping while retaining the other assumptions requires leaving the pointwise class through divergent coincident variance.
- (QFT | the limit): Smeared correlators of a non-pointwise infinite-width limit are computable with controlled error, but singular composite operators are not.At controlled numerical accuracy, this QFT-limit computation is indistinguishable from computing an effective theory.
- (EFT | finite N): The finite-width EFT interpretation is not ruled out, but the paper finds no demonstrated scale separation placing positivity violations outside the EFT domain.The theorem identifies a negative direction without fixing its momentum scale, while non-perturbative finite-width effects may remain relevant.
- (EFT | the limit): For an EFT defined and regulated at infinite width, finite-width errors can be pushed below the cutoff-induced systematic errors, making the network a legitimate correlator-computation method.Here finite N is a truncation of the limiting theory rather than the defining theory itself.
4 One Dimension
In one dimension, the infinite-width theory can be reflection positive, but finite-width corrections still violate reflection positivity for the cosine network under broad finite-mass conditions.
- Infinite-width limit: The obstruction absent in one dimension allows a finite coincident propagator, G(x, x) = 1/(2m), so the infinite-width limit can be reflection positive.This differs from d ≥2, where the theorem rules out the finite-width QFT interpretation for typical architectures.
- Infinite-width limit: For a bounded, non-constant, completely monotone kernel with finitely many masses, the infinite-width Gaussian theory is reflection positive.The kernel is a positive combination of finitely many decaying exponentials, and the associated matrix A∞ is positive semi-definite.
- Cosine-network construction: The cosine architecture realises every stationary kernel through the choice of ρ, including one-dimensional free-theory kernels, while remaining exactly stationary at finite width.In one dimension, stationarity of the two-point function is equivalent to Euclidean invariance.
- Finite-width corrections: The cosine network violates reflection positivity at every finite width N for targets with finitely many masses and at least one positive mass.The defect is at least order 1/N^2, and becomes order 1/N when ker A∞ contains a direction v with v^T Dv < 0.
- Finite-width corrections: Therefore, the finite-width QFT interpretation is false in every dimension: in d ≥2 for typical architectures and in d = 1 for the cosine network with finitely many masses.In one dimension, any defect must arise from finite-width corrections because the infinite-width theory is reflection positive but degenerate.
- Cosine-network construction: A cosine network can reproduce the free-field two-point function exactly at every width while its full finite-width measure still violates reflection positivity.Choosing ρ(w) = m/[π(w^2 + m^2)] gives K(t) = e^−m|t|; the violation is therefore not simply caused by network smoothness.
5 Two Ways Forward
The paper considers two ways to evade the obstruction: abandoning finite pointwise variance or exact rotation invariance. Each escape preserves some computational possibilities but introduces a distinct limitation on controlled continuum calculations.
- 5.1 Escape one: give up finite variance at a point.: Relaxing finite pointwise variance permits interacting limits, but computation must identify observables with finite-variance estimators and controlled regulator removal.Coincident-point divergences may leave smeared observables computable, while pointwise sampling ceases to be finite-variance.
- 5.1 Escape one: give up finite variance at a point.: Safe smeared observables retain shrinking Monte Carlo errors as draws increase and support convergence checks by increasing N.This requires divergences confined to coincident points and a well-defined two-point function at separated points.
- 5.1 Escape one: give up finite variance at a point.: Heavy-tailed feature-level divergences remain at separated points, preventing finite error bars even after smearing.Singular composite operators can become exponentially expensive as the cutoff is removed, with cost ∼e^Λ^(d−2) for d ≥3.
- 5.2 Escape two: give up exact rotation invariance.: Relaxing exact rotation invariance allows finite-variance architectures that reproduce discrete rotation symmetry at the level of two-point functions.A product of independent cosine networks satisfies reflection positivity across coordinate hyperplanes.
- 5.2 Escape two: give up exact rotation invariance.: Unlike the lattice spacing, conventional NN-FT hyperparameters do not independently control rotational-symmetry breaking while holding the target kernel fixed.Consequently, these constructions lack a demonstrated controlled continuum limit, although the limitation does not apply to neural-network representations in the broad universality class.
6 Conclusion
The conclusion applies the no-go theorem to the four NN-FT interpretations and identifies which versions remain computationally viable. It also clarifies that the two escapes relax different assumptions while preserving different kinds of control.
- 6 Conclusion: In d ≥2, pointwise Euclidean-invariant, mean-square-continuous, reflection-positive ensembles have constant two-point functions and are therefore trivial.The obstruction depends only on the two-point function, not on higher correlators.
- 6 Conclusion: Finite-width NN-FT cannot be an exact reflection-positive QFT in d ≥2 under the theorem’s finite-variance assumptions.In one dimension, the theorem is absent, but reflection positivity still fails at every finite width for the cosine network.
- 6 Conclusion: The infinite-width QFT interpretation is only partially computable: smeared correlators have controlled errors, but singular observables do not.At finite accuracy, its computation is indistinguishable from a regulated effective theory.
- 6 Conclusion: The finite-width EFT interpretation lacks an established scale separation that would place reflection-positivity violations outside its low-energy domain.The required separation would need independent scale dependence and validation on intended low-energy observables.
- 6 Conclusion: The infinite-width EFT interpretation remains a legitimate computational framework, with finite-width effects—including non-perturbative 1/N errors—below cutoff-imposed systematic errors.The network computes correlators of an effective theory defined and regulated at infinite width.
- 6 Conclusion: The two proposed escapes are to give up finite variance at a point or exact rotation invariance, each shifting rather than eliminating the computational challenge.The first requires identifying safe observables; the second lacks an independently tunable continuum parameter in conventional architectures.
A Proof of Theorem 2.1
Under finite pointwise variance, Euclidean invariance, continuity, and reflection positivity, the proof forces every spatially smeared two-point function to be constant in Euclidean time. For d ≥2, rotation invariance then forces the entire momentum measure to zero momentum, making the field spatially constant and the theory trivial.
- Step 2: reflection positivity implies convexity: Reflection positivity makes each spatially smeared Euclidean two-point function convex in Euclidean time.Applying reflection positivity to differences of smeared time slices yields midpoint convexity; continuity upgrades this to convexity.
- Step 3: rotational invariance makes the initial slope finite: For d ≥2, rotational invariance and spatial smearing give a finite initial slope at Euclidean time zero.On large momentum spheres, smearing restricts contributions to polar caps whose relative area scales as r^−(d−1).
- Step 4: convexity and a finite initial slope force the spectrum to zero momentum: A bounded convex two-point function with finite initial slope is non-decreasing and therefore constant.The two-point function is bounded above by its value at the origin, while convexity makes its derivative non-decreasing.
- Step 4: convexity and a finite initial slope force the spectrum to zero momentum: The nonnegative spectral integrand then forces the smeared momentum measure to be supported at p0 = 0.The support conclusion follows because every nonzero p0 contributes positively for some Euclidean time.
- Step 5: assembly: Repeating the argument for every reflection direction restricts the momentum measure to {0}, so K is constant and the field is spatially constant almost surely.The intersection of all rotated hyperplanes {p0 = 0} is the single point {0} when d ≥2.
B Reflection positivity fails at every finite width (d = 1) for the cosine network
For the cosine network in one dimension, the finite-width Osterwalder–Schrader form can be analyzed through an exact 1/N expansion on low-degree observables. The two-point block has no width correction, while higher moments create the reflection-positivity defect.
- Finite-width structure: The cosine-network two-point function is exact at every width because cross-neuron terms vanish after bias averaging.Consequently, the degree-one block receives no finite-width correction.
- Finite-width structure: Uniform bias averaging makes odd single-feature moments vanish and block-diagonalizes the Osterwalder–Schrader form by degree parity.The defect therefore lies in the odd sector, where degree-one and degree-three observables suffice.
- Finite-width structure: Finite width affects reflection positivity through higher-point functions even though the two-point data remain unchanged.For independent features, surviving corrections are organized by partitions of field insertions into same-neuron blocks.
- Finite-width structure: The exact expansion on the relevant observable space terminates at order 1/N^2 and uses only iid features with finite moments through sixth order.The termination and parity arguments do not depend on one dimension or a particular activation, although the theorem’s cosine-network application does.
B.1 Proof of Theorem 4.1
The proof constructs a null direction of the infinite-width Osterwalder–Schrader form for finitely supported spectral measures, then shows finite width couples it to a cubic observable. This coupling makes a two-dimensional restriction indefinite, proving reflection positivity fails at every finite width.
- Step 1: a null direction at infinite width: For a target with finitely many masses, r + 1 distinct times produce a reflected two-point matrix of rank at most r and a nontrivial kernel.The kernel vector satisfies separate identities for each positive mass.
- Step 2: the width correction cannot affect the two-point data: The null direction has no 1/N diagonal defect because the exact two-point function leaves the degree-one block unchanged.Indeed, v^T Dv = 0 for every degree-one direction.
- Step 3: finite width couples the null direction to a degree-three observable: Finite width couples the null direction to a degree-three observable through a connected single-feature four-point function.Kernel identities cancel terms with one time dependence, leaving a nonzero contribution for suitable choices of the total time.
- Step 4: the two-dimensional restriction is indefinite: A negative determinant in the restricted 2×2 Osterwalder–Schrader matrix gives one positive and one negative eigenvalue.The diagonal degree-three entry changes the magnitude of the negative eigenvalue but not its existence.
- Step 4: the two-dimensional restriction is indefinite: Reflection positivity therefore fails for every finite width, while the construction does not directly extend to representing measures with infinite support.For infinite support, the reflected two-point matrix may be strictly positive definite on every finite collection of times, leaving the theorem’s status open there.