Source-linked AI summary

S-matrix informed neural networks for amplitude analysis

Wyatt A. Smith, Arkaitz Rodas, Marius D. Thomas, César Fernández-Ramírez, Giorgio Foti, Lin Qiu, Adam P. Szczepaniak, Alessandro Pilloni

arXiv:2608.23750v1hep-phcs.LGnucl-th

TL;DR

Scattering amplitudes must be inferred from finite, noisy, and mutually inconsistent measurements while preserving first-principles constraints. This paper introduces SINNs and constrained ensemble responses for amplitude learning and compatible-data selection, applying them to ππ scattering. The resulting workflow produces direct amplitudes with correlated uncertainties and shows negligible sensitivity to model architecture in the reported validation studies.

  • Problem

    Scattering amplitudes must be reconstructed from finite, noisy, and mutually inconsistent measurements, while reliable derived quantities require accurate amplitudes and faithful uncertainties.

  • Method

    SINNs impose S-matrix constraints during optimization, use constrained ensemble responses to select compatible experiments, and provide direct amplitude representations without a fixed functional ansatz.

  • Results

    The framework produces low-energy ππ amplitudes and correlated uncertainties, with negligible impact of model architecture on central observable values.

  • Takeaways & Limitations

    The workflow unifies physics-constrained representation learning, compatible-data selection, and uncertainty quantification for ππ scattering.

  • Takeaways & Limitations

    The selected data set implicitly depends on the architecture used for spectral response clustering, and this coupled systematic has not been exactly quantified.

Abstract

from arXiv · show

Reconstructing scattering amplitudes from finite, noisy, and mutually inconsistent measurements is an ill-posed inverse problem common to many reactions relevant to particle physics. We introduce S-matrix informed neural networks (SINNs), and demonstrate their ability to learn scattering amplitudes directly from data while respecting first principles. We further develop a novel data selection procedure, which uses the response of constrained neural network ensembles to identify a set of experiments compatible with first principles, and with each other. We apply this framework to $ππ$ scattering, producing reusable amplitudes and correlated uncertainties without relying on a fixed functional form. We validate our results against residual model dependencies and training biases through closure tests and ablations. We find negligible impact of model architecture on our results. Our workflow unifies physics-constrained representation learning, data selection, and uncertainty quantification. Our strategy is transferable to other scattering processes, and other constrained physics problems limited by inconsistent data.

I. INTRODUCTION

Scattering amplitudes are essential for extracting hadronic information relevant to precision and New Physics studies, but low-energy analyses must reconstruct them from inconsistent data while enforcing S-matrix principles. The paper introduces SINNs to provide flexible, constrained amplitude representations and compatible-data selection for ππ scattering.

  • Motivation: Scattering amplitudes encode resonance properties, hadronic contributions to muon g −2, and final-state interactions relevant to CP-violation measurements.Accurate amplitudes and faithful uncertainty estimates are required for these derived quantities.
  • Problem: Low-energy amplitudes must be reconstructed from measurements because nonperturbative QCD has no analytic solution, while representations must respect unitarity, analyticity, and crossing symmetry.The inverse problem requires a flexible functional representation constrained by these principles.
  • Problem: Traditional parametrizations span limited regions of admissible amplitudes, making precision conditional on representation choice, experimental quality, and extrapolation prescriptions.These dependencies contribute to epistemic and stochastic uncertainties, especially for broad light scalars.
  • Approach: Neural representations remove an explicit fixed ansatz, allowing data to determine the amplitude shape while making architecture, selection criteria, and physics inputs reproducible analysis choices.The framework treats the analysis as a systematically testable learning problem.
  • Contribution: SINNs impose S-matrix constraints during optimization, select compatible experiments from an incompatible global set, and produce a direct low-energy ππ amplitude representation with uncertainty.The framework targets ππ scattering, whose measurements and global data set carry substantial model dependence and historical inconsistency.
  • S-matrix constraints: Analyticity permits continuation from real-axis information to the complex plane, where resonance poles and branch cuts impose global constraints on allowed amplitude representations.Ignoring this structure can produce large systematic errors in extracted resonance poles, especially for broad states or poles far from the physical region.
  • S-matrix constraints: Crossing symmetry relates the three s-channel isospin amplitudes to a single scalar amplitude evaluated in different kinematic channels.The relations connect s-, t-, and u-channel descriptions algebraically.
  • S-matrix constraints: In phenomenological analyses, analyticity and crossing are often restricted or neglected, whereas unitarity can be imposed directly on partial-wave parametrizations.This motivates incorporating multiple first-principles constraints into the learning framework.

B. Roy equations

Roy equations supply the dispersive implementation of analyticity, unitarity, and crossing symmetry used to constrain the learned ππ partial waves. Their subtraction terms, low-energy integrals, and high-energy driving term define a finite low-energy domain in which coupled-channel information regularizes the amplitude.

  • B. Roy equations: Roy equations provide the quantitative form in which analyticity, unitarity, and crossing symmetry enter the SINN learning problem.They derive from fixed-t dispersion relations and relate amplitudes to integrals over discontinuities across physical cuts.
  • B. Roy equations: Crossing symmetry fixes the subtraction polynomial, expressed for ππ scattering through the S-wave scattering lengths a0 0 and a2 0.These scattering lengths are both physical observables and subtraction constants.
  • B. Roy equations: The matching point sh separates the region directly constrained by partial-wave data from the high-energy region needed to complete the dispersion relation.The fixed-t integral is split at this matching point.
  • B. Roy equations: The Roy representation combines a subtraction term, low-energy partial-wave contributions, and a high-energy driving term supplied by fixed Regge input above √sh = 1.5 GeV.The low-energy partial waves are learned by SINN, while the asymptotic contribution completes the dispersion relation.
  • B. Roy equations: The analysis imposes Roy residuals on the real axis only for 2mπ < √s < 1.1 GeV, within the standard low-energy validity domain.Outside the finite applicability region, additional assumptions about the amplitude and high-energy behavior enter.
  • B. Roy equations: Roy kernels mix isospin and angular-momentum channels, so each reconstructed wave receives information from imaginary parts of multiple waves and the left-hand cut.Diagonal kernels require a principal-value prescription on the real axis.
  • B. Roy equations: GKPY equations are a related once-subtracted alternative that reduces subtraction-constant weight at higher energies and can produce smaller errors for some observables.Other dispersive equations can extend applicability to processes such as πK and πN scattering.

C. High-energy constraints

The framework combines high-energy Regge input with a SINN architecture that represents multiple partial waves while enforcing hard threshold and unitarity constraints and imposing analyticity and crossing numerically through Roy equations.

  • C. High-energy constraints: Above the matching energy, Regge theory replaces the finite partial-wave expansion because many partial waves contribute comparably and the truncated expansion fails to converge.The resulting high-energy amplitude is projected into the Roy driving terms.
  • C. High-energy constraints: Roy equations twice-subtract the high-energy contribution, strongly suppressing it in the low-energy training region, so the Regge input is kept fixed.This fixed input supplies the asymptotic contribution used in the dispersive constraints.
  • C. High-energy constraints: The production fit uses seven partial waves, with inelasticity outputs for S0, P1, S2, D0, and F1, while D2 and G2 remain elastic.The fitted dataset contains phase-shift data for all seven waves.
  • C. High-energy constraints: The SINN takes energy as input, uses a shared trunk and wave-specific heads, and applies postprocessing maps that enforce selected constraints exactly.The shared representation reflects correlations among partial waves, while specialized heads retain wave-specific behavior.
  • C. High-energy constraints: Threshold factors and inelasticity maps guarantee the required low-energy behavior and enforce 0 ≤ η ≤ 1, with elastic unitarity below each inelastic threshold.The inelastic-channel assignments are wave dependent, including K̄K, πω, and an effective S2 threshold.

B. Loss design

The loss combines experimental agreement with Roy-equation and threshold-consistency objectives, while a curriculum and auxiliary subtraction constants stabilize the multi-objective training.

  • B. Loss design: The total loss combines data, Roy-equation, scattering-length, and consistency terms: Ltotal = λdataLdata + λRoyLRoy + λSLLSL + λconLcon.These terms encode agreement with measurements and the required dispersive and threshold constraints.
  • B. Loss design: The data weight wχ interpolates between unnormalized least squares at wχ = 0 and standard χ2 at wχ = 1.This parameter controls how experimental uncertainties enter the data objective.
  • B. Loss design: Roy residuals are evaluated on 75 fixed-energy points for the low-energy dominant S0, P1, and S2 waves, while higher waves enter the dispersive integrals.The residuals are differences that vanish when the Roy equations are exactly satisfied.
  • B. Loss design: Threshold consistency requires auxiliary subtraction constants in the Roy equations to coincide with scattering lengths derived from the same SINN amplitude.The associated consistency loss stabilizes the dispersive representation at threshold.
  • B. Loss design: The DIRAC measurement supplies a soft constraint on the S-wave scattering-length difference, weighted as one effective data point through λSL = 1/Ndata.The asymmetric quoted uncertainty is used according to the sign of the residual.

C. Curriculum training

Training proceeds through a six-phase curriculum that gradually activates competing objectives, while constrained ensembles and spectral response clustering support uncertainty estimation and compatible-data selection.

  • C. Curriculum training: The training minimizes 17 competing objectives: 12 data terms, 3 Roy terms, and 2 threshold scattering-length objectives.Because the total loss is non-convex, the objectives cannot all be activated at the start.
  • C. Curriculum training: The first curriculum phase combines scattering-length and data losses with wχ = 0.25, allowing observables to learn their shapes before small experimental errors dominate.Later phases increase wχ to 1 and gradually activate Roy and consistency penalties.
  • C. Curriculum training: The final phase removes auxiliary subtraction constants and uses SINN-derived values directly in the Roy equations.The consistency loss is fixed to zero in this phase.
  • C. Curriculum training: UPGrad combines active objective gradients separately to project conflicting gradients toward a common descent direction.AdamW, initialization, clipping, decay, and patience-based transitions are also used during optimization.
  • C. Curriculum training: Gaussian replicas combine bootstrap data shifts with independently trained networks, and retained replicas share hard threshold and unitarity maps and a common loss cut.The resulting intervals are empirical distributions over constrained amplitudes and include data and initialization variation.
  • C. Curriculum training: Spectral response clustering uses constrained SINN ensembles to measure relative experiment compatibility and identify coherent data subpopulations.Two selection rounds were used to construct the production ensemble.

A. Spectral response clustering

Spectral response clustering uses constrained SINN ensembles to quantify experiment compatibility, embeds response patterns, clusters experiments, and iteratively removes groups with excessive tension.

  • Group testing and selection: +0.3 is the maximum surviving between-group response after selection.The final visualization distinguishes retained and rejected groups and marks tension and compatibility with red and blue edges, respectively.
  • Response construction: Each target experiment is upweighted in a SINN ensemble, and the resulting models are evaluated using uniform-weight χ2 on other experiments.This produces directed responses that are later symmetrized because upweighting one experiment and evaluating another need not give the transposed response.
  • Response construction: Positive symmetrized responses indicate tension, whereas negative responses indicate compatibility between experiments.The responses form the experiment response matrix used for subsequent embedding and clustering.
  • Embedding and clustering: The response matrix is embedded into a reduced Euclidean space using a truncated spectral decomposition, representing each experiment by leading spectral loadings.Ward linkage then hierarchically merges experiments according to distances between their retained embedding centroids.
  • Embedding and clustering: Ward clustering cuts the dendrogram at its largest merge-height gap, producing ten candidate experiment groups.The cut leaves N−m groups, with candidate group counts restricted to 2 ≤ ng ≤ N.
  • Group testing and selection: Groups are tested in the original response space using mean within-group and between-group responses, which retain the tension-versus-compatibility interpretation.Groups are removed iteratively until no positive surviving between-group average exceeds the cutoff C̄cut = 1.

B. Retained experimental input

The selection procedure removed incompatible experiments and retained a mutually compatible ππ dataset for constrained SINN training. The resulting amplitudes reproduce expected partial-wave behavior, satisfy dispersive tests, and yield scattering lengths consistent with established determinations.

  • Data selection: Two rounds of spectral response clustering selected the production input by removing incompatible experiment groups.The first round immediately removed the Protopopescu (VI, XIII) datasets because their S0 phase shifts disagreed strongly with the rest of the data above 1 GeV.
  • Data selection: 18 experiments and 621 measurements were retained over 0.28 ≤√s ≤1.5 GeV.Across the two rounds, median total loss decreased from 7.2 to 3.1, driven mainly by the data term falling from 6.9 to 2.9.
  • Partial-wave amplitudes: The retained data produced phase shifts and inelasticities with expected σ/f0(500), f0(980), ρ(770), and repulsive I = 2 features.Most selected data points lie within the central 68% uncertainty band, while D0 and F1 have wider bands because they rely on Hyams 73 data alone.
  • Partial-wave amplitudes: The S0 inelasticity depletion above the K ¯K threshold is milder because the selection removes deep-dip inputs and the SINN phase representation remains smooth there.The available data are not dense enough in that region to resolve an explicit derivative discontinuity.
  • Dispersive fulfillment: The SINN and Roy-reconstructed amplitudes agree within small residuals throughout the Roy validity range √s < 1.1 GeV.The comparison is nonlocal because reconstructed waves depend on coupled partial waves, subtraction constants, and high-energy Regge input.
  • Scattering lengths: The extracted scattering lengths agree within errors with chiral-Roy and Roy-equation determinations and lie inside the universal and DIRAC-allowed bands.They are determined by the amplitude at threshold without model parameters associated with the scattering lengths.

VII. ROBUSTNESS AND VALIDATION

Robustness studies tested architecture dependence, closure, and the role of Roy constraints. They support representative, statistically sound SINN uncertainties, while identifying the deliberately unmodeled K ¯K cusp as the closure-test failure.

  • Validation strategy: Closure tests, architecture comparisons, and removal of Roy constraints jointly supported representative sampling of physically allowed amplitudes.The ablation removing Roy equations left the data description untouched.
  • A. Architectural dependence: The architecture-ensemble scattering-length distributions were compared with the production ensemble across 10,000 and 1,000-network replicas.The figure encodes joint distributions, central 68% ellipses, and an added architecture-median spread for the production interval.
  • A. Architectural dependence: A 500-trial sweep varied 21 hyperparameter dimensions, including network geometry, optimizer settings, weight decay, and the Roy weight λRoy.The sweep was designed to test hidden model dependence from implicit neural regularization.
  • A. Architectural dependence: Alternative architecture ensembles produced statistically indistinguishable observable medians, with the largest shift equal to 0.42 δx.Because architecture variation did not shift central values materially, the authors neglect the associated systematic uncertainties.

B. Closure test

The closure test reconstructs known ππ amplitudes from pseudo-data and finds excellent agreement across the primary waves, with threshold observables recovered within central uncertainties. Ablation shows that Roy-equation constraints regularize threshold behavior without reducing data satisfaction, while the deliberate omission of the S0 cusp remains the main closure discrepancy.

  • Closure test: Pseudo-data generated from the CFD parametrization were passed through the full reconstruction pipeline to test bias in the production ensemble.The generating amplitude and published uncertainties were known, enabling direct comparison with the reconstructed ensemble.
  • Closure test: All three scattering-length combinations satisfy |z| < 0.4, with every target inside its asymmetric central 68% interval.Target values were recalculated from the same Roy-equation implementation used for the closure extraction.
  • Closure test: The reconstructed S0, P1, and S2 waves agree excellently with the CFD amplitude across nearly the full fitted range.Uncertainty bands widen where the pseudo-data inherit sparse or imprecise real-data coverage.
  • Ablation: Removing Roy and consistency penalties leaves data satisfaction nearly unchanged but substantially increases Roy residuals and shifts scattering lengths.a0 0 falls from 0.222 to 0.166, while a2 0 shifts from −0.041 to −0.097, outside the production intervals and universal acceptable band.
  • Implications: The workflow’s scope includes constrained uncertainty bands, reusable line shapes, and possible extensions to other scattering processes and off-real-axis observables.The authors identify πK and πN scattering as direct next tests and provide production replicas, medians, intervals, and an example notebook.
  • Limitations: The only closure failure is the S0 phase near the K ¯K threshold, where the CFD amplitude has a cusp omitted by the production representation.The production representation retains the K ¯K inelasticity opening but does not impose the phase cusp.

Appendix A: Numerical Roy implementation

The numerical Roy implementation evaluates constrained partial-wave residuals using principal-value treatment on the real axis and ordinary quadrature after continuation off the real axis. Fixed Regge input supplies the high-energy completion and is strongly suppressed at low energy by the twice-subtracted kernels.

  • Real-axis implementation: Roy residuals are evaluated for the constrained S0, P1, and S2 channels on the numerical mesh.The higher-wave set H = {D0, D2, F1, G2} enters the dispersive integrals as auxiliary input.
  • High-energy completion: The fixed-t dispersion relation splits at √sh = 1.5 GeV, projecting the low-energy part into Roy kernels and evaluating the high-energy part with fixed Regge input.Only the s′ > sh contribution uses the fixed Regge amplitudes in the driving term.
  • Principal-value treatment: On the real axis, the diagonal-kernel pole at s′ = s is handled by subtracting the numerator at the pole.The regular term is integrated with Gauss-Legendre quadrature, while the singular contribution is evaluated analytically.
  • Complex continuation: After continuation off the real axis, the displaced pole permits the same quadrature without principal-value subtraction.The change follows because the pole no longer lies on the integration contour.
  • High-energy completion: Twice-subtracted Roy kernels strongly suppress the fixed high-energy contribution at low energy, leaving a smooth asymptotic completion.The Regge parameters and t-dependence are held fixed, with only imaginary parts entering the dispersive integrals.

Appendix C: Spectral clustering and group rejection

Spectral response clustering embeds experiment-response patterns into a low-dimensional eigenvector space, forms Ward groups, and iteratively removes groups with the strongest incompatibilities. The second round retains 16 experiments with a maximum surviving between-group average response of +0.3.

  • Spectral embedding: The second selection round begins with 21 experiments and retains k = 11 eigenmodes capturing 90% of the cumulative absolute spectral weight.The response matrix is ordered by Ward group and used for the spectral decomposition.
  • Ward clustering: Experiments with similar leading-eigenvector loading patterns are merged early in the Ward hierarchy, while distinct patterns remain separate.The loadings identify response-pattern similarity before the final pruning step.
  • Ward clustering: Ward clustering yields ng = 10 candidate groups, including seven singleton groups and three multi-member clusters.The dendrogram is cut at its largest merge-height gap.
  • Group rejection: At each removal iteration, the group with the most partner violations above +1 is removed, with ties resolved by the largest violation and then smaller group size.The Ward partition is not recomputed after removals.
  • Group rejection: After five removal steps, 16 clustered experiments remain across G3, G4, G5, G6, and G7, with maximum surviving between-group average response +0.3.The rejected groups include cases with maxima of +32.2, +69.9, +15.0, +12.8, and +3.5 against other groups.

Appendix D: Auxiliary partial waves

The auxiliary partial waves complete the Roy-system input while the primary physics targets remain S0, P1, and S2. Training diagnostics show convergent retained losses, but the nominally overparameterized network relies on imposed physics and closure validation to constrain its effective function space.

  • Auxiliary waves: D0, F1, D2, and G2 provide non-negligible input to the Roy system up to √sh, although they are auxiliary rather than primary physics targets.The primary targets are the S0, P1, and S2 channels.
  • Training diagnostics: All production loss components converge across the ensemble with approximately Gaussian distributions.The production ensemble contains 10,000 replicas satisfying the common total-loss cut.
  • Training diagnostics: When Roy and scattering-length terms activate, their losses briefly spike before relaxing within the new curriculum phases.The retained loss components stabilize during the final phase of the representative run.
  • Overparameterization: The production network has 1,156 trainable weights for 621 retained data points, making the fit nominally overparameterized.A weight distribution peaked near zero is consistent with active weight decay but is not by itself evidence against overfitting.
  • Overparameterization: Threshold factors, the inelasticity map, and coupled Roy residuals restrict admissible functions beyond the raw parameter count.Closure recovery of threshold observables and replica spread provide the stated checks on remaining freedom.
  • Auxiliary-wave limitations: The F1 partial is underspecified above threshold because data coverage is poor and Roy equations do not control it there.Its large 95% uncertainty-band fluctuation is attributed to truncation of the dispersive integral.

Appendix G: Architecture sweep protocol

The appendix describes a 500-trial sweep over 21 architecture, physics, optimizer, and constraint-weight dimensions, followed by closure-test validation of reconstructed amplitudes and derived quantities.

  • Architecture sweep: 500 trials explored 21 hyperparameter dimensions spanning trunk geometry, per-channel heads, physics switches, and optimizer or constraint weights.The sweep used Ray Tune; learning rate, weight decay, and final-phase Roy weight ranges were specified for the optimizer and constraint group.
  • Architecture sweep: 14 of the 21 sweep dimensions controlled independent per-channel head structure, nonlinearity, and capacity.Total parameter counts ranged from 414 to 4,266, with a median of 1,649, compared with 1,156 for the production architecture.
  • Architecture sweep: All tested activations were smooth because piecewise-linear activations can interfere with partial-wave smoothness and spoil fulfillment of the Roy equations.The stated rationale connects smoothness requirements to analyticity and causality.
  • Closure test: The closure test trained 1,000 retained replicas on independent Gaussian fluctuations around CFD central values, retaining replicas with Ltotal ≤4.5.Six channels used CFD central values, while the absent G2 wave used zero phase shifts and unit inelasticities.
  • Closure test: Reconstructed D0, D2, and F1 waves agreed with the generating CFD amplitude, with uncertainty bands widening where pseudo-data had sparse or imprecise coverage.The reconstructed a0^0–a2^0 correlation also reproduced the generating amplitude’s correlation, and the target lay near the center of the 68% ellipse.
Loading 2608.23750v1…