Source-linked AI summary

Scaling laws for amplitude surrogates

Henning Bahl, Victor Bresó-Pla, Anja Butter, Joaquín Iturriza Ramirez

arXiv:2601.13308v1hep-phcs.LG

TL;DR

The paper asks how much data, compute, and network capacity amplitude surrogates need to reach target precision. It studies their scaling laws across particle-physics processes and finds power-law behavior whose coefficients are empirically related to the number of external particles, supporting systematic surrogate training.

  • Problem

    Amplitude surrogates require predictable data, compute, and network-size budgets to reach desired precision, but interpolation quality can stop improving beyond a threshold.

  • Method

    The paper systematically studies scaling with training-set size, compute, and network size across amplitude-surrogate architectures, losses, and processes.

  • Results

    Power-law scaling appears across the studied processes, and the scaling coefficients show empirical evidence of a relation to the number of external particles.

  • Takeaways & Limitations

    Scaling laws provide a practical recipe for systematic and predictable surrogate training toward predefined precision targets.

  • Takeaways & Limitations

    The cross-process conclusions are constrained by the studied training ranges, and two-particle final states are outliers associated with singular amplitude structure.

Abstract

from arXiv · show

Scaling laws describing the dependence of neural network performance on the amount of training data, the spent compute, and the network size have emerged across a huge variety of machine learning task and datasets. In this work, we systematically investigate these scaling laws in the context of amplitude surrogates for particle physics. We show that the scaling coefficients are connected to the number of external particles of the process. Our results demonstrate that scaling laws are a useful tool to achieve desired precision targets.

2 Scaling laws in deep learning

The paper formulates neural-network performance as power laws in network size, training-set size, and compute, linking their scaling to intrinsic dimension and process degrees of freedom. These laws provide a way to predict precision targets while clarifying assumptions and saturation limits.

  • Scaling-law formulation: Power-law scaling in N, Dtrain, and C can predict the training budget needed for a target performance and characterize saturation.The generic law decreases with the scaled variable until reaching a plateau determined by the other variables.
  • Network-size scaling: Under piecewise-linear approximation and Lipschitz continuity, network-size scaling follows from the number of linear pieces, with αN ≃ 4/d for generic targets.For simpler targets, the exponent estimate can become a lower bound, αN ≳ 4/d.
  • Dataset-size scaling: Training-set scaling assumes noise-free, sufficiently large networks and uniformly distributed samples on a d-dimensional manifold, but the resulting estimate can underestimate the observed exponent.The paper notes that nonlinear effects near training points and non-i.i.d. datasets complicate the estimate.
  • Compute scaling: Compute scaling is modeled through network size and dataset size, with a threshold Cthres above which additional compute no longer reduces the loss.The compute estimate is rough because the ratio of N to Dtrain and batch size can affect performance.
  • Scaling-law formulation: The intrinsic dimension of amplitude data manifolds matches the number of interaction-amplitude degrees of freedom, making process-dependent scaling coefficients predictable.The paper defines intrinsic dimension through the minimum variables needed to parameterize a local manifold environment and plans numerical checks of this relation.
  • Representation dimension: The learned representation dimension provides a potential diagnostic of prediction quality and matches the number of degrees of freedom in the studied amplitude datasets.The paper uses the twoNN method to estimate the dimension of the final hidden-layer representations.

3 Computational setup

The computational setup defines amplitude surrogates, compares Lorentz-invariant MLP-I and permutation-equivariant LLoCa-Transformer architectures, and scans data, network, and compute resources under a practical training-time limit. It also evaluates heteroscedastic uncertainty modeling alongside standard losses and uses maximal-update parametrization to transfer learning rates across widths.

  • Computational setup: Amplitude surrogates predict squared interaction amplitudes from phase-space points, with scaling studied across network size, training-set size, and compute.The setup specifies both architecture and training strategy for extracting these scaling laws.
  • Architectures: The MLP-I uses momentum-invariant inputs, while the LLoCa-Transformer uses learned local Lorentz frames and additionally handles particle permutations equivariantly.Both architectures are Lorentz-invariant by construction; permutation equivariance is especially useful for identical external particles.
  • Training ranges: The study limits training to at most 3 days on an NVIDIA H100 GPU, bounding the explored network-size and compute ranges.Scaling scans vary one factor while selecting representative values for a secondary factor and fixing the third at its budget-limited maximum.
  • Uncertainty modeling: Heteroscedastic regression predicts both A(x) and σ(x), penalizing inflated uncertainty while recovering MSE when σ is constant.The learned σ(x) represents systematic uncertainty from data noise and limited surrogate expressivity, not statistical uncertainty from finite training data.
  • Uncertainty modeling: Systematic-uncertainty calibration is tested with a pull distribution that should be unit Gaussian when uncertainties are calibrated and statistical uncertainty is negligible.This test also probes the validity of the Gaussian likelihood ansatz.
  • Maximal Update Parametrization: Under µP, the optimal learning rate remains constant across hidden dimensions, enabling tuning on a small network and zero-shot transfer to larger networks.The study does not use this transfer strategy for regularization parameters.

4 Case study: q ¯q →t ¯t H

The q¯q →t¯tH case study finds power-law scaling with compute, dataset size, and network size, while identifying dataset size as a key source of performance saturation. The extracted intrinsic dimension matches the five degrees of freedom of the 2 →3 process, and heteroscedastic loss preserves the observed scaling behavior.

  • Scaling using MSE loss: Increasing compute reduces test loss logarithmically, but small datasets produce visible plateaus that limit performance.For the largest dataset, loss continues decreasing beyond the compute budget used in the earlier study.
  • Scaling using MSE loss: Test loss follows power-law scaling with dataset size, with saturation mainly at NMLP = 10^2 and only slight gains from larger networks.The training protocol uses 10^4 epochs to avoid reducing the number of epochs as dataset size grows.
  • Scaling using MSE loss: Network-size scaling also follows a power law, while larger datasets improve performance even for small networks and prevent dataset-limited saturation.
  • The intrinsic dimension is five for q¯q →t¯tH, and all fitted scaling exponents are compatible with the predicted lower bound αX = 0.8.Networks recover this intrinsic dimension for NMLP ≥ 10^4 and Dtrain ≳ 10^2.
  • Scaling using heteroscedastic loss: Heteroscedastic loss yields the same compute and dataset-size scaling behavior as ordinary MSE loss.
  • Scaling using heteroscedastic loss: Learned uncertainties shift lower with more compute or training data, while their bulk is well calibrated and only the tails show slight overestimation.Calibration improves with additional compute, especially through larger training datasets.
  • Scaling using LLoCa-Transformer: The LLoCa-Transformer provides no significant advantage over MLP-I and has less predictable scaling because it is harder to optimize.The authors attribute its irregular scaling partly to suboptimal hyperparameters for medium and large datasets.

5 Scaling laws across processes

Across processes, amplitude-surrogate losses exhibit similar power-law scaling, with dataset size the main driver and scaling exponents tied mainly to final-state particle multiplicity. Architecture affects scaling and performance, while two-particle final states are notable outliers.

  • Scaling behavior: Dataset size is the main performance driver, compute scaling is constrained by limited data, and increasing network size beyond 10^4 parameters provides no benefit.These observations hold across the studied processes using the MLP-I architecture.
  • Architecture comparison: LLoCa-Transformer matches MLP-I's α_D exponent on q¯q →Z +4g while reaching improved performance at smaller datasets and maintaining flatter gains across D_train.The differing behavior is associated with the process's complexity and permutation-invariance structure, but requires further study.
  • Scaling exponents: The fitted dataset-size and compute exponents agree with theoretical bounds and depend mainly on final-state particle multiplicity, not particle species or symmetry patterns.The intrinsic dimension is related to final-state multiplicity by d = 3n_f − 4.
  • Outliers: Processes with two final-state particles are outliers, plausibly because forward/backward scattering singularities place more samples in low-precision regions.Those regions can violate the assumed negligible-error linear dependence used to derive the scaling relations.
  • Resource prediction: A low-cost surrogate can determine scaling-law parameters and provide a conservative estimate of resources needed to reach a target precision.The procedure assumes numerical noise in evaluating the true amplitude is negligible.

6 Conclusions

The paper finds power-law scaling across architectures, objectives, and eleven particle-physics processes, with coefficients related to external-particle count. This supports predictable resource planning for surrogate precision targets and reveals that learned representations estimate intrinsic dimension.

  • Conclusions: For q¯q →t¯tH, dataset size is the strongest resource factor, followed by compute, while network size has only a minor role within the studied ranges.The authors caution that these conclusions are strictly valid within the investigated training-variable ranges.
  • Conclusions: Power-law scaling appears across the q¯q →t¯tH study and ten additional processes, while permutation-equivariant architectures improve performance without significantly changing scaling coefficients.The broader processes vary in external-particle number, particle types, and interactions.
  • Conclusions: The number of external particles predicts scaling coefficients and enables conservative resource estimates for any desired surrogate precision target.A low-cost surrogate supplies an anchor point for the scaling curves.
  • Conclusions: The learned representation dimension closely approximates the processes' degrees of freedom, allowing amplitude surrogates to estimate the intrinsic dimension of the input manifold.The paper presents this as a by-product and notes its relevance to data spaces with strong symmetry structure.
  • Conclusions: Universal power-law scaling and the precision-targeting recipe support systematic surrogate training for reliable next-generation event generators.This is presented as the practical significance of the observed scaling behavior.

A.1 Scaling with other precision metrics

Scaling laws persist when precision is measured with L1-type metrics rather than mean-squared error. The corresponding slopes are reduced consistently with the loss's dependence on the underlying error.

  • A.1 Scaling with other precision metrics: L1 test error exhibits analogous scaling laws, with curve slopes roughly half those of MSE scaling.This agrees with the relation that L1 loss scales linearly with error, whereas MSE scales quadratically.
  • A.1 Scaling with other precision metrics: The ε metric used in Ref. [23] is equivalent to an L1 loss apart from normalization and therefore has the same loss-curve slope.The curves reach saturation above D_train ≲ 10^4 in the cited comparison.

A.2 Jet-associated Z production

For jet-associated Z production, test loss follows scaling with training dataset size, while network size is a minor limitation in the studied regime. The higher-dimensional Z g g g g process is less precise and has a shallower loss slope than Z g.

  • A.2 Jet-associated Z production: The training dataset size is a major limiting factor for both Z g and Z g g g g surrogates, whereas NN size has only a minor role.
  • A.2 Jet-associated Z production: The Z g g g g surrogate is less precise than Z g, with a smaller loss-curve slope due to its higher-dimensional phase space.
  • A.2 Jet-associated Z production: Figure 13 measures MSE test loss against spent training compute for three training-dataset sizes in both processes.
  • A.2 Jet-associated Z production: Compute-scaling results again confirm that training dataset size is a major limiting factor.

A.3 Jet associated di-photon production

Jet-associated di-photon surrogates exhibit the same scaling pattern: training dataset size is important, while network size is comparatively minor. Compute-scaling curves confirm this behavior for both studied amplitudes.

  • A.3 Jet associated di-photon production: Figure 14 compares MSE test loss versus training dataset size for g g →γγ + g and g g →γγ + g g across different NN sizes.
  • A.3 Jet associated di-photon production: Figure 15 compares MSE test loss versus training iterations for the two di-photon amplitudes at three training-dataset sizes.
  • A.3 Jet associated di-photon production: Training dataset size is a major limiting factor for both di-photon surrogates, while NN size plays only a minor role.
  • A.3 Jet associated di-photon production: Compute-scaling results confirm the dataset-size limitation and the minor role of NN size.

A.4 Electroweak multi-boson production

Electroweak multi-boson surrogates show that training dataset size is an important limitation, whereas NN size is minor in the considered regime. Compute-scaling curves support this pattern, with one small increase attributed to learning-rate optimization.

  • A.4 Electroweak multi-boson production: Figure 16 shows MSE test errors versus training dataset size for q¯q →W Z, q¯q →WW Z, q¯q →W Z g, and q¯q →W Z g g at different NN sizes.
  • A.4 Electroweak multi-boson production: Figure 17 shows compute-scaling curves for the four electroweak multi-boson processes and confirms the dataset-size limitation.
  • A.4 Electroweak multi-boson production: Training dataset size is an important limiting factor for the electroweak multi-boson surrogates, while NN size plays only a minor role.
  • A.4 Electroweak multi-boson production: A slight final-point test-loss increase on one curve results from an insufficiently fine learning-rate scan and would be removed by better optimization.

B.1 Method

The method estimates intrinsic dimension from nearest-neighbor distances in hidden-layer activations. It uses shell-volume ratios and a likelihood maximization based on the resulting distribution.

  • B.1 Method: The twoNN method estimates the intrinsic dimension of the network representation.
  • B.1 Method: The method constructs k-nearest-neighbor lists and computes hyperspherical-shell volumes between successive neighbors.
  • B.1 Method: Assuming locally constant point density, shell volumes follow an exponential distribution, enabling a distribution for ratios of shell volumes.
  • B.1 Method: The resulting ratio variable µ follows a Pareto distribution, whose joint likelihood is maximized to infer the intrinsic dimension.
  • B.1 Method: For each activation point, nearest and second-nearest neighbor distances are used to approximate the ratio distribution and estimate d by maximizing P(µ|d).
  • B.1 Method: The activations used are taken from the last hidden layers over a subset containing at least 10^4 phase-space points.

B.2 Results

The twoNN method recovers the theoretically expected intrinsic dimensions across Z-plus-gluon, di-photon, and electroweak multi-boson surrogates, with larger networks and datasets improving agreement.

  • Z-plus-gluon surrogates reasonably recover the true phase-space dimensions when network and training dataset sizes are sufficiently large.Deviations are sizeable only for the Z g g g and Z g g g g surrogates, and remain small overall.
  • Di-photon surrogates show good overall agreement with theoretical intrinsic dimensions, although the higher-dimensional g g →γγg g process has greater difficulty reaching its minimal dimension.
  • Electroweak multi-boson surrogates exhibit excellent convergence toward their theoretically expected intrinsic dimensions.
  • The twoNN method reliably recovers the number of degrees of freedom across a wide range of processes, including when trained on 4-momenta rather than momentum invariants.This indicates that full knowledge of process symmetries is not required for intrinsic-dimension estimation.

C Power law fit results

The paper reports fitted power-law coefficients for all studied processes and model variants, with fit uncertainties intended to capture rerun variation but not gains from further hyperparameter optimization.

  • Fitted power-law coefficients are provided for all processes and models studied, covering MLP-I, LLoCa-Transformer, and multiple loss settings.The reported tables include t¯tH, Z-plus-gluon, di-photon, and electroweak multi-boson surrogates.
  • The coefficient uncertainties assume a 10% relative uncertainty on test losses and are intended to represent variation from rerunning identical networks and hyperparameters.
  • The fitted uncertainties do not include improvements that might result from further hyperparameter optimization.
Loading 2601.13308v1…