Source-linked AI summary
Constraining Effective Field Theories with Machine Learning
Johann Brehmer, Kyle Cranmer, Gilles Louppe, Juan Pavez
TL;DR
The paper addresses the challenge of constraining EFT parameters from high-dimensional LHC data when likelihoods are intractable and existing methods do not scale well. It extracts additional simulator information to train neural networks for likelihood-ratio estimation, achieving stronger dimension-six operator constraints while retaining scalability and fast evaluation.
Problem
High-dimensional LHC processes make likelihood-based EFT constraints difficult because simulations involve intractable integrals, while existing methods do not scale well to many parameters and observables.
Method
The methods use simulator-derived information, including scores and joint likelihood ratios, to train neural networks that estimate likelihood ratios without approximating shower or detector effects.
Results
The Rascal technique produces significantly stronger constraints on two dimension-six operators, with expected limits virtually indistinguishable from the theoretical optimum.
Takeaways & Limitations
The techniques scale to many observables and high-dimensional parameter spaces, evaluate likelihood ratios in microseconds, and could improve LHC legacy measurements.
Takeaways & Limitations
Sally performs best near the Standard Model, while its local approximation can yield weaker bounds farther away.
Abstract
from arXiv · showhide
We present powerful new analysis techniques to constrain effective field theories at the LHC. By leveraging the structure of particle physics processes, we extract extra information from Monte-Carlo simulations, which can be used to train neural network models that estimate the likelihood ratio. These methods scale well to processes with many observables and theory parameters, do not require any approximations of the parton shower or detector response, and can be evaluated in microseconds. We show that they allow us to put significantly stronger bounds on dimension-six operators than existing methods, demonstrating their potential to improve the precision of the LHC legacy constraints.
INTRODUCTION
LHC legacy constraints on dimension-six SMEFT operators must resolve subtle signatures across high-dimensional phase spaces. The paper addresses the limitations of hand-picked variables and existing likelihood-based methods with machine-learning estimators built from information extracted from Monte-Carlo simulations.
- Dimension-six SMEFT constraints probe subtle kinematic signatures involving many EFT coefficients in high-dimensional phase spaces.
- Hand-picked kinematic variables discard information and can constrain some parameter directions precisely while leaving other directions weakly constrained.
- The fully differential cross section improves multi-parameter sensitivity, but Matrix Element and Optimal Observables methods approximate or neglect shower and detector effects.
- The proposed techniques extract likelihood dependence on theory parameters from Monte-Carlo simulations and use augmented data to train neural networks estimating likelihood ratios.
- The methods are demonstrated on weak-boson-fusion Higgs production in the four-lepton decay mode.
Learning likelihood ratios
The paper exploits tractable parton-level information in particle-physics simulators to learn likelihood ratios despite intractable shower and detector integrations. Rascal combines joint likelihood-ratio and score information in neural networks, enabling rapid event-level evaluation.
- Simulator structure: The likelihood factorizes into a theory-dependent parton-level process followed by parton-shower, detector, and reconstruction effects.The parton-level density p(z|θ) is tractable, while the later conditional densities map z to observables x.
- Simulator structure: Millions of simulator random variables make direct integration over shower and detector degrees of freedom infeasible, rendering the likelihood ratio intractable.Monte-Carlo simulators sample the sequential distributions using their Markov structure.
- Simulator structure: Accessible parton-level momenta allow simulations to provide joint likelihood ratios and joint scores in addition to generated observables.The parton-level density can be evaluated for arbitrary momenta and theory parameters using matrix-element and parton-density-function calculations.
- Learning likelihood ratios: Although these joint quantities involve unobserved parton-level momenta, they define functionals whose extrema recover the observable-level likelihood ratio.This connection avoids requiring direct evaluation of the intractable components in the likelihood ratio.
- Learning likelihood ratios: Rascal trains a deep neural network with loss functions based on joint likelihood-ratio and score information, replacing expensive numerical integrals with regression and microsecond evaluations.The network is optimized by stochastic gradient descent, and the approach is described as a machine-learning version of the Matrix Element Method.
Local approximation
The local approximation uses score-based neural estimators to summarize likelihood information near the Standard Model. It supports compact inference in high-dimensional parameter spaces, but can lose sensitivity farther away.
- Local model: The local model treats score functions as sufficient statistics that retain all parameter information near the Standard Model.This makes the estimated score a machine-learning analogue of Optimal Observables.
- Sally: The Sally method estimates the score with a neural network, forming the basis for local likelihood-ratio estimation.The score estimator uses joint score information available from the simulator.
- Sallino: Sallino compresses high-dimensional observations to one scalar without losing sensitivity between two parameter points under the local approximation.The scalar product captures the discrimination power between θ0 and θ1, even with hundreds of theory parameters.
- Limitations: Sally and Sallino work very well close to the Standard Model, while the local approximation may deteriorate farther away.Approximation error reduces sensitivity and weakens bounds rather than producing overly optimistic results.
EXAMPLE PROCESS
The methods are demonstrated on weak-boson-fusion Higgs production in the four-lepton mode, using an idealized parton-momentum setup and expected constraints on two operators. Rascal closely matches the theoretical optimum and outperforms the histogram analysis in the studied region.
- EXAMPLE PROCESS: The example studies SMEFT constraints from weak-boson-fusion Higgs production in the four-lepton mode, focusing on two sensitive operators.The event samples use MadGraph 5 and MadMax, with exactly measured parton momenta in the idealized setup.
- EXAMPLE PROCESS: 16% larger reach in the new physics scale or 90% more collected data is achieved by Rascal than by the histogram analysis in one parameter region.The comparison uses the hardest-jet transverse momentum and the azimuthal angle between the two jets.
- EXAMPLE PROCESS: Rascal limits are virtually indistinguishable from the true likelihood contours after 36 observed events.Sally and Sallino produce nearly optimal bounds close to the Standard Model.
- EXAMPLE PROCESS: All new techniques impose significantly tighter parameter bounds than the doubly differential histogram analysis.At the 95% CL level, slightly weaker constraints indicate breakdown of the local model approximation.
CONCLUSIONS
The paper develops machine-learning techniques that use additional Monte-Carlo information to estimate likelihood ratios for effective-field-theory constraints. Rascal and Sally scale to high-dimensional LHC analyses, with Rascal producing substantially stronger operator bounds and Sally remaining effective near the Standard Model.
- The techniques extract additional information from Monte-Carlo simulations to train neural networks that estimate arbitrary likelihood ratios for limit setting.
- Rascal produces significantly stronger constraints on two dimension-six operators in weak-boson-fusion Higgs production, with expected exclusion limits virtually indistinguishable from the theoretical optimum.
- Sally compresses observations into a low-dimensional score without loss of sensitivity near the Standard Model and performs very well in that region.
- Sally yields only slightly weaker constraints farther from the Standard Model.
- Both approaches scale to many observables and high-dimensional parameter spaces, require no approximations of hard-process, shower, or detector effects, and evaluate likelihood ratios in microseconds.