Source-linked AI summary
A Guide to Constraining Effective Field Theories with Machine Learning
Johann Brehmer, Kyle Cranmer, Gilles Louppe, Juan Pavez
TL;DR
Collider EFT inference is hindered by intractable likelihood integrals over latent particle-physics variables. The paper develops machine-learning estimators that exploit simulator structure, finding precise, scalable likelihood-ratio methods that yield tighter EFT constraints than traditional kinematic analyses.
Problem
Intractable integrals over latent spaces prevent direct evaluation of likelihoods and likelihood ratios needed to constrain EFT parameters.
Method
The paper develops machine-learning likelihood-ratio estimators trained with additional information from particle-physics simulators, including the score and latent-space structure.
Results
In weak-boson-fusion Higgs production, the new methods provide precise likelihood-ratio estimates and tighter EFT-coefficient constraints than traditional kinematic histograms.
Takeaways & Limitations
The methods scale to many observables and high-dimensional parameter spaces, require no approximations of hard-process, shower, or detector effects, and evaluate likelihood ratios in microseconds.
Abstract
from arXiv · showhide
We develop, discuss, and compare several inference techniques to constrain theory parameters in collider experiments. By harnessing the latent-space structure of particle physics processes, we extract extra information from the simulator. This augmented data can be used to train neural networks that precisely estimate the likelihood ratio. The new methods scale well to many observables and high-dimensional parameter spaces, do not require any approximations of the parton shower and detector response, and can be evaluated in microseconds. Using weak-boson-fusion Higgs production as an example process, we compare the performance of several techniques. The best results are found for likelihood ratio estimators trained with extra information about the score, the gradient of the log likelihood function with respect to the theory parameters. The score also provides sufficient statistics that contain all the information needed for inference in the neighborhood of the Standard Model. These methods enable us to put significantly stronger bounds on effective dimension-six operators than the traditional approach based on histograms. They also outperform generic machine learning methods that do not make use of the particle physics structure, demonstrating their potential to substantially improve the new physics reach of the LHC legacy results.
I. INTRODUCTION
Collider measurements of SMEFT parameters face high-dimensional kinematic structure, intractable likelihoods, and limitations of traditional low-dimensional analyses. The paper develops simulation-based machine-learning methods that exploit particle-physics structure to estimate likelihood ratios efficiently and improve EFT constraints.
- Motivation: SMEFT measurements may involve tens of parameters with subtle signatures in high-dimensional data, challenging traditional analysis techniques.Dimension-six operators provide a parameterization of indirect physics beyond the Standard Model, but the relevant constraints are difficult to extract from complex collider observables.
- Traditional methods: Traditional analyses restrict data to a few discriminating variables, often weakening constraints in other parameter-space directions.Fully differential information can improve sensitivity, but existing matrix-element-based approaches require approximations to shower or detector effects.
- Contribution: The paper develops particle-physics-tailored methods that extract additional simulator information to train neural networks estimating likelihood ratios.The methods target many observables and high-dimensional parameter spaces while supporting parton showers, backgrounds, and full detector simulations.
- Evaluation: The methods are evaluated on two dimension-six operators in weak-boson-fusion Higgs production with a four-lepton decay at the LHC.Part of the study uses an idealized setting where the true likelihood is available as ground truth for comparing methods.
- Likelihood-free inference: Generic likelihood-free machine-learning methods use simulated samples but do not exploit the structure of particle-physics processes.This limits their use of the information available from the event-generating simulator.
- Statistical challenge: The EFT measurement problem includes intractable integrations over latent shower and detector variables, preventing direct evaluation of the observable likelihood.These integrals arise because reconstruction-level observables are generated through high-dimensional simulator stages.
2. Operator morphing
Operator morphing exploits the finite amplitude structure of EFT predictions to reconstruct parameter-dependent distributions from basis evaluations. The paper applies this framework to weak-boson-fusion Higgs production, where interference and multivariate kinematics carry information beyond standard one- and two-dimensional observables.
- Operator morphing: EFT likelihoods can be decomposed into finite amplitude components with parameter-dependent weights and momentum-dependent basis functions.The component functions need not individually be positive definite or normalized.
- Operator morphing: Choosing as many invertible basis points as amplitude components rewrites the EFT prediction as a mixture of basis densities.The analytically known weights then allow the full parton-level likelihood to be extracted from finite basis evaluations.
- Operator morphing: Morphing can also estimate conditional distributions when exact shower- and detector-level likelihoods are intractable.For example, histograms or other density estimators for individual observables can be combined through the morphing structure.
- Example process: The weak-boson-fusion Higgs-to-four-leptons example probes dimension-six Higgs-gauge operators through precisely reconstructed jets and leptons.The final-state momenta span a 16-dimensional phase space, providing many potentially informative observables.
- Example process: Leading-jet transverse momentum and dijet azimuthal separation probe different parameter directions, while interference creates non-trivial kinematic features.The leading-jet momentum correlates with momentum transfer, and an amplitude driven through zero can produce a depleted distribution region near 350 GeV.
- Information content: The full multivariate distribution contains significantly more information than one- and two-dimensional marginal distributions of standard kinematic variables.This motivates retaining correlations rather than relying only on traditional jet observables.
2. Sample generation
Sample generation exploits a morphing structure in which parameter-dependent weights combine phase-space-dependent components to construct likelihoods. The resulting estimator strategies range from point-by-point scans to parameterized models and score-based local compression, with numerical stability remaining a concern.
- Sample generation: The process likelihood approximately follows a mixture model because its amplitude factorizes into parameter-dependent factors and phase-space-dependent amplitudes.The small total-width effect from OW and OWW is treated as irrelevant at experimental resolution, allowing the morphing weights wc(θ) to be calculated.
- Sample generation: 5.5·10^6 parton-level events were generated, with likelihoods evaluated at 15 basis parameter points before calculating morphing weights.These calculations provide the true parton-level likelihood function for each generated phase-space point.
- Sample generation: |wc| ≲ O(100) in some parameter-space regions, with large cancellations between positive and negative weights that challenge numerical stability.Other basis choices produced comparable or larger morphing weights.
- Inference strategies: Point-by-point estimators scan parameter values separately, optionally using a fixed or composite reference hypothesis to improve numerical stability.Only final results are interpolated between scanned θ0 values, while alternative denominators can be obtained from ratios of estimated ratios.
- Inference strategies: Point-by-point scans discard parameter-space smoothness, become expensive in high dimensions, and may introduce interpolation uncertainties.Parameterized estimators instead learn the dependence on observables and theory parameters directly, avoiding final interpolation and assumptions about the likelihood-ratio form.
- Score and local model: The score provides a local likelihood equivalent and sufficient statistics for the local model, reducing high-dimensional observables without information loss within that approximation.Locally, the likelihood ratio depends only on the scalar product of the score with θ0 − θ1, permitting scalar compression.
B. Available information and its usefulness
The simulator supplies ordinary events plus particle-physics-specific latent information, including joint likelihood ratios and scores. Regression and classification methods can use these quantities to estimate observable-level likelihood ratios while retaining exact shower and detector treatment.
- Available information: Generic likelihood-free inference uses simulated observables, whereas particle-physics simulators additionally expose latent parton-level momenta and process structure.This extra structure enables access to joint likelihood ratios and scores even when the observable-level likelihood is intractable.
- Likelihood-ratio estimation: Classifier decision functions can be transformed into likelihood-ratio estimators, with calibration helping when the classifier learns only a monotonic function of the ratio.The likelihood-ratio trick is illustrated by the Gaussian toy example.
- Particle-physics structure: The score is the tangent vector measuring relative likelihood change under infinitesimal parameter variations, and its tractable joint form can be extracted from the simulator.Intractable shower, detector, and reconstruction factors cancel in the joint construction.
- Particle-physics structure: The methods do not require simplifying approximations and remain applicable with reducible backgrounds, higher-order matrix elements, matched parton showers, and full detector simulations.These claims concern the available simulator information and the validity of the joint constructions.
- Likelihood-ratio estimation: Regression on joint likelihood ratios recovers the observable-level likelihood ratio when events are sampled according to the denominator hypothesis.The same conditional-regression principle applies to scores when events are sampled according to the corresponding parameter value.
- Available information: Table II contrasts quantities available in general likelihood-free inference with those available from particle-physics structure, including quantities regressable from joint information.Asterisks mark quantities that are not immediately available but can be regressed from corresponding joint quantities.
C. Strategies
The paper combines estimator structures with simulator-accessible information into several likelihood-ratio strategies, from histogram and approximate-computation methods to calibrated classifiers and neural density estimators. These approaches differ in how much parameter dependence and particle-physics structure they exploit.
- Strategy design: The strategies are organized by combining estimator dependence on θ with the samples, likelihood ratios, and scores available from the generator.The overview introduces the main ideas, while technical details are provided separately.
- Histogram methods: Histogram estimators approximate likelihoods from generated samples of manually chosen low-dimensional kinematic variables.The implementation scans θ0 point by point with a fixed reference and uses variable subsets similar to traditional histogram analyses.
- Approximate computation: Approximate Frequentist Computation adapts Approximate Bayesian Computation to frequentist particle-physics inference using kernel density estimates and summary statistics.Its likelihood estimate uses generated events, a kernel bandwidth, and a low-dimensional summary-statistics construction.
- Classifier methods: Calibrated classifiers estimate likelihood ratios from classifiers separating samples generated under θ0 and θ1.The paper implements both point-by-point and agnostic parameterized classifiers, including a morphing-aware variant.
- Neural density estimation: Neural conditional density estimators model target densities through invertible transformations or related autoregressive constructions.A detailed treatment and implementation for the example process are deferred to future work.
2. Particle-physics structure
Particle-physics simulators expose joint likelihood-ratio and score information beyond generated observables. The paper uses this structure to build parameterized estimators and score-based summaries for scalable inference.
- Simulator information: Particle-physics simulations provide joint likelihood ratios and scores that can support inference strategies tailored to the simulator structure.These quantities use latent simulator information in addition to observed samples.
- Likelihood-ratio estimators: Ratio regression directly learns the likelihood ratio, while differentiable parameterized models can also derive an estimated score.The parameterized approaches include agnostic and morphing-aware versions.
- Score-augmented training: Score-augmented training combines classifier or ratio losses with squared error against the simulator-provided joint score.The score term is weighted alongside the primary likelihood-ratio or classification objective.
- Local score methods: Local score regression treats the score at a reference point as sufficient statistics in the neighborhood of that point.The score estimator can then be used for density estimation and likelihood-ratio inference.
- Local score methods: Sallino further compresses high-dimensional observations into one scalar without losing local sensitivity, making it suitable for many theory parameters.Its likelihood-ratio estimation uses univariate density estimators.
- Morphing-aware methods: Morphing-aware strategies exploit basis-point structure, but a tested latent-basis strategy can suffer huge variance because its joint ratios and scores span a large range.The resulting convergence is extremely slow for that strategy.
D. Calibration
Calibration is an additional post-training step intended to improve raw likelihood-ratio estimators. The paper considers probability calibration and expectation calibration.
- Calibration: Calibration transforms a trained raw estimator ˆr_raw through a function C so that ˆr_cal better estimates the true likelihood ratio.The procedure is applied after the initial estimator has been trained.
- Calibration: The paper considers probability calibration and expectation calibration as two ways to define the calibration function.Both approaches target improvement of the raw estimator.
- Probability calibration: Classifier decision functions may require calibration because good class separation does not ensure a direct probabilistic interpretation of the likelihood ratio.The decision function can be approximately monotonic in the ratio without equalizing it.
2. Expectation calibration
Expectation calibration uses the likelihood-ratio expectation under the denominator hypothesis as a normalization condition. Its implementation and evaluation rely on finite simulated samples and neural estimators of event-level ratios.
- Expectation condition: A likelihood-ratio estimator should have expectation one when evaluated under the denominator hypothesis.This property motivates expectation calibration and provides a normalization target.
- Expectation calibration: The expectation is numerically estimated from N events and can be used to rescale an estimator that violates the condition.The calibration factor is computed from simulated events.
- Expectation calibration: For a perfect estimator, the variance of the numerical expectation depends on the finite number of events used to calculate it.The paper derives this variance under the perfect-estimator assumption.
- Caveat: Rare events with large estimated ratios can dominate expectation calibration and degrade performance on the bulk distribution when misestimated.This is a practical failure mode of the calibration procedure.
- Estimator implementation: The example uses fully connected neural networks for classifier, score, and ratio regression, with point-by-point, agnostic parameterized, and morphing-aware setups.The morphing-aware model uses separate networks for basis components.
- Local methods: For the local methods, the score reference point is the Standard Model, with two-dimensional histograms for Sally and one-dimensional histograms for Sallino.Both methods estimate densities point by point in the numerator parameter.
- Morphing uncertainty: Morphing weights can increase estimator errors substantially, especially away from basis points.The uncertainty is associated with combining basis components.
F. Challenges and diagnostics
Finite training data, interpolation, model complexity, and morphing weights create estimator uncertainties. The paper proposes closure tests and ensemble-based diagnostics, while noting that passing such tests does not guarantee an optimal likelihood-ratio estimator.
- Sources of uncertainty: Finite training and calibration samples make likelihood-ratio estimators deviate from the truth, requiring additional modeling uncertainty.Different strategies experience different finite-statistics variance sources.
- Sources of uncertainty: Point-by-point methods incur training-sample and interpolation uncertainties, whereas morphing-aware methods can amplify small basis-estimator fluctuations through large weights and cancellations.The trade-off is between local data scarcity, learned parameter dependence, and physics-structured reuse of training data.
- Morphing uncertainty: For morphing-aware estimators, errors on the combined log ratio can be 100 times larger than errors on individual component estimators in some parameter regions.The effect is particularly pronounced far from basis points.
- Diagnostics: Closure tests can correct estimators, define uncertainty bands, or identify estimators for rejection.The paper discusses expectation, reweighting, ensemble, and reference-hypothesis checks.
- Diagnostics: Ensemble variance measures uncertainty from training data and random seeds, while reference-hypothesis variation checks stability under the choice of denominator reference.Reference variation has no proper likelihood-based statistical interpretation.
- Diagnostics: The expected ratio under the denominator hypothesis should be close to one and can serve both as a calibration target and a diagnostic.The paper notes that calibration based on this expectation can reduce performance in some situations.
- Interpretation: Passing closure tests does not guarantee a good likelihood-ratio estimator, although statistically correct exclusion limits can still be derived.Such limits may be non-optimal but are not wrong.
B. Neyman construction
The paper compares asymptotic and toy-experiment constructions for setting limits, emphasizing computational trade-offs and estimator behavior. In the WBF Higgs example, score-informed methods outperform traditional histograms, while Neyman-based contours preserve statistical correctness.
- B. Neyman construction: Toy experiments are more computationally expensive than asymptotic limits but are useful outside the asymptotic regime or when estimator uncertainties are unreliable.The resulting constraints may be non-optimal, but they remain conservative.
- B. Neyman construction: The profile log likelihood ratio requires maximizing separately for every toy experiment, substantially increasing computation in high-dimensional parameter spaces.A fixed-hypothesis likelihood ratio provides a faster alternative, with the Standard Model as the natural EFT reference.
- B. Neyman construction: Cascal and Rascal provide very accurate likelihood-ratio estimates, while all machine-learning strategies outperform traditional one- or two-dimensional histograms and Approximate Frequentist Computation.The comparison uses an idealized WBF Higgs setup with access to the true likelihood ratio.
- B. Neyman construction: Morphing-aware estimators perform comparably to or worse than the two-dimensional histogram because basis-estimator errors are amplified by large weights and cancellations.Probability calibration generally improves results, whereas expectation calibration often increases prediction variance.
2. Efficiency and speed
The methods are designed to reduce training and evaluation costs while retaining strong likelihood-ratio performance. Score-informed approaches achieve especially favorable training efficiency, fast evaluation, and tighter physics constraints than histogram analyses, although asymptotic contours can be slightly too tight.
- 2. Efficiency and speed: Approximately 100 000 training events give Rascal exceptional performance, whereas Carl requires a sample two orders of magnitude larger for comparable performance.Rascal uses more simulator information; Carl is the more general method but does not exploit those additional quantities.
- 2. Efficiency and speed: Cascal or Rascal leads to the best likelihood-estimation and classification metrics during training.The learning-curve metrics are not directly comparable to the prior-weighted metrics in Table IV and Fig. 13.
- 2. Efficiency and speed: All tested algorithms estimate the likelihood ratio for 50 000 events in around one second or less for fixed hypotheses.Local score regression is particularly efficient because its score estimator is evaluated once, with only density estimation repeated across θ0 values.
- 3. Physics results: The new machine-learning strategies produce visibly tighter Wilson-coefficient constraints than the doubly differential histogram, with score-based estimators close to the true-likelihood contours.Traditional jet-pT and Δφjj histograms are overly conservative, especially where destructive interference dominates.
- 3. Physics results: Asymptotic estimated contours can be slightly too tight, but profiling estimator uncertainties can mitigate this issue.Neyman-based contours avoid this specific undercoverage problem, though they may be non-optimal.
B. Detector effects
With detector smearing, the true likelihood becomes intractable, but Neyman constructions still provide statistically correct exclusion limits. Likelihood-ratio estimators remain more powerful than histograms, with Cascal and Rascal strongest in the example.
- Detector effects: Detector smearing makes the true likelihood intractable, preventing direct validation against true likelihood contours.The observed ratio is stochastically smeared relative to the joint parton-level ratio.
- Detector effects: Neyman constructions with toy experiments provide statistically correct expected exclusion contours under smearing.The construction is guaranteed to cover even when the true likelihood contours are unavailable.
- Detector effects: Likelihood-ratio estimators produce robust bounds that are clearly more powerful than traditional histogram limits.This ordering persists in the smeared setup despite the unavailable true-likelihood baseline.
- Detector effects: Cascal and Rascal yield the strongest limits among the evaluated methods.The methods exploit simulator information beyond ordinary observable samples.
- Detector effects: The methods scale to many observables and high-dimensional parameter spaces without approximating hard-process, shower, or detector effects.Their likelihood-ratio evaluations can be performed in microseconds.
- Detector effects: Score-based methods provide sufficient statistics locally and can compress observations to a scalar without losing local information.Sallino compresses the estimated score vector independently of the number of theory parameters.
Appendix A: Appendix
The appendix documents the toy smearing model and the implementation of histogram and kernel-density likelihood-ratio methods. These approaches require selecting and estimating low-dimensional summary statistics, with bandwidth choices controlling the trade-off between sample requirements and smoothing.
- Detector model: The toy detector model smears lepton momenta and models jet properties through transfer-function distributions.The jet-resolution parameters are based on MadWeight defaults.
- Implementation: The documented analysis separates training, evaluation, calibration, and parameter choices for comparing the techniques.Histograms were not additionally calibrated, while AFC calibration was left for future work.
- Histogram analysis: Histogram likelihood ratios are obtained by binning selected observables separately under numerator and denominator hypotheses.For equal binning, the estimate is the ratio of corresponding bin contents.
- Histogram analysis: The example compares one- and two-dimensional histograms built from leading-jet transverse momentum and the dijet azimuthal separation.The variants range from one-dimensional histograms to coarse, medium, and fine two-dimensional binnings.
- AFC: AFC estimates numerator and denominator densities with kernel density estimation in a chosen summary-statistics space.It can be used point by point or in a morphing-aware parameterized form when the required structure holds.
- AFC: AFC performance depends strongly on summary-statistics choice and kernel bandwidth.Small bandwidths require large training samples, whereas large bandwidths oversmooth and lose information.
c. Calibrated classifiers ( Carl)
Carl estimates likelihood ratios by training classifiers to distinguish samples from two hypotheses, then transforming and optionally calibrating the classifier output. Related regression methods instead train directly on simulator-provided joint likelihood ratios.
- Carl: Carl exploits invariance under monotonic transformations of the likelihood ratio to derive a ratio estimator from a binary classifier.The classifier is trained on numerator and denominator samples and its decision function is transformed using the likelihood-ratio relation.
- Carl: Carl supports point-by-point, agnostic parameterized, and morphing-aware estimators when the morphing condition holds.The parameterized versions include theory parameters as inputs or use component-network weights.
- Carl: The classifier is trained with binary cross-entropy to discriminate numerator and denominator samples.The numerator and denominator labels are y = 0 and y = 1, respectively.
- Calibration: Classifier outputs can be calibrated with isotonic regression before evaluation as likelihood-ratio estimates.Calibration curves are shown for the classifier-based and regression-based estimators.
- Rolr: Rolr regresses directly on joint likelihood ratios available from matrix-element information at parton-level phase-space points.It requires a generator that can evaluate squared matrix elements for specified phase-space points and theory parameters.
- Parameterized estimators: Parameterized ratio estimators evaluate a learned function of the observation and theory parameter, with optional calibration at evaluation time.The same parameterized structure can be implemented agnostically or through morphing-aware component networks.
e. Carl + score regression ( Cascal)
Cascal combines classifier training with score regression, while Rascal combines likelihood-ratio and score regression. Both use simulator-provided derivative information to enrich parameterized likelihood-ratio estimators.
- Cascal: Cascal combines Carl-style classification with squared error between the estimated score and the simulator-provided joint score.The score term is weighted by the hyperparameter α.
- Cascal: The classifier term captures a fixed-hypothesis likelihood ratio, whereas the score term captures its relative change with theory parameters.These components provide complementary information during training.
- Cascal: Cascal requires access to joint score information and therefore supports only parameterized estimators.Its agnostic and morphing-aware forms use the parameter dependence to derive the estimated score.
- Implementation: Both methods include optional isotonic calibration and use α to weight the score contribution in their combined losses.The documented neural-network settings use 100-unit tanh layers and Adam optimization.
- Rascal: Rascal combines squared errors on the likelihood ratio and score in a single parameterized regression loss.The likelihood-ratio term handles a fixed comparison, while the score term describes changes with the numerator parameter.
- Rascal: Rascal requires generators that provide both joint likelihood ratios and joint scores through squared matrix-element evaluations.Like Cascal, it can be used in agnostic or morphing-aware parameterized forms.
g. Local score regression and density estimation ( Sally)
Sally estimates the score from simulator events and uses it to construct likelihood-ratio estimates. It exploits joint simulator information while requiring access to the joint score or squared matrix elements.
- Local information: The score at a reference parameter point is sufficient for the local likelihood approximation and acts as a set of optimal observables near that point.For the example process, the Standard Model is used as the natural reference hypothesis.
- Score construction: Particle-physics generators provide parton-level momenta and squared matrix elements, allowing the joint score to be calculated for generated events.The derivatives can be evaluated numerically, or obtained from morphing weights when the process has the required morphing structure.
- Requirements: Sally requires a generator with access to the joint score, which in particle physics means evaluating squared matrix elements at phase-space points and theory parameters.This requirement limits applicability to processes whose generators expose the needed matrix-element information.
- Structure and training: Sally regresses the score using generated observables and corresponding joint scores, then performs density estimation in estimated score space.The score regression uses a vector-valued neural network trained with squared error; multidimensional histograms estimate densities for tested parameter points.
- Evaluation: For an observation, Sally evaluates the score regressor, extracts numerator and denominator histogram contents, and calculates the likelihood ratio.The estimate is evaluated separately for each tested parameter pair using Eq. (38).
h. Local score regression, compression to scalar, and density estimation ( Sallino)
Sallino compresses the estimated score to a scalar aligned with the tested parameter-point difference, then estimates likelihood ratios in that one-dimensional space. The accompanying results compare estimator variants and show genuine overlap between benchmark hypotheses.
- Local compression: Near the Standard Model, likelihood ratios depend on the scalar product of the score and the parameter-point difference.Sallino defines this scalar product as h-hat and uses it as an observable for likelihood-ratio estimation.
- Requirements: Sallino requires access to the joint score, which requires evaluating squared matrix elements for specified phase-space points and theory parameters.The method therefore depends on generator information beyond reconstructed observables alone.
- Method: Sallino regresses the score from generated observables and joint scores, then fills one-dimensional histograms of the compressed h-hat observable.The score regression is independent of the tested hypothesis, while density estimation is performed point by point in parameter space.
- Evaluation: For each tested parameter pair, the estimated score is multiplied by the parameter difference, and histogram contents provide the estimated likelihood ratio.The evaluation extracts numerator and denominator bin contents after computing h-hat for the observation.
- Additional results: The reported estimator comparisons use expected mean squared error on the log likelihood ratio, with both untrimmed and trimmed versions.The tables compare alternative versions of the listed techniques and identify estimators included in the main analysis.
- Additional results: The benchmark ROC curves show poor separation because the two probability distributions genuinely overlap, matching the true-likelihood ROC curve.The true-likelihood ROC AUC for the benchmark scenarios is 0.6276.