Source-linked AI summary
Extending the Bump Hunt with Machine Learning
Jack H Collins, Kiel Howe, Benjamin Nachman
TL;DR
The paper addresses limited coverage and simulation dependence in resonance searches by extending bump hunting with a model-agnostic CWoLa classifier trained directly on data. The method uses mass-defined signal and sideband samples plus mass-uncorrelated auxiliary features, and it identifies a 3 TeV resonance while selecting signal-consistent event populations in realistic studies.
Problem
Dedicated searches have reduced sensitivity away from target processes, while simulation mismodeling can limit machine-learning classifiers trained on simulated data.
Method
The method trains a classifier to distinguish a mass-defined signal region from a sideband using auxiliary observables decorrelated from mass under the background-only hypothesis, then performs a thresholded bump hunt.
Results
A clear bump develops at 3 TeV under stronger classifier thresholds, and the classifier selects event populations consistent with the injected signal.
Takeaways & Limitations
The extended bump hunt can identify potential signal-enhanced samples directly from data without requiring the signal process to be known in advance.
Takeaways & Limitations
The method requires a signal localized in one variable with a smooth background and is vulnerable to statistical fluctuations unless training, validation, and test samples are separated.
Abstract
from arXiv · showhide
The oldest and most robust technique to search for new particles is to look for `bumps' in invariant mass spectra over smoothly falling backgrounds. We present a new extension of the bump hunt that naturally benefits from modern machine learning algorithms while remaining model-agnostic. This approach is based on the Classification Without Labels (CWoLa) method where the invariant mass is used to create two potentially mixed samples, one with little or no signal and one with a potential resonance. Additional features that are uncorrelated with the invariant mass can be used for training the classifier. Given the lack of new physics signals at the Large Hadron Collider (LHC), such model-agnostic approaches are critical for ensuring full coverage to fully exploit the rich datasets from the LHC experiments. In addition to illustrating how the new method works in simple test cases, we demonstrate the power of the extended bump hunt on a realistic all-hadronic resonance search in a channel that would not be covered with existing techniques.
1 Introduction
The paper extends bump hunting with a data-driven classifier that uses auxiliary observables to identify signal-like events without requiring a known signal process. This addresses limited coverage from dedicated searches and simulation-dependent classifiers.
- 1 Introduction: Dedicated searches can miss signals because sensitivity diminishes away from target processes and analyzing every possible topology is infeasible.Global searches also rely heavily on simple objects and simulation for background estimation.
- 1 Introduction: Modern machine-learning classifiers exploit subtle jet-radiation correlations, but simulation mismodeling can make classifiers sub-optimal on data.Existing multivariate classifiers may require large post-hoc mis-modeling corrections.
- 1 Introduction: The proposed method trains a classifier to distinguish a resonance region from a mass sideband using observables decorrelated from mass under the background-only hypothesis.A bump hunt is then performed on the mass distribution after thresholding the classifier output.
- 1 Introduction: Because the method uses Classification Without Labels, it learns directly from data and is insensitive to simulation mismodeling while remaining agnostic to the signal process.This can provide sensitivity to signatures without dedicated searches.
- 1 Introduction: The approach combines auxiliary-feature information with the extended bump hunt, sharing with sPlot the use of discriminating variables while requiring no auxiliary-feature distribution as input.Unlike sPlot, the extended bump hunt uses machine learning to identify signal-like regions of phase space.
2 Bump Hunting using Classification Without Labels
CWoLa hunting constructs a signal region and sideband in invariant mass, trains a classifier on auxiliary information, and then performs a thresholded bump hunt. Its benefit depends on auxiliary features being similarly distributed for background in both regions but more signal-like in the resonance region.
- 2 Bump Hunting using Classification Without Labels: The invariant mass is smooth for background but localized near m0 for signal, while Y represents the event information beyond the resonance variable.This auxiliary information can include properties of the objects and their surroundings.
- 2 Bump Hunting using Classification Without Labels: CWoLa hunting trains a classifier to distinguish a mass-defined signal region from a sideband using auxiliary observables, then applies a thresholded bump hunt across mass hypotheses.The sideband width is chosen so the auxiliary-feature distribution is nearly matched between regions under background-only conditions.
- 2 Bump Hunting using Classification Without Labels: The auxiliary feature improves significance when q > √p, with significance scaling as qNs/√Nbp after selecting Y = 1.Here p and q are the probabilities for Y = 1 under background and signal, respectively.
- 2 Bump Hunting using Classification Without Labels: For Nb = 1000 and Ns = 20, rejection probability increases especially when p is small and q is close to 1, while q → 0 returns the 0.05 false-positive probability.The p = q = 1 case corresponds to the standard search without additional information.
- 2 Bump Hunting using Classification Without Labels: The simplified model captures the method’s central practical questions: how to identify useful auxiliary information and how to incorporate it into the bump hunt.The paper uses later sections to address both questions with neural networks and a complete procedure.
3 Illustrative Example: Learning to Find Auxiliary Information
The illustrative two-dimensional example shows that CWoLa can learn useful auxiliary-information structure without truth labels and recover much of an ideal tagger’s discriminating power. Performance depends on the classifier threshold and training stability, with the strongest shown significance reaching 10.8σ.
- Ideal benchmark: The ideal tagger has expected significance 15σ, while CWoLa aims to recover much of this power by learning the signal-region shape from pseudodata.The ideal decision region is the square centered at the origin, and the optimal classifier is based on the likelihood ratio.
- CWoLa procedure: CWoLa trains on auxiliary variables from target-window and sideband events, using mass-defined mixed samples rather than truth labels.The toy setup uses two-dimensional auxiliary information Y = (x, y), with the network distinguishing sideband from signal-region events.
- Training stability: Training can overfit statistical fluctuations or fail to converge on the signal region, motivating multiple classifiers and procedures that mitigate training fluctuations.Overtrained regions admit more background after thresholding, reducing classifier effectiveness.
- Thresholded bump hunt: 10.8σ is the maximum shown significance at a 5% classifier efficiency, compared with 13.9σ for the ideal classifier on the same pseudodataset.The reported thresholds give significances of 3σ, 9.4σ, 10.8σ, and 3.4σ for no threshold, 10%, 5%, and 1% efficiency; the 0.2% threshold removes statistical significance.
- Ensemble behavior: Narrower signal regions make training less reliable: about 50% of networks find the signal at ws = 0.1, but only about 5% do so at ws = 0.05.These cases maintain an expected ideal significance of 15σ while changing the signal-region width and signal count.
4 Full Method
The full method uses nested cross-validation to train classifiers on signal-region and sideband data without reusing events for selection, then performs a bump hunt on the selected invariant-mass distribution. It also specifies statistical interpretation and highlights computational and global-significance considerations.
- 4 Full Method: The procedure requires limited modeling of auxiliary features to ensure their correlations with mres remain minimal under the background-only hypothesis.That model may come from simulation, theory, or a sufficiently signal-devoid data sample.
- 4 Full Method: Nested cross-validation prevents events from being selected by classifiers trained on those same events while retaining all data for the bump hunt.The dataset is split into five subsets, with each subset serving once as test data and the remaining subsets used for training and validation.
- 4 Full Method: For each test subset, twenty neural networks are trained per training-validation split, the best model is selected by validation performance, and the four selected models are averaged.Signal-region events receive label 1 and sideband events label 0; validation performance compares the true-positive rate at a specified false-positive rate.
- 4 Full Method: Selected events from all test subsets are merged into a new mres histogram, where a localized excess is evaluated with standard bump-hunting methods.The selection threshold retains a chosen fraction of the most signal-like events, after which a smooth background is fitted with the signal region masked.
- 4.1 Interpreting the Results: The primary output is a local p-value, while global p-values require accounting for scanned mass windows and neural-network threshold fractions, potentially through computationally expensive pseudo-experiments.Bonferroni correction is straightforward for fixed, non-overlapping mass bins but can be over-conservative when the mass-bin width is scanned.
5 Physical Example
A realistic all-hadronic dijet search applies CWoLa hunting to a benchmark W′→WX, X→WW signal that existing dedicated taggers are not designed to target. The classifier identifies signal-like jet-substructure populations and, after selection, produces a clearer resonance bump while requiring closure tests for residual mass correlations.
- 5.3 Results: A clear 3 TeV bump develops after applying stronger neural-network thresholds, with the background estimated by fitting a smooth function outside the signal region.The analysis uses a profile-likelihood counting experiment with fit parameters treated as nuisance parameters.
- 5.3 Results: The method requires auxiliary features to avoid sculpting artificial mJJ bumps, so simulation closure tests and checks for residual kinematic correlations are necessary.Residual kinematic correlations can arise when features correlate with jet transverse momentum and mJJ; mixed-event samples are proposed as a data check.
- 5.3 Results: CWoLa taggers improve with increasing statistics but do not reach the performance of a fully supervised tagger trained on labelled signal and background.The supervised tagger provides a measure of the maximum performance achievable with the selected variables.
- 5 Physical Example: The benchmark probes a signal topology outside the design target of standard W/Z taggers, because the X jet rarely passes their cuts at εB ∼10^-4.CWoLa can identify unexpected signals when the signal-to-background ratio is sufficiently high, but underperforms when that ratio is too low.
6 Conclusions
The paper presents CWoLa hunting, a model-agnostic anomaly-detection method that combines a smooth resonance variable with auxiliary features for classifier-based signal enhancement. Its applicability depends on a localized bump, useful additional features, and weak background correlations between those features and the resonance variable.
- 6 Conclusions: CWoLa hunting trains a classifier on auxiliary observables to distinguish a signal region from a sideband, then performs a bump hunt after thresholding its output.The method uses mixed samples whose backgrounds should have nearly identical characteristics under the background-only hypothesis.
- 6 Conclusions: When a distinctive signal is present, the classifier output becomes an effective signal-background discriminant, whereas no signal produces no clear output pattern.Threshold selection therefore preserves a smooth mass spectrum under the null and produces a bump when a signal is present.
- 6 Conclusions: The demonstrated dijet application augments dijet mass with jet-substructure information and trains a deep neural network without requiring a known signal topology.Related resonance variables include single-jet mass and the average mass of pair-produced objects.
- 6 Conclusions: The approach requires a smooth bump variable, auxiliary features with potential signal-background discrimination, and weak background correlations across the resonance width.Correlations can alternatively be removed by transforming variables or penalized during classifier training, with closure tests used for validation.
- 6 Conclusions: CWoLa hunting and related weakly supervised strategies may help uncover beyond-the-Standard-Model signals in existing LHC datasets.The conclusion frames this as a potential benefit of applying modern machine learning directly to data.
A Statistical Analysis
The statistical analysis selects classifier-enhanced events, fits their sidebands with a smooth background function, and evaluates a profile-likelihood significance. Toy studies find no apparent distortion from cross-validation correlations, while the reported p-value remains local.
- A Statistical Analysis: The selected dataset is binned in dijet mass, the signal region is masked, and a smooth three-parameter function is fitted to estimate the background.The signal-region yield is predicted by summing fitted expectations in its three bins, with fit-parameter uncertainties propagated to the prediction.
- A Statistical Analysis: The hypothesis test uses a profile likelihood based on the total signal-region count because the signal shape is unknown a priori.The likelihood combines a Poisson factor for the observed count with a Gaussian constraint for the background nuisance parameter.
- A Statistical Analysis: Asymptotic formulae convert the test statistic into significance Z = √q0 and p0 = 1 − Φ(Z).Here µ denotes the signal rate, while θ represents the nuisance parameter associated with background systematic uncertainty.
- A Statistical Analysis: Cross-validation event rates need not be uncorrelated because events used for training are applied to other samples.If strong correlations are expected, reliable p-values require calibration with many fresh toy experiments or bootstrapped samples, which is computationally expensive.
- A Statistical Analysis: The toy study found no apparent deviation between NN-trained, randomly selected, and asymptotic test-statistic distributions, indicating no observed cross-validation distortion.The study used 103 NN toys and 105 random-selection toys, with the distributions compared in Figure 12.
- A Statistical Analysis: The procedure computes only a local p-value, so a low value calls for a follow-up analysis or a global correction for the trials factor.The paper mentions an orthogonal dataset for follow-up and Bonferroni correction as one route to a global p-value.
B Dijet Mass Scans
The dijet mass-scan analysis compares invariant-mass distributions before and after tagger cuts across the full scan range and displays the resulting p-values.
- B Dijet Mass Scans: Figures 13 and 14 show dijet invariant-mass distributions before and after tagger cuts over the full mass-scan range, with corresponding p-values shown in Figure 8.The passage identifies the distributions and p-value display but does not state a specific scan outcome.