Source-linked AI summary
The LHC Olympics 2020: A Community Challenge for Anomaly Detection in High Energy Physics
Gregor Kasieczka, Benjamin Nachman, David Shih, Oz Amram, Anders Andreassen, Kees Benkendorfer, Blaz Bortolato, Gustaaf Brooijmans, Florencia Canelli, Jack H. Collins, Biwei Dai, Felipe F. De Freitas, Barry M. Dillon, Ioan-Mihail Dinu, Zhongtian Dong, Julien Donini, Javier Duarte, D. A. Faroughy, Julia Gonski, Philip Harris, Alan Kahn, Jernej F. Kamenik, Charanjit K. Khosa, Patrick Komiske, Luc Le Pottier, Pablo Martín-Ramiro, Andrej Matevc, Eric Metodiev, Vinicius Mikuni, Inês Ochoa, Sang Eon Park, Maurizio Pierini, Dylan Rankin, Veronica Sanz, Nilai Sarda, Urous Seljak, Aleks Smolkovic, George Stein, Cristina Mantilla Suarez, Manuel Szewc, Jesse Thaler, Steven Tsan, Silviu-Marian Udrescu, Louis Vaslin, Jean-Roch Vlimant, Daniel Williams, Mikaeel Yunus
TL;DR
HEP needs complementary, less model-dependent searches because conventional approaches can leave unconventional signatures unexplored, while anomaly detection requires dedicated collider-oriented methods. The paper introduces and reviews the LHC Olympics 2020, using an R&D dataset and blinded black boxes to develop and test methods. The challenge engaged many teams and methods, while the review identifies shared scaling challenges and the need for continued method development and combinations of approaches.
Problem
Model-dependent searches can leave unconventional new-physics signatures and substantial phase space unexplored, motivating complementary anomaly-detection approaches for HEP.
Method
The LHC Olympics uses labeled R&D data and three black boxes with different simulated backgrounds and potential anomalies to develop and test anomaly-detection methods.
Results
Many teams tested diverse techniques on the black boxes before knowing their signals, refining strategies through workshop feedback and gaining practical experience.
Takeaways & Limitations
The reviewed methods offer a snapshot of a rapidly developing field, but reaching the data’s full physics potential will require combined approaches and further development.
Takeaways & Limitations
Non-resonant searches lack a general background-estimation approach, while high-dimensional comparisons can exceed the available systematic-uncertainty covariance information.
Abstract
from arXiv · showhide
A new paradigm for data-driven, model-agnostic new physics searches at colliders is emerging, and aims to leverage recent breakthroughs in anomaly detection and machine learning. In order to develop and benchmark new anomaly detection methods within this framework, it is essential to have standard datasets. To this end, we have created the LHC Olympics 2020, a community challenge accompanied by a set of simulated collider events. Participants in these Olympics have developed their methods using an R&D dataset and then tested them on black boxes: datasets with an unknown anomaly (or not). This paper will review the LHC Olympics 2020 challenge, including an overview of the competition, a description of methods deployed in the competition, lessons learned from the experience, and implications for data analyses with future datasets as well as future colliders.
1 Introduction
Traditional LHC searches are largely model-dependent, leaving unconventional signatures and substantial phase space unexplored. The paper motivates machine-learning anomaly detection and introduces the LHC Olympics as a benchmark for model-agnostic searches.
- Current search paradigm: Top-down LHC searches select specific signal models, even when their reported limits are model-independent.Event selection and background estimation remain strongly model-dependent.
- Motivation: Model dependence may create blind spots, leaving much phase space and many possible signals unexplored.The paper presents complementary search strategies as necessary for broader exploration.
- Existing model-independent searches: Generic bump hunts require few signal assumptions but usually use only resonant structure, limiting their sensitivity.More differential model-independent searches face look-elsewhere and simulation-systematics challenges.
- Machine-learning anomaly detection: HEP anomaly detection targets ensemble-level over-densities rather than isolated off-manifold events common in industrial anomaly detection.This distinction motivates dedicated approaches for collider data.
- LHC Olympics 2020: The LHC Olympics 2020 provides mostly Standard Model black boxes in which participants search for new physics and identify its properties.The challenge restricts attention to hadronic final states where sidebands can estimate backgrounds.
2 Dataset and Challenge
The LHC Olympics combines a labeled R&D sample with simulated black boxes containing mostly Standard Model events and potentially unknown anomalies. The datasets support method development, blinded testing, and characterization of diverse signal and background settings.
- Challenge format: Contestants submitted a black-box number, method abstract, null-hypothesis p-value, inferred new-physics properties, and estimated signal yield.Plots and Jupyter notebooks were encouraged, and the datasets were publicly downloadable.
- R&D Dataset: The R&D dataset contains one million QCD dijet background events and 100,000 Z′ →XY signal events with specified particle masses.Events were generated with Pythia and Delphes without pileup or multiparton interactions.
- Background simulation: R&D events use Pythia defaults, whereas Black Boxes 1 and 3 use modified Pythia settings; Black Box 2 uses Herwig++.The table summarizes these generator-setting differences across datasets.
- Black Box 1: Black Box 1 preserves the R&D signal topology but changes the particle masses and includes 834 signal events among one million total events.Its background uses the same generators with altered Pythia and Delphes settings.
- Black Box 3: Black Box 3 contains a 4.2 TeV KK graviton with dijet and trijet decay modes, including 1,200 dijet and 2,000 trijet events.The event counts were chosen so that finding only one mode would not produce a significant excess.
- Simulation caveat: A modified Pythia setting produced a bump-like multijet-background feature that was flagged as anomalous by one analysis.The paper labels this identification “Human NN.”
Individual Approaches
The paper organizes participating approaches by supervision level and distinguishes blinded black-box tests, unblinded black-box results, and R&D-only studies. This taxonomy frames the methods and the evidentiary roles of their results.
- Method categories: The approaches are grouped into unsupervised, weakly supervised, and (semi)-supervised categories according to label information used during training.Unsupervised methods learn directly from background-dominated data and typically seek events with low p(background).
- Result categories: Results are classified as blinded black-box contributions, unblinded black-box results, or studies using only the R&D dataset.Each category provides insight but serves a different purpose, with blinded tests approximating a real analysis.
- Summary of contributions: The paper summarizes the methods and result types in Table 2, giving precedence to workshop results when overlapping contributions exist.Some workshop results and later section results differ in their blinding status.
3 Unsupervised
This section presents a VRNN that models jets as variable-length constituent sequences for anomaly detection, then converts jet anomalies into event-level scores. Validation and black-box results show sensitivity to certain anomalous substructures, while also exposing scope limitations.
- 3.1.1 Method: A VRNN replaces a standard RNN encoder-decoder with a VAE, combining sequence modeling and variational inference for jet-level anomaly detection.The sequence architecture accommodates variable-length constituent lists and avoids direct loss contributions from zero padding.
- 3.1.1 Method: The model processes constituent pT, η, and φ sequences, reconstructs each constituent, and assigns jet Anomaly Scores from reconstruction and KL-divergence information.Jets are preprocessed to share reference mass, energy, and orientation, leaving substructure as the main difference between inputs.
- 3.1.1 Method: Kt-ordered constituent sequences consistently traverse separate prongs, making non-QCD-like substructure easier to model and producing a significant performance boost.The ordering is contrasted with pT-sorted sequences, which do not expose the prong structure as directly.
- 3.1.1 Method: Event Scores use the more anomalous leading or sub-leading jet score, are transformed to have mean 0.5, and increase toward more anomalous events.This event-level construction aligns the jet-level method with the challenge’s event-level search task.
- 3.1.2 Results on LHC Olympics: A 0.65 Event Score selection increased the R&D signal excess from 0.18σ to 2.2σ without significantly sculpting the background shape.The validation used a 0.5% contaminated sample containing 895113 background and 4498 signal events.
- 3.1.2 Results on LHC Olympics: Black Box 1 showed an mJJ enhancement just below 4000 GeV, consistent with its 3800 GeV Z′ signal, whereas Black Box 2 showed no significant excess and similar background-shape effects.The Black Box 2 result is consistent with its QCD-only composition.
- 3.1.2 Results on LHC Olympics: The method was insensitive to Black Box 3 because that signal involved varied final states beyond the two-prong large-R-jet substructure it targets.The stated sensitivity is therefore tied to anomalous substructure within large-R jets.
- Lessons learned: The authors identify preprocessing, especially kt ordering, as highly influential and expect constituent-based learning to have less jet-mass correlation than standard substructure variables.The possible reduction in mass correlation is presented as an expectation requiring further characterization.
3.2 Anomaly Detection with Density Estimation5
ANODE searches for resonant new physics by estimating data and background densities in discriminating features within a signal region. Its likelihood-ratio score separates signal-like events, achieves competitive proof-of-concept performance, and faces challenges at low background densities and higher dimensions.
- 3.2.1 Method: ANODE estimates where signal and background densities differ in feature space x to improve a resonance search in feature m without labels or a signal model.The signal region is scanned around candidate m0 values, with sidebands supplying background information.
- 3.2.1 Method: In the ideal mixture model, R(x|m) is the optimal test statistic, while no signal implies R(x|m) = 1.Signal produces density away from one when its feature density differs from the background density.
- 3.2.1 Method: The method learns pdata(x|m) in the signal region and pbackground(x|m) from sidebands, interpolating the latter into the signal region with conditional density estimation.Classification then uses their likelihood ratio R(x|m).
- 3.2.2 Implementation: Normalizing flows provide the proof-of-concept density estimator, although the ANODE procedure itself is general with respect to the density-estimation algorithm.The implementation uses conditional masked autoregressive flows optimized with a log-likelihood loss.
- 3.2.3 Results: Background events concentrate near R(x|m) = 1, while signal events form a higher-R tail in relatively rare background regions, matching the expected over-density pattern.The relevant signal region has R > 1 and pbackground(x|m) ≪ 1.
- 3.2.3 Lessons Learned: Density-ratio estimates are narrower around one at high background densities and more dispersed where background densities are low, indicating reduced estimation accuracy in sparse regions.This behavior is attributed to the relative availability of nearby data points.
- 3.2.3 Results: ANODE obtains an AUC of 0.82 and performance comparable to CWoLa hunting, while CWoLa performs better at high and especially low signal efficiencies.The comparison reflects supervised CWoLa’s direct likelihood-ratio approach versus ANODE’s unsupervised density estimation.
- 3.2.3 Lessons Learned: ANODE’s main practical challenges are obtaining precise background densities at very small S/B and extending density estimation to higher dimensions.Alternative neural density estimators are identified as possible avenues for improvement.
3.3 BuHuLaSpa: Bump Hunting in Latent Space8
BuHuLaSpa uses a VAE to encode collider events into a latent space while explicitly representing invariant mass as a latent dimension, enabling mass-aware anomaly classification. Studies found that optimizer choice, latent-vector norm, and early stopping strongly affect classification performance.
- BuHuLaSpa assumes events arise from a probabilistic generative model and uses a VAE to approximate likelihood and posterior distributions.
- Invariant mass as latent dimension: The method makes invariant mass an explicit latent dimension by sampling it around the reconstructed event mass with uncertainty σ(m_i) = 0.1m_i.
- Invariant mass as latent dimension: Explicit mass conditioning lets the decoder treat similar latent representations with different invariant masses differently, reducing pressure on other latent dimensions to encode mass.
- Invariant mass as latent dimension: Latent-space scans over z_i and m̃_i provide an explicit visualization of learned correlations between observables and invariant mass.
- Optimization and classification: Adagrad and Adadelta consistently outperform momentum-based optimizers such as Adam and Nadam for this VAE anomaly-detection task.
- Optimization and classification: Classification peaks and later declines during training; monitoring per-observable reconstruction-loss variance identifies an early-stopping point correlated with peak performance.
3.4 GAN-AE and BumpHunter9
GAN-AE combines an autoencoder with an adversarial discriminator, while BumpHunter evaluates mass-spectrum deviations using a background reference. The method discriminates signals on the R&D dataset, but background-shape mismodeling prevents meaningful black-box p-values.
- Full analysis workflow: The workflow combines two independent anomaly-detection algorithms to reduce background with GAN-AE and evaluate a global p-value with BumpHunter.
- GAN-AE: GAN-AE alternates training an MLP to expose reconstruction weaknesses and an autoencoder to reconstruct events while misleading the MLP.
- GAN-AE: The autoencoder loss combines mean Euclidean reconstruction distance, the MLP loss, and a distance-correlation term intended to decorrelate reconstruction error from invariant mass.
- BumpHunter: BumpHunter scans a data distribution against a reference with variable-width windows and uses pseudo-experiments to obtain a global p-value and significance.
- Results on LHC Olympics: On the R&D dataset, the background-only-trained GAN-AE obtained good discrimination between background and signals, although Euclidean distance remained correlated with dijet mass.
- Results on LHC Olympics: Black-box Euclidean-distance distributions were shifted relative to the R&D background, indicating sensitivity to background modeling and causing poorly fitted references.
- Results on LHC Olympics: Because the constructed reference backgrounds did not fit the post-cut data, the analysis could not evaluate a meaningful p-value for a potential signal.
3.5 Gaussianizing Iterative Slicing (GIS): Unsupervised In-distribution Anomaly Detection through Conditional Density Estimation10
GIS performs in-distribution anomaly detection by estimating conditional densities near a parameter of interest and comparing local signal density with interpolated background density. Applied to jet observables conditioned on dijet mass, it identified a concentrated anomaly near 3750 GeV and characterized its decay structure.
- The method searches for excess density in a narrow invariant-mass region rather than for outliers in the tails of data distributions.
- The analysis conditions on dijet invariant mass and uses jet masses, mass differences, and n-subjettiness observables selected from leading jets.
- GIS is a conditional normalizing flow that iteratively Gaussianizes one-dimensional marginalized distributions and models conditional dependence by binning and interpolation.
- The anomaly workflow estimates p_signal at each point, interpolates neighboring conditional densities as p_background, and computes α = p_signal/p_background.
- Results on LHC Olympics: The anomaly score strongly peaks around MJJ ≈ 3750 GeV.
- Results on LHC Olympics: Events selected with α > [1.5, 2.5, 5.0] cluster in leading-jet mass, jet-mass difference, and small τ21, indicating an overdensity absent at neighboring MJJ values.
- Results on LHC Olympics: For α > 2.0 events, the inferred particle mass is 3772.9 ± 8.3 GeV, with daughter masses M1 = 727.8 ± 3.8 GeV and M2 = 374.8 ± 3.5 GeV.
3.6 Latent Dirichlet Allocation11
The LDA method models collider events as mixtures of latent themes and uses jet-substructure representations to select signal-like events before a bump hunt. On Black Box 1, it found a dijet excess near the injected resonance, but its inferred decay-product masses were inaccurate and no compelling candidates appeared in Black Boxes 2 or 3.
- Method: LDA represents collider events as mixtures of latent theme distributions over binned measurements.The model infers event-level mixing proportions and theme multinomial parameters from unlabelled data.
- Method: Data representation and binning must balance signal-background discrimination against sufficient measurement co-occurrences for latent-distribution extraction.The authors considered mass-basis and Lund-basis jet-substructure observables, with coarse binning often needed to preserve co-occurrences.
- Method: The pipeline optimizes LDA hyperparameters in overlapping invariant-mass bins, selects signal and background themes, then performs a bump hunt on selected events.The uncut invariant-mass distribution provides the background template, normalized using sidebands.
- Results on LHC Olympics: 1.8σ and 3.8σ were reported for the simulated background and Black Box 1, respectively, with the inferred dijet resonance mass compatible with 3.8 TeV.The decay-product masses were estimated as 732 and 378 GeV, substantially above the LDA estimates.
- Results on LHC Olympics: The method found no compelling new-physics candidates in Black Boxes 2 or 3.The authors attribute the inaccurate decay-product estimates possibly to binning and sculpting effects.
- Lessons Learned: A realistic LDA implementation should test multiple data representations and binnings, include more jets, and incorporate global event variables.The authors note that focusing on the two leading jets missed a rare signal lacking rich jet substructure there.
3.7 Particle Graph Autoencoders15
Particle graph autoencoders represent jets as fully connected particle graphs and use edge convolutions to encode and reconstruct particle four-momenta. The method detected the injected Black Box 1 resonance with MSE loss, while performance and interpretation remain limited by loss-function and background-modeling choices.
- Method: PGAEs encode particle jets as graphs to exploit particle relationships for unsupervised anomaly detection in multijet events.Each particle is a node, with edges connecting every particle pair.
- Method: The encoder reduces particle four-momenta to two-dimensional representations, and the decoder reconstructs each particle’s four-momentum.Both stages use edge convolution layers with message passing and node-level aggregation.
- Method: MSE and Chamfer losses were compared because MSE depends on matching input and output particle order, whereas Chamfer loss is permutation invariant.The R&D study found better discrimination for an unseen signal with MSE loss.
- Results on LHC Olympics: The anomaly search applies a dijet bump hunt to events whose two leading jets both exceed the 90% reconstruction-loss quantile.Background in the outlier region is estimated from nonoutlier data using a fourth-order-polynomial transfer factor.
- Results on LHC Olympics: 2.1σ at 3.9 TeV was obtained for Black Box 1 with MSE, while Black Box 2 gave 0.8σ at 3.3 TeV with the same loss.The MSE Black Box 1 excess lies near the injected resonance at 3823 GeV.
- Lessons Learned: Further work is needed to develop a permutation-invariant loss that performs better for anomaly detection and a more general resonance-search procedure.The proposed extensions include multidimensional fits across dijet, trijet, and single-jet masses.
3.8 Regularized Likelihoods16
Regularized likelihood methods combine manifold reconstruction with density estimation to build anomaly scores from generative models. Although the combined metric outperformed its components on the R&D dataset, black-box modeling differences produced mass bias and prevented reliable hidden-signal detection.
- Method: M-flows combine autoencoder-like reconstruction error with the tractable density estimation of normalizing flows.The model learns a lower-dimensional data manifold and the density over that manifold.
- Method: The M-flow maps latent variables into data space, projects onto the manifold by discarding off-manifold variables, and learns manifold density with a normalizing flow.Training alternates between minimizing reconstruction error and negative log likelihood.
- Method: The anomaly score combines manifold density and reconstruction error, and the R_mjj score divides manifold likelihood by marginal dijet-mass likelihood estimated with KDE.The mass term was added to reduce the score’s bias toward high-mass events.
- Results: The combined anomaly metric performed better on the R&D dataset than its components and a basic normalizing flow.Performance was assessed using ROC curves.
- Results on LHC Olympics: Raising the R_mjj threshold from the 50th to the 70th percentile selected slightly higher-mass events without revealing a sharp resonance peak.The observed mass bias persisted across threshold choices.
- Lessons Learned: The method could not reliably find the hidden black-box signal, with this behavior persisting across R_mjj thresholds.The authors link the challenge to differences in background modeling between datasets.
- Lessons Learned: The challenge showed that neural networks alone cannot achieve good anomaly detection without a good background model.The authors emphasize that analysis details beyond the neural network require comparable attention.
3.9 UCluster: Unsupervised Clustering17
UCluster learns event embeddings through a particle-level mass-classification task supplemented by a clustering objective, then groups events with similar representations. On the R&D dataset, anomalous events concentrated in clusters and signal-to-background improved substantially, but identifying useful clusters was not conclusive for the challenge signals.
- Method: UCluster reduces event dimensionality while adding a clustering objective to encourage similar events to occupy nearby embedding regions.The method uses per-particle jet-mass classification to create the embedding and a clustering loss to organize events.
- Method: UCluster combines focal classification and clustering losses, with pretraining and K-Means initialization used to stabilize cluster-center learning.The combined loss weights the two components through hyperparameters.
- Method: The embedding is built with ABCNet, a graph-based architecture that treats reconstructed particles as nodes and learns node importance through attention mechanisms.The implementation uses graph aggregation layers based on nearest-neighbor relationships.
- Results on LHC Olympics: Anomalous BSM events concentrated in a common cluster, while the signal-to-background ratio increased from 1% to 2.5%.Increasing the number of clusters further enhanced the maximum cluster signal-to-background ratio.
- Results on LHC Olympics: The signal-to-background ratio increased to around 28% as the number of clusters increased.For initial significances between 2 and 6, approximate significance enhancements by factors of 3–4 were observed.
- Lessons Learned: No conclusive challenge results followed because identifying interesting clusters was not fully investigated and only one decay mode had distinguishing jet substructure.The authors suggest adapting the classification task or using a more general event embedding.
4 Weakly Supervised
The section compares weakly supervised and unsupervised anomaly-detection methods in hadronic dijet resonance searches, emphasizing their dependence on signal abundance and their complementarity. It also reports applications to LHC Olympics black boxes, including a significant BB1 excess and difficulties controlling background sculpting in BB3.
- CWoLa Hunting: CWoLa Hunting trains a classifier to distinguish a hypothesized resonance window from adjacent sidebands using features orthogonal to the resonant variable.If background distributions match between signal and sideband regions, classifier performance should be poor without a signal; a signal-region excess can instead drive discrimination.
- Results on LHC Olympics: On BB1, the analysis found a large 5σ excess near a resonance mass of 3500 GeV, with selected events showing two-pronged jet structure.The substructure study found clusters in jet masses around 400 and 750 GeV and evidence for a two-pronged structure in both jets, but no strong clustering in τ32.
- Lessons learned: CWoLa SALAD variants applied to BB3 heavily sculpted the background mJJ distribution, indicating that a more global decorrelation strategy is needed.Rescaling jet momenta by the total mass did not prevent the sculpting in these attempts.
- CWoLa and autoencoder comparison: CWoLa reaches AUC above 0.90 at large S/B, while the autoencoder maintains solid performance across the tested S/B range.CWoLa approaches the 0.98 AUC of a fully supervised classifier at large S/B, whereas reduced signal contamination limits its test performance.
- CWoLa and autoencoder comparison: CWoLa enhances excess significance by 3σ−8σ above S/B ∼3 · 10^-3, while the autoencoder gains at least 2σ−3σ below that range.The reported crossing shows that the preferred method depends on the amount of signal in the resonant region.
- Results on LHC Olympics: Tag N’ Train was competitive on the R&D dataset, but insufficient signal at 0.3% and 0.1% made it significantly worse than the autoencoder.Adding a dijet-mass cut made TNT comparable to CWoLa at larger signal fractions, while at 0.1% it outperformed CWoLa but not the autoencoder.
4.4 Simulation Assisted Likelihood-free Anomaly Detection21
This section studies Salad, which reweights background simulation using sideband data and interpolates the correction into the signal region. The method improves anomaly sensitivity and background prediction, while retaining limitations from residual bias and imperfect reweighting.
- Method: The reweighted simulation can estimate events passing a classifier threshold even when the classifier g(x) is correlated with the resonant feature m.This avoids the bump-sculpting restriction that applies to sideband fits.
- Method: Salad learns a sideband reweighting function that makes simulated background resemble data, then interpolates it conditionally on m_jj into the signal region.The signal region must be large enough that sideband signal contamination does not bias the reweighting.
- Results on LHC Olympics: Salad remains effective to about S/B ≲0.5%, whereas the Herwig-only tagger provides useful discrimination only down to about S/B ∼1%.At S/B ∼O(1), Salad performs similarly to a fully supervised classifier; as S/B approaches zero, Pythia curves approach a random classifier.
- Results on LHC Olympics: With interpolated reweighting, the background prediction is accurate within a few percent down to about 1% data efficiency.Without Dctr reweighting, the prediction is too low by a factor of two or more below 10% data efficiency.
- Lessons Learned: Salad’s numerical results do not fully reach a fully supervised classifier trained with inside knowledge about the data.The authors identify hyperparameter scans and calibration techniques as potential improvements, while residual-bias estimation remains difficult for direct background prediction.
- Lessons Learned: The simulation-augmented methods mitigate classifier sensitivity to dependencies between the resonant feature and classification features.SA-CWoLa completely recovers the ideal CWoLa performance in the studied setting, while Salad is comparable to SA-CWoLa and significantly better than random.
5 (Semi)-Supervised
The section reviews supervised and semi-supervised approaches that classify, model, or constrain anomalous collider events. Results show strong performance on development data, but generalization can fail when black-box signals or backgrounds differ from training.
- Supervised approaches: The ResNet-34 reached 92% accuracy and supplied event-level signal scores that were appended to tabular kinematic data.The combined representation was then used by a BDT for final event classification and metric estimation.
- Factorized topic modeling: The factorized-topic model classified events with AUCs of 0.88 and 0.81 using a likelihood-ratio discriminant.Its generative model uses a two-dimensional histogram and alternating minimization to obtain a locally optimal solution.
- Semi-supervised approaches: CWoLa can be more sensitive when a signal region exists because localization supplies an implicit signal prior, but selection biases must be controlled.The approach embeds assumptions about where a signal is expected, whereas the section also considers methods designed to preserve model-independent searches.
- QUAK and signal priors: Adding approximate signal priors to QUAK improved anomaly sensitivity without degrading model-independent sensitivity, sometimes approaching or exceeding supervised discrimination.A substantial gain for a 3-prong signal arose from adding a 2-prong signal prior, and incorrect priors could remain useful.
- Supervised approaches: Supervised models performed well on the R&D dataset but relatively poorly on black boxes, especially when signal structure differed from training.Black Box 3 produced outputs similar to pure background, while Black Box 1 yielded low-confidence and potentially spurious signal assignments.
6 Discussion
The discussion compares anomaly-detection submissions on blinded and unblinded black boxes rather than selecting a universally superior method. Several methods found the first resonance, while the second and especially third boxes exposed false-positive and generalization challenges.
- Lessons learned: The challenge did not identify one universally best method and instead aimed to foster novel unsupervised anomaly-detection tools.The authors emphasize that complementary evidence across methods is more informative than declaring a single winner.
- Challenge results: The challenge evaluated diverse methods chronologically across blinded submissions, workshop results, and later improvements.At least nine blind approaches were submitted for Black Box 1, with additional contributions reviewed in the paper.
- Black Box 1: Four Black Box 1 submissions identified the correct resonance mass within stated tolerances, but only Density Estimation accurately predicted the other observables.The correct mass was obtained within the claimed error by PCA or within ±200 GeV by LSTM, Tag N Train, and Density Estimation.
- Black Box 2: Black Box 2 contained no injected signal, yet several methods reported resonances, highlighting vulnerability to anomalies in statistical tails.Predicted resonance masses ranged from 4.2 to 5 TeV, while LDA reported no dijet-resonance signal.
- Black Box 3: No approach detected Black Box 3’s true 4.2 TeV resonance with two competing decay modes.Methods reported different resonance interpretations or no signal, and the signal structure differed from the shared development and Black Box 1 topology.
- Lessons learned: The first black-box resonance was detected repeatedly before unblinding, with the closest mass estimates mainly coming from likelihood-based methods or a signal-likelihood model.These methods likely benefited from the matching topology between the development signal and Black Box 1.
7 Outlook: Anomaly Detection in Run 3, the HL-LHC and for Future Colliders
The outlook identifies organizational, methodological, computational, and interpretive challenges for deploying anomaly detection at the LHC and future colliders. It emphasizes complementary searches, realistic background treatment, reproducibility, and broader future challenge designs.
- Organization and analysis strategy: A model-agnostic search program would benefit from a coherent analysis group or subgroup connected to statistics and machine-learning communities.Existing ATLAS and CMS structures are organized primarily around physics models, creating an organizational mismatch.
- Background estimation: Classifiers dependent on resonant features can sculpt artificial background bumps, complicating higher-dimensional anomaly searches.The LHC Olympics focused on resonant signals partly because sideband methods provide a natural background-estimation scheme.
- Background estimation: Non-resonant anomaly detection remains difficult because no general background-estimation approach exists, although data–simulation comparisons are promising in suitable final states.Combining anomaly-sensitive methods with reliable background strategies remains an open research problem.
- Online searches: Online anomaly detection could access phase space discarded by real-time triggers, but identifying events is insufficient without quantifying their level of strangeness statistically.In HEP, discoveries generally require an over-density in phase space rather than a declaration based on one collision.
- Interpretation: Model-agnostic searches make sensitivity difficult to quantify because they can respond to many models simultaneously, especially in high-dimensional spaces.This contrasts with model-dependent limits, which can be stated for a specified signal model and cross section.
- Interpretation: Reanalysis and reinterpretation require preserving both data-dependent event selection and the optimization procedure that produced it.Automated training and statistical analysis could support a RECAST-like workflow, but reproducibility must include optimization.
- Future of the challenge: The challenge lacked a single winner metric, preventing direct method comparison, and attracted few ML experts outside HEP without a broad platform or prize.Accessibility could also improve through more complete data-reading, jet-clustering, and dimensionality-reduction guidance.
- Future directions: Future Olympics could add signal models, final-state topologies, detector effects, and observables such as tracking and vertex information.The authors also point toward physics-informed learning and broad, theoretically motivated priors as possible routes to robust sensitivity.
8 Conclusions
The conclusions present the LHC Olympics as a community benchmark for developing anomaly-detection methods in realistic unlabeled collider-like data. The diverse results support continued development of complementary approaches rather than a single definitive algorithm.
- Conclusions: The LHC Olympics provided R&D data and three black boxes with different backgrounds and potential anomalies for testing anomaly-detection methods.Many teams implemented at least 18 methods in the challenge.
- Conclusions: Teams tested unsupervised, semisupervised, and supervised methods on black boxes before knowing their signal content.Several strategies were refined after the first black box was unveiled and continued toward collider-data applications.
- Conclusions: Methods have distinct advantages and disadvantages, share challenges such as scaling to higher dimensions, and may ultimately need to be combined.The authors state that further method development is required to reach the data’s full physics potential.
- Conclusions: The LHC Olympics is framed as a starting point for machine-learning-enabled exploration of high-dimensional collider data at the LHC and beyond.The conclusion connects the challenge to future physics results from current and future datasets.