Source-linked AI summary
Hadronic Mono-Z Dark Matter Sensitivity with Flow Matching on CMS Open Data
Hitesh Rasineni, Bhavishya Chebrolu
TL;DR
The paper studies projected dark-matter sensitivity in the hadronic mono-Z channel, where larger event rates accompany complex backgrounds and missing-object features. It models backgrounds with a conditional flow-matching continuous normalizing flow and applies held-out scoring with safeguards against preprocessing and optimisation biases. The baseline yields expected significances of 2.89σ, 7.62σ, and 7.41σ, while removing detailed extra-jet kinematics reduces sensitivity by 53–71%.
Problem
The study addresses dark-matter searches in the higher-rate hadronic mono-Z channel, whose complex multijet background and undefined extra-jet features require dedicated modelling.
Method
A conditional flow-matching continuous normalizing flow models the background, with sentinel imputation, held-out reweighted scoring, a signal-side offline trigger proxy, and a B ≥20 working-point floor.
Results
2.89σ, 7.62σ, and 7.41σ are obtained for three simplified-model benchmarks, while removing detailed extra-jet kinematics reduces expected significance by 53–71%.
Takeaways & Limitations
Extra-jet topology carries substantial discriminating power in the hadronic mono-Z channel beyond jet multiplicity alone.
Takeaways & Limitations
The results are projected expected sensitivities rather than observed excesses, and no signal-region unblinding was performed.
Abstract
from arXiv · showhide
We present a projected sensitivity study for hadronic mono-$Z$ dark-matter production using CMS Run~2015D HTMHT open data corresponding to 2.256382381~\invfb, from which 1{,}439{,}523 events satisfy the hadronic mono-$Z$ selection. Backgrounds are modelled with a conditional flow-matching continuous normalizing flow trained on the selected HTMHT events and evaluated on a held-out validation split reweighted to the full selected population. To mitigate artifacts from missing-object features and avoid in-sample scoring bias we apply sentinel imputation for undefined angular features, persist the train/validation split indices, and enforce a minimum reported background yield of 20 events when selecting the working point. A signal-side offline trigger proxy is applied to the simulated signal before scoring. Under this procedure the baseline analysis yields expected significances of 2.89$σ$, 7.62$σ$, and 7.41$σ$ for three simplified-model benchmarks. An ablation study that removes the detailed extra-jet kinematics reduces the expected significance by 53--71\%, indicating that extra-jet topology carries substantial discriminating power in the hadronic mono-$Z$ channel. These results are projected sensitivities (no unblinding performed); the limitations and reproducibility of the study are discussed in Sections limitations and reproducibility.
1 Introduction
The study extends density-based dark-matter searches to the higher-rate but more complex hadronic mono-Z channel using a continuous flow-matching background model. It addresses missing-object features, in-sample scoring, trigger emulation, and working-point selection while reporting projected sensitivities for three benchmarks.
- The hadronic mono-Z signature offers a larger Z branching fraction than the leptonic channel but faces more complex multijet backgrounds.
- The analysis replaces closed-form neural spline flows with a conditional flow-matching continuous normalizing flow for background-density modelling.
- Undefined extra-jet features are preserved as an explicit does-not-exist state rather than zero-imputed into the physical continuum.
- Held-out validation scoring, population reweighting, a signal-side offline trigger proxy, and B ≥20 working-point selection address key analysis biases and simulation limitations.
- 2.89σ, 7.62σ, and 7.41σ are the expected significances for the three simplified-model benchmarks under the post-fix baseline procedure.
2 Data and simulation
The study combines CMS Run 2015D HTMHT open data with simulated dark-matter signals for a hadronic mono-Z sensitivity analysis. It uses matched feature definitions, trigger handling, benchmark models, and leading-order signal normalization, while noting higher-order corrections as a scope limitation.
- Data selection: 2.256382381 fb−1 of CMS Run 2015D HTMHT open data were processed for the analysis.The dataset spans runs 256630–260627 and was traversed across 444 MINIAOD files.
- Data selection: The hadronic mono-Z selection requires two jets forming a candidate with 70 < mjj < 110 GeV and ΔRjj < 2.0, plus an additional-jet b-tag veto.Events pass one of five unprescaled trigger paths before offline selection.
- Signal simulation: The signal samples use MadGraph5 aMC@NLO, Pythia8, parameterised Delphes, and anti-kT jets for the hadronic-Z and extra-jet topology study.The anti-kT choice provides soft-resilient, conical jet boundaries for isolated hard particles.
- Signal simulation: The three retained benchmarks are s-channel spin-1 mediator models selected from the ATLAS/CMS Dark Matter Forum framework.They cover axial threshold and on-shell points and a vector on-shell heavy-mediator point.
- Normalization and scope: The benchmark signal rates use leading-order cross sections, which should be treated as lower estimates because expected NLO corrections are moderate, K ∼1.1–1.4.The cited NLO study indicates that the Emiss_T and jet pT shapes are largely preserved at NLO.
- Signal simulation: The simulated signal receives an offline proxy for the five data triggers before scoring, with 100% trigger-proxy efficiency after offline selection for every benchmark.Data and signal share a locked 41-branch feature schema despite different hadronic-Z candidate construction rules.
3 Methods
The analysis models the selected background with a conditional flow-matching continuous normalizing flow and evaluates event densities numerically. Its preprocessing, validation design, feature exclusions, and scoring procedure are structured to maintain consistent data–signal inputs and avoid in-sample background evaluation.
- Background model: The background density is modelled with a conditional flow-matching continuous normalizing flow mapping Gaussian noise to standardised feature vectors.Here, conditional refers to conditioning the training objective on the target data point, not on auxiliary covariates.
- Background model: The CNF evaluates densities by numerical ODE integration, trading closed-form density evaluation for a fully continuous-time model.The implementation uses a fixed-step RK4 solver with step size 0.05 and Hutchinson trace estimation.
- Training and evaluation: The vector field is trained with simulation-free conditional flow matching as mean-squared-error regression on velocities along a straight-line Gaussian-to-data probability path.The density objective is evaluated separately during inference through the probability-flow ODE.
- Preprocessing: 32 of the shared 41 features are used for training, while audit, trigger, veto, and data–simulation-incompatible columns are excluded.The exclusions include bookkeeping fields, trigger and b-tag fields, and variables with differing Delphes semantics.
- Preprocessing: Undefined extra-jet angular observables are mapped to the sentinel value −999 before standardisation rather than median-imputed into the physical continuum.The same sentinel handling is applied to background, validation, and signal samples.
- Training and evaluation: The implementation uses a residual MLP with sinusoidal time embeddings, four residual blocks, hidden width 192, and dropout 0.05.The architecture maps the time-conditioned input back to the 32-dimensional training feature space.
- Scoring: Background scoring uses only the persisted held-out validation split, reweighted to the full selected population, while signal is scored after the offline trigger proxy.Per-event negative log-likelihoods from Eq. (3) provide the background–signal discriminant.
4 Validation
Validation finds generally good closure for the fitted flow on held-out data, while sentinel imputation removes the dominant undefined-feature artifact. Residual mismodelling remains for continuous models representing zero-spike extra-jet pT features.
- 4.1 Closure plots (representative set): Held-out closure shows the flow reproduces the hadronic-Z mass peak and tracks validation data without order-of-magnitude mismatches.The Z-peak structure is reproduced inside the 70–110 GeV window, with a localized ∼50% discrepancy near mjj ≈91.6 GeV.
- 4.1 Closure plots (representative set): Good agreement is observed for EmissT, ∆RZjj, HT, reconstructed Z pT, hadronic recoil, EmissT/√HT, and |mjj − mZ|.These derived-observable closures indicate that the flow captures relevant Z-kinematic and recoil correlations used in the density score.
- 4.2 Before/after imputation-fix comparison: Median imputation of four undefined extra-jet angular features created a spurious point mass by merging the no-extra-jet state with physical values.The affected features are jet1 eta, jet2 eta, dphi met jet1, and dphi met jet2.
- 4.2 Before/after imputation-fix comparison: The −999 sentinel fix separates the no-extra-jet state from physical values and substantially reduces the resulting closure artifacts.Post-fix panels show an isolated sentinel bar and a smooth physical-value distribution, although some sentinel-bar mismatches remain.
- 4.3 Training convergence and validation NLL: Validation-NLL curves descend smoothly and plateau, with best checkpoints at epoch 245 for the baseline and epoch 165 for the ablation.The reported validation NLL values are approximately −64.97 for the baseline and −43.63 for the ablation.
- 4.4 Known residual limitation: jet pT delta-plus-continuum: Continuous normalizing flows cannot exactly represent a discrete zero-spike alongside a continuous extra-jet pT distribution.The jet1 pT closure retains a mild zero-spike undershoot, with flow density ∼0.071 versus data ∼0.142.
5 Sensitivity methodology
The analysis converts background-density scores into projected counting significances using an Asimov formula, with held-out background scoring, signal-yield normalization, and a constrained threshold scan. Sentinel closure checks and trigger-proxy treatment support the procedure, while continuous-flow smoothing remains a limitation for physical point masses.
- Asimov significance: The negative log-likelihood discriminant defines a counting region by selecting events above a threshold, where signal-like events populate the high-NLL tail.Signal and background yields above the threshold are denoted S and B.
- Validation and limitations: Closure panels reproduce the sentinel bar and physical continuum for undefined angular features, while extra-jet pT closure retains a smoothed pT = 0 spike.The latter is a residual limitation of representing a discrete point mass and continuous distribution with a continuous normalizing flow.
- Asimov significance: Expected significance is computed with the Asimov counting formula from weighted signal and background yields above the chosen NLL threshold.The result is a projected median significance based on simulated signal and the background model, not an observed-data comparison.
- Signal-yield normalization: The signal yield is normalized using the leading-order cross section, integrated luminosity, trigger-proxy efficiency, and generated-event count.The per-event signal weight is multiplied by the number of simulated signal events exceeding the threshold.
- Discriminant scan and working-point selection: The reported threshold is the best scan point satisfying B ≥20, preventing rare-tail fluctuations and look-elsewhere effects from determining the headline significance.The unconstrained optimum is retained only diagnostically; the background floor supports the asymptotic approximation.
- Background estimation: Held-out validation events are reweighted to the full selected background population, avoiding the positive bias from scoring training-set rows.The baseline contains 1,439,523 selected events and 287,905 held-out validation events.
- Signal trigger treatment: The signal-side offline trigger proxy retains 116, 487, and 848 events for the three benchmarks, with trigger-proxy efficiency 100% for each.The proxy uses HT, missing transverse momentum, and leading-jet requirements because Delphes does not emulate the CMS high-level trigger.
6 Results
The post-fix analysis reports projected Asimov sensitivities from a conditional flow-matching background model trained on the selected HTMHT sample and evaluated with held-out, reweighted background scores. The heavier benchmarks reach discovery-level projected sensitivity, while removing detailed extra-jet information substantially reduces significance; no unblinding was performed.
- Interpretation and scope: The study reports projected Asimov expected significances rather than an observed result from unblinded data.The calculation combines simulated signal with a background density model trained on held-out HTMHT data.
- Results setup: The study uses a conditional flow-matching background model trained on 1,439,523 selected events from the CMS Run 2015D HTMHT open dataset.The dataset corresponds to 2.256382381 fb−1 and was extracted from 444 MINIAOD files.
- Sensitivity procedure: The reported working point uses B ≥20, held-out validation events reweighted to the full population, and an offline trigger proxy applied before signal scoring.The unconstrained optimum is retained for comparison but excluded from the headline results.
- Baseline sensitivity: 2.89σ, 7.62σ, and 7.41σ are the expected significances for the light axial, heavier axial, and vector benchmarks, respectively.The heavier axial and vector benchmarks provide projected discovery-level sensitivity, while the light axial point is moderately sensitive.
- Baseline sensitivity: B ≃25 weighted background events is the working point for all three benchmarks, with relative Poisson uncertainty 1/√B ≃0.20.Signal efficiency at the working point is 0.20–0.25.
- Working-point robustness: The unconstrained diagnostic optima reach approximately 4.96, 10.81, and 10.70 but are excluded because they occur at very small background yields and are unstable under validation fluctuations.These values are not the principal reported results.
7 Robustness: extra-jet feature ablation
Removing detailed extra-jet kinematics substantially weakens discrimination across all three benchmarks, with the largest loss for the light axial point. The retained extra-jet information therefore provides important sensitivity in the hadronic mono-Z analysis.
- Validation: The ablation is not explained by undertraining: closure validation shows every retained feature closes as well as in the baseline.The ablation was trained from scratch with the same setup, and its best checkpoint occurred at epoch 165.
- Ablation result: 53–71% relative drops in expected significance occur across all three benchmark points after removing six detailed extra-jet features.The ablation reduces the feature dimension from D = 32 to D = 26 while retaining the same architecture, hyperparameters, and held-out scoring procedure.
- Ablation result: 0.83σ replaces 2.89σ for the light axial benchmark after ablation, with the B ≥20 floor selecting B ≃50 at lower signal efficiency.The higher background yield and lower signal efficiency indicate that the signal NLL distribution has lost separation from the background.
- NLL separation: The axial mχ = 10 GeV, mmed = 20 GeV signal nearly overlaps the background NLL peak in the ablation model, collapsing its significance below 1σ.The other two benchmarks retain a visible, though reduced, tail beyond the background peak.
- Interpretation: Detailed extra-jet kinematics carry genuine discriminating information, whereas the extra-jet count alone does not recover the lost power.The study attributes the likely source to topology differences in initial-state-radiation jets between signal and generic multijet HTMHT background.
8 Limitations
The reported sensitivities are constrained by projection-only inference, approximate signal triggering and detector simulation, omitted background systematics, and unresolved signal-side diagnostics. A continuous density model also cannot exactly represent the padded extra-jet pT mixture.
- Scope of inference: The results are projected expected significances, not observed excesses, because no signal-region unblinding was performed.Background scores come from held-out real data, while signal and background yields are combined for the expected calculation.
- Trigger treatment: The offline trigger proxy has 100% efficiency for all three benchmarks and may overestimate true trigger acceptance for lighter signals.A complete reinterpretation would require trigger-efficiency parameterization and detector systematic uncertainties.
- Uncertainties: Only statistical Poisson background uncertainty is propagated; uncertainties from the density model, feature choices, ODE tolerance, and finite validation data are omitted.The study is an expected-sensitivity calculation rather than a full profile-likelihood limit-setting analysis.
- Detector simulation: Delphes signal simulation omits pileup overlay and uses b-tag semantics incompatible with data, so corresponding features are excluded from training.A fuller treatment would require pileup overlay and a compatible b-tag discriminator output.
- Open limitations: The light axial benchmark has an unresolved per-stage signal cutflow diagnosis, and continuous flows cannot exactly represent the padded extra-jet pT point mass.The latter remains a structural caveat despite sentinel imputation fixing the largest undefined-angular-feature artifact.
9 Conclusion
This study applies a conditional flow-matching density model to hadronic mono-Z dark-matter sensitivity using CMS open data and finds substantial baseline projected sensitivity. The conclusion emphasizes extra-jet topology while identifying several steps needed for a fuller experimental interpretation.
- Conclusion: 1,439,523 selected background events from 2.256382381 fb−1 of CMS Run 2015D HTMHT data are modeled with a conditional flow-matching continuous normalizing flow.Per-event negative-log-likelihood scores discriminate background from simulated dark-matter signal.
- Conclusion: 2.89σ, 7.62σ, and 7.41σ are the post-fix baseline expected significances for the three retained simplified-model benchmarks.The benchmarks are axial mχ = 10/mmed = 20, axial mχ = 50/mmed = 200, and vector mχ = 1/mmed = 500 GeV.
- Conclusion: 53–71% reductions in expected significance after removing detailed extra-jet kinematics indicate substantial discriminating power beyond jet multiplicity alone.The interpretation is plausibly connected to initial-state-radiation topology differences in the hadronic mono-Z channel.
- Methodological contribution: The work replaces the earlier leptonic analysis’s closed-form flow and likelihood-ratio strategy with continuous flow matching and a held-out background-density discriminant.The revised approach avoids the high-Emiss T tail-modelling residual identified in the earlier leptonic interpretation.
- Future work: Future work requires full CLs or profile-likelihood treatment, background systematics, data-driven trigger efficiencies, pileup overlay, cutflow diagnosis, and direct flow-model comparison.These extensions define the boundary between the current projected study and a fuller experimental reinterpretation.
Declarations
The paper declares no funding or conflicts and relies on publicly available CMS Open Data. It also documents the trigger, selection, feature schema, audit counts, and excluded training variables used for reproducibility.
- Declarations: No funding was received, and the authors declare no competing interests; analyzed datasets are publicly available CMS Open Data.Ethics, consent to participate, and consent to publish declarations are not applicable.
- Trigger configuration: Five primary HLT paths define the background trigger sample and the signal-side offline trigger proxy.The listed paths combine PFMET/PFMHT requirements with a MonoCentralPFJet path.
- Trigger audit: 91.3% of 20,679,437 raw events pass at least one HLT path, with trigger decisions reconstructed from filter labels because pathNames is often empty.The audit identifies 18,875,596 passing events through the filter-label fallback.
- Feature schema: The 41-branch schema is reduced to 32 training features by excluding bookkeeping, trigger, b-tag mismatch, and pileup-mismatch columns.The schema and extraction formulas are tabulated in Table 5, while the offline selection parameters are documented separately.
C.1 Data-side cutflow
The data-side cutflow applies the unprescaled trigger paths and hadronic mono-Z requirements, while the accompanying material documents closure checks across the selected observables. The signal-side yields are reported, but their per-stage counters remain unreviewed.
- Data-side selection: The data-side selection applies five unprescaled HLT paths followed by jet, η, m_jj, ΔR_jj, and extra-jet b-tag requirements.The overall efficiency is defined as the final selected population divided by the raw event count.
- Signal yields: 116, 487, and 848 events are the final offline post-selection yields for the three signal benchmarks.These yields are reported before trigger-proxy application and NLL scoring.
- Signal cutflow: The per-stage signal-generation counters have not been reviewed, so the signal cutflow is not tabulated.The reported final yields remain available despite this review gap.
- Closure diagnostics: The closure documentation covers missing-momentum, angular, recoil, Z-daughter, extra-jet, and jet-multiplicity observables.Figures 8–14 collect the baseline closure plots for these observable groups.
E Reproducibility and code availability
The study makes its analysis configuration, processing records, split indices, logs, diagnostics, and outputs available with code. Reproducibility is further specified through explicit software, training, solver, and imputation settings, while signal cutflow counters remain under review.
- Code availability: The accompanying analysis code provides the full configuration, per-file counts, train/validation indices, logs, closure diagnostics, and sensitivity outputs.The implementation uses Python 3.12 with uproot, awkward, numpy, PyTorch, and torchdiffeq, while signal samples use MadGraph5_aMC@NLO.
- Training configuration: SEED=20260730, batch size 2048, learning rate 2e-4, and maximum epochs 500 define key flow-matching training settings.The configuration also specifies early stopping, validation-NLL checks, an rk4 solver, and density-batch size 16.
- Numerical configuration: The sentinel imputation value is -999, with rk4 step size 0.05 and solver tolerances of 10^-3.The configuration uses Hutchinson divergence estimation for density modelling.
- Reproducibility caveat: The signal-generation per-stage CSV counters have not been reviewed to diagnose the lower raw selection efficiency of the axial benchmark.The final offline yields are 116, 487, and 848 events before trigger-proxy application and NLL scoring.