Source-linked AI summary

Localizing Global Discrepancies: Marginal Contributions and Contextual Anomaly Detection

Tommaso dorigo

arXiv:2608.28375v1stat.MLcs.LGphysics.data-an

TL;DR

The paper addresses how to identify observations driving a global discrepancy and whether surrounding events contain additional class information. It unifies conditional and marginal localization, connects the construction to classical projections and MMD witnesses, and develops efficient estimators. The pair estimator recovers the direct empirical MMD witness with correlation 0.9993, while shared latent structure enables class information unavailable from isolated events.

  • Problem

    Global discrepancy statistics can detect sample-level departures without identifying which observations contribute to them, and context may or may not contain class information beyond an event’s features.

  • Method

    The paper assigns observations conditional or marginal contributions across random contexts and relates these scores to U-statistic projections, influence functions, MMD witnesses, and shared-latent ensemble information.

  • Results

    The pair estimator reaches correlation 0.9993 with the direct empirical MMD witness, while shared latent structure yields additional class information unavailable from an event in isolation.

  • Takeaways & Limitations

    Context serves two distinct roles: efficiently estimating a global-discrepancy localizer and supplying genuinely new class information when alternatives contain shared structure.

  • Takeaways & Limitations

    The framework’s extensions remain important for data-adaptive features, intrinsically higher-order discrepancies with vanishing first projections, nuisance-aware or learned references, and non-IID data.

Abstract

from arXiv · show

Global goodness-of-fit and discrepancy statistics can establish that a sample departs from a reference distribution without identifying which observations drive the departure. We develop a framework for this localization problem by assigning to each observation its conditional or marginal contribution across random statistical contexts. This connects resampling diagnostics and data valuation to projection theory and event-level anomaly detection. For symmetric statistics, fixed-size replacement is exactly equivalent to centered conditional localization. For U-statistics, the addition score equals the first Hoeffding/Hájek contribution; for smooth distributional functionals it is related at leading order to the influence function; and for unbiased known-background MMD it reduces exactly to the MMD witness. This viewpoint also yields more efficient estimators. Matched-context subtraction removes fluctuations unrelated to the observation, while for pairwise MMD the event-containing terms give a simple localizer. On the LHC Olympics anomaly-detection benchmark, the pair estimator converges to the direct empirical MMD witness with the predicted 1/(Rm^2) scaling, where m is batch size and R the number of batches. At m=1000 and R=5x106 it reaches correlation 0.9993 with essentially identical AUC. We also ask when context contains information beyond an event's own features. In a shared-latent toy model, the full single-event signal and background distributions are identical by construction, forcing isolated-event AUC=0.5. Discriminating information survives only in cross-event dependence induced by the shared latent parameter; the ensemble recovers this information, whereas an independent-latent control does not. This separates two roles of context: efficient localization of a global discrepancy and genuinely additional class information when the alternative contains shared structure.

1 Introduction

The paper distinguishes sample-level discrepancy detection from identifying which observations drive that discrepancy. It develops a unified localization framework based on conditional and marginal contributions across random contexts, then tests it in analytic, collective-information, and collider settings.

  • Motivation: Sample-level anomaly detection does not by itself identify the observations responsible for a distributional departure.An ordinary-looking event may belong to an anomalous population, while an isolated extreme event may not support a coherent discrepancy.
  • Framework: For symmetric statistics, replacing an ordinary member of a random batch by x exactly reproduces the centered conditional localizer.Matched-context subtraction also cancels fluctuations unrelated to the identity of x.
  • Framework: The framework interprets conditional localization and marginal contributions as event-level projections of a global discrepancy.It connects resampling diagnostics, data valuation, statistical projections, influence functions, and anomaly detection.
  • Framework: For MMD, the unbiased known-background estimator is a second-order U-statistic whose conditional localizer is proportional to the MMD witness, with analytically isolable event-containing terms.This provides an exactly tractable setting for comparing population localizers, finite-sample witnesses, and Monte Carlo estimators.
  • Empirical studies: The experiments separately validate localization, test collective information from shared latent structure, and measure efficient localizer recovery on collider-scale data.In the LHC Olympics benchmark, the pair estimator reaches correlation 0.9993 with the direct empirical witness and essentially identical AUC.

2 Related work

The paper situates discrepancy localization among established attribution, projection, influence, MMD, and anomaly-detection methods. Its contribution is a common formulation that also yields variance-efficient estimators and separates localization from collective class information.

  • Existing foundations: Marginal attribution, resampling diagnostics, influence measures, low-order projections, and MMD witnesses provide established antecedents for the framework.These methods differ in conditioning structure, objective, or statistical object.
  • Present formulation: The paper targets the expected contribution of a fixed observation across random contexts to a chosen global discrepancy from a reference distribution.This differs from ordinary deletion diagnostics and related subset-based attribution objectives.
  • Present formulation: The subset value is specialized to a sample-level discrepancy, with context size m treated as a localization and estimation parameter.The construction reframes data-valuation ideas around discrepancy localization rather than training-data valuation.
  • Statistical connections: For U-statistics, the local score is the first Hoeffding/Hájek projection, while smooth-functional localization connects asymptotically to the influence function.These connections provide the classical statistical structures underlying the event-level construction.
  • MMD and anomaly detection: Known-background MMD supplies an exactly soluble framework for testing finite-m event-containing estimators and their Monte Carlo scaling.The MMD identity is used as a tractable member of the broader localization framework.
  • MMD and anomaly detection: The paper extends the Inverse Bagging perspective by giving small-batch goodness-of-fit attribution a statistical target and separating raw conditional averages from lower-variance estimators.Its broader scope includes conditional subset scores, marginal contributions, projections, and event-containing terms.

3 From global discrepancy to local contribution

The section defines localization as an observation’s association with a sample-level discrepancy, expressed through conditional averages or marginal changes. Fixed-size replacement gives an exact marginal interpretation, while addition and replacement forms share the same population ranking.

  • Conditional localization: For a symmetric statistic T_m on batches of size m, conditional localization fixes one observation to x and compares the resulting expected statistic with its population expectation.The centering term is independent of x, so conditional and uncentered scores induce the same ranking.
  • Conditional localization: A conditional localizer measures association with a sample-level departure, not the posterior probability that x belongs to a signal class.It ranks observations by alignment with the ensemble’s discrepancy rather than by isolation under the reference model.
  • Marginal contributions: Fixed-size replacement marginal contribution is exactly the expected change when an ordinary member of a random batch is replaced by x.This establishes an exact marginal-contribution interpretation for any symmetric statistic with finite expectation.
  • Marginal contributions: Addition marginal contribution compares a batch containing x with the same context before x is added, and differs from conditional localization only by an x-independent term.For unbiased statistic families including U-statistics, the two forms preserve the same population ranking.
  • Estimation: Conditional localization and marginal contribution represent the same population quantity but can have different estimator variances.This distinction motivates matched Monte Carlo estimators and variance-efficient computation.

4 Statistical structure of local contribution

The framework identifies event-level contributions to global discrepancies through conditional localization and marginal constructions. For U-statistics this quantity is the first Hoeffding/Hájek projection, while for MMD it is exactly connected to the witness function; matched subtraction improves estimation efficiency.

  • U-statistics: A U-statistic averages a symmetric fixed-order kernel over distinct observation tuples and estimates the corresponding population expectation.
  • U-statistics: Conditioning on one observation separates kernel terms containing it from context-only terms, yielding an exact first-order localization proportional to r/m.The identity holds for every m ≥ r, and contextual averaging isolates the first-order component observation by observation.
  • U-statistics: Changing batch size m only rescales the population localizer, so event ordering is independent of m.If h1 ≡ 0, single-observation localization vanishes and pairs or higher-order groups are required.
  • Influence functions: For smooth functionals, contextual localization is asymptotically related to the influence function; for U-statistics, IFD(x; P) = rh1(x) gives the exact finite-sample counterpart.
  • MMD: For known-background MMD, the first Hoeffding projection is the centered witness function, so conditional resampling and the MMD witness produce exactly the same event ranking.The global squared discrepancy scales as O(ϵ2), whereas the witness and its class-mean separation scale as O(ϵ).
  • Efficient estimators: Near a degenerate null, suppressed first-order terms can produce pre-asymptotic O(m−1) behavior rather than the generic O(m−2) asymptote.

5 When does the ensemble contain additional class information?

The ensemble does not generally add class information once a fully specified IID mixture and an event’s features are known. Additional contextual information arises when observations share an unknown signal parameter, creating cross-event dependence that the ensemble can learn.

  • IID mixtures: Under a fully specified IID mixture, the posterior class probability satisfies P(Zi = S | Xi, X−i) = P(Zi = S | Xi).The Bayes-optimal event ranking is determined by s(x)/b(x).
  • IID mixtures: Consequently, the remaining sample contains no additional information about an event’s class beyond its own features under the complete IID model.
  • Shared latent structure: Genuinely contextual class information requires an unknown ingredient or dependence; the paper focuses on an unknown alternative parameter shared by all signal events.
  • Shared latent structure: When the shared parameter is unknown, other observations update its posterior and change the predictive signal density used to evaluate the target event.The ensemble contains additional class information when learning from X−i changes the relevant predictive signal model.
  • Shared latent structure: The ensemble recovers part, but never more than all, of the class information available if the hidden alternative were revealed by an oracle.
  • Independent-latent control: If each signal event has an independent latent parameter, marginalization restores IID draws and the other observations provide no information about the target event’s private parameter.
  • Gaussian toy model: In the shared-latent Gaussian construction, isolated observations contain no class information, while information survives in dependence between signal events.The paper states that no transformation of an isolated observation can distinguish its class.
  • Gaussian toy model: The relevant direction must be learned from the observed ensemble and localized back onto its members because no fixed experiment-independent event score solves the shared-latent problem.

6 Estimation from finite samples

Finite-sample localization estimates population event contributions from resampled contexts, with forced inclusion, global batches, and direct empirical witnesses providing complementary implementations. MMD offers a calculable convergence target, while leave-one-out, cross-fitting, and null calibration delimit reliable use.

  • Resampling estimators: Inverse Bagging generates small batches, evaluates a goodness-of-fit statistic, and assigns batch information back to participating events.In the large-resampling limit, the continuous score estimates the conditional average over contexts containing each event.
  • Resampling estimators: Forced-inclusion estimation places xi exactly once in each context and samples the remaining m−1 companions from the observed sample.Sampling with or without replacement differs only by a finite-population effect when m/N is small.
  • Resampling estimators: Global-batch and forced-inclusion constructions use different meanings of R, so their Monte Carlo scalings must not be compared without accounting for that distinction.Global batches also permit bootstrap multiplicities, whereas controlled forced-inclusion studies avoid this ambiguity.
  • Finite-sample limits: For fixed data, resampling error decreases as R−1/2, but increasing R cannot remove finite-N fluctuations of the empirical distribution.Estimator convergence on a fixed dataset is distinct from convergence of empirical localization to its population limit.
  • MMD estimation: For MMD, the infinite-resampling event ordering equals the empirical witness ordering, so explicit resampling is unnecessary for ranking but useful for validating stochastic estimators.
  • Self-influence control: Leave-one-out removes an event’s direct self-contribution and is sufficient when the kernel, representation, bandwidth, and background model are fixed independently.If scoring ingredients are estimated from the same data, cross-fitting or sample splitting is required to control indirect self-use.
  • Scalable approximations: Exact LOO kernel sums cost O(N2), motivating random Fourier features for scalable MMD estimation.After feature evaluation, all scores can be obtained in O(ND) time, with D controlling the computational–approximation trade-off.
  • Estimator comparison: Matched and event-containing estimators preserve population localization information while differing substantially in Monte Carlo efficiency.For MMD, the exact LOO witness is the finite-data target for measuring stochastic errors.

7 Numerical studies

Three experiments validate the localization identities, characterize finite-sample and Monte Carlo behavior, and test estimator efficiency in controlled and collider settings.

  • Experiment I: Experiment I validates population localization identities, finite-sample behavior, and variance cancellation for marginal and event-containing estimators.The study uses a fixed known alternative and separates population, finite-sample, and Monte Carlo effects.
  • Population identity and resampling: Explicit resampling reproduces the analytic conditional projection, with residuals decreasing at the expected R^-1/2 Monte Carlo rate.This confirms resampling recovers the conditional projection rather than defining an independent anomaly score.
  • Population identity and resampling: MMD localization responds to coherent distributional support rather than merely to low background density, assigning an extreme isolated background-tail event a predicted zero first-order score.The null test contrasts isolated improbability with contribution to a sample-level discrepancy.
  • Event ranking: The population MMD witness nearly matches Bayes-optimal event ordering, while the empirical witness suffers correlated finite-sample distortions that can reduce AUC to 0.691.For one representative pseudoexperiment, the exact likelihood-ratio oracle, population witness, and empirical LOO witness have AUC values 0.948, 0.947, and 0.935, respectively.
  • Batch size: The localization separation follows exact 1/m scaling, leaving population event ordering unchanged as batch size varies.The tested batch sizes are m = 5, 10, 20, 40, 80, 160, 320, 640.
  • Batch size: At fixed R = 3000, relative RMS error decreases from 0.396 at m = 5 to approximately 0.362 at m = 80, then rises through m = 640.The U-shaped behavior marks a crossover from near-degenerate second-order fluctuations toward the generic first-order Hoeffding regime.
  • Marginal and event-containing estimation: With R = 1000 contexts, raw-estimator uncertainty is approximately independent of m, whereas marginal and event-containing estimators decrease approximately as m^-1/2.The efficient estimators’ scaling follows from context variance O(m^-3) and localization signal O(m^-1).
  • Marginal and event-containing estimation: Both variance ratios decrease approximately as m^-1 because raw MMD variance remains approximately O(m^-2), while matched estimators achieve approximately O(m^-3) variance.This behavior differs from the generic m^-2 ratio expected after the raw statistic reaches its nondegenerate regime.

7.2 Experiment II: zero isolated-event information

The shared-latent experiment makes isolated-event classification impossible by construction, so any systematic discrimination must arise from information shared across observations. The contextual witness recovers this collective information, whereas the independent-latent control remains at chance.

  • Signal and background have identical one-event marginals, forcing every isolated-event classifier to AUC = 0.5.
  • The independent-latent control removes shared cross-event structure while preserving each signal event’s one-event marginal.
  • Increasing N helps only when additional observations constrain structure shared with the event being classified.
  • The empirical contextual witness recovers about 94% of the MMD classification advantage available when the hidden alternative is known.
  • Only the shared-context empirical witness improves systematically with ensemble size N; the independent-µi control remains consistent with chance.
  • For N = 5000 and m = 100, both Monte Carlo estimators approach the empirical LOO target at AUC = 0.7584, while the true-µ MMD witness reaches AUC = 0.8448.In the independent-µi control, the estimators converge to the accidental finite-sample LOO value AUC ≃0.5177, while the ensemble-average control remains consistent with chance.

8 Discussion

The paper unifies several localization constructions as ways to assign a sample-level discrepancy to observations, then shows how matched estimators improve computation and distinguish localization from genuinely collective information. It also identifies scope limits involving non-IID data, adaptive features, higher-order discrepancies, and post-selection inference.

  • A unified view of discrepancy localization: Deletion, insertion, marginal-contribution, conditional-projection, influence-function, and kernel-witness methods can all localize an observation’s association with a sample-level discrepancy.The framework places these established constructions in a common discrepancy-localization setting.
  • A unified view of discrepancy localization: For symmetric statistics, fixed-size replacement is exact, while unbiased U-statistic addition has the same expectation as conditional localization.For U-statistics, the localized quantity is proportional to the first Hoeffding/Hájek projection; for known-background MMD, it equals the witness function.
  • Computational consequence: canceling irrelevant context: Matched marginal subtraction removes context fluctuations, and U-statistic estimators can accumulate only event-containing terms.These constructions preserve the observation-specific contribution while avoiding fluctuations unrelated to its identity.
  • Computational consequence: canceling irrelevant context: The event-containing estimator gains the predicted additional power of m in absolute variance, including in the near-degenerate MMD regime.In that regime, the variance ratio follows an m^-1 law rather than the generic m^-2 asymptote.
  • Computational consequence: canceling irrelevant context: qρ = C(Rm^2)^-α with α = 0.9995 and C = 6.80 × 10^9, matching the predicted 1/(Rm^2) convergence behavior.At the largest sampled exposure, the pair estimator reaches correlation 0.9993 with the direct witness and essentially saturates its AUC.
  • Localization is not the same as collective information: Context can provide class information beyond isolated features only when the alternative contains shared structure; otherwise, localization and collective information remain distinct.The shared-latent ensemble recovers about 94% of the available MMD classification advantage at N=20000, whereas the independent-latent control remains at chance.
  • Limitations and statistical inference: The framework remains to be extended to nuisance-parameter or learned backgrounds, non-IID observations, adaptive features, and higher-order discrepancies requiring pair- or group-level localization.Localized scores are descriptive associations, not posterior signal probabilities; formal inference after selecting high-score events requires calibration or data splitting.
  • Computational consequence: canceling irrelevant context: When an analytic localizer exists, direct evaluation or random-feature approximations may be cheaper than explicit resampling.The marginal formulation is most valuable when no closed-form localizer is available.

9 Conclusions

The paper unifies event-level localization of global discrepancies, clarifying both the attributed quantity and efficient estimation. It also identifies when surrounding observations provide information beyond an event’s isolated features.

  • Fixed-size replacement makes conditional localization and marginal contributions coincide, providing a unified projection view of global departures onto observations.
  • Matched marginal differences cancel context fluctuations, while event-containing U-statistic terms isolate the relevant contribution more directly.
  • Correlation 0.9993 with the direct empirical MMD witness demonstrates that the pair estimator recovers the same event-level target more efficiently than raw batch averaging.
  • For shared-latent alternatives, the ensemble can provide class information absent from an individual event by constraining the latent structure.
  • The framework bridges global discrepancy detection and interpretable localization, with extensions proposed for degenerate discrepancies, nuisance-aware references, learned models, and non-IID data.
Loading 2608.28375v1…