Source-linked AI summary
Testing for Outliers with Conformal p-values
Stephen Bates, Emmanuel Candès, Lihua Lei, Yaniv Romano, Matteo Sesia
TL;DR
The paper studies nonparametric outlier testing with many new observations, where false-discovery control is complicated by dependence among conformal p-values. It develops conformal methods using calibration data and high-probability bounds, showing both dependence-related failures and stronger conditional guarantees. The resulting tools support multiple testing and confidence bounds for false-positive rates, with stronger guarantees sometimes reducing power.
Problem
The problem is testing which new observations are outliers while controlling false discoveries when many conformal p-values share calibration data.
Method
The paper uses conformal p-values built from black-box scores, positive-dependence analysis, and high-probability bounds to obtain calibration-conditional procedures.
Results
20.5% marginal type-I error and 71.7% 90th-percentile conditional type-I error occur at α = 5% and γ = 3 for Fisher’s combination test.
Takeaways & Limitations
The methods provide finite-sample nonparametric outlier inferences, independent conditional p-values for multiple testing, and uniform false-positive-rate bounds for raw-score thresholds.
Takeaways & Limitations
Marginal conformal p-values may be anti-conservative for particular calibration sets, and conformal p-values cannot be smaller than 1/(n + 1).
Abstract
from arXiv · showhide
This paper studies the construction of p-values for nonparametric outlier detection, taking a multiple-testing perspective. The goal is to test whether new independent samples belong to the same distribution as a reference data set or are outliers. We propose a solution based on conformal inference, a broadly applicable framework which yields p-values that are marginally valid but mutually dependent for different test points. We prove these p-values are positively dependent and enable exact false discovery rate control, although in a relatively weak marginal sense. We then introduce a new method to compute p-values that are both valid conditionally on the training data and independent of each other for different test points; this paves the way to stronger type-I error guarantees. Our results depart from classical conformal inference as we leverage concentration inequalities rather than combinatorial arguments to establish our finite-sample guarantees. Furthermore, our techniques also yield a uniform confidence bound for the false positive rate of any outlier detection algorithm, as a function of the threshold applied to its raw statistics. Finally, the relevance of our results is demonstrated by numerical experiments on real and simulated data.
1 Introduction
The paper frames outlier detection as testing new observations against a reference distribution while controlling false discoveries across many tests. It develops conformal p-values around black-box one-class scores, addressing dependence and weak calibration-conditional guarantees.
- Problem statement and motivation: Outlier detection tests which new observations were not drawn from the reference distribution, using clean i.i.d. data and applications including diagnostics, fraud detection, and intrusion monitoring.
- Problem statement and motivation: Conformal inference wraps a one-class classifier’s black-box score function to produce finite-sample-valid, nonparametric p-values for inlier hypotheses.Smaller learned scores are treated as evidence that a point may be an outlier, while calibration scores estimate the unknown score distribution.
- Problem statement and motivation: Multiple testing matters because large outlier screens can generate costly false discoveries, making FDR control a meaningful error criterion.The paper defines FDR as the expected proportion of true inliers among reported outliers.
- Preview of contributions: Classical conformal p-values are marginally valid but dependent across test points because they share the same random calibration data.Conditioning on a particular calibration set can make them anti-conservative, so average validity may be unsatisfactory for practitioners using one calibration set.
- Preview of contributions: The paper introduces calibration-conditional validity, yielding p-values independent conditional on calibration data and suitable for downstream procedures that assume independence.The stronger guarantee is intended to provide confidence for a particular calibration set, rather than only average validity across researchers or datasets.
- Preview of contributions: The results also provide a uniform upper confidence bound for an outlier detector’s false positive rate as a function of its raw-score threshold.Numerical experiments on simulated and real data compare marginal and calibration-conditional methods, showing stronger guarantees can reduce power.
2 Marginal conformal inference for outlier detection
Marginal conformal p-values are dependent across test points, invalidating naive Fisher aggregation but retaining useful structure for FDR control. PRDS-based results justify BH control, while Storey’s correction restores control at the nominal level despite positive dependence.
- Marginal validity and dependence: Marginal conformal p-values are valid for individual inlier hypotheses but are dependent across test points through the shared calibration data.This dependence can make them anti-conservative conditional on the observed data set.
- Global testing: Fisher’s combination test can fail for marginal conformal p-values because their positive dependence inflates the combination statistic’s variance.The failure is characterized asymptotically when the number of test points is proportional to the calibration-set size.
- Global testing: For α = 5%, Fisher’s marginal type-I error reaches 20.5% when γ = 3 and approaches 50% as γ →∞.The conditional 90th-percentile type-I error reaches 71.7% when γ = 3 and approaches 100% as γ →∞.
- FDR control: Marginal conformal p-values are PRDS on the set of inliers when inlier test points are jointly independent of one another and of the observed data.This dependence structure supports applying BH rather than treating the p-values as independent.
- FDR control: The BH procedure controls FDR at π0α, where π0 is the proportion of true nulls, under the PRDS result.This control is marginal because the guarantee averages over both calibration data and future test points.
- FDR control: Storey’s BH correction controls FDR at the target level α for marginal conformal p-values, despite their positive correlation.The result uses λ = K/(n + 1) for an integer K and applies under the stated continuity assumptions.
3 Calibration-conditional conformal p-values
The paper adjusts marginal conformal p-values to obtain validity conditional on calibration data, using simultaneous bounds on order statistics. It compares Simes, DKWM, asymptotic, and Monte Carlo adjustments, showing distinct finite- and large-sample power trade-offs across testing settings.
- 3.2 A generic strategy to adjust marginal conformal p-values: Marginal conformal p-values can be anti-conservative conditional on calibration data, motivating adjustment functions that satisfy calibration-conditional validity.The required guarantee is uniform rather than pointwise in the testing level.
- 3.3 Simes adjustment of marginal conformal p-values: The generalized Simes inequality provides a simultaneous order-statistic bound whose adjustment yields calibration-conditional valid p-values.For δ = 0.1 and n = 1000, the smallest marginal p-value is mapped to approximately 0.0046, versus a much larger DKWM correction.
- 3.5 Monte Carlo adjustment of marginal conformal p-values: The Monte Carlo adjustment is finite-sample valid, preserves very small p-values relatively closely, and tracks the asymptotic envelope for larger p-values.It is designed to combine the finite-sample validity of Simes with the large-p-value behavior of the asymptotic approach.
- 3.6.1 Testing a single hypothesis: For a single hypothesis, the Monte Carlo adjustment approaches the asymptotically efficient method as n grows and can be more powerful when n is small.The Simes adjustment behaves similarly at small sample sizes but does not reach the nominal level in the large-n limit.
- 3.6.2 FWER control with a single strong signal: For FWER control with one strong signal, Simes is more powerful asymptotically, whereas Monte Carlo has nearly the same power as the asymptotic correction.The DKWM adjustment incurs a large additive inflation and corresponding power loss.
4 Extensions beyond conformal p-values
The paper extends its conformal framework to uniform false-positive-rate confidence bounds and stronger conditional validity for prediction sets. These extensions provide threshold-specific FPR guarantees and simultaneous coverage guarantees conditional on calibration data.
- 4.1 Simultaneous confidence bounds for the false positive rate: Uniform confidence bounds upper-bound the false positive rate of any one-class classifier as a function of its detection threshold.For score threshold t, the bound is h(F̂_n(z)), where F̂_n(z) is the calibration-score empirical CDF.
- 4.1 Simultaneous confidence bounds for the false positive rate: 90% confidence bands are designed to lie above the true FPR curve across reporting thresholds, with empirical FPR shown separately.Figure 6 compares adjustment methods for an isolation forest classifier on simulated data; the right panel zooms into small score values associated with likely outliers.
- 4.1 Simultaneous confidence bounds for the false positive rate: The DKWM confidence bound is loose near zero, motivating adjustment functions that remain close to the identity for small empirical CDF values.This matters because small p-values are the most relevant for detecting outliers and multiple-testing rejections.
- 4.2 Extensions to conformal prediction: CCV p-values yield prediction sets that are simultaneously valid for every α conditional on the calibration data.With high probability, a new observation lies in the set with probability at least 1 − α for all α ∈ (0, 1), unlike the usual single-level marginal guarantee.
5 Numerical experiments
The experiments compare marginal and calibration-conditional conformal p-values for individual and batch outlier detection on simulated and real data. Calibration-conditional methods generally control conditional FDR for at least 90% of practitioners, while marginal methods control FDR only marginally and can lose validity under batch testing.
- Experimental setup: The experiments evaluate conditional FDR, marginal FDR, and power across independent practitioners and future test scenarios.Performance is defined conditional on each practitioner’s training and calibration data, with test sets treated as random future scenarios.
- Individual outlier detection: In individual detection, calibration-conditional p-values control conditional FDR for at least 90% of practitioners, whereas marginal p-values do not.All methods control marginal FDR; Monte Carlo and Simes conditional calibration yield slightly higher power than the asymptotic approximation.
- Batch outlier detection: In batch detection, marginal p-values can produce noticeably excessive conditional FDR, while simultaneous calibration is conservative without much power loss.The simulated setting has power close to 1; Monte Carlo and asymptotic conditional calibration have higher power among the conditional alternatives.
- Batch outlier detection: Under the global null, marginal p-values become especially invalid as batch size increases, while calibration-conditional tests remain valid.The global null is tested by combining batch p-values with Fisher’s method.
- Real and benchmark data: On credit-card and benchmark data, simultaneous calibration controls conditional FDR for at least 90% of applications, but incurs some power loss.Both calibration methods control marginal FDR.
6 Discussion
The discussion positions conformal p-values as finite-sample, model-agnostic tools for multiple outlier testing, while emphasizing dependence and power limitations. The proposed high-probability calibration yields mutually independent conditional p-values, and the paper identifies several extensions for future work.
- Discussion: Conformal inference can wrap black-box one-class classifiers to produce finite-sample-valid, nonparametric outlier p-values under clean training data and an i.i.d. assumption.Its agnosticism avoids requiring accurate models but limits how small conformal p-values can be.
- Discussion: Conformal p-values cannot be smaller than 1/(n + 1), which may reduce power relative to likelihood-based methods when signals are strong but sparse.The paper presents this as a limitation or strength depending on whether clean data or accurate models are available.
- Discussion: Mutual dependence among marginal conformal p-values can invalidate Fisher’s combination test and complicate the validity of procedures such as BH.The paper highlights positive dependence through its PRDS result, whose practical verification is often difficult.
- Discussion: High-probability bounds produce calibration-conditional conformal p-values that are mutually independent and can be used directly in multiple-testing procedures.The bounds are simultaneous and may also support a posteriori significance-threshold tuning.
- Discussion: Future work includes alternative hold-out methods, relaxing i.i.d. assumptions, connections to two-sample testing, and applications of the bounds beyond p-value calibration.The theory requires calibration and test inliers to be exchangeable and mutually independent, while test outliers may depend on one another.
A Technical proofs
The technical proof analyzes dependence among marginal conformal p-values through rank permutations under the global null. It establishes the joint rank probabilities needed to characterize the inflated variance of combination statistics relative to independent p-values.
- Dependence structure: Under the global null, conformal p-values for multiple inlier test points are exchangeable.The proof represents score ranks as a uniformly distributed permutation when scores are independent draws from a non-atomic distribution.
- Variance consequence: For any transformation G with finite second moment, dependence inflates the combination-statistic variance by a factor of 1 + γ relative to independent p-values.The proof compares Var[G(p_1)] with the covariance structure induced by shared calibration data.
- Rank calculations: The probability that two test ranks take adjacent values is twice the probability of a specified non-adjacent ordered pair.These probabilities follow from uniformity over all permutations of the ranks.
A.2 Failure of type-I error control with combination tests
The section analyzes why combining marginal conformal p-values can fail to control type-I error, despite marginal validity, because calibration induces dependence and conditional anti-conservatism. It develops asymptotic corrections and establishes positive dependence properties relevant to alternative multiple-testing procedures.
- Combination-test analysis: Theorem 6 derives asymptotic type-I-error behavior for general adjusted combination tests under a global-null regime with m = ⌊γn⌋.The proof uses conditional normal approximations, Berry–Esseen bounds, and concentration arguments.
- Combination-test analysis: For Fisher’s combination test, G(u) = −2 log u, and its regularity conditions support the general combination-test analysis.The paper notes that G(U) follows χ2(2) and verifies the required monotonicity conditions.
- Adjusted tests: Choosing ξ = √(1 + γ) makes the limiting marginal type-I error equal to α, while a larger adjustment controls conditional error with probability at least 1 − δ asymptotically.The conditional guarantee uses ξ = 1 + √γz1−δ/z1−α.
- Adjusted tests: Monte Carlo simulations compare unadjusted ξ = 1 and adjusted ξ = √(1 + γ) Fisher tests across γ values using n = 10^5 and 10^4 samples.Figure A1 reports simulated and asymptotic type-I errors for both procedures.
- Dependence structure: Randomized marginal conformal p-values remain PRDS when conformity scores are not continuously distributed.This extends the positive-dependence result beyond the continuous-score case.
A.4 Storey’s correction does not break FDR control
This section studies Storey’s correction under PRDS conformal p-values and derives a bound showing that the correction preserves false discovery rate control. It also relates the analysis to calibration-conditional p-value adjustments and uniform confidence constructions.
- Storey correction: Storey’s procedure uses ordered p-values, a target FDR level α, and tuning parameter λ ∈ (0, 1) to define its rejection set.The paper notes that λ is often chosen as 0.5, α, or 1 − α.
- PRDS FDR control: The paper introduces a novel FDR bound for Storey’s procedure when the p-values are PRDS.The argument assumes null p-values are super-uniform and have an almost-sure lower bound pmin.
- PRDS FDR control: The derivation conditions on rejection-related quantities and uses super-uniformity of each null p-value to bound the relevant error contribution.The proof organizes possible values of the Storey ratio into a finite set.
- PRDS FDR control: The proof exploits PRDS monotonicity together with the decreasing behavior of the factor 1/(1 + A) in the p-values.This supports the conditional expectations needed for the FDR bound.
- Calibration-conditional adjustments: Calibration-conditional conformal adjustments are analyzed through stochastic dominance of transformed calibration-score order statistics.The resulting construction yields a uniform upper confidence band for the score distribution.
- Calibration-conditional adjustments: The section evaluates Fisher’s combination test applied to calibration-conditional conformal p-values under several adjustment functions.The experiments use k = 0.5n for the relevant adjustment.
B.1.3 Effective α-level
This section characterizes the effective α-level of Fisher’s combination test after calibration-conditional adjustment, focusing on regimes where the number of test points scales with or grows more slowly than the calibration size.
- Approximation strategy: The effective-level analysis is derived from a distributional approximation for Fisher’s combination statistic.The section states that the proof of this approximation is involved and is provided elsewhere in the paper.
- Linear regime: Theorem 9 analyzes the type-I error when m = γn for fixed γ ∈ (0, ∞).The result gives an asymptotic approximation for Fisher’s test applied to the adjusted p-values.
- Linear regime: In the linear regime, the effective type-I error is determined by γ, α, and the adjustment’s asymptotic scaling.The displayed approximation includes e(γ) = π^2γ/[4(1 + γ)].
- Sparse regime: Theorem 10 treats the sparse regime m → ∞ with m = o(n / log log n).It gives a separate asymptotic type-I-error characterization for this growth rate.
B.1.4 Proof of Lemma 8
The proof of Lemma 8 establishes a high-probability approximation for Fisher’s statistic conditional on the calibration data, combining concentration, moment bounds, and normal-approximation arguments.
- Concentration bounds: The argument uses sub-exponential concentration and Bernstein inequalities to control conditional moments and sums.The proof identifies variables with sub-exponential parameters (2, 2).
- Conditional statistic: The proof treats transformed p-values through G(p) = −log(h_a(p)) and analyzes conditional sums of these variables.The notation suppresses the global-null conditioning throughout the argument.
- Normal approximation: A Berry–Esseen bound controls the Kolmogorov distance between the conditional statistic and a normal distribution.One displayed bound is 16CA/√m + 6C log^6 n/√n on the high-probability event.
- Normal approximation: The proof also compares conditional variances and normal approximations using Wasserstein-distance bounds and coupling arguments.These steps combine triangle inequalities with moment and inverse-Gamma calculations.
- Proof assembly: The resulting approximation is assembled from the preceding concentration and moment estimates to complete Lemma 8.The proof explicitly combines the intermediate bounds before concluding.
B.2.4 Effective α-level
The section analyzes the effective type-I error of Fisher’s combination test when applied to marginal conformal p-values under specified asymptotic regimes. The results provide bounds involving sample sizes, logarithmic factors, and constants depending on model parameters.
- Theorem 12 considers Fisher’s combination test applied to adjusted marginal conformal p-values under regimes relating m and n.The stated cases include m = γn and m →∞ with m = o(n/log^2 n).
- Theorem 13 assumes k = ⌈ζn⌉ for a constant ζ > 0 and establishes a corresponding asymptotic result.
- Lemma 16 provides a finite-sample bound involving a universal constant and a threshold n0 that depends only on δ.The bound applies for n ≥ n0(δ).
- Lemma 17 gives another global-null bound with terms involving n, m, t, and logarithmic factors.
- Theorem 14 assumes m/log n →∞ and analyzes the type-I error of Fisher’s combination test applied to another adjusted marginal-p-value construction.Its bound includes a constant depending on ζ and δ.
C Numerical comparisons of different adjustment functions
The section compares adjustment functions based on generalized Simes, DKWM, and Dempster boundary-crossing bounds. The generalized Simes adjustment is best in most examined settings, while DKWM becomes tighter for large n and moderately larger marginal p-values.
- The comparison includes generalized Simes, DKWM, and Dempster exact linear-boundary crossing adjustment functions.The Dempster construction uses a boundary-crossing probability and linear boundaries admit analytical treatment.
- The generalized Simes adjustment function is best in most scenarios examined.
- When n = 10000 and the marginal conformal p-value exceeds 0.03, the DKWM bound is tighter.The comparison focuses on small marginal p-values within [0, 0.05].
- For multiple testing, p-values above 0.03 would rarely be expected to be significant.
- The figures compare adjustment functions at n = 1000 and n = 10000 with δ = 0.1.Figures A2–A4 report these parameter settings.
D.1 Outlier detection on simulated data
The simulated-data section evaluates outlier detection through FDR and power across sample sizes, calibration procedures, p-value calibration methods, signal strengths, and global-testing combinations.
- Figure A5 reports FDR and power as functions of the number of samples, using half the data for calibration.
- Figures A6–A8 evaluate simulated outlier detection with the BH procedure using Storey’s correction.Figure A8 applies conditional calibration with δ = 0.25 instead of δ = 0.1.
- Figure A9 evaluates simultaneously calibrated conformal p-values as a function of the Simes parameter n/k at signal strength 2.
- Figure A10 compares p-value calibration methods across different signal strengths in a simulated outlier batch detection problem.
- Figure A11 compares methods for combining p-values from the same batch for global testing.
D.2 Outlier detection on real data
The real-data section reports outlier and outlier-batch detection performance across data sets, machine-learning models, nominal FDR levels, and alternative BH corrections.
- Tables A1 and A2 report outlier detection performance on real data with different data sets, models, and nominal FDR levels.Table A1 uses Storey’s correction, whereas Table A2 omits it.
- The reported tables include continuation pages for the real-data performance results.
- Tables A3 and A4 report outlier batch detection performance on real data using Storey’s correction.The tables vary data sets, machine-learning models, and nominal FDR levels.