Source-linked AI summary
Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets
Jia Huang, Yankai Wan, Yangjun Ou
TL;DR
Detector-defined datasets can have precision governed by deployment-pool prevalence rather than detector quality alone, but direct measurement is often unavailable because false positives are hidden inside accepted data. This paper uses an independent official index to identify phantoms across three pools, then derives exact contamination and estimator results. It finds that naive precision transfer fails badly, Bayes predictions remain accurate, contamination forms a structured phantom signal, and estimator choice can reverse its apparent direction.
Problem
False positives inside detector-defined datasets are difficult to identify, leaving limited evidence about how pool prevalence governs precision after deployment.
Method
The paper compares one magnitude-threshold detector across three candidate pools using an independent official index to classify detected items as real or phantom.
Results
Naive cross-pool precision transfer errs by +422%, whereas the Bayes expression predicts measured precision within 3.3%; contamination also forms a structured phantom component and changes direction across estimators.
Takeaways & Limitations
Contamination is a second detector-inherited signal rather than additive noise, and its effect cannot be assumed to attenuate conclusions because estimator choice can reverse the direction.
Takeaways & Limitations
The paper supports relative conclusions between datasets and estimators, not absolute statements about response speed.
Abstract
from arXiv · showhide
Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom component (residual 1.1e-16), with phantoms outnumbering true events 473 to 308, so contamination is a second signal with detector-inherited shape, not additive noise. Contamination direction depends on the estimator: on identical windows one statistic is diluted and another inflated because its denominator is also contaminated. A common normalization turns the estimator into a mean of ratios whose expectation need not exist; on the same 335 events it returns 0.40 where the well-defined estimator returns 0.10.
1 Introduction
Detector-defined datasets inherit precision from the prevalence of true events in the deployment pool, making cross-pool transfer unreliable. The paper measures this failure with recoverable phantoms and shows that contamination can alter conclusions differently across estimators.
- Prevalence-driven contamination: Precision depends on pool prevalence π, while detector operating characteristics TPR and FPR transfer across pools.Validation at π ≈1 provides little information about false-positive behavior at low prevalence.
- Prevalence-driven contamination: 81.7% / 9.0% / 0.0% phantom rates across three pools made naive precision transfer err by +422%, whereas Bayes predictions stayed within 3.3%.The study uses one detector across structurally different pools and measures against an independent official index.
- Structured contamination: A convex-mixture decomposition separates contaminated estimates into true-event and phantom components, with the phantom component reproducible from the detector’s acceptance rule.This establishes contamination as a structured component rather than unstructured additive noise.
- Estimator dependence: Contamination dilutes a denominator-free statistic but inflates a terminally normalized statistic on identical windows.The direction changes because contamination affects the estimator’s own scale.
- Estimator dependence: Per-item normalization produces a mean of ratios with no finite expectation and returns 0.40 versus 0.10 for the well-defined estimator on the same events.The pathology can yield a stable-looking but non-convergent statistic.
2 Setup
The setup distinguishes externally observed events from detector-accepted windows and defines the estimators used to measure their response curves. It also identifies prevalence, detector operating point, and contamination intensity as key quantities for interpreting detector-defined datasets.
- Observed and detected indices: A detector-defined index is formed by applying a deterministic accept/reject rule A to candidate windows C, while users typically observe only the accepted count.The setup separates the candidate pool, detector, and resulting detected index.
- Reporting quantities: Five quantities are proposed for reporting alongside detector-defined datasets, including deployment prevalence and the detector operating point.The supplied passages specifically identify prevalence and operating-point components.
- Reporting quantities: Precision equals 1 − c, while the contamination intensity ρ measures phantom loudness relative to real events and jointly determines downstream damage with c.Contamination fraction alone does not characterize the estimator’s distortion.
- Response curves: The response-curve setup compares denominator-free S, pooled-scale D, and per-item normalized eD estimators.S is reported in basis points, D uses a pooled scale, and eD normalizes each item by its reference return.
3 Theory
The theory shows that detector contamination has an exact mixture structure and is not identifiable from the contaminated dataset alone. It also explains why contamination can resemble a real signal and why per-item normalization may lack a finite expectation.
- Prevalence and deployment: High-prevalence validation carries little information about low-prevalence deployment because precision is determined by π through Bayes despite fixed TPR and FPR.At π ≈1, even a poor detector can appear nearly perfect.
- The mixture identity: The mixture relation is an identity for any detector, horizon, and data, but users holding only the contaminated curve cannot invert it.Knowing the contamination fraction, relative intensity, and component curves would identify the published curve; those quantities are unavailable internally.
- The mixture identity: Contamination is not additive noise: detector selection gives the phantom component its own shape, which can resemble the target phenomenon.The correct phantom null passes pure noise through the same acceptance rule and estimator.
- Non-integrable normalization: The ratio expectation diverges logarithmically when the signed reference return has positive density near zero.This makes per-item normalization non-integrable under the stated conditions.
- Non-integrable normalization: The finite-sample mean of the ratio estimator can look stable while failing to converge and being dominated by items with the smallest reference returns.Those items correspond to windows where little happened.
4 A dataset with recoverable phantoms
The empirical test pairs a detector-defined index with an independent official event index on the same futures series, enabling item-level identification of real events and phantoms. Its scope is deliberately narrow: one instrument, detector family, and official-index class.
- Price data: The price data are 1-second best-bid-offer observations for E-mini S&P 500 futures in windows spanning t0 ± 2 hours.The stated data source and total spend define the instrument-level testbed.
- Official index: The official index contains 555 announcement windows from CPI, non-farm payroll, and FOMC releases over 2010–2026.Release dates and times are sourced from official calendars and press releases.
- Detector and pools: The detector accepts windows when a 2-minute move exceeds ten times a pre-window per-second standard deviation and is run separately on three candidate pools.The pool structure enables comparison of one detector under different prevalence conditions.
- Testbed and scope: The testbed combines a price-threshold detector with an independently generated official index, making phantom identification credible.The two indices arise from causally independent processes and were not constructed with each other in view.
- Testbed and scope: The study’s conclusions are limited to one instrument, one detector family, and one class of official index.The authors state this scope boundary explicitly.
5 Results
Across three candidate pools, the same detector produces sharply different contamination rates, while Bayes-based precision predicts cross-pool performance far better than naive transfer. The contaminated response is an exact mixture of true and phantom components, and estimator choice determines whether contamination dilutes or inflates results.
- 5.1 One detector, three pools, three datasets: 473 phantoms versus 308 true events: phantoms form a majority of the screened published event list.The pooled phantom rate is 59%, but aggregation conceals the fact that one deployment is majority-synthetic.
- 5.2 Cross-pool transfer: naive +422%, Bayes −3.3%: 0.955 versus 0.183: naive cross-pool precision transfer errs by +422%, whereas the Bayes expression predicts measured precision within 3.3%.The naive transfer averages precisions from the two high-prevalence pools; the Bayes calculation uses each pool’s prevalence, TPR, and FPR.
- 5.3 Mixture decomposition: the contaminated curve, exactly: Exact to machine precision: the detected curve is a convex mixture of true-event and phantom components, with no residual to interpret.The phantom component has a detector-inherited shape rather than behaving like flat or unconditional random-walk noise; its level remains partly unexplained.
- 5.4 The direction of contamination is a property of the estimator: 0.082 versus 0.068 at 5 s: terminal normalisation inflates contamination, while denominator-free displacement reverses the ordering and dilutes it.At 300 s, the corresponding terminally normalised values are 0.317 versus 0.271, whereas denominator-free displacement is higher for official than legacy data.
- 5.5 A normalisation that returns 0.40 where the estimator returns 0.10: 0.3997 versus 0.1020: per-item normalisation produces a mean of ratios with no finite expectation on the same 335 events.Quiet windows dominate the ratio average because signed reference returns create the largest inverse-magnitude weights.
6 Related work
The paper situates its contribution within work on prevalence-dependent precision, weak and distant supervision, label noise, dataset documentation, and event studies. It distinguishes detector-defined contamination from random label corruption and measures its consequences in a deployed dataset.
- Base rates and precision under imbalance: The paper’s contribution is an end-to-end measurement of prevalence-driven contamination with recoverable ground truth, extending prior work on prevalence-dependent precision.The authors distinguish measuring the mechanism in a deployed dataset from introducing the mechanism itself.
- Weak, distant, and self-supervision: Weak- and distant-supervision methods construct labels by applying rules, but label-model approaches typically assume pool-stable accuracies that the paper’s results challenge.
- Label noise: Unlike standard label-noise models based on random corruption, detector-generated phantoms are correlated with features and carry their own signal shape.The paper argues that class-conditional random-flip methods do not address this form of contamination.
- Dataset quality and documentation: The protocol specializes dataset documentation to detector-defined datasets by reporting prevalence, detector rates, contamination, and phantom-to-event weighting quantities.
- Event studies: Event-study research supplies the realistic detector and downstream estimand used to study announcement responses.
7 A reporting protocol for detector-defined datasets
The reporting protocol requires analysts to measure detector performance in the deployment pool, report prevalence and contamination quantities, use selection-conditioned nulls, and avoid unstable signed-denominator normalizations. It also recommends oracle power checks before comparing methods.
- 7 A reporting protocol for detector-defined datasets: Precision must be measured or derived in the deployment pool, because cross-prevalence transfer is not permitted and high-prevalence validation gives near-zero evidence about FPR.
- 7 A reporting protocol for detector-defined datasets: Report π, TPR, FPR, c, and ρ, because precision alone conceals which quantity changed and ρ determines whether contamination is tolerable.
- 7 A reporting protocol for detector-defined datasets: Declare the time scale and reference window for every z-like statistic and normalization.
- 7 A reporting protocol for detector-defined datasets: Use an analytic null matched to the detector’s acceptance rule and estimator, selecting the selection-conditioned rather than unconditional null for screened subsets.
- 7 A reporting protocol for detector-defined datasets: Avoid per-item normalization by signed, mean-zero denominators because such estimators have no finite expectation; use a pooled denominator or none.
- 7 A reporting protocol for detector-defined datasets: An oracle power check should establish whether a configuration can distinguish methods before method comparisons are interpreted.
8 Limitations
The empirical magnitudes are constrained to one detector family, instrument, asset class, and official-index family, with incomplete historical coverage and an incomplete selection-conditioned null. The paper therefore supports relative conclusions rather than absolute claims about response speed.
- 8 Limitations: The study covers one magnitude-threshold detector family; its equations are general, but the numerical results do not transfer to other detector classes.
- 8 Limitations: The reported magnitudes concern one instrument, one asset class, and one official-index family, so the mechanism is general but the magnitudes are not.
- 8 Limitations: 415 clean official windows are effectively available from 2013 onward, limiting cross-decade comparability because 68 earlier windows lack usable one-second data.
- 8 Limitations: The selection-conditioned null reproduces the phantom component’s shape but not its level, leaving the null incomplete.
- 8 Limitations: The paper supports relative conclusions between datasets and estimators, not absolute statements about response speed.
Reproducibility statement
The authors release the event indices, phantom labels, screens, curves, simulator code, traceability map, decision log, and accident log. Because the price data cannot be redistributed, exact request manifests enable reconstruction of an identical dataset.
- Reproducibility statement: The release includes event indices, phantom labels, screens, curves, simulator code, traceability information, decision logs, and an accident log.
- Reproducibility statement: Although the price data cannot be redistributed under the vendor’s license, exact request manifests support reconstruction of an identical dataset.
A Six accidents in this paper’s own pipeline
The paper documents six pipeline accidents and resolves them through pre-registered fixes, showing how implementation choices can distort event timing, screening, estimation, and evaluation. It also reports a concrete estimator pathology and a limitation in the revision-only control.
- Accident 1: the unscaled z-score: 0.361 false-positive rate from an unscaled z-score caused 487 phantom CPI detections and an observed 81.7% phantom rate.The error was found by deriving the closed-form null and checking it against simulation.
- Accident 5: the same indexing bug, in a different file: 126 bp post-event movement was deleted by a row-offset screening bug in the November 2022 CPI example, whose true pre-window maximum was 3.3 bp.The repeated indexing pattern let the pre-window slice cross t0, causing announcement reactions to be classified as rolls.
- Accident 6: a factor-of-four estimator artefact: 4× disagreement on identical events exposed per-item normalization as an estimator pathology rather than a mere implementation mishap.The paper reports this resolution in the main text because it treats the pathology as a finding.
- Oracle checks: 0.024 CRPS oracle gap made an earlier alignment comparison unable to detect method differences above the seed-noise floor.The oracle-minus-baseline gap bounds the total effect any alignment or denoising method can produce in that configuration.
- The revision-only negative control: n = 3 usable revision-only control windows was insufficient for interpretation, leaving price coverage as the cheapest stated follow-up improvement.The control was pre-registered but included only for completeness because adequate coverage could not be obtained within the project’s data budget.
E Detector specification and screens
The detector scans one-second front-month ES mid-price windows and accepts two-minute moves exceeding a threshold based on preceding volatility. Timestamp-based screens then remove roll-contaminated and degraded-quote windows.
- Detector specification: The detector accepts windows when |∆p2m| exceeds ten per-second standard deviations estimated over the preceding five minutes.Windows span t0 ± 2 h on bid-ask mid prices.
- Screens: Screens exclude pre-window one-second returns above 20 bp and windows with more than 5% missing bid or ask observations.The roll screen is evaluated strictly before t0 −1 s using timestamp masks.