Source-linked AI summary

Point-in-Time Audit Before Alpha: Public-Archive Availability and a Negative Matched-Budget Study on BTC Perpetual Futures

Baocheng Zeng, Jinhao Yang, Peilin Han, Kangnan He

arXiv:2608.25348v1cs.SE

TL;DR

The paper asks whether public cryptocurrency archives support point-in-time-valid factor research when observations must be available and executable at each decision time. It audits Binance BTCUSDT perpetual-futures data with timestamp-aware admission, deterministic leakage auditing, matched-budget search evaluation, and one-time historical holdout access, finding a scoped negative result rather than profitability or general agent-superiority evidence.

  • Problem

    Public archives can appear usable while records are unavailable, unfinished, backfilled, or otherwise unusable at the decision time, so factor research must establish when observations could have been used.

  • Method

    The study audits Binance BTCUSDT USD-M perpetual data using event, publication, and availability times, then separates deterministic auditing, evaluation, and one-time holdout access.

  • Results

    Under matched valid-candidate budgets, the audited adaptive agent tied random search, while all evaluated historical-holdout runs had positive IC but negative net Sharpe under primary costs.

  • Takeaways & Limitations

    The evidence supports an audit-focused negative-result paper rather than a profitability or general agent-superiority claim.

  • Takeaways & Limitations

    Evidence is limited to Binance BTCUSDT USD-M, the retrospective holdout is not prospective validation, and the audit benchmark covers known injected rules rather than unknown or adversarial leakage implementations.

Abstract

from arXiv · show

Public cryptocurrency archives may appear usable when files exist, although factor research requires observations available and executable at each decision time. We audit public Binance BTCUSDT USD-M perpetual-futures data using event, publication, and availability times and separate proposal from deterministic auditing, evaluation, and holdout access. An initial gapless five-minute requirement for trade, mark, index, and open interest failed: the longest unrepaired intersection was 304.5729166666667 days. A disclosed revision made trade, mark, index, and realized funding the core streams and made open interest optional because its publication time was unverified. The revised mask retained 727 complete UTC days and supported a 436/145/146-day train, validation, and historical-holdout split. On 80 frozen known-rule templates, the auditor detected 40/40 violations and rejected 0/40 legal templates. Across ten null-signal paths, full auditing reduced mean false passes from 0.2910 to 0.0625. Under matched valid-candidate budgets, the audited adaptive agent tied random search and did not establish superiority. In the one-time historical holdout, all evaluated runs had positive IC but negative net Sharpe under primary costs. We therefore report a scoped negative result rather than a profitability or agent-superiority claim.

1. Introduction

The paper argues that perpetual-futures factor research must establish when observations became usable, then audits Binance BTCUSDT data under a disclosed revised admission rule.

  • A valid perpetual-futures study must establish both what an observation represents and when it could have been used.Candidates can appear predictive when records were unavailable, bars unfinished, funding moved backward, or gaps filled with future information.
  • 304.5729166666667 days was the longest unrepaired intersection after the initial continuous five-minute requirement failed.The requirement covered trade, mark, index, and open interest on one exact grid for at least 365 days.
  • The revised core uses trade, mark, index, and realized funding aligned by availability time, while open interest is optional because its publication time is unverified.The revision preserved the failed exact-grid result instead of silently repairing the archive.
  • The study asks whether revised streams support sufficient point-in-time data, whether auditing classifies legality, whether null false passes decline, whether adaptive search outperforms baselines, and whether holdout value persists.These five questions cover data admission, auditing, false discovery, search comparison, and historical economic evaluation.
  • The contributions are deliberately narrow: timestamp-aware auditing, executable-plan checks, null controls, append-only accounting, and one-time holdout access.The paper reports negative search and economic findings without selecting a post-holdout rescue specification.

2. Related Work

Related work includes multiple formulaic factor-search agents, but the paper narrows its contribution to perpetual-specific point-in-time auditing and matched search evaluation.

  • RiskMiner, AlphaGen, AlphaForge, FAMA, AlphaSAGE, and Navigating the Alpha Jungle represent Monte Carlo, reinforcement-learning, joint-mining, neural-symbolic, GFlowNet, and LLM-guided approaches.Their reported emphases include formula discovery, factor combination, exploration, convergence, runtime, and search feedback.
  • Recent public work also covers constrained cryptocurrency factor agents, safe DSL execution, executable program evolution, digital-asset feedback, and execution-governed perpetual tuning.These works narrow any broad novelty claim for agentic factor mining.
  • The paper’s distinguishing question is whether a perpetual-specific point-in-time audit changes false passes and whether an adaptive proposer survives an exactly matched valid-evaluation comparison.This framing differs from asking only whether an agent can emit formulas.

3. Method

The method constructs a masked Binance archive with explicit timing assumptions, deterministic auditing, frozen candidate gates, and matched-budget search and holdout evaluation.

  • Data and timing: The archive contains 826 official-checksum-verified ZIP files covering trade, mark, index, funding, and daily metrics/OI data.Historical OKX probes were unusable for the tested public route because old funding and OI responses were empty, so the study used Binance.
  • Data and timing: A decision at time t may read only records with availability_time <= t; bar availability is close_time + 1 ms, while funding uses event time plus a five-minute research assumption.Admission sensitivity tests use 0, 5, 10, and 15 minutes because the archive lacks a distinct publication-time field.
  • Data and timing: Missing observations are removed through an explicit mask, with interpolation, backward filling, timestamp snapping, duplicate deletion, and venue mixing prohibited.Open interest is disabled as a factor input because its publication time is unverified.
  • Data and timing: The retired rule measured an exact five-minute intersection, whereas the revised rule measures complete decision days for trade, mark, index, and native-frequency funding after masking.These are different sample definitions rather than two estimates of one quantity.
  • Data and timing: 727 complete UTC days yielded 436 training, 145 validation, and 146 historical-holdout days with a 60-minute purge and embargo.The hash-locked final partition is called a one-time historical holdout rather than prospective validation.
  • Auditing and candidate qualification: The deterministic evaluator parses expressions, constructs lineage, applies availability masks, audits execution, computes metrics, performs selection, and writes the ledger.It rejects future source shifts, same-bar execution, future realized funding, incorrect OI alignment, and backward fill.
  • Auditing and candidate qualification: The conformance benchmark contains 40 illegal and 40 legal executable templates, with acceptance requiring at least 38/40 illegal detections and no more than 2/40 legal rejections.Wilson intervals accompany but do not replace the finite-template rule.

4. Experiments

The experiments are organized around five research questions and evaluate data integrity, audit classification, false discovery, search efficiency, and historical economic performance.

  • Data and audit experiments: The data audit measures census coverage, duplicates, ordering defects, exact-grid gaps, funding cadence, and lag sensitivity.
  • Data and audit experiments: The template audit evaluates confusion-matrix and reason-code accuracy on 80 frozen programs.The false-discovery controls use 10 primary null paths and 270 audit-by-gap scenarios.
  • Search and holdout experiments: The search comparison evaluates 3 methods × 5 seeds × 2 horizons × 100 valid candidates.The historical holdout uses frozen candidates, primary costs, 20 fee–slippage combinations per run, one-bar delay, extreme-observation deletion, and volatility-state checks.
  • Search and holdout experiments: The principal outcomes include complete days, violation recall, legal false rejection, false-pass rate, PBO, candidate counts, test IC and RankIC, and economic performance metrics.Direction accuracy, trade count, and evaluations to first discovery were not emitted upstream and are not reconstructed.

5. Results

The results establish a revised point-in-time data admission rule, strong performance on known-rule auditing and null controls, but no demonstrated superiority of the audited search agent or economic value in the historical holdout.

  • 5.1 Public-archive availability: 304.5729166666667 days was the longest unrepaired exact-grid intersection under the retired trade/mark/index/OI requirement.The revised analysis instead retained 727 complete UTC days under availability masking for trade, mark, index, and realized funding.
  • 5.3 Null controls and missing-policy ablation: 0.0625 mean false passes under the full audit compared with 0.2910 without auditing at 5% missingness and availability masking.The full audit improved all 10 null paths, with a relative reduction of 78.5%.
  • 5.3 Null controls and missing-policy ablation: The supported null-control result is the joint effect of the full known-rule audit pipeline, not an independent effect of masking.Exact-grid deletion and availability masking were identical in the executed simulation, while backward fill was not uniformly worse.
  • 5.4 Search accounting and efficiency adjudication: The audited agent tied random search on qualified candidates and therefore did not establish superiority over both baselines.Raw attempts were lower for the agent, but evaluations to first qualified factor and resource-normalized runtime were not frozen outcomes.
  • 5.5 Historical-holdout prediction and economics: 0/460 executed fee–slippage cells had positive Sharpe, despite positive IC across the evaluated runs.The best observed Sharpe values were -6.3148 for the agent, -5.5009 for random search, and -5.4938 for tree GP.

6. Discussion

The strongest evidence is procedural: auditing removed invalid candidates and sharply reduced false passes, while predictive gains did not translate into economic value or agent superiority.

  • Auditing removed candidates that should not enter statistical selection and sharply reduced false passes in executed null controls.The benchmark result is procedural rather than evidence of profitability.
  • The benchmark is not an unseen-adversary test, and its finite-template intervals are wide.
  • Positive holdout IC coexisted with very high turnover and uniformly negative cost-adjusted results.Changing the combiner, holding period, or cost model after opening the holdout would invalidate the one-time-access design.
  • The audited adaptive proposer tied random search under matched valid-candidate budgets and did not establish the preregistered discovery-efficiency claim.First-discovery and resource measurements were not emitted.

7. Limitations

The study’s conclusions are bounded by its single-market historical setting, assumptions about funding publication, finite known-template audit, incomplete search comparisons, and execution approximations.

  • Evidence is limited to Binance BTCUSDT USD-M and cannot establish cross-market validity.
  • The holdout interval ended before protocol lock, so one-time access is not prospective validation.
  • The archive lacks a separate funding publication timestamp, making the five-minute timing assumption an admission sensitivity rather than a predictive-outcome test.
  • The audit benchmark covers known injected rules, not unknown or adversarial leakage implementations.
  • Incomplete search curves, missing baselines, and aggregated-bar execution assumptions limit discovery-efficiency, nonlinear-baseline, capacity, and live-tradability claims.

8. Conclusion

The paper concludes that point-in-time auditing exposed data-validity problems and supported a narrow procedural negative result. The audited proposer did not beat random search, and positive historical-holdout IC did not survive economic evaluation.

  • 8. Conclusion: 304.5729166666667 days was the failed original exact-grid intersection, while the revised core definition yielded 727 complete days without filling.
  • 8. Conclusion: The deterministic auditor passed a finite known-template benchmark and reduced false passes across executed null controls.
  • 8. Conclusion: Under matched valid-candidate budgets, the audited adaptive proposer tied random search and did not establish superiority.
  • 8. Conclusion: Positive historical-holdout IC failed every primary and sensitivity economic test.
  • 8. Conclusion: The evidence supports an audit-focused negative-result paper, not a profitability or general agent-superiority claim.
Loading 2608.25348v1…