Source-linked AI summary
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng
TL;DR
AI research agents need evidence that useful results were produced independently of withheld research history and that feedback contributed to reaching them. DCP turns these questions into executable recovery and randomized feedback tests, with Core and Evidence decisions replayed from frozen records. In two controlled audits, it found zero recoveries in 96 episodes and 30 truthful versus zero neutral recoveries in each paired study, while additional cases exercised alternative decisions.
Problem
AI research results need evidence connecting measured outcomes to the research that produced them, including whether matched agents can recover them and whether feedback helps.
Method
DCP audits one numerical outcome through sealed validation, registered matched-agent recovery tests, and randomized truthful-versus-neutral feedback comparisons.
Results
Two controlled audits recorded zero recoveries in 96 episodes and 30 truthful recoveries versus zero neutral recoveries in paired studies, with decisions reproduced by a deterministic verifier.
Takeaways & Limitations
DCP provides a portable, replayable evidence language that separates useful outcomes, alternative routes, and feedback effects across research domains.
Takeaways & Limitations
DCP certificates are scoped to a registered model, information boundary, budget, probability bound, and replayable record, with captured Web data subject to privacy, licensing, and security requirements.
Abstract
from arXiv · showhide
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 gives matched agents the registered starting information and observed Web content while withholding the target research history. Every valid method reaching the numerical target supplies a recovery witness and triggers the Core veto. DCP Core requires adequate controls, zero observed recoveries, and a finite-sample bound on recovery in one fresh registered episode. Optional Gate 3 measures the average effect of truthful feedback relative to a specified neutral policy from a shared checkpoint. DCP Evidence adds this effect after independent null calibration and a registered effect margin. Two controlled audits exercise the complete protocol in SQLite optimization and virtual catalyst control under different models. Each produced zero recoveries in 96 episodes, with an upper bound of 0.0468. Each paired study yielded 30 truthful recoveries and zero neutral recoveries, with passing 60-pair null studies. Additional cases exercise Core, recovered, and audit-incomplete decisions. A deterministic, LLM-free verifier reproduces the decisions from frozen evidence. DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research.
1 Introduction
DCP frames AI research evaluation as three distinct, executable questions about one numerical outcome: usefulness, recoverability by matched agents, and feedback effects. Its contributions combine alternative-route auditing, finite-sample recovery bounds, randomized feedback comparisons, and replayable decisions.
- AI research agents produce useful artifacts through experimentation and revision, so measured results need evidence connecting outcomes to the research that produced them.
- DCP separates sealed utility validation, matched-agent recovery with withheld run history, and feedback effects as three measurable questions about one outcome.
- Every valid alternative route to the same numerical outcome can supply a recovery witness under registered information, resources, provenance, and executable validity conditions.
- Core combines qualified recovery witnesses, finite-sample recovery bounds, and zero recoveries, while Evidence adds randomized truthful-versus-neutral feedback effects.
- The protocol is calibrated across software optimization, virtual experimental control, recovered and incomplete cases, with portable evidence bundles and a deterministic verifier.
2 Related Work
Related work spans AI research agents, iterative feedback and experimental design, and evaluation integrity. DCP extends these lines by testing recovery and feedback under registered information boundaries and randomized comparisons.
- AI research agents and their evaluation: AI research agents combine search with objective evaluation, while broader systems automate research and hypothesis development.
- Iterative feedback and experimental design: Prior studies examine tool use, search, reflection, feedback, controlled revision, and skill refinement in iterative research workflows.
- Evaluation integrity and process attestation: Evaluation-integrity research studies contamination, memorization, fixed-corpus retrieval, query initialization, and executable outputs as influences on measured performance.
3 Method
DCP audits useful outcomes, alternative recovery routes, and feedback effects under registered information, execution, and statistical boundaries. Its Core and Evidence decisions combine executable validity checks, recovery controls, finite-sample bounds, and randomized feedback comparisons.
- The outcome and its information boundary: DCP defines an outcome audit around a useful result, its available information, and alternative recovery under a registered boundary.The audit applies across programs, models, data products, and experimental recipes.
- Registered audit boundary: Registration fixes the task, model, information, interface, tools, budgets, baseline, validity rules, selection, stopping, and statistical analysis before the target run.A symmetric evaluator checks baseline, target, and controls after candidate generation closes.
- Gate 2 tests recovery: Every valid method reaching the registered score threshold qualifies for recovery, while the verifier computes the target score after artifacts are committed.The tolerance ε must remain below the minimum useful gain and the target’s improvement over baseline.
- Recovery evidence: A qualified recovery witness establishes an alternative route and triggers the Core veto, whereas pB estimates recovery probability across fresh challenger episodes.A witness and a low recovery probability answer different evidential questions.
- Gate 3 feedback effect: Gate 3 estimates the conditional average effect of truthful versus neutral feedback from a shared checkpoint using fresh paired branches.Evidence additionally requires independent null calibration and a lower confidence bound above the registered effect threshold plus calibration margin.
- Verification and decisions: A deterministic verifier separates evidence production from decision checking and recomputes numerical and record consistency from frozen bundles.Independent readers can check the same evidence across agent implementations.
4 Experimental Audit Cases
The audit exercises DCP on SQLite optimization and virtual catalyst control, with additional calibration cases testing recovery and incomplete decisions. The primary audits combine sealed improvement, matched recovery tests, and paired feedback studies under registered models and controls.
- Audit design: Two complete three-gate audits and three diagnostic cases calibrate DCP across controlled information boundaries with committed candidates and a shared evaluator.The diagnostic cases exercise Core, recovered, and incomplete decisions.
- SQLite-Web: DeepSeek-v4-flash’s SQLite agent scored 0.8855 on the sealed workload, an 88.55% reduction from the no-index baseline.The final plan selected the four high-traffic query families after public SQLite documentation access.
- SQLite-Web: Gate 2 withheld the SQLite target plan, score, traffic measurements, and reasoning trace while exposing the registered workspace and observed Web bytes.Gate 3 compared truthful traffic-weighted feedback with neutral uniform-family feedback from a pre-experiment checkpoint.
- Virtual catalyst optimization: Gate 2 withheld the catalyst target plate and downstream reasoning, while Gate 3 compared truthful measurements with matched non-directional measurements from the frozen anchor checkpoint.A separate 60-pair family tested whether the neutral channel itself changed performance.
5 Results
The complete audits passed useful-outcome validation and found no recoveries under their registered scopes, while feedback studies supported Evidence decisions. Additional cases demonstrate that qualified alternative routes trigger recovery vetoes and that the protocol distinguishes Core, recovered, and incomplete outcomes.
- Primary decisions: All four outputs passed Gate 1, while SQLite-Web and catalyst each recorded 0/96 recoveries.Their positive controls passed 45/45, yielding a recall lower bound of 0.8889 and a zero-hit upper bound of 0.0468 < ρ.
- Primary decisions: 45/45 positive controls passed, with a recall lower bound of 0.8889 above the registered 0.8 minimum.The deterministic verifier returned DCP Core in all three scopes.
- Recovery decisions: 0.9363 and 0.9356 qualified knapsack artifacts exceeded the 0.9329 recovery line, triggering the Core veto.These artifacts supplied constructive recovery witnesses rather than probability-only evidence.
- Recovery decisions: SQLite-Web challengers reached at most 0.6734 against a 0.8805 recovery target, while catalyst challengers reached at most 0.8146 against 0.95.Figure 3 presents the two model-specific audits separately.
- Feedback evidence: 30/30 truthful recoveries versus 0/30 neutral recoveries produced an estimated binary policy effect of 1.0.The conservative exact paired-binary 99% interval was [0.6379, 1.0], and both Evidence decisions passed null calibration.
- Feedback evidence: The Evidence decisions measure the average truthful-versus-neutral policy effect, distinct from the neutral-recovery upper bound of 0.1619.This distinction applies after the registered checkpoint.
- Decision coverage: Five cases exercise Evidence, Core, recovered, and audit-incomplete outcomes across distinct information and execution conditions.SQLite and catalyst support Evidence; device calibration supports Core; knapsack supplies qualified recovery; affine remains audit incomplete.
6 Discussion
DCP reports research contributions through reusable records that expose thresholds, scopes, budgets, witnesses or bounds, and feedback intervals for inspection and replay. Its prospective and retrospective modes distinguish registered present audits from present-day recoverability.
- Reporting a research contribution: DCP evaluates recovery across every admissible challenger implementation, so recombination, cross-domain transfer, or a different program can supply the same numerical witness.The knapsack case demonstrates this outcome-level rule directly.
- Reporting a research contribution: A reusable report records the outcome threshold, model and information scope, episode budget, recovery witnesses or probability bound, and any feedback interval.This profile lets reviewers inspect alternative routes and replay the numerical decision.
- Reporting a research contribution: These fields turn discovery reporting into a common experimental record that authors can submit, auditors can challenge, and readers can verify.Core and Evidence preserve the quantities needed for comparison.
- Deployment and retrospective recovery: Prospective audits begin with approved registration and sealed evaluation, followed by evidence collection, controls, offline replay, and independent countersigning.The SQLite-Web and catalyst audits used 435 and 507 recorded sessions and cost 56.40 and 61.17 USD, respectively.
- Deployment and retrospective recovery: Retrospective audits distinguish pre-publication-equivalent recovery scope from present-day recoverability based on post-publication knowledge.Positive Core and Evidence decisions require prospective registration and complete evidence collection.
7 Conclusion
DCP makes evidence about AI research contributions executable by separating utility, alternative routes, and feedback effects into checkable decisions. Two controlled audits and additional cases show that this evidence language is portable across research domains and reproducible from frozen records.
- Conclusion: DCP makes evidence about AI research contributions executable through sealed utility validation, matched-agent recovery tests, and optional randomized feedback comparisons.Core requires zero recoveries and a finite-sample bound; Evidence adds a checkpoint-conditional truthful-feedback effect, null calibration, and a registered margin.
- Conclusion: Two controlled audits recorded zero recoveries in 96 episodes and 30 truthful recoveries against zero neutral recoveries, with passing 60-pair null calibration.They used distinct models in SQLite optimization and virtual catalyst control.
- Conclusion: A deterministic verifier reproduced complete decisions from frozen records, while device, knapsack, and affine cases exercised Core, recovered, and incomplete outcomes.The resulting protocol is presented as a portable experimental language across research domains.
Reproducibility Statement
Frozen machine-readable bundles and evidence records support deterministic offline replay of the numerical claims.
- Reproducibility Statement: Frozen machine-readable bundles contain task construction, registrations, contracts, evaluation units, thresholds, statistics, budgets, costs, identifiers, and replay tooling.The release includes dcp-audit, dcp-harness, example bundles, and figure-generation scripts.
Ethics Statement
DCP supports accountable reporting by binding each decision to a registered model, information boundary, budget, probability bound, and replayable record. Deterministic replay gives authors, reviewers, and auditors a common record for examining research claims.
- DCP binds each decision to a registered model, information boundary, budget, probability bound, and replayable record.
AI Use Statement
Generative AI tools supported a wide range of research activities, including design, implementation, testing, evidence inspection, literature work, and manuscript editing. GPT, Codex, Claude, DeepSeek models, and image-generation tools were used in specified roles.
- Generative AI tools supported ideation, protocol design, statistical design, review, experiment design, implementation, testing, evidence inspection, literature work, and editing.
- GPT and Codex assisted these activities, while Claude supported review and the registered CLI execution layer.
- DeepSeek-v4-flash and DeepSeek-v4-pro served as audited agents and matched challengers.
A.1 Main result
The main result section specifies DCP’s statistical and procedural safeguards, then reports complete audits with zero recoveries, validated controls, and positive truthful-feedback effects. It also records recovered and incomplete cases, showing that the verifier distinguishes these outcomes.
- A.1 Main result: DCP estimates uncertainty separately for scores, episode recovery, and randomized feedback, using registered sampling units.
- A.2 Candidate decisions: Candidate hits require an adjusted lower confidence bound at least −ε, while misses require an adjusted upper bound below −ε; remaining candidates are unresolved.
- A.1 Main result: 90 episodes are required for αrecovery = 0.01 and ρ = 0.05; the complete audits registered 96 episodes.
- A.1 Main result: 0 recoveries occurred in 96 matched challenger episodes for both complete audits, yielding an upper bound of 0.0468381.
- A.1 Main result: 30/30 truthful recoveries versus 0/30 neutral recoveries produced a paired exact 99% interval of [0.6379275, 1] in each target feedback study.
- A.1 Main result: A recovered case contained qualified witnesses, whereas the affine parity run remained incomplete because control or challenger and interface records were inadequate.