Source-linked AI summary

What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals

Xi Qin

arXiv:2608.19269v3cs.SEcs.AI

TL;DR

The paper addresses the gap between an evaluation artifact’s reported metric and the scientific claim attached to it. It formalizes claim replay over frozen evidence and grounded semantic alternatives, then audits 124 Inspect Evals units. The census finds 110 typed stops before deterministic inference, while completed audits expose claim-specific instability and stable structure across primary and review families.

  • Problem

    Evaluation artifacts do not necessarily license their attached claims because the historical evidence and alternative semantics needed for replay may be unbound.

  • Method

    The paper defines claim replay using a frozen substrate D, a grounded family F, and a claim query q, then enumerates the resulting identified set across Inspect Evals.

  • Results

    110 of 124 mechanically eligible units stop before deterministic inference because required historical evidence or semantic grounding is unavailable.

  • Takeaways & Limitations

    The audit returns typed stops, instability witnesses, and stable substructure instead of forcing one evaluator meaning or one robust/not-robust label.

  • Takeaways & Limitations

    The census is limited to the pinned eligible frame and does not estimate ecosystem prevalence, benchmark defects, or prospective release-process effects.

Abstract

from arXiv · show

Evaluation artifacts specify a forward computation: a task, scorer, and reported metric. They do not necessarily license the claim attached to that metric because the historical evidence and alternative semantics needed to replay it may be unbound. We formalize this missing claim-replay layer through a frozen substrate D, a grounded family F, a claim query q, and the resulting identified set. We then census all 124 mechanically eligible Inspect Evals units at a pinned commit. Every unit receives a terminal disposition; 110 stop before deterministic inference because required historical evidence or semantic grounding is unavailable. Where execution closes, exact values, winners, complete orders, and pairwise relations separate by claim resolution and by primary versus review family. The audit therefore returns typed stops, instability witnesses, and stable substructure rather than forcing one evaluator meaning or one robust/not-robust label.

1 INTRODUCTION: THE MISSING INFERENCE LAYER

The paper argues that evaluation artifacts report forward computations but do not by themselves license the scientific claims attached to their metrics. It introduces a claim-replay object and applies it in a complete Inspect Evals census.

  • The missing inference layer: Evaluation artifacts report what ran, how outputs were scored, and which metric was reported, while scientific claims require replay under grounded alternative readings.The paper frames claim replay as an additional contract binding historical observations, admissible semantic alternatives, and the queried claim.
  • Contributions: The claim-relative object (D, F, q) separates primary conclusions from sensitivity to disputed semantic readings.
  • Contributions: The complete census covers 124 mechanically eligible units and records where historical claim replay becomes executable.
  • Contributions: The executable audit distinguishes exact values, directions, winners, complete orders, and stable pairwise structure.
  • Paper organization: The paper proceeds from defining the object, through the evidence gate and completed audits, to the executable contract.A reproducibility package maps reported results to machine-readable records and verification checks.

2 CLAIM-RELATIVE IDENTIFICATION

Claim-relative identification evaluates a query over a frozen historical substrate and a grounded family of admissible specifications. The framework distinguishes primary and review interpretations, then reports identified claims, witnesses, and stable relations.

  • Inputs and families: The frozen substrate D contains endpoint-specific outputs and relevant sample, judge, state, support, and precision bindings.
  • Inputs and families: Primary specifications are admitted candidates, while the review family adds disputed candidates and excludes candidates enter neither family.
  • Inputs and families: Admission changes the inferential domain, so primary conclusions need not survive the wider review envelope.
  • Identification and resolution: A claim is identified when its identified set has one element; queries may target exact values, directions, thresholds, winners, top sets, or complete weak orders.
  • Identification and resolution: Distinct outcomes provide witnesses for non-identification, while singleton pairwise relations form a stable backbone.
  • AgentDojo example: AgentDojo retains a unique primary order and Claude as winner, but disputed readings add a review-order alternative while preserving Claude’s two pairwise relations.The paper treats the displayed percentages as per-reading values because eligibility and aggregation differ across the ordinal family.

3 COMMIT-BOUND CENSUS

The commit-bound census evaluates all 124 eligible units and assigns each a terminal disposition. Most units stop before deterministic claim analysis because historical evidence or semantic grounding is unavailable.

  • Finite frame and terminal accounting: 124 eligible units receive terminal records at pinned commit 32b79a2, with 14 reaching deterministic claim analysis and 110 stopping earlier.The 110 stops include 103 in the outcome-blind census-review path and seven in the frozen random draw.
  • Finite frame and terminal accounting: 47 stops involve evidence-binding failure, 35 missing comparative observations, and 9 missing judge traces.The remaining twelve concern endpoint or target typing, missingness, claim selection, or terminal state.
  • Terminal accounting: A stop identifies the first unavailable object and the minimum artifact needed for claim replay, rather than representing a benchmark failure or zero score.A fresh model, judge, agent, or sandbox run changes the historical substrate and does not replace stopped units.
  • Three evidence streams: The census, frozen random draw, purposive audits, and exploratory cases license different statements and are never pooled.The 3/10 figure is an equal-stratum-weighted design statement, not a unit-weighted estimate for the 124-unit frame.
  • Native capacity versus claim replay: Claim replay additionally requires a bound historical substrate, grounded semantic family, and typed query beyond native task, scorer, and metric objects.The audited records show task and scorer availability often preceding missing observations, judge decisions, state, support joins, or semantic registry information.

4 WHEN EXECUTION CLOSES

When the historical substrate and semantic family are executable, enumeration separates claim resolutions and primary from review conclusions. The completed audits retain stable winners or pairwise structure even when finer values or orders are non-identified.

  • Random complete panel: All three completed random draws are winner- and order-non-identified in both primary and review scopes.Their mechanisms differ in credit definition, aggregation, and typed target; the panel describes the frozen draw rather than stop frequency.
  • Purposive mechanism panel: AgentDojo and AutoML retain unique primary orders but disputed readings reverse one lower pair, while BEIR reverses its two nonwinning BM25 systems.All three retain the winner, showing that admission can alter order identification without changing the top result.
  • Exploratory landscape panel: Fine-resolution non-identification can coexist with an identified direction or winner and a large stable pairwise backbone.Reporting only flips discards decision-relevant stable structure, while reporting only the review envelope can hide a primary-family singleton.
  • Contracted interpretation: The added contract preserves non-executable cases, types family admission, and returns stable coarsenings alongside instability.Its purpose is not to create a larger multiverse but to distinguish evidence availability, semantic admission, and claim resolution.

5 EXECUTABLE AUDIT CONTRACT

The executable audit contract binds human-defined claim and admission decisions before deterministic analysis. It reports typed stops, scoped identified sets, instability witnesses, and stable coarsenings instead of a single robustness label.

  • Human boundary: Human auditors define the claim, endpoint type, comparable support, candidate origin, and A/D/X admission before prospective outcomes.The contract does not automate discovery of a uniquely correct evaluator meaning or certification that the family is complete.
  • Contract contents: The contract records claim and query identifiers, typed endpoints, candidate provenance, reliability, support, aggregation, missingness, exposure, and evidence-horizon information.These fields bind the semantic and historical decisions needed to interpret the resulting claim image.
  • Deterministic execution: After admission is frozen, the analyzer evaluates primary and review families deterministically and constructs values, top sets, weak orders, pairwise relations, and minimum witnesses.Missing observations, support decisions, orientation, or admission produce a typed STOP with an actionable unblock rather than imputation.
  • Inference outputs: The audit reports a resolution ladder from value image through direction, winner or top set, complete order, and pairwise backbone.Every non-singleton level carries a two-specification witness, while stable coarsenings record what survives.
  • Reporting rules: Results must name (D, Fscope, q), separate primary from review conclusions, give every STOP its missing object and minimum unblock, and retain selection and exposure state.The main-text interface is the scoped claim, witness, and stable structure; machine-readable realization is appendix-facing.

6 IMPLICATIONS, LIMITS, AND RELATED WORK

Evaluation releases should bind a claim-replay interface that connects historical evidence, semantic alternatives, and the queried claim. The paper frames this interface as a claim-relative diagnostic whose retrospective findings do not establish prospective utility or ecosystem-wide prevalence.

  • Implications: A claim-replay interface should bind historical observations, evaluator configuration, support and missingness decisions, typed aggregation, candidate admission, and the claim query.The interface either closes the claim or returns a typed stop identifying what is missing.
  • Limits: The retrospective census does not estimate whether prospective adoption would reduce stops or change downstream decisions.A prospective test would need preregistered release units and claims, required replay fields, and outcome comparisons.
  • Implications: The paper’s contribution connects finite-frame accounting, evidential executability, grounded semantic families, and claim-level identified sets.This chain makes the gap between benchmark execution and claim licensing inspectable.
  • Implications: A scalar sensitivity range can conflate semantic variation, sampling uncertainty, and reversible representation changes that the contract keeps separate.The contract enumerates semantic variation on fixed evidence and removes purely reversible representation changes before claim comparison.
  • Limits: The complete census contains 124 mechanically eligible units from 129 strict-path candidates after five rule-based exclusions.No evaluation outcome enters the eligibility predicate.
  • Limits: The frozen random panel reports 3/10 completion under equal-stratum-weighted sampling, not a unit-weighted frame rate or non-identification rate.Seven selected units stop before any claim image exists, and purposive audits are not pooled with the random panel.

B.3 PRESERVATION LAYERS

The audit treats preservation as claim-relative: replaying a scorer may require historical evidence, intermediate state, support joins, and admitted semantics beyond the retained computation. It assigns typed first-stop dispositions and, when execution closes, reports identified winners, orders, pairwise relations, and stable backbones rather than a single evaluator verdict.

  • Preservation layers: Claim replay requires layered preservation, including historical outputs, mediator decisions, state, support joins, and an admitted semantic registry.A retained scorer is only the first preservation layer.
  • Preservation layers: First-stop accounting follows registration and gates G1–G8, from source binding through synthesis and verification.Simultaneously missing requirements remain secondary modes.
  • Stopping dispositions: 110 stops are preservation diagnoses, not zero-valued performance results or evidence that the benchmarks are defective.Each terminal record names the first reason, fresh-run boundary, and minimum unblock.
  • Claim-relative results: Across complete audits, claim resolution and primary versus review scope determine whether winners, orders, and pairwise relations are identified.AgentDojo and AutoML illustrate primary orders that are identified while review-envelope orders are not.
  • Stable structure: Complete-order non-identification still permits stable-backbone analysis: random completions retain no stable pair, whereas all three purposive audits retain an identified winner.Two purposive audits also retain most lower relations in the review envelope.
  • Audit design: The frozen stratified sample contains 10 objects—3 complete audits and 7 stopping records—while 3 purposive objects are excluded from the 3/10 completion denominator.The random and purposive panels have different acquisition licenses and are not pooled.

E.1 THEAGENTCOMPANY: COMPLETE FINITE-FAMILY AUDIT

The audit separates executable claim replay from unavailable historical evidence and semantic grounding across finite evaluation families. Closed cases show that values, winners, orders, and pairwise relations can resolve differently across primary and review families.

  • Audit scope: The audit evaluates every admitted member of each declared finite family but does not certify family completeness.A fresh run would create a new evaluation substrate rather than reconstructing the frozen one.
  • Claim resolution: inspect_haiku45 and original_sonnet4 reverse their winner under mean-checkpoint-progress versus strict-task-completion.The corresponding primary-family pairwise relation is unstable: 0/1 method pairs are stable.
  • Claim resolution: inspect_reproduction and original_paper likewise reverse their winner across aggregation and zero-handling specifications.Primary and review envelopes both have 0/1 stable method pairs.
  • Claim resolution: The three-system audit yields non-identified winner and order claims, with the winner image changing across admissible metric specifications.The primary winner image includes gpt35_turbo, claude3_opus, and a three-way tied set.
  • Protocol stopping records: 110 of 124 units stop before deterministic inference because required historical evidence or semantic grounding is unavailable.The audit records typed stopping modes and minimum unblocking artifacts rather than imputing missing outcomes.

G.2 AUTOML BENCHMARK: PURPOSIVE DEEP AUDIT

The purposive AutoML audit enumerates five systems under five metric specifications and separates stable winners from unstable complete orders. H2OAutoML remains the winner in the review envelope, while the full order changes.

  • Frozen substrate: The audit fixes a 149-row results table spanning 5 frameworks, 3 tasks, and 29 shared task-fold keys.AUC values and the missingness mask are fixed, with ranking as the endpoint.
  • Primary family: H2OAutoML is the primary-family winner, with complete order H2OAutoML > autosklearn > TPOT > AutoWEKA > oboe.The primary family contains three admitted specifications.
  • Review envelope: The review envelope preserves H2OAutoML as winner but makes the complete order non-identified.It additionally admits H2OAutoML > autosklearn > AutoWEKA > TPOT > oboe.
  • Pairwise stability: 10/10 primary-family pairs are stable, compared with 9/10 review-envelope pairs.The available_fold_micro_auc versus worst_task_auc specifications witness the sole review-envelope reversal.
  • SciFact comparison: SciFact identifies bm25_titlex2_k1_1_2_b_0_75 as winner but not a unique complete order across five metric specifications.The two order images differ only in the relative placement of the other BM25 runs.
  • SciFact comparison: 2/3 SciFact system pairs remain stable across the five admitted specifications.The bm25_k1_0_9_b_0_4 versus bm25_k1_1_2_b_0_75 pair reverses by metric.

H.1 CASE-LEVEL MECHANISM CARDS

The mechanism cards show that exploratory cases can preserve winners while changing scalar values, orders, or pairwise relations. They also distinguish endpoint typing from genuine instability and restrict claims when support is limited.

  • VQA-RAD: VQA-RAD preserves the GPT-5.2 winner and sole pair relation despite differing exact values.An equal closed/open macro is transparent but is not source-declared as the overall.
  • MORU: MORU changes only the Gemini/Grok relation under equal dimension weighting, while GPT-5.2 remains winner and 9/10 pairs stay stable.Equal dimension weighting is audit-proposed rather than source-declared.
  • InstrumentalEval: InstrumentalEval exhibits scalar sensitivity only because one bound system makes winner, order, and pair claims inapplicable.The evaluator declares a prompt-global convergence fraction and reports six task-type metrics.
  • ANIMA: ANIMA keeps GPT-5-mini as winner, swaps Claude and Gemini, and preserves 5/6 pairwise relations.The dimension-normalised average is reconstructable from thirteen displayed means, while raw prompts, grader replies, and weights are absent.
  • AIR Bench: AIR Bench changes the exact delta while preserving its tested-versus-paper direction.The fixed source reports a 5,694-item overall delta and sixteen category rows.
  • Reconstruction boundary: The exploratory chains reparse fixed source-table cells independently, validating source-table reconstruction rather than raw trajectory replay.A second implementation reconstructs the reported orders from the same displayed cells.

I.1 MORU

The MORU, ANIMA, and BFCL audits show that order sensitivity ranges from a lower-pair change with a fixed winner to a support-limited subset-winner reversal.

  • MORU: MORU preserves the winner while changing one lower pair.
  • ANIMA: ANIMA preserves GPT-5.2 as the identified winner, with 9/10 pairs stable despite a Gemini–Grok swap.
  • ANIMA: ANIMA’s exactness is relative to the displayed table because raw prompts, grades, and weights are unavailable.
  • BFCL: BFCL reverses the order on a 3,399-item published-category subset rather than the 3,981-item headline support.
  • BFCL: The BFCL reversal is exploratory because the equal-group overall is not source-declared, so it is not a headline winner flip.

J EXECUTABLE RECORD AND DETERMINISTIC ENUMERATION

The executable record turns admission and claim replay into deterministic set construction and enumeration, while preserving typed stops when critical bindings are missing.

  • Executable record: Admission maps candidates into primary, review-only, or excluded families through a(f) = A, D, or X.
  • Audit outcomes: MORU changes one lower pair while keeping its winner fixed, whereas BFCL has a subset-winner reversal whose support does not close.
  • Audit outcomes: Figure 10 localizes each order change by independent recomputation: MORU and ANIMA alter one lower pair, while BFCL remains support-limited.
  • Deterministic enumeration: Figure 11 shows that human judgment fixes type and admission before deterministic family projection and identified-set enumeration.
  • Scope: The proposed release contract is retrospective diagnostic evidence and does not claim how many future stops it would prevent.
  • Scope: Controls outside the Inspect census hold predictions or support fixed while varying a declared semantic coordinate.

L.2 SYMMETRIC TERMINAL CONTROLS

Symmetric terminal controls separate effect typing from ranking structure: contrast-comparable effects can collapse to zero, while rankings may remain non-identified alongside stable winners and relations.

  • Symmetric terminal controls: Every PageRank-minus-Earliest adoption effect is exactly zero under symmetric terminal handling across four backgrounds and three disposition policies.
  • Effect typing: The asymmetric effects are all negative, but adding symmetric zeros makes strict negativity non-identified while preserving non-positivity.
  • Ranking structure: Exclude-both yields five weak orders but Mini remains the unique winner in every cell.
  • Ranking structure: In the original eight-contract family, two weak orders coexist with 20/21 stable method pairs and unchanged top-two positions for Mini and Terminal Event Only.
  • Effect typing: The fixed-support identified set is {−7.72, −3.46} pp, so exact effect is non-identified while its negative direction remains identified.
  • Cross-corpus scope: The baseline design reports −2.40 pp for AgentRx, −10.42 pp for TELBench, and a +8.02 pp frozen interaction, while these corpora are heterogeneous sources rather than independent replications.
Loading 2608.19269v3…