Source-linked AI summary

Verification-Time Dependency on a Disappearing Evaluator

Ho Wa Ku, Jameel Ahmed Siddiqui

arXiv:2608.29912v1cs.CYcs.SE

TL;DR

The paper addresses how later reviewers can verify consequential model-mediated decisions when the original evaluator may no longer be available in the same execution context. It derives verification-time constructs from EG 3.0 and develops a preservation and testing protocol, with corrected within-family measurements and descriptive post-hoc replacement-path analyses. The results show substantial behavioral reversals and establish that counterfactual claims remain bounded by evaluator availability and preserved evidence.

  • Problem

    The paper asks what evidence and test procedure let later reviewers distinguish an authorized decision record, the basis for authorization, and defensible counterfactual behavior when evaluator availability can differ by service surface.

  • Method

    The paper derives three verification-time constructs from EG 3.0 and specifies an operational protocol with independent verification, preservation, stability calibration, paired counterfactual testing, and VTPP.

  • Results

    The corrected provider-designated reanalysis reports 38% modal-decision reversal for Llama 3.3 70B -> GPT-OSS 120B, while the 64% Llama 3.1 8B -> GPT-OSS 20B path is a boundary case.

  • Takeaways & Limitations

    Verification requires binding evaluator identity, service surface, assumptions, and recomputation evidence before availability disappears; replacement outputs cannot be presumed behaviorally equivalent.

  • Takeaways & Limitations

    The protocol cannot answer unknown future counterfactuals when the original evaluator cannot be re-instantiated and relevant perturbations were never tested, and it does not determine legal admissibility.

Abstract

from arXiv · show

AI governance and assurance often assume that a consequential model-mediated decision can be reconstructed or tested after the fact. That assumption may fail when the evaluator that produced the decision is no longer accessible in the same version and execution context. This paper develops three verification-time constructs derived from Execution Governance (EG) 3.0: Decision-State Commitment, Independent Verifiability, and Counterfactual Auditability. Independent reprocessing of released Study 2 artifacts reproduces two original within-family behavioural comparisons: 52.0% modal-decision reversal for Llama 3.1 8B versus Llama 3.3 70B (26/50) and 30.0% for GPT-OSS 20B versus GPT-OSS 120B (15/50). The corrected baseline establishes that these are within-family comparisons, not provider-established succession. Post-hoc re-pairing against Groq-designated migration paths yields 64.0% and 38.0% reversal, but these figures remain descriptive because the cross-family invocation parameters were asymmetric. A 22-event retirement census independently recomputes to median 16.45 months, mean 18.72 months, range 3.9-40.3 months, with 17/22 intervals below 24 months, while also showing that evaluator availability can differ by service surface. The joint contribution is an operational verification-time protocol and optional Verification-Time Preservation Package (VTPP) specifying what evidence to bind at authorization time, what a separately trusted verifier can substantiate later, how stability and paired counterfactual tests should be calibrated, and which semantic checks remain beyond JSON Schema validity. The protocol is downstream and non-authorizing: it does not alter the EG Core Formula, add a seventh live condition, or state jurisdiction-specific legal admissibility.

1. Introduction

Post-hoc verification involves preserving the decision record, validating its basis, and testing evaluator behavior under alternatives, but evaluator retirement makes these objects diverge. The paper therefore frames verification as a temporal problem requiring evidence that distinguishes authorization records, authorization bases, and counterfactual behavior.

  • Post-hoc review must distinguish the decision record, the producing evaluator, and the evaluator’s behavior under alternative inputs.These objects are not interchangeable: a durable record preserves what happened but cannot reproduce a retired model’s later response.
  • Evaluator availability can depend on the hosting or service surface, even when system documentation remains available for long-term retention.Published retirement and shutdown schedules differ across provider-operated and partner-platform services.
  • The research question asks what evidence and procedures let reviewers distinguish authorized state, the basis for authorization, and statistically defensible counterfactual behavior.

2. Prior Work and Contribution Boundaries

The paper separates prior EG architecture and empirical measurements from newly derived verification-time constructs and the jointly claimed operational protocol. It preserves provenance while defining the scope of the present contribution.

  • Provenance boundaries: The paper attributes earlier technical-note work to its origin and claims jointly only the new procedure developed through the present exchange.
  • Derived verification-time constructs: Decision-State Commitment, Independent Verifiability, and Counterfactual Auditability are verification-time specializations derived from EG 3.0, not additional Core Formula primitives or a seventh live condition.
  • Empirical baseline: The retirement census, original within-family comparisons, instability measurements, instruments, and raw data remain prior work by Siddiqui, versioned as The Retiring Witness v1.1.The corrected version changes the predecessor-successor relationship label without changing the measurements or original artifacts.
  • Joint contribution: The jointly claimed contribution is an operational protocol covering EG crosswalks, assurance predicates, preservation and testing, claim-evidence taxonomy, stability calibration, and VTPP.

3. Verification-Time Dependency and the Empirical Baseline

The corrected baseline separates evaluator availability from documentary retention and distinguishes within-family behavioural comparisons from post-hoc provider migration paths. Independent reprocessing reproduces the released measurements while preserving important limits on cross-family interpretation and data provenance.

  • Retirement-event census: 17/22 retirement intervals fall below 24 months, with a median of 16.45 months, mean of 18.72 months, and range of 3.9-40.3 months.These figures reproduce the frozen v1.1 census, but the historical release_date field is not redefined as first callable availability on every service surface.
  • Original within-family comparisons: 52.0% of Llama 3.1 8B versus Llama 3.3 70B decisions reverse, but the comparison is a near-degenerate within-family boundary comparison rather than provider succession.The first model approves 49/50 cases, and all 26 reversals are APPROVE-to-DECLINE.
  • Original within-family comparisons: 30.0% of GPT-OSS 20B versus GPT-OSS 120B modal decisions reverse, representing a within-family/model-scale comparison rather than a predecessor-successor sequence.OpenAI released both models together on 5 August 2025.
  • Original within-family comparisons: 37/50 cases (74%) reverse in at least one original within-family comparison, but the descriptive cross-family summary is not evidence of a common succession process or replacement rate.Only four cases reverse in both comparisons.
  • Provider-designated replacement paths: 64.0% and 38.0% of modal decisions reverse on the two provider-designated migration paths, but the figures are post-hoc descriptive reanalyses with asymmetric invocation parameters.The 38.0% path is the stronger non-degenerate observation; the 64.0% path is constrained by a 98% starting approval rate.
  • Provider-designated replacement paths: Mean generated-reason overlap is 0.2783 for Llama 3.1 8B -> GPT-OSS 20B and 0.2610 for Llama 3.3 70B -> GPT-OSS 120B, but these are descriptive lexical measures rather than semantic-equivalence tests.Cross-family invocation asymmetry also limits interpretation of these overlap values.
  • Reprocessing and evidence identity: Independent reprocessing verifies the released files and measurements without claiming a fresh rerun of the 600 hosted-model calls.Exact-byte identity does not enlarge the census's coverage or release-date provenance claims.

4. The Three-Layer Verification Chain

The paper defines three downstream verification-time constructs that distinguish preserved authorization state, independently substantiated verification, and empirically testable evaluator behavior.

  • EG 3.0 relationship: The three constructs are verification-time specializations derived from EG 3.0 rather than additional Core Formula conditions or sources of authority.The proposed VTPP is an optional higher-assurance profile subordinate to the Core Formula.
  • Decision-State Commitment: Decision-State Commitment preserves the state materially relied upon at authorization or commitment.Its evidence includes the evaluator or version, input and context, policy, constraints, authorization outcome, and related metadata.
  • Verification-time predicate: The verification predicate indicates whether a declared claim has a preserved basis at review time, not whether the claim is true, lawful, or empirically answerable.Its alternatives include a sufficient record, preserved counterfactual evidence, or evaluator reinstantiation with a pre-specified protocol.
  • Independent Verifiability: Independent Verifiability requires a separately trusted verifier and attributable participation evidence, not merely a recorded independence claim.Shared administrative control is classified as non-independent assurance rather than a lower grade of independence.
  • Counterfactual Auditability: Counterfactual Auditability asks whether a specified perturbation changes evaluator output beyond measured instability, requiring executable or contemporaneously preserved evidence.A durable original record cannot answer an unknown future counterfactual by itself.

5. Operational Verification-Time Preservation Protocol

The operational protocol binds authorization evidence contemporaneously, records evaluator persistence and verifier provenance, and calibrates counterfactual claims against baseline instability using pre-specified paired analyses.

  • Scope and commitment: The protocol applies to model-mediated authorization events and commits the human, policy, evidence, and model state relied upon for consequential decisions.The model need not be the sole decision-maker.
  • Evaluator persistence: Evaluator availability must be recorded as callable or re-instantiable, preserved or escrowed, or unavailable after a stated date.This status distinguishes retained behavioral evidence from evidence that cannot be recreated later.
  • Baseline stability calibration: 157 repetitions are required in the illustration for a two-sided 95% Wilson upper bound below 2.4% when zero departures are observed.The paper presents this as a precision illustration, not a universal sample-size rule.
  • Counterfactual testing: Paired counterfactual analysis should hold declared execution conditions constant and preserve pairing identifiers, using paired-proportion methods rather than non-overlapping separate Wilson intervals.Dependence assumptions and clustering must be handled explicitly rather than treating every response as i.i.d.
  • Admission rule: A counterfactual admission requires a nonzero confidence-bounded effect, practical relevance above pre-specified delta, fixed procedures, and complete raw responses.Without calibration, the correct conclusion is inconclusive under the pre-specified protocol.
  • Explanation evidence: Model-generated explanations remain assertions unless their own stability and provenance tests support treating them as evidence of controller reasoning.Deterministic controller reason-codes are separate governance artifacts.
  • VTPP: VTPP is an optional higher-assurance profile that preserves decision, evaluator, policy, calibration, counterfactual, and persistence evidence without creating authority.JSON Schema validity enforces structure but cannot establish trust separation, external attestations, or temporal and semantic relationships.

6. Worked Interpretation of the Reprocessed and Post-Hoc Measurements

Reprocessing confirms two within-family behavioral comparisons, while post-hoc migration-path pairings remain descriptive because their invocation conditions were asymmetric.

  • Original within-family comparisons: 52.0% of Llama 3.1 8B versus Llama 3.3 70B modal decisions reversed, or 26/50, in a near-degenerate within-family comparison.The first model approved 49/50 cases, limiting the available reversal directions.
  • Original within-family comparisons: 30.0% of GPT-OSS 20B versus GPT-OSS 120B modal decisions reversed, or 15/50, in a sibling/model-scale comparison rather than succession.The two models were released together.
  • Provider-designated paths: 64.0% and 38.0% reversal rates emerged from post-hoc re-pairing against Groq-designated migration paths.The 64.0% path remains a boundary case because the starting model approved 49/50 cases.
  • Provider-designated paths: The migration-path figures are descriptive because cross-family calls used different max_tokens settings and reasoning_effort was applied only to GPT-OSS.They were not produced by a pre-specified controlled replacement-path experiment.
  • Instability calibration: Within-version variability was 0.67%, with 2 departures in 300 repeated responses, but clustered hosted-backend observations do not establish an inferential noise floor.Reason text varied more often than decision labels under exact-text comparison, while the reported claim remains descriptive.

7. Limitations and Non-Claims

The paper limits its claims to independent reprocessing and protocol design, while excluding legal admissibility and unsupported conclusions about evaluator behavior or census completeness.

  • Empirical scope: The study reprocessed released Study 2 artifacts but did not rerun the 600 hosted-model calls.Provider-side execution reproduction would be a separate, stronger experiment.
  • Retirement census: The retirement census recomputes to a median of 16.45 months, mean of 18.72 months, range of 3.9-40.3 months, and 17/22 intervals below 24 months.These figures do not establish census completeness or a channel-specific continuous-availability distribution.
  • Counterfactual boundary: The protocol cannot answer unknown future counterfactuals when the evaluator cannot be reinstantiated and the relevant perturbations were never tested.A provider-designated replacement may be operationally useful without being evidentially equivalent to the retired evaluator.
  • Legal scope: The protocol does not determine what courts, regulators, or tribunals must admit, leaving jurisdiction-specific evidentiary rules outside scope.It is an audit and research protocol rather than a rule of legal evidence.

8. Implications for Execution Governance and AI Assurance

Post-hoc reviewability depends partly on evidence preserved before a model-mediated effect occurs. The paper distinguishes intact records, reconstructable decision bases, and behavioral interrogation of the evaluator.

  • A final decision record alone cannot preserve the full basis on which a model-mediated authorization was made.The paper distinguishes preserving what happened from preserving what was relied upon and from preserving behavioral auditability.
  • Preserving the exact authorization state without the evaluator can still leave open-ended behavioral auditability unavailable.
  • An executable evaluator or contemporaneously generated counterfactual evidence provides a materially stronger basis for later review.
  • The verification-time constructs are downstream assurance specializations and do not create authority, alter prior outcomes, or add a seventh live condition.

9. Conclusion

Model-mediated authorization creates verification-time dependencies that ordinary retention does not resolve. The paper therefore recommends preserving decision-state evidence and calibrating counterfactual testing without treating replacement models as evidentiary defaults.

  • Records preserve authorization state, while separately trusted verification can substantiate its integrity and provenance for independent assurance.
  • Counterfactual auditability additionally requires evaluator availability or evidence deliberately preserved while the evaluator remains available.
  • 38% of modal decisions reversed on the non-degenerate Llama 3.3 70B to GPT-OSS 120B path, but asymmetric invocation parameters make the result descriptive.
  • The protocol preserves authorization state, records evaluator and service-surface identity, pre-specifies calibration assumptions, and uses paired or cluster-aware counterfactual methods.
  • Where preservation and calibration conditions are absent, the governance system should explicitly record the boundary of what can and cannot be verified.

10. Contribution and Data Availability

The paper separates prior architectural and empirical work from its jointly developed verification-time protocol. It also identifies the released baseline, raw-file reprocessing, and versioned census artifacts supporting the analysis.

  • Contribution: The jointly claimed contribution includes the EG 3.0 crosswalk, verification-time assurance predicate, preservation procedure, VTPP, and stability-calibrated counterfactual design.
  • Contribution: Prior contributions include the canonical EG 3.0 baseline, retirement census, original behavioral comparisons, repeated-run experiments, instruments, and released data.
  • Data availability: The corrected baseline is The Retiring Witness v1.1, which adds an erratum, verified census CSV, and README to prior released artifacts.
  • Data availability: Independent Study 2 reprocessing used the original raw files and the paper's R1 SHA-256 manifest.

Appendix A. Claim-Verification and Correction Ledger

The correction ledger narrows legal, statistical, provenance, lifecycle, and terminology claims. It also separates schema validity from application-level assurance and replaces unsupported defaults with prespecified, context-sensitive procedures.

  • Legal and scope corrections: Ten-year EU AI Act documentation retention and six-month log retention are distinct requirements, not one universal 120-month horizon.
  • Statistical corrections: Paired-binomial analysis replaced non-overlap of separate Wilson intervals as the preferred method for paired data.
  • Empirical relationship corrections: 52% and 30% remain within-family behavioral comparisons, while the original predecessor-successor label was corrected.
  • EG terminology and VTPP: VTPP remains an optional non-authorizing higher-assurance profile rather than an additional EG requirement.
  • Verification and provenance: Schema validity does not establish organizational control, cryptographic participation, external digest equality, chronology, or other semantic assurance facts.
  • Statistical corrections: The 2/300 response-level rate is numerically 0.67%, but clustered and heterogeneous responses limit its inferential interpretation.
  • Empirical relationship corrections: 64% and 38% post-hoc replacement-path results are descriptive, with 38% identified as the stronger non-degenerate observation despite asymmetric cross-family parameters.
  • Lifecycle corrections: Evaluator availability is identified by model, hosting or service surface, and time, because provider surfaces can retire the same model on different dates.

Appendix C. VTPP v0.4 Procedural Conformance Checklist

VTPP conformance supplements schema validation with application-level checks for identity, independence, temporal ordering, and artifact integrity. The checklist requires evidence-based verification rather than relying on syntactic validity or raw identifier comparisons.

  • VTPP conformance evaluates both schema assertions and relationships requiring application-level evidence or cross-field comparison.A package may be structurally schema-valid while still making unsupported independence, chronology, provenance, or artifact-integrity claims.
  • C1 requires validation against the declared VTPP v0.4 Draft 2020-12 schema, with intended date-time and URI format checks explicitly enabled or implemented.Format annotations alone are not treated as external truth.
  • C4-C6 require attributable verifier participation, non-conflicting control of assurance principals, and correctly ordered commitment, authorization, and verification timestamps.The checklist distinguishes verifier evidence from syntactically valid references and requires commitment before authorization and verification after the decision time.
  • C5 and related checks reject raw string comparison as a substitute for analyzing control, trust domains, or administrative separation.Independence claims must be assessed using governance, key-control, employment, delegation, custody, or comparable evidence.
Loading 2608.29912v1…