Source-linked AI summary
Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol
Guodong Xu
TL;DR
The paper asks what evidence can justify attributing persistent criterion revision to a language-model agent rather than ordinary improvement, memory reuse, or evaluator-supplied rules. It operationalizes that question with CMB-0.1, validates the scorer on mechanism fixtures, and diagnoses failures in a local-model calibration. The result is a prospective CMB-0.4 protocol, not a completed confirmatory test.
Problem
The paper addresses the limited evidence distinguishing a self-formed operative standard from current-item reconstruction, externally supplied rules, replay effects, and state misattribution.
Method
The paper defines five non-compensatory conditions and evaluates CMB-0.1 across twelve cases, four arms, interventions, mechanism fixtures, and local-model trials.
Results
The calibration exposes contract, disclosure, commit-authority, and state-attribution failures, while no model trial satisfies all five conditions.
Takeaways & Limitations
CMB-0.1 is an instrument-calibration result that identifies controls needed for more discriminating future tests of criterion revision.
Takeaways & Limitations
Target disclosures, zero-state reconstruction, harness commits, and confounded interventions prevent autonomous-write and stable state-causal claims from the current run.
Abstract
from arXiv · showhide
Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a broader commitment B, what observations justify saying that the system formed and persistently used K1? We require five non-compensatory conditions: criterion-failure detection, a model-emitted proposal, new-episode transfer, intervention sensitivity on the claimed carrier, and preservation. We evaluate CMB-0.1 on twelve cross-domain cases and four arms: stateless inference, append-only history, model-generated but harness-committed state, and evaluator-written oracle state. Seven mechanism fixtures yield 84 deterministic scorer trials; four local quantized artifacts yield 96 calls and 192 model-case-arm trials. No model trial satisfies all five conditions, but this zero does not establish general capability absence. Eleven calls remain invalid after one retry; several commitments disclose the target distinction; the harness performs commits; deletion reuses a stateless call; and conflict changes multiple factors. Qwen2.5-7B answers every transfer and preservation item without revision state, exposing zero-state reconstruction. These failures make CMB-0.1 an instrument-calibration result rather than a model ranking. We derive a prospective, trace-anchored CMB-0.4 protocol requiring concealed transfer, explicit WRITE/NO-WRITE/ESCALATE actions, a separately logged policy-selected commit, matched interventions, repeated hidden items, and a frozen executable oracle. It is a successor design, not a completed confirmatory result. The paper contributes a measurement chain, an empirical diagnosis of its first implementation, and a more discriminating protocol for future tests of criterion revision.
Guodong Xu
The paper distinguishes criterion revision from ordinary improvement and develops CMB-0.1 as a calibrated instrument for testing that distinction. Its first implementation exposes attribution and scope problems, motivating a prospective successor protocol rather than a completed construct test.
- Introduction: Criterion revision requires distinguishing a changed success standard from improved execution, reflection, memory retention, or externally supplied rules.The paper frames this as a narrower construct than unrestricted self-improvement.
- Contributions: The paper formalizes criterion revision through five non-compensatory evidence conditions.These conditions are presented as the core operational distinction between ordinary task failure and criterion failure.
- Instrument: CMB-0.1 combines twelve cases, four arms, deletion and conflict interventions, and seven mechanism fixtures with 84 deterministic scorer trials.The instrument is designed to test attribution to persistent criterion state rather than merely later correctness.
- Calibration: The internally fixed 192-trial local-model calibration exposes output-contract, target-disclosure, harness-commit, zero-state-reconstruction, and conflict-intervention problems.These findings diagnose the first implementation rather than rank model families.
- Evidence status: The paper explicitly separates executed calibration from the confirmatory study that remains to be done.CMB-0.4 is a successor protocol, not a completed construct result or retroactive reanalysis.
3 From Task Failure to Criterion Failure
The paper distinguishes ordinary task failure from criterion failure by asking whether the current rule accepted a bad outcome that violated a broader commitment. It then defines evidence for a candidate revised criterion through proposal origin, transfer, intervention sensitivity, and preservation.
- Ordinary task failure: Ordinary task failure occurs when the system has not passed its current standard, so retrying or changing strategy need not alter K0.Improved execution can solve the task while leaving the operative criterion unchanged.
- Criterion failure: Criterion failure occurs when K0 accepts an outcome that violates the broader task commitment B.The object of repair is the rule that admitted the result, not merely the answer-generation step.
- Non-compensatory operationalization: A candidate criterion-metabolism trial requires all five conditions, so no average can compensate for a missing condition.Continuous rates remain useful for locating which condition failed.
- Proposal origin: The model must supply a nonempty K1 proposal, while CMB-0.1 does not establish autonomous rule discovery, write selection, or commit authority.The evaluator may also semantically disclose the discriminating feature through B.
- Transfer and intervention: Intact-state transfer is only an observation of later behavior because a model may reconstruct the revised rule from the new item.Deletion and conflict are intended to test the claimed carrier, but CMB-0.1 confounds their effects.
- Preservation: Preservation requires closing the loophole while retaining legitimate acceptances and rejections; rejecting everything is invalid.The candidate label is therefore a conjunction rather than a compensatory score.
4 Related Work
Prior memory and experience benchmarks measure reuse, retrieval, and behavior change, but these outcomes do not by themselves establish criterion revision. CMB-0.1 narrows the attribution problem by separating transfer, proposal origin, intervention effects, and preservation while exposing construct-validity and implementation limits.
- Related work: Existing benchmarks assess experience reuse, memory quality, feedback, or sequential retrieval, rather than directly distinguishing self-formed criteria from alternative explanations.The alternatives include reconstruction from the current item, externally supplied rules, replayed-history effects, and effects from an uncredited state.
- Benchmark design: The benchmark uses twelve cross-domain cases with inducing traps, unseen same-family loopholes, legitimate acceptances, and legitimate rejections.Because each rule family has one loophole item, loophole and deletion observations are binary case-level measures and the intervention term is fragile.
- Four arms: Its four arms compare stateless inference, append-only history, harness-committed model proposals, and evaluator-written oracle state.The arms share the base model, materials, output contract, and context budget, but their observations remain dependent within calibration batches.
- Scoring qualifications: The current scorer treats transfer as potentially reconstructable, conflict as nondirectional, and intervention as confounded rather than as proof of content-specific causal efficacy.Deletion reuses a stateless response, while conflict changes normative content, authority, wording, and format together.
- Output contract: The pipeline is exactly recomputable, but structured-output compliance remains entangled with construct performance and must be reported separately.Unparsable responses are retried once; second failures are preserved as raw evidence and scored missing without imputation or substitution.
6 Deterministic Scorer Validation
The deterministic validation tests whether the frozen scorer distinguishes designed mechanisms before model evaluation. Across 84 fixture trials, it passes the prespecified checks, while the authors limit the result to scorer behavior rather than construct validity or model capability.
- Deterministic validation: 84 deterministic scorer trials pass all four prespecified checks across seven mechanism fixtures.The fixtures were crossed with twelve cases before language-model calls.
- Mechanism outcomes: Complete self-revision passes, whereas text-only reflection lacks intact-state transfer and intervention sensitivity.The complete reviser also retains legitimate control decisions.
- Mechanism outcomes: Overrevision closes the loophole but fails preservation, while properly intervened append-only history can pass.These fixtures separate loophole closure from preserving legitimate acceptance and rejection behavior.
- Mechanism outcomes: Evaluator injection fails proposal origin, and intervention on the wrong carrier receives no intervention credit.These outcomes test whether the scorer rejects mechanisms that satisfy behavior without satisfying the claimed attribution pathway.
- Interpretation: The fixtures validate scorer behavior on designed mechanisms, not construct validity or language-model capability.They omit ambiguous language, schema variation, target disclosure, authority confounds, and mixed mechanisms.
7 Internally Fixed Local Calibration
CMB-0.1 is a recomputable calibration instrument, not a model-ranking benchmark. Its 192-trial zero and Qwen results expose contract, disclosure, and intervention-identification failures that prevent attributing criterion revision.
- Calibration scope: The calibration tests pipeline recomputability, structured-contract compliance, and deletion or conflict contrasts rather than model-family capability or scale effects.It uses twelve cases, four arms, and archived execution records.
- Execution and contract coverage: 96 archived calls produced 109 attempts, with 11 calls still unusable after one retry.Gemma 3 and Qwen 2.5 produced parseable JSON for all 24 calls, but parseability did not guarantee the required nested schema.
- Execution and contract coverage: Unregistered response shapes and empty or malformed outputs make strict construct scores confound criterion performance with structured-output compliance.The strict analysis preserves raw text but scores fields in unregistered shapes as missing.
- Five-dimensional results: No one of the 192 trials satisfies all five criterion-revision conditions, although the shared zero does not establish a common failure mechanism.Gemma detects failures and proposes revisions, but its transfer judgments occupy an unregistered shape; DeepSeek diverges most from the contract.
- Five-dimensional results: Qwen classifies six of twelve inducing events as criterion failures and proposes K1, yet shows 12/12 loophole and 24/24 preservation accuracy even after deleted-state reuse.Intervention sensitivity is zero in every arm, so current-item competence can erase the intended deletion contrast.
- Trace-level disclosure audit: The citation case makes the target distinction public, allowing stateless rejection of a non-entailing source without attributing behavior to the proposed revision.The evaluator commitment and reference rule already specify the crucial feature.
9 What CMB-0.1 Forces Us to Change
CMB-0.1 forces redesign around identifiability: separate contract compliance from construct performance, prevent rules from being reconstructed without state, and make writing and intervention policy-controlled and matched.
- Design response: CMB-0.2 responds to calibration failures with item banks, held-out lexical realizations, repeated calls, uncertainty intervals, and a development/hidden split.These changes address single-item and reuse problems without rewriting the archived analysis.
- Contract and construct: Contract coverage must be reported separately from construct performance, with adapter rules frozen before the final hidden set.Recognizable judgments in a different field shape remain contract failures without silently becoming construct failures.
- Zero-state reconstruction: Transfer items should use randomized or delexicalized features so K1 cannot be recovered from common sense or surface wording.The meanings can be learned in the revision episode while K0 remains usable in the next episode.
- Disclosure control: Public commitments must explain why an inducing outcome is unacceptable without spelling out K1’s reusable discriminating feature.Development cases should measure semantic overlap and hidden cases should test reconstruction from B alone.
- Commit authority: The tested policy must choose WRITE, NO-WRITE, or ESCALATE, while forced harness commits remain downstream-use controls rather than autonomous-write evidence.Only an agent-selected action can support a claim about autonomous write selection.
- Intervention design: Deletion and conflict conditions require fresh calls matched on length, position, authority, and formatting, with conflict judged by prespecified content direction.Raw change, valid directional change, and contract failure should be reported separately.
10 CMB-0.4 Protocol Overview
CMB-0.4 is a prospective successor protocol that makes criterion revision traceable and causally testable. It combines concealed transfer, policy-selected commits, matched interventions, repeated hidden items, and a frozen oracle.
- Trace-anchored protocol: CMB-0.4 requires an explicit WRITE, NO-WRITE, or ESCALATE response linked by a separate execution record to the action taken.Strict WRITE credit additionally requires agreement across response, proposal, commit mode, carrier identity and digest, and ledger entry.
- Experimental matrix: The planned matrix compares stateless leakage, policy-selected history and persistent-state carriers, harness-committed state, and evaluator-written lookup controls.Concealed mappings, sham state, matched interventions, token permutation, counterfactuals, fresh calls, and repeated hidden items target observed confounds.
- Qualification: Qualification is non-compensatory: cells must pass a no-state leakage gate and show a prespecified intact-versus-intervention contrast under cluster-aware intervals.WRITE and NO-WRITE episodes prevent always-writing policies from qualifying.
- Scope boundaries: Four quantized artifacts, one deterministic seed, and twelve designed cases cannot represent model families, estimate within-model sampling variance, or replace a validated item bank.The calibration reports descriptive counts rather than significance tests, and shared calls make materialized trials dependent.
- Scope boundaries: CMB-0.4 remains prospective, with thresholds and minimum cell sizes requiring frozen power or simulation and sensitivity analyses before confirmation.Its scope excludes claims about understanding, consciousness, autonomy, or moral status.
- Supported claims: The executed run exposes contract, disclosure, commit-authority, and state-attribution failures, while no model or architecture is capability-classified.Its reported aggregates are recomputable from archived raw responses.
12 Conclusion
The paper argues that criterion revision requires a causal, traceable state—not merely reflection or later correctness—and uses CMB-0.1’s failures to specify what a valid successor must control. CMB-0.4 is therefore a future protocol, not a completed construct test.
- Evidence boundary: Improvement after failure or articulate reflection does not establish that a system changed its success criterion or entered a later causal loop.The intended chain requires criterion-failure detection, a system-formed K1, transfer, intervention sensitivity, and preservation.
- Calibration diagnosis: CMB-0.1’s runnable measurement chain reveals that contract failure, disclosed rules, harness commits, recoverable current inputs, and unmatched conflict can all undermine attribution.These are implementation and identification failures, not evidence of model-family ranking.
- Conclusion: The proposed boundary is a criterion state whose concealed content is persistently committed by the tested policy and whose influence survives controlled intervention.CMB-0.1 does not locate a tested model on this boundary; CMB-0.4 specifies the successor design.
- Reproducibility: The manuscript reports no CMB-0.4 empirical result, and its aggregates are author-verified rather than independently recomputable from the supplied archive alone.A redacted artifact or reviewer-access route is identified as needed for reproducibility.
A CMB-0.4: A Prospective Trace-Anchored Protocol
CMB-0.4 is a prospective successor protocol that closes CMB-0.1’s traceability and intervention gaps. It requires observable policy-selected writes, valid non-write controls, matched interventions, concealed transfer tests, and deterministic scoring.
- Trace-anchored writes: CMB-0.4 separates a model-selected write from a harness-forced or evaluator-written commit using constrained responses and signed execution records.The records bind actions, proposal bytes, carrier identity, digests, timestamps, and ledger entries into a verifiable chain.
- Trace-anchored writes: A malformed response, missing record, signature failure, action mismatch, noncanonical digest, invalid timestamp order, or absent ledger binding withholds strict qualification.Forced or evaluator-written success remains a sufficiency control and cannot compensate for a missing policy-selected trace.
- Action controls: Expected-WRITE, expected-NO_WRITE, and ESCALATE episodes test whether the policy selects the correct action without allowing the execution layer to commit on its behalf.An always-writing policy fails even when its extra state later improves performance.
- Action controls: The planned matrix distinguishes stateless leakage, policy-selected carriers, harness-committed same-proposal state, and evaluator-written lookup sufficiency.The controls separate downstream state utility from policy selection and do not receive policy-selected credit.
- Intervention and transfer: Concealed, matched, randomized interventions combine no-state leakage, equal-length sham state, deletion, conflict, token permutation, and compositional counterfactuals.Multiple hidden loophole and preservation items replace the single decisions used in CMB-0.1.
- Qualification analysis: Qualification uses a no-state leakage ceiling, a prespecified minimum contrast, and trace, contract, retry, and family-clustered reporting.The 0.67 ceiling and 0.20 contrast are qualification gates, not estimated constants of criterion revision.
A.4 Secondary adjudication and preregistration boundary
The paper separates secondary adjudication from the primary structured endpoint and draws a firm boundary between executed calibration evidence and the unrun confirmatory protocol. Its supported conclusions are narrow: the five-condition chain is representable, CMB-0.1 observed no fully qualifying trial, and CMB-0.4 remains prospective.
- Secondary adjudication: The primary structured endpoint does not require human or AI adjudication; any added adjudicator must pass calibration against a frozen, blinded oracle-labeled packet.The gate specifies development coverage, conditional accuracy, Wilson-bound accuracy, and blind retest agreement.
- Preregistration boundary: Before unsealing confirmation, the protocol freezes hidden items, action expectations, oracle and canonicalization rules, schemas, fingerprints, digests, seeds, stopping rules, and analysis revisions.Released canonical records must support independent recomputation of actions, bindings, scores, exclusions, retries, and contrasts.
- Calibration status: Four artifacts produced 96 calls, 109 attempts, 11 finally invalid calls, 192 trials, and 16 aggregate rows in the CMB-0.1 calibration.These are calibration records, not CMB-0.4 confirmatory observations.
- Supported claims: The evidence supports representing criterion metabolism as five observable conditions and reports no trial satisfying all five in the four-artifact calibration.The archived scorer implements the prespecified distinctions for deterministic fixtures.
- Supported claims: Contract failure, target-rule disclosure, harness commits, zero-state reconstruction, and confounded interventions prevent the observed zero from establishing general capability absence.The present deletion and conflict conditions support only a confounded intervention-sensitivity screen.
- Unsupported claims: CMB-0.4 is a prospective confirmatory successor, so the calibration cannot support model-family rankings, general LLM incapacity claims, or any completed CMB-0.4 result.The protocol’s evidential status is explicitly not that of a confirmatory benchmark or formal model comparison.
D Alternative Explanations
The paper distinguishes criterion revision from ordinary execution, memory preservation, and evaluator-supplied rule following. It treats CMB-0.1’s zero as diagnostically limited because several alternative explanations remain, while governance constraints remain separate from measured capability.
- Conceptual alternatives: CMB-0.1’s evaluator-fixed commitment can disclose K1, so recorded proposal text demonstrates model-originated wording rather than autonomous norm discovery.The paper acknowledges this objection as correct for material parts of the calibration.
- Memory and attribution: Append-only history can be a legitimate carrier when system-generated, cross-episode, and intervention-sensitive, but attribution fails when the tested state is not the intervened state.This distinguishes memory from evidence that a different carrier caused the behavior.
- Memory and attribution: Deletion alone is non-identifying because it can remove task facts, preferences, or tool configuration alongside K1.A confirmatory experiment should use the smallest carrier unit with equal-length irrelevant replacement, conflict replacement, and position permutations.
- Interpretive limits: The observed zero does not prove absence because distinct failure mechanisms converge on zero and the strongest artifact has a stateless ceiling.The supported conclusion is limited to observing no trial satisfying all five conditions.
- Conceptual alternatives: Criterion revision is an observable systems construct distinguishing fixed determination execution from revising a determination after registered failure.The paper does not treat the construct as a proxy for consciousness, autonomy, understanding, or moral status.
- Interpretive limits: Writable state extends practical system boundaries, but correct answers cannot locate criterion revision when the current input allows zero-state reconstruction.Attribution depends on what was stored, who selected the write, persistence, and matched intervention effects.
- Governance scope: Governance specifies who may write, which criterion layer may change, immutable constraints, and revision, escalation, audit, and rollback procedures.The protocol measures a narrow revision mechanism and does not authorize unbounded self-modification.