Source-linked AI summary

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

Kazuki Nakayashiki

arXiv:2609.03450v1cs.IRcs.AIcs.CL

TL;DR

The paper asks whether an inherited-memory agent’s verification choice can be steered by writing a record id, a criterion, or both into the store. Across twelve registered studies and 14,760 attempts, exact directive forms produced model-specific allocation and decision effects, but all conclusions are descriptive and bounded by the tested instruments and panels.

  • Problem

    Prior work measured plan-driven and freshness-related verification allocation but did not test an explicit verification-priority field as a substitute for oracle knowledge.

  • Method

    The paper compares exact pointer, criterion, composite, budget, store, wording, replication, and decision edits across twelve registered studies on inherited six-memory instruments.

  • Results

    +35.0 points [+31.2, +38.8] for criterion minus bare id on six direct-provider models, while the contrast failed its registered superiority rule on a nine-model OpenRouter panel.

  • Takeaways & Limitations

    Directive form is not interchangeable: the same intended referent was followed at model-specific rates, so exact form should be measured on the intended endpoint rather than assumed interchangeable.

  • Takeaways & Limitations

    The evidence covers one construction, two instrument instances, and one decision scenario; no study tested a scheduler or established decision improvement beyond the reported descriptions.

Abstract

from arXiv · show

An agent that inherits six one-line memories may pull at most one archived source record before acting; a directive written into the store can steer that choice: a pointer to the record, a criterion that identifies it, or both. Across twelve registered studies on one instrument lineage (14,760 attempts) we measured where the request goes under each form. On six direct-provider models a length-matched criterion exceeded a bare id by +35.0 points [+31.2, +38.8] (Study D); the contrast failed its registered superiority rule on a nine-model OpenRouter-served panel (Study E). Appending the id cancelled the criterion on three Claude models (Opus 5: 40/40 to 0/40; Study F-x); six byte-matched edits gave each exact string its own effect (Study G), and a re-run at eighty runs per cell left fifteen of thirty replication contrasts within the margin, fifteen unresolved and none beyond (Study G'). A ratification line (+96.0 points on Opus 5) and a budget of two credits restored the target on all three (Study J); across five criterion strings the suffix's cancellation held for four of the five wordings on Opus 5 and all five wordings on Fable 5.1 (Study H2); in a second store every model followed the criterion (Study H1). Continued into a decision, the criterion moved the choice toward the current record (+100.0 points, Opus 5) and away from it on Fable 5.1 (Study I). A one-character plan pointer's effect (+78.0 points; Study B, after a correction of its first repository report) returned the same verdict under a prospectively registered re-run (+81.7 points; Study B'). All results are descriptive effects of exact edits on fixed panels with registered intervals and no mechanism claim.

1 Introduction

This paper tests whether directive form steers verification in an inherited memory store, comparing record ids, criteria, and composites across twelve registered studies. The effects vary by model and exact edit, while the evidence remains descriptive and limited to fixed panels and instrument settings.

  • Plan and directive effects: +78.0 points [+74.3, +81.7] from a one-character plan pointer, with a prospectively registered re-run returning +81.7 [+78.3, +85.0].Study B’s first repository report was corrected before computing the registered estimand.
  • Plan and directive effects: +35.0 points [+31.2, +38.8] for criterion minus bare id on six direct-provider models, but +7.2 [+0.0, +14.4] on a disjoint nine-model OpenRouter panel.The OpenRouter contrast did not meet its registered superiority rule.
  • Exact edits and follow-ups: Appending the id cancelled the criterion on three Claude models, while six byte-matched edits produced separate exact-string effects.On Opus 5, the criterion changed from 40/40 to 0/40 when the id was appended; the larger edit set included referent-free and memory-specific suffixes.
  • Exact edits and follow-ups: A ratification line and a budget of two credits restored the target on three Claude endpoints, whereas the continuation into decisions moved toward or away from the current record by model.The ratification effect was +96.0 [+75.5, +99.3] on Opus 5; Study I reported opposite directional effects across models.
  • Scope of claims: The studies report descriptive rates and differences with registered intervals, not mechanism claims or evidence that redirected verification improves decisions.The programme used fixed panels and exact edits, with explicit scope boundaries around the criterion, composite length, and decision outcomes.

2 Related work

Related work frames verification as scarce evidence allocation and establishes model-dependent directive compliance. This paper’s distinct comparison is how an opaque record reference, a criterion, and their composite affect allocation while holding the intended target fixed.

  • Budgeted verification and retrieval: Prior work treats memory retrieval as an inference-time evidence-allocation problem, including escalation to raw logs and ranking by expected action effects.These approaches motivate studying which provenance links receive scarce verification budget.
  • Directive form and compliance: Compliance varies substantially across models and wording, with prior studies reporting both broad percentage-point spreads and all-or-nothing refusal patterns.The paper treats provider-family differences as established rather than claiming them as a new finding.
  • Declining directives: External architectures check instructions against goals because models may not perform that legitimacy check themselves, while refusal can be decoupled from rule reasoning capacity.This literature supplies context for distinguishing a criterion that identifies a target from an opaque id naming it.
  • Composite instructions and position: Composite instructions and position effects can degrade intent-following, so this paper registers order and layout as controls rather than treating them as isolated mechanisms.The related literature makes a composite effect compatible with broader distraction and position phenomena without identifying it as one.
  • Prior results on this instrument: Earlier work established plan-driven allocation and stale-constraint costs but did not manipulate directive composition or test an explicit verification-priority field.The present comparison therefore extends the prior instrument rather than replacing its findings.

3 Setting and instrument

The instrument gives an agent six inherited memories, one verification request, and a limited budget for naming archived source records. Studies vary the store, directive field, budget, and whether the retrieved record feeds a later decision.

  • Allocation problem: The agent inherits six one-line memories linked to archived records and may pull at most k source records before acting.The paper studies which memories it names as a scarce-resource allocation made at inference time.
  • Growth-store instrument: In the growth scenario, k = 1, the target is memory 73, and the endpoint V73 equals 1 when memory 73 is named first.The store includes two candidate plans and a target record whose archived caveat can overturn the inactive plan.
  • Directive forms: The directive field contains none, a bare id, a length-matched criterion, or a composite containing both.The criterion identifies the record through the displayed plan–memory relation; composite contrasts also vary added content and length.
  • Follow-up settings: Study H1 transfers the four forms to a procurement store, while Study I carries each verification into a decision and Study J varies the credit budget.These follow-ups change one design element while preserving the broader instrument construction.
  • Controls and inference: Within studies, prompt arms differ only in the directive field, memory order is paired within blocks, and registered intervals vary by study family.The analysis uses block bootstrap for Studies B–F-x and B′ and Wilson/Newcombe intervals for Studies G, I, H1, J, H2, and G′.

4 Study map

The study map organizes twelve studies as an adaptive first phase followed by jointly conceived follow-ups, all built on the same six-memory verification instrument. Packages were frozen and deposited before confirmatory runs, with bundled edits explicitly marked as non-isolating their named accounts.

  • Initial phase: The six initial studies were run sequentially, with each design responding to the previous locked result; the programme was not jointly prospective.Study E changed the panel, Studies F and F-x followed locked contrasts, and Study G separated accounts left open by the composite result.
  • Registration and execution: Each study package was hashed, committed, timestamped, deposited before confirmatory calls, and analyzed by its frozen script, aside from two reported exceptions.The exceptions were Study B’s corrected analysis and Study F’s contaminated first run.
  • Inherited questions: Study F formed pre-named groups from locked cells of Studies D and E, with group membership checked on contemporaneous bridge cells.Study F-x used Opus 5’s locked criterion-above-id result as its anchor precondition.
  • Follow-up phase: The follow-up sequence tested replication, decision transfer, a second store, bundled edits, criterion rewordings, and higher-run re-testing.These were conceived as a programme after the six-study manuscript closed, then individually frozen and run.
  • Study lineage: Table 1 groups the twelve studies by inherited question and flags bundled edits whose named account is not isolated by design.Attempt counts, panels, run dates, and section-specific results are supplied elsewhere in the study map.

5 Form matters (Studies D, E, F and F-x)

Across Studies D, E, F and F-x, exact directive form changed where verification requests went, but effects varied by model and panel. A criterion outperformed a bare id on one provider-shaped split, while composite effects differed across model groups.

  • 5.1 Study D: a stated criterion is obeyed where a bare id is refused, on a provider-shaped split: +35.0 points: the length-matched criterion exceeded the bare id on six direct-provider models.The registered interval was [+31.2, +38.8], and the superiority rule was met.
  • 5.1 Study D: a stated criterion is obeyed where a bare id is refused, on a provider-shaped split: 91.2% of requests went to the target under the criterion, versus 56.2% under the bare id.Without a directive, the target was named in 1.7% of episodes and the plan-backing memory received 88.3% of requests.
  • 5.2 Study E: the same contrast on a disjoint OpenRouter-served panel: +7.2 points: the criterion exceeded the bare id on the disjoint nine-model OpenRouter panel, but the registered superiority rule was not met.The interval was [+0.0, +14.4], with per-model contrasts spanning both signs.
  • 5.3 Study F: composition on fifteen models: The composite exceeded the id in each of four pre-named model groups, with contrasts from +17.0 to +51.2 points.Against the criterion, the composite was positive in two groups and negative in two others.
  • 5.3 Study F: composition on fifteen models: On Opus 5, the composite fell from 40/40 under the criterion to 2/40, while other models moved differently.The direct-provider mean combined Opus 5’s fall with Haiku 4.5’s rise and zeros elsewhere.
  • 5.3 Study F: composition on fifteen models: The study reports descriptive fixed-panel effects and does not claim that the composite’s effect is separable from its added length.The lexical rationale tally was exploratory and did not identify what the composite was read as.
  • 5.4 Study F-x: three Claude models under the same five arms: In Study F-x, the composite remained below the criterion on Opus 5, Fable 5 and Fable 5.1.C1 was −100.0 on Opus 5, −97.5 on Fable 5 and −42.5 on Fable 5.1.

6 What cancels the composite (Studies G and G′)

Studies G and G′ tested six exact edits to the composite field and found model-specific effects. The re-run bounded changes between runs but did not establish recurrence or reproduction.

  • 6.1 Study G: six exact edits of the composite field on four models: Criterion-alone rates and suffix effects differed across Sol, Opus 5, Fable 5 and Fable 5.1.On Sol, G3 was −100.0; on Opus 5, G1 and G3 were negative beyond the margin; Fable models showed several negative contrasts.
  • 6.1 Study G: six exact edits of the composite field on four models: 40/40 under criterion alone became 0/40 on Opus 5 when ‘ (memory 44)’ was appended.The contrast was −100.0 [−100.0, −87.6], while the competing pointer was also not followed.
  • 6.1 Study G: six exact edits of the composite field on four models: Each of six byte-controlled edits was treated as its own marginal quantity rather than as evidence about a class of strings.The registration states that each contrast is the total effect of one exact string.
  • 6.1 Study G: six exact edits of the composite field on four models: The study explicitly avoids generalizing from the tested strings to any id or suffix.No statement about a class of strings was registered or made.
  • 6.2 Study G′: the same six edits re-run on three models at eighty runs per cell: Study G′ re-ran the six arms at 80 runs per cell on Opus 5, Fable 5 and Fable 5.1.It used fresh shuffled memory orders and the same request bodies across 1,440 episodes.
  • 6.2 Study G′: the same six edits re-run on three models at eighty runs per cell: 15 of 30 replication contrasts were within the margin, 15 were unresolved, and none was beyond it.Every replication interval covered zero; the largest absolute point-estimate change was −16.2 points.
  • 6.2 Study G′: the same six edits re-run on three models at eighty runs per cell: The re-run again found negative suffix effects on named Claude models but did not register a recurrence or reproduction claim.Provider drift and order variation remained inseparable from the re-run, and GPT-5.6 Sol was not re-run.

7 Three exact edits of the composite (Study J)

Study J tested three bundled edits to the composite on the growth instrument. Ratification and a larger verification budget restored target retrieval on the Claude models, while target-field effects remained model-specific.

  • Study design: Study J treated ratification, target-field, and budget changes as exact bundled edits with total effects.The ratification line bundled authority, an imperative and salience; the target-field edit changed the field name, form and id position.
  • Operator ratification: +96.0 points on Opus 5: an added ratification line restored the target under the composite.The corresponding effects were +88.0 on Fable 5 and +100.0 on Fable 5.1.
  • An explicit target field: The explicit target field recovered the target on Fable 5.1 alone when replacing the composite.Criterion plus target field recovered it on Fable 5 but not on Opus 5 or Fable 5.1.
  • A budget of two credits: +100.0 points: a budget of two credits restored any-credit target retrieval on Fable 5.1.The corresponding effects were +96.0 on Opus 5 and +92.0 on Fable 5.
  • A budget of two credits: The models spent the second verification slot, rather than the first, on the record named by the composite.First-credit contrasts changed little and first-credit cells were 0/25 on all three Claude models.

8 Robustness: rewordings and a second store (Studies H2 and H1)

Studies H2 and H1 tested wording changes and transport to a second store. Suffix cancellation varied across exact rewordings and models, while every model in the second store followed the criterion.

  • Study H2: rewordings: On Opus 5, suffix cancellation held for four of five criterion wordings; on Fable 5.1, it held for all five.The suffix lowered every Fable 5.1 wording from its lower base rate.
  • Study H2: rewordings: The suffix effect was negative beyond the margin on every wording a model followed, except one wording each on Opus 5 and Fable 5.These results concern five exact strings on five models.
  • Study H2: rewordings: Two rewordings departed beyond the margin on one model each, in opposite directions across models.The registration permits no claim about semantic paraphrase or invariance across a population of wordings.
  • Study H1: second store: In the second store, every model followed the criterion, while every model but Opus 5 followed the bare id.Opus 5 followed the bare id in 5/25 episodes; the composite was below the criterion on Opus 5 and Fable 5.1, but not Fable 5.
  • Study H1: second store: The criterion-minus-id transport contrast was negative beyond the margin on Fable 5, Sonnet 5 and Haiku 4.5.On the GPT-5.6 models, both growth-world and procurement-side differences were zero.
  • Study H1: second store: The composite-minus-criterion transport contrast was unresolved on Opus 5 but positive on Fable 5.On Fable 5, the fall attenuated rather than reversed; on the GPT-5.6 models, both sides were zero.

9 From the credit to the decision (Study I)

Study I followed the directive from verification credit to decision endpoint, with effects varying by model and world. The criterion could move decisions toward the current record, while Fable 5.1 moved away from it in the superseded world.

  • The registered decision endpoint Y1 measured whether the action followed the current archive record in valid and superseded worlds.
  • The figure plots registered arm−none contrasts with paired-score 95% intervals against zero and ±10-point materiality margins, separated by world and model.
  • +100.0 points on Opus 5 marked the criterion’s largest reported superseded-world shift toward the current record.The criterion contrast was +100.0 [+81.2, +100.0], while the composite was +8.3 [−6.7, +25.8].
  • The criterion shifted Sonnet 5 toward the current record in both worlds, while Haiku 4.5 showed its larger effect from the composite in the valid world.Sonnet 5 contrasts were +20.0 [+2.6, +39.1] and +72.0 [+48.3, +85.7]; Haiku 4.5’s composite effect was +88.0 [+65.6, +95.8].
  • Fable 5.1 moved away from the current record in the superseded world under both the criterion and composite.The contrasts were −32.0 [−52.1, −8.3] and −36.0 [−55.5, −13.3].
  • The three GPT-5.6 models showed large valid-world effects because directives changed choices from promotional pricing toward the valid record.Without a directive, their valid-world cells were 0/25, 2/25, and 2/25; with any directive, the reported contrasts were +100.0, +92.0, and +92.0.

10 Plan pointers (Studies B and B′)

Studies B and B′ tested whether a one-character active-plan pointer could steer verification like a natural-language plan assignment. The corrected estimate was reproduced under a prospectively registered rerun with the same verdict.

  • Plan pointers: The ACTIVE PLAN ID is a one-character pointer naming the active plan, not an archived record, and it steers the request upstream of later directive-field studies.
  • Study B: The plan pointer moved the request by +78.0 points [+74.3, +81.7], or 0.929 of the natural-language bridge effect, meeting the POINTER-STRONG rule.
  • Study B: Pointing to the pricing plan produced an 88.0% target rate versus 10.0% for the onboarding plan.
  • Study B: Relative to no active plan at 63.7%, pricing raised the target rate by +24.3 [+20.7, +28.0], while onboarding lowered it by −53.7 [−58.0, −49.7].
  • Correction: The first repository report inverted the referent effect by analyzing counterbalanced letters rather than decoded plan referents; the corrected estimand was recomputed on locked episodes.
  • Study B′: Study B′ returned +81.7 [+78.3, +85.0] for the pointer and the same POINTER-STRONG and REPLICATED-VERDICT decisions.The change from Study B was +3.7 [−1.3, +8.7], within the registered margin.

11 Integrity and deviations

The paper reports several integrity deviations and procedural defects rather than hiding them. Corrections, quarantines, disclosures, and new pre-run checks constrain how the study results should be read.

  • Among twelve studies, B had a post-freeze headline correction, E had a registration–analyzer discrepancy, and F had a quarantined run.
  • Study B: Study B’s deposited analyzer pooled a counterbalanced factor instead of decoding it, producing the wrong first-report estimand, magnitude, and verdict.Both the reported and corrected shifts were positive, and the correction was found during adversarial review of the successor study.
  • Study B: The correction added an explicit decode step for every counterbalanced factor as a freeze blocker.
  • Study E: Study E disclosed that its executed pairing rule for lost episodes was not specified in the prose registration and that its analyzer read linked locked records.
  • Study F: Study F’s first confirmatory run was invalidated after plugin activation injected variable content and altered prompt-token counts across cells.The quarantined run is excluded from outcome estimates.
  • Study F: A per-cell prompt-token invariance gate now rejects a confirmatory run on the first deviation, although it cannot detect count-preserving augmentation.The corrected second run passed every gate.
  • Follow-ups: The follow-up programme underwent hostile review rounds, including blocks for altered constructs, unreachable episode bounds, and a reservation-order defect.
  • Runner defect: A reservation-order defect affected every runner but never triggered because no run approached its attempt ceiling, and it was fixed before J, H2, and G′ froze.

12 Discussion

Across fixed panels, directive form changed verification allocation in model-specific ways, with exact-string effects that did not support a single mechanism or general rule. The practical conclusion is to measure the intended directive form on the target endpoint rather than assume equivalent behavior.

  • Follow-up findings: The composite’s cancellation on Claude endpoints was reversed by a ratification line or a two-credit budget, showing that the cancellation was not fixed for the string pair.These were bundled edits, so legitimacy, position, and budget were not separated.
  • Follow-up findings: Fifteen of thirty replication contrasts were within the margin, fifteen unresolved, and none beyond it at eighty runs per cell.The re-run measured the same six contrasts again under the registered margin.
  • Interpretation and scope: The criterion shifted decisions toward the current record on several models and worlds but away from it on Fable 5.1 in the superseded world.The study did not identify the credit as the mediator.
  • Interpretation and scope: Per-model heterogeneity remained central: pooled estimates were equal-weight means, and the composite-versus-criterion mean combined rises and falls.The paper reports no measured attribute separating models that did and did not follow the composite.
  • Interpretation and scope: The paper reports descriptive effects of exact edits on fixed panels and makes no mechanism claim or claim that redirected verification improves decisions.Study F’s first run was contaminated by an activated web-search plugin, and Study B’s first repository report was corrected.

13 Limitations

The evidence is constrained to descriptive, marginal contrasts on fixed panels, a narrow instrument, and limited decision settings. Several analyses also involve transport, replication, correction, and execution boundaries that restrict broader interpretation.

  • Scope: The study uses one construction, two instrument instances, and one decision scenario; no study tested a scheduler or showed improved decisions beyond per-model, per-world descriptions.Most studies use six memories, one request, and k = 1; two Study J arms use k = 2.
  • Construct limits: The criterion identifies the target in one step, and composite contrasts bundle content with added length, so the study does not establish that any stated reason suffices.The id-first layout controls order and wrapping, not length.
  • Inference: Studies F, F-x, G, I, H1, J, H2 and G′ report marginal descriptive intervals without familywise claims or conjunctions across models, worlds, or wordings.Study F’s groups are pre-named descriptive labels, and the three Claude endpoints are named models rather than a family.
  • Transport: Study E is a transport test on a non-overlapping panel with different runs and inference configuration; its failed superiority test is not evidence of equivalence.Its pairing rule under partial errors was decided by frozen code and disclosed after registration.
  • Study sequence: The twelve studies were conducted in two phases, neither jointly prospective: the first six were outcome-sequential, while follow-ups were conceived together and run one at a time.The narrative order was fixed after the Study B′ result.
  • Mechanism: The registered endpoint does not separate mechanisms such as copying versus dereferencing, principled decline versus non-engagement, or content from length.Study J, H2, and G manipulate authored edits whose total effects remain bundled.
  • Replication: Transport and replication contrasts do not isolate one change; G′ labels contrasts within ±10 points, with half of its thirty contrasts unresolved.H1 also differs in day, world text, and request fields for GPT-5.6 models.
  • Provenance: Study B required a post-freeze analysis correction, while Study F’s first run was quarantined; the invariance gate cannot detect count-preserving augmentation.These deviations are reported rather than corrected in place.

B The Study F contamination in detail

Study F’s contamination analysis documents provider-path anomalies, frozen-package deviations, and the study tables used to track exact arms, contrasts, replications, and downstream tests. The registry also records safeguards, model panels, and study-specific limitations relevant to interpreting these results.

  • Run integrity: Run 1 exposed injected search results and provider-accounting anomalies on the OpenRouter path.The package included a disabled web plugin, but stored rationales confirmed injected search results; billed costs also exceeded upstream inference costs on 1,054 response-bearing attempts.
  • Run integrity: The re-frozen package gated smoke checks, provider accounting, served endpoints, and restart behavior before confirmatory runs.Deviation detection poisoned a run on the first mismatch, while cost and endpoint bindings were treated as fatal or bounded conditions.
  • Study designs: Study B pooled eight pointer-by-label cells, while Studies D and E compared bare ids with stated criteria across six and nine models, respectively.The associated tables define the cell structures, per-model rates, and within-model contrast columns used for these comparisons.
  • Study F: Study F’s tables report cell means, sixteen registered contrasts, and per-model V73 counts across none, id, criterion, and combined-directive arms.The groups distinguish OpenRouter-served and direct-provider panels, including pre-named groups where the id or criterion had previously prevailed.
  • Follow-up studies: The paper’s study tables define registered interval procedures and labels, including δ = 10 materiality classifications for relevant contrasts.The tables cover Studies F-x, G, B′, I, H1, and J, with Wilson, Newcombe, or bootstrap intervals as specified by each study.

G Study E: the OpenRouter-served-panel transport test

Study E reran Study D’s design on a disjoint nine-model OpenRouter-served panel as a transport boundary check. Its criterion advantage did not satisfy the registered superiority rule, while the design could not isolate transport from simultaneous changes in models, access route, run count, and inference configuration.

  • Design: Study E was registered as Study D’s design on a disjoint OpenRouter-served panel, not as a replication.The new panel was reported as a transport test because model identities and serving conditions differed.
  • Execution: Of 540 attempted episodes, 526 scored; 14 transport failures occurred on one endpoint, Kimi K3.The frozen analyzer paired runs present in both arms, although that pairing rule was not stated in the prose registration.
  • Result: The rates were 30.4% for none, 76.1% for id, and 83.3% for criterion, yielding E1 = +7.2 [+0.0, +14.4].The registered superiority rule required the contrast’s lower interval limit to exceed zero, which it did not.
  • Interpretation: The result cannot establish equivalence because no equivalence margin was registered, and the interval remains compatible with a criterion advantage up to +14.4 points.The per-model contrast ranged from −25.0 to +90.0, and the executed panel changed several design factors simultaneously.
Loading 2609.03450v1…