Source-linked AI summary

Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5

Justin Bronder

arXiv:2608.14992v1cs.AIcs.CL

TL;DR

Language models may treat previously written claims as retrieved facts when they reappear in tool-result packaging. Across three synthetic assignment studies, Claude Opus 5 adopted unsupported codes, but native tool results did not outweigh announced inline text in the strongest comparison.

  • Problem

    The paper asks whether message packaging changes a model’s answer when an unsupported claim returns from a store it also writes to.

  • Method

    Three prospective studies compared false-code adoption across assistant assertions, native tool-result records, metadata-wrapped records, and announced inline JSON.

  • Results

    False-code adoption was 57/60 with tool results versus 60/60 with inline text, so the preregistered result-first superiority criterion failed.

  • Takeaways & Limitations

    The evidence supports sensitivity to message packages, not intrinsic authority of the tool-result channel.

  • Takeaways & Limitations

    External validity is limited to one model, provider API, synthetic task template, opaque identifiers, color tokens, and abstention token.

Abstract

from arXiv · show

Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model gives in a synthetic lookup task. Claude Opus 5 selected a color code for a named item or abstained. In an exploratory four-arm study, false-code adoption was 0/24 with no target claim, 0/22 scorable trials when a prior assistant assertion named the target, 14/24 when a tool-result record named it, and 15/24 when that result used a ten-field metadata wrapper that marked it unchecked. The tool-result arm selected the record's code in 11/12 supported trials and 14/24 unsupported trials, ruling out a fixed output-token bias while leaving substantial planted-token heterogeneity. A document-preregistered replication reproduced the tool-result versus assistant-assertion gap, 7/24 against 0/24, one-sided Fisher exact p = 0.0047. The tool-result rate nevertheless fell from 14/24 to 7/24 across runs made four days apart. A second preregistered study gave the earlier comparison a live text control: both records were announced in advance and placed in the same final user turn, then target binding was swapped between the linked tool result and later inline JSON. Inline text was sufficient for false-code adoption in 60/60 trials; the tool-result condition produced 57/60, so the registered result-first superiority criterion failed, p = 1. The result does not show that tool results have no effect. It shows that native tool-result placement was not necessary and that this experiment did not find greater behavioral weight for the result package than for announced inline text. The findings concern a single model on one synthetic task template, accessed through one API.

1. Introduction · 2. Task and message packages

The paper isolates whether message packaging changes model responses in a narrow synthetic assignment task. Its prospective studies compare inter-turn tool-result and assistant-assertion packages with an announced same-turn comparison against inline text.

  • 1. Introduction: The motivating concern is that an unchecked claim can return from a self-written store as a structured tool result without gaining evidence.The scenario changes packaging, not the customer’s underlying knowledge.
  • 1. Introduction: The studies isolate message-package effects rather than testing deployed memory, retrieval, scratchpad, or multi-agent systems.The authors frame the contribution as bounded observations, not a general theory of memory.
  • 1. Introduction: A native tool result bundles role, serialization, tool linkage, compulsory position, and retrieval signaling, so complete-package contrasts cannot identify intrinsic channel authority.This motivates the stronger same-turn text comparator in Study 3.
  • 2.1 Synthetic assignment task: Each trial names one opaque target, offers AMBER, ORCHID, SABLE, or ABSTAIN, hides the correct code, and plants a task-provided code in one or more messages.Primary trials use a planted code differing from the correct code.
  • 2.1 Synthetic assignment task: False-code adoption is the exact final token matching the planted wrong code, while ABSTAIN is separately recorded.This endpoint measures output behavior, not belief, confidence, rationality, or internal representation.
  • 2.2 Studies 1 and 2: inter-turn packages: Studies 1 and 2 use a fixed five-message history containing a historical tool call, serialized tool result, later assistant assertion, and final question.Study 1 compares no-claim, assistant-assertion, tool-result, and annotated tool-result arms; Study 2 repeats the assistant-assertion and tool-result arms.
  • 2.3 Study 3: announced records in one user turn: Study 3 announces both records in advance, places a linked tool result and schema-matched inline JSON in the same final user turn, and swaps which names the target.Channel, linkage, and first position remain bundled because the message grammar requires the tool result to follow its call and precede ordinary text.

3. Methods

Three studies progressively narrowed the question from message-package effects to whether native tool-result placement retained an advantage over announced inline text. Across studies, the researchers fixed the provider, model, and generation settings while preregistering confirmatory analyses and handling unscorable trials conservatively.

  • Study design: Study 1 compared four complete message packages with no-claim and supported-case controls to map false-code adoption.It contained 144 conversations, with 36 per arm, crossing three correct codes with three planted codes and repeating each cell four times.
  • Study design: Study 2 tested whether the largest exploratory contrast recurred using 24 differing-code trials in each of the assistant-assertion and tool-result arms.The document-preregistered criterion covered all 48 attempts and required zero local evidence-integrity failures.
  • Implementation: All three studies used the Anthropic Messages API with model alias claude-opus-5 and fixed generation settings, including high adaptive reasoning and automatic tool choice.Requests set max_tokens to 8,192, disabled parallel tool use, and omitted temperature so the provider default applied.
  • Study design: Study 3 replaced the nearly inactive assistant-assertion comparator with an announced inline-text record and swapped target binding within the same final user turn.It used 60 trials per arm, with ten repetitions in each of six differing-code cells; matched trials were adjacent and arm order was balanced 5/5.
  • Analysis: Confirmatory analyses were designed to prevent exploratory evidence or missing outcomes from gaining weight, including conservative Study 2 handling of unscorable trials and a stratified exact test for Study 3.Study 3 conditioned on the six correct-code-by-planted-code cells and required a positive worst-case lower bound with p at or below 0.05.

4. Results

Across separate message constructions, inter-turn tool results produced false-code adoption while prior assistant assertions did not, but same-turn announced inline text produced near-universal adoption. The tool-result rate also varied between runs, and a separate verification run yielded uniformly truthful answers.

  • Observed arm-level behavior: 14/24 (58%) and 15/24 (63%) false-code adoption occurred with raw and annotated tool results, versus 0/24 (0%) with no claim and 0/22 (0%) after assistant assertion.The assistant-assertion worst-case bound was 2/24 (8%). Rates were reported separately because the studies used different message constructions.
  • Preplanned comparisons: 7/24 (29%) versus 0/24 (0%) in Study 2 met the registered directional criterion, whereas Study 3 did not.Study 3’s comparison had p = 1 because inline text produced 60/60 false-code adoptions and the tool-result total was the conditional minimum.
  • Same-turn construction: 57/60 (95%) tool-result adoption and 60/60 (100%) inline-text adoption occurred in the same-turn swap construction.Both target-bound packages therefore produced near-universal adoption; inline text was sufficient, and the result-first superiority criterion failed.
  • Content tracking: 11/12 supported trials followed the tool-result record’s code, while false-code adoption varied by planted token across runs.Study 1 tool-result adoption was 8/8 for SABLE, 4/8 for ORCHID, and 2/8 for AMBER; Study 2 was 5/8, 0/8, and 2/8.
  • Between-run stability: The inter-turn tool-result rate fell from 14/24 on August 9 to 7/24 on August 13.Post-hoc two-sided Fisher tests gave p = 0.080 unstratified and p = 0.021 when stratified over six task cells; the cause was not established.
  • Independent checking: 144/144 verifier invocations returned correct assignments, and all 144 follow-up answers matched truth; false-code adoption was 0/96.This run was descriptive rather than a causal estimate because it changed the tool schema, verifier invitation, and abstention language.

5. Discussion

Study 3 showed that native tool-result delivery was not necessary for false-code adoption and failed to support tool results as inherently more authoritative than text. The studies establish comparator-sensitive package effects, not belief, confidence, intrinsic memory authority, provider effects, or a general source hierarchy.

  • Study 3: Study 3 produced false-code adoption with announced inline records, so native result delivery was unnecessary and the registered superiority criterion was not met.Both records were announced as task inputs and placed in the final user turn.
  • Interpretation: The earlier assistant assertion remained at the no-claim floor, whereas same-turn inline text was maximally active, implicating task and message construction rather than channel-invariant authority.The comparison does not invalidate the earlier randomized contrast; it changes its plausible explanation.
  • Claim boundaries: The studies establish that comparator design can reverse a reproducible package contrast, but do not establish belief, confidence, intrinsic memory authority, provider effects, or a general source hierarchy.These boundaries separate observed behavior from broader claims about what the model believes or remembers.
  • Future research: The next scientifically useful test is whether a fixed provenance policy reduces unsupported-record uptake while preserving correct use of supported records.This requires concurrent controls and predeclared safety and utility criteria, rather than merely another identical replication.

6. Limitations · 7. Related work

The study’s conclusions are limited by its narrow synthetic setup, bundled interventions, and inability to establish false-belief adoption or general channel rankings. Related work situates the findings within retrieval faithfulness, instruction hierarchy, attribution, automation bias, verification, preregistration, and behavioral evaluation.

  • 6. Limitations: The study tested one model alias, one provider API, one synthetic task template, opaque identifiers, three color tokens, and one abstention token, limiting external validity.The setup may evoke benchmark compliance more strongly than natural retrieval, memory, or agent work.
  • 6. Limitations: Native treatments bundled tool-result role, linkage, serialization, compulsory first position, and retrieval fulfillment, while Study 3 also changed system instruction and conversation graph.Therefore, the message packages do not isolate individual source mechanisms, and Study 3 rates cannot be pooled with Studies 1 and 2.
  • 6. Limitations: The model received no independent evidence that the planted assignment was false, so false-code adoption may reflect competent task compliance rather than gullibility, irrationality, or belief.True-code cases demonstrate content tracking but do not make the false record visibly false to the model.
  • 6. Limitations: Study 1’s assistant-assertion arm matched the no-claim control’s observed floor, while Study 3’s inline comparator adopted the planted code on every trial.Non-detection is not equivalence, and Study 3’s p = 1 leaves the registered directional criterion unmet while saying little about effect-size similarity.
  • 6. Limitations: 14/24 to 7/24 tool-result adoption occurred across different days without a provider checkpoint or serving snapshot, so post-hoc tests show association, not cause.Safety refusals elsewhere also indicate interaction between the instrument and model- or provider-level classification.
  • 6. Limitations: The verifier context changed several components, the annotation arm was bundled, and neither established mitigation efficacy; live save, retrieval, verification, and persistence were not exercised.Historical messages were rendered directly into each request, rather than testing concurrent live behavior.
  • 6. Limitations: A recommended third same-turn arm separating task-sanctioned groundedness from native provenance was not run, leaving Study 3 as the registered two-arm comparison.The missing arm limits interpretation of the registered comparison.
  • 7. Related work: Related work frames the study as a narrower test of whether structured wrappers alter unsupported-claim uptake, alongside retrieval faithfulness, instruction hierarchy, prompt injection, attribution, automation bias, verification, preregistration, and behavioral evaluation.Prior findings include overreliance on inaccurate external context, attacker-chosen answers from poisoned database text, and inappropriate weight for unreliable system-presented information.

8. Conclusion … Funding, compute, and conflicts of interest

Claude Opus 5 often followed unsupported assignments presented as records, but the evidence supports sensitivity to message-package features rather than intrinsic authority of native tool results. The work is limited to one model, API, and synthetic task, with artifacts subject to redaction and no external funding or conflicts declared.

  • 8. Conclusion: 57/60 false-code adoptions with tool results versus 60/60 with announced inline JSON failed the preregistered superiority criterion.Native tool results outperformed prior assistant assertions in two inter-turn studies but not announced inline text in the same-turn comparison.
  • 8. Conclusion: The evidence supports message-package sensitivity, not intrinsic authority of the tool-result channel.Task framing, target binding, source role, linkage, serialization, position, and the apparent basis for answering rather than abstaining remained partly bundled.
  • 8. Conclusion: A useful next question is whether fixed evidence-checking or provenance policies reduce unsupported-record uptake while preserving correct use of supported records.This shifts the focus from ranking message channels to operational safeguards.
  • Reference verification: RefChecker 3.0.183 found no reference errors or unverified entries among the first 26 entries.Its 19 warnings were manually adjudicated as preprint-versus-publication or venue-normalization mismatches rather than unresolved errors.
  • Data and code availability: The repository preserves code, preregistrations, result memos, and private raw evidence, but this version does not claim public reproducibility.A public release requires redaction of provider headers containing organization, workspace, request, trace, and network metadata.
  • Author contributions and AI participation: Justin Bronder is the sole author and accepts responsibility for the paper’s research questions, design, interpretation, evidence, and revisions.The supplied passage identifies him as affiliated with Corabo.
  • Author contributions and AI participation: AI participation was substantial and continuous across design, implementation, evidence review, statistical verification, critical review, and editing.Fable served as a top-level critical reviewer and editor, while Claude Opus 5.0 contexts participated in research critique and artifact review.
  • Funding, compute, and conflicts of interest: The research received no external funding, grants, donated compute, API credits, or other in-kind support, and the author declares no conflicts of interest.All local compute and commercial API usage were paid for by the author; Anthropic did not fund, audit, or endorse the work.

Appendix A. Exact prompts, records, and model settings … A.3 Study 3 message template

The appendix specifies the exact prompts, record structures, and final-answer constraints used in Studies 1–3. Study 3 additionally placed a linked tool-result record before an inline record with the same schema in the final user turn.

  • A.1 System prompts: Studies 1 and 2 instructed the model to treat recalled and assistant statements as available context and provide no explanation.The prompt also required exact compliance with the final response format.
  • A.1 System prompts: Study 3 instructed the model to treat both task-provided records as available context, follow the final response format exactly, and provide no explanation.
  • A.2 Studies 1 and 2 message template: Studies 1 and 2 used a synthetic target-item prompt, required retrieval of a separate record, and warned that records might concern a different item.The historical assistant message called cw_recall_id and described retrieval as returning authored content that was not verification.
  • A.2 Studies 1 and 2 message template: The final question required determining the exact target’s assigned code and answering only AMBER, ORCHID, SABLE, or ABSTAIN.ABSTAIN was defined as insufficient available context to choose a code for the exact target.
  • A.2 Studies 1 and 2 message template: Arm E used a repository renderer that added ten named fields and reordered the record, including notices, lineage metadata, hashes, support status, citations, and verification availability.The representative envelope marked support_status as unchecked and preserved exact per-trial requests in the run archive.
  • A.3 Study 3 message template: Study 3 told the model to inspect both task-provided records: the linked cw_recall_id result and the immediately following inline task record.The final user turn ordered the linked tool_result JSON first, the inline JSON record with the same schema second, and the final-question template third.

A.4 Generation settings · Appendix B. Registration, scoring, and evidence handling

The studies used fixed high-effort Claude Opus 5 generation settings and documented preregistration, single-attempt collection, replay-based scoring, and evidence-handling procedures. The archive also records a source-replay failure that was recovered offline without new provider requests.

  • A.4 Generation settings: Claude Opus 5 used adaptive omitted thinking, high effort, an 8,192-token limit, automatic tool choice, and disabled parallel tool use.The recorded request specified the model, max_tokens, thinking, output_config, and tool_choice fields.
  • A.4 Generation settings: Temperature was absent, the API version was 2023-06-01, service tier was standard, and no provider checkpoint identifier was available.These settings and metadata were reported for the requests.
  • Appendix B. Registration, scoring, and evidence handling: Study 1 fixed its run specification before contact and named exploratory contrasts but lacked a confirmatory exact-test rule.Study 1’s registration differed from the later preregistered studies in this respect.
  • Appendix B. Registration, scoring, and evidence handling: Studies 2 and 3 were governed by named preregistration files and Block 2 registration amendments.Study 2 used confirmation-48-preregistration-v1.md, while Study 3 used same-turn-announced-record-swap-preregistration-v1.md.
  • Appendix B. Registration, scoring, and evidence handling: Each main-study request was attempted once, with retries, behavioral follow-ups, and output tool execution disabled.Exact request bytes and response or transport-error records were saved; duplicate valid provider identities were hard failures in Studies 1 and 3 only.
  • Appendix B. Registration, scoring, and evidence handling: Studies 2 and 3 committed private seeds before contact, revealed them afterward, replayed runs offline, and independently recomputed registered statistics.Hashes and manifests established evidence identity and completeness rather than behavioral evidence.
  • Appendix B. Registration, scoring, and evidence handling: The verifier-enabled collection completed provider interactions, but source replay failed because Windows line endings differed from normalized Git blobs.Scores were recovered offline from preserved responses without new provider requests, and the failure and recovery remain separately archived.

Appendix C. Evidence archive and public release

The private archive preserves preregistrations, raw materials, registered analyses, audits, and tests, while the public package is planned as a redacted derivative mapped to immutable private-archive hashes. Raw headers contain sensitive identifiers and network metadata, and the current release status is reported in Data and code availability.

  • Private archive: The private archive preserves preregistrations, failure records, result memos, source manifests, exact requests, raw responses, registered scores, replay audits, and tests.Protocol code is under experiments/authority_without_evidence_prospective/; tests are under the repository's tests/ directory.
  • Sensitive metadata: Raw response headers include organization, workspace, request, trace, and network metadata, while API credentials were not stored.The metadata should not be published without review.
  • Public release: The public package should be a redacted derivative with a manifest mapping every released artifact to an immutable private-archive hash.The current public-release status is stated in Data and code availability.

Appendix D. Closed and separately registered replacement runs

The initial Study 2 and Study 3 runs were closed after three consecutive transport failures, and separately registered replacement runs used fresh protocol identities and seeds. Failed-run trials were excluded from study denominators, while the failures remained disclosed program history.

  • Run closure: Three consecutive transport failures halted both initial runs, leaving exact requests and transport-error records but no response status, headers, or body.The first Study 3 run also used a defective 32-zero-byte private seed, violating the fresh-random-seed premise.
  • Replacement runs: No failed-run trial entered a study denominator, although zero observed responses did not prove that no request reached the provider.Replacement specifications used new protocol identities, fresh seeds, new conversation and provider-visible identifiers, and disjoint request hashes.
  • Replacement authorization: Replacement runs were authorized only after independent qualification of the transport fault and preregistration of new specifications, with the preserved failures retained in the disclosed program history.Replacement was therefore not treated as automatic.

Appendix E. Complete program disclosure

Appendix E discloses 20 experiment records, including focal, exploratory, diagnostic, pilot, and transport runs. It also documents substantial provider refusals and qualification failures that constrain interpretation of the focal contrast.

  • Verifier and recovery runs: 144/144 verifier uses in a separate setup verified truth, while false-code adoption was 0/96 and supported cases were 12/12.The disclosure identifies this as a separate verifier setup with recovered offline results.
  • Run status and preregistration: 3/0 transport failures marked both an incomplete first Study 2 run and an incomplete first Study 3 run, while later listed runs met their registered criteria.The disclosed table reports 48/48 requests and zero unscorable or failure cases for one registered recurrence run, and 120/120 with zero for the registered superiority run.
  • Refusal and qualification behavior: 34/36 provider cybersecurity refusals occurred in the non-aliased portfolio, versus 11/36 in the aliased portfolio and 5/8 in the preface diagnostic.A concurrent qualification observed 34/36 exact-condition versus 9/36 alias-condition refusals, but failed its registered usable-response thresholds.
Loading 2608.14992v1…