Source-linked AI summary

Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout -- and the factorizations of it that failed

Nicolás Vera Zúñiga

arXiv:2608.21315v1cs.CL

TL;DR

Prompt–model interaction is known in task performance, but task readouts cannot show whether it is a property of task machinery or the conditional distribution. The paper therefore measures deterministic fixed-point structure in a short-window argmax map, finding large conditioning effects that resist five proposed factorizations. The supported unit of explanation is the prompt–model pair, within a readout that disappears at longer windows and whose nearest mechanistic account is outside its tested regime.

  • Problem

    Task-performance evidence does not settle whether prompt–model interaction belongs to task machinery or the conditional distribution itself.

  • Method

    The paper iterates a deterministic short-window argmax map from 96 starts and censuses its fixed-point fraction and structural classes.

  • Results

    Nine conditioning tokens move the fixed-point fraction across most of its range and change a four-way class, whereas +60.5 IFEval points move the class by zero; five proposed factorizations fail widening.

  • Takeaways & Limitations

    On this readout, the supported unit of explanation is the prompt–model pair rather than prefix-, model-, or nearest-mechanism-only factors.

  • Takeaways & Limitations

    The fixed-point structure is a short-window property: by W=16, four of six models have raw ϕ=0.000, and the probe uses uniformly random starts outside the mechanistic account’s regime.

Abstract

from arXiv · show

That a prompt's effect is not a property of the prompt is established: prompts optimised for one model degrade on another, and rankings reorder under neutral reformatting. That evidence is about task accuracy, which cannot say whether the interaction is a fact about task machinery or about the conditional distribution itself. We ask on a readout with no task in it: the fixed-point structure of the short-window argmax map x_{t+1} = argmax_x p(x | x_{t-1}, x_t), censused from 96 starts. It is deterministic, so nothing can be helped or hurt, and it exists only at short windows -- four of six models lose it entirely by window 16 -- so everything here concerns how a model reads a fragment. Two results. First, the interaction reaches this readout at full magnitude: nine tokens of conditioning move the fixed-point fraction across most of its range, change a four-way structural class, and reorder models, while instruction tuning worth 60.5 IFEval points moves the class by zero. Second, nothing we proposed carries it. Prefix length fails: the effect is not monotone. Four phenomenological factors -- prose-versus-markup, a universal direction, bidirectionality, instruct-resistance -- were each withdrawn within one run of being proposed, dissolved by widening the sample. And the nearest mechanistic account, attention-sink dominance of early tokens, predicts the sign of the shift on 2 of 5 models -- chance -- while a length-by-content cross shows it holds on real text and fails on our probe's uniformly random input, so we are outside its regime, not against it. One fixed nine-token prefix drives four models toward 0 and two toward 1; the bidirectionality survives in-distribution starts. On this readout the unit of explanation is the prompt-model pair. The recurring error it caught in us has a name: a criterion with a shape applied to a quantity with no room to vary.

1 Introduction

The paper asks whether prompt–model interaction reflects task machinery or the conditional distribution itself, using a task-free structural readout. It finds a large interaction that resists prefix-, model-, and mechanistic factorizations, leaving the prompt–model pair as the supported explanatory unit.

  • Motivation: Task-performance evidence cannot distinguish interaction in task machinery from interaction in the conditional distribution itself.
  • The discipline: The paper evaluates candidate factorizations by pre-registered widening and concludes that five natural factors did not survive, without claiming that no factorization exists.
  • The claim: Nine tokens move the fixed-point fraction across most of its range and change a four-way structural class, while +60.5 IFEval points move the class by zero.
  • Prefix-side factors: Prefix-side explanations fail because the effect is non-monotone in length and proposed content factors dissolve when the sample widens.
  • Model-side factors: Model-side explanations also fail: a universal direction, bidirectionality as a model property, and instruct-resistance dissolve, while one prefix drives four models toward 0 and two toward 1.
  • Mechanistic factor: Attention-sink strength agrees with fixed-point shifts on 2 of 5 models, and the account holds on real text but fails on uniformly random probe input.

2 Setup

The study measures fixed-point structure by iterating a deterministic argmax map from random starts and classifying where trajectories land. Its scope is intrinsically short-window, because widening the state makes the readout disappear on most models.

  • Estimator: The estimator iterates the conditional argmax map from 96 random two-token starts and computes fixed-point fraction ϕ plus four trajectory classes.The classes are funnel, none, fragmented, and borderline.
  • Domains: Prefixes are domains prepended before every forward pass, while the estimator remains unchanged; tested kinds are raw, bos, text, template, and struct.
  • Scope: At W=16, raw ϕ is 0.000 on four of six models and ≤0.10 on a fifth, leaving measurable fixed-point structure on only one model.At W=2, raw ϕ ranges from 0.22–0.70.
  • Comparative framing: Table 2 compares conditioning effects on this structural readout with accuracy-related magnitudes, while its class rows have no accuracy analogue.
  • Scope: The readout concerns how a model reads a fragment, not task performance, because the map is deterministic and has no answer, format, verbalizer, or sampling noise.

3 The locus: the interaction reaches a task-free structural readout

Conditioning reaches the task-free structural readout at substantial magnitude, while instruction tuning leaves the structural class unchanged. The contrast isolates a selective conditional-distribution effect rather than generic sensitivity to intervention.

  • +60.5 IFEval points from instruction tuning move the structural class by zero, while nine conditioning tokens move it completely.
  • The readout is not sensitive to every intervention: instruction tuning leaves it unchanged even as a short prefix substantially changes its fixed-point structure.

4 Prefix-side factors fail

Prefix length and content do not provide stable explanations for the interaction. Length effects are non-monotone, and proposed prose-versus-markup and no-text narratives dissolve when model and corpus coverage widen.

  • Length: Three of four models with enough span are non-monotone across raw, bos, text, and template conditions.One sequence is 0.948 →0.005 →1.000 →0.000, with nine prose tokens restoring a perfect funnel after BOS nearly annihilates structure.
  • Length: Length-matched text versus template isolates prefix kind, while the structural readout retains the largely non-monotonic shape known from accuracy-level prompt effects.
  • Content: Two prefix-content factors die on widening: prose-versus-markup sign flips do not survive mid-range models, and the no-text effect fails on a wider corpus.
  • Content: Table 3 holds one fixed prefix to nine tokens across six models and tests whether bidirectionality survives in-distribution starts.

5 Model-side factors fail, and what remains is pairwise

Model-side explanations did not survive widened sampling: apparent universal direction, instruct-resistance, and related content effects dissolved, leaving the prompt-model pair as the supported residual.

  • Seven base models produced two immediate bidirectional counterexamples to the proposed universal direction.
  • One C-source header raised gemma-2-2b-it from 0.714 to 0.917, disconfirming the hoped-for instruct-resistance null.
  • Instruct models rose on 1 of 24 text-model units versus 11 of 48 for base models, but the authors decline to promote this gap as a factor.
  • With prefix-side and model-side factors unsupported, the residual explanation is the prompt-model pair.
  • Table 3 reports a structural analogue of model-specific prompt effects: one model rises to 0.979 while SmolLM-1.7B remains at 0.000.

6 The mechanistic factor: where the sink account predicts magnitude, we observe sign

The attention-sink account explains behavior in its own regime but does not explain this probe: sink strength predicts the fixed-point shift’s sign on only 2 of 5 models.

  • The sink account is positional and content-independent, attributing early-token attention concentration to SoftMax normalization rather than token meaning.
  • A single BOS token moves Falcon3-1B-Base from 0.214 to 0.906 while driving other models toward zero, so the shift’s sign is not shared.
  • Sink strength and fixed-point fraction move together on only 2 of 5 models under length-matched BOS comparison, which is chance.
  • Across lengths 2–512, sink concentration rises with length on all six models for both uniformly random tokens and real text.
  • At lengths ≥8, no resolved model shows sink decrease under BOS on real text, whereas some do so at every length for random tokens.
  • The probe therefore lies outside the sink account’s described regime, making its failure to predict the sign unexplained rather than contradictory.
  • Because the proposed mechanism is architecture-general and content-free, reversing signs across models creates a real tension beyond a definitional mismatch.

7 The discipline

The paper’s discipline is to test only quantities with room to vary, separating distinct questions and checking whether pooled or apparent effects survive widening and unit-specific noise.

  • Validity checks: A criterion with a shape is invalid when the measured quantity has no room to vary.Examples include flat sequences, floored arms, and ceiling baselines that could not meaningfully disagree.
  • Validity checks: A universal inferred from n = 2 may be counterexample-possible yet still too weakly tested to put the claim at risk.The paper treats widening before claiming as a cheap safeguard against both vacuity and implausible certainty.
  • Validity checks: The protocol screens for headroom, gates each question separately, and compares each shift with its own noise.It also verifies anti-vacuity and records tests that cannot fail informatively at the available sample size.
  • Pooling: Pooling high-markup and prose results produced 7/18 versus 3/18, but a reversed model made the pooled predictive rule uninterpretable.A consistency gate rejected the pooled number because units pointed in different directions.

8 Limits

The limits are substantive: the evidence is narrow, the estimator disappears at long windows, several confounds remain, and proposed explanations survive only within carefully bounded regimes.

  • Scope: The largest single comparison uses six models and twelve texts from three English sources at one prefix length, so general claims are not supported.The authors therefore scope boundaries to these models and texts rather than the general form.
  • Estimator boundary: The window sweep bounds the estimator rather than the cohort: Qwen2.5-1.5B-Instruct retains |∆ϕ| = 0.547 at W=4 versus 0.521 at W=2 before collapsing by W=8.This single-model observation is recorded as a knife-edge check, not promoted as a rate.
  • Confounds: gemma-2-2b-it confounds fragmentation with cohort membership, while Falcon3-3B-Instruct is excluded because seed variation of 0.615–0.792 swamps its domain shifts.These constraints prevent clean interpretation of class or direction effects for those models.
  • Input distribution: Uniform random token-pair starts raise an OOD concern, but the domain effect survives on 6 of 6 models with real-text adjacent starts.The sign and rough magnitude persist, including one model rising to 0.964 while the rest collapse.
  • Interpretive cautions: Four models reach ϕ ≈0 under BOS in both regimes, so their ∆ϕ is floor-bounded and cannot be compared across regimes.A baseline shift of −0.271 on one model also makes the comparison cross-regime rather than a within-regime robustness check.
  • Mechanistic scope: The sink account cannot be connected to this readout where it applies because ϕ is unmeasurable at long context.The paper bounds the account’s regime without claiming that the account is contradicted.
  • Additional factorization: A sixth endpoint-token factorization failed in strong form: ϕ = 0.995 occurred although the predicted token’s margin never became positive.Across nine models per arm, modal endpoint-token agreement reached only 4 of 9, leaving another model–prefix interaction.

9 Conclusion

The conclusion is that prompt–model interaction reaches a deterministic, task-free structural readout, while every tested factorization dissolves under widening and the prompt–model pair remains explanatory.

  • Conclusion: The fixed-point structure reaches the task-free readout, but prefix, content, model-side, and mechanistic factorizations all fail after widening.The readout is deterministic and therefore does not encode helped-or-hurt task performance.
  • Conclusion: Five proposed factors did not survive the pre-registered widening criterion, so the unit of explanation remains the prompt–model pair.A factor survives only when widening that it did not choose fails to dissolve the effect.
  • Conclusion: The forward-looking discipline is to refuse criteria with a shape whenever the quantity has no room to vary.This rule is presented as the main methodological practice to carry forward.
Loading 2608.21315v1…