Source-linked AI summary

Relation Before Entity: Deferred Commitment in Language Model Factual Recall

Divyansh Agarwal

arXiv:2609.17537v1cs.CL

TL;DR

The paper asks whether relation-type and entity-specific information become causally active at the final token at the same depth during recall. It applies four causal diagnostics across four decoder-only models and eight prompt families, finding that relation information controls generation before entity information does. Entity information is available early at the entity-token position but is routed to the final token and committed later.

  • Problem

    Prior work did not quantify a layer-wise causal onset gap across architectures or distinguish early entity availability from later final-token commitment.

  • Method

    The study combines transfer-curve patching, both-change competition, entity-token patching, and steering across four decoder-only models and eight controlled prompt families.

  • Results

    Relation onset precedes entity onset by 10–16 tested layers at threshold 0.4, with the ordering holding across all 16 model-threshold combinations for thresholds 0.2–0.5.

  • Takeaways & Limitations

    Recall is temporally factorized: relation information becomes final-token generation-controlling before entity information, whose commitment is deferred despite early availability.

  • Takeaways & Limitations

    The experiments use controlled fill-in-the-blank prompts and greedy first-answer generation, while natural-language QA, longer contexts, and free-form reasoning remain untested.

Abstract

from arXiv · show

We ask whether relation-type information (e.g., capital-of) and entity-specific information (e.g., France to Paris) become causally active at the final-token position at the same depth during recall. Using four complementary causal diagnostics across four decoder-only models and eight prompt families, we find a robust temporal asymmetry: relation information becomes generation-controlling before entity information does. Relation onset precedes entity onset by 10-16 tested layers (31-44% of network depth) at threshold 0.4, with the ordering holding across all 16 model-threshold combinations for thresholds 0.2-0.5. Critically, entity information is not absent early: entity-token patching succeeds at 90-100% in early layers. Instead, entity commitment to generation is deferred: entity information is available at the entity-token position but becomes generation-controlling at the final token only after being routed there.

1. Introduction

The paper asks when relation-type and entity-specific information become generation-controlling at the final token, addressing limits in prior layer-wise factual-recall evidence. Across four decoder-only models and eight prompt families, it uses complementary causal diagnostics to test whether these signals become active at the same depth.

  • Motivation: The study measures when relation and entity signals become generation-controlling at the final token using direct causal diagnostics.It addresses prior work that located factual associations or enriched subject representations without quantifying an onset gap across architectures or separating entity availability from final-token commitment.
  • Terminology: The paper distinguishes relation type from the entity supplied to that relation, with the entity-token position denoting where that input item appears.Examples include capital-of and France, past-tense and a verb, or symbolic transformations and an element.
  • Scope: The experiments cover four decoder-only models and eight fill-in-the-blank families spanning factual, morphological, lexical, and symbolic transformations.The models range from 28 to 36 tested layers, and prompt items are programmatically defined with greedy-generation sanity checks.
  • Core claim: The central hypothesis is temporal factorization: relation information becomes final-token active in middle layers, whereas entity commitment is deferred to late layers.This ordering is tested across four architectures, eight prompt families, and thresholds 0.2–0.5.
  • Implications: The proposed distinction matters because layer-targeted interventions and monitoring may affect or measure relation-level and answer-identity computation differently.The paper also identifies possible relevance to chain-of-thought reasoning and multi-hop settings, without testing those settings here.

3. Experiment 1: Transfer Curves and Onset

Transfer-curve patching tests whether relation or entity information controls the final-token state at each layer. Relation transfer rises earlier than entity transfer, with a stable onset gap and evidence that mid-layer transfer reflects relational structure rather than donor-answer copying.

  • Design: Transfer curves compare prompts differing in relation but sharing an entity against prompts sharing a relation but differing in entity.Relation transfer tests whether the final-token state carries relation type, whereas entity transfer tests control by specific answer information.
  • Onset definition: Onset is the first tested layer exceeding 0.4 for two consecutive tested layers, measuring stable causal commitment rather than a single-layer spike.Threshold sensitivity from 0.2–0.5 is reported separately.
  • Results: Relation transfer rises before entity transfer in every model, while entity transfer remains near 0% for multiple layers as relation transfer exceeds 0.4–0.8.Figure 1 uses shading for SEM and dotted lines for onset layers.
  • Controls: 0.79–0.86 relation-only transfer coexists with donor-answer copying of ≤0.11 at each model’s peak relation-only layer.Donor-answer copying rises only later, when entity commitment takes over, indicating that mid-layer patches carry relational structure rather than merely copying donor answers.

4. Experiment 2: Both-Change Competition

Both-change competition places relation and entity information in direct causal conflict across depth. Relation wins in middle layers, entity wins late, and the crossover tracks the independently measured entity onset.

  • Results: Relation wins dominate middle layers, whereas entity wins dominate late layers across all four models.The direct-conflict design uses different relations and entities, classifying outputs as original, relation-win, entity-win, or mixed.
  • Results: The crossover aligns with Experiment 1 entity onset within ≤2 tested layers in every model.This links the transition from relation to entity dominance with the layer where entity transfer becomes active.
  • Controls: Noise-patch controls produce near-zero structured wins at ≤0.008, while self-patches retain the original output at ≥99.2%.Alternating donors preserve the qualitative pattern, and unrelated relation-wins remain near zero at ≤0.030.

5. Experiment 3: Entity-Token vs. Final-Token Patching

Experiment 3 separates entity information from its causal control at the generation position, showing early entity availability but late final-token commitment.

  • 5. Experiment 3: Entity-Token vs. Final-Token Patching: The experiment compares donor-state patching at the final-token position with patching at the entity-token position for same-relation, different-entity pairs.The design tests whether early entity-token patching succeeds even when final-token patching fails.
  • 5. Experiment 3: Entity-Token vs. Final-Token Patching: 90–100% entity-token patching in early and middle layers contrasts with ≈2.4% final-token transfer, while late layers reverse this pattern.The crossover matches entity onset from Experiments 1–2 in every model, supporting delayed routing rather than delayed knowledge.
  • 5. Experiment 3: Entity-Token vs. Final-Token Patching: Entity information is available early at the entity-token position but becomes generation-controlling at the final token only after being routed there.This rules out the explanation that entity information is simply absent until late layers.

6. Experiment 4: Steering Temporal Asymmetry

Experiment 4 uses relation- and entity-specific steering directions as an independent check on the temporal asymmetry in causal control.

  • 6. Experiment 4: Steering Temporal Asymmetry: Relation steering is effective in middle layers, whereas entity steering is substantially weaker mid-layer and strongest in late layers.The directions are constructed from same-entity/different-relation and same-relation/different-entity prompt pairs, respectively.
  • 6. Experiment 4: Steering Temporal Asymmetry: Random-direction baselines are ≤ 0.015, supporting the direction-specificity of the steering effects.
  • 6. Experiment 4: Steering Temporal Asymmetry: The steering analysis complements causal-tracing, activation-patching, transformer-circuits, and prior work on subject-position factual recall and relation representations.Related work links subject-token computations to factual recall and distinguishes decodable information from information that causally mediates behavior.

8. Discussion

The discussion frames recall as a staged process: entity information is initially available, relation information reaches the final token, entity commitment follows, and late representations become overwrite-sensitive.

  • 8. Discussion: Entity-token availability precedes relation activation at the final token, entity commitment there, and late overwrite sensitivity.This four-stage sequence summarizes the converging results across the paper.
  • 8. Discussion: 90–100% early entity-token patching coexists with ≈2.4% final-token transfer; middle layers activate relations, while late layers route and commit entities.The latest layers also show broad overwrite sensitivity, including unrelated donors inducing late entity-like overwrite while unrelated relation-wins remain near zero.
  • 8. Discussion: Monitoring and intervention methods should distinguish information that is represented from information that causally controls generation.The transition zone offers a concrete target for future editing and steering analyses, although weight editing is not directly tested.
  • 8. Discussion: The evidence is limited to controlled fill-in-the-blank prompts, greedy first-answer generation, and openweight decoder-only models in the 3B–8B range.Natural-language QA, longer contexts, free-form reasoning, and the routing mechanisms themselves remain untested or unlocalized; the ordering, not absolute layer indices, is central.

9. Conclusion

The conclusion reports a robust temporal factorization of recall in controlled prompt families, with relation information controlling the final token before entity information.

  • 9. Conclusion: Relation onset precedes entity onset by 10–16 tested layers at threshold 0.4, or 31–44% of network depth.The ordering holds across all 16 model-threshold combinations for thresholds 0.2–0.5.
  • 9. Conclusion: Entity information is available at the entity-token position from the earliest layers but reaches generation control at the final token only after routing.Four complementary causal diagnostics converge on the same transition layers.
  • 9. Conclusion: The work aims to improve mechanistic understanding of recall and may support safety monitoring, steering, and model editing.

A. Both-Change Competition Figure

Both-change competition shows relation signals dominating middle layers and entity signals dominating late layers, with crossover layers aligned to entity onset. Steering results likewise show relation control in middle layers and entity control late, above near-zero random baselines.

  • A. Both-Change Competition Figure: Relation wins dominate middle layers, whereas entity wins dominate late layers across all four models.The crossover layers align with entity onset from Experiment 1.
  • A. Both-Change Competition Figure: Relation directions steer effectively in middle layers, while entity directions are weaker there and strongest late.Random-direction baselines are at or below 0.015.
  • A. Both-Change Competition Figure: Relation-before-entity ordering holds across all 16 model-threshold combinations for thresholds 0.2–0.5.The table reports pair-balanced onset across the tested thresholds.

D. Additional Controls

Additional controls distinguish genuine mid-layer relation dominance from donor-answer copying, generic overwrite sensitivity, and patching artifacts. They converge on a relation-dominant middle-layer regime followed by late entity commitment.

  • D. Additional Controls: Peak relation-only transfer reaches 0.79–0.86 while donor-answer copying remains ≤0.11 at relation-dominant layers.Donor-answer copying rises substantially only later, when entity commitment takes over at the final token.
  • D. Additional Controls: The peak wrong-entity layer matches the peak relation-wins layer in every model, providing convergent evidence for the same middle-layer regime.This alignment rules out donor-answer copying as the explanation for mid-layer relation transfer.
  • D. Additional Controls: The controls include near-zero structured wins from noise patches and preservation of the original output under self-patches.Table 7 defines these as diagnostic checks for structured effects and patch fidelity.
  • D. Additional Controls: Unrelated donors produce high entity-like overwrite only late, while unrelated-donor relation wins remain ≤0.030 throughout.Thus generic patching does not explain mid-layer relation dominance.

E. Prompt Family Details

The study uses controlled prompt banks spanning multiple transformation types, with fixed item banks and audited greedy-generation statistics to document prompt quality.

  • E. Prompt Family Details: Table 8 lists prompt families, templates, and record counts for the controlled prompt banks.The prompt families cover the study’s controlled evaluation materials.
  • E. Prompt Family Details: Prompt items are programmatically defined from fixed item banks, and greedy generation is audited using rank and logit-margin statistics.These audits identify ambiguous or unstable items and document prompt quality.
Loading 2609.17537v1…