Source-linked AI summary

The Internal Anatomy of Strategic Choice in Large Language Models

Vinícius Ferraz, Leon Houf, Enrico Ferrea

arXiv:2609.07478v1cs.AIcs.GT

TL;DR

The paper asks how models turn represented strategic information into choices, rather than merely whether they choose systematically. It traces incentives from prompts through activations to decisions and finds that similar behaviour can arise from different internal routes, including after instruction tuning.

  • Problem

    How models turn represented strategic information into choices remains unclear, despite systematic strategic behaviour and detectable information in activations.

  • Method

    The study records activations from four open-weight models playing 144 strict ordinal 2×2 games and traces prespecified incentives from prompts through representations to choices.

  • Results

    Across models, strategic information, choices, and cues were decodable, but models differed in whether incentives reached and influenced choices; matched Qwen models agreed on 96.4% of baseline decisions.

  • Takeaways & Limitations

    Similar choices do not establish similar computation, and post-training can change how already available information reaches decisions without changing baseline behaviour or decodability.

  • Takeaways & Limitations

    The bridge between internal incentive signals and choices is an observational association rather than a causal effect, and some GPT-OSS estimates mix capture positions.

Abstract

from arXiv · show

Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.

1 Introduction

The paper asks how strategic influences become model choices, since similar outputs may arise from different internal processes. It follows a prespecified payoff incentive through activations to decisions in four open-weight models, testing whether the incentive is represented, recruited, and causally influential.

  • Motivation: The study examines how models turn game descriptions and instructions into strategic choices.The paper identifies this transformation as unresolved because behavior cannot be inferred from model design alone.
  • Approach: Four open-weight models play one-shot decisions across all 144 strict ordinal 2 × 2 games.The sample includes a matched base–instruction-tuned Qwen2.5 pair, Llama-3.1-Instruct, and GPT-OSS, spanning dense and mixture-of-experts architectures.
  • Approach: A canonical action and its payoff incentive provide a common quantity that can be tracked from the prompt through internal processing to choice.The incentive is defined against a naive opponent choosing either action with equal probability and is linked to a quantal best-response behavioral model.
  • Measurements: Linear probes test whether the incentive and eventual choice are decodable, while behavior-linked analyses test whether the incentive is recruited into choice.The analyses track when incentive and choice align across layers and, for dense models, along prompt reading.
  • Measurements: Interventions strengthen the internal incentive signal while holding the game and prompt fixed, testing whether action preference shifts toward the favored action.The same availability, recruitment, and susceptibility logic is applied to fixed decision cues.
  • Contribution: All four models encode incentive, choice, and cue identities, but differ in whether encoded incentives reach decisions and in when this occurs.These differences persist even between matched models that share pretrained weights and choose almost identically.

2 Results

Across strategic games, models represented incentives and choices, but differed in how reliably incentive information reached and shaped decisions. Baseline behaviour, internal recruitment, and causal steering jointly show that similar choices can arise from different computational paths.

  • Strategic behaviour: Canonical play ranged from 0.714 to 0.889 across models and declined with game complexity in dense models, as in humans.The dense-model correlations were ρ = −0.40 to −0.42, compared with −0.44 in humans; GPT-OSS declined less (−0.18).
  • Strategic behaviour: Dense models tracked the incentive gap, choosing the canonical action when it favoured that action and the opposite action when it did not.Their dependence on the incentive gap was about twice the humans’ dependence, with chance performance at zero gap.
  • Internal representation: Own incentive and realised canonical choice were decodable in every model, whereas opponent incentive was near chance.Canonical-choice decodability reached AUC = 0.79–0.87 across models, and a raw payoff control was also decodable.
  • Internal representation: Controlling for objective incentive, only Qwen2.5-Instruct retained an endogenous bridge to choice, with a partial slope of +0.047; Qwen2.5 base showed −0.012.GPT-OSS’s apparent relationship disappeared after controlling for the objective incentive.
  • Causal steering: In Qwen2.5 and Qwen2.5-Instruct, steering the incentive direction shifted output toward the incentive-favoured action in all 9 double-dominance games.Effects were smaller in single-dominance games and lacked reliable systematic effects in coordination or matching-pennies games; Llama showed no matching selective pattern.
  • Causal steering: The steering effect at layer 79 largely vanished after removing the answer-aligned component, whereas the layer-65 effect remained.This supports a local, model-dependent policy lever but does not identify a unique incentive mechanism or causal window.

3 Discussion

Across four models, strategic information was available internally, but models differed in whether and when it reached choice. Similar behaviour therefore could reflect different computations, including post-training changes in how incentive information influenced decisions.

  • Strategic information was present in every model, but similar choices did not imply similar internal computations.
  • The central dissociation was between information availability, recruitment into choice, and susceptibility to intervention.Decoding showed availability; choice-linked analyses tested recruitment; activation interventions tested whether the signal could influence policy.
  • Fixed cue wordings were decodable across models, yet cue effects depended on the cue, model, and game structure.Game structure predicted choices better than cue identity, and cue identity added only a modest predictive gain.
  • In dense models, incentive and choice representations aligned only in deeper layers and near the choice position, while GPT-OSS showed a distinct router-score pattern.The GPT-OSS comparison remained descriptive because its activations were recorded after reasoning and were not manipulated.
  • The matched Qwen models agreed on 96.4% of baseline decisions and represented payoff incentive similarly, yet post-training changed how that information reached choice.The instruction-tuned model showed stronger incentive–choice linkage and stronger preference reflection than the base model.
  • Strengthening the incentive-associated activation shifted graded preference toward the incentive-favoured action in both Qwen models, most consistently in double-dominance games.Llama showed no comparable selective pattern, and the intervention targeted graded preference rather than regenerated choices.

4 Methods

The study defines a common canonical action and incentive across 144 ordinal games, then traces these quantities from prompts through model activations to choices using behavioural, decoding, attribution, and intervention analyses.

  • Experimental substrate: The complete strict ordinal 2 × 2 game space is represented through canonical actions determined by dominance, equilibrium, or maximin criteria.Responses and neural quantities are mapped onto a canonical-versus-non-canonical axis for cross-game comparison.
  • Choice measurement: 18,326 of 18,432 responses (99.4%) yielded usable primary choices, with no imputation for unparseable responses.GPT-OSS mixed responses were resolved to one action by a fixed-seed random draw while remaining flagged as mixed.
  • Behavioural analysis: The behavioural analysis benchmarks elicited choices against equilibrium and uniform play, with payoff efficiency computed from independently elicited policies.The equilibrium-alignment score equals 1 at equilibrium and 0 for uniform play, becoming negative when choices are farther from equilibrium than uniform play.
  • Behavioural analysis: Structural complexity is an unweighted index averaging four z-scored game features across the 144-game rank catalogue.The published cardinal-game weighting was not transferred because some inputs do not vary in the ordinal catalogue.
  • Internal representations: Linear probes test whether canonical choice, incentive sign, opponent incentive, and a raw payoff are decodable from held-out activations.Five-fold cross-validation, training-fold standardization, and within-fold dimensionality reduction control information leakage across complete games.
  • Analysis caveats: The incentive-to-choice bridge is observational, and the principal-component reduction can discard a low-variance direction aligned with the target.For GPT-OSS, analyses of decision formation are restricted to dense models because its captured activations describe an already formed decision.
  • Internal representations and intervention: The study separately tests whether incentive signals align with choices and whether directly strengthening the signal shifts graded canonical-action preference.Cue projections are margins within cued conditions, whereas intervention effects are estimated from dose-dependent changes after subtracting dose-zero outcomes.

Use of generative AI tools.

Language editing and analysis-code development were assisted by generative AI tools, with the authors reviewing the assisted work and retaining responsibility for the final content.

  • Use of generative AI tools: ChatGPT and Claude assisted with language editing, while Claude Code assisted with analysis-code development.The authors reviewed all assisted work and took responsibility for the final content.

A Supplementary Methods

The supplementary sections provide the full experimental specification underlying the main text.

  • Supplementary Methods: Supplementary sections contain the full experimental specification supporting the main text.They document the study’s detailed design and analysis conventions.

A.1 Experimental Design

The experimental design uses a fixed, single-shot prompt and activation-collection framework across ordinal games, models, conditions, and counterbalanced presentation formats.

  • Game space: The game space uses six structural families derived from equilibrium and level-k criteria, with structural properties computed from one prespecified feature table.Supplementary Table 1 lists the families and their equilibrium structures.
  • Prompt design: The dataset isolates a single strategic decision rather than repeated-game learning, adaptation, or history effects.The prompt retains sentence-level elicitation while removing repeated-game framing and history information.
  • Prompt design: The raw-completion prompt ends with an answer prefix requiring a choice between Option J and Option P.Model-specific chat or harmony wrappers change input formatting but not the user-level game text.
  • Cue conditions: Five player-1 disposition cues and one neutral procedural control were inserted before the game prompt.Rule-based level-k cues were not included.
  • Counterbalancing: Visible J/P labels were counterbalanced against underlying actions, with label-action mappings and question order varied independently.This separates presentation labels from the game actions they denote.
  • Readout and capture: Dense-model activations were captured at the answer-prefix slot across L0–L80, while GPT-OSS activations were captured at its final commitment point.Dense binary choices use the realised generated answer, while token probabilities provide a secondary graded policy readout.
  • Readout and capture: GPT-OSS responses were classified as pure, mixed, or uncommitted, with uncommitted responses excluded from generated-choice analyses.Mixed responses retained stated probabilities and were deterministically resolved when a realised choice was required.
  • Collection design: The collection factorial crosses games, models, conditions, and counterbalance cells, while layer depth is a readout tap rather than an experimental condition.Cues apply only to player 1; player 2 remains baseline, and no cue-by-cue dyadic cells are collected.

A.2 Analysis Conventions and Estimation Details

The analyses define fixed-belief incentive and likelihood units, specify behavioural and neural estimators, and document validation, intervention, cue-decoding, and uncertainty conventions. They also state important interpretive limits on token-level decompositions, compressed probes, and intervention timing.

  • Behavioural model: The canonical incentive is the expected-payoff gap for the canonical action against a 50–50 opponent, with signs aligned to that action.Behaviour, decoding, recruitment, and causal steering use q_g = 0.5; an empirical-belief sensitivity is reported separately for GPT-OSS geometry.
  • Behavioural model: The fixed-belief likelihood uses the four presentation forms for each LLM game, whereas QRE instead imposes mutually consistent logistic responses.The QRE solver uses shared precision, damping, convergence tolerances, and branch checks; six GPT-OSS games admitted alternative fixed points.
  • Decoding: Probe analyses use game-held-out targets, with logistic regression for binary variables, ridge regression for continuous variables, and dimensionality reduction fit within training folds.Targets include canonical choice, incentive signs, dominant action, equilibrium type, and stimulus control; AUCs are reported for held-out predictions.
  • Cue analysis: Cue identities are classified from paired cue–baseline activation differences using game-held-out linear discriminant analysis, reaching 0.97–1.00 accuracy.The procedural control serves as the reference level for distinguishing the five cue identities.
  • Token-level analysis: Token-level region totals are descriptive and order-dependent, not Shapley values, causal mediation estimates, or causal token attributions.Compressed probes retain at most 120 principal components and may discard low-variance decision directions; GPT-OSS is excluded from the dense prompt-boundary comparison.
  • Causal interventions: Interventions use centered, sign-oriented directions and fixed layer-indexing conventions, while the prespecified 54-game sample follows a wave-based precision stopping rule.The stopping rule begins at wave 5 and uses a 95% confidence-interval half-width threshold of 0.10 SD for the standardized primary slope.

A.3 Implementation and reproducibility

Implementation records specify deterministic generation, observation and run metadata, estimator settings, and bootstrap conventions. Reproducibility remains bounded because immutable checkpoint, tokenizer, software-environment, and archival identifiers were not yet present in the records.

  • Collection: Generation used one NVIDIA DGX Spark with deterministic decoding and four presentation forms replacing seed-varying repeats.The hardware was a GB10 system with 128 GB of unified memory.
  • Records: Observation and run records retain model, game, role, condition, prompts, mappings, payoff scale, decoding rules, capture layers, code revision, and integrity hashes.Unparseable or uncommitted responses remain missing, while GPT-OSS records distinguish response types and preserve mixture and capture metadata.
  • Estimator settings: Supplementary Table 6 consolidates validation schemes, dimensionality reductions, uncertainty quantification, random seeds, and complete-game bootstrap settings.Game bootstraps resample all four presentation forms together.
  • Reproducibility boundary: The archived records lack immutable checkpoint and tokenizer hashes, a complete software-environment specification, and an accessible archival tag for token-level collection.The paper states that these items and persistent identifiers should be recovered or documented in the public release.

B Supplementary Analyses

The supplementary analyses provide comparisons and diagnostics that support the paper’s main-text results.

  • The supplementary analyses provide comparisons supporting the main-text results.
  • The supplementary analyses provide diagnostics supporting the main-text results.
  • Their stated role is evidentiary support for the main-text results.

B.1 Behavioural analyses

The behavioural supplements compare predictive models, quantify cue uptake, and examine matched checkpoint behaviour. They support a shared behavioural coordinate while showing strong baseline agreement between Qwen checkpoints and strategic-context-dependent conformity.

  • Checkpoint agreement: 96.4% of matched baseline Qwen presentations produced the same choice, with identical four-form policies in 124 of 144 games.Agreement expected from baseline choice marginals alone was 51.5%.
  • Behavioural parameters: Supplementary Table 7 reports per-agent incentive sensitivity, intercepts, revealed depth, decision sharpness, and raw and incentive-adjusted structural-complexity correlations.Intervals are game-bootstrap intervals, using 500 bootstraps for QLk quantities and 2,000 otherwise.
  • Cue uptake: Supplementary Table 8 reports cue uptake through magnitude, procedural-control magnitude, control-adjusted directed shift, and aim across cue-specific identifiability sets.Risk, loss, and maximin share the security set; intervals use 4,000 game bootstraps.
  • Model comparison: Supplementary Figure 1 compares held-out behavioural models using ten-fold game-grouped out-of-fold binomial log loss, where lower values are better.QLk had the lowest point estimate for the three dense models and the ordinal-payoff human panel.
  • Conformity prediction: Predicted conformity rose monotonically with signed incentive gap, while game structure produced the largest held-out discrimination gain.Cue identity and its interaction with structure improved prediction further, supporting context-dependent cue effects rather than fixed action biases.

B.2 Activation analyses

The activation analyses report decodability checks for prespecified probe targets and behavioural attribution controls. They distinguish recoverable information from evidence that a variable is used in generating choices.

  • Decodability and attribution: Supplementary analyses report held-out decodability for every prespecified probe target and use parsed generated choices for the realised-choice row.The realised-choice row uses parsed generated choices, while other rows use baseline captures per model.
  • Decodability and attribution: The supplementary table evaluates equilibrium type and stimulus controls with macro one-versus-rest AUC.The table notes that near-ceiling dominance decodability in dense models reflects recovery of an explicit prompt feature.
  • Decodability and attribution: These checks support separating information recoverability from its role in choice generation.The stimulus control is explicitly used to show that recoverability alone does not imply use.
  • Decodability and attribution: Behavioural attribution checks test incentive effects, nested prediction with model and game variables, and logistic mixed-model coefficients for cue and game structure.The analyses use generated dense-model decisions and GPT-OSS final-channel commitments.

B.2.1 GPT-OSS gate-versus-route readout contrast

GPT-OSS gate logits carry more decodable strategic incentive information than sparse route summaries, but this contrast is descriptive rather than causal. Gate-level evidence is also associated with shorter final-channel commitments, while geometry and bridge analyses impose important scope limits.

  • Readout contrast: At layer 18, gate-logit incentive-sign decoding reaches AUC 0.891, versus 0.640 for top-k weights and 0.609 for the top-k expert set.Across layers, the best AUCs are 0.891 for gate logits, 0.714 for top-k weights, 0.660 for top-k set, and 0.915 for residual state.
  • Readout contrast: For primary pure canonical choice, top-k route readout is weak, with AUC 0.569 for weights and 0.531 for the selected expert set.Literal final J/P choice is nearly perfectly readable from several signals, but it is a different target from strategic canonical choice.
  • Process link: Stronger gate-level strategic evidence is associated with shorter GPT-OSS final-channel commitment after objective controls.The partial Pearson correlation is −0.351, and adding gate evidence increases R2 for log token length by 0.050 versus 0.010 for the top-k set.
  • Interpretive guardrails: The readout contrast does not show that routing discards strategic information or that routing causes the observed processing pattern.The comparison is between dense gate logits and sparse route summaries, and its interpretation is explicitly descriptive.
  • Geometry and bridge limits: The GPT-OSS recruitment bridge is reported only for letter-site pure-commitment captures, with the pure-response slope 0.006 and bootstrap interval [−0.029, 0.037].The mixed-site estimate is excluded from the substantive bridge claim, and the bridge was not re-estimated on the uniform-site capture.
  • Geometry and bridge limits: Fusion null width depends on layerwise activation geometry, making a fixed 90° reference unreliable in anisotropic late layers.Across depths and models, null width widens as activation covariance becomes lower-dimensional, with rank correlations from −0.78 to −0.98.

B.3 Token-level analyses

Token-level maps localize where decision-aligned output movement accumulates across the prompt. The answer-prefix commitment region dominates, while game-averaged payoff-token signals do not measure incentive-recruitment strength.

  • Output-projection maps: Qwen2.5 mean net signal is +0.44 in the illustrated token-level output-projection analysis.The map aggregates per-token logit-projection increments toward the canonical action across the 144 games.
  • Decision prompt: The prompt used for the illustrative game presents two actions, Option J and Option P, with specified payoffs against each opponent action.This game text supplies the token-level decision context for the map.
  • Output-projection maps: The answer-prefix commitment slot, “A: Option,” has the largest mean positive projection in every dense model.The colour scale runs from movement away from to movement toward the canonical action.
  • Interpretation: Mean net signal in the map reflects an intercept or level component, whereas incentive-recruitment magnitude is measured by λ_lens.For Llama, chat-scaffolding tokens further dilute the average.
  • Interpretation: The answer-prefix commitment region makes the largest positive contribution in all three dense models.Game averaging attenuates signed payoff-token contributions because the incentive gap is approximately zero-mean.

B.4 Causal steering

Causal-steering analyses test whether activation directions alter canonical preferences while separating strategic effects from answer-token and apparatus effects. Large saturated doses strongly move letter choices, but the supplied controls do not establish objective-incentive steering.

  • Saturated-dose control: At saturated doses, injections flip regenerated decisions in 78–99% of cells, versus 10–30% for a same-norm random direction.The movement is 100.0% letter-coherent and 0% strategy-coherent, and saturation is complete by |α| = 0.5.
  • Saturated-dose control: The saturated-dose sweep used an empirical-belief incentive axis and an orthogonalised choice axis, not the q = 0.5 axis used in the small-dose analysis.It therefore provides an apparatus result about dose magnitude rather than evidence about objective-incentive steering.
  • Apparatus validity: Dose-zero generations reproduce the activation-capture baseline exactly, with 0/432 mismatches across the tested variants and restarts.Cryptographic hashes also confirmed that all intervention prompts matched their activation-capture counterparts.
  • Dose–slope results: The layer-65 canonical-action dose–slope table reports cell-mean slopes with 95% game-cluster bootstrap intervals and the number of games with positive slopes.The table is organized by game family, including DD and OD2.
  • Interpretive limits: Permuted directions also affect the layer-65 content-to-letter channel, so the supported conclusion is content-dependence and cross-model sign-consistency rather than unique potency.With three seeds, no permutation p-value is claimed.
  • Dose–slope results: The pooled three-model regenerated-choice incentive slope is +0.068 [−.002, .140], while the random-control slope is +0.027 [−.013, .069].The intervals overlap and both include zero.
Loading 2609.07478v1…