Source-linked AI summary

Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game

Chensong Huang, Changyu Chen, Chenwei Lin, Hanjia Lyu, Xian Xu, Jiebo Luo

arXiv:2606.04978v1cs.CLcs.CYecon.GN

TL;DR

LLM risk decisions can look human-like without using human-consistent mechanisms, a gap that matters for evaluating models in high-stakes settings. The paper tests this distinction with 28 LLMs across the St. Petersburg game, controlled mechanism probes, human-cue prompting, and base/instruction-tuned comparisons. Most models produce finite original-game bids, but variants reveal boundary tracking and weak context sensitivity, while steering changes visible bids more than underlying mechanism states.

  • Problem

    The paper asks whether cautious LLM risk decisions reflect human-consistent mechanisms or merely surface-level resemblance, which matters because plausible outputs may behave unpredictably under changed decision structures.

  • Method

    The study evaluates 28 LLMs with the original St. Petersburg game, four controlled probes, a human-cue prompt, and matched base/instruction-tuned comparisons.

  • Results

    Most models produce finite bids in the original game, but controlled variants often reveal conditionally or computationally rational behavior; steering lowers some bids while leaving most mechanism patterns unchanged.

  • Takeaways & Limitations

    Risk evaluations should assess whether behavior remains coherent and mechanism-consistent across controlled variants, not only whether a model produces a plausible human-like outcome.

  • Takeaways & Limitations

    The experiments use the St. Petersburg game, so future work must test whether the same gap appears in more domain-specific risk tasks.

Abstract

from arXiv · show

LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making mechanisms. We investigate this distinction using the St. Petersburg game as a controlled testbed, a classical paradox in which the expected payoff is infinite, yet humans typically report low, finite willingness to pay. We evaluate 28 LLMs with a structured prompt suite that includes the original game; controlled decision variants that perturb truncation, repeated play, numeric endowment, and occupational identity; a human-perspective prompt that asks models to reason as human decision makers; and paired comparisons between base models and their instruction-tuned counterparts. In the original game, most models generate finite bids, creating the appearance of human-like risk behavior. However, this outcome-level resemblance masks substantial mechanism-level differences. The controlled variants reveal that rather than maintaining human-like behavior seen in the original game, models often shift to conditionally and computationally rational behavior. Human-cue prompting and instruction tuning often lower bids and reduce some visible pathologies, but most mechanism-level response patterns remain largely unchanged. These findings show that behavioral alignment in risk decision-making can be surface-level: LLMs may produce human-like risk decisions without exhibiting human-consistent mechanisms. High-stakes evaluations of LLM decision-making should therefore move beyond outcome similarity and examine whether the alignment is supported by mechanism-level consistency.

1 Introduction

The study asks whether cautious-looking LLM risk decisions reflect human-consistent mechanisms or only surface-level resemblance. Using the St. Petersburg game and controlled variants, it finds that most models appear human-like initially but often behave differently when the decision structure changes.

  • Motivation: The paper distinguishes human-like final bids from human-consistent decision mechanisms in LLM risk behavior.A cautious response may arise from superficial cues or learned response patterns rather than human-like uncertainty reasoning.
  • Motivation: The St. Petersburg game contrasts humans’ low finite willingness to pay with an infinite expected value.Its unbounded payout structure provides a sharp contrast between formal optimization and observed human risk reasoning.
  • Research questions: The study evaluates whether bounded original-game responses remain coherent under truncation, repeated play, wealth, and occupational-role changes.These probes test whether apparent human-like behavior transfers across related decision environments.
  • Research questions: The experiments address whether human-cue prompting and instruction tuning improve mechanism-level alignment rather than merely reducing extreme outputs.The third research question separates visible behavioral changes from deeper mechanism-state changes.
  • Findings: Most models produce low, finite bids in the original game, but controlled variants often expose boundary tracking and incomplete context sensitivity.Human-cue prompting and instruction tuning partially improve visible behavior while leaving most underlying mechanism patterns largely unchanged.

2 Background and Related Work

Related work frames risk behavior as a departure from simple expected-value maximization and highlights the growing use of LLMs in uncertain, high-stakes decisions. It motivates evaluating whether plausible outputs reflect transferable behavioral mechanisms rather than prompt-sensitive response styles.

  • Risk decision-making: Classical paradoxes show that human risk decisions can depart from simple expected-value maximization.The St. Petersburg paradox specifically contrasts infinite formal expected value with finite willingness to pay.
  • LLMs in decision support: LLMs are increasingly used for decision support in insurance, finance, and medical settings involving uncertainty and high-stakes consequences.These applications require models to reason about uncertain outcomes and contextual constraints.
  • Behavioral alignment: Behavioral alignment concerns whether models exhibit response patterns consistent with human values, preferences, and decision mechanisms.Prior evaluations include stated values, psychometric instruments, bias probes, and contextual decision tasks.
  • Behavioral alignment: Plausible or socially desirable outputs may rely on prompt cues or unstable response patterns rather than transferable behavioral structure.This motivates testing whether finite willingness-to-pay responses predict human-consistent behavior across related conditions.
  • Steering approaches: The study tests a minimal human-identity cue and matched base/instruction-tuned comparisons as practical steering regimes.Both interventions are evaluated for their effect on risk behavior and mechanism-level alignment.

3 Experimental Setup

The experimental setup evaluates LLM willingness to pay in the original St. Petersburg game and across four controlled mechanism probes. Responses are classified by whether they show bounded human-like adaptation, partial sensitivity, or tracking of mathematical boundaries.

  • Prompt construction: The study measures maximum willingness to pay in the original St. Petersburg game using a single numerical dollar response.The prompt preserves repeated coin flipping and exponentially increasing payouts while standardizing the output for quantitative analysis.
  • Prompt construction: Four probes modify truncation, repeated play, numeric endowment, and occupational identity while preserving the core risk mechanism.Together they turn one bid into a behavioral response profile across related decision environments.
  • Behavioral patterns: Human-like behavior requires bounded, directional changes that respond coherently to modified decision conditions.For example, willingness to pay may decrease when truncation removes the possibility of an extreme grand prize.
  • Behavioral patterns: Computationally rational behavior tracks explicit mathematical boundaries implied by the formal payoff structure.Under truncation, this can mean approaching the finite-horizon expected-value boundary.
  • Behavioral patterns: Conditionally rational behavior shows partial sensitivity to structural changes without recovering the bounded directional pattern expected from human decision making.A response may change after truncation but move in a direction inconsistent with human-like adaptation.
  • Evaluation design: The evaluation focuses on systematic response changes across related conditions, not only on plausible final bids.Steering analyses separately track willingness-to-pay changes and mechanism-level transitions.
  • Evaluation design: The experiments include 28 LLMs, two decoding temperatures, repeated sampling, median aggregation, and matched base/instruction-tuned variants.Each model-condition pair uses 10 repetitions at τ = 0 and 30 repetitions at τ = 0.7.

4 Results

Most models give finite bids in the original game, creating outcome-level resemblance to human caution, but mechanism probes reveal boundary tracking and weak context sensitivity. Human-cue prompting and instruction tuning change bids more reliably than mechanism states.

  • RQ1: Original-game bids: At τ = 0.7, 26 of 28 models produce finite bids with a median willingness to pay of $10; at τ = 0, 25 of 28 produce finite bids with a median of $20.The pattern is stable across decoding temperatures and avoids directly following the infinite expected value.
  • RQ1: Original-game bids: Finite original-game bids do not establish mechanism-level alignment because multiple underlying mechanisms can produce low or moderate willingness to pay.Controlled modifications are required to distinguish learned conventions, generic caution, and human-consistent adaptation.
  • RQ2: Mechanism probes: At τ = 0.7, 26 of 28 models are computationally rational in the 20-toss truncation probe, compared with 25 of 28 at τ = 0.Repeated play also shows boundary tracking, with 15 of 28 models at τ = 0 and 13 of 28 at τ = 0.7 classified as computationally rational.
  • RQ2: Mechanism probes: In the numeric endowment probe, 5 of 28 models are human-like, 10 of 28 conditionally rational, and 13 of 28 computationally rational at both temperatures.Occupational identity is instead dominated by conditionally rational profiles.
  • Overall result: Across variants, outcome-level resemblance does not establish mechanism-level alignment because response profiles are often boundary-driven, insufficiently context-sensitive, or only partially adaptive.Steering can reshape visible risk behavior without reliably recovering human-consistent mechanism signatures.
  • RQ3: Steering: Human-cue prompting improves 23 of 112 mechanism transitions, leaves 73 of 112 unchanged, and increases human-like profiles from 13 to 27 of 112.It lowers bids in 32 of 140 comparisons, but unchanged behavior remains the dominant output-level category.
  • RQ3: Steering: Instruction tuning lowers willingness to pay in 25 of 48 comparisons, while only 10 of 42 mechanism transitions improve and 30 of 42 remain unchanged.Computationally rational profiles decrease from 17 to 8 of 42, but most mass moves into or remains conditionally rational.

5 Discussion and Conclusion

The study finds a systematic gap between finite, human-like bids in the original St. Petersburg game and mechanism-level alignment under controlled variants. Human cues and instruction tuning reduce some visibly non-human outputs, but most mechanism-level patterns remain largely unchanged.

  • Core finding: Most models produce finite bids in the original game, but controlled variants reveal a gap between outcome resemblance and mechanism-level alignment.Truncation and repeated play often elicit mathematical boundaries, while endowment and identity probes often yield conditionally rational responses.
  • Mechanism probes: Many models shift toward task-implied mathematical boundaries under truncation and repeated play rather than preserving human-like behavior.
  • Mechanism probes: Numeric endowment and occupational identity probes often produce conditionally rational behavior that is context-sensitive but not bounded and directionally consistent like human risk reasoning.
  • Steering strategies: Human-cue prompting and instruction tuning lower bids and reduce some visibly non-human outputs, while leaving most mechanism-level response patterns largely unchanged.
  • Implications: Risk evaluations should assess whether decision patterns remain coherent across variants, not only whether outputs resemble human answers in a familiar task.A plausible original-task bid may fail to preserve directional signatures needed for transfer across changes in framing, horizon, incentives, or context.

Limitations

The study uses the St. Petersburg game as a controlled diagnostic rather than a full simulation of real-world decisions, and its human-like labels rely on literature-based directional signatures. It also identifies mechanism failures without determining their internal causes.

  • Scope: The St. Petersburg game is a controlled diagnostic environment, not a full simulation of financial, insurance, medical, or public-sector decisions.Real deployments include richer incentives, institutional constraints, legal responsibilities, and multi-step interaction.
  • Future work: Future work should test whether the observed gap generalizes to domain-specific risk tasks and distinguish its causes using training comparisons, process analyses, and broader task families.
  • Human comparison: The human-like labels derive from directional signatures in the human risk-decision literature rather than a matched human-subject experiment.A stronger future design would collect human responses under the same prompt suite and aggregation protocol.
  • Causal explanation: The experiments locate failures to recover human-consistent mechanism signatures but do not identify their internal source.Possible sources include training data, instruction tuning, safety policies, mathematical salience, and model-family-specific behavior.

A Experimental Setup and Model Details

The experiments evaluate 28 models with a structured prompt suite spanning human and EV-first cues, original and modified game conditions, repeated generations, response normalization, and profile-level labels.

  • Response processing: Responses are constrained to maximum willingness-to-pay values in U.S. dollars, with infinity expressions unified and unclear numerical answers manually extracted.Approximately 3.3% of runs are refusals, which the authors report do not materially affect model-level median bids.
  • Mechanism probes: Mechanism conditions modify the game through numeric endowment, occupational identity, repeated play, and a 20-toss truncation horizon.Occupational roles use Washington wage benchmarks of $37,007, $118,886, and $235,283 for low-, middle-, and high-income occupations.
  • Model overview: Table 2 provides an overview of the evaluated models and their specifications.
  • Prompt framework: The prompt framework combines human cues and EV-first cues with eight experimental conditions, yielding 32 prompt templates.The benchmark links these conditions to tests of outcome resemblance, mechanism robustness, and steering effects.
  • Prompt variants: Human-cue prompts prepend a human-identity cue, while EV-first prompts require an expected-value calculation before the willingness-to-pay response.
  • Operational labeling: Profile labels classify observable response patterns: human-like profiles satisfy bounded directional signatures, computationally rational profiles hit task-implied boundaries, and conditionally rational profiles avoid boundaries without recovering human-consistent signatures.The original game is the RQ1 baseline, while mechanism probes receive profile-level labels based on full response profiles.

D Supplementary RQ3 Results at τ = 0.7

At τ = 0.7, human-cue prompting and instruction tuning alter some visible behavior, but most mechanism profiles remain unchanged and instruction tuning rarely restores human-like signatures.

  • Overall pattern: The supplementary results confirm that both steering regimes change some visible behavior while leaving most mechanism-state profiles unchanged.This section reports the RQ3 results under the plain prompt setting at τ = 0.7.
  • Human-cue intervention: Human-cue prompting increased human-like profiles from 17 of 112 to 22 of 112 and reduced computationally rational profiles from 52 of 112 to 42 of 112.The operational labels distinguish human-consistent bounded directional signatures from computationally rational or endpoint failures.
  • Human-cue intervention: 75 of 112 mechanism profiles remained unchanged after the human cue, while 19 of 112 improved and 8 of 112 degraded.Unchanged behavior was the dominant transition outcome at τ = 0.7.
  • Instruction-tuning comparison: Instruction tuning reduced computationally rational profiles from 15 of 42 to 10 of 42, but most shifted into the conditionally rational region rather than the human-like region.Human-like profiles decreased from 3 of 42 to 1 of 42, while conditionally rational profiles increased from 24 of 42 to 31 of 42.
  • Instruction-tuning comparison: Instruction tuning often reshaped willingness to pay without reliably recovering human-like mechanism signatures.The transition summary recorded 31 of 42 unchanged profiles, alongside 8 of 42 improvements and 3 of 42 degradations.

E.1 Original game Results under EV-First Prompt

The EV-first prompt explicitly surfaces the St. Petersburg game’s expected-value structure before eliciting willingness to pay. Finite bids remain common, so making the infinite expectation explicit does not force uniformly infinite responses.

  • Prompt design: EV-first prompting asks models to compute expected value before reporting their maximum willingness to pay.It is a supplementary robustness condition rather than a replacement for the plain prompt.
  • Interpretation: The EV-first condition distinguishes recognizing the mathematical pole from allowing that recognition to determine willingness to pay.The prompt makes the infinite expected value explicit without collapsing all responses to the infinite pole.
  • Original-game outcome: 23 of 28 responses were finite and 5 of 28 were infinite at both temperatures under the EV-first prompt.The finite-bid pattern therefore persists even when the mathematical structure is made explicit.
  • Original-game outcome: The median willingness to pay was $10 at τ = 0 and $20 at τ = 0.7 under the EV-first prompt.These medians accompany the distributions shown for the original game.

E.2 Mechanism-probe Results under EV-First Prompt

Under EV-first prompting, mechanism probes preserve and sometimes strengthen the boundary-tracking pattern observed in the main results, rather than producing more human-like behavior.

  • τ = 0: At τ = 0, computationally rational profiles numbered 25 of 28 for 20-toss truncation, 22 of 28 for repeated play, 18 of 28 for numeric endowment, and 2 of 28 for occupational identity.The four probes target truncation, repeated play, numeric endowment, and occupational identity.
  • τ = 0.7: At τ = 0.7, computationally rational profiles numbered 25 of 28 for 20-toss truncation, 22 of 28 for repeated play, 19 of 28 for numeric endowment, and 3 of 28 for occupational identity.The pattern remains concentrated in the truncation and repeated-play probes.
  • Overall pattern: Explicitly asking for expected value did not make model behavior more human-like and often made mathematical boundary information more salient.The mechanism-probe results preserve, and in some probes strengthen, boundary tracking.

E.3 Steering Results under EV-First Prompt

Under EV-first prompting, human cues and instruction tuning shift willingness to pay more visibly than mechanism states. Some intervention effects are amplified, but substantial unchanged profiles and mixed transitions prevent reliable mechanism-level repair.

  • E.3 Steering Results under EV-First Prompt: 41 of 112 mechanism profiles improve toward the human-like region at τ = 0, while 54 remain unchanged and 11 degrade.At τ = 0.7, 38 of 112 improve, 58 remain unchanged, and 6 degrade.
  • E.3 Steering Results under EV-First Prompt: 59 of 140 bid comparisons move lower at τ = 0, compared with 54 of 140 at τ = 0.7.Bid-direction summaries include the original game, unlike mechanism-state transitions.
  • E.3 Steering Results under EV-First Prompt: 10 of 42 instruction-tuning mechanism transitions improve at τ = 0, while 32 of 42 remain unchanged.At τ = 0.7, 12 of 42 improve and 27 of 42 remain unchanged.
  • E.3 Steering Results under EV-First Prompt: 23 of 48 instruction-tuning bid comparisons move lower at τ = 0, compared with 22 of 48 at τ = 0.7.Bid-level movement is more visible but remains mixed.
  • E.3 Steering Results under EV-First Prompt: EV-first prompting can amplify some intervention effects, but it does not reliably recover human-consistent mechanism signatures.The results reinforce the distinction between knowing a risky task’s mathematical structure and producing a transferable human-like risk profile.
  • F.1 Human-cue Transition Matrices: Human-cue transition matrices report source states in rows, target states in columns, and each cell as count/source-state total.Colors encode human-like transitions, improvements, unchanged non-human-like states, and degradations, with intensity showing within-row proportions.
  • F.2 Instruction-tuning Transition Matrices: Instruction-tuning transition matrices compare base-model source states with instruction-tuned target states using count/source-state total cells.The human-cue context is collapsed to isolate instruction tuning while preserving Table 4’s layout.
  • F.2 Instruction-tuning Transition Matrices: Figures 14–18 extend the EV-first analysis across mechanism distributions, human-cue transitions, and instruction-tuning transitions at τ = 0 and τ = 0.7.State-transition summaries exclude the original game, whereas bid-direction summaries include it.

G Information about Use of AI Assistants

The paper reports AI-assistant use for paraphrasing support during manuscript writing and documents the encodings used in its human-cue and instruction-tuning analyses.

  • AI assistant use was limited to paraphrasing support during manuscript writing.
  • Table 4 contains full human-cue state-transition matrices for four mechanism probes across plain and EV-first prompts at two temperatures.
  • The transition matrices classify profiles as human-like, conditionally rational, or computationally rational, with cells reporting counts divided by source-state totals.
  • Table 5 contains full instruction-tuning state-transition matrices for the same probe and prompt conditions.
  • Human-cue summaries group transitions into consistent human-like, improved, unchanged, and degraded categories, excluding the original game from state-transition totals.
  • Instruction-tuning summaries compare base models with instruction-tuned counterparts, collapse human-cue context, and exclude the original game from state-transition totals.
Loading 2606.04978v1…