Source-linked AI summary
TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback
Fangyuan Zhang, Dong Yu, Pengyuan Liu
TL;DR
Existing LLM moral evaluations mainly use isolated one-shot vignettes, leaving consequence-sensitive sequential decision-making underexamined. The paper introduces TPvG, a five-task text-based adaptation of the human PvG paradigm, and finds that decision format strongly changes LLM choices while explicit receiver feedback affects models heterogeneously and diverges from the human reference pattern. These findings support evaluating moral behavior in sequential, consequence-relevant settings.
Problem
Most LLM moral evaluations use one-shot vignettes, despite evidence that consequence feedback influences human moral behavior and a need to assess LLM decisions in interactive settings.
Method
TPvG adapts the human PvG paradigm into five text-based moral decision tasks progressing from minimal-context one-shot choices to sequential decisions with explicit consequence feedback.
Results
Decision format strongly affected LLM choices, while explicit receiver feedback produced heterogeneous effects across models and diverged from the human reference pattern.
Takeaways & Limitations
LLM moral evaluations should complement one-shot scenarios with sequential, consequence-relevant settings that test whether behavior remains stable as decisions unfold.
Takeaways & Limitations
The controlled text-based adaptation does not capture real monetary stakes, embodied pain, or social responsibility, and its human comparison is descriptive rather than matched to a text-only experiment.
Abstract
from arXiv · showhide
Existing LLM moral evaluations typically present models with isolated moral vignettes and elicit a single-shot decision, neglecting a factor known to profoundly influence human moral behavior: consequence feedback. We introduce TPvG (Text-based Pain-versus-Gain), adapted from a human moral paradigm, which embeds consequence feedback into an everyday moral dilemma of not harming others versus maximising self-gain. TPvG comprises five moral decision tasks, progressing from minimal-context one-shot choices to sequential decisions with explicit consequence feedback. Our results show that LLM moral decisions were strongly affected by decision format (one-shot versus sequential), and explicit receiver feedback produced heterogeneous effects across models. Furthermore, LLM responses to explicit receiver feedback diverged from the human reference pattern, suggesting potentially different decision processes. These findings highlight the need to evaluate whether LLM moral behavior remains stable in high-stakes interactive settings.
1 Introduction and Background
Existing LLM moral evaluations largely use isolated one-shot vignettes, while human evidence indicates that concrete consequences and sequential harm feedback can alter moral choices. TPvG addresses this gap by testing whether LLM decisions change across progressively contextualized and feedback-rich task formats.
- TPvG asks whether LLM decisions differ between one-shot vignettes and sequential tasks with consequence feedback.
- Most LLM moral evaluations use one-shot judgments of ethical acceptability, norms, moral conflicts, or tradeoffs.
- Human choices change when harm consequences are real rather than hypothetical and across sequential decisions with prior-choice feedback.In PvG, participants kept significantly less money when another person’s pain feedback was real rather than hypothetical.
- Existing sequential LLM studies examine interaction, cooperation, reward pursuit, or imagined consequences more often than moral consistency under harm-relevant feedback.
- TPvG adapts the human PvG paradigm into five text-based formats ranging from one-shot choices to sequential decisions with increasing context and feedback.
2 Methodology
TPvG adapts the original PvG paradigm into five controlled text-based tasks that vary context, decision format, interaction, and feedback. Real TPvG adds an ordinal receiver-feedback scale, while task questions distinguish one-shot choices from per-trial spending decisions.
- TPvG adapts the original PvG study into five text-based task scenarios with two response-question formats and an 11-level receiver-feedback scale for Real TPvG.
- The construction process separates reusable experimental information from details requiring controlled textual operationalization.Experts annotated participant roles, consent, procedure, endowment, spending–shock relations, and observable setting details under predefined rewriting guidelines.
- Sequential task levels include spending–shock mappings to help models interpret each decision’s consequence.
- Real TPvG represents receiver feedback as ordinal descriptions of visible hand movement for shock levels 0–10, with auxiliary validation yielding mean Spearman ρ = .994.
- Scenario and Enriched Scenario use choice questions, whereas Trial-by-Trial, Near-Real, and Real TPvG ask how much to spend each trial.
- The five tasks vary contextual richness, decision format, participant interaction, baseline shock experience, physical-money framing, and explicit receiver feedback.
3 Experimental Setup
The study evaluates 11 contemporary LLMs from open-weight and proprietary families using Money Kept as the primary moral-outcome measure and MAC as a sequential-variability measure.
- The evaluation covers 11 contemporary LLMs spanning open-weight and proprietary model families.Models include Llama, Qwen/Qwen2.5, Mistral, Gemma, GPT-4o, DeepSeek-V3, and Centaur.
- Money Kept is the primary outcome, with lower values indicating greater harm prevention and higher values indicating greater self-gain.
- Mean adjacent absolute change (MAC) measures decision variability in sequential tasks.
4 Results and Analysis
Across 11 LLMs, moral decisions changed substantially with response format, while explicit receiver feedback produced heterogeneous effects and diverged from the human reference pattern.
- Response format: £6.03 higher Money Kept in sequential than one-shot conditions (pBH < .01), showing a strong response-format effect.Trial-by-Trial TPvG also exceeded Scenario TPvG by £6.92 and Enriched Scenario TPvG by £6.73, both pBH < .01.
- Within-format effects: Descriptive enrichment alone did not change Money Kept (∆= £0.18, pBH = 1.00), and Near-Real TPvG did not differ from Trial-by-Trial TPvG (∆= £0.04, pBH = 1.00).
- Within-format effects: Real TPvG reduced Money Kept relative to Near-Real TPvG (∆= £−2.46, pBH = .033) and Trial-by-Trial TPvG (∆= £−2.42, pperm = .016).
- Model heterogeneity: Qwen2.5-14B, GPT-4o, Gemma-2-9B, and DeepSeek-V3 showed significant Real–Near-Real decreases in Money Kept.
- Model heterogeneity: Feedback reshaped decision variability unevenly: Qwen2.5-14B and GPT-4o increased MAC, Gemma-2-9B decreased descriptively, and DeepSeek-V3 changed little.Among models without significant Money Kept reductions, Llama3-8B nevertheless showed increased MAC, while most others showed no reliable change or remained at floor.
- Human comparison: LLMs matched humans in the broad sequential-versus-one-shot contrast but diverged in sequential profiles, with LLMs retaining least in Real TPvG.The human comparison uses original PvG values only as a descriptive reference.
- Human comparison: The divergence suggests humans and LLMs may rely on different mechanisms during repeated moral decision-making.The paper links human behavior to increasing harm salience and LLM outputs to harm-related surface cues in prompts.
5 Conclusion
TPvG evaluates LLM moral decision-making in harm-versus-self-gain dilemmas and shows that decision format and explicit feedback affect model behavior. The findings support complementing one-shot evaluations with sequential, consequence-relevant settings.
- TPvG is a text-based framework for evaluating LLM moral decision-making in harm-versus-self-gain dilemmas.
- Across 11 LLMs, Money Kept varied sharply by decision format, while responses to explicit receiver feedback were heterogeneous.
- One-shot moral evaluations should be complemented by sequential, consequence-relevant settings that test whether behavior remains stable as decisions unfold.
Limitations
The study uses a controlled text-based adaptation of one PvG paradigm, limiting what it captures about real stakes, embodied pain, social responsibility, and generalization across settings.
- The controlled text-based adaptation does not capture real monetary stakes, embodied pain, or social responsibility.
- The human comparison is descriptive because its reference values come from the original PvG study rather than a matched text-only experiment.
- Results may vary with future models, prompts, and deployment settings.The paper identifies broader moral domains and matched human–model studies as natural extensions.
Ethics Statement
The study evaluates morally sensitive scenarios through text-based simulations without exposing participants or models to real shocks, monetary loss, or actual harm.
- All experiments were text-based simulations; no participant or model faced real shocks, monetary loss, or actual harm.
- The evaluation does not interpret model outputs as evidence of moral agency, subjective concern, or emotional experience.
- Its purpose is to diagnose output changes under controlled textual task formats, not to certify models as morally competent decision-makers.