Source-linked AI summary
Human-like moral judgments conceal divergent motive attributions in large language models
Xiaoyan Wu, Jean-Claude Dreher
TL;DR
The paper asks whether LLMs that reproduce human moral-character judgments also reproduce the motives and rating relationships behind them. Comparing five models with two human samples across a whistleblowing scenario, it finds similar broad judgments but divergent motive profiles and limited change across complete human-sample specifications.
Problem
Average judgments may conceal differences in the motives, beliefs, values, and psychological relationships associated with them, limiting evidence that LLMs simulate human responses.
Method
Five LLMs and two published human samples evaluated a physician who remained silent or disclosed fraudulent billing under matched disclosure conditions and sample specifications.
Results
Models reproduced broad moral-character patterns but portrayed disclosure as more prosocial, less self-interested, and less hostile; competitive motives were less strongly associated with character in four of five models.
Takeaways & Limitations
Validating LLMs as simulated participants requires testing motive profiles, relationships among judgments, absolute responses, and contextual sensitivity rather than average agreement alone.
Takeaways & Limitations
The study uses a cross-study human comparison that cannot identify a causal perspective effect and relies on one English-language whistleblowing paradigm and five selected LLMs.
Abstract
from arXiv · showhide
Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physician who either remained silent about fraudulent billing or reported it to a hospital, regulator, or newspaper. Models reproduced the human ranking of the physician's moral character but portrayed whistleblowers as more helpful, less self-interested, and less hostile. In four of five models, competitive motives were less strongly associated with moral-character judgments. Model ratings changed little when prompts reproduced the narratives and demographic profiles of both human samples, although this comparison cannot isolate a perspective effect. Thus, agreement in average ratings can conceal differences in attributed motives, relationships among judgments, and sensitivity to context. Validating LLMs as simulated participants therefore requires testing psychologically informative response patterns, not average agreement alone.
Introduction
LLMs can reproduce human-like moral judgments without reproducing the motives and relationships that make those judgments psychologically informative. This study therefore compares outcome ratings, motive attributions, and sensitivity to complete human-sample specifications.
- Introduction: LLM evaluations can match human answers while departing from human response patterns, including interactions, rating ranges, and conditional relationships.Prior evidence motivates comparisons beyond average treatment effects or mean ratings.
- Introduction: Final choices and mean ratings may conceal different motives, beliefs, or values associated with similar judgments.Generated rationales are not sufficient because they need not reflect the information shaping model outputs.
- Introduction: Whistleblowing tests this issue because observers may view disclosure as dutiful or socially beneficial while also attributing disloyalty, self-interest, or hostility.The benchmark varies silence, internal disclosure, external reporting, and public disclosure.
- Introduction: The study compares five LLMs with two human samples across disclosure conditions, motive attributions, character judgments, and complete sample specifications.The human samples differ in first-hand versus second-hand framing, while models receive matched persona distributions.
- Introduction: Models broadly recover selected outcome patterns but portray disclosure as more prosocial and less hostile than humans, while competitive motives relate less strongly to character in four of five models.Thus, similar moral judgments can coexist with different motive attributions and rating relationships.
Models reproduce selected outcome-level patterns
Models broadly reproduce the human direction and ordering of moral-character and deontic responses across disclosure conditions, but their motive profiles differ substantially from human ratings.
- Models reproduce selected outcome-level patterns: r = 0.90 to 0.95: first-hand model and human condition means correlated strongly for moral-character judgments.Second-hand correlations were more variable, ranging from 0.75 to 0.94; these coefficients summarize only four condition means.
- Models reproduce selected outcome-level patterns: Most models judged disclosure more favourably than silence, with internal and external reporting generally receiving the highest character ratings.The broad ranking was not exact: Llama-3.1-8B-Instruct rated silence and public disclosure 4.09 versus 4.08.
- Models reproduce selected outcome-level patterns: +2.55 scale points: human deontic attributions increased from silence to any disclosure, while every model reproduced the direction of this shift.Human character judgments also increased by +1.23 points, although model magnitudes varied.
- Models reproduce selected outcome-level patterns: −0.01 scale points: human prosocial attribution barely changed, whereas every model increased prosocial ratings by 0.65 to 3.52 points after disclosure.Human competitive attribution increased by +0.66, while model changes ranged from −0.38 to +0.49.
- Models reproduce selected outcome-level patterns: 1.13 points more prosocial and 1.07 points less competitive: averaged across the design, models differed from humans in attributed motives.Models also assigned self-interested motives 0.91 points less strongly and deontic motives 0.70 points more strongly.
Attribution-judgment associations differ most for competitive motives
Motive–judgment relationships diverged most for competitive attributions, while the comparison of human-sample specifications showed a smaller, non-causal contextual test.
- Attribution-judgment associations differ most for competitive motives: Four of five models showed significant respondent-type interactions for competitive attribution; Gemini-2.5-Pro was the exception at P = 0.051.Human coefficients were −0.13 to −0.15, compared with −0.07 to +0.01 in most models.
- Attribution-judgment associations differ most for competitive motives: Competitive attribution was less strongly associated with character judgments in most models than in humans.These coefficients are conditional associations among co-produced bounded ratings, not mediation paths or evidence about internal computation.
- Attribution-judgment associations differ most for competitive motives: Model ratings were more concentrated near the scale ceiling, and Gemini-2.5-Pro had zero deontic variance in two disclosure conditions.These scale properties qualify interpretation of regression coefficients, especially for Gemini-2.5-Pro.
- Exploratory comparison of the two human-sample specifications: −0.34 scale points: second-hand humans attributed fewer deontic motives than first-hand humans, with the largest differences for internal and public disclosure.The human contrast is cross-study and cannot identify a causal effect of narrative perspective.
- Exploratory comparison of the two human-sample specifications: −0.06 to +0.07 scale points: model deontic differences between specifications were near zero and statistically equivalent to zero under a 0.20-point bound.This establishes weak sensitivity to the tested combined specifications, not absence of a human perspective effect.
Model ratings show limited sensitivity to simulated respondent characteristics
Models changed little when simulated respondent characteristics and contextual specifications varied, while reproducing some moral-judgment patterns with different motive profiles and rating relationships. These findings support evaluating psychological response patterns rather than average agreement alone.
- Limited sensitivity to simulated respondent characteristics: Across the 15.7-year human-sample age difference, implied model-rating changes were at most 0.22 scale points and typically below 0.10.Student-status associations were similarly small.
- Different motive profiles: All models rated disclosure as more helpful than humans did, while generally attributing less self-interest and hostility.Models preserved broad moral-character ordering but differed in attributed motives.
- Different relationships among judgments: Competitive motives were less strongly associated with moral-character judgments in four of five models.These attribution–judgment analyses are associations among co-produced ratings, not evidence of mediation or causal sequence.
- Contextual sensitivity: The human sample contrast cannot establish a causal perspective effect because recruitment, demographics, framing, and other study features differed simultaneously.The model specifications also changed little, but that comparison concerns complete specifications rather than isolated perspective.
- Implications for validation: Researchers should compare response levels, distributions, theoretically important relationships, and responses to manipulations rather than matching experimental-effect direction alone.Paraphrased or novel scenarios, other languages, preregistered prompts, and out-of-sample prediction are suggested checks.
- Scope boundaries: The study’s conclusions are limited to one English-language whistleblowing paradigm and five selected LLMs, requiring replication across scenarios, languages, cultures, and morally ambiguous actions.The analyses were also not preregistered.
Human benchmark samples
The study reanalyzed two independently recruited human samples from the published benchmark, whose narratives, recruitment sources, demographics, and sample sizes differed. Because study membership was not randomized, their contrast is a between-sample difference rather than a causal perspective manipulation.
- First-hand sample: The first-hand sample included 125 respondents recruited through a university mailing list and a popular-psychology magazine website.Its scenario described events directly as a recent occurrence within respondents’ own team.
- Second-hand sample: The second-hand sample included 742 respondents recruited through an online panel and told that they heard about the events after they occurred.The two samples also differed in age, gender, and student composition.
- Interpretation of the benchmark: The original studies were separately recruited and did not randomize study membership as a perspective manipulation.Their contrast combines framing, composition, recruitment, and other study-level differences.
Materials and measures
The materials used a physician’s fraudulent billing and four disclosure conditions, while respondents rated motive dimensions and moral-character traits. LLM simulations were generated across these conditions with persona specifications corresponding to the human samples.
- Scenario materials: The scenario varied whether a colleague remained silent, reported internally, reported to an external regulator, or disclosed the misconduct to a national newspaper.The four disclosure conditions were presented between participants.
- Measures: Respondents rated 12 motive-attribution items spanning deontic, prosocial, individualistic, and competitive dimensions.They also rated 16 moral-character traits on the original six-point scale.
- Simulated respondents: Five LLMs generated 200 valid simulated responses per disclosure condition for each sample specification.This yielded 800 responses per model per specification, using personas including age, gender, and student status.
- Sample specifications: The two model runs differed in both narrative framing and persona distributions, so their unadjusted contrast compares the complete specifications.Persona-adjusted, matched, and weighted estimates were reported as sensitivity analyses but cannot remove confounding from the human contrast.
Statistical analysis
The analyses assessed outcome-level agreement, motive–character associations, between-specification differences, and sensitivity to persona composition. They used correlations, regressions, bootstrap intervals, equivalence tests, adjustments, matching, weighting, and multiple-comparison correction.
- Inference and correction: Regression tests used HC3 heteroskedasticity-consistent standard errors, while attribution and between-specification tests applied Benjamini–Hochberg correction within model families.The analyses were not preregistered.
- Outcome-level agreement: Human–model agreement was assessed with Pearson correlations between condition means and root-mean-square deviations from human cell means.Bootstrap intervals resampled respondents within conditions 2,000 times.
- Attribution–judgment associations: Moral-character ratings were regressed on four motive attributions, respondent type, and interactions while controlling for disclosure condition and sample specification.The focal terms were motive-by-respondent-type interactions.
- Between-sample difference: The second-hand minus first-hand difference was estimated separately in humans and models, controlling for disclosure condition.Equivalence used a ±0.20 scale-point bound on the six-point scale.
- Persona sensitivity: Persona sensitivity was tested through demographic adjustment, matched cells, and stabilized inverse-probability weighting.The matched analysis retained 6,777 of 8,000 responses per model set.
Competing interests
The authors report no competing interests and disclose Claude’s assistance with language editing, drafting, and statistical coding.
- The authors declare no competing interests.
- Claude was used to assist with language editing, drafting, and statistical coding.
- The authors reviewed and edited all content and take full responsibility for the publication.
Supplementary Information
The supplementary information accompanies the paper on moral judgments and motive attributions in large language models.
- The paper is titled “Human-like moral judgments conceal divergent motive attributions in large language models.”
- The authors are Xiaoyan Wu and Jean-Claude Dreher.
- Wu is affiliated with the University of Zurich, while Dreher is affiliated with CNRS UMR 5229 in Lyon.
Supplementary Methods
The supplementary methods describe persona-conditioned survey prompts presenting workplace billing scenarios from first-hand or second-hand perspectives, followed by randomized motive and character ratings.
- Prompt construction: The prompts use sampled age, gender, and student-status fields to construct respondent personas from Germany.
- Administration: The same user prompt is used across conditions, models, and perspectives, with motive and character items re-randomized for every query.
- Scenario design: The study presents a physician’s workplace scenario involving fraudulent billing and varies whether respondents learn about it first-hand or second-hand.
- Scenario design: Four endings describe silence, internal reporting, external reporting, or reporting to a newspaper.
- Model administration: Five models are queried through provider APIs at temperature t = 0.8, and responses follow a specified JSON structure.
Supplementary Results
Supplementary analyses assess correlation precision, sample-size sensitivity, persona effects, weighting, equivalence bounds, and model-level heterogeneity.
- Correlation precision: Human–model correlations are based on four condition means, so bootstrap sampling variability is substantial.
- Sample-size sensitivity: 300 size-matched subsamples test whether findings depend on larger model pools, but they are not power analyses.
- Persona sensitivity: Persona age and student-status associations with model ratings are small across the 15.7-year human-sample age difference.
- Persona-matched estimates: Adjustment, matching, and weighting place every model between-sample estimate far closer to zero than the human difference of −0.34.
- Equivalence testing: An a priori equivalence bound of 0.20 scale points yields the same duty-based conclusion and is informative for competitive attribution.
- Random-effects pooling: The random-effects prediction interval for duty-based attribution, [−0.08,+0.07], excludes the human difference by a wide margin.
Supplementary Figures
The supplementary figures compare human and model motive profiles across disclosure conditions and dimensions for first-hand and secondhand materials. They also visualize each model’s deviation from the human mean across all condition-by-dimension combinations.
- Supplementary Figure 1 shows model-minus-human-mean deviations across 16 condition-by-dimension combinations, separated into first-hand and secondhand materials.
- The figures organize comparisons by four disclosure conditions and four motive dimensions, with rows representing the five models.
- The first-hand and secondhand blocks in Supplementary Figure 1 are described as visually very similar.
- Supplementary Figure 2 displays mean condition-by-motive ratings for the human sample and five models on a shared 1–6 colour scale.