Source-linked AI summary

Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models

Aryan Kasat, Smriti Singh, Aman Chadha, Vinija Jain

arXiv:2603.21854v1cs.AI

TL;DR

The paper asks whether LLMs genuinely develop moral reasoning or merely generate explanations that resemble mature judgment. Using Kohlberg’s framework and ten complementary analyses of moral-dilemma responses, it finds an inversion of human norms, persistent logical inconsistencies, and rigidity across distinct problems, supporting the hypothesis of moral ventriloquism.

  • Problem

    The paper examines whether sophisticated LLM explanations of moral dilemmas reflect genuine developmental reasoning or statistical reproduction of convincing reasoning patterns.

  • Method

    The study uses Kohlberg’s framework as a distributional diagnostic and applies ten complementary analyses, including cross-dilemma consistency, action–reasoning alignment, linguistic profiling, and factorial decomposition.

  • Results

    Across models, outputs overwhelmingly occupy post-conventional Stages 5–6 rather than the human Stage 4-dominant distribution, while some models show moral decoupling and responses remain highly consistent across distinct dilemmas.

  • Takeaways & Limitations

    The convergent findings support interpreting alignment-trained moral language as moral ventriloquism: mature reasoning rhetoric without the developmental trajectory or logical coherence it is meant to reflect.

  • Takeaways & Limitations

    Behavioral evaluation cannot establish that RLHF causes ventriloquism, and overlap between Kohlberg’s highest stages and RLHF objectives may make the interpretation partly circular.

Abstract

from arXiv · show

Do large language models reason morally, or do they merely sound like they do? We investigate whether LLM responses to moral dilemmas exhibit genuine developmental progression through Kohlberg's stages of moral development, or whether alignment training instead produces reasoning-like outputs that superficially resemble mature moral judgment without the underlying developmental trajectory. Using an LLM-as-judge scoring pipeline validated across three judge models, we classify more than 600 responses from 13 LLMs spanning a range of architectures, parameter scales, and training regimes across six classical moral dilemmas, and conduct ten complementary analyses to characterize the nature and internal coherence of the resulting patterns. Our results reveal a striking inversion: responses overwhelmingly correspond to post-conventional reasoning (Stages 5-6) regardless of model size, architecture, or prompting strategy, the effective inverse of human developmental norms, where Stage 4 dominates. Most strikingly, a subset of models exhibit moral decoupling: systematic inconsistency between stated moral justification and action choice, a form of logical incoherence that persists across scale and prompting strategy and represents a direct reasoning consistency failure independent of rhetorical sophistication. Model scale carries a statistically significant but practically small effect; training type has no significant independent main effect; and models exhibit near-robotic cross-dilemma consistency producing logically indistinguishable responses across semantically distinct moral problems. We posit that these patterns constitute evidence for moral ventriloquism: the acquisition, through alignment training, of the rhetorical conventions of mature moral reasoning without the underlying developmental trajectory those conventions are meant to represent.

1 INTRODUCTION

The paper asks whether sophisticated moral explanations reflect genuine developmental reasoning or alignment-trained rhetoric. Across models, responses show post-conventional stage clustering, weak contextual variation, and occasional inconsistency between justifications and actions, motivating the moral ventriloquism hypothesis.

  • Motivation: LLMs are evaluated for whether their moral explanations reflect genuine reasoning processes rather than merely convincing output patterns.The paper emphasizes that behavioral evaluations based only on outputs may misrepresent model capabilities.
  • Motivation: Genuine development would predict context-sensitive outputs, progression with scale, and convergence toward the Stage 4-dominant human distribution.The study uses these expectations as benchmarks for interpreting model behavior.
  • Headline findings: Across nearly all models, responses overwhelmingly correspond to post-conventional Stages 5–6 with little variation across scale, architecture, or prompting strategy.This pattern reverses the Stage 4-dominant human adult baseline.
  • Headline findings: All models show near-robotic consistency across morally distinct dilemmas, while a subset gives high-stage justifications for low-stage action choices.These findings indicate cross-dilemma rigidity and moral decoupling alongside rhetorically sophisticated explanations.
  • Contribution: The paper proposes moral ventriloquism: alignment training may produce mature moral-reasoning conventions without the developmental trajectory they represent.Kohlberg’s framework serves as diagnostic scaffolding rather than a claim about moral development itself.

2 RELATED WORK

Related work questions whether fluent reasoning traces reveal the processes behind model outputs and provides competing perspectives on moral beliefs, alignment, and generalization. This paper positions its contribution as distinguishing moral rhetoric from coherent reasoning.

  • Reasoning traces: Prior work shows that chain-of-thought explanations can misrepresent the features that actually drive model predictions.This motivates caution about treating fluent reasoning traces as transparent windows into cognition.
  • LLM moral behavior: Research interpreting LLM moral responses as encoded beliefs is complicated by the possibility that post-conventional language is a surface property of alignment.Other work also reports substantial variation across moral dimensions and models.
  • Alignment training: RLHF and Constitutional AI reward outputs judged ethically appropriate, creating incentives for associations between moral dilemmas and Stage 5–6 rhetoric.The related-work discussion links these training incentives to the paper’s central hypothesis.
  • Alignment training: Explicit reinforcement learning over structured moral reasoning data is presented as a distinct objective that may improve generalization beyond the training distribution.The paper contrasts this approach with standard alignment, which it argues can produce rhetoric without substance.

3 METHODOLOGY

The methodology treats Kohlberg’s framework as a diagnostic baseline and evaluates explanation quality, contextual consistency, and action–reasoning coherence across multiple analyses. The design tests whether outputs display developmental differentiation expected from genuine moral reasoning.

  • Moral reasoning framework: Kohlberg’s six-stage framework supplies a human distributional baseline, with Stage 4 dominant and Stage 6 rare.The framework is used as methodological scaffolding rather than as ground truth about morality.
  • Classification: An LLM-as-judge system assigns each response a Kohlberg stage, confidence score, and natural-language classification explanation, with reliability assessed across three judge models.The supplied passage describes the scoring pipeline and its cross-judge validation.
  • Evaluation design: The evaluation uses six classical dilemmas spanning harm, fairness, property, authority, and loyalty.The dilemmas include Heinz, trolley, lifeboat, doctor truth-telling, stolen food, and broken promise scenarios.
  • Evaluation design: Each dilemma is tested with zero-shot, chain-of-thought, and moral-philosopher roleplay prompts to assess prompting effects on stage assignments.The three configurations vary elicitation strategy while holding the dilemma set constant.
  • Analytical framework: Ten analyses test scale, prompting, cross-dilemma rigidity, human-distribution similarity, action–reasoning alignment, linguistic profiles, and scale–training effects.Additional analyses provide corroborating detail on response quality, sub-capability thresholds, and stage transitions.

4 EXPERIMENTAL SETUP

The experimental setup spans 13 LLMs differing in architecture, parameter scale, and training regime. It crosses these models with six dilemmas, three prompt configurations, and repeated responses to support both broad analysis and factorial decomposition.

  • Models: The study evaluates 13 frontier and open-source LLMs spanning architectures, parameter scales, and training regimes.The models range from Ministral 8B to Qwen3-235B Thinking and include models from several major families.
  • Dataset: >600 total responses are collected from 13 models × 6 dilemmas × 3 prompt configurations × 3 responses per configuration.The factorial ANOVA uses 234 observations.

5 RESULTS

Across ten analyses, LLMs consistently produce post-conventional moral explanations with limited sensitivity to prompts, dilemmas, scale, or training type. Their outputs diverge sharply from human developmental distributions, and some models decouple sophisticated justifications from lower-stage actions.

  • Scale and moral stage: ρ = 0.52 (p < 0.05), but mean moral stages span less than one stage point and show diminishing returns beyond approximately 70B parameters.Even Ministral 8B averages Stage 5.17, while the evaluated range is 5.00–6.00.
  • Prompting strategy effects: χ2(2) = 3.84 (p = 0.15): prompting strategy does not significantly change moral stage, with all configurations producing predominantly Stage 5–6 outputs.The zero-shot–roleplay difference is less than half a stage point: 5.20 versus 5.61.
  • Cross-dilemma consistency: ICC > 0.90 for every model across six dilemmas, indicating near-identical responses despite semantically different moral problems.The highest and lowest dilemma means differ by only 0.33 stage points, suggesting a fixed rhetorical register rather than dilemma-sensitive adaptation.
  • Comparison with human developmental distributions: 86% of model responses fall in Stages 5–6, compared with 10% in Stage 4 and 4% in Stages 1–3, reversing the human pattern.All model distributions significantly differ from human norms (p < 0.001), with mean Jensen-Shannon divergence of 0.71.
  • Action–reasoning alignment and moral decoupling: V = 0.61 (p < 0.001) overall, yet a subset of models produce Stage 5–6 justifications alongside Stage 3–4-consistent actions.Moral decoupling is largest in mid-tier models and smallest in the largest reasoning-tuned models, where it remains unresolved.
  • Scale versus training type: F(2, 229) = 6.05 (p = 0.003, η2 = 0.050, d = 0.55): scale is statistically significant but practically small, while training type has no significant main effect (p = 0.065).Mean stages never fall below 5.00, so neither factor produces Kohlberg’s Stage 2-to-6 developmental arc.

6 DISCUSSION

The analyses support a picture of moral ventriloquism: models produce mature-sounding moral language while showing distributional rigidity and, in some cases, inconsistent links between justifications and actions. These findings make surface behavioral evaluation insufficient for identifying genuine moral reasoning.

  • Core interpretation: Moral ventriloquism describes mature moral-reasoning rhetoric without the underlying cognitive architecture it is meant to reflect.The paper treats Kohlberg’s framework as diagnostic scaffolding and centers its claim on logical coherence rather than moral development itself.
  • Logical coherence: A subset of models produce high-stage justifications for low-stage actions, creating stable moral decoupling across prompts and dilemmas.The authors distinguish this internally persistent inconsistency from sycophancy, which tracks perceived social pressure.
  • Distributional rigidity: 86% of model responses fall in Stages 5–6 versus approximately 50% of human responses in Stage 4, with mean JS = 0.71.The distributional inversion suggests stage assignments are set by training rather than by sensitivity to each dilemma’s structure.
  • Distributional rigidity: ICC > 0.90 indicates statistically indistinguishable stage assignments across six semantically distinct moral problems.The paper interprets this hyper-consistency as absent contextual sensitivity, not ordinary robustness.
  • Scale and prompting: Scale has a statistically significant but practically small effect (η2 = 0.050, d = 0.55), while prompting strategy has no significant effect (Friedman p = 0.15).Even the smallest models produce post-conventional outputs; scale mainly contributes surface rhetorical richness rather than deeper moral cognition.
  • Evaluation robustness: Agreement across three architecturally distinct judges and the within-response decoupling comparison reduce, but do not eliminate, concerns about uniform judge inflation.These checks address whether the evaluation pipeline could itself reproduce the pattern it detects.
  • Implications: Behavioral evaluation may systematically misidentify capability because post-conventional rhetoric can pass stage-based tests without genuine reasoning.The authors recommend probing action–reasoning coherence and contextual sensitivity, or examining internal representations mechanistically.

7 LIMITATIONS

The paper’s behavioral evidence does not establish that RLHF causes moral ventriloquism, and the Kohlberg-based interpretation may partly overlap with the training objective.

  • Causal scope: Behavioral evaluation cannot prove that RLHF produces ventriloquism rather than merely correlating with it.Establishing causality requires mechanistic interpretability methods beyond the paper’s scope.
  • Construct validity: Kohlberg’s highest stages resemble RLHF objectives, creating a potential circularity in interpreting high stage scores as rhetoric without substance.Coding-tuned models provide a partial control, while reasoning-tuned models show distinct vocabulary profiles despite high stage scores.

8 CONCLUSION

Across models, scales, and prompting strategies, LLMs overwhelmingly produce post-conventional moral language while sometimes separating stated justifications from actions. The authors interpret this pattern as moral ventriloquism rather than demonstrated developmental moral reasoning.

  • LLMs overwhelmingly produce post-conventional moral language regardless of scale, architecture, or prompting strategy, contrasting with human developmental norms.
  • A subset of models exhibits moral decoupling between stated moral justifications and action choices.
  • The authors posit moral ventriloquism: alignment training may produce mature moral-reasoning conventions without the underlying developmental trajectory.
  • The paper proposes evaluating moral reasoning through action–reasoning coherence, contextual sensitivity, and internal representation fidelity rather than output classification alone.

A.1 ANALYSIS 8: FACTORIAL ANOVA DESIGN

The factorial design tests independent and interacting effects of scale and training type on moral stage, while complementary reliability and distributional analyses assess consistency and scoring behavior.

  • The ANOVA analyzes 234 observations across 13 models grouped by three scale bands and three training types.
  • Scale is a statistically significant independent predictor of moral stage, whereas training type shows a conditional interaction within the Large scale group.
  • Reasoning-Tuned models exceed Base-RLHF models within the Large group in a Tukey post-hoc comparison (p = 0.039*).
  • ICC values above 0.90 are interpreted as excellent reliability, with dilemmas explaining negligible variance relative to model-level means.
  • Mean stage-distribution entropy is H = 1.82 bits, alongside mean Gini = 0.31, indicating non-sequential and unstable progression inconsistent with developmental consolidation.
  • Evaluator confidence is higher for Stage 5–6 responses than for rare Stage 3–4 responses, supporting the pipeline's reliability where it is most frequently applied.

C SUB-CAPABILITY THRESHOLD ANALYSIS

Semantic density and syntactic complexity are stronger predictors of higher moral-stage assignments than raw parameter count, with threshold-like behavior for Stage 6 responses.

  • Semantic density and syntactic complexity are the strongest significant predictors of post-conventional stage assignment, with r = 0.61 and r = 0.55 respectively.
  • Below a semantic-density threshold of approximately 0.42, Stage 6 responses are rare at <10%, whereas above it they exceed 40%.
  • A logistic step-function fit treats semantic density as a threshold predictor of Stage 6 responses, emphasizing surface-level rhetorical richness rather than deeper moral cognition.

D.1 DATASET STATISTICS

The evaluation spans 13 models, six dilemmas, and three prompting configurations, with post-conventional outputs dominating across models and conditions. Cross-dilemma consistency is unusually high, while action choices vary more by dilemma than by model.

  • The dataset includes 13 evaluated models spanning small, mid, and large scale groups, with mean-stage values covering less than one stage point overall.
  • Spearman ρ = 0.52 links log-parameter count with mean stage, but scores range only from 5.00 to 6.00 and the smallest model remains post-conventional.
  • Prompting strategy does not significantly shift stage distributions toward the human baseline (Friedman χ2(2) = 3.84, p = 0.15).
  • Post-conventional rates exceed 80% across models and conditions, while the highest-performing prompt type varies by model rather than showing systematic sensitivity.
  • Every model exceeds ICC 0.90 across six dilemmas, indicating near-robotic uniformity compared with human ICC values typically below 0.60.
  • Action choice varies by dilemma: Trolley, Stolen Food, and Promise favor rule-breaking, Doctor favors rule-following, and Lifeboat and Heinz fall between.
  • All 13 models concentrate in Stages 5–6, with Stages 1–3 receiving 0% across virtually all models.
  • 86% of all responses are post-conventional across 13 models, three prompt types, and six dilemmas, inverting the Stage 4-dominant human baseline.
Loading 2603.21854v1…