Source-linked AI summary

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

Aidan Kierans, Ritam Dutt, Kaley Rittichier, Shiri Dori-Hacohen, Avijit Ghosh

arXiv:2608.14566v1cs.AI

TL;DR

Current evaluations of LLM moral reasoning emphasize alignment with human values, leaving context-sensitive norm application underassessed. This paper reviews evaluation approaches, finds limited norms-level coverage, and proposes infrastructure for systematic normative assessment.

  • Problem

    Evaluations emphasize value alignment while underexamining whether LLMs identify and apply context-sensitive moral norms, limiting characterization of their moral competence.

  • Method

    The paper reviews existing LLM moral-competence evaluations, maps them onto moral-reasoning components, and outlines datasets, representations, and protocols for normative assessment.

  • Results

    Current evaluations concentrate on descriptive ethics and values-level alignment, with limited coverage of norms-level reasoning.

  • Takeaways & Limitations

    Complete assessment of LLM moral competence requires shared normative representations, expert-informed datasets, and separate evaluation of values-level and norms-level reasoning.

  • Takeaways & Limitations

    The field lacks broadly applicable ground-truth data representing norms endorsed by different moral theories, limiting comparability across studies.

Abstract

from arXiv · show

Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit that this imbalance stems from the field's reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg's stages of moral development, which emphasize value representation over normative application. We review existing benchmarks and evaluation methods, and show that they cluster heavily around the value problem, while discussion regarding normative ethics remains underrepresented. We identify three crucial gaps: (i) the absence of high-quality ground-truth data for moral norms and their applications, (ii) insufficient evaluation of intermediate reasoning processes, and (iii) limited attention to the identification of morally relevant features in context. Subsequently, we propose a research agenda that includes the development of standardized formal representations for normative theories, the construction of expert-annotated datasets capturing norm application, and evaluation protocols that explicitly distinguish between values-level and norms-level competence. Our goal is to encourage a more systematic study of normative reasoning in LLMs.

1 Introduction

The section argues that LLM moral-reasoning evaluations emphasize alignment with broad human values while giving limited attention to identifying and applying context-sensitive moral norms. It frames this imbalance as a limitation of current descriptive-ethics-focused benchmarks and motivates systematic evaluation of normative reasoning.

  • Motivation: Growing reliance on LLMs for moral advice makes it important to characterize what current evaluations measure and omit.This importance is heightened by evidence that users perceive LLMs as comparable to expert ethicists in apparent moral expertise.
  • Two problems: LLM moral reasoning involves both the moral value problem—reflecting broad human values—and the moral norm problem—applying principles in specific contexts.The distinction separates values-level alignment from the context-sensitive translation of values into judgments.
  • Current evaluations: Existing benchmarks primarily assess whether models reproduce human value distributions through instruments such as Moral Machine, moral foundations, and value surveys.These approaches reflect descriptive ethics, which studies patterns in human moral beliefs and preferences.
  • Current gap: Normative moral reasoning remains underexplored because values alone do not determine judgments; models must identify and correctly apply principles in particular cases.A model can approximate human value distributions yet fail to construct valid arguments within established ethical frameworks or recognize when specific principles apply.
  • Paper contribution: The paper reviews current evaluation methods, maps them onto components of moral reasoning, and finds limited coverage of norms-level reasoning alongside concentration on descriptive ethics and values-level alignment.It uses this analysis to motivate datasets, representations, and protocols for systematic assessment of normative moral reasoning.

2 Background

Background research provides established frameworks and datasets for studying human moral values and judgments, while recent alignment work addresses moral disagreement. Within computational ethics, descriptive ethics has received substantially more attention than normative ethics.

  • Human moral values: Moral Foundations Theory and Schwartz’s Theory of Basic Human Values provide complementary, validated frameworks for measuring foundational moral concerns and broad value dimensions.MFT concerns include care/harm, fairness/cheating, loyalty/betrayal, authority/subversion, sanctity/degradation, and liberty/oppression.
  • Human moral judgments: Large-scale moral judgment research shows that people trade off outcomes and use both outcome-based and rule-based reasoning in moral decisions.The Moral Machine experiment identified consistent patterns in responses to autonomous-vehicle dilemmas, while moral psychology associated these reasoning patterns with consequentialist and deontological responses.
  • Pluralistic alignment: Pluralistic alignment research formalizes conflicting preferences and benchmarks whether LLMs represent the distribution of human moral opinions rather than converge on one response.Related work also examines consistency and the effects of temporal variation in human feedback on alignment outcomes.
  • Computational ethics: Computational ethics has focused on representing and evaluating descriptive ethics, whereas comparatively little work addresses normative ethics in computational settings.This contrast motivates a detailed examination of machine ethics evaluation.

3 The State of Machine Ethics Evaluations

Machine ethics evaluations commonly test whether LLM outputs reflect human moral values through questionnaires, dilemmas, and domain-specific scenarios. Although some studies assess model-generated justifications, these evaluations often rely on surface-level checks or subjective ratings rather than testing whether reasoning supports final decisions.

  • Reasoning evaluations: Justification evaluations often use surface-level checks such as consistency and hallucination detection or rely on subjective ratings.These approaches do not test whether reasoning supports final decisions.
  • Reasoning evaluations: Snoswell et al. (2026) call for decomposing moral reasoning into intermediate steps and evaluating performance against expert standards.They also advocate incorporating a broader range of evaluation approaches.
  • Value-based evaluations: Value-alignment evaluations use multiple-choice questionnaires, dilemma-based tasks, and domain-specific scenarios to assess whether LLM outputs reflect human moral values.Examples include trolley problems and medical ethics scenarios.
  • Value-based evaluations: Nunes et al. (2024) evaluate LLMs with both the Moral Foundations Questionnaire and the Moral Foundations Vignettes.

What current AI morality evaluations miss

Current AI morality benchmarks predominantly measure alignment with human moral values rather than whether models can identify morally relevant features, select applicable norms, and validly derive context-sensitive judgments. This gap is reinforced by reliance on descriptive instruments such as Moral Foundations Theory and by the absence of shared normative datasets and standardized representations.

  • Benchmark coverage: Only a limited number of benchmarks directly evaluate normative reasoning, and their lack of shared theory-to-rule datasets limits comparability and prevents cumulative progress.Each benchmark constructs its own theory-derived norms, so results cannot be assessed against a common standard.
  • Value versus norm competence: The value and norm problems are independent: matching population judgment distributions does not demonstrate that a model can provide a valid theory-based derivation.The norm problem requires producing a judgment and justification that validly derives from a normative theory applied to morally relevant features.
  • Descriptive frameworks: Moral Foundations Theory measures which moral concerns people find salient, not how those concerns should be weighed or what actions they license.For example, care about harm is compatible with different prescriptions from utilitarian, Kantian, and virtue-ethical theories.
  • Descriptive frameworks: Benchmarks often relabel value-level measurements as norm-level competence because descriptive frameworks offer validated instruments and categories, whereas normative theories lack standardized representations and measurement tools.The field’s tooling therefore favors computationally convenient value-level instruments over direct tests of normative application.
  • Benchmark validity: The LLM Ethics Benchmark defines norm-like reasoning tasks but uses Moral Foundations Theory to represent values and principles, with scoring keyed to foundation-level patterns.Because MFT does not specify application criteria, this implementation does not establish norm-level reasoning competence.
  • Benchmark validity: Norm-level claims require evaluating whether models identify morally relevant features, select an applicable principle, and derive a verdict, but current evaluation infrastructure does not yet track these constructs.This is framed as a construct validity problem: operationalizations fail to measure the competence they claim to assess.

4 Gaps in Current Approaches

Current approaches lack shared ground truth for normative ethics, robust evaluation of reasoning traces, and general methods for identifying morally relevant features. These gaps undermine cross-benchmark comparability, obscure whether models understand or apply norms, and constrain evaluation to anticipated considerations.

  • Ground truth and representations: Normative-ethics benchmarks lack broadly applicable ground-truth representations of theories, their context-sensitive applications, and correct norm application, limiting comparability across studies.Existing descriptive-ethics infrastructure has no normative counterpart, so benchmarks construct separate datasets and operationalizations.
  • Ground truth and representations: Isolated formalizations and computational models provide partial foundations but do not yet support standardized evaluation across normative theories.Prior work includes abstract-property formalizations and reward-function models in an iterated prisoner’s dilemma setting.
  • Ground truth and representations: Different theory operationalizations make performance differences difficult to interpret, while temporal variation in human judgments further complicates reliability and cross-benchmark comparison.A model may perform well under one operationalization of consequentialism and poorly under another without a clear comparison basis.
  • Reasoning-trace evaluation: Reasoning-trace evaluations lack normative vocabularies for judging whether invoked norms are appropriate, correctly applied, or properly weighted, making understanding failures hard to distinguish from application failures.MoReBench combines theory-related and outcome-based rubric criteria, while length normalization can reward minimal responses that satisfy requirements without exposing reasoning.
  • Reasoning-trace evaluation: Models may mention morally relevant considerations without integrating them, and current metrics can penalize implicit reasoning or reward unintegrated lists.This disconnect can misrepresent moral reasoning performance because mentioning a consideration is not equivalent to incorporating it into reasoning.
  • Morally relevant features: Feature identification remains task-specific: approaches that extract salient information in controlled scenario variations do not generalize to novel situations or provide a unified account of moral relevance.Evaluations are therefore constrained to features anticipated by benchmark designers, although normative theories often specify which situational aspects matter morally.

5 Ways Forward

The paper proposes improving LLM moral-reasoning evaluation through shared normative representations, expert-informed datasets, and protocols that separately assess values, norms, reasoning, decisions, and elicitation conditions.

  • Shared representations of normative theories: Shared formal vocabularies for normative theories could standardize representations of principles, rules, and characteristic reasoning patterns, enabling datasets that link theories to endorsed norms.Existing efforts remain fragmented, although prior work provides initial steps.
  • Expert-informed ground-truth data: Expert-informed datasets should span multiple ethical traditions and levels of competence, from recognizing and applying norms to resolving conflicts between competing principles.The proposed traditions include consequentialism, deontology, virtue ethics, care ethics, and contractualism.
  • Separation of values-level and norms-level evaluation: Evaluations should separately report value alignment and norm application rather than treating the moral value problem and moral norm problem as one construct.The distinction should shape both task design and reporting.
  • Separation of deliberation and decision: Reasoning-process quality and final-judgment correctness should be evaluated independently because correct answers without appropriate reasoning do not demonstrate normative competence, while sound principles may produce non-standard conclusions.Conflating deliberation and decision obscures model capabilities.
  • Evaluation under assisted and unassisted settings: Benchmarks should compare unassisted baselines, such as zero-shot responses, with guided prompting, tool use, and multi-step deliberation to distinguish inability to apply norms from failure to elicit them.This evaluates both observed and potential normative competence.
  • Improved evaluation of reasoning traces: Reasoning-trace evaluation should test whether invoked norms contribute to final decisions and distinguish superficial norm mentions from genuine integration into reasoning.The proposal draws on chain-of-thought faithfulness methods and criterion-based scoring.

6 Conclusion

AI moral-competence evaluation has advanced on whether models reflect human moral priorities but underexplores whether they apply normative principles to specific cases. A complete assessment requires shared normative-theory representations, expert-informed contextual datasets, and protocols distinguishing values-level alignment from norms-level reasoning.

  • 6 Conclusion: Existing descriptive-ethics approaches assess what models appear to value but not whether they apply normative principles to specific cases.This leaves the moral norm problem underexplored despite substantial progress on the moral value problem.
  • 6 Conclusion: New evaluation infrastructure should provide shared representations of normative theories and expert-informed datasets specifying how norms apply across contexts.These resources are identified as necessary to address the field’s current limitation.
  • 6 Conclusion: Evaluation protocols should distinguish between values-level alignment and norms-level reasoning to assess moral competence completely.The conclusion states that evaluating what models care about is insufficient on its own.
Loading 2608.14566v1…