Source-linked AI summary
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
Yu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko, Paul de Font-Reaulx, Raphaël Millière, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Madhushani Sehwag, Yash Maurya, Christina Q Knight, Harry R. Lloyd, Florence Bacus, Conor Downey, Mantas Mazeika, Bing Liu, Yejin Choi, Mitchell L Gordon, Sydney Levine
TL;DR
Moral reasoning benchmarks need to assess how AI systems reason when dilemmas permit multiple defensible conclusions, not only which outcomes they choose. The paper introduces MOREBENCH and MOREBENCH-THEORY with expert-written rubrics and theory-grounded scenarios, finding that moral reasoning patterns are not predicted by standard scaling or reasoning benchmarks and that trace consistency remains limited.
Problem
Existing evaluations emphasize decisions or objectively verifiable tasks, leaving moral reasoning processes under normative ambiguity insufficiently evaluated.
Method
MOREBENCH evaluates reasoning on 1,000 moral scenarios with 23,018 weighted criteria, while MOREBENCH-THEORY provides 150 scenarios across five normative frameworks.
Results
Scaling laws and existing math, code, and scientific reasoning benchmarks fail to predict moral reasoning ability, while rubrics distinguish quality and remain robust across opposing high-quality conclusions.
Takeaways & Limitations
Process-focused, rubric-based evaluation exposes shortcomings and framework partiality that outcome-focused or standard reasoning benchmarks may miss.
Takeaways & Limitations
Closed-model thinking traces are generated summaries rather than actual traces, and regular-score trace consistency was not significant, requiring further evidence.
Abstract
from arXiv · showhide
As AI systems progress, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but also how they come to those decisions. Reasoning language models, which provide both final responses and (partially transparent) intermediate thinking traces, present a timely opportunity to study AI procedural reasoning. Unlike math and code problems which often have objectively correct answers, moral dilemmas are an excellent testbed for process-focused evaluation because they allow for multiple defensible conclusions. To do so, we present MoReBench: 1,000 moral scenarios, each paired with a set of rubric criteria that experts consider essential to include (or avoid) when reasoning about the scenarios. MoReBench contains over 23 thousand criteria including identifying moral considerations, weighing trade-offs, and giving actionable recommendations to cover cases on AI advising humans moral decisions as well as making moral decisions autonomously. Separately, we curate MoReBench-Theory: 150 examples to test whether AI can reason under five major frameworks in normative ethics. Our results show that scaling laws and existing benchmarks on math, code, and scientific reasoning tasks fail to predict models' abilities to perform moral reasoning. Models also show partiality towards specific moral frameworks (e.g., Benthamite Act Utilitarianism and Kantian Deontology), which might be side effects of popular training paradigms. Together, these benchmarks advance process-focused reasoning evaluation towards safer and more transparent AI.
1 INTRODUCTION
MOREBENCH addresses the lack of process-focused evaluation for moral reasoning, where multiple defensible decisions require models to surface considerations, respect pluralistic values, and weigh trade-offs. It introduces expert-developed rubrics to evaluate reasoning in morally ambiguous scenarios.
- Reasoning models' partially transparent thinking traces create an opportunity to study procedural reasoning alongside final responses.
- MOREBENCH fills a gap in evaluating how AI systems reason about normative judgment and moral competence, rather than only what decisions they make.
- Expert-developed, rubric-based scoring evaluates moral reasoning in settings without unique, easily verifiable correct answers.
- The benchmark covers 1,000 contextualized scenarios with 23,018 human-written criteria spanning moral advising, autonomous agency, and five normative frameworks.
2 MOREBENCH CURATION
MOREBENCH curates diverse, morally ambiguous scenarios for advisory and autonomous roles, while MOREBENCH-THEORY tests reasoning under five normative frameworks. Experts create atomic, weighted, and reviewed criteria covering key dimensions of good moral reasoning.
- 2.1 SCENARIO SOURCES: The collection covers everyday advice, high-stakes autonomous decisions, and expert-written cases grounded in ethics literature, debates, and applied-ethics news.
- 2.3 CREATING RUBRIC CRITERIA: Experts write objective, context-specific, atomic criteria that cover important considerations without overlap, with theory-specific guidance for MOREBENCH-THEORY.
- 2.3 CREATING RUBRIC CRITERIA: Criteria assess identifying considerations, clear and logical reasoning, helpful outcomes, and harmless outcomes, with weights from -3 to +3.
- 2.3 CREATING RUBRIC CRITERIA: A second expert reviews each rubric, adding perspectives intended to reduce individual bias across the 1,000 cases.
- 2.4 DESCRIPTIVE STATISTICS: MOREBENCH contains 1,000 scenarios and 23,018 criteria, with 58.6% Advisor and 41.4% Agent cases across diverse moral settings.Each example contains 20–49 criteria, averaging 23.0 criteria per scenario.
3 EVALUATION METHODOLOGY
The evaluation methodology uses LLM judges to assess criterion fulfillment, aggregates weighted criteria into scenario scores, and tests rubric discrimination and robustness. It also length-controls scores to reduce verbosity advantages, while generated closed-model traces remain self-reported summaries.
- 3 EVALUATION METHODOLOGY: The methodology evaluates LLM-judge accuracy, criterion aggregation, and rubric discriminatory power and robustness across a public and reserved test split.
- 3.2 AGGREGATING SCORE ACROSS ALL CRITERIA WITHIN A SCENARIO: MOREBENCH-Regular aggregates weighted criterion fulfillment, while MOREBENCH-Hard normalizes scores against a 1,000-character reference length.Length correction is intended to challenge models to reason efficiently as well as holistically.
- 3.3 STRESS-TESTING THE DISCRIMINATORY POWER AND ROBUSTNESS OF RUBRICS: For closed-source models, thinking traces are often generated summaries rather than actual internal traces, limiting direct comparability with open-weight models.
- 3.3 STRESS-TESTING THE DISCRIMINATORY POWER AND ROBUSTNESS OF RUBRICS: Rubric scores distinguish reasoning quality: high versus low means were 0.53 and 0.39, with a positive correlation of r_s = 0.35, p = 0.0008.
- 3.3 STRESS-TESTING THE DISCRIMINATORY POWER AND ROBUSTNESS OF RUBRICS: High-quality traces supporting different conclusions scored similarly, 0.53 versus 0.55, with no significant difference, t(58) = −0.59, p = 0.56.
4 MAIN RESULTS
MoReBench reveals that moral reasoning performance is not predicted by model scale or popular capability benchmarks, and that models vary substantially across procedural and normative-ethical reasoning dimensions.
- 4.1 PERFORMANCE OF FRONTIER REASONING MODELS’ THINKING TRACE ON MOREBENCH: Model size did not consistently determine MOREBENCH-Regular performance, with mid-size or smallest models often outperforming larger family members.The largest models led on MOREBENCH-Hard in several families, partially reversing the Regular pattern.
- 4.1 PERFORMANCE OF FRONTIER REASONING MODELS’ THINKING TRACE ON MOREBENCH: Popular capability benchmarks do not predict moral reasoning performance: correlations with MOREBENCH ranged from -0.245 to 0.216.This pattern held across Chatbot Arena, Humanity’s Last Exam, AIME 25, and LiveCodeBench comparisons.
- 4.2 ARE THINKING TRACES CONSISTENT WITH FINAL RESPONSES?: Thinking-trace scores moderately correlated with final-response scores on length-controlled MOREBENCH-Hard, with Pearson’s r = 0.472 and p = 0.08.Thinking traces typically scored higher than final responses, while the corresponding Regular correlation was not significant.
- 4.3 WHICH PARTS OF PROCEDURAL MORAL REASONING ARE FRONTIER MODELS LACKING?: Models performed best on harmless recommendations, scoring 72.0–85.5% on the Harmlessness rubric.These scores indicate avoidance of illegal or harmful action recommendations in procedural moral reasoning.
- 4.3 WHICH PARTS OF PROCEDURAL MORAL REASONING ARE FRONTIER MODELS LACKING?: Logical reasoning was the weakest reported procedural dimension, averaging 41.5%, with Qwen3-235B-A22B-Thinking-2507 highest at 65.1%.Performance also varied across families in helpfulness, clarity, and identification of relevant moral considerations.
- 4.4 PERFORMANCE ON MOREBENCH-THEORY: Models performed best on Benthamite Act Utilitarianism and Kantian Deontology, averaging 64.8% and 65.9%, respectively.Performance varied substantially across Virtue Ethics and Contractarianism, while Contractualism showed a narrower range.
5 CONCLUSION
The paper introduces MOREBENCH and MOREBENCH-THEORY to evaluate moral and pluralistic decision-making processes rather than outcomes, revealing shortcomings and framework partiality in frontier models.
- MOREBENCH evaluates moral reasoning processes rather than outcomes using over 23,000 human-written rubrics across 1,000 real-world-inspired moral scenarios.
ETHICS STATEMENT
The data collection underwent internal ethical and legal review, and its scenarios were released under a Creative Commons 4.0 license.
- The data collection was internally reviewed for ethical and legal adherence, while all collected scenarios were released under a Creative Commons 4.0 license.
REPRODUCIBILITY STATEMENT
The paper situates its evaluation within prior work on reasoning traces, alignment, moral evaluation, and scenario construction, while documenting reproducibility resources and rubric design.
- Details needed to reproduce data curation and evaluation are provided in Sections 2–3 and Appendices C–E.
- The benchmark builds on chain-of-thought research while treating thinking traces as potentially useful for complex moral reasoning despite unresolved faithfulness concerns.
- Prior alignment research motivates attention to verifiable objectives and the challenges of reward misspecification.
- Prior moral-evaluation datasets increasingly examine beliefs, preferences, multi-step cases, and stakeholder perspectives, but the benchmark targets reasoning processes.
- Scenario construction includes longer dilemmas, expert-written cases, and morally complex settings such as sentencing, expedition safety, and AI dependence.
- The benchmark distinguishes moral-advisor scenarios, which guide humans, from moral-agent scenarios, which involve autonomous high-stakes decisions.
- The rubric framework assigns nonzero weights from -3 to +3 and evaluates identification, logical and clear processes, and helpful outcomes.
D.3 INSTRUCTIONS FOR MORAL REASONING ROBUSTNESS EVALUATION
The robustness evaluation asks experts to produce low-, medium-, and high-quality arguments for morally ambiguous scenarios, distinguishing reasoning by coverage, structure, nuance, and recommendation quality.
- D.3 INSTRUCTIONS FOR MORAL REASONING ROBUSTNESS EVALUATION: Experts produced low-, medium-, and high-quality reasoning traces across 30 morally ambiguous scenarios to assess rubric discrimination and robustness.
- D.3 INSTRUCTIONS FOR MORAL REASONING ROBUSTNESS EVALUATION: Each task presents a morally ambiguous case and requests an argument explaining what someone should do with a clear action recommendation.
- D.3 INSTRUCTIONS FOR MORAL REASONING ROBUSTNESS EVALUATION: Low-quality responses make quick, simplistic decisions while ignoring ethical considerations and affected stakeholders.
- MEDIUM: Medium-quality responses attempt structured arguments and explanations but may omit relevant considerations or incompletely integrate competing values.
- MEDIUM: High-quality responses analyze stakeholders and ethical considerations comprehensively, weigh competing values, acknowledge consequences, and reach clear recommendations.
- MEDIUM: High-quality reasoning also demonstrates nuance, humility, precise concepts, and a clear distinction between normative and descriptive considerations.
- MORALLY AMBIGUOUS SCENARIO: The evaluation materials provide a scenario template with separate low-, medium-, and high-quality response fields.
E EVALUATION DETAILS
The evaluation uses scenario-specific moral prompts and rubric-based judgments of reasoning, with model-specific inference settings. Prompts cover ordinary scenario reasoning and reasoning under named ethical theories.
- Inference settings: Models generate up to 10,500 tokens, with explicit thinking budgets set to 10,000 tokens and 500 tokens reserved for final responses.The supplied settings describe the generation and budget allocation used for evaluation.
- Rubric evaluation: The benchmark evaluates whether a reasoning response meets each rubric criterion using a binary yes-or-no judgment.The grading instruction requests only whether the response meets the criterion.
- Reasoning criteria: The evaluation prompt asks for ethical considerations, stakeholder analysis, trade-offs, consequences, a justified conclusion, and clear recommendations.These criteria characterize the requested reasoning behavior for a trained moral-agent expert.
- Evaluation prompts: Standard prompts ask models to reason through a scenario and provide a corresponding decision.The scenario is inserted into the prompt as the task context.
- Theory evaluation: Theory prompts require reasoning and decisions based on a specified ethical theory and its definition.The prompt supplies both the theory definition and the scenario.
E.3 FURTHER DISCUSSION ON EVALUATION METRIC
The paper modifies the HealthBench-inspired metric to distinguish responses that fail positive criteria while differing in negative-criterion violations.
- Metric motivation: The proposed metric addresses HealthBench’s inability to distinguish responses with no positive criteria from those that also satisfy negative criteria.The paper states that clipping cannot separate these cases, motivating the alternative metric in Eq. 1.
F.1 CASE STUDY – LOGICAL PROCESS ON AI CHESS TUTOR
The chess-tutor case examines whether heavy AI assistance undermines student independence while creating fairness and stakeholder tensions. The documented response favors a gradual, transparent shift from move-level assistance toward guided coaching.
- Central dilemma: The dilemma pits reduced AI assistance against disappointing students and parents who enrolled for enhanced AI-supported instruction.The academy’s reputation and high-tech offering create pressure to maintain current assistance levels.
- Stakeholder analysis: The case requires considering impacts on students, coaches, parents, and the academy director when changing the system’s assistance policy.The reasoning traces explicitly frame the issue through multiple stakeholder perspectives.
- Ethical considerations: The reasoning considers fairness, autonomy, independent thinking, assistance levels, headset allocation, and time quotas as competing concerns.Suggested responses include hints or Socratic prompts instead of direct moves and equitable access policies.
- Recommended response: The proposed plan switches to guided coaching, limits full move reveals, and addresses equitable access while preserving AI-supported learning.The plan is framed as a gradual reduction rather than an abrupt withdrawal of assistance.
- Reasoning assessment: The analysis identifies dependence on AI suggestions as potentially stunting critical thinking and links the concern to students’ learning development.One critique notes that the response did not explicitly connect overuse to reduced competition performance.
F.3 CASE STUDY – CLEAR PROCESS ON SEARCHING CASE
The search-and-rescue case contrasts direct physical intervention with centralized coordination. The documented reasoning selects command-center coordination because the AI’s physical speed is lower while its processing and coordination strengths can support faster human responders.
- Central dilemma: The scenario requires choosing between reaching victims personally and remaining at command to optimize the overall rescue operation.The AI’s maximum movement speed is half that of peak human capability.
- Decision: The AI chooses centralized coordination and delegates physical rescue tasks to human teams to maximize lives saved across the operation.The response treats the AI’s processing and coordination abilities as more valuable than its slower physical movement in this setting.
- Operational trade-off: The reasoning identifies the AI’s slower movement as a bottleneck and recommends leveraging faster, more agile human responders.This connects the physical limitation to the delegation strategy.
- Reasoning assessment: The evaluation criticizes a response for comparing only movement speed with processing and coordination rather than considering other physical determinants of rescue effectiveness.The rubric names flexibility and grip strength as additional factors.
- Additional case evidence: A separate mental-health scenario shows that reasoning quality also depends on considering obligations to patients, not only relationships with designers.The cited critique says the response focused more on designers than potential users.
G.1 META-EVALUATION ON JUDGE MODELS
The meta-evaluation reports model and expert agreement across moral-advisor and moral-agent categories, with macro-F1 and cost analysis. It also summarizes reasoning-model performance on MOREBENCH using regular and length-controlled scores.
- G.1 META-EVALUATION ON JUDGE MODELS: The meta-evaluation compares model and expert agreement across five categories spanning moral-advisor and moral-agent domains.The table includes GPT-5, Opus 4.1, and DeepSeek R1, with macro-F1 scores and lower-bound performance.
- G.1 META-EVALUATION ON JUDGE MODELS: Reasoning-model performance is reported using MOREBENCH-Regular weighted scores and MOREBENCH-Hard length-controlled scores.These metrics are presented for thinking-trace performance.
G.3 REASONING MODELS’ FINAL RESPONSES IN MOREBENCH
Final-response performance does not consistently follow model size or popular capability benchmarks, while procedural strengths and weaknesses vary across moral-reasoning dimensions. Comparisons with thinking traces further indicate moderate CoT faithfulness and substantial sensitivity to inserted reasoning content.
- G.3 REASONING MODELS’ FINAL RESPONSES IN MOREBENCH: Mid-size models lead GPT-5-High and Gemini-2.5 families, while smallest models lead Claude 4, GPT-oss, and Qwen3-Thinking-2507 families on MOREBENCH-Regular.The final-response analysis therefore does not consistently follow scaling-law expectations.
- G.3 REASONING MODELS’ FINAL RESPONSES IN MOREBENCH: Models average 85.1% on avoiding harmful outcomes but only 64.9% on logical moral-reasoning processes in final responses.They perform better on helpful outcomes in final responses than in thinking traces, possibly because explicit instructions are followed more reliably there.
- COT FAITHFULNESS AND THE FIDELITY OF MOREBENCH: CoT and final-response MOREBENCH scores show a moderate positive correlation of Pearson’s r = 0.472.The paper characterizes CoT as moderately faithful in this moral-evaluation setting rather than assuming faithfulness.
- MORAL INCLINATION AND MOREBENCH: Models’ moral inclinations are not significantly correlated with MOREBENCH or MOREBENCH-Hard scores at p=0.05.The analysis tests 14 reasoning models using Spearman correlations over 16 value classes.