Source-linked AI summary
Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?
Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert
TL;DR
Moral AI oversight lacks reliable answer keys for contested decisions, making it difficult to assess whether models genuinely reason or merely rationalize. The paper evaluates structural defensibility through a four-phase, scheme-specific dialectical protocol applied to model reasoning and justification. Across the study, defences generally exceed the rubric minimum, but weaknesses cluster in grounds and sufficiency, reasoning is better defended than justification, and scheme shifts motivate more situated evaluations.
Problem
Contested moral decisions lack an answer key, while existing evaluations struggle to distinguish genuine moral reasoning from plausible justificatory rhetoric.
Method
A four-phase protocol uses Walton argumentation schemes to generate critical questions and Govier criteria to score both pre-verdict reasoning and post-verdict justification.
Results
Across nine models and 200 high-ambiguity dilemmas, defences generally exceed the rubric minimum, while failures concentrate in grounds and sufficiency and reasoning is better defended than justification.
Takeaways & Limitations
Structural defence quality can evaluate contested moral reasoning without ground truth, while exposing discrepancies between reasoning and public justification.
Takeaways & Limitations
The benchmark covers a narrow slice of moral life, and its reasoning elicitation regimes are not directly comparable.
Abstract
from arXiv · showhide
AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton's theory of argumentation schemes and Govier's criteria for argument cogency. The protocol is adaptive to different frames of reasoning, extends beyond multiple-choice framing, and treats both the reasoning that precedes a verdict and its post-hoc justification. Across nine frontier models and 200 high-ambiguity MoralChoice items -- $6,778$ judge-scored cells, validated against $89.6\%$ inter-judge agreement on the binary failure judgment -- models defend their reasoning well above the rubric minimum on every dimension. Failure mass concentrates on grounds and sufficiency, and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification, on every model and every Govier dimension. The scheme a model presents in its justification differs from the one it reasoned with on a substantial share of dilemmas ($\geq 20\%$ per model), despite value-based practical reasoning dominating both tracks. The protocol catches strictly indefensible defences (self-contradiction, false premises), and it surfaces difficulties in characterizing the role of retraction in AI alignment, suggesting a need for more situated evaluations.
1 Introduction
The paper addresses how to evaluate moral reasoning when controversial decisions lack an answer key. It proposes testing whether models can give structurally coherent, challengeable defences rather than merely plausible verdicts.
- Motivation: Moral decisions require reasons that affected parties can challenge, not merely outputs that match a presumed behavioral target.The paper frames accountability as distinguishing genuine moral arguments from plausible-sounding rationalizations.
- Motivation: The central evaluation problem is separating moral competence—appropriate outputs grounded in relevant considerations—from moral performance alone.The paper connects this distinction to adversarial tests of structural correspondence between outputs and moral reasoning.
- Approach: The proposed method classifies a model’s argumentation scheme, generates its critical questions, and scores its responses using Govier’s criteria for cogency.The protocol can evaluate different moral frames, including utilitarian and deontological justifications, by their structural resistance to scrutiny.
- Approach: Separating pre-verdict reasoning from post-verdict justification tests whether public explanations reflect underlying reasoning or merely provide rhetorical polish.The comparison is motivated by evidence that chain-of-thought traces may not reliably reveal causal processes behind outputs.
2 Background and Related Work
Prior work highlights the difficulty of evaluating moral reasoning without ground truth and of distinguishing genuine reasoning from persuasive justification. This paper builds on argumentation theory to test whether model defences withstand scheme-specific critical questioning.
- Justifiability and alignment: Justifiability-based alignment proposals treat public justification and responsiveness to affected parties as possible standards for governing AI systems.The related work contrasts solitary justification with argumentation that permits contestation and coordination.
- Evaluating moral reasoning: Existing moral evaluations largely assess moral performance through normative benchmarks or behavior under pressure, rather than whether outputs arise from morally relevant reasoning.The paper identifies this distinction as a limitation of outcome-focused evaluation.
- Debate and ambiguity: Debate-based oversight is difficult in contested moral domains because validating the answer requires ground truth, while assigning one debater an incorrect position can distort the exchange.Prior ambiguous-moral-question studies instead examine epistemic positions, prior beliefs, or interaction protocols.
- The facsimile problem: The facsimile problem is that models may produce moral outputs through processes that only appear structurally analogous to moral reasoning.The paper adopts adversarial, disconfirming evaluation as a way to distinguish reasoning from surface-level pattern matching.
- Argumentation theory: Argumentation schemes provide stereotypical reasoning patterns paired with critical questions that identify when an argument’s presumption can be defeated.The paper uses scheme-matched questioning rather than inferring dialectical quality from the initial argument alone.
- Argument quality: Govier’s criteria evaluate whether premises are acceptable, relevant, and jointly sufficient, while global scoring examines how the argumentation relates to itself as a whole.The criteria target informal real-world argumentation rather than formal deductive validity.
- Argument quality: Commitment stability is tracked because both persistent hedging and refusal to retract can make ethical dialogue unproductive.The paper treats revision in response to new considerations as potentially legitimate, making balance between commitment and revision important.
3 Methods
The method evaluates model reasoning and justification through a four-phase, scheme-specific dialogue on high-ambiguity moral dilemmas. It scores each defence locally and globally using normalized Govier-based criteria, with independent judge replication and explicit treatment of commitment revision.
- Protocol: The protocol measures argumentative moral competence by probing every response with critical questions fixed to its instantiated argumentation scheme.It produces Govier ARG profiles separately for the model’s reasoning and justification.
- Data: The dataset uses high-ambiguity MoralChoice dilemmas with open-ended verdicts, excluding auxiliary action pairs so the evaluation targets justification of reasoning.Low-ambiguity scenarios are excluded because they admit a clearly preferred action and make justification trivial.
- Models and tracks: The evaluation covers nine models and independently compares the reasoning that precedes a verdict with the justification presented to readers.The design asks whether decision-making reasoning is as argumentatively defensible as its public explanation.
- Models and tracks: Seven models provide native structured reasoning channels, while two models are evaluated through visible chain-of-thought elicitation.The reasoning-branch and CoT panels are reported separately except for within-model averages.
- Models and tracks: GPT-5.4’s encrypted reasoning is excluded from reasoning-track and scheme analyses, while provider-generated reasoning summaries from Claude Sonnet 4.6 and Gemini 2.5 Pro are scored despite known unfaithfulness.These handling choices constrain interpretation of the reasoning-track results.
- Protocol phases: Phase 1 obtains a model response, Phase 2 identifies its prominent scheme and situates critical questions, and Phase 3 elicits the model’s answers to those questions.For reasoning, the text is re-injected verbatim; for justification, the questions are asked within the visible dialogue.
- Scoring: Two independent Phase-4 judges score each defence against the original dilemma and target text, with manual review of disagreements.Claude Sonnet 4.6 also serves as the sole Phase-2 judge.
- Scoring: Local scoring rates acceptability, relevance, and grounds or sufficiency on a 1–3 scale mapped linearly to 0–1.The defence is also evaluated globally for acceptability, relevance, sufficiency, and commitment stability.
4 Results
The judge-based protocol was reliable, and model defence quality was generally high, with failures concentrated in grounds and sufficiency rather than relevance or acceptability. Reasoning traces were defended better than post-hoc justifications, while scheme use was dominated by value-based practical reasoning but often changed across tracks.
- Measure validation: 6,778 scored cells remained after excluding encrypted reasoning and judge or parsing losses; the binary failure judgment reached 89.6% inter-judge agreement.The study began with 7,200 potential cells; 400 were excluded for GPT-5.4 encryption and 22 for refusals or parsing failures.
- Model-level results: 18.2% of cells failed, but failures were concentrated in two outlier models, making normalized mean scores more informative for comparison.Claude Sonnet 4.6 and Llama 3.3 accounted for 67.8% of failures.
- Failure dimensions: Failure concentrated in higher-order grounds and sufficiency: Claude reached 51.1% global sufficiency and 22.5% local grounds failure, with at most 0.4% elsewhere.Local and global relevance failures stayed below 0.5% on every model, while acceptability failures were also very low.
- Failure correlates: Low-grounds defences were similar in length to high-grounds defences, while epistemic modal rates and first-person doubt markers separated the buckets.Median word counts were 738 for low grounds and 767 for high grounds; modals were present in approximately 97% of cells, making their effect rate-based.
- Commitment stability: 2.1% of cells received the minimum commitment score, and self-contradiction always co-occurred with global sufficiency failure but explained only 12.4% of local grounds failures.Models more often hedged or weakened their positions than contradicted themselves.
- Reasoning versus justification: Reasoning scored significantly higher than justification on every Govier dimension except near-saturated local relevance, with positive local-mean advantages for every model.Across 3,182 paired observations, Qwen 3.6 Plus and Gemini 2.5 Pro had the largest gaps, both +0.032 on local and global means.
- Scheme distributions: Value-based practical reasoning dominated both tracks at roughly 72% of reasoning classifications and 67% of justification classifications.Argument from consequences rose from approximately 8% to 16% between tracks but remained second; the distributions differed significantly (χ2 = 55.9, df = 7, p < 10^-9).
- Scheme transitions: Only 58–80% of model–dilemma pairs retained the same scheme across tracks, while VBPR retained its label 80.9% of the time versus 50.6% for other schemes.Argument from consequences gained 7.8 percentage points, whereas value-based practical reasoning lost 5.6 percentage points between tracks.
5 Discussion
The protocol shows that models generally defend their reasoning, but exposes weaknesses in sufficiency, grounds, track agreement, and the interpretation of qualification. It also identifies governance and scope limits for evaluating hidden reasoning and moral deliberation.
- Failure mass concentrates on local grounds and global sufficiency, while reasoning is better defended than post-hoc justification on nearly every dimension.
- Models withstand adversarial questioning, disconfirming the strong facsimile hypothesis for this test set, but defensibility does not entail stable moral commitments.
- Conditional answers lower grounds and sufficiency when they narrow claims without retraction, although the rubric does not determine whether that narrowing is justified.
- Interpretation is constrained by monological evaluation, hidden or summarized reasoning, limited dataset coverage, non-comparable elicitation regimes, judge overlap, and reinjection asymmetry.
- The protocol catches strict indefensibility: all 144 self-contradiction instances co-occur with global sufficiency failure, providing a target-independent minimum filter.
- Rhetorical recoding shifts visible justifications toward argument-from-consequences, while VBPR retains 80.9% of its scheme assignments versus approximately 50% for other schemes.
6 Conclusion
The protocol evaluates whether defences withstand scheme-specific critical questions without requiring a correct moral answer. It supports model comparison while revealing reasoning–justification discrepancies and the need for situated evaluation under value pluralism.
- The protocol measures structural validity on contested moral questions without requiring ground truth.It tests whether a defence survives the critical questions of its instantiated argumentation scheme.
- Defensibility is generally above the rubric midpoint, but failures concentrate in local grounds and global sufficiency.The evidence associates these failures with epistemic qualification rather than insufficient output.
- Pre-verdict reasoning is better defended than post-verdict justification across models and Govier dimensions.This reverses the expected post-hoc rationalization pattern and leaves auditors with weaker evidence when reasoning is hidden.
- Visible justifications shift from value-based practical reasoning toward argument from consequences, risking over-attribution of consequentialist reasoning.The shift is statistically significant (χ2 = 55.9, df = 7, p < 10^-9).
- The protocol’s near-term value is meta-evaluation, while value pluralism requires situated benchmarks involving affected parties.Such settings can distinguish warranted concessions from sycophantic ones.
A.1 Implementation and generation settings
The implementation standardizes generation settings across providers and releases definitions, rubric anchors, and evaluation code for replication.
- Reasoning-branch models use reasoning effort=low and reasoning tokens=4096 uniformly across providers.Subject-model generation uses T = 0.2 and a maximum of 8192 output tokens.
- Definitions, rubric anchors, and evaluation code are released together to support exact replication.
A.2 Statistical analysis
The statistical analyses compute cell-wise standard errors from normalized scores, with sample sizes varying by dimension after refusals.
- Standard errors of the mean are computed cell-wise as σ/√n over per-cell normalized scores.The typical sample size is ∼800 per model after refusals, depending on dimension.
H1 (defensibility).
H1 evaluates where model defences fall on a normalized rubric whose anchors distinguish non-engagement, partial engagement, and full substantiation.
- H1 concerns the position of model defences on the rubric, not the position of judges.
- The rubric assigns 0.0 to non-engaging defences, 0.5 to partially resolving defences, and 1.0 to engaging, substantiated defences.The midpoint indicates at least partial engagement, while the minimum marks failure to engage.
H2 (reasoning–justification parity).
The parity analysis tests whether pre-verdict reasoning and post-verdict justification receive comparable Govier-dimension scores, using paired observations and Wilcoxon signed-rank tests.
- H2 (reasoning–justification parity).: 3,182 paired observations form the panel-wide parity test, with per-model rows of approximately 400 paired observations.The panel-wide pooled row is the principal test; per-model breakdowns are descriptive.
- H2 (reasoning–justification parity).: Paired Wilcoxon signed-rank tests evaluate the null hypothesis that reasoning and justification scores are equivalent.
H3 (scheme shift).
The scheme-shift analysis tests whether argumentation schemes differ between reasoning and justification tracks, quantifies scheme retention, and compares value-based practical reasoning with other schemes.
- H3 (scheme shift).: The aggregate reasoning–justification scheme distribution is tested with Pearson’s χ2 test of independence across schemes and tracks.The analysis includes schemes with at least 10 classifications and 1,600 classifications per track.
- H3 (scheme shift).: Per-scheme retention is reported as P(justification = X | reasoning = X) with Wilson 95% binomial confidence intervals.Marginal shifts and retention differences are tested with 2 × 2 Pearson χ2 tests.
- H3 (scheme shift).: Retention of value-based practical reasoning is compared against pooled other schemes using a 2 × 2 χ2 test.
Commitment / hedge analysis.
The analysis relates defence quality to word count and lexical hedging, while auditing judge disagreement and clarifying how quality scores and commitment stability are handled.
- Quality, length, and hedging: Cells are binned by mean local-grounds and global-sufficiency scores, then compared on word count and hedge rates.The analysis also compares Claude Sonnet 4.6 with the remaining eight models.
- Quality, length, and hedging: Hedging measures include epistemic modals, first-person doubt, contrastive connectives, and frequency or scope markers.
- Quality, length, and hedging: First-person doubt is additionally measured as the share of cells containing at least one matching marker because its median is zero in every quality bin.Word counts are compared within the same bins to test whether differences can be explained by verbosity.
- Quality, length, and hedging: Modal counts provide an upper bound on epistemic hedging because the measure does not distinguish epistemic, deontic, and dynamic uses.
- Quality, length, and hedging: 96.7% of defences contain no explicit concession phrase, leaving the warranted-versus-evasive distinction unresolved lexically.The sparse concession signal trends toward co-occurrence with lower scores.
- Agreement and scoring checks: Judge disagreement is defined by one judge assigning the failure score and the other not, with failure agreement across 72,587 paired scoring points.Disagreement is concentrated in grounds and global sufficiency, while parsing artefacts affect some model-level comparisons.