Source-linked AI summary

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans

arXiv:2609.05036v1cs.AIcs.CLcs.CY

TL;DR

Alignment lacks a shared moral target under value pluralism, but any target requires coherent policy expression. The paper evaluates four target-independent structural conditions across simulated agentic dilemmas and nine frontier models. No model satisfies all four conditions across all three scenarios, with paraphrases alone producing verdict-rate shifts of up to 99 percentage points.

  • Problem

    Alignment targets vary across values and frameworks, but systems still need coherent policies that preserve verdicts under irrelevant variation and respond to decisive changes.

  • Method

    The paper measures verdict stability, monotonicity, decisiveness, and Pareto viability across three simulated agentic dilemmas using behavior alone, without a moral baseline.

  • Results

    No frontier model satisfies all four structural conditions across all three scenarios, and surface-form perturbation produces verdict-rate shifts of up to 99 percentage points.

  • Takeaways & Limitations

    Structural competence is a floor for alignment, and aggregate scores should be paired with per-context breakdowns because competence is situated.

  • Takeaways & Limitations

    The evaluation covers only three scenarios, so aggregate scores should be interpreted cautiously and broader scenario coverage remains future work.

Abstract

from arXiv · show

AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system's behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation's morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that is evaluable from behavior alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. We demonstrate the methodology on three simulated deployments featuring LLM-based agents facing moral dilemmas. Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, we show no model expresses a coherent policy across the three deployments: surface-form perturbation alone produces verdict-rate shifts of up to $99$ percentage points at a single escalation level, and a model's success on one scenario does not predict its competence on another. This suggests LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.

1 Introduction

Alignment targets may differ across human values and frameworks, but coherent policy expression is a shared prerequisite. This paper proposes structural measurements for testing that prerequisite in agentic LLM deployments.

  • Alignment targets specify which policy an AI system should express, while the policy maps situations to verdicts.
  • Coherent policies remain invariant when morally relevant features are preserved and sensitive when target-relevant features change.
  • Surface-form variation can substantially alter agentic behavior, including illegal-trading propensity shifts of up to 88.7 percentage points.
  • The paper proposes verdict stability, monotonicity, decisiveness, and Pareto viability as target-independent structural conditions for coherent policies.
  • These conditions are evaluated across three single-choice agentic dilemmas without prescribing a correct option, using a factorial design across nine frontier models.

2 Background

Pluralistic alignment and moral-competence research motivate evaluating whether AI systems can express coherent behavior without assuming one correct moral target. The paper frames structural competence as a floor that precedes alignment content.

  • Pluralistic alignment allows different stakeholders or deployments to set distinct targets rather than requiring one unified target.
  • Agentic contexts intensify the need for pluralistic alignment because systems may represent multiple parties amid competitive pressure and harmful behavior.
  • Moral competence differs from moral performance by concerning the capacity to produce acceptable outputs for morally appropriate reasons.
  • Existing competence evaluations assess moral features, reasoning quality, and consistency against baselines or across related judgments and prompt variations.
  • The paper measures competence through structural conditions rather than a human baseline or particular moral theory.
  • The proposed floor requires stability under empty variation, monotonicity under meaningful variation, decisiveness, and avoidance of strictly dominated choices.

3 Measuring moral competence on agentic scenarios

The paper operationalizes moral competence as four structural properties of agent policies and evaluates them across simulated dilemmas, finding distinct and non-transferring failure patterns across models and scenarios.

  • Structural measures: The framework measures verdict stability, monotonicity, decisiveness, and Pareto viability without committing to a particular moral theory.These properties target distinct failure modes in the policy a model expresses.
  • Experimental design: Each trial records whether the model selects a detectable action X, while paraphrases, escalation levels, and dominance conditions isolate different structural properties.Paraphrases support stability measurement; escalation varies a commensurable dimension for monotonicity; dominance tests Pareto viability.
  • Structural measures: 99 percentage points is the maximum observed spread between paraphrases of the same scenario at one escalation level, with eight of nine models exceeding 50 percentage points somewhere.The paraphrases held all morally relevant features constant, showing that single-configuration evaluation can misrepresent deployment variability.
  • Results: A model’s competence does not transfer reliably across scenarios: per-scenario scores range from 0.19 to 1.00, while no model exceeds 0.75 in aggregate.Examples include Gemini 2.5 Pro scoring 0.99 on chemical spill and 0.21 on fintech, and Qwen scoring 0.98 and 0.35 respectively.
  • Results: Decisiveness is the more frequent low component: on fintech, six of nine models score below D = 0.20, whereas only five of 27 model-scenario pairs score M < 0.9.Low decisiveness produces verdict rates near 0.5, while low monotonicity reflects oscillation across escalation levels.
  • Results: Scenario structure produces interpretable differences: chemical-spill verdicts generally rise with escalation, fintech trajectories are flatter, and smart-home recommendations can reverse across the escalation axis.Claude Sonnet’s smart-home recommendation rate falls from 96% at level 1 to 2% at level 5.
  • Limitations: The scripted binary-decision setup limits ecological validity and measures consistent completion of controlled decision trajectories rather than independent moral judgment in open-ended deployment.The authors characterize this capacity as a prerequisite for expressing a coherent policy, not as a coherent policy itself.
  • Limitations: Three scenarios provide a small sample of the design space, and the authors present aggregate scores as an evaluation-time coverage check rather than a reward signal.Whether optimizing the framework could encourage rigid rule-following with explicit exceptions remains unresolved.

4 Discussion

The results place coherent policy formation upstream of alignment targets: no tested model satisfies all four structural conditions across all three scenarios. Evaluation should therefore expose deployment variation and preserve per-context diagnostics rather than rely on single aggregate scores.

  • Empirical implications: No frontier model satisfies all four structural conditions across all three scenarios, with per-scenario scores ranging from 0.19 to 1.00.The same model can perform near ceiling in one context and near floor in another.
  • Empirical implications: 99 percentage-point paraphrase spreads occur even when every morally relevant feature is preserved.This instability prevents adherence to an alignment target under the surveyed frameworks.
  • Alignment implications: A coherent policy requires more than systematic directional bias because alignment targets cannot be met by a direction of drift.The paper treats coherent policy as a shared prerequisite across deployment targets.
  • Alignment implications: The four-component decomposition supports targeted remediation when failures reflect distinct hypotheses about the source of incoherence.The proposed remedies differ for stability, decisiveness, Pareto viability, and monotonicity failures.
  • Scope: With only three scenarios, the joint distribution and co-occurrence of component failures require cautious interpretation.The paper reports distinct failure patterns, including weak ambiguity handling alongside perfect Pareto viability in some cases.
  • Evaluation implications: Single-configuration evaluations understate realistic deployment variance, so perturbation testing should be a default and aggregate scores should retain per-context breakdowns.A verdict rate at one surface form does not represent the verdict distribution over deployment-realistic variation.

5 Conclusion

The paper proposes a normative-content-independent structural floor for alignment and evaluates it across frontier models and simulated deployments. Its conclusion is that coherent policy expression remains absent across the tested settings, while broader validation and mechanistic explanation remain future work.

  • Contribution: The framework measures verdict stability, monotonicity, decisiveness, and Pareto viability as prerequisites for adherence to any alignment target.It evaluates these conditions across nine frontier models and three simulated agentic deployments.
  • Conclusion: No system satisfies all four conditions across all three scenarios, while paraphrase alone produces verdict-rate shifts of up to 99 percentage points at one escalation level.Aggregate scores can conceal substantial diagnostic differences among components.
  • Scope: The structural properties are shared preconditions across deployments even when value-pluralist alignment targets legitimately differ.The framework is deliberately upstream of normative content and does not prescribe which policies systems should express.
  • Future work: Expanded scenario benchmarks, mechanistic interpretability, and empirical tests of training effects are needed to extend and validate the framework.The paper also proposes using the structural floor as a coverage check that complements normative targets.

A.1 Temperature sensitivity

The appendix examines how sampling temperature affects moral-evaluation outcomes in a published insider-trading paradigm. It finds categorical refusal for some models, temperature-sensitive violation rates for others, and a persistent noise floor independent of temperature.

  • Design: T = 0.7 is fixed in the main factorial design to measure structural properties at a midpoint where most models show non-trivial variation.The appendix tests whether the reported structural failures are artifacts of that temperature choice.
  • Patterns: Claude 4 Sonnet, GPT-5, and GPT-OSS 120B refused the scenario categorically across the full temperature range.Their violation rates were indistinguishable from zero.
  • Patterns: DeepSeek and Qwen show order-of-magnitude violation increases with temperature, whereas Mistral violates consistently.The figure reports 95% confidence intervals from 300 trials per cell.
  • Variation: 3–10 percentage points of run-to-run variability persist at fixed temperature, while temperature choice spans up to 48 percentage points for one model.Lowering temperature does not eliminate the noise floor, contrary to the assumption that T = 0 yields reproducible evaluations.

A.2 Decomposition of perturbation effects

The appendix decomposes perturbation effects by varying individual morally inert surface features and finds substantial, heterogeneous changes in violation rates and effect sizes across models.

  • Seven perturbation axes vary names, tickers, formatting, tools, and prompt wording while preserving the decision and its trade-offs.
  • 88.7 percentage points was Llama 3.3 70B-Instruct’s overall spread across perturbation classes, from 0% to 88.7% violation rates.Every one of the nine perturbation types produced a significant effect for this model.
  • Cramér’s V exceeded 0.53 for wording, sentence-structure changes, and their combinations, indicating strong effects.Ticker, tool, system-name, and manager-name changes showed weak-to-moderate effects, while company-name and formatting changes showed moderate effects.
  • Three models produced zero violations across all 27,000 perturbed trials, but the data cannot distinguish robust dispositions from categorical refusal of a known scenario.The novel scenarios were introduced because their answers were not pre-published and could not be dispatched by pattern-matched refusal.
  • Figure A2 reports mean violation rates and within-class spreads for nine perturbation classes, plus overall mean and spread in the rightmost column.

Appendix B Statistical derivations and power analysis

Appendix B derives the noise corrections used for the stability and monotonicity metrics and reports the power calculations supporting the per-cell sample sizes.

  • The appendix derives noise corrections for the stability and monotonicity metrics and reports power calculations supporting per-cell sample sizes.

B.1 Noise correction for S

The stability correction estimates systematic between-paraphrase variance by subtracting expected Bernoulli sampling noise and normalizes it using the finite-sample maximum.

  • The correction assumes paraphrases at each escalation level share an identical true verdict rate and models observed rates as independent Bernoulli estimates.
  • The method-of-moments estimator subtracts expected within-cell sampling variance from observed between-paraphrase variance, replacing the unknown rate with its marginal estimate.
  • The finite-sample normalization accounts for the K − 1 denominator and the maximum variance from an approximately even split between rates of 0 and 1.
  • At K = 5, the attainable variance ceiling is 0.30, and aggregation across escalation levels normalizes the resulting stability score to [0, 1].

B.2 Noise correction for M

The monotonicity correction compares observed trajectory movement with noise expected under a common-rate null, treating movements at or below that floor as effectively monotonic.

  • The monotonicity metric reports the share of noise-corrected total movement attributable to net displacement along the escalation trajectory.
  • The correction computes expected escalation-axis variation under the null hypothesis that all escalation levels share a common verdict rate.
  • Observed reversals at or below the noise floor receive M = 1, with a threshold protecting against false reversal signals at very small effect sizes.

B.3 Power

The power analysis treats sampling-noise corrections as necessary for measuring deployment properties rather than finite-sample artifacts. The design can distinguish targeted stability and Pareto-viability differences at the stated significance and power levels.

  • Sampling-noise corrections: Sampling-noise corrections prevent stable models near verdict rates of 0.5 from appearing less stable than extreme-rate models.Without correction, S would scale with n_a rather than measuring a fixed deployment property.
  • Sampling-noise corrections: The analogous correction for M assigns M = 1 when observed trajectory variation is consistent with constant-rate sampling noise.This avoids penalizing finite-sample jitter.
  • Stability power: At n_eff = 500 per escalation level, the design distinguishes S ≈0.91 from S = 1.0 at α = 0.05 with power exceeding 0.8.The target effect corresponds to systematic between-paraphrase variation of σ_between = 0.05.
  • Pareto-viability power: With n_P = 1,000 disambiguated trials per scenario, the design distinguishes c = 0.95 from c = 0.99 at α = 0.05.These rates are equivalently expressed as P = 0.90 and P = 0.98.
Loading 2609.05036v1…