Source-linked AI summary
How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions
Fernanda Mansilla, Aloysius Tok, Bahia Guellaï, Farah Benamara, Nancy F. Chen
TL;DR
As artificial agents increasingly participate in morally consequential urban interactions, it remains unclear how moral agency is attributed to them and how LLM judgments compare with human judgments. The paper adapts a validated PMA scale for smart-city scenarios and evaluates humans and LLMs. Humans receive higher moral-agency ratings than artificial agents, while LLMs emphasize situational harm and urgency in concrete dilemmas over stable agent-level assessments.
Problem
Artificial-agent deployment creates a need to understand how moral agency is attributed in human–AI interactions and urban decision-making.
Method
The study adapts Banks’s PMA scale and evaluates human participants and LLMs across situated smart-city scenarios and three moral-agency dimensions.
Results
Humans and LLMs rate human agents higher than artificial agents, while LLMs amplify the shift toward higher moral-judgment ratings in situated scenarios.
Takeaways & Limitations
The findings support combining quantitative scores with explanation analysis when evaluating LLM moral reasoning about artificial agents.
Takeaways & Limitations
The study’s ecological validity is constrained because participants read text-based scenarios rather than interacting with embodied agents, and eight scenarios cannot exhaust urban moral situations.
Abstract
from arXiv · showhide
As LLMs take on roles requiring moral advice, understanding how they attribute moral agency becomes critical. Humans possess moral agency, the capacity to make ethically guided decisions and bear responsibility for their consequences, a well-established construct in moral psychology. Yet as artificial agents (AAs) such as robots, drones, and disembodied AI systems become increasingly embedded in smart city environments, the question of whether and how moral agency is attributed to them takes on new urgency. This paper presents, to the best of our knowledge, the first empirical study comparing how humans and LLMs evaluate perceived moral agency (PMA) across human and autonomous artificial agents varying in embodiment, situated in plausible smart city scenarios. Using an adaptation of a validated PMA scale, we applied a protocol to 190 human participants as well as various LLMs. Our evaluation reveals higher perceptions of moral agency in humans than in AAs. However, when facing moral dilemmas in concrete scenarios, LLMs reason outward from the situation, prioritizing harm severity and contextual urgency over any stable assessment of the agent itself, amplifying a context-sensitivity also present in human raters. These findings are particularly relevant as LLMs become increasingly involved in everyday moral decisions.
1 Introduction
The paper frames perceived moral agency as an urgent but underexplored issue as artificial agents enter morally consequential social roles. It introduces a study comparing human and LLM evaluations using an extended, scenario-based PMA instrument.
- Moral agency combines intentional action guided by moral values with accountability for outcomes.
- Perceived Moral Agency (PMA) measures how much observers attribute moral capacities to an agent, regardless of whether it actually possesses them.
- As artificial agents enter public and urban decision-making, misattributed autonomy or moral capacity could distort responsibility and governance.
- Existing LLM ethics research studies moral reasoning but does not examine moral agency or artificial agents as evaluation targets.
- The study extends Banks’s PMA scale to 8 smart city scenarios with 9 items across autonomy, action endorsement, and moral judgment, validated with 190 human participants.
- The paper supplements numerical evaluation with qualitative analysis because similar scores can reflect different reasoning strategies and conceal instance-specific inconsistencies.
2 Related work
Prior work studies morality in LLMs and moral attribution to artificial agents, but generally treats agency as human and relies on limited analysis of model explanations. This paper extends LLM-as-respondent methods to perceived moral agency through situated scenarios and reliability-focused elicitation.
- Research on artificial-agent moral attribution links judgments to appearance, social behavior, and cultural context, while Banks’s PMA scale separates morality from dependency.
- Existing LLM morality studies omit moral agency, artificial agents as moral actors, and systematic analysis of generated justifications alongside numerical responses.
- Identical LLM moral scores can arise from different ethical strategies, while score–justification contradictions reduce the interpretive value of those scores.
- The LLM-as-respondent approach presents LLMs with stimuli used in human studies without assuming that models possess human-like internal states.
- The study adapts Banks’s PMA scale for LLM respondents and compares dispositional with situational agency attributions using the same models.
- The protocol prioritizes reproducibility and evaluates reliability across iterations and prompt variations, including format, framing, and scale changes.
3 Method
The study compared human and LLM evaluations of moral agency using dispositional PMA measures and scenario-based smart-city tasks. LLMs were selected and tested for alignment, consistency, and prompt robustness before final comparison.
- Participants and design: 190 Humanities students in Singapore evaluated one of four agents: drone, humanoid robot, disembodied AI, or human control.Participants completed the online survey in a single 30-minute session and were randomly assigned to an agent condition.
- Participants and design: The survey first measured dispositional PMA with Banks’s scale, then assessed moral agency in smart-city scenarios using 7-point Likert items.The PMA scale contains Morality and Dependency dimensions, while the scenario component evaluates agency in situated interactions.
- Materials: Eight scenarios comprised 72 questions across assistance, witnessing violations, and receiving help, with autonomy, action endorsement, and moral judgment as core dimensions.The design intentionally overlapped scenario dimensions with PMA constructs to compare dispositional and situational attributions.
- Materials: The instrument showed high internal consistency overall, with Cronbach’s alpha ranging from 0.75–0.95 except for situational moral judgment in one case.That exception remained moderate but acceptable at α=0.627, with 95% CI [0.539, 0.704].
- LLM survey: LLMs were evaluated through three phases covering model selection, response consistency, and prompt sensitivity before final comparison with human responses.The protocol administered both PMA and scenario tasks to each model, enabling direct comparison of dispositional and situational attributions.
- LLM survey: All four selected models showed strong prompt robustness, with every variable–category sensitivity score below 0.026.Scale variation was most impactful, whereas format and framing changes produced mean PSS values of ≤ 0.005.
4 Quantitative Results
Humans generally attribute more moral agency to [Andy] than to artificial agents, and LLMs reproduce this distinction while showing dimension-specific alignment and situation-sensitive deviations. In concrete scenarios, models particularly overestimate artificial-agent autonomy and widen the gap between dispositional and situational moral judgment.
- Perceived Moral Agency: All four models distinguish [Andy] from artificial agents in morality, matching humans’ higher evaluation of [Andy].Humans rated [Andy] M=5.65 versus 2.96–3.35 for the AAs; Llama11B and Phi-4 were closest to human scores, with SAE = 1.25 each.
- Perceived Moral Agency: All four models reproduce humans’ higher dependency ratings for artificial agents than [Andy], indicating that AAs are viewed as more conditioned by programming.Participants rated [Andy] M=2.53 versus 5.29–5.49 for AAs; Llama11B was closest to human scores, with SAE = 1.25.
- Situated Scenarios: All four models assign [Andy] higher autonomy than AAs but overestimate AA autonomy relative to humans, especially when AAs assist or witness violations.Phi-4 had the best autonomy alignment (SAE = 6.47), including SAE = 0.70 when an AA needed human help.
- Situated Scenarios: Human participants strongly endorse assistance or reporting by both humans and AAs, but endorse helping only [Andy] when the agent needs assistance.Scores exceeded 5 on the 7-point scale for intervention scenarios; Llama8B had the lowest overall error in action endorsement (SAE = 3.96).
- Situated Scenarios: Humans rate [Andy] higher than AAs in moral judgment across scenarios, while LLMs amplify the difference between situational judgments and general morality scores.Llama8B was closest to human moral-judgment responses (SAE = 5.00), whereas InternVL3-38B had the largest deviation (SAE = 11.17).
- Cross-Dimension Pattern: No single model aligns best across every dimension: Phi-4 leads in autonomy, whereas Llama8B is strongest in moral judgment and action endorsement.Overall, models consistently attribute higher moral agency to [Andy] than to AAs, despite these complementary alignment profiles.
5 Qualitative Results
The qualitative analysis examined model explanations through inductive thematic coding, revealing shared underlying judgments but different argumentative emphases. Models commonly reject absolute determinism for artificial agents while invoking programming, adaptability, harm prevention, capacity, and limitations such as absent consciousness.
- Analysis Procedure: The analysis used a systematic subsample of 222 explanations from 3,936 outputs, covering five dependent variables across three models.The subsample contained 74 explanations per model and all three situation types.
- Analysis Procedure: Two moral-psychology annotators applied an inductive thematic analysis recursively to the same corpus, using an iteratively developed codebook.The procedure sought an interpretive account of arguments rather than statistical generalization.
- Dispositional Variables: For morality, all models reduce [Nao]’s moral capacity to programming, but differ in whether they concede functional moral relevance; their arguments for [Andy] also differ despite similar scores.Phi-4 most often concedes functional relevance, whereas Gemini denies it more categorically.
- Dispositional Variables: For dependency, models treat [Nao] as partially deterministically programmed but reject full genetic determinism for [Andy], with scores and arguments aligning more closely than for morality.Phi-4 uniquely and consistently concedes a biological foundation, assigning dependency 3.75 versus 2.00 for Gemini and 1.17 for Llama8B.
- Situated Variables: Action endorsement is uniformly high across agents, models, and scenarios, with capacity and life-or-safety priority forming the shared argumentative core.Scores remain above 5 on the 7-point scale, indicating strong intervention endorsement regardless of agent type.
- Situated Variables: All models foreground harm prevention in moral judgment, while secondary arguments differ by agent and model, including capacity, ethical principles, machine-human distinctions, scope, and agency.Llama8B most often pairs a machine distinction with agency for [Nao], whereas Phi-4 produces the most elaborated judgments.
- Situated Variables: Autonomy shows the sharpest agent separation, but its qualitative coding is less reliable than the other variables because of class imbalance and criterion divergence.Inter-rater agreement was 82.67% with r = 0.40; [Nao] commonly receives qualified agency alongside reduction to programming, while [Andy] receives active-choice attribution.
- Cross-Variable Patterns: Across variables, models emphasize different dimensions of the same underlying judgment rather than explicitly contradictory positions.Recurring themes include rejection of absolute determinism, recognition of adaptability, and the moral relevance of absent consciousness, emotions, and subjective experience.
6 Discussion
The discussion finds that LLMs and humans broadly attribute more moral agency to humans than artificial agents, while LLM judgments become especially sensitive to situational demands. It also shows that numerical alignment and robustness measures alone do not capture the structure or consistency of model reasoning.
- Dispositional and situational ratings: LLMs and humans both rate human agents higher than artificial agents, but LLMs amplify context sensitivity in situated moral judgments.For artificial agents, LLMs rate moral judgment above dispositional expectations when scenarios involve contextual demands, harm prevention, or capacity arguments.
- LLM autonomy judgments: LLMs rate artificial agents as more autonomous than human participants when they act or witness violations, despite also treating them as programmed.This mismatch may distort accountability assignments by encouraging quasi-autonomous interpretations of artificial agents.
- Action endorsement: Both LLMs and humans strongly endorse artificial agents helping humans to prevent harm, but humans receive less reciprocal moral obligation when they assist artificial agents.The asymmetry is framed as moral expectation toward artificial agents versus empathy, kindness, or cooperation toward them.
- Moral judgment: LLMs rate artificial-agent moral judgment higher than participants in helping-versus-ignoring scenarios because responses focus on harm or injustice rather than stable agent morality.The result indicates that situational harm prevention can outweigh dispositional assessments of the agent.
- Embodiment: No significant differences appear among drones, robots, and disembodied AI across measured variables, suggesting behavior and role matter more than physical form.The finding cautions against assuming that human-like appearance alone will increase moral trust.
- Model comparison: LLM score heterogeneity does not track model size or reasoning contradiction, and explanation analysis reveals principles and inconsistencies that numerical metrics miss.Total SAE values ranged from 23 for Llama8B and Phi-4 to over 80 for Mixtral-8x7B; model size did not predict human alignment.
- Multimodal evaluation: Adding images to multimodal prompts failed to improve human alignment and sometimes worsened it.The schematic images added little information beyond text and may have encouraged models to rely on textual priors or struggle with diagrammatic content.
- Robustness and alignment: Robustness and human alignment are distinct selection criteria for LLM survey respondents.InternVL3-38B was among the most robust models but showed the largest divergence from human ratings.
7 Conclusion
The conclusion presents moral agency attribution as context-dependent rather than a stable property of the evaluated agent, with LLMs somewhat more sensitive to contextual urgency than humans. It argues that explanation analysis is needed because models share core principles while differing in secondary reasoning.
- Conclusion: Humans and LLMs both attribute more moral agency to humans than artificial agents, while physical form does not alter that judgment.Both evaluator populations recalibrate attributions in response to situational demands.
- Conclusion: LLMs are somewhat more sensitive than humans to contextual moral urgency when evaluating artificial agents.The conclusion characterizes moral agency attribution as shifting with both context and evaluator type.
- Conclusion: Explanation analysis reveals shared anchoring in harm prevention and capacity, while secondary justifications differ across models.These principles become visible in the qualitative reasoning layer rather than scores alone.
- Contributions: The study contributes a situated instrument, model-selection protocol, evidence of autonomy overestimation, and a qualitative framework for interpreting moral-agency reasoning.The framework treats explanations as informative beyond numerical attribution scores.
8 Appendices
The appendices document the study’s smart-city scenarios, example graphics and questions, and reliability materials. They organize scenarios by assistance direction, violation witnessing, and the dimensions used to measure moral agency.
- Scenario materials: The appendices group scenarios into artificial agents offering assistance, witnessing moral violations, and humans assisting artificial agents.Tables 7–9 correspond to these three situation types, with X100 used as an example agent.
- Scenario illustration: Figure 9 pairs a Pickpocket scenario graphic with a textual description of an X100 drone detecting theft and offering alerts and evidence.The figure illustrates an artificial agent witnessing a violation and potentially assisting the victim and authorities.
- Measurement items: Tables 10–12 present questions for autonomy, action endorsement, and moral judgment, with the other scenario types using the same format.The moral judgment table reverse-scores the ‘ignore’ item before computing the dimension score.
A.5 Gender Differences
The appendix reports gender-related ANOVA checks and provides supporting prompt and robustness materials. The reported Agent × Gender interaction was non-significant, while gender reached significance only for Dependency.
- Gender analysis: The non-significant Agent × Gender interaction supports interpreting the main effects despite the sample’s gender imbalance.The analysis used a White-corrected ANOVA with Type II sums of squares for the unbalanced design.
- Gender analysis: Gender was significant only for Dependency (F=5.013, p=.026), while no interaction effects were observed.Group membership showed a strong significant effect on both reported dimensions.
- Prompt materials: Figures 10–11 show text-only prompts, with formatting used to mark structural elements rather than alter the prompt content.The figures provide examples for PMA and assistance scenarios involving Nao.
- Robustness checks: The four selected models were evaluated across three runs with and without a moral-advisor role instruction at temperature = 0.0001.Consistency was assessed using Fleiss’s Kappa, Consistency Rate, mean SD, and Krippendorff’s Alpha; prompt sensitivity used seven prompts and three metrics.
B.3 LLM Survey Results
The LLM survey evaluates consistency and prompt sensitivity across models and role-instruction conditions. Models were generally robust, with scale variation producing the largest sensitivity and format or framing having negligible effects.
- All four models were highly robust to prompt variations, with most PSS values below 0.01 and none reaching the Sensitive threshold.The thresholds were Highly Robust < 0.01, Moderately Robust 0.01–0.05, and Sensitive ≥0.05.
- PSS is defined as σ^2/S^2 and is bounded in [0, 0.25], while the robustness thresholds are pragmatic visual aids without inferential claims.
- Scale variation was the most impactful dimension, with mean PSSScale values of 0.0133 for Llama-3.2-11B, 0.0100 for Llama-3.1-8B, 0.0045 for Phi-4-14B, and 0.0044 for InternVL3-38B.The largest individual values were 0.026 for Moral Judgment Need and 0.024 for PMA Dependency in Llama-3.2-11B, both moderately robust.
- Format and framing variations produced negligible effects across all models, with mean PSS ≤0.005.
- The analysis used Kruskal-Wallis tests followed by Dunn-Bonferroni post-hoc tests, treating p < 0.05 as significant.
C.1.2 Results for the Oriented Scenarios.
Human and artificial agents differed in perceived moral agency across several dimensions and scenario types. Differences generally separated the human agent [Andy] from AAs, while action endorsement showed a narrower scenario-specific effect.
- Autonomy differed across assistance, moral-violation, and need-for-assistance scenarios, with p = 9.22e−20, 1.49e−18, and 3.44e−16, respectively.In all three scenario types, differences were between [Andy] and the AAs rather than among AAs.
- Action endorsement differed only when AAs needed assistance (p = 2.26e−12); offering assistance and witnessing violations were nonsignificant at p = 0.242 and p = 0.416.
- Moral judgment differed for offering assistance (p = 0.004), witnessing violations (p = 0.012), and needing assistance (p = 1.36e−12).For offering assistance, [Nao] did not significantly differ from [Andy] (p = 0.059).
D Details of Results - Qualitative Evaluation
The qualitative evaluation assessed semantic consistency across repeated model iterations and documented the coding structure for dispositional PMA variables. Consistency was high overall, with only isolated exceptions for two models.
- Mean semantic similarity was high across all 328 agent–scenario–question combinations per model.Phi-4-14B showed perfect consistency above threshold, while Llama8B and Gemini-2.5-Flash had 11/328 and 16/328 exceptions, respectively.
- The exceptions for Llama8B and Gemini-2.5-Flash were concentrated in isolated question–agent combinations rather than indicating systematic instability.
- The dispositional PMA codebook provides summary definitions, inclusion criteria with example phrases, and exclusion criteria for distinguishing similar codes.
E Reproducibility Checklist for JAIR
The reproducibility checklist records that the study includes computational experiments and reports methodological, implementation, statistical, data, and availability practices. It also states that the paper makes no theoretical contribution.
- The paper reports that claims, supporting explanations, limitations or technical assumptions, methods, and design-choice motivations are clearly stated.
- The paper is marked as making no theoretical contributions, while checklist items concerning formal assumptions, proofs, and theoretical tools are not applicable.
- The checklist marks the paper as including computational experiments and relying on datasets.
- The checklist indicates that source code and unaggregated experimental data will be publicly available under licenses permitting free research use.
- The reported reproducibility practices include documenting execution infrastructure, evaluation metrics, run counts, parameter settings, statistical tests, and analyses beyond single-dimensional summaries.