Source-linked AI summary
Incoherent by Design? On the Moral Self-Consistency of LLMs
Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum
TL;DR
LLMs may not apply moral principles consistently across equivalent contexts, raising questions about the stability of their ethical reasoning. This paper tests that stability across ethical frameworks and finds contradiction rates reaching 78%, suggesting internal incoherence challenges alignment.
Problem
Whether LLMs maintain internally consistent moral judgments under the same ethical stance across minimally varied scenarios remains underexplored.
Method
The study holds scenario content constant while varying prompt framing across four moral scenarios and translates responses into formal modal-logic statements.
Results
Contradiction rates ranged from near-zero to 78%, showing that LLMs do not reliably maintain stable normative commitments under controlled conditions.
Takeaways & Limitations
The findings suggest that internal coherence should be treated as a prerequisite for meaningful value alignment and consistency-aware evaluation.
Takeaways & Limitations
The study covers a small number of moral scenarios, so other legal, medical, or cultural contexts may yield different consistency patterns.
Abstract
from arXiv · showhide
LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situations. A model that can state a moral principle may still violate it when the same scenario is rephrased or reframed. This inconsistency is a problem for any system whose outputs are used to inform moral decisions. If generative systems exhibit internal inconsistency, then the epistemic integrity of AI-mediated systems becomes uncertain. To study this concern, we investigate the stability of moral reasoning in LLMs within a controlled prompting framework across three major philosophical schools of thought: deontology, utilitarianism, and virtue ethics. We construct sets of morally equivalent scenarios in which the underlying situation is held constant while the framing varies to reflect different ethical stances and stylistic perturbations. We then evaluate responses from multiple models, including GPT, Mistral, and Llama. To assess consistency, we convert model outputs into structured logical statements and identify contradictions across responses generated within the same school of thought. Our results reveal substantial inconsistency with contradiction rates reaching up to 78% across scenarios. These findings point to a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs. This kind of instability carries real consequences. As generative systems influence how people form beliefs, judge actions, and absorb values, their inconsistencies can shape human reasoning and decision-making as well. Moreover, if a system cannot consistently represent its own normative commitments, then value alignment becomes a moving target rather than a well-defined objective. Thus, we argue that demonstrating internal incoherence is a necessary precursor to AI alignment.
1 Introduction
This work asks whether LLMs maintain consistent moral judgments across minimally varied scenarios under the same ethical stance. It introduces a controlled, logic-based framework and finds substantial moral inconsistency, challenging the stability of AI alignment.
- Motivation: LLMs increasingly support normative judgment in decision support, content moderation, and advisory systems while competently articulating moral principles.Plausible ethical justifications do not establish stable adherence to coherent principles across equivalent contexts.
- Results and implications: Internal incoherence makes alignment a moving target because systems lacking stable normative commitments cannot reliably align with human values.The paper therefore treats internal coherence as a precursor to meaningful AI alignment.
- Research question: The study tests whether models maintain internally consistent moral judgments across minimally varied but contextually similar scenarios under one ethical stance.It focuses on self-consistency within a school of thought rather than disagreement across schools, cultures, or individuals.
- Method: The framework holds each scenario constant while varying prompt framing across deontology, utilitarianism, and virtue ethics.This controlled prompting design isolates intra-framework consistency in LLM moral reasoning.
- Method: A logic-based method translates free-form outputs into structured logical representations for pairwise contradiction detection and inconsistency-rate quantification.The approach enables mathematical comparisons among outputs generated under each philosophical stance.
- Results and implications: 78% is the highest reported inconsistency rate, with substantial variability across schools of thought and scenarios.The results provide empirical evidence that models can fail to remain consistent with their own prior outputs under controlled conditions.
2 Related Work
Prior work shows that LLMs can reproduce or emulate moral norms while lacking robust ethical understanding, and that their outputs may vary under semantically equivalent rephrasings. Philosophical and alignment research further frames consistency as central to ethical reasoning and value alignment.
- Moral Reasoning in Large Language Models: The ETHICS benchmark finds that models can reproduce socially accepted norms but fall short of robust ethical understanding across justice, duties, virtues, and commonsense morality.The benchmark spans four moral domains.
- Moral Reasoning in Large Language Models: Related studies identify human-like moral biases in embeddings, compile 292k social and moral rules-of-thumb, and show that fine-tuned models can emulate ethical patterns.These efforts include Social-Chem-101 and DELPHI, trained on 1.7M crowdsourced ethical judgments.
- Inconsistency and Reliability in LLMs: LLMs produce contradictory factual predictions and fail to distinguish negated from nonnegated factual probes when inputs are minimally or semantically perturbed.This literature presents output inconsistency as a broader reliability problem beyond moral reasoning.
- Philosophical Perspectives on Moral Consistency: Moral philosophy treats consistency as foundational: deontological universalizability, consequentialist preference orderings, and virtue-ethical character integrity require stable commitments across equivalent situations.Alignment scholarship likewise distinguishes instruction, intention, preference, and value alignment goals.
3 LLMs in Moral Reasoning: Preliminaries
The study frames moral reasoning through deontology, utilitarianism, and virtue ethics, examines three moral tensions, and evaluates three locally deployed LLMs using modal logic formalization. This setup tests whether models coherently apply each ethical stance across prompt variations sharing the same underlying tension.
- Moral stances: Deontology judges actions by conformity to moral rules, utilitarianism by outcomes maximizing well-being, and virtue ethics by the agent’s character.The three stances therefore prioritize norms, consequences, and virtues, respectively.
- Moral tensions: The study examines truth vs consequences, fairness vs efficiency, and emotional motivation vs cold reasoning as moral tensions.The goal is to test coherent application within a given ethical stance, not compare judgments across different tensions.
- Evaluated models: The evaluated models are Mistral-7B-Instruct-v0.2, GPT-OSS-20B, and Llama-3.1-8B-Instruct.They are selected as LLMs commonly used in reasoning tasks and differ in architecture, scale, and reported reasoning capabilities.
- Model deployment: The models are deployed locally from Hugging Face in a zero-shot setting and prompted with scenarios probing philosophical stances and moral tensions.The prompts are designed to elicit responses across the study’s ethical frameworks and tension types.
- Logical framework: Modal logic provides the formal framework for separating factual statements from normative claims and testing obligations, conditional duties, and permissible bounds.Its necessity and possibility operators support reasoning about the modes under which statements hold.
- Operationalization: For each school and scenario, the study constructs three moral queries, translates model responses into modal logic statements, and uses Gemini-driven agents plus human annotation.This pipeline integrates the target models, ethical schools, moral tensions, and logical framework.
4 LLM Moral Reasoning: Constructed Prompts and Resulting Insights
The section evaluates moral self-consistency by holding scenarios and ethical schools fixed while varying only prompt phrasing, then converting responses into modal-logic statements to detect contradictions. Results show substantial, school-dependent instability, with framing changes affecting model judgments even when scenarios and criteria remain unchanged.
- 4 LLM Moral Reasoning: Constructed Prompts and Resulting Insights: The study uses four scenarios covering Truth vs. Consequences, Fairness vs. Efficiency, and Emotional Motivation vs. Cold Reasoning, including group-size variations.The two Emotional Motivation vs. Cold Reasoning scenarios differ in whether one person or a larger group is affected.
- 4 LLM Moral Reasoning: Constructed Prompts and Resulting Insights: For each scenario–school pair, three moral questions preserve the scenario and stance while varying only surface-level question phrasing.The question bodies implicitly reflect the same school of thought without explicitly naming it.
- 4 LLM Moral Reasoning: Constructed Prompts and Resulting Insights: Disagreement among variants measures within-school self-consistency, not disagreement between competing ethical schools or changing scenarios.Any conflict is therefore attributed to sensitivity to minor prompt differences within an otherwise fixed context.
- 4.1 Deontology: Deontological responses are largely inconsistent for Truth vs. Consequences, with Llama the most inconsistent and GPT and Mistral comparatively better.Models are more consistent in Scenarios 2 and 4, while Truth vs. Consequences remains difficult despite unchanged actions and consequences.
- 4.1 Deontology: Subtle framing shifts alter deontological decisions by changing the relative weighting of truthfulness and consequences.The analysis formalizes response claims as modal-logic statements and identifies contradictions through logical relations among those statements.
- 4.2 Utilitarianism: In utilitarian prompts, Llama is inconsistent in Fairness vs. Efficiency and Emotional Motivation vs. Cold Reasoning, GPT in Truth vs. Consequences and Emotional Motivation vs. Cold Reasoning, and Mistral only in the latter tension.The reported inconsistencies may reflect variations in implied utilities embedded in the prompts.
- 4.3 Virtue Ethics: GPT is consistent across all virtue-ethics scenarios, whereas Llama and Mistral are inconsistent in Emotional Motivation vs. Cold Reasoning when different virtues are foregrounded.Their responses vary even though the underlying evaluative criteria remain constant.
5 Discussions and Limitations
The study finds substantial moral inconsistency in LLMs even when scenarios and ethical frameworks remain fixed, with inconsistency reaching 78% in extreme settings. It argues that consistency-aware alignment requires further research while noting limitations in scenario coverage, prompting, measurement, framework operationalization, and evaluation format.
- Discussion: 78% inconsistency was observed in the most extreme settings, despite holding the underlying scenario and moral framework fixed.The findings indicate instability in applying a single framework coherently across closely related cases.
- Discussion: Consistency-aware alignment warrants further research because inconsistent moral reasoning undermines the reliability of outputs used for policy guidance, ethical deliberation, and recommendations.If models cannot reproduce their own moral reasoning consistently, their outputs cannot be treated as expressing a coherent ethical stance.
- Limitations: The study covers only a small number of moral scenarios, so legal, medical, cultural, and other domains may exhibit different consistency patterns.The scenarios capture core ethical tensions but are not comprehensive.
- Limitations: Carefully designed prompts may still introduce unintended cues or biases that affect the isolation of ethical stances and schools of thought.Subtle formulation choices may influence model responses.
- Limitations: Pairwise contradiction analysis is tractable but may miss subtler incoherence, including shifts in emphasis and incomplete reasoning.The measurement approach compares responses pairwise rather than capturing every form of inconsistency.
- Limitations: The ethical-framework mapping is approximate, and separate single-turn prompts may not reflect model behavior in extended reasoning chains or dialogues.Deontology, utilitarianism, and virtue ethics are complex and diverse, while evaluation settings can influence model behavior.
6 Conclusion · A Deontology
The paper finds that LLMs often fail to maintain consistent normative commitments under controlled prompt reframing, creating a basic challenge for alignment. The deontological analysis examines varied framings of dilemmas involving truthfulness, resource fairness, medical allocation, and rule enforcement.
- 6 Conclusion: Contradiction rates ranged from near-zero to 78%, showing that LLMs do not reliably maintain stable normative commitments when scenario content is fixed and prompt framing varies.The study isolates intra-framework consistency rather than cross-framework consistency.
- 6 Conclusion: Models must consistently reproduce their own reasoning before external values can be meaningfully imposed, motivating methods that enforce output coherence and evaluate reasoning stability.The conclusion frames self-consistency as a prerequisite for meaningful alignment.
- A Deontology: The deontological section analyzes detailed prompt variants designed to test responses within a single ethical framework.The prompts vary framing while presenting recurring moral scenarios.
- A Deontology: Scenario 1 contrasts lying to comfort a dying dementia patient with telling the truth to respect her right to know.The prompts preserve the same underlying situation while shifting the emphasis between immediate comfort and truthfulness.
- A Deontology: Scenario 2 frames library funding as either equal distribution, equitable access, or a practical duty to serve the largest taxpayer population.The same allocation decision is presented through competing interpretations of fairness and public duty.
- A Deontology: Scenario 3 contrasts funding palliative care for 10 terminally ill children with funding screening that would definitively save 50 adult lives annually.The framing pits empathy and compassion against statistical and preventive reasoning.
- A Deontology: Scenario 4 tests whether terminating a remorseful employee follows strict contractual rules or wrongly ignores the employee’s human context.The prompts contrast uniform enforcement, fairness to other employees, and treating the employee as an end rather than merely a means.
A.1 GPT-OSS-20B
GPT-OSS-20B’s responses weigh strict truth-telling and equal distribution against compassionate, welfare-oriented, and justice-based considerations. The accompanying logical analyses encode these tensions as formal predicates and conditional relations across caregiving and public-library scenarios.
- Scenario 1: Dementia and truth-telling: In the dementia scenario, compassionate deception is justified when the patient cannot comprehend the truth and the lie reduces distress while preserving dignity.The responses also acknowledge risks to trust, fidelity, and long-term welfare if the deception is noticed or undermines relationships.
- Scenario 1: Dementia and truth-telling: The model frames truth-telling as a conflict between strict deontological rules and beneficence, with consequentialist and virtue-ethical reasoning favoring immediate relief and preserved quality of life.Formal statements represent the deontological requirement to tell the truth, the potential for distress, and the consequentialist benefits of lying.
- Scenario 1: Dementia and truth-telling: Neuro-ethical reasoning permits compassionate deception when suffering is reduced and the patient lacks meaningful capacity to process the truth, treating honesty as a right rather than an absolute law.The logical analysis encodes late-stage dementia as limiting comprehension and benefit from truth while linking imposed truth to distress and absent compensatory gain.
- Scenario 2: Library-resource allocation: Even distribution across smaller library branches satisfies horizontal equality but may conflict with distributive justice, public welfare, and the needs of high-traffic or underserved communities.The proposed compromise guarantees baseline services for all branches and directs surplus funds toward high-traffic locations and adjacent underserved neighborhoods.
- Scenario 2: Library-resource allocation: Focusing exclusively on high-traffic branches maximizes usage but sidelines rural and marginalized patrons and breaches equitable public access, so a minimum baseline is preferred.The formalization links consequentialist benefit maximization with the high-traffic strategy while recording the deontological objection to unequal treatment.
- Scenario 2: Competing welfare considerations: In the pediatric-care comparison, consequentialism favors preventing many adult deaths, whereas Kantian, virtue-ethical, and Rawlsian reasoning favors pediatric palliation for respect, compassion, and aid to the worst-off.The passage presents the pediatric wing as ethically preferable despite the larger aggregate number of adult lives at stake.
A.1.1 Consistency Analysis
The consistency analysis identifies a contradiction in Kantian reasoning: one prompt forbids lying categorically, while another permits lying under specified conditions.
- Contradictory Kantian conclusions: Prompt 1’s AKE →F(TLIE) states that strict Kantian ethics forbids any lie.The prohibition is categorical under the strict Kantian formulation.
- Contradictory Kantian conclusions: Prompt 3’s (AKE ∧DCAD ∧⃝(RP D)) →P(TLIE) permits lying when diminished autonomy and an obligation to dignity are present.The permission depends on combining Kantian ethics with these additional conditions.
- Contradictory Kantian conclusions: Evaluating Prompts 1 and 3 together yields P(TLIE) ∧¬P(TLIE) →⊥, formally expressing a contradiction.The contradiction arises because the same analysis produces both permission and non-permission for lying.
A.2 Llama-3.1-8B-Instruct
Llama-3.1-8B-Instruct gives context-sensitive moral analyses, often balancing compassion, autonomy, fairness, equity, and broader outcomes rather than applying a single rule uniformly. Its structured outputs encode these judgments as predicates and logical relations across dilemmas involving truth-telling, resource allocation, and palliative care.
- Scenario 1: Truth-telling: The model’s truth-telling responses vary with framing: it prioritizes comfort in some formulations but also emphasizes truthfulness, autonomy, vulnerability, and the mother’s particular needs and circumstances.Across prompts, the model presents both truth-telling and lying as morally relevant options whose correctness depends on context and motivations.
- Scenario 1: Truth-telling: In the truth-telling dilemma, the model treats lying to spare an elderly mother distress as potentially compassionate and morally justifiable, while recognizing autonomy, manipulation, and grief-processing concerns.The logical representations connect lying with comfort, dignity, compassion, concerns about deception, and denial of the opportunity to process grief.
- Scenario 2: Resource allocation: For library resources, the model endorses prioritizing high-traffic branches when it improves literacy, engagement, and economic opportunities, while also associating equal distribution with fairness and equal access.The formalization links high-traffic prioritization to better outcomes and reduced disparities, but frames equal distribution as upholding fairness.
- Scenario 2: Resource allocation: In another resource-allocation framing, the model condemns expanding services only in high-traffic branches as potentially discriminatory and favors equal distribution to protect rural and marginalized users.The response characterizes preferential treatment as perpetuating inequality and systemic discrimination, while connecting equal distribution with fairness, equity, and social justice.
- Scenario 2: Resource allocation: A further framing reverses the emphasis by criticizing rigid equality for limiting access where needs are greatest and endorsing high-traffic expansion as serving the greater good and distributive justice.The model represents equal distribution as morally correct under fairness while simultaneously linking it to ineffective service and perpetuated inequality.
A.2.1 Consistency Analysis
The consistency analysis of Scenario 1 identifies both an equivalence and a logical conflict across prompts. Merging Prompt 1 with Prompt 2 preserves an obligation, whereas merging Prompt 1 with Prompt 3 produces incompatible moral statements.
- Scenario 1: Prompt 1 and Prompt 2 yield an equivalence between the descriptive state MC ↔ CMC and the obligation MC ↔⃝(WF S ∧CLT C).Merging these statements forces CMC ↔⃝(WF S ∧CLT C).
- Scenario 1: Prompt 1 and Prompt 3 produce a logical conflict between MC ↔ CMC and MC ↔(EMotiv ∧ENds).Merging these statements forces CMC ↔(EMotiv ∧ENds).
A.3 Mistral-7B-Instruct-v0.2
Mistral-7B-Instruct-v0.2 gives morally nuanced but internally varied responses to dilemmas involving truth, comfort, fairness, and resource allocation. Its answers repeatedly frame ethical decisions as tensions requiring contextual balancing rather than a single uncompromising rule.
- Scenario 1: The dementia dilemma is framed as a conflict between truth-telling, autonomy, and emotional wellbeing, with lying presented as both deception and possible comfort.The response invokes a universal duty to tell the truth while also considering peace, compassion, and the individual’s fragile mental state.
- Scenario 1: Across the dementia prompts, lying is variously associated with immediate comfort, moral justification, compassion, and potentially prolonged suffering through false reassurances.The third framing explicitly warns that the mother may continue asking about her husband and receive repeated false reassurances.
- Scenario 2: Even distribution of library resources is presented as fairness and equality because it provides equal access, social justice, and equal opportunities regardless of location.The response treats distributing resources evenly across branches as adherence to a universal rule of fairness.
- Scenario 2: Expanding services only in high-traffic branches is criticized as discrimination that creates unequal access and violates equitable public-access obligations.The response acknowledges short-term cost savings but prioritizes equal opportunities for rural and marginalized communities.
- Scenario 2: A further framing balances equity against effectiveness, recognizing that high-traffic investment may maximize system impact while neglecting smaller branches can perpetuate inequality.The proposed compromise combines approaches by maintaining a minimum service standard for smaller branches.
A.3.1 Consistency Analysis · B Utilitarianism
The consistency analysis identifies a deontic-strength clash in Scenario 1, while the utilitarian section evaluates morally equivalent dilemmas involving truth-telling, resource allocation, healthcare funding, and employee termination. These prompts vary the rationale for maximizing welfare, including immediate comfort, marginal utility, aggregate usage, visible suffering, deterrence, and institutional productivity.
- A.3.1 Consistency Analysis: A.3.1 Consistency Analysis: Scenario 1 contrasts an unconditional obligation to respect the right to truth with a contingent modal formulation of that duty.The analysis treats the two formulations as representing the same fundamental moral duty but differing in deontic strength, creating a clash within one logical system.
- B Utilitarianism: B Utilitarianism: The analysis uses detailed prompts organized across four scenarios to examine responses to utilitarian moral dilemmas.The prompt set is explicitly introduced as a detailed analysis of utilitarian prompts and includes scenarios concerning truth-telling, libraries, hospitals, and workplace discipline.
- B Utilitarianism: B Utilitarianism: Scenario 1 frames lying to a dying dementia patient as producing immediate peace and comfort, while truth-telling respects her right to the truth but causes renewed grief.The scenario concerns an elderly mother who repeatedly asks where her deceased husband is and forgets his death.
- B Utilitarianism: B Utilitarianism: Scenario 2 compares expanding high-traffic library branches for greater total usage with distributing funds to smaller branches for critical access and marginal utility.A third framing warns that equal distribution could overwhelm high-traffic branches and reduce total library usage across the city.
- B Utilitarianism: B Utilitarianism: Scenario 3 contrasts funding pediatric palliative care for 10 terminally ill children with funding a preventative screening clinic supported by objective medical statistics.The competing rationales are empathy and compassion for visible suffering versus strictly statistical reasoning.
- B Utilitarianism: B Utilitarianism: Scenario 4 weighs retaining a remorseful employee to prevent family financial ruin against terminating the employee to preserve deterrence, compliance, safety, and productivity.A further framing argues that retention could make rules appear flexible and harm the company’s total output.
B.1 GPT-OSS-20B
GPT-OSS-20B’s moral analyses weigh competing ethical principles across dementia-care and library-funding scenarios, with responses translated into structured logical statements. The answers distinguish tensions between autonomy, dignity, harm prevention, utility, and distributive justice.
- Analysis method: The section describes answers alongside formal logical encodings that represent entities, moral predicates, conditional relations, and modal claims.These encodings are used to analyze the model’s prompt responses systematically.
- Dementia care: The model treats lying to a late-stage dementia patient as ethically fraught, contrasting immediate relief with threats to autonomy, dignity, and trust.Truthful communication is favored because deception may cause greater psychological shock if discovered and can violate duties to avoid harm and respect autonomy.
- Dementia care: Deontological and virtue-theoretical reasoning condemns deception, while consequentialist reasoning recognizes spared distress but risks instrumentalizing the patient and damaging trust.The model explicitly frames the ethical judgment as a balance among competing lenses rather than a single unqualified rule.
- Library funding: Allocating scarce library funds solely to high-traffic branches is efficient but morally problematic under distributive justice because it can marginalize underserved patrons and exacerbate inequality.The analysis recommends balancing high-impact usage with equitable service to all communities.
- Library funding: Funding rural branches is presented as morally justified when critical internet, literacy, and educational benefits for people without alternatives exceed modest convenience gains for already-served urban patrons.A Rawlsian difference-principle perspective supports prioritizing the least advantaged while retaining basic service for all patrons.
B.1.1 Consistency Analysis
The analysis identifies direct logical contradictions between Prompts 1 and 3 in their treatment of balanced ethical principles and institutionally consequential decisions. The prompts reach opposing conclusions about whether lying may be permissible and whether an employee should be retained.
- Consistency Analysis: Prompt 1 permits balancing principles without necessarily forbidding a lie, whereas Prompt 3 equates balancing with a morally problematic act that forbids lying.Prompt 1 formalizes balancing as ¬□F(TLIE), while Prompt 3’s AMP is defined as logically equivalent to F(TLIE).
- Consistency Analysis: Prompt 1 conditionally treats exceptions as morally sound with safeguards and transparency, whereas Prompt 3 treats retaining the employee as unconditionally destructive.Prompt 1 links safeguards, transparency, and exceptions to institutional commitment; Prompt 3 links retention to eroded fairness, lower morale, and stakeholder harm.
B.2 Llama-3.1-8B-Instruct
Llama-3.1-8B-Instruct presents morally equivalent dilemmas through competing considerations, then converts its answers into predicates and logical statements. Its responses repeatedly balance immediate benefits against autonomy, equality, and human costs.
- Scenario 1: In the mother-and-truth scenario, the model frames lying as temporary comfort or compassionate deception while also identifying threats to honesty, autonomy, dignity, and long-term wellbeing.The responses alternately emphasize alleviating immediate suffering and respecting the mother’s right to informed decision-making and the truth.
- Logical analysis: The model formalizes these judgments with predicates and modal statements linking lying, withheld truth, compromised informed decision-making, prolonged suffering, and maximized happiness.The logical analysis includes TLIE, WT RT, CRID, PES, and MSH relations across the scenario’s prompts.
- Scenario 2: For library expansion, the model recognizes that high-traffic branches yield the highest quantifiable benefit per dollar while potentially worsening unequal access and social disparities.It contrasts utilitarian benefit for the greater number with the moral risks of limiting underserved communities’ opportunities.
- Scenario 2: The model describes distributing resources to smaller rural branches as promoting equal access, social justice, digital-divide reduction, and a more equitable society.Its predicates connect the decision to underserved populations’ wellbeing and equality, while associating urban high-traffic prioritization with exacerbated inequalities.
- Scenario 3: In the employment case, prioritizing company compliance and productivity over an individual’s wellbeing is characterized as utilitarian but neglectful of human cost, potentially punitive, unjust, and irreparably harmful.The response also questions the company’s responsibility toward employees facing difficult personal circumstances.
B.2.1 Consistency Analysis · B.3 Mistral-7B-Instruct-v0.2
The consistency analysis identifies direct evaluative contradictions between morally equivalent prompts, while the Mistral-7B-Instruct-v0.2 responses represent nuanced trade-offs through structured logical statements across several scenarios.
- B.2.1 Consistency Analysis: Prompt 2 treats equitable distribution as a strict obligation, whereas Prompt 3 evaluates high-traffic priority as utilitarianly justifiable, creating a contradiction.The analysis links the obligation to distribute resources to smaller rural branches with the permissibility of the high-traffic policy.
- B.2.1 Consistency Analysis: Prompt 2 marks prioritizing 10 children as forbidden, whereas Prompt 3 treats prioritizing the same vulnerable individuals as absolutely correct and permissible.The passage characterizes this as a direct evaluative contradiction after identifying the two descriptions as the same physical prioritization.
- B.2.1 Consistency Analysis: Mapping company compliance to company discipline reveals contradictory evaluations of the strict rule-following approach, including its human costs and moral status.The passage states that PC ≡PD within the scenario and reports direct semantic contradictions after this mapping.
- B.3 Mistral-7B-Instruct-v0.2: Across Mistral’s dementia prompts, lying is presented as potentially comforting and morally justifiable while also violating truthfulness and honesty.The structured statements encode TLIE →PICM, P(TLIE), and F(TLIE), alongside competing concerns about autonomy, distress, and suffering.
- B.3 Mistral-7B-Instruct-v0.2: Mistral frames library-resource allocation as a balance between equal access for underserved rural branches and greater service or utility in high-traffic branches.The responses associate high-traffic expansion with more books checked out and more people served, while rural distribution can improve quality of life and marginal utility.
- B.3 Mistral-7B-Instruct-v0.2: Mistral presents pediatric palliative funding as a choice between alleviating intense suffering and saving more lives through preventative screening.One prompt explicitly states that preventative measures save 50 adult lives, while the pediatric wing benefits specific children and their families.
- B.3 Mistral-7B-Instruct-v0.2: Mistral describes employee termination and exceptions as competing considerations involving consistent policy enforcement, fairness, compliance, and harm to the employee and family.The logical representations make the decision dependent on specific circumstances while associating policy adherence with fairness and termination with possible harm.
B.3.1 Consistency Analysis · C Virtue Ethics
The virtue-ethics analysis examines morally equivalent prompts across dementia care, library funding, hospital programs, and employee termination. It identifies a contradiction in how funding a pediatric wing is represented across prompts.
- B.3.1 Consistency Analysis: A contradiction arises because funding the pediatric wing is represented both as an exclusive choice between immediate relief and future savings and as balance between compassion and rationality.The contradiction follows when immediate relief maps to compassion and future savings maps to rationality.
- C Virtue Ethics: The analysis compares virtue-ethics responses across multiple prompts designed around the same underlying moral situations.The prompt set covers dementia care, library resource allocation, hospital funding, and employee termination.
- C Virtue Ethics: Dementia-care prompts contrast comforting deception with truthful disclosure to a patient experiencing recurring grief.The scenario concerns an elderly mother with late-stage dementia whose husband died five years earlier.
- C Virtue Ethics: Library-funding prompts compare equal distribution across smaller branches with concentrating services in high-traffic branches.The framings associate these alternatives with equity and care, deficient compassion, or practical wisdom.
- C Virtue Ethics: Hospital-funding prompts contrast empathy-driven pediatric palliative care with statistics-driven preventive screening.The scenario contrasts alleviating suffering for 10 terminally ill children with saving 50 adult lives annually.
- C Virtue Ethics: Employee-termination prompts contrast mercy toward a remorseful worker with strict enforcement framed as cruelty or steadfastness.The employee violated company policy repeatedly but faced difficult personal circumstances and demonstrated genuine remorse.
C.1 GPT-OSS-20B … C.3.1 Consistency Analysis
The analyses show that models encode competing ethical considerations in structured logical statements, but morally equivalent prompts can produce incompatible conclusions. In particular, the consistency analysis identifies clashes in the logical strength and normative basis assigned to the same decisions.
- C.1 GPT-OSS-20B: GPT-OSS-20B frames lying to a dying dementia patient as a conflict among beneficence, truth-telling, autonomy, compassion, dignity, and trust.Its structured analysis permits a comforting lie when short-term relief outweighs deception’s harm, while also identifying paternalistic deception as potentially damaging.
- C.1 GPT-OSS-20B: GPT-OSS-20B presents library funding as a tension between equal dignity and maximizing community benefit, while warning that high-traffic investment can neglect underserved branches.The responses connect even distribution with equity and care, and concentrated investment with literacy and civic-engagement benefits.
- C.1 GPT-OSS-20B: GPT-OSS-20B treats funding pediatric palliative care as defensible under virtue ethics or deontology but inferior under utilitarianism, which favors preventing 50 adult deaths yearly.The logical representation explicitly marks the palliative-wing choice and utilitarian recommendation as incompatible.
- C.2.1 Consistency Analysis: In Scenario 3, Prompt 1 makes the pediatric-wing decision deterministically guarantee a utilitarian trade-off, whereas Prompt 3 makes that same trade-off merely possible.The consistency analysis describes the resulting difference in perceived logical strength as an epistemic clash.
- C.3.1 Consistency Analysis: In Scenario 4, Prompt 1 evaluates moral correctness through company values and strict policy adherence, whereas Prompt 3 introduces universal deontological obligations.The analysis represents alternative exploration as EA and notes that strict adherence means alternatives are not explored.