Source-linked AI summary
Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers
Taenyun Kim, Edyta Bogucka, Daniele Quercia
TL;DR
Voting-based moral AI alignment may appear to recover public values, but developers shape which features, voters, and questions enter the process. This paper empirically examines these choices across three contexts and finds that elicited preferences vary with each upstream configuration.
Problem
Whether voting-based moral alignment neutrally recovers public values remains underexamined because developers choose the features, participants, and wording presented for evaluation.
Method
Across two phases, the study examines moral preferences and question framing across three use cases, three framing conditions, and two phases.
Results
Elicited preferences are contingent on configuration: morally relevant features vary by context, ideology shifts roughly one-third of feature judgments, and framing alters judgments and foundation associations.
Takeaways & Limitations
Preference-aggregation systems should disclose and audit feature definitions, voter samples, and question framing because elicited preferences are contingent rather than fixed.
Takeaways & Limitations
The US-based sample limits generalizability across cultural contexts.
Abstract
from arXiv · showhide
As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices in the moral AI elicitation pipeline: feature scoping, voter sampling, and question framing. In other words, they decide which features go to a vote, which voters to include, and how to present the question. These choices are often opaque, undocumented, and treated as technical details rather than normative ones. We examine each of these choices within a common empirical study and show that each can shape the preferences produced by moral AI elicitation. Across two phases (N = 809) in three deployment contexts (i.e., AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of the deceased), we examine the three main stages of the moral AI elicitation pipeline. First, morally relevant features shift across contexts. This suggests that feature schemas should not be assumed to transfer across deployment domains. Second, preferences differ by political ideology for roughly one-third of features, with some differences reversing direction. The ideological composition of the voter pool can therefore affect the resulting aggregated preference profile. Third, the wording of the elicitation question can narrow or widen ideological gaps by up to a full scale point. The framing conditions also change how moral foundations are associated with participants' judgments. Taken together, these findings suggest that voting-based alignment cannot deliver fair or transparent AI by aggregation alone; at minimum, each stage of the moral AI elicitation pipeline should be audited and disclosed.
1 Introduction
Moral preference elicitation is shaped by developers’ upstream choices about which features are voted on, whose preferences are represented, and how questions are framed. These choices are often undocumented and normative, challenging the assumption that aggregation alone produces fair, accountable, and transparent alignment.
- Study design: The study examines feature scoping, voter sampling, and question framing across AI kidney allocation, agents simulating absent workers, and generative AI depictions of the deceased.Participants with different political views were classified on whether features were relevant, irrelevant, or divisive.
- Main findings: Question framing can narrow or widen ideological gaps by up to one scale point and changes how moral foundations relate to participants’ judgments.The study compared a control condition with World-You-Want and Could-Be-You framing across three AI use cases.
- Main findings: Across contexts, morally relevant features differ, so feature scoping defines the moral boundaries of AI elicitation rather than merely implementing a technical design.KIDNEY emphasizes distributive justice and medical utility, WORK emphasizes accountability and legitimacy, and GEN centers on consent and dignity.
- Main findings: Political ideology shapes preferences for roughly one-third of features, meaning voter-pool composition can change the aggregated moral preference profile.Some ideological differences reverse direction, so aggregation does not simply recover a fixed public preference.
2 Related Work
Moral preference elicitation aggregates responses to hypothetical dilemmas into AI policies, presenting alignment as democratization through public input. However, developers still shape the process through feature scoping, voter sampling, and question framing, each of which can influence moral judgments and challenge neutrality.
- Voting-Based AI Alignment: Moral preference elicitation collects responses to hypothetical dilemmas and aggregates them into decision policies that models apply at scale, replacing developers’ private judgments with public input.
- Developer Choices: Developers determine what participants evaluate, who participates, and how questions are posed, making feature scoping, voter sampling, and question framing normative choices.Claims of neutral aggregation assume features transfer across contexts, voter composition has limited influence, and wording reveals rather than shapes values; aggregation may also marginalize minority views.
- Feature Scoping: Feature scoping may be top-down or bottom-up, but forced-choice designs can narrow ethical problems by preselecting people, outcomes, and trade-offs while omitting uncertainty and likelihood.Prior work calls for documenting and opening these pre-vote decisions to scrutiny.
- Voter Sampling: Moral judgments differ across demographic, AI-literacy, and political groups, challenging systems that seek a single set of universal moral preferences.Moral Foundations Theory links ideological differences to distinct emphases, including care and equality among progressives and proportionality, authority, loyalty, and purity among conservatives.
- Question Framing: Question framing changes which moral considerations participants weigh, while response framing constrains which preferences they can express.Identical outcomes can elicit different responses when framed as lives saved rather than lives lost, and offering an equal-treatment option can substantially alter apparent support for unequal treatment.
3 Methods
The study examines moral preferences across three AI use cases, three question-framing conditions, and two phases. It combines participant elicitation, LLM-assisted feature coding, and analyses of ideology, moral foundations, and framing.
- Study design: The two-phase design first identifies candidate moral features and then evaluates their importance under Control, World-You-Want, and Could-Be-You framing conditions.The three use cases span rare, high-impact kidney allocation; common workplace simulation; and generative depictions of the deceased, selected by harm impact and likelihood of encounter.
- Question framing: World-You-Want prompts participants to consider long-term societal effects and AI policy, whereas Could-Be-You invokes perspective-taking through Rawls’s veil of ignorance; Control adds no prompt.The World-You-Want prompt asks participants to consider the world created if their answers shaped AI company policies.
- Participants: Phase 1 recruited 449 participants across use cases using quota sampling to balance conservatives, moderates, and progressives, while Phase 2 recruited 360 participants with equal conservative and progressive representation per framing condition.Phase 1 sample sizes were NKIDNEY = 150, NW ORK = 149, and NGEN = 150; participants were recruited via Prolific at a rate of at least 8 USD/hour.
- Feature elicitation: Participants listed five features that should and should not be considered, justified each feature, and supplied example levels for features deemed morally relevant.This procedure was conducted after consent, a use-case introduction, and adoption of a “moral point of view.”
- Analysis: GPT-4o-mini extracted and labeled features from responses, with author–LLM agreement of 85% for feature names and 97% for relevance labels; authors then calibrated names across five rounds.Features were retained using a prevalence threshold: the top 30 to 35 features, excluding those mentioned by fewer than 4 participants, a choice that may exclude minority perspectives.
4 Results
Across three deployment contexts, morally relevant features were highly context-specific, while most features reached consensus but many remained divisive. Political ideology and question framing also shifted elicited preferences, with framing effects reaching one full scale point.
- Consensus and disagreement: Most features reached consensus on moral relevance, but substantial minorities disagreed about direction, especially for features with unclear or mixed implications.Consensus covered 65% of KIDNEY, 57% of WORK, and 70% of GEN features, while divisive features comprised 31%, 43%, and 30%, respectively.
- Context-specific moral features: Morally relevant features differed sharply by context: KIDNEY emphasized health status and survival, WORK motivation and policy, and GEN consent and intended purpose.Participants focused on health status (47%) and survival chance (35%) in KIDNEY, request reason (32%) and company policy (27%) in WORK, and deceased’s consent (28%) and intended purpose (40%) in GEN.
- Political ideology: Ideological differences affected roughly one-third of features and sometimes reversed direction, so changing the voter pool’s ideological composition could alter aggregated preferences.Examples included conservatives favoring expected full recovery in KIDNEY and opposing paid-but-not-working requests in WORK.
- Question framing: Question framing systematically changed ideological gaps, reducing some while increasing others by up to one full scale point.Could-Be-You increased conservatives’ support for grief and memorial requests in GEN (b = 2.32, 95% CI [0.94, 3.71], p < .02), while World-You-Want altered foundation effects in WORK and GEN.
5 Discussion
The discussion shows that feature scoping, voter sampling, and question framing are normative choices that can materially shape elicited moral preferences. It recommends auditing each stage while recognizing that aggregation cannot resolve deep value conflict.
- 5.1 Feature scoping: Feature sets differed across deployment contexts, so developers delimit what moral elicitation can express before voting begins.Participants generally rejected sociodemographic features as morally relevant unless necessary for the decision, such as age in KIDNEY.
- 5.2 Voter sampling: Ideological preferences differed for roughly one-third of features, sometimes reversing direction, so sample composition can change the aggregated preference profile.The aggregated profile reflects the values of the sampled population.
- 5.3 Question framing: Question framing systematically shifted preferences and moral-foundation associations, reducing ideological differences in some cases but increasing disagreement in others.In GEN, one framing increased disagreement, and overall framing could not be assumed to produce consensus.
- 5.4 Recommendations: Figure 4 proposes a sensitivity audit requiring developers to document and stress-test pipeline choices under plausible alternatives, reporting consequential changes to resulting policies.The recommendations specifically address feature scoping, voter sampling, and question framing, alongside aggregation, disagreement, and system implementation.
- 5.5 Limits of aggregation: Even full transparency cannot make aggregation neutral or resolve deep value conflict, because pluralistic groups prioritize different values and averaging risks privileging one set.The discussion calls for explicit trade-offs and transparency rather than simple averaging.
6 Conclusion
Moral preference elicitation produces contingent rather than stable values because developers shape outcomes through feature scoping, voter sampling, and question framing. Voting-based alignment therefore relocates human judgment into pipeline design and requires disclosure and sensitivity assessment.
- 6 Conclusion: Across three deployment contexts, elicited preferences depend on how feature scoping, voter sampling, and question framing are configured.Feature scoping is context-dependent because morally relevant features differ across use cases, limiting transfer across domains; voter sampling also shapes outcomes.
- 6 Conclusion: Moral preference elicitation does not produce a single, stable set of values; deployed “public morality” partly reflects how developers construct the pipeline.The output depends on upstream design choices made by developers.
- 6 Conclusion: Voting-based alignment relocates rather than removes human judgment, making elicited preferences contingent and requiring disclosure of features, samples, framing, and output sensitivity.Without such disclosure, claims of neutrality or fairness are difficult to evaluate.
- 6 Conclusion: Because preferences shift with context, population, and framing, alignment requires normative judgment and institutional design beyond aggregation alone.Suggested directions include structured deliberation, stakeholder representation, accountability mechanisms, additional domains, and alternative elicitation methods.
Endmatter Statements
The authors disclose ethical safeguards, limitations in participant sampling and use-case framing, and the dual-use risk of their sensitivity-audit findings. They also report using ChatGPT-4 and Gemini 2.0 for editing, summarization, coding support, and response analysis, but not to generate original publication text.
- Ethical considerations: The study protected participants through organizational approval, anonymization, restricted data access, compensation of at least 8 USD/hour, and the right to withdraw.Participants were recruited via Prolific, and no personal identifiers were collected.
- Ethical considerations: The sensitive kidney-allocation and deceased-person depiction scenarios used expert consultation and standardized wording, but the authors do not consider the conditions normatively neutral.The Control condition provided only a baseline without additional framing.
- Limitations: Political sampling balanced conservative, moderate, and progressive participants, while acknowledging that this simplified categorization may exclude diverse, especially non-Western, perspectives.
- Limitations: Because the findings could enable manipulation of elicitation pipelines, the authors frame their contribution as a sensitivity audit and recommendations rather than prescriptive design rules.They also acknowledge that their positionality may have influenced aspects of the research.
- AI-use disclosure: ChatGPT-4 and Gemini 2.0 supported editing, summarization, figure and table structuring, coding, and response analysis, but not original publication text generation.The authors state that humans retained responsibility for final text, interpretations, and conclusions.
Supplementary Materials … C Language Model Prompt for Identifying Moral Features from Participant Responses (Phase 1)
The supplementary materials document participant demographics, the Phase 2 moral-feature rating interfaces and scenarios, and the Phase 1 language-model procedure for standardizing moral features and judgments. The prompt converts participant responses into structured annotations containing features, values, reasoning summaries, and moral-relevance labels.
- A Demographic Characteristics of Phase 1 and Phase 2 Study Participants: Phase 1 and Phase 2 participant demographics are reported across the three use cases, including political ideology, age, gender, race/ethnicity, education, and additional Phase 2 characteristics.Phase 1 reports age as mean ± SD; Phase 2 additionally reports AI literacy, religiosity, and area of residence.
- B.2 Phase 2 Instructions: Evaluating the Moral Weight of Features Under Three Framing Conditions: Phase 2 participants rated the moral importance of contrasting feature values on a 7-point scale from -3 to +3 under Control, World-You-Want, or Could-Be-You framing.The kidney interface illustrates age contrasts, while the other interfaces apply the same rating structure to work-related emergencies and respect toward the deceased.
- B.2 Phase 2 Instructions: Evaluating the Moral Weight of Features Under Three Framing Conditions: The WORK scenario asks whether an AI should simulate activity for absent remote workers or reject the request, including automated keystrokes or responses.The scenario concerns employees who ask an AI to appear active during work hours despite being absent.
- B.2 Phase 2 Instructions: Evaluating the Moral Weight of Features Under Three Framing Conditions: The deceased-content scenario asks whether generative AI should create a realistic video of a deceased person or reject the request, considering why the artist wants it.Participants evaluate features such as whether the artist is respectful rather than disrespectful toward the deceased.
- C Language Model Prompt for Identifying Moral Features from Participant Responses (Phase 1): The Phase 1 prompt instructs a language model to analyze participant answers about which features should or should not be morally considered in AI decisions.The prompt applies to kidney allocation, absent-worker simulation, and other listed use cases, with responses containing feature mentions, moral-consideration judgments, explanations, and feature levels.
- C Language Model Prompt for Identifying Moral Features from Participant Responses (Phase 1): The model must standardize explicit or implied features into single real-world feature names and identify at least two concrete values for each feature.Values are taken from level1 and level2 when available, otherwise cautiously inferred; vague, implausible, unsupported, or overgeneralized features and values are excluded.
- C Language Model Prompt for Identifying Moral Features from Participant Responses (Phase 1): The annotation procedure summarizes each participant’s reasoning, labels moral relevance as “should,” “should not,” “mixed,” or “unclear,” and returns the results in a strict JSON schema.The required output includes the response identifier, information field, features, values with explicit/inferred types, and templates per feature; the prompt requires JSON-only output.
D Moderated Mediation Path Model: Political Ideology and Moral Foundations Moderated by Question Framing (Phase 2)
Phase 2 uses a moderated mediation path model to examine how political ideology and six moral foundations relate to perceived moral importance under different question framings.
- Moderated mediation path model: The model tests whether political ideology influences six moral foundations, whose associations with perceived moral importance are moderated by question framing.The six foundations are Care, Equality, Proportionality, Loyalty, Authority, and Purity; the analysis uses Hayes’s 2017 Model 15.
E Phase 1 and Phase 2 Results for Use Case 1: AI Kidney Allocation (KIDNEY)
For AI kidney allocation, Phase 1 identified moral features that participants considered relevant or irrelevant, while Phase 2 elicited their moral relevance and importance. Moderated mediation analyses then examined how political ideology shaped moral preferences across framing conditions.
- Phase 1: Phase 1 catalogued moral features participants said should be considered in AI kidney allocation.Features were color-coded by mention rate, with features mentioned by fewer than 2.5% of participants omitted.
- Phase 1: Phase 1 also catalogued moral features participants said should not be considered in AI kidney allocation.The table uses the same mention-rate thresholds and omits features mentioned by fewer than 2.5% of participants.
- Phase 2: Phase 2 measured the moral relevance and importance of features for AI kidney allocation.The table reports significance tests for coded relevance against 0.5 and importance against 0.
- Phase 2: Moderated mediation models examined political ideology’s direct and indirect effects on moral preferences across three framing conditions.The analysis reports conditional direct effects for conservatives versus progressives and an index of moderated mediation (IMM).
F Phase 1 and Phase 2 Results for Use Case 2: AI Agents Simulating Absent Workers (WORK)
For WORK, Phase 1 identified moral features that should or should not be considered, while Phase 2 assessed their moral relevance and importance across framing conditions.
- Phase 1: Phase 1 identified moral features to consider for AI agents simulating absent workers, omitting features mentioned by fewer than 2.5% of participants.The features were organized by mention-rate thresholds of at least 25%, at least 10%, and below 10%.
- Phase 1: Phase 1 also identified moral features that should not be considered for the WORK use case, omitting features mentioned by fewer than 2.5% of participants.These features were color-coded by mention rates of at least 10% or below 10%.
- Phase 2: Phase 2 elicited each feature’s moral relevance and importance for AI agents simulating absent workers and tested both measures for statistical significance.Mean (Coded) was tested against 0.5, and Mean was tested against 0, using two-sided one-sample t-tests.
- Phase 2: Moderated mediation models examined direct and indirect effects of political ideology on WORK moral preferences across three framing conditions.The analysis reported conditional direct effects for conservatives versus progressives and an index of moderated mediation.
G Phase 1 and Phase 2 Results for Use Case 3: Generative AI Content of the Deceased (GEN)
For the generative-AI-deceased-content use case, the study identified which moral features should or should not be considered and then assessed their moral relevance, importance, and ideological variation. Conservatives were more likely than progressives to support granting an artist’s request under several specified conditions.
- Phase 1 Results: Phase 1 separately catalogued moral features that participants said should and should not be considered for generative AI content of the deceased.Features mentioned by fewer than 2.5% of participants were omitted; the tables use mention-rate thresholds to color-code remaining features.
- Phase 2 Results: Phase 2 elicited each feature’s moral relevance and importance using tests against 0.5 for coded means and 0 for means.
- Phase 2 Results: Moderated mediation results examined political ideology’s direct and indirect effects on moral preferences across three framing conditions.The analysis reported conditional direct effects and the index of moderated mediation (IMM).
- Phase 1 and Phase 2 Results: Conservatives are more likely than progressives to consider that an AI agent should grant an artist’s request when the artist is a close friend or relative, the video is intended for adults, or it respects the deceased’s tradition.