Source-linked AI summary
Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment
Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia
TL;DR
AI systems that affect multiple people create a social-choice problem that anonymous RLHF does not directly resolve. The paper reformulates alignment as linear welfare optimization over a convex impact space, then derives strategyproof and welfare-constrained protocols. It applies these protocols to human-preference data across several allocation, language, and moral-decision domains.
Problem
When AI decisions affect multiple people, the unresolved problem is how to aggregate divergent individual preferences rather than pool them anonymously.
Method
The paper represents model welfare consequences as impacts and reformulates alignment as linear social choice and optimization over a convex impact space.
Results
The framework yields strategyproof and unanimous voting-by-issues and random-dictatorship mechanisms, plus protocols maximizing utilitarian welfare under individual or group welfare constraints.
Takeaways & Limitations
Impact-space analysis provides one representation for comparing utilitarian, strategyproof, harm-bounded, and public-spirit alignment protocols through their welfare consequences.
Takeaways & Limitations
The framework relies on preferences being linear in model features, and its strategyproofness results assume the planner broadcasts individual selection probabilities in advance.
Abstract
from arXiv · showhide
When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model? The standard approach to aligning frontier AI models$\unicode{x2013}$reinforcement learning from human feedback$\unicode{x2013}$largely sidesteps this question and has poor social choice guarantees. However, it remains unclear what alternative should replace it. We show that, by focusing directly on an algorithm's welfare consequences, the alignment problem can be reformulated as linear optimization over a convex impact space, which makes it amenable to the standard toolkit of welfare economics and mechanism design. This reformulation clarifies how alignment protocols translate into welfare consequences and, conversely, how a social planner's desired constraints on welfare consequences can be translated back into alignment protocols. We apply this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous. Demonstrating the reverse direction, we also apply the impact representation to derive a family of alignment protocols that maximize utilitarian social welfare subject to various social desiderata, such as bounds on individual or group harm. We illustrate the welfare implications of these alignment protocols empirically using real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.
1 Introduction
The paper reframes alignment for multi-person AI decisions as a social-choice problem and proposes impact-space optimization as a constructive alternative to anonymous RLHF.
- AI systems affecting multiple people must reconcile potentially conflicting individual preferences when selecting among alternatives.
- Anonymous RLHF pools annotator feedback but does not directly address alignment’s intrinsic social-choice problem.
- More granular preference elicitation leaves open how individual preferences should be combined into one deployed model.
- A decision-making algorithm can be represented by an impact vector, reducing utilitarian alignment to linear optimization over a convex polytope.
- The main theorem gives constructive procedures to reformulate alignment as linear social choice and map solutions back to deployable model parameters.
- The paper studies strategyproof mechanisms and welfare-constrained alignment protocols, including voting-by-issues, random dictatorships, and harm bounds.
2 Related Work
Related work identifies distributional and strategic shortcomings of RLHF and develops alternatives for heterogeneous preferences; this paper places these objectives in a common geometric framework.
- Recent work studies how RLHF’s welfare distortion affects participants or groups with diverse preferences.
- Prior research reports that standard RLHF can be strategically manipulated, motivating approximately strategyproof alternatives.
- Within the paper’s restricted linear-reward domain, strategyproof mechanisms are characterized as option-set mechanisms, with random dictatorship as a focal example.
- Alternative approaches model latent preference types or optimize heterogeneous preferences under selected desiderata, whereas this paper uses a more general impact-space representation.
3 Preliminaries
The preliminaries model stochastic pairwise choices, linear utilities over feature embeddings, weighted utilitarian preferences, and welfare relative to random choice.
- The model assumes finite contexts and actions, agent-specific utilities, and stochastic binary choices following a Bradley–Terry–Luce model.
- Each context–action pair is embedded in a feature space, with each agent’s utility represented as the feature vector’s inner product with that agent’s preference parameter.
- Weighted utilitarian social preferences aggregate individual utilities using weights over agents and the same feature representation.
- The framework justifies weighted sums of utilities through Harsanyi’s aggregation theorem under expected-utility and Pareto-indifference conditions.
- Pairwise comparisons depend on feature differences, while training and deployment are described by separate query distributions.
- Individual welfare measures utility generated by deployment relative to randomly choosing between the two available options.
- Utilitarian social welfare is defined as the sum of individual welfares.
4 Framework: Alignment as Linear Optimization in Impact Space
The framework replaces nonlinear parameter-space comparisons with linear welfare optimization over a convex impact set, enabling social-choice analysis and welfare-specific constraints.
- 4 Framework: Alignment as Linear Optimization in Impact Space: Parameter-space alignment obscures comparisons because welfare is a nonlinear function of model parameters and anonymous RLHF can violate distributional welfare criteria.
- 4 Framework: Alignment as Linear Optimization in Impact Space: The key insight is to analyze the set of welfare implications directly rather than the model parameterization.
- 4 Framework: Alignment as Linear Optimization in Impact Space: Under linear utility assumptions, each model’s welfare effects are summarized by an impact vector, making welfare linear in that vector.
- 4 Framework: Alignment as Linear Optimization in Impact Space: The feasible impact set is generated by deployment queries and is convex, centrally symmetric, and contained in their span.
- 4 Framework: Alignment as Linear Optimization in Impact Space: Every interior feasible impact has a unique deployment parameter in the deployment-query subspace.
- 4 Framework: Alignment as Linear Optimization in Impact Space: Scaling a model direction benefits agents whose preferences align with that direction across deployment queries and harms agents whose preferences covary negatively.
- 4 Framework: Alignment as Linear Optimization in Impact Space: The utilitarian optimum is a boundary impact in the direction of the weighted-average preference.
- 4 Framework: Alignment as Linear Optimization in Impact Space: The impact-space formulation treats externalities as objects to regulate when designing alignment protocols.
5 Strategyproof Alignment in Impact Space
The impact-space framework places alignment within social choice theory, yielding strategyproof mechanisms and connecting anonymous RLHF to voting and random dictatorship.
- Framework: Impact-space alignment models protocols as mechanisms mapping preference reports to distributions over feasible impact vectors.This reformulation makes welfare consequences the direct object of analysis.
- Option-Set Mechanisms: Strategyproof mechanisms can be represented by menus determined by other agents’ reports, with each agent receiving their most-preferred available impact.The menu cannot be enlarged by the agent whose choice it serves.
- Voting by Deployment Query: Voting-by-issues mechanisms aggregate directional votes independently across deployment queries and are strategyproof and unanimous.Theorem 2 establishes both properties for every voting-by-issues mechanism.
- Random Dictatorships: Random dictatorship is impact-equivalent to voting-by-issues with vote-average aggregators and can be implemented by one static parameter exactly or in the limit.Exact implementation requires the target impact to lie in the feasible impact set’s relative interior; boundary points are attainable as limits.
- Random Dictatorships: Random dictatorship can have arbitrarily poor worst-case welfare relative to the utilitarian optimum when welfare weights are non-degenerate.The adverse profile concentrates strong minority preferences against weak majority preferences, creating outsized minority harm.
- Random Dictatorships: Anonymous RLHF is impact-equivalent to sampling a labeler and response for each deployment query, becoming voting-by-issues in the deterministic labeling limit.With query-independent labeling weights, the limiting mechanism is random dictatorship.
6 Linear Constraints Can Enforce Externality Bounds and Trade-offs
Linear welfare constraints become halfspace constraints in impact space, allowing alignment protocols to trade utilitarian welfare against individual, group, or externality protections.
- General Constraints: Linear welfare guarantees turn the alignment problem into linear optimization over a convex feasible impact space.The formulation supports individual floors, harm-to-benefit ratios, and participation-externality bounds.
- General Constraints: Only binding constraints alter the effective welfare direction, and the optimal utilitarian value decreases concavely as thresholds tighten.Binding constraints receive nonnegative multipliers that tilt the objective.
- Welfare Floors: Harm-bounded alignment imposes individual welfare floors, giving additional effective weight to agents who would otherwise fall below them.The construction can also use group preferences and group welfare floors.
- Public-Spirit Constraints: Public-spirit constraints permit personal harm when sufficient social gain justifies it, reducing to individual rationality when γ = 0.The comparison is made against a reference impact.
- Welfare Floors: At high choice precision, tightening welfare floors toward zero removes harmed agents’ left tail and compresses welfare upward, while reducing mean welfare.More constraints bind as the floor tightens.
- Participation Externalities: Externality bounds limit how much including one agent can reduce the aggregate welfare of everyone else relative to a reference impact.When the bound is active, the agent’s preference is attenuated while the utilitarian direction is amplified.
7 Empirical Illustrations
The framework is evaluated on four real human-preference datasets, comparing utilitarian, RLHF, strategyproof, and harm-bounded mechanisms through welfare and externality measures.
- Setup: The empirical study covers kidney allocation, charitable food distribution, trolley problems, and LLM responses using fitted linear participant preferences.Feature spaces contain between 5 and 20 interpretable features.
- Mechanisms and Measures: The comparison measures each agent’s welfare relative to the utilitarian optimum and person-on-person participation externalities.Mechanisms include anonymous RLHF, uniform random dictatorship, and harm-bounded alignment at multiple floors.
- Findings: Across Community Alignment, harm-bounded alignment prevents welfare from falling below its chosen floor, but mean welfare declines as the floor tightens.The cost is associated with a small number of agents whose harm constraints bind.
- Findings: Preference disagreement concerns direction in Moral Machine and Community Alignment but mainly magnitude in kidney and food-rescue data.The strategyproof mechanism therefore loses only a fraction of optimal welfare in these datasets.
8 Limitations and Discussion
The framework depends on linear, interpretable preference representations and assumptions about query distributions and planner communication; the discussion positions impact space as a broader alignment lens.
- Limitations: The framework assumes preferences are linear in the model’s feature space, whose interpretability is nontrivial for LLM alignment.The LLM demonstration imports an external feature-extraction pipeline.
- Limitations: Some theoretical assumptions about query distributions may not hold in practice, and strategyproofness requires broadcasting selection probabilities in advance.The optimal strategyproof solution may additionally require a fixed-point procedure learned over time.
- Discussion: The discussion treats welfare impacts as a first-order object for designing strategyproof and welfare-constrained alignment mechanisms.It proposes extending the framework to demographic-group impacts, externality characterization, and other alignment protocols.
A Proofs for Sections 3 and 4
The proofs characterize the attainable impact set as a convex, origin-symmetric zonotope and show that the impact map covers its relative interior. Support-function arguments identify welfare-maximizing impacts, including unique vertices under generic directions.
- Technical scope: The closed impact space is optimized even when the inverse impact map is used only on the attainable relative interior.The proofs optimize over the closure while restricting ψ^-1 to the region where the impact map is invertible.
- Impact-space characterization: The attainable impact set is an origin-symmetric zonotope generated by deployment queries, with single-parameter impacts equal to its relative interior.The impact map is a bijection from the deployment subspace onto that relative interior.
- Impact-space characterization: The impact map is the gradient of a strictly convex log-partition potential on the deployment subspace.Positive-definite Hessian and Legendre-type properties yield the required bijection and smooth inverse.
- Welfare maximization: For any direction, the support function gives the maximum attainable welfare-weighted impact over the closed impact space.The proof evaluates this maximum as 1/2 E_z∼q_dep[|α_z^⊤θ|].
- Welfare maximization: Under generic directions, the maximizing impact is unique and forms an exposed vertex of the impact polytope.The maximizer is determined by the signs of α_z^⊤θ; at an agent’s ideal direction, it is that agent’s ideal impact.
B Proofs for Section 5
These proofs connect strategyproof alignment to menu mechanisms and establish strategyproofness and unanimity for structured mechanisms. They also show that random dictatorship can be welfare-reproduced by a single parameter, while its welfare ratio can become arbitrarily small relative to first-best welfare.
- Structured mechanisms: Voting-by-issues is strategyproof and unanimous because truthful reports maximize each voter-controlled issue simultaneously.Monotonicity of each issue rule makes truthful voting optimal term by term, while unanimous votes recover the ideal impact.
- Structured mechanisms: Random dictatorship’s expected impact can be reproduced by a single parameter whenever that mean lies in the relative interior of the impact space.The reproducing parameter is ψ^-1(ψ̄_λ), and it yields identical welfare for every agent.
- Welfare limitation: The first-best welfare in the constructed profile is U⋆ = ε h_Ψ̄(d2) > 0, while random dictatorship’s welfare ratio satisfies U_sp/U⋆ = o(1) as ε ↓ 0.This result holds without restrictions on the random-dictatorship weights λ.
- Welfare limitation: The low-ratio construction requires no continuity or strict-convexity assumption on the impact space and applies to the paper’s zonotopes.Upper hemicontinuity and singleton maximizers at selected directions suffice for the limiting argument.
C.1 The Identity and Its Mechanism Consequences
The identity links RLHF training statistics to deployed impact, revealing when deterministic RLHF implements strategyproof voting-by-issues and when finite precision permits manipulation. It also shows that anonymous RLHF is constrained in recoverable welfare outcomes, whereas individual-level weighting can recover broader welfare objectives.
- Impact identity: The deployed impact depends on training data through per-query label frequencies, enabling an equivalent random-labeler protocol.The protocol answers each query using the corresponding population label frequency and reproduces the same expected impact.
- Mechanism consequences: In the deterministic labeling limit, convex aggregation of votes yields a strategyproof and unanimous voting-by-issues mechanism.When query weights are agent-independent, the rule becomes a weighted aggregation of individual votes.
- Finite-precision manipulation: At finite labeling precision, agents can manipulate RLHF by intensifying reported preferences, and no optimal report exists.For generic preferences, reporting νθn with ν > 1 strictly improves utility, while the supremum is approached only as ν →∞.
- RLHF efficiency: RLHF stationarity requires marginal welfare effects to cancel on average, whereas utilitarian efficiency requires them to vanish source by source.When individual source effects are nonzero, reweighting positive-effect and negative-effect sources improves deployed welfare to first order.
- Recoverability: Individual-level weighting can recover every welfare-weighted parameter in the population preference hull, but anonymous RLHF is pinned to at most one identifiable parameter.The resulting welfare gap equals the utilitarian preference’s value on the impact distortion and is nonnegative.
D.1.1 Participation Alignment Externality
Participation alignment externality measures how including one agent’s data changes other agents’ welfare, providing a mechanism-design lens for regulating training effects. Welfare-floor mechanisms constrain these externalities and protect binding agents, but their inclusion can impose substantial costs on others.
- Definition: Participation alignment externality is the welfare impact of including an agent’s data relative to excluding that agent.It is computed from the difference between full-population and leave-one-out deployed impacts.
- Welfare interpretation: An agent benefits another whenever that agent’s preference aligns positively with the participation-induced impact shift, and harms them when alignment is negative.The aggregate externality uses the leave-one-out utilitarian preference of the remaining population.
- Optimization: Welfare-floor constraints can be formulated as a finite linear program over a convex impact polytope.Linear-programming duality shows that only binding constraints receive positive multipliers in the effective welfare direction.
- Externality regulation: Participation externalities connect welfare floors to the effects of agents’ inclusion on others, allowing constrained mechanisms to regulate the externality system directly.The welfare-floor condition requires participation effects to lift protected agents above their floors given their baseline without the relevant participant.
- Public spirit: Public-spirit constraints shift binding agents’ influence from personal protection toward the social objective as the public-spirit parameter increases.Small values emphasize individualized compensation, whereas large values make the solution more utilitarian.
- Empirical implications: Under the per-agent harm bound, a small set of binding agents accounts for almost all participation externalities.Removing a binding agent eliminates that constraint and lets others move closer to the utilitarian regime, benefiting the remaining population.
E.5 Four Dataset Comparisons
The four datasets differ in preference disagreement and in how the strategyproof mechanism affects expected welfare, welfare dispersion, and the welfare of agents at the extremes.
- Moral Machine and Community Alignment show high preference disagreement, with many preferences on both sides of 0.
- Kidney Allocation and 412 Food Rescue show less disagreement, with more unanimous preference direction despite disagreement about magnitude.
- Figure 8 compares expected welfare across the four datasets using bootstrap means and standard deviations for the strategyproof mechanism.It also includes welfare calculated through sequential addition across 100 permutations of agents.
- The strategyproof mechanism produces the least welfare dispersion across the four datasets, while RLHF is second for larger values of β.
- Figure 10 compares the welfare of the best- and worst-off agents across the four datasets for three model parameters.