Source-linked AI summary

A Roadmap to Pluralistic Alignment

Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, Yejin Choi

arXiv:2402.05070v3cs.AIcs.CLcs.IR

TL;DR

AI systems must serve people with diverse values and perspectives, but pluralistic alignment remains an open question. The paper proposes a roadmap using LLMs, formalizing three model definitions and three benchmark classes. Its evidence suggests current alignment techniques may reduce distributional pluralism, motivating new measurement and alignment methods.

  • Problem

    AI alignment must address widely varying human values and perspectives, but how to define pluralistic systems and measure them remains unresolved.

  • Method

    The paper uses LLMs as a testbed to formalize three forms of pluralistic models and three classes of pluralistic benchmarks.

  • Results

    Current alignment techniques are associated with reduced distributional pluralism, including lower similarity to target human distributions after alignment across two datasets.

  • Takeaways & Limitations

    Pluralistic evaluation and alignment require methods that explicitly represent diverse objectives, perspectives, and human distributions.

  • Takeaways & Limitations

    The definitions are not universally desirable or jointly satisfiable; for example, distributional pluralism may be unsuitable for controlled environments and may conflict with Overton pluralism.

Abstract

from arXiv · show

With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, aligning models to serve pluralistic human values remains an open research question. In this piece, we propose a roadmap to pluralistic alignment, specifically using language models as a test bed. We identify and formalize three possible ways to define and operationalize pluralism in AI systems: 1) Overton pluralistic models that present a spectrum of reasonable responses; 2) Steerably pluralistic models that can steer to reflect certain perspectives; and 3) Distributionally pluralistic models that are well-calibrated to a given population in distribution. We also formalize and discuss three possible classes of pluralistic benchmarks: 1) Multi-objective benchmarks, 2) Trade-off steerable benchmarks, which incentivize models to steer to arbitrary trade-offs, and 3) Jury-pluralistic benchmarks which explicitly model diverse human ratings. We use this framework to argue that current alignment techniques may be fundamentally limited for pluralistic AI; indeed, we highlight empirical evidence, both from our own experiments and from other work, that standard alignment procedures might reduce distributional pluralism in models, motivating the need for further research on pluralistic alignment.

1. Introduction

The paper frames pluralistic alignment as designing AI systems that serve people with diverse values and perspectives. It proposes explicit definitions, benchmarks, and evidence that current alignment techniques may reduce distributional pluralism.

  • AI systems must serve people whose goals, intentions, and values vary widely across tasks and populations.
  • The paper uses LLMs as a testbed while arguing that its pluralism concepts can generalize to other AI systems.
  • It formalizes Overton, steerable, and distributional pluralism as three ways to operationalize pluralistic models.
  • It defines multi-objective, trade-off steerable, and jury-pluralistic benchmarks, each suited to different evaluation needs.
  • Initial findings indicate that current alignment techniques reduce distributional pluralism, motivating further pluralistic evaluation and alignment research.

2. Arguments for Pluralism in AI Systems

The paper argues that pluralism matters because AI systems must support diverse users, values, and objectives. Treating human variation as signal can improve customization, technical understanding, and evaluation of generalist systems.

  • Customization requires systems that can reflect diverse use cases and values within applicable guardrails.
  • Pluralism treats human variation as signal rather than noise, potentially clarifying the relationship between model decisions and their sources.
  • Pluralistic evaluations measure generalist systems across varied objectives and users instead of optimizing only averaged human preferences.
  • The paper presents pluralism as a value because modern societies may regard competing values and perspectives as intrinsically important.
  • Reflecting human diversity can expose people to diverse ideas and help address unfairness from algorithmic monocultures.

3. Pluralism for AI Models/Systems

The paper distinguishes three model-level forms of pluralism—covering reasonable answers, steerable perspectives, and population distributions—and connects them to different applications and evaluations. It also outlines benchmark designs for objectives, trade-offs, and human distributions.

  • Distributional Pluralism: Distributional pluralism matches a model’s answer distribution to that of a target population, but requires a predefined population and suitable frequency data.
  • Overton Pluralism: An Overton window contains answers with suggestive but inconclusive evidence or substantial population support, subject to restrictions such as safety.
  • Overton Pluralism: Overton pluralism means presenting all potentially reasonable answers rather than selecting one idiosyncratically.
  • Overton Pluralism: Overton pluralism can be operationalized by clustering surveyed responses and filtering candidates through polling or support thresholds.
  • Overton Pluralism: Overton evaluations face risks from difficult expert scaling and false balance when reasonable-answer boundaries are poorly defined.
  • Steerable Pluralism: Steerable pluralism requires faithful responses to specified attributes, such as cultures, philosophical schools, or values.
  • Distributional Pluralism: Distributional evaluations compare distributions rather than averages, although stochasticity is undesirable when model behavior must be tightly controlled.

4. Pluralism for Benchmarks

Pluralistic benchmarks evaluate models across multiple objectives, explicit trade-offs, steerable objective combinations, and diverse human welfare preferences rather than a single averaged objective.

  • Multi-Objective Benchmarks: Pluralistic benchmarks measure more than one objective separately, unlike monistic benchmarks focused on a single objective.
  • Multi-Objective Benchmarks: Explicit multi-objective evaluation makes trade-offs such as helpfulness versus harmlessness visible for application-specific model selection.
  • Trade-Off Steerable Benchmarks: Trade-off steerable benchmarks test whether one model can maximize objectives while steering toward different commensurating functions at inference time.
  • Trade-Off Steerable Benchmarks: These benchmarks support deployment-time customization by measuring whether a single model represents solutions across a spectrum of objectives.
  • Jury-Pluralistic Benchmarks: Jury-pluralistic benchmarks explicitly model annotators’ utilities and combine them through a welfare function into an overall evaluation objective.
  • Jury-Pluralistic Benchmarks: Jury-based alignment can clarify which users are represented and potentially produce fairer outcomes, but estimating individual juror functions may require substantial data.

5. Current Alignment Approaches and Pluralism

The paper examines how alignment approaches relate to pluralism and hypothesizes that current techniques can reduce distributional pluralism. Initial experiments across model classes and datasets support this concern, while broader validation remains necessary.

  • Current approaches: Alignment methods include supervised fine-tuning and RLHF, whose resulting pluralism depends on annotator representativeness, dataset richness, and reward-model factors.Prior work also argues that monistic RLHF can violate democratic properties and underweight outliers.
  • Hypothesis: Current LLM alignment techniques can reduce distributional pluralism relative to internet-user populations.This is stated as the paper’s central hypothesis.
  • Empirical approach: The study compares pretrained and aligned LLaMA(2), Gemma, and GPT-3 models on GlobalOpinionQA and the 120-question Machine Personality Inventory.Target distributions include Japan, the US, and a global population; similarity is measured with Jensen-Shannon distance averaged over five prompts.
  • Results: Almost all pre-aligned models have lower Jensen-Shannon distance to target human distributions than post-aligned models on both datasets.The only reported exception is GPT-3 on MPI, whose results are difficult to verify because the original model versions and procedures are unavailable.
  • Results: Post-alignment also reduces entropy, consistent with prior findings of reduced distributional variance across domains.The authors therefore suggest alignment may limit pluralism when target populations are diverse.
  • Open questions: More comprehensive testing requires large-scale experiments across broader domains and further investigation of entropy’s role.The paper also notes that steerable pluralism through prompting remains insufficiently evaluated.

6. Discussion

The discussion frames the work as a set of pluralism definitions and evaluation frameworks rather than a prescription for whom models should serve. It emphasizes unresolved operational, desirability, and generalization questions.

  • Scope: The paper formalizes frameworks for aligning models to values, characteristics, or perspectives without specifying exactly whom or what to align.Its stated goal is clearer, more pluralistic alignment approaches.
  • Limitations: Several pluralism definitions are difficult to operationalize, including describing the Overton window and selecting a target population.The authors regard this precision as necessary for measuring pluralism and call for careful justification of design choices.
  • Limitations: The GPT-3 result is the sole reported exception to the main similarity pattern and should be interpreted cautiously because the original model and procedure are difficult to confirm.This limitation specifically affects the MPI comparison.
  • Desirability and trade-offs: Distributional pluralism may be useful for cultural or creative modeling but undesirable in controlled environments such as customer support.The paper also notes that Overton and distributional pluralism may conflict within a single model.
  • Broader systems: The LLM-focused definitions are presented as broadly applicable to other AI systems and modalities, especially subjective tasks with diverse users or objectives.The query/response framework can cover actions, images, audio, and other inputs and outputs.

7. Conclusion

The conclusion calls for more precise attention to pluralism in AI alignment and summarizes three model definitions alongside three benchmark forms. It recommends better evaluations, normative discussion, and new alignment techniques.

  • Conclusion: The paper formalizes three definitions of pluralistic models and three forms of pluralistic benchmarks.These frameworks organize how pluralism can be measured and operationalized.
  • Recommendations: The authors recommend fine-grained pluralistic evaluations, continued discussion of alignment targets and customization bounds, and additional techniques for creating pluralistic models.These recommendations are presented as broad directions for future work.

Impact Statement

The impact statement presents pluralistic alignment as potentially beneficial for systems serving diverse people while acknowledging risks from dual use. Because the work is theoretical, the authors judge marginal dual-use potential to be minimal.

  • Potential impact: The paper aims to encourage AI systems that work better with diverse people while acknowledging limitations and potential dual use.The authors specifically mention aligning systems to harmful attributes as a possible risk.
  • Risk assessment: The authors believe the theoretical work’s positive contribution to pluralism discussions outweighs its marginal dual-use potential.They characterize that potential as minimal.

A. Experimentation Details

The experiments compare pre- and post-aligned models across two diverse opinion and personality datasets, using model answer distributions and distributional metrics.

  • Datasets: The experiments use GlobalOpinionQA and the Machine Personality Inventory to evaluate model distributions against human responses.GlobalQA aggregates cross-national surveys, while MPI contains personality-trait questions.
  • Datasets: GlobalQA includes 741 questions with responses from both the United States and Japan, while MPI contains 120 questions and 600K responses from 240 countries.GlobalQA samples required at least 1,200 nationally representative adult respondents per country.
  • Models: The study compares three model classes, each with pre- and post-aligned versions: LLaMA, LLaMA2, and GPT-3.Post-alignment includes fine-tuning, RLHF, or unknown alignment types.
  • Model Distribution: Model distributions are formed from next-token probabilities for each answer choice, averaged across five prompted evaluations.In-context learning steers pretrained models to output the selected answer letter as the next token, and answer choices are randomized.
  • Evaluation Metrics: The primary metric is Jensen-Shannon distance to the target human distribution, where lower values indicate greater similarity; entropy is also computed.The values are averaged across questions.

A.1. Further Analysis

Further analysis finds that pre-aligned models are generally closer to target human distributions than post-aligned models, with larger distributional variation and entropy.

  • Distributional Similarity: Almost all pre-aligned models have lower Jensen-Shannon distance to target human distributions than post-aligned models on both datasets.The comparison is reported for both GlobalQA and MPI.
  • Distributional Similarity: The pre- versus post-alignment gap more than doubles from LLaMA to LLaMA2, while LLaMA2 model size has little apparent impact between 7B and 13B.The larger gap is associated with more training data and higher context length.
  • Distributional Variation: Pre-aligned models produce more variable answer distributions, whereas post-aligned models concentrate probability mass on one or two answer choices.The qualitative pattern is confirmed by average entropy comparisons.
  • Distributional Variation: Pre-aligned models have 100% more average entropy than post-aligned models across the analyzed distributions.The paper presents higher entropy as reflecting greater distributional spread.
  • Related Evidence: Prior studies similarly report lower entropy, poorer calibration, and reduced textual diversity in aligned models than in reference populations or pretrained models.These findings provide related evidence for reduced distributional variation after alignment.

B. Additional Experimentation

Additional experiments reproduce the pre-aligned models’ advantage in distributional similarity and test whether lower entropy alone explains that pattern.

  • Results: Across GlobalQA and MPI, pre-aligned models are closer to target human distributions than post-RLHF models.The experiments compare model distributions on multiple-choice questions against target human populations.
  • Entropy Analysis: Pre-aligned models show more variable answer-choice distributions, while post-aligned models exhibit more sharply concentrated distributions.The entropy analysis reports higher average entropy for every pre-aligned model.
  • Entropy Analysis: The shuffled-distribution analysis preserves model entropy while randomizing answer labels to estimate how much entropy contributes to human-model similarity.The resulting shuffled distributions are compared with the same human distributions using Jensen-Shannon distance.
  • Entropy Analysis: Shuffled distributions generally show greater similarity scores, indicating that entropy explains some, but not all, of the similarity between model and human distributions.The authors state that further investigation is needed to substantiate this interpretation.
Loading 2402.05070v3…