Source-linked AI summary

Artificial Intelligence, Values and Alignment

Iason Gabriel

arXiv:2001.09768v2cs.CY

TL;DR

AI alignment requires deciding whose values increasingly autonomous systems should embody, not merely solving how to encode them. This paper connects technical methods with normative choices, distinguishes several alignment targets, and argues that fair, pluralism-sensitive principles and processes are needed because indirect value-learning still requires moral evaluation.

  • Problem

    AI alignment must determine which values or principles to encode in systems, despite disagreement about morality and uncertainty that technical methods can remain value-neutral.

  • Method

    The paper analyzes the relationship between technical alignment methods and normative choices, compares alignment targets, and examines fair processes for pluralistic value alignment.

  • Results

    The paper concludes that technical and normative alignment are interrelated, indirect value-learning does not remove moral evaluation, and alignment should target more than instructions or revealed preferences alone.

  • Takeaways & Limitations

    AI alignment should be pursued as a combined technical and normative research agenda oriented toward fair principles rather than a single supposedly true moral theory.

  • Takeaways & Limitations

    Pluralistic alignment faces inconsistent preferences, impossibility results, divergent interpretations, and unresolved questions about representation and implementation.

Abstract

from arXiv · show

This paper looks at philosophical questions that arise in the context of AI alignment. It defends three propositions. First, normative and technical aspects of the AI alignment problem are interrelated, creating space for productive engagement between people working in both domains. Second, it is important to be clear about the goal of alignment. There are significant differences between AI that aligns with instructions, intentions, revealed preferences, ideal preferences, interests and values. A principle-based approach to AI alignment, which combines these elements in a systematic way, has considerable advantages in this context. Third, the central challenge for theorists is not to identify 'true' moral principles for AI; rather, it is to identify fair principles for alignment, that receive reflective endorsement despite widespread variation in people's moral beliefs. The final part of the paper explores three ways in which fair principles for AI alignment could potentially be identified.

1 Introduction

AI alignment raises questions about whose values should guide increasingly autonomous systems and how to choose principles in a pluralistic world. The paper distinguishes technical alignment from the normative task of deciding what values or principles AI ought to embody.

  • The alignment challenge: AI alignment asks whose values powerful, increasingly autonomous systems ought to follow.Proposed targets include happiness, universalizable principles, human instructions, intentions, interests, rights, and values.
  • The alignment challenge: Human instructions and desires may require constraints when following them could enable harm, imprudence, or self-destruction.Rights or objective interests are presented as possible limits on what AI may permissibly do.
  • The normative problem: Choosing alignment principles is difficult because people hold competing conceptions of value and may otherwise impose their views on others.The paper frames this as a question of both which principles to encode and who has authority to decide.
  • Two dimensions: The technical challenge concerns formally encoding values and evaluating agents whose abilities may exceed human cognitive capacities.The normative challenge asks which values or principles ought to be encoded.
  • Paper approach: The paper focuses on normative alignment, examining alignment targets and arguing for fair processes rather than identifying one true moral theory.It explores three potentially fair processes for determining which values to encode across moral disagreement.

2 Technical and Normative Aspects of Value Alignment

Technical methods and normative choices in AI alignment are interdependent: how agents are built may shape which values they can encode. Approaches that infer values from conduct or data still require moral evaluation, so alignment demands clarity about its goal and principles.

  • Technical and normative interdependence: The simple thesis holds that technical alignment could be solved before loading any preferred values or principles.The paper questions whether machine-learning systems can remain fully compatible with arbitrary later choices.
  • Reinforcement learning: Reinforcement learning trains agents to maximize numerical reward signals, making it naturally suited to consequentialist objectives.RL systems include policies, reward signals, value functions, and environment models; non-consequentialist goals such as satisficing and rights are less straightforward.
  • Technical and normative interdependence: The methods used to build artificial agents may influence which values or principles they can encode.This motivates value-open design that remains compatible with a wide range of perspectives and values.
  • Learning values indirectly: Inverse reinforcement learning and imitation-based approaches infer rewards or conduct from experts, observed behavior, or large datasets instead of specifying principles upfront.These approaches could learn virtuous conduct or aggregate values and preferences from many people.
  • Limits of technical solutions: Indirect approaches do not eliminate moral evaluation because deployment still requires deciding which experts, behaviors, or data should represent value.The paper identifies unresolved questions about excluding unethical behavior and ranking moral agents.
  • Limits of technical solutions: Facts about human behavior and beliefs cannot by themselves determine what AI ought to do.The paper concludes that any technical approach still requires greater clarity about alignment goals and appropriate moral principles.

3 The Goal of Alignment

AI alignment can target instructions, intentions, preferences, interests, or values, but each target has distinct benefits and limitations. The paper argues that guiding principles can combine these aims with objective constraints and avoid relying on any single imperfect proxy.

  • Different goals for alignment: Alignment goals differ substantially: AI may follow instructions, expressed intentions, revealed or informed preferences, human interests, or values.The paper treats these as distinct targets rather than interchangeable formulations of alignment.
  • Instructions and intentions: Literal instruction-following can produce outcomes that satisfy a request while missing the principal’s actual objective.The King Midas example and CoastRunners show how proxy optimization can achieve the specified target while defeating the intended goal.
  • Instructions and intentions: Intention alignment requires AI to understand implied meaning, language, culture, institutions, and practices, while human intentions may still be irrational, misinformed, or harmful.The paper therefore treats intention-following as useful but potentially constrained by broader considerations.
  • Preferences: Revealed and informed preferences provide accessible behavioral evidence but struggle with reward-function degeneracy, unobserved situations, adaptive preferences, and harmful ends.Informed preferences may reduce errors from ignorance and poor reasoning, yet they require filtering and do not resolve self-harming or unethical preferences.
  • Interests and values: Aligning AI with human interests may reduce imprudence, self-harm, and harm to others, but interests alone do not determine what people are morally entitled to do.The paper presents safety and respect for human interests as boundary conditions, not a complete account of acceptable action.
  • Interests and values: A values-based, principle-oriented approach can combine principal direction with objective constraints, while incorporating justice and other considerations in group decisions.The paper argues that alignment should use guiding principles anchored in evaluative judgments rather than seek one universally correct proxy.

4 Principles for Alignment

The paper argues that alignment should seek fair principles compatible with reasonable moral pluralism rather than a supposedly true moral theory. It examines overlapping consensus, human-rights-based principles, and social-choice or democratic procedures, while identifying substantial limitations in each approach.

  • The central task is selecting alignment principles compatible with diverse, reasonable, and contrasting beliefs about value.
  • A single true moral theory would not resolve alignment because reasonable pluralism makes reliably communicating that alleged truth difficult.
  • Human-rights-congruent AI could avoid domination and value imposition if human rights constitute a global overlapping consensus.
  • Human-rights alignment faces a trade-off: negative rights have broad support but limited scope, whereas positive rights offer richer guidance with less global support.
  • Existing AI principles remain difficult to operationalize because people diverge over their meaning, scope, implementation, and implications for processes versus outcomes.
  • An overlapping consensus may be premature unless alignment principles are intercultural, inclusive, concrete, and stable after implementation.
  • Social-choice approaches aggregate individual views fairly, but inconsistent preferences and impossibility theorems limit systematic collective rankings.
  • Democratic processes may confer legitimacy, yet designing voting procedures and scaling individual values to collective decisions remains theoretically and practically difficult.

5 Conclusion

The conclusion links technical methods and alignable values, rejects alignment with instructions or revealed preferences alone, and favors fair principles supported across moral disagreement. It recommends investigating consensus, veil-of-ignorance, and democratic approaches through procedurally fair, concrete, stable, comprehensive, and inclusive processes.

  • Machine-learning techniques and alignable values are interdependent, so normative questions should form part of a combined AI-alignment research agenda.
  • Properly aligned AI should not follow instructions, expressed intentions, or revealed preferences alone, because it must also address unethical or imprudent behavior.
  • The paper frames alignment as political rather than metaphysical and recommends principles supported by global consensus, veil-of-ignorance reasoning, or democratic processes.
  • Fair alignment procedures should avoid arbitrary advantage, provide concrete guidance, remain stable and robust, cover relevant cases, and include willing participants.
  • Alignment should account for possible widespread moral error rather than tethering AI too closely to present-day morality.

Compliance with Ethical Standards

The paper reports no conflicts of interest and states that all contributions are the author's own.

  • The author declares no conflicts of interest.
  • All contributions are identified as the author's own.
  • The article is published under a Creative Commons Attribution 4.0 International License, subject to attribution and license conditions.
Loading 2001.09768v2…