Source-linked AI summary
In conversation with Artificial Intelligence: aligning language models with human values
Atoosa Kasirzadeh, Iason Gabriel
TL;DR
Conversational agents can produce useful language but also false, offensive, or irrelevant outputs, raising the question of how they should be aligned with human values. The paper develops a principle-based account grounded in philosophical and linguistic analysis, emphasizing pragmatic norms and domain-specific discursive ideals. It concludes that these norms can complement harm reduction and inform context-sensitive conversational-agent design, within an English-focused scope.
Problem
Conversational-agent alignment must address which norms and values should guide agents, because reducing particular harms does not by itself specify ideal or beneficial communication.
Method
The paper analyzes linguistic communication through philosophy, linguistics, speech-act theory, pragmatic norms, Gricean maxims, and discursive ideals across three conversational domains.
Results
The paper proposes that pragmatic norms and domain-specific discursive ideals can operationalize a principle-based approach to designing better value-aligned conversational agents.
Takeaways & Limitations
The proposal supports context-sensitive evaluation and fine-tuning of language models for more productive conversational exchange.
Takeaways & Limitations
The analysis focuses on English textual language and draws primarily on pragmatic and speech-act traditions, leaving other languages, communication modes, and theoretical traditions for further study.
Abstract
from arXiv · showhide
Large-scale language technologies are increasingly used in various forms of communication with humans across different contexts. One particular use case for these technologies is conversational agents, which output natural language text in response to prompts and queries. This mode of engagement raises a number of social and ethical questions. For example, what does it mean to align conversational agents with human norms or values? Which norms or values should they be aligned with? And how can this be accomplished? In this paper, we propose a number of steps that help answer these questions. We start by developing a philosophical analysis of the building blocks of linguistic communication between conversational agents and human interlocutors. We then use this analysis to identify and formulate ideal norms of conversation that can govern successful linguistic communication between humans and conversational agents. Furthermore, we explore how these norms can be used to align conversational agents with human values across a range of different discursive domains. We conclude by discussing the practical implications of our proposal for the design of conversational agents that are aligned with these norms and values.
1 Introduction
The paper frames conversational-agent alignment as a social and ethical problem that cannot be addressed solely by eliminating harms. It develops a principle-based, pragmatic account of ideal communication and applies it to context-specific conversational design.
- Motivation: Conversational agents expand natural-language interaction across domains, while their false, offensive, or irrelevant outputs create potential harms.These capabilities and risks motivate questions about which norms and values agents should follow.
- Motivation: Existing alignment efforts mainly mitigate specific harms, but removing errors alone may not establish what makes an agent good or beneficial.The paper therefore distinguishes harm reduction from a broader principle-based approach.
- Approach: The paper proposes identifying principles that specify ideal linguistic communication across contexts and incorporating those properties into conversational-agent design.This principle-based approach is presented as complementary to harm reduction.
- Approach: Its analysis examines syntactic, semantic, and pragmatic requirements, emphasizing pragmatic norms because they have received less attention in language-model scholarship.The paper uses these requirements to analyze successful communication between humans and conversational agents.
- Scope and implications: The paper develops discursive ideals for scientific discourse, democratic debate, and creative exchange, then considers implications for value-aligned conversational-agent research.It also asks whether the proposed approach captures all or the most important values underlying successful conversational-agent design.
- Scope and implications: The analysis is primarily grounded in speech-act theory, English-language textual communication, and pragmatic traditions, without directly engaging alternative theoretical traditions or other languages.The authors identify these as limitations and open questions for further research.
2 Evaluating human-conversational agent interactions
The paper evaluates human–conversational-agent interaction through syntax, semantics, and pragmatics. It argues that contextual meaning and pragmatic norms are essential because grammaticality and literal meaning alone do not determine successful communication.
- Communication framework: Conversation is treated as a cooperative activity governed by norms and implicit goals that shape its flow and content.Understanding these norms is presented as important for enabling agents to cooperate with human interlocutors across domains.
- Three evaluative lenses: The paper uses syntax, semantics, and pragmatics as distinct but complementary lenses for evaluating linguistic communication.It briefly discusses syntactic and semantic norms before focusing on domain-specific pragmatic norms.
- Three evaluative lenses: Syntax concerns sentence structure and grammatical rules, which are necessary for constructing intelligible linguistic expressions.Syntactic quality is relevant to language-agent evaluation but is not the paper’s main focus.
- Three evaluative lenses: Semantics concerns literal meaning and the mapping of statements to truth conditions, so grammatical sentences can still lack clear meaning.“Colourless green ideas sleep furiously” illustrates semantic incomprehensibility despite grammatical form.
- Pragmatics: The example “do whatever makes you feel better” shows that semantic analysis alone cannot determine whether an agent’s advice permits morally dubious actions.Its interpretation depends on contextual assumptions about what the speaker means and what actions are relevant.
- Pragmatics: Pragmatics interprets utterances through shared presuppositions, contextual information, common ground, and other features such as implicatures and cultural conventions.These contextual elements help determine what an utterance communicates beyond its literal wording.
- Pragmatic analysis: The paper organizes pragmatic analysis around utterance categories, Gricean conversational maxims, and domain-specific norms for guiding cooperative interaction.These schemas are introduced as tools for assessing which expressions are appropriate for conversational agents.
3 Utterances and maxims: towards value-aligned conversational agents
The paper classifies utterances by illocutionary act and argues that their validity depends on the kind of speech act and its context. It then proposes using cooperative conversational maxims to guide agent design, while recognising that their application varies across domains.
- Utterances and validity criteria: The paper uses illocutionary acts to classify utterances and examine what makes agent speech appropriate.This classification includes assertives, directives, expressives, performatives, and commissives.
- Utterances and validity criteria: Assertives represent the world, whereas directives aim to make listeners act, so their validity cannot be assessed by the same criterion.Assertives are evaluated by whether their content corresponds to the world; directives are evaluated in relation to their action-guiding function.
- Utterances and validity criteria: Expressives, performatives, and commissives create additional design problems because they involve mental states, changes in reality, or commitments to future action.The paper questions whether agents possess the relevant mental states, notes that performatives often fail to change reality, and observes that commissives may fail without memory or the capacity to act.
- Utterances and validity criteria: Because utterances have different validity criteria, conversational agents should be designed to make some kinds of statements while restricting others.The paper identifies an asymmetry in which utterances agents are able or allowed to produce.
- Conversational maxims: Gricean maxims of quantity, quality, relation, and manner provide general norms for cooperative dialogue, but their content depends on context.The maxims concern appropriate information, truth, relevance, and clarity; contextual variation means they cannot by themselves determine ideal speech in every domain.
4 Discursive ideals for human-conversational agent interaction
The paper develops domain-relative discursive ideals for scientific, democratic, and creative conversations between humans and conversational agents. These ideals connect conversational success to cooperative goals, subject matter, evaluative criteria, and contextual relationships.
- Domain-specific discursive ideals: The analysis applies pragmatic norms to scientific, democratic, and creative discourse, showing how domain-specific ideals guide appropriate agent behaviour.The domains are presented as illustrations of how cooperative goals and domain-relative information shape conversational norms.
- Scientific discourse: Scientific discourse prioritises epistemic virtues such as empirical adequacy, simplicity, coherence, consistency, and predictive accuracy.These virtues support scientific goals of identifying true or reliable knowledge, although they may need to be balanced against one another.
- Scientific discourse: Scientific agents should distinguish empirical claims from opinions and avoid misrepresenting claims to knowledge.Scientific conversation rests primarily on assertive utterances that promise some correspondence to the world.
- Scientific discourse: Scientific validity incorporates value judgements, so agents should articulate relevant values when needed; determining which virtues to respect requires broader interdisciplinary input.The paper states that scientific choices involve ethical beliefs, perceived social goods, and other value judgements, while the appropriate balance cannot be settled by existing norms alone.
- Democratic discourse: Democratic discourse places particular importance on civility, respect, tolerance, and consideration, including constraints against toxic or discriminatory speech.Conversational agents should evidence democratic qualities and respect democratic constraints rather than receive a direct mapping of citizens’ rights or prerogatives.
- Scope and caveats: The three domains are analytical distinctions, while real deployments require attention to hybrid discourse, intended audiences, topic familiarity, agent roles, and power relationships.Conversations may cross domain boundaries, and relational considerations add nuance to practical deployments.
5 Implications and consequences for conversational agent design
The paper translates its account of conversational norms into practical design implications for aligned agents. It emphasises contextual, domain-sensitive evaluation and careful treatment of utterance validity, anthropomorphism, transparency, and human interaction.
- Research implications: The paper identifies seven practical implications for future research on designing aligned conversational agents.These implications concern how the proposed analysis can inform future design and evaluation work.
- Operationalising norms: Gricean maxims can guide cooperative conversations, but their interpretation—especially quality—depends on conversational context and aims.Quantity may admit relatively uniform interpretation across domains, whereas other maxims vary more substantially.
- Operationalising norms: Different utterance types require different validity criteria, evidence, and substantiation methods rather than one universal standard of truthfulness or accuracy.The paper connects validity to the type of utterance and, for declarations, to whether the utterer has relevant authority.
- Context and meaning: Contextual meaning should inform agent design, requiring research on literal versus contextual meaning and improved handling of varied conversational settings.The paper identifies data annotation as one area where contextual information matters.
- Domain and culture: Domain-specific discursive ideals require interdisciplinary work to specify their content and account for variation across cultural backgrounds.The paper treats these ideals as appropriately contested rather than fixed by analysis of existing norms alone.
- Anthropomorphism and constraints: Because conversational agents are not moral agents, design may need to constrain their language and avoid anthropomorphic implications that could misattribute authority or responsibility.The paper discusses constraints on performatives and conditional use of expressives when transparency and value alignment support them.
- Human interaction: Context construction and elucidation may make conversations more robust and respectful while increasing users’ awareness of discourse goals when agents are transparent.The proposed approach would have agents provide relevant contextual information and clarify how conversational goals can be pursued.
- Evaluation: The framework could support further refinement of human and automatic evaluation of conversational-agent performance.The paper presents this as a possibility requiring further research.
6 Conclusion
The paper proposes principle-based alignment of conversational agents with human values by combining pragmatic norms, Gricean maxims, and domain-specific discursive ideals. It also identifies context-sensitive evaluation and further critical investigation as necessary for operationalising these ideals.
- The paper uses philosophy and linguistics to identify communicative components and pragmatic norms relevant to designing ideal conversational agents.
- It maps discursive ideals across three conversational domains and combines them with Gricean maxims as a possible principle-based design approach.
- The proposed norms require further technical and non-technical investigation before they can be fully interpreted and operationalised.
- The paper recommends context-sensitive evaluation and fine-tuning of language models for future research on aligned conversational agents.
- The analysis focuses on English and selected communicative norms, while drawing primarily on pragmatics and speech act theory.