Source-linked AI summary

ProsocialDialog: A Prosocial Backbone for Conversational Agents

Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, Maarten Sap

arXiv:2205.12688v2cs.CL

TL;DR

Existing dialogue systems may ignore or agree with unsafe user utterances, motivating a dataset and models for norm-guided prosocial responses. ProsocialDialog combines human-AI-created dialogues with rules-of-thumb and safety labels, and introduces Prost and Canary. Prost produces more socially acceptable responses than other state-of-the-art models, while Canary improves prosocial responses from language models, though the approach remains limited by dataset and model failure cases.

  • Problem

    Existing conversational systems can produce or agree with unsafe content, while evasive strategies may disrupt dialogue or block safe discussions.

  • Method

    The paper creates ProsocialDialog with human-AI collaboration, grounds responses in rules-of-thumb, and introduces Prost for dialogue and Canary for safety guidance.

  • Results

    Prost generates more appropriate responses than state-of-the-art language and dialogue models, while Canary guides large language models toward significantly more prosocial responses.

  • Takeaways & Limitations

    The results support using socially grounded datasets and safety guidance to steer conversational agents toward more prosocial responses in problematic contexts.

  • Takeaways & Limitations

    The dataset reflects mostly US-based, liberal-leaning, white crowdworkers, and Canary and Prost can still produce irrelevant, incoherent, inappropriate, biased, or toxic responses.

Abstract

from arXiv · show

Most existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them. To address this issue, we introduce ProsocialDialog, the first large-scale multi-turn dialogue dataset to teach conversational agents to respond to problematic content following social norms. Covering diverse unethical, problematic, biased, and toxic situations, ProsocialDialog contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-of-thumb, RoTs). Created via a human-AI collaborative framework, ProsocialDialog consists of 58K dialogues, with 331K utterances, 160K unique RoTs, and 497K dialogue safety labels accompanied by free-form rationales. With this dataset, we introduce a dialogue safety detection module, Canary, capable of generating RoTs given conversational context, and a socially-informed dialogue agent, Prost. Empirical results show that Prost generates more socially acceptable dialogues compared to other state-of-the-art language and dialogue models in both in-domain and out-of-domain settings. Additionally, Canary effectively guides conversational agents and off-the-shelf language models to generate significantly more prosocial responses. Our work highlights the promise and importance of creating and steering conversational AI to be socially responsible.

1 Introduction

Existing conversational systems may produce or agree with unsafe content, while evasive safety strategies can disrupt dialogue or exclude safe discussions. ProsocialDialog addresses this gap with norm-grounded dialogues, safety labeling, and models that generate constructive responses to problematic contexts.

  • Motivation: State-of-the-art conversational systems risk producing or agreeing with toxic, unethical, rude, or dangerous content.The paper attributes overly agreeable behavior partly to predominantly positive or agreeable training data.
  • Motivation: Mechanical avoidance strategies can disturb conversation flow and accidentally block safe discussions about sensitive topics such as gender or race.The paper identifies a need for responses guided by social norms rather than simple evasion.
  • Dataset and approach: PROSOCIALDIALOG is a 58K-dialogue dataset in which speakers respond prosocially to potentially unsafe situations by following social norms.The dialogues begin with potentially unsafe content and provide constructive, respectful guidance.
  • Dataset and approach: Human-AI collaboration uses GPT-3 to generate potentially unsafe utterances and crowdworkers to provide prosocial responses.This framework addresses the lack of large-scale human prosocial dialogue corpora while avoiding asking workers to write problematic utterances.
  • Models and tasks: The dataset supports prosocial response generation and fine-grained safety detection grounded in rules-of-thumb, with Prost and Canary released for these tasks.At each turn, the system determines safety, infers relevant RoTs, and generates constructive feedback.
  • Results: Prost generates more appropriate responses than state-of-the-art language and dialogue models, while Canary guides large language models toward more prosocial zero-shot responses.These findings are reported for problematic contexts and include both quantitative and qualitative evaluation.

2 Prosociality and Receptiveness in Conversational Agents

The paper frames prosocial conversational behavior as constructive, norm-guided feedback that remains receptive to users and distinguishes response needs across context types. Its design also acknowledges that crowdsourced norms can be culturally narrow and may not equal moral correctness.

  • Prosociality: Prosocial behavior is defined as actions that benefit others or society, including helping others and following societal norms.The paper introduces prosociality as a basis for handling problematic conversations directly.
  • Prosocial responses: Agents should infer appropriate social rules and guide users toward them while adapting to rules that differ across cultures and time.The dataset grounds constructive feedback in RoTs and dialogue context so feedback can be customized to new rules.
  • Receptiveness: PROSOCIALDIALOG promotes receptiveness through asking questions first, combining feedback with empathy, and showing better alternatives.These strategies are intended to help interlocutors adjust their behavior toward prosociality.
  • Safety labeling: The three-way safety schema classifies contexts as Needs Caution, Needs Intervention, or Casual based on the response or action an agent should produce next.This differs from classifying the safety or toxicity of the context itself.
  • Safety labeling: Needs Intervention covers situations such as medical issues or imminent danger where seeking help from real humans may be required, while Casual covers non-problematic interactions.Needs Caution covers potentially problematic, unethical, rude, toxic, or biased situations requiring careful responses.
  • Limitations: Crowdsourced social norms may privilege dominant opinions and remain incomplete across cultures, so they are not equivalent to moral correctness.The authors release individual annotations and use bias-focused data, while calling for further study of effects on marginalized groups.

3 PROSOCIALDIALOG

PROSOCIALDIALOG is collected through a human-AI pipeline that combines GPT-3-generated problematic dialogue with crowdworker-written, RoT-grounded constructive feedback and layered validation. The resulting dataset provides large-scale, richly annotated, dynamically changing safety contexts with more negative and prosocial content than conventional dialogue datasets.

  • Data Collection: PROSOCIALDIALOG uses GPT-3 to draft problematic dialogue openings, while crowdworkers provide constructive feedback grounded in rules-of-thumb.Dialogues can continue for up to six turns, with workers revising utterances for coherence and soundness.
  • Data Collection: Workers select or write RoTs, guide interlocutors toward prosocial behavior, and may respond freely when no problematic behavior is present.The annotation process explicitly connects feedback to communicative social norms while allowing non-problematic contexts to remain unforced.
  • Dataset Analysis: PROSOCIALDIALOG provides five context labels that preserve disagreement, ranging from CASUAL through graded caution categories to NEEDS INTERVENTION.The initial three-way annotation is refined into five labels to retain subjective annotator differences.
  • Dataset Analysis: 58,137 dialogues contain 331,362 utterances, 160,295 unique RoTs, and 497,043 safety annotations and reasons.The labels show Krippendorff’s α=0.49, and 42% of utterances are labeled Needs Caution.
  • Dataset Analysis: 42% of utterances are labeled Needs Caution, while safety states can shift across turns from casual discussion to intervention-level contexts or back to casual after accepted feedback.The dataset also reports that intervention contexts do not transition to CASUAL, whereas some caution contexts de-escalate after interlocutors accept feedback.
  • Dataset Analysis: Compared with other dialogue datasets, PROSOCIALDIALOG contains richer negative and constructive-feedback content rather than predominantly agreeable utterances.The comparison uses sentiment-classifier analysis and reports 166K utterances annotated by three workers with free-form rationales.

4 Building Socially Responsible Dialogue Agents with PROSOCIALDIALOG

The paper builds Canary as a safety module that jointly generates dialogue safety labels and RoTs, and Prost as a dialogue agent trained to produce responses with or without explicit RoT generation. The models are trained with complementary dialogue data so agents can navigate both problematic and casual contexts.

  • System Design: The paper separates Canary from Prost so the safety module can be updated when social norms or safety criteria change without retraining the entire dialogue agent.This modular design connects safety reasoning to response generation while preserving independent update paths.
  • 4.1 Canary: Canary generates both a safety label and relevant rules-of-thumb from a potentially problematic dialogue context.Unlike binary classification, RoT generation explains what is problematic and grounds the agent’s prosocial communicative intent.
  • 4.1 Canary: Canary is trained to model p(s, r|c) by generating a safety token followed by one or more RoTs, with only the safety token for CASUAL contexts.The target text concatenates the label and RoTs into one generated sequence.
  • 4.1 Canary: Canary uses T5-large and has variants pretrained on Social Chemistry, MIC, and Commonsense Norm Bank data, supplemented with casual dialogue datasets.The additional dialogue data accommodates diverse safe contexts.
  • 4.2 Prost: Prost is trained as the guiding speaker to generate either both RoTs and a response, p(u, r|c), or only a response, p(u|c).Both variants use maximum-likelihood training.
  • 4.2 Prost: Prost combines PROSOCIALDIALOG with six large-scale dialogue datasets to balance its deliberately negative feedback content against predominantly positive conversational data.The training mixture is intended to support navigation across diverse contexts rather than only objectionable ones.

5 Experiments on PROSOCIALDIALOG

The experiments evaluate Canary for dialogue safety classification and rules-of-thumb generation, and Prost for prosocial response generation on PROSOCIALDIALOG. Across automatic and human evaluations, Canary's social-norm knowledge and Prost's rules-of-thumb-conditioned generation improve performance.

  • 5.1 Dialogue Safety Classification & Rule-of-thumb Generation: Canary generally outperforms vanilla T5, while Delphi-based Canary outperforms all evaluated models on dialogue safety classification and rules-of-thumb generation.The results attribute this advantage to Delphi's knowledge of common patterns in human moral sense for short snippets.
  • 5.2 Response Generation via Prost: Human evaluation compares response models across prosociality, engagement, respect, coherency, and overall quality using 400 randomly sampled test examples.Judges compare two model responses and may select a tie.
  • 5.2 Response Generation via Prost: Prost with RoT & Response generally performs better than Response only on automatic and human evaluations of PROSOCIALDIALOG.Providing gold RoTs improves automatic evaluation further, suggesting that RoTs help guide prosocial responses.
  • 5.2 Response Generation via Prost: Prost performs better than GPT-3 and Instruct GPT-3 across all reported human-evaluation metrics on PROSOCIALDIALOG.The dataset is unseen by GPT-3 models, whereas Prost is trained on PROSOCIALDIALOG, contributing to the observed performance gap.

6 Generalizability of Prost and Canary

The paper tests Prost on real-world toxic language and Canary as a steering module for large pretrained language models. Prost more often disagrees with toxic content, while Canary improves model responses and narrows the gap between GPT-3 and Instruct GPT-3.

  • 6.1 Generalizability of Prost: On ToxiChat, both Prost models produce more disagreeing responses than other evaluated dialogue agents.BlenderBot 1 and GPT-3 show much higher rates of responses agreeing with toxic content than Prost and other models.
  • 6.1 Generalizability of Prost: Prost (RoT & Response) generates more toxic or offensive responses than Prost (Response), likely because its disapproving outputs contain offensive implications.The paper notes that neural models may mistake such disagreement for offensiveness because of lexical correlations and weak handling of negation.
  • 6.1 Generalizability of Prost: BlenderBot 2 and Instruct GPT-3 produce 95.3% and 90% neutral responses, compared with 61.8% and 70.2% for BlenderBot 1 and GPT-3.The paper cautions that neutrality can still be perceived as condoning unacceptable behavior, especially in toxic contexts.
  • 6.2 Steering Language Models with Canary: Responses generated with Canary are preferred over responses without Canary by approximately 2–3× on prosociality and overall quality.Responses with Canary RoTs are better or at least as good on the other evaluated dimensions.
  • 6.2 Steering Language Models with Canary: GPT-3 equipped with Canary is on par with Instruct GPT-3 on overall quality and better on prosociality.The paper reports that Canary can close the gap despite Instruct GPT-3 receiving substantially more additional training.

7 Related Work

Prior dialogue-safety work largely detects problematic contexts or uses evasive response strategies. ProsocialDialog instead provides multi-turn conversations that directly model disagreement with unethical and toxic content using safety labels and rules-of-thumb.

  • 7 Related Work: Existing dialogue-safety research has primarily focused on binary or ternary detection of problematic contexts.Other work also develops classifiers for toxic agreement, fine-grained safety labels, canned non-sequiturs, toxicity steering, and apologies.
  • 7 Related Work: ProsocialDialog directly targets responses to unsafe content through conversations where speakers disagree with problematic utterances using safety labels and social norms.The paper presents it as the first large-scale multi-turn dialogue dataset focused on prosocial feedback to unethical and toxic contexts.

8 Conclusion

The paper introduces ProsocialDialog, Prost, and Canary for constructive, rules-of-thumb-grounded responses across problematic contexts. Experiments show that Prost navigates such contexts more prosocially and Canary improves large language model responses.

  • 8 Conclusion: ProsocialDialog is an English dataset providing constructive feedback aligned with commonsense rules-of-thumb across diverse problematic contexts.Its three-tier safety schema distinguishes situations requiring human intervention from those requiring careful responses.
  • 8 Conclusion: Prost navigates problematic contexts in a more prosocial manner, while Canary outputs relevant rules-of-thumb when dialogue is detected as not casual.The conclusion describes Prost as trained on the dataset and Canary as a dialogue safety model.
  • 8 Conclusion: Human evaluation shows that Canary significantly improves the prosociality and overall quality of large language model responses to objectionable contexts.Canary supplies rules-of-thumb that guide responses toward prosocial behavior.

9 Societal and Ethical Considerations

The paper addresses worker safety, dataset-release risks, cultural scope, and potential misuse while positioning prosocial dialogue agents within a broader socio-technical ecosystem.

  • Worker protection: Workers were protected through age restrictions, advance warnings, feedback access, compensation, and institutional ethics-board approval.Workers received approximately $15 per hour on average, and feedback received responses within 24 hours.
  • Dataset-release risks: Releasing the dataset creates a misuse risk because problematic interlocutor utterances could train agents to generate disturbing, troublesome, or dangerous content.The paper argues that agents nevertheless need exposure to such inputs to navigate them according to social rules.
  • Cultural scope: The dataset’s mainly US-based rules-of-thumb may not transfer universally across cultures or remain appropriate as social consensus changes over time.Applying the rules unchanged elsewhere or in the distant future could produce socially unacceptable outputs.
  • Annotation scope: The RoT set is only a subset of US social rules, and MTurk annotators may share group characteristics that bias annotations in a particular direction.The paper cautions that the resource should not be treated as a complete representation of general US social rules.
  • Training considerations: Training solely on the dataset can produce a negativity-prone chatbot because the data deliberately counterbalances positivity-biased dialogue corpora with more negative situations.The authors encourage using the dataset alongside other dialogue data.
  • Socio-technical context: The paper recommends human handover when needed and supports improved regulation of AI and dialogue-system misuse.This recommendation is tied specifically to the Needs Intervention label and the wider socio-technical ecosystem.

10 Limitations

The limitations concern cultural and annotator representativeness, subjective feedback, data-construction quality, and the remaining errors of Canary and Prost.

  • Representativeness: The dataset’s English-speaking MTurk workforce, concentrated in the US and with predominantly liberal-leaning and White participants, limits coverage of North American and other cultural norms.The authors note that some RoTs may therefore be debatable for readers with different backgrounds.
  • Normative validity: Crowdsourced social norms are not equivalent to moral correctness and may privilege dominant opinions over marginalized groups.The paper identifies this as a risk because dominant normative values have historically supported oppression.
  • Subjectivity: Constructive feedback is subjective and can vary widely, so some responses may be questionable or accusatory despite guidelines and multiple verification steps.The authors ground guidelines in social-science research and use verification to reduce this issue.
  • Data construction: The source data include selected material from Social Chemistry, ETHICS, and SBIC, with GPT-3 converting short situations into first-person narratives for some sources.SBIC posts are retained in their original form, while Social Chemistry and ETHICS situations are transformed through prompting.
  • Dialogue construction: Reflective-listening questions are deliberately rephrased rather than reduced to short inquiries, while GPT-3 is prompted to continue problematic roles from the initial utterance.These design choices aim to create engaged openers and realistic problematic continuations.
  • Response annotation: Workers provide constructive feedback grounded in RoTs, but may respond freely when no problematic behavior is found.Instructions emphasize grounding responses, advising socially accepted behavior, and explaining better alternative outcomes.
  • Quality control: Dialogue proofreading is necessary because feedback can be harsh or accusatory and GPT-3 responses may lack consistency and coherency.The paper treats tone verification as important for maintaining objectivity.
  • Safety-label aggregation: Safety labels preserve intervention votes conservatively: a single Needs Intervention vote makes the context NEEDS INTERVENTION, while CASUAL requires unanimity.The remaining vote combinations determine the intermediate category.

A.5 Additional Dataset Statistics

Additional statistics describe RoT diversity, source composition, worker demographics, model-training settings, evaluation criteria, and Canary’s effect on response preferences.

  • Dataset statistics: RoTs average 9.5 words, each dialogue includes 3.3 RoTs on average, and newly written RoTs outnumber selected candidates 6 to 4.These figures characterize RoT length, density, and originality.
  • Dataset statistics: 160,296 unique RoTs comprise 74% of 217,321 total, with 27% unique 3-grams versus 23% in Social Chemistry.The paper reports comparable RoT uniqueness of 73% for Social Chemistry.
  • Dataset composition: The problematic-situation sources are 62% Social Chemistry, 21% SBIC, and 17% ETHICS, producing 42,304 / 7,132 / 8,701 train / valid / test dialogues.The splits follow the corresponding source-dataset splits.
  • Worker statistics: 212 workers participated in annotation, with reported demographics covering gender, age, race, geography, education, class, and political stance.Almost all workers had lived in the US for more than 10 years, and 73% identified as White.
  • Worker statistics: Workers’ mean assertiveness and conflict-aversiveness scores were 2.79 and 3.63, respectively.The measures used five-point scales ranging from not at all to very much.
  • Human evaluation: The evaluation compares responses on prosociality, engagement, respect, coherency, and overall suitability.These criteria assess both social qualities and contextual response quality.
  • Canary pipeline: Canary supplies sampled RoTs from dialogue context, then inserts them into a prompt asking a pretrained language model to explain the rule gently.The resulting prompt is used to generate the model’s response to the conversation.
  • Additional evaluation: Canary-guided GPT-3 responses were preferred 55.7% of the time over responses guided by irrelevant or random RoTs, preferred 28.4% of the time.The comparison supports the importance of selecting appropriate RoTs for controlling language models.

E Dialogue Dataset Descriptions

The paper situates ProsocialDialog among datasets emphasizing casual, knowledge-grounded, empathetic, persona-based, and mixed dialogue skills, while emphasizing its less-positive tone.

  • Existing dialogue datasets: Many large-scale dialogue datasets prioritize positive conversational elements such as emotion, persona, empathy, knowledge, commonsense, or combinations of these skills.ProsocialDialog is contrasted with this prevailing emphasis on positive or casual interaction capabilities.
  • Dataset descriptions: DailyDialog contains casual conversations collected from English-learning websites.It represents everyday conversational practice rather than problematic-content handling.
  • Dataset descriptions: TopicalChat provides knowledge-grounded conversations across eight popular topics.Its conversations organize dialogue around topical information such as fashion, books, sports, and music.
  • Dataset descriptions: Wizard of Wikipedia pairs a learner with a knowledgeable speaker in Wikipedia-grounded conversations, while PersonaChat uses specified personas for getting-to-know-you dialogue.These datasets target knowledge use and persona-conditioned interaction, respectively.
  • Dataset descriptions: EmpatheticDialogues centers on one speaker showing empathy toward another emotional speaker, and BlendedSkillTalk combines persona, empathy, and knowledge skills.Both datasets emphasize supportive or multi-skill casual interaction.
  • ProsocialDialog: PROSOCIALDIALOG is much less positive in tone than the comparison situations and conversations, making toxic or unsafe utterances less out-of-domain for trained models.The paper presents this tonal difference as a distinguishing dataset characteristic.
Loading 2205.12688v2…