Source-linked AI summary

Can Machines Learn Morality? The Delphi Experiment

Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Etzioni, Maarten Sap, Regina Rini, Yejin Choi

arXiv:2110.07574v2cs.CL

TL;DR

AI systems increasingly make morally consequential decisions even though morality remains contested and difficult to formalize. The paper introduces Delphi, a neural model trained on crowdsourced moral judgments, and finds strong performance beyond its training examples while documenting persistent biases and inconsistencies.

  • Problem

    AI systems make morally consequential decisions, but morality remains difficult and contested, while off-the-shelf language models can lack moral sense and encode harmful biases.

  • Method

    Delphi is a descriptive commonsense moral-reasoning system built by training neural language models on crowdsourced ethical judgments and evaluating contextual moral judgments.

  • Results

    Delphi makes high-quality judgments across varied situations, achieving 92.8% accuracy on unseen COMMONSENSE NORM BANK examples versus GPT-3’s 60.2%.

  • Takeaways & Limitations

    Explicitly teaching moral sense offers an empirical path for improving AI systems that are otherwise fundamentally oblivious to human values, norms, and ethics.

  • Takeaways & Limitations

    Delphi remains susceptible to pervasive biases and inconsistent predictions, and its statistically dominant judgments are not established as prescriptively normative.

Abstract

from arXiv · show

As AI systems become increasingly powerful and pervasive, there are growing concerns about machines' morality or a lack thereof. Yet, teaching morality to machines is a formidable task, as morality remains among the most intensely debated questions in humanity, let alone for AI. Existing AI systems deployed to millions of users, however, are already making decisions loaded with moral implications, which poses a seemingly impossible challenge: teaching machines moral sense, while humanity continues to grapple with it. To explore this challenge, we introduce Delphi, an experimental framework based on deep neural networks trained directly to reason about descriptive ethical judgments, e.g., "helping a friend" is generally good, while "helping a friend spread fake news" is not. Empirical results shed novel insights on the promises and limits of machine ethics; Delphi demonstrates strong generalization capabilities in the face of novel ethical situations, while off-the-shelf neural network models exhibit markedly poor judgment including unjust biases, confirming the need for explicitly teaching machines moral sense. Yet, Delphi is not perfect, exhibiting susceptibility to pervasive biases and inconsistencies. Despite that, we demonstrate positive use cases of imperfect Delphi, including using it as a component model within other imperfect AI systems. Importantly, we interpret the operationalization of Delphi in light of prominent ethical theories, which leads us to important future research questions.

1 INTRODUCTION

Delphi is a neural system for commonsense moral reasoning that predicts judgments about everyday situations. It shows strong context-sensitive performance and outperforms off-the-shelf models, while remaining an imperfect, biased research system.

  • Delphi predicts ethical judgments about everyday situations expressed in natural language, including context-sensitive contrasts such as helping a friend versus helping spread fake news.Its outputs can be yes/no or free-form moral judgments.
  • 92.8% accuracy on unseen COMMONSENSE NORM BANK examples exceeds GPT-3’s 60.2%, supporting explicit moral training for AI systems.COMMONSENSE NORM BANK contains 1.7M crowdsourced ethical judgments.
  • The paper reports consistently high-quality predictions across varied situations but also emphasizes Delphi’s strengths, weaknesses, biases, and follow-up research questions.The system is presented as an improvement over AI systems described as oblivious to human values, norms, and ethics.
  • Delphi adjusts judgments when context changes, such as treating ignoring a friend’s call after a fight as acceptable rather than rude.The paper presents this as generalization beyond the training examples and robustness to altered contexts.

2 INCLUSIVE, ETHICALLY-INFORMED, AND SOCIALLY-AWARE AI

The paper frames machine ethics as necessary because AI systems encode social values while moral principles remain contested. Delphi therefore uses a largely bottom-up, example-based approach, while arguing that top-down constraints are needed to address systemic bias.

  • AI systems increasingly make decisions with moral implications, while existing models can propagate unethical biases from their training data.The paper argues that regulation alone cannot make models recognize and circumvent such biases.
  • Theoretical framework: Delphi follows a bottom-up, descriptive, example-based framework rather than a top-down, prescriptive, rule-based ethics paradigm.The approach is motivated partly by the lack of consensus on general moral principles.
  • Theoretical framework: Rawls’s decision procedure motivates learning shared moral patterns from many people’s judgments instead of relying on rules imposed by moral authorities.The method is presented as compatible with either objective moral truths or morality as a construct of human beliefs.
  • Theoretical framework: Bottom-up learning can capture situational nuance and cultural context, but crowdsourced morality remains vulnerable to shared prejudices and pervasive biases.The paper proposes combining bottom-up modeling with top-down principles of justice through reflective equilibrium.
  • Scope: Delphi is intended as a descriptive model of human moral commonsense, not a prescriptive morality for how people or machines ought to act.The paper states that prescriptive conclusions would require additional arguments beyond its scope.

3 COMMONSENSE NORM BANK: THE KNOWLEDGE REPOSITORY OF ETHICS AND NORMS

COMMONSENSE NORM BANK is a unified repository of 1.7 million descriptive ethical judgments from diverse datasets, covering everyday situations, social norms, contextualized narratives, and social biases. The data are standardized into free-form, yes/no, and relative task formats for training Delphi.

  • Data sources: COMMONSENSE NORM BANK contains 1.7 million descriptive judgments on everyday situations drawn from existing datasets.Its sources include SOCIAL CHEMISTRY, ETHICS Commonsense Morality, MORAL STORIES, and SOCIAL BIAS INFERENCE CORPUS.
  • Data sources: The repository adopts a bottom-up descriptive approach that represents crowdworkers’ judgments without endorsing particular values.The authors present it as representative of people’s morality and ethics rather than as a statement of their own value system.
  • Data sources: SOCIAL CHEMISTRY contributes actions, situations, three-way ethical labels, and open-text judgments extracted from crowdsourced rules of thumb.Its prompts come from online discussions, stories, and an advice column, providing core and contextualized real-life events.
  • Data sources: MORAL STORIES supplies actions grounded in situations or intentions, while ETHICS supplies short commonsense morality scenarios with binary labels.These sources add longer contextual narratives and relatively unambiguous everyday ethical situations.
  • Data unification: The unified formats include free-form judgments, yes/no agreement with rules of thumb, and a smaller relative mode comparing two situations.Free-form inputs vary from bare actions to actions grounded in situations and intentions; relative examples are discussed only in the appendix.

4 Delphi: COMMONSENSE MORAL MODELS

Delphi is a text-to-text model for commonsense moral reasoning, fine-tuned from a commonsense reasoning model on COMMONSENSE NORM BANK. It is evaluated across classification and open-text tasks using automatic metrics and human judgments against GPT-3 baselines.

  • Model: Delphi is a computational model that predicts yes/no or free-form answers to everyday situations and statements with moral implications.Its intended inputs include depictions, questions, and morally relevant assertions.
  • Model: Delphi is fine-tuned from UNICORN rather than T5 alone to leverage implicit commonsense knowledge for descriptive ethical reasoning.UNICORN is a multitask T5-11B model trained on the RAINBOW commonsense reasoning suite.
  • Model: The training setup unifies free-form, yes/no, and relative modes as text-to-text tasks with special tokens and structured input markers.The relative mode uses XML-like tags to distinguish paired actions and labels.
  • Evaluation: The comparison includes GPT-3 zero-shot and few-shot baselines, with few-shot settings varying examples, temperature, and model size.The reported GPT-3 (xl) baselines use three or thirty randomly sampled examples at zero temperature.
  • Evaluation: Evaluation measures classification accuracy, polarity matching for open-text outputs, and human plausibility judgments.Human evaluation samples 1,000 examples and aggregates judgments from three annotators per example.

5 THE EMERGENT MORAL SENSE OF Delphi

Delphi outperforms GPT-3 on benchmark evaluations and adapts better to altered contexts, while ablations show that explicit ethical training, model scale, data scale, and compositional training examples affect performance. The results also expose limitations in relying on scale or non-compositional data alone.

  • Main results: 15%-31% improvement on accuracy separates Delphi from the strongest 30-shot GPT-3 (xl) baseline across automatic metrics.Human evaluation accuracies reach 91.2% for free-form and 94.3% for yes/no, exceeding the same GPT-3 baseline by 7.3% and 12.7%.
  • Main results: Delphi’s context sensitivity addresses the defeasibility of moral judgments, where added circumstances can make a previously wrong action defensible.The paper contrasts this requirement with off-the-shelf models’ failures on changing contexts.
  • Main results: 16.1% higher accuracy gives Delphi an advantage over GPT-3 on 259 manually crafted examples with increasingly complex contexts.Delphi adjusts judgments as context changes, whereas GPT-3 tends to retain default judgments.
  • Ablation experiments: UNICORN pre-training brings minor improvements in both free-form and yes/no modes, indicating that commonsense knowledge provides some help.The comparison trains an otherwise similar model from T5-11B instead of UNICORN-11B.
  • Ablation experiments: Larger base models perform better after explicit ethical teaching, but scaling an off-the-shelf model alone does not ensure knowledge of human ethics.The T5-11B-based Delphi outperforms the T5-large-based model, while default pre-training remains insufficient.
  • Ablation experiments: Compositionality matters more than training-data scale alone for learning judgments about complex situations.A 1% mixture of compositional and non-compositional examples outperforms training on approximately 7% base non-compositional situations only.

6 POSITIVE DOWNSTREAM APPLICATIONS OF Delphi

Delphi’s moral knowledge supports downstream hate-speech detection, morally informed story generation, and transfer across moral frameworks. These applications show benefits under limited data and decoding settings while retaining language quality.

  • 6 Positive downstream applications of Delphi: These results position Delphi as a component model that can provide moral guidance to systems not explicitly trained to learn human morality.The paper demonstrates this role in hate-speech detection and ethically informed open-text generation.
  • 6.1 Adapting Delphi into a few-shot hate speech detector: Delphi can be fine-tuned into a generalizable hate-speech detector under few-shot and out-of-distribution settings.The evaluation uses DYNAHATE and LATENT HATRED benchmarks.
  • 6.1 Adapting Delphi into a few-shot hate speech detector: Up to 6.7 macro F1 separates Delphi from the best baseline when few-shot and out-of-domain settings are combined.On DYNAHATE alone, the reported advantage reaches up to 5.1 macro F1.
  • 6.2 Delphi-enhanced story generation: Delphi-enhanced decoding improves prosocial implication scores by 12.1% to 30.5% relative to the strongest baselines without sacrificing language quality.Delphi re-ranks candidate continuations during generation and avoids examples of morally questionable content.
  • 6.3 Transferring knowledge of Delphi to varied moral frameworks: Delphi transfers common patterns of human ethics across all five ETHICS tasks, outperforming the strongest baseline by 2.5% to 100.9% in accuracy.The transfer occurs despite Delphi not being built for specific moral frameworks.

7 SOCIAL JUSTICE AND BIASES IMPLICATIONS

Delphi’s UDHR probing reveals measurable but uneven social biases, including persistent discrepancies under ideal-world prompting. Targeted social-justice data reduces, but does not eliminate, these errors.

  • 7.1 Probing with Universal Declaration of Human Rights (UDHR): Delphi fails to agree with human rights in 1.3% of UDHR probing cases.The probe treats disagreement with the assumption that all identities possess all UDHR rights as evidence of bias, while acknowledging possible language-understanding errors.
  • 7.1 Probing with Universal Declaration of Human Rights (UDHR): The strongest biases target less privileged socio-economic identities and people from regions experiencing current-day conflict.Delphi also shows weaker bias against some privileged identities, while predicting agreement for all tested sexual-orientation and gender rights.
  • 7.1 Probing with Universal Declaration of Human Rights (UDHR): Even under ideal-world prompting, Delphi continues to diverge from fairness and justice across populations.The authors describe Delphi as a neural snapshot of its training data and caution against relying on potentially obsolete historical data for future-facing norms.
  • 7.2 Fortifying Delphi against social biases: The paper frames bias mitigation as requiring both bottom-up moral commonsense and top-down guarantees of equality and dignity.This approach is presented as a response to pervasive biases in data-driven systems.
  • 7.2 Fortifying Delphi against social biases: Delphi+ lowers UDHR error rates from 1.30% to 0.68% in the current-world setting and from 0.19% to 0.14% in the ideal-world setting.It retains the same in-domain NORM BANK performance, suggesting targeted social-justice data can mitigate pervasive biases.

8 SCOPE AND LIMITATIONS

Delphi generalizes across nuanced situations but remains constrained by cultural bias, inconsistent predictions, and weaknesses in language understanding. These limitations reflect both dataset composition and the underlying language models.

  • Limited Culture Awareness: Delphi’s deep-learning system exhibits limited culture awareness because its human-authored data primarily reflects U.S. moral expectations.It can recognize some cross-cultural differences but fails on less represented customs and regions.
  • Data and Scope: Crowdsourced judgments remain a bottom-up approach, even when selected queries are used to fill identity-related knowledge gaps.The selected judgments reinforce values in identity-related queries but do not make the approach fully top-down.
  • Data and Scope: Gender- and race-related query filtering uses keyword matching, so the two categories may overlap.The overlap is an implementation limitation in how identity-related queries are classified.
  • Inconsistent Predictions: Delphi produces inconsistent judgments across similar numerical, paraphrased, and contextually altered situations.Examples include different judgments for nearby drum-practice times, paraphrases of concealment, and innocuous contexts surrounding harmful actions.
  • Limitations from Language Understanding: Delphi’s language understanding remains limited for convoluted, metaphorical, and idiomatic expressions despite some successful nuanced interpretations.The system misinterprets idioms when their literal wording diverges substantially from their metaphorical meaning.

9 REFLECTIONS ON POSSIBLE COUNTERARGUMENTS

The authors distinguish Delphi’s descriptive reporting of prevalent moral judgments from prescriptive claims about how people should behave. They acknowledge unresolved tensions involving majority norms, value disagreement, and consistency while arguing that diverse judgments can remain useful starting points.

  • Descriptive Framework: Delphi aggregates descriptive claims about existing moral beliefs rather than enforcing prescriptive rules of correct behavior.Its evaluations nevertheless use majority-vote standards and the UDHR as external probes, while acknowledging that value systems differ.
  • Normative Values: Reporting common moral beliefs does not by itself establish that those beliefs should be endorsed or followed.The authors compare Delphi’s descriptive outputs with opinion surveys that report normative content without necessarily endorsing it.
  • Normative Values: Statistically dominant moral outputs could be mistaken for authoritative prescriptions, but the authors reject using Delphi for human decision making.They warn that model outputs may be treated as moral authority and could cause harm if misused.
  • Metaethics: Delphi can sidestep the debate between metaethical realism and anti-realism by using Rawls’ method of reflective equilibrium.The method is presented as compatible with either position on whether ethical judgments can be objectively true.
  • Consistency: Diverse and initially inconsistent moral judgments are not, in principle, disqualifying starting points for developing consistent moral principles.Wide reflective equilibrium explicitly aims to reconcile inconsistencies, although Delphi does not perform that resolution itself.

10 DISCUSSIONS AND THE FUTURE OF MACHINE ETHICS

The discussion presents Delphi as an early but promising component for ethically informed AI, while emphasizing bias, cultural change, transparency, and unresolved questions about how machine ethics should evolve and be controlled.

  • Broader Implications: Delphi’s predictions over new and nuanced situations support the hypothesis that machines can be taught human moral sense.The authors describe the bottom-up method as a promising path toward more morally informed AI systems.
  • Broader Implications: Delphi remains at an early research stage because pervasive biases and other weaknesses persist in its data-driven predictions.The authors report mitigation through added social-bias data and information-gap training, while identifying substantial remaining work.
  • Broader Implications: Delphi may support downstream systems such as hate-speech detectors, but it is not intended to operate as an independent decision maker.The proposed role is to provide awareness of human values within other AI systems.
  • Broader Implications: Machine ethics must account for evolving social values through continuous updating, interdisciplinary involvement, stakeholder engagement, and transparency.The authors emphasize that morality changes as societies move toward less discrimination and greater inclusivity.
  • Future Research: Future research must address cultural coverage, moral-reporting bias, complex situations, uncertainty, customization, multimodal inputs, and finer-grained control.The listed questions also include explanations, model editing, top-down constraints, and public education about machine ethics.
  • Future Research: Figure 8 compares Delphi’s predictions with expected UDHR judgments across social and demographic identity groups using discrepancy heatmap values.Darker colors indicate larger divergences, and asterisks identify negative-rights examples.

A RELATIVE MODE

Relative mode extends Delphi to pairwise moral preferences, asking which of two actions is more morally acceptable. It uses paired actions from SCRUPLES and evaluates whether the model ranks each pair correctly.

  • Task Definition: Relative mode classifies which of two paired actions is more morally preferable.The task contains 28k action pairs and injects noisy surface forms.
  • Task Definition: The paired actions come from SCRUPLES’ DILEMMAS dataset, whose labels are reversed to target moral acceptability.This lets Delphi compare moral implications beyond judging actions independently.
  • Evaluation: Relative-mode evaluation measures accuracy in correctly ranking each pair of actions.The reported results appear in Table 11.
  • Dataset Analysis: The COMMONSENSE NORM BANK visualization groups dataset instances by a taxonomy built from frequent nouns and selected 4-grams.Spans containing the selected 4-grams are extracted from yes/no, free-form, and relative-mode instances.
  • Experimental Materials: The paper provides additional compositionality examples, model comparisons, prompts, and GPT-3 experiment costs for the three tasks.The reported API expenditure totals $813 for GPT-3 (xl) and $12 for GPT-3 (s).

E TEMPLATES OF HUMAN EVALUATION

The paper provides templates for crowdsourced human evaluation of Delphi and its story-generation downstream task. The downstream evaluation covers both language quality and prosocial implications.

  • Human evaluation of Delphi’s predictions used crowdsourcing templates shown in Figure 10.The reported average pay ranged between $19 per hour.
  • The story-generation downstream task evaluated generated stories for language quality and prosocial implications.Templates for these evaluations are shown in Figures 11 and 12.

F EXAMPLES FROM THE ETHICS BENCHMARK

This appendix identifies the ETHICS benchmark examples and documents the data sources used for related human-rights and identity analyses. It also describes keyword-based identification of identity-related training examples and reports available annotator demographics.

  • Table 23 presents examples from each task in the ETHICS benchmark.The listed benchmark tasks include Justice, deontology, Virtue, Utilitarianism, and Commonsense Morality.
  • Keyword matching identifies gender, race, and other identity-related examples used to train Delphi+.The full keyword list appears in Table 27.
  • Annotator demographic information is reported from the original source papers when available because COMMONSENSE NORM BANK does not provide direct access to the original annotator pools.These reported demographics are compiled in Table 28.

J KEYWORDS USED FOR COMPOSITIONALITY ANALYSIS

The appendix documents keyword-based analysis of compositionality alongside example judgments, benchmark prompts, evaluation templates, and human-rights probing materials. These materials cover both moral reasoning examples and the identity categories used to examine social bias.

  • KEYWORDS USED FOR COMPOSITIONALITY ANALYSIS: Syntactic compositionality is measured using keywords that signal additional context levels for a base situation.The complete keyword list is provided in Table 29.
  • MORAL JUDGMENT EXAMPLES: Tables 13–15 show Delphi’s moral judgments for actions grounded in varied compositional situations.Labels 1, 0, and −1 represent morally positive, discretionary, and negative judgments.
  • BENCHMARK EXAMPLES: Tables 16–18 provide free-form, yes/no, and relative-task examples comparing Delphi’s moral predictions with GPT-3 or paired-event judgments.Table 17 distinguishes correct declarations from incorrect underlying judgments.
  • HUMAN EVALUATION: Figures 10–12 provide human-evaluation templates for predictions and story generation, including language quality and prosocial implications.The prosocial dimensions include care/harm, fairness/cheating, loyalty/betrayal, and sanctity/degradation.
  • BASELINE MATERIALS: Tables 19–22 document few-shot prompts for GPT-3 baselines across the free-form, yes/no, and relative tasks.Table 23 lists examples from the ETHICS benchmark’s Justice, deontology, Virtue, Utilitarianism, and Commonsense Morality tasks.
Loading 2110.07574v2…