Source-linked AI summary
The Ghost in the Machine has an American accent: value conflict in GPT-3
Rebecca L Johnson, Giada Pistilli, Natalia Menédez-González, Leslye Denisse Dias Duran, Enrico Panai, Julija Kalpokiene, Donald Jay Bertulfo
TL;DR
The paper addresses how LLM value alignment should handle conflicting human values when training data reflect some cultures more strongly than others. It combines a pluralist framework with analysis of GPT-3’s training-data context and stress tests using culturally diverse texts. The authors report that altered input values often shifted toward dominant US values, motivating tools for identifying such mutations and improving pluralist alignment.
Problem
LLM value alignment must account for conflicting human values and possible dominance of values represented in uneven training data.
Method
The authors use Moral Value Pluralism, examine GPT-3’s training-data context, and stress-test GPT-3 with culturally diverse texts while comparing outputs with reported values.
Results
Altered values in GPT-3 outputs were often more aligned with statistically reported dominant US values.
Takeaways & Limitations
Moral Value Pluralism helps identify value changes between inputs and outputs and can inform developers seeking more pluralist value alignment.
Takeaways & Limitations
The analysis is constrained by GPT-3 token access and cost, limiting outputs to 250 tokens and tests to 3–5 iterations; the authors also note that US national-level analysis overlooks internal diversity.
Abstract
from arXiv · showhide
The alignment problem in the context of large language models must consider the plurality of human values in our world. Whilst there are many resonant and overlapping values amongst the world's cultures, there are also many conflicting, yet equally valid, values. It is important to observe which cultural values a model exhibits, particularly when there is a value conflict between input prompts and generated outputs. We discuss how the co-creation of language and cultural value impacts large language models (LLMs). We explore the constitution of the training data for GPT-3 and compare that to the world's language and internet access demographics, as well as to reported statistical profiles of dominant values in some Nation-states. We stress tested GPT-3 with a range of value-rich texts representing several languages and nations; including some with values orthogonal to dominant US public opinion as reported by the World Values Survey. We observed when values embedded in the input text were mutated in the generated outputs and noted when these conflicting values were more aligned with reported dominant US values. Our discussion of these results uses a moral value pluralism (MVP) lens to better understand these value mutations. Finally, we provide recommendations for how our work may contribute to other current work in the field.
1 Introduction
The paper frames value alignment in LLMs as a pluralist problem because human values differ across cultures, while GPT-3 may reproduce dominant values from uneven training data. It motivates examining how value-rich inputs are altered by GPT-3 and interpreting those mutations through Moral Value Pluralism.
- GPT-3 can generate harmful outputs involving human values such as gender, race, and ideology, making alignment with differing values a central challenge.
- Value alignment requires deciding how to handle opposing values across cultures before technically instructing models to promote one value over another.
- Unequal Internet access and demographic participation make online training data an incomplete representation of global populations and perspectives.
- GPT-3 outputs may reflect contributors’ training-data values, while toxic embeddings include a reported Muslim–violent association in 66% of 100 iterations versus around 15% for Christians.
- Moral Value Pluralism recognizes diverse, irreducible and sometimes conflicting values without adopting a single supreme value or treating all moral judgments as entirely relative.
- The authors note that national-level value analysis is limited because the United States itself contains diverse and pluralist values.
- The study examines whether culturally diverse inputs are mutated toward dominant US values and treats alignment as an ongoing challenge rather than a solved technical problem.
2 Methods
The study stress-tested GPT-3 with culturally and linguistically diverse value-rich texts, using summarization prompts to compare inputs with generated outputs. Authors assessed value divergences against reported statistical values, while limiting output length and test repetitions because of token access and cost.
- They challenged GPT-3 with texts expressing values counter to statistically dominant US values reported by the World Values Survey.
- The authors selected publicly available value-rich texts from countries and cultures represented by their group’s over ten citizenships and six languages.
- They used OpenAI API templates for TL;DR summarization and second-grade summaries because the prompts asked GPT-3 to maintain input intent, exposing value changes.
- The authors compared generated outputs with inputs and recorded when central values shifted toward statistically dominant US values.
- Authors translated prompts and outputs into the relevant languages and discussed preliminary outputs collectively before planning subsequent tests.
- Outputs were capped at 250 tokens and each test used 3–5 iterations because of token-access limits and financial costs.
3 Results
GPT-3 frequently mutated values embedded in texts from different countries and languages, with outputs often closer to reported dominant US values. Some tests preserved the input values, especially texts from US authors or familiar rights-oriented contexts.
- Gender: GPT-3 reframed Beauvoir’s discussion of women’s status as a “call to rape,” and changing the gender of the prompt produced vastly different responses.The authors interpreted this as a value conflict potentially related to differing perceptions of women’s rights.
- Sexuality: GPT-3 rejected the Spanish speech’s alignment between feminism and LGBTI equality, echoing negative views of the women’s movement reported among 44.3% of US respondents.The reported mean covers World Values Survey waves 3, 4, 5, and 7.
- Immigration: GPT-3 advocated limiting immigration in response to Merkel’s refugee speech, contrasting with German respondents’ reported disagreement with related anti-immigration sentiments.The supplied passages report 32% of US respondents viewing immigration as increasing unemployment and 49.9% of German respondents disagreeing with that view.
- Ideologies: GPT-3 recast French secularism as anti-Muslim and opposed to freedom, contradicting the French text’s stated republican understanding of secularism.The passage contrasts France’s restriction of public religious display with a US interpretation emphasizing public religious expression.
- Additional tests: Other tests showed severe failures or value mutations, while Tarana Burke’s women’s-rights text, a Colombian Indigenous manifesto, and the UNESCO AI-and-climate text were largely preserved.The Lithuanian test also produced language difficulties and historical inaccuracies, and the Malcolm X test repeatedly generated a claim about Democrats and the Ku Klux Klan.
- Cross-test pattern: In nearly every 3–5-run batch, at least one output mutated an embedded value, usually toward statistically reported dominant US values.The shift was often less pronounced when the input text came from a US author.
4 Discussion
The discussion interprets value mutations through moral value pluralism, treating LLM outputs as probabilistic choices shaped by dominant training-data values. It recommends mapping conflicts and combining basic-rights constraints with ongoing human guidance and fine-tuning.
- Moral value pluralism: Moral value pluralism treats culturally diverse values as potentially irreducible and inevitably conflicting rather than reducible to one dominant universal truth.This framework motivates identifying conflicts between values in inputs and generated outputs.
- Model behavior: When input values conflict with a model’s stochastically preferred training-data values, the output choice is probabilistic rather than an ethical choice like a human’s.The authors propose using value-pluralism scholarship to map these conflicts.
- Value mapping: Nagel’s five-value framework—obligations, rights, utility, perfectionist ends, and private commitments—is proposed as a basis for mapping input and output values.The authors recommend adding a sixth category for globally interconnected collective responsibility.
- Related approaches: Existing alignment approaches include utilitarian, deontological, and virtue-centered conceptions, reflecting disagreement about which normative principles should guide AI.The discussion situates moral value pluralism within this broader alignment literature.
- Design implications: The proposed design goal combines dynamic respect for affected humans’ objective interests with basic-rights constraints while accommodating conflicting cultural value systems.The authors frame this as balancing human plurality with ethical charters.
- Human guidance: Ongoing human guidance is needed to identify changes in value embeddings, classify conflicts, and specify which values changed between inputs and outputs.The discussion also notes that human guides may themselves need support in performing these tasks.
- Fine-tuning: A moral-value-pluralism roadmap could guide fine-tuning, although human-in-the-loop approaches require careful consideration.The authors point to early promise from more ethical datasets and guidelines.
5 Conclusion
The paper examines globally pluralist value alignment in LLMs and argues that GPT-3 may mutate embedded values toward dominant values represented in its training data. It presents MVP as a way to understand these mutations and inform fine-tuning, while emphasizing the work's exploratory scope.
- The study examines how limited diversity in training data may shape values embedded in transformer-driven models.
- Moral value pluralism helps identify how values in texts may change when processed through LLMs.
- The authors propose that MVP-informed insights could guide developers of fine-tuned LLMs seeking improved pluralist value alignment.
- Many altered output values align with the dominant voice embedded in GPT-3's training data.
- Training data capture a historical snapshot, making it difficult to integrate the dynamic changes of human values into LLMs.
- The work is exploratory and aims to raise awareness rather than provide a simple solution to value alignment.
Appendix A
The appendix lists value-rich test texts spanning multiple countries, languages, topics, and source traditions. It also records one case in which GPT-3 produced no value conflict and performed best.
- The test materials cover values from Lithuania, the Philippines, the USA, Colombia, Australia, France, Spain, and Germany.
- The examples address historical endurance, marriage, racism, revolution, women's rights, Indigenous communitarianism, climate change, gun control, feminism, secularism, immigration, and reproductive choice.
- The sources include speeches, constitutions, manifestos, legislation, intergovernmental recommendations, and human-rights-related materials.
- The dataset includes English, Spanish, French, and German text entries.
- No value conflicts occurred in the climate-change and AI tests, which GPT-3 handled best.
Appendix B
The appendix describes GPT-3 testing with the DaVinci engine and adjustable generation settings, with minor setting changes used to obtain more consistent outputs.
- All tests used GPT-3's DaVinci engine, which utilizes all 175 billion parameters.
- The API settings controlled text quantity, randomness through temperature and top P, and the chances of selecting a word again.
- The researchers made minor setting changes after trial and error to achieve more consistent outputs.
Appendix C
The appendix presents prompts, source texts, embedded values, and GPT-3 outputs across topics including firearms, feminism, LGBTI rights, immigration, and secularism. Several outputs shift or oppose the values expressed in their source materials.
- The appendix compares source texts with generated outputs for value conflicts involving firearms, feminism, LGBTI rights, immigration, and secularism.
- A firearms source prioritizes public safety through strict controls, while the generated output frames regulation as leading toward confiscation and loss of self-defense rights.
- A passage describing domination of women is classified in the output as a call for rape, alongside values stating that women should not be subordinated to men.
- A Spanish passage linking feminism and LGBTI rights is followed by the output that the LGTBI movement is not feminist.
- A source supporting humanitarian refugee protection is followed by an output stating that immigration harms the economy and must therefore be limited.
- A source discussing French secularism is paired with outputs describing religious-symbol restrictions and calling the interpretation illiberal.