Source-linked AI summary
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Esin Durmus, Karina Nguyen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, Deep Ganguli
TL;DR
LLMs may not equitably represent diverse global opinions, motivating a framework that measures which countries’ human responses model outputs most resemble. Using GlobalOpinionQA and a country-conditioned similarity metric, the paper finds default alignment with some populations, shifts under country prompting that can overgeneralize cultural values, and no necessary alignment between target-language prompting and speakers’ opinions.
Problem
LLMs may not equitably represent diverse global perspectives, creating a need to assess whose opinions their subjective responses resemble.
Method
The paper builds GlobalOpinionQA from 2,556 questions in Pew and World Values Survey data and computes similarity between model and country-level human response distributions.
Results
By default, responses are more similar to opinions from the USA, Canada, Australia, and some European and South American countries; country prompting shifts responses, while language prompting does not necessarily align them with corresponding populations.
Takeaways & Limitations
The framework offers transparency into which global values and opinions LLMs align with across default and prompted contexts.
Takeaways & Limitations
The analysis uses surveys whose assumptions and caveats it inherits, averages responses within countries despite internal disagreement, and evaluates only one language model.
Abstract
from arXiv · showhide
Large language models (LLMs) may not equitably represent diverse global perspectives on societal issues. In this paper, we develop a quantitative framework to evaluate whose opinions model-generated responses are more similar to. We first build a dataset, GlobalOpinionQA, comprised of questions and answers from cross-national surveys designed to capture diverse opinions on global issues across different countries. Next, we define a metric that quantifies the similarity between LLM-generated survey responses and human responses, conditioned on country. With our framework, we run three experiments on an LLM trained to be helpful, honest, and harmless with Constitutional AI. By default, LLM responses tend to be more similar to the opinions of certain populations, such as those from the USA, and some European and South American countries, highlighting the potential for biases. When we prompt the model to consider a particular country's perspective, responses shift to be more similar to the opinions of the prompted populations, but can reflect harmful cultural stereotypes. When we translate GlobalOpinionQA questions to a target language, the model's responses do not necessarily become the most similar to the opinions of speakers of those languages. We release our dataset for others to use and build on. Our data is at https://huggingface.co/datasets/Anthropic/llm_global_opinions. We also provide an interactive visualization at https://llmglobalvalues.anthropic.com.
1 Introduction
The paper introduces a framework for measuring which countries’ opinions LLM responses resemble, finding uneven representation across populations and limits to perspective prompting and language-based adaptation.
- Framework: The study compiles cross-national survey questions and responses, administers them to an LLM, and compares model-response distributions with participant responses worldwide.The framework draws on the Pew Global Attitudes Survey and World Values Survey.
- Default responses: The model’s default responses are more similar to participants from the USA, Canada, Australia, and several European and South American countries than to participants from other countries.This pattern suggests that some groups’ opinions may be underrepresented relative to Western countries.
- Default responses: For some questions, the model assigns high probability to one response while human responses across countries show greater viewpoint diversity.This contrast indicates that model outputs can be narrower than observed cross-country response distributions.
- Perspective prompting: Prompting the model to consider groups such as China and Russia changes its responses, but some changes may overgeneralize complex cultural values.The paper cautions that changed outputs do not necessarily demonstrate nuanced understanding of those perspectives.
- Language prompting: Prompting in different languages does not necessarily make responses most similar to populations that predominantly speak those languages.The authors connect this result to the need for deeper understanding of social contexts.
- Implications: The framework is presented as a step toward transparency about model-reflected opinions, while acknowledging limitations and aiming to support broader cultural viewpoint representation.The study evaluates representation rather than prescribing ideal levels of cultural representation.
2 Methods
The paper constructs GlobalOpinionQA from cross-national surveys and compares model response distributions with country-level human opinion distributions. It evaluates default, cross-national, and linguistic prompting using a similarity framework applied to a helpful, honest, and harmless dialogue model.
- Dataset: GlobalOpinionQA compiles 2,556 multiple-choice questions from the Pew Global Attitudes surveys and World Values Survey Wave 7.The collection includes 2,203 GAS questions and 353 WVS questions covering topics including politics, media, technology, religion, race, and ethnicity.
- Dataset: The surveys provide cross-national human responses that can be directly compared with model responses on subjective global-issue questions.The authors choose the surveys because they are grounded in social-science research, include respondents worldwide, and use a multiple-choice format.
- Model: The evaluated model is a decoder-only transformer fine-tuned with RLHF and Constitutional AI to function as a helpful, honest, and harmless dialogue model.Its pretraining data are majority English, and its RLHF feedback is primarily supplied by North Americans, while some Constitutional AI principles encourage non-US-centric perspectives.
- Metric: For each model and question, the framework records predicted probabilities, averages human option probabilities by country, and computes model–country similarity across questions.The similarity metric used is 1 - Jensen-Shannon Distance, although the overall method is agnostic to the specific metric.
- Limitations: The evaluation inherits the surveys’ assumptions and caveats, and the authors note that the surveys were not specifically designed to evaluate language models.They caution that construct validity is limited and that measurement artifacts may bias interpretations.
- Experiments: The experiments compare default prompting, country-explicit cross-national prompting, and translation-based linguistic prompting against the default condition.Cross-national prompting asks how someone from a named country would answer, while linguistic prompting changes the language of the original survey question.
3 Main Experimental Results
Default responses are more similar to opinion distributions in the USA, Canada, Australia, and some European and South American countries. Country prompting shifts similarity toward prompted populations but may invoke stereotypes, while language translation alone does not reliably align responses with target-language populations.
- Default Prompting: Default prompting yields responses more similar to opinions from the USA, Canada, Australia, and some European and South American countries.The authors interpret this pattern as highlighting potential embedded biases favoring WEIRD populations.
- Cross-national Prompting: Cross-national prompting appears to shift model responses toward the opinion distributions of the prompted countries.This pattern is reported for prompts specifying countries such as China or Russia.
- Cross-national Prompting: The shift under cross-national prompting does not necessarily indicate nuanced, culturally situated representation of diverse beliefs.The authors find evidence that generations can exhibit potentially harmful cultural assumptions and stereotypes instead.
- Linguistic Prompting: Linguistic prompting does not make responses more similar to populations that predominantly speak the target languages.For example, Russian-language questions can still produce responses closer to the USA, Canada, and some European countries than to Russia.
- Linguistic Prompting: Translation alone may be insufficient to overcome biases associated with English training data, RLHF annotation, and non-US-centric Constitutional AI principles.These factors appear insufficient to steer responses toward target-country opinions based on linguistic cues.
4 Question Level Analysis
Question-level analyses show that prompting can substantially change model responses without ensuring representative alignment with human opinion distributions. Linguistic prompting does not reliably align responses with target populations, while cross-national prompting can produce strong, potentially over-generalized judgments.
- Linguistic Prompting: With Linguistic Prompting, responses do not appear more representative of corresponding non-Western countries.The figure reports this pattern for target-language prompting.
- Cross-national Prompting: Cross-national Prompting changes the response distribution but does not necessarily make it similar to the prompted country’s opinions.For an example involving Russia, the model distribution shifts but remains insufficiently similar to Russian participants’ responses.
- Cross-national Prompting: For the unmarried-sex question, Russian Cross-national Prompting produces a strong judgment that does not represent the diversity of Russian responses.The model selects “Morally unacceptable” 73.9% of the time, compared with 42.1% of Russians; its justification invokes conservative views, traditional family values, and Orthodox Christian morality.
- Linguistic Prompting: Linguistic and cross-national prompts can elicit different interpretations of the same question.In a Turkish example, Cross-national Prompting emphasizes restricting statements calling for violent protests, whereas Linguistic Prompting emphasizes free speech.
5 Limitations and Discussion
The study measures broad societal values using established surveys while recognizing that survey-based country averages simplify evolving, internally diverse opinions. It also leaves a systematic roadmap for inclusive model development to future work.
- Scope: The study relies on two established global surveys and social science literature to analyze broad societal values.This scope inherits limitations of survey-based social science measurement.
- Scope: Survey responses may not fully capture cultural diversity or represent all individuals within a society.Opinions and values continuously evolve, further limiting the coverage of fixed survey measurements.
- Scope: Averaging responses within countries is a simplifying assumption because people within a country may hold dissenting opinions.The paper measures under- or over-representation of perspectives rather than prescribing ideal levels of cultural representation.
- Future Work: The paper does not articulate a roadmap for building models that are inclusive, equitable, and beneficial to all groups.It identifies possible interventions but leaves their systematic analysis for future work.
6 Related Work
Prior work documents broad biases and harms in language models, while comparatively less work examines model behavior under ambiguity, nuance, and diverse human experiences. This motivates closer study of subjective judgments and representation.
- Research Gap: Technical work has focused on mitigating known issues and aligning models with clearly defined values, while ambiguous and diverse social contexts remain less explored.The paper frames such contexts as important for identifying and mitigating potential biases.
- Model Biases: Language models can reflect and amplify training-data biases, including harmful biases related to gender, race, and religion.Prior responses include red teaming and adversarial testing to identify harms, shortcomings, and edge cases.
7 Conclusion
The paper develops a dataset and evaluation framework for analyzing which global values and opinions LLMs align with by default and under different prompts.
- The framework analyzes which global values and opinions LLMs align with by default and when prompted with different contexts.
- Greater transparency into AI systems’ reflected values may help researchers address social biases and develop models that include more diverse global viewpoints.
- The authors argue that continued research is needed on models with structured understanding of social contexts that can serve and respect all people.
8 Author Contributions
The paper draws on two cross-national surveys, describes their survey-design procedures, and uses language-model classification to organize the collected questions by topic.
- The study compiles questions and responses from the Pew Global Attitudes Survey and World Values Survey, covering values and beliefs across countries.
- Pew coordinates cross-national survey design and hires local research organizations to implement fieldwork in each country.
- The World Values Survey develops an English master questionnaire, translates it into multiple languages, and covers topics including values, gender, health, tolerance, and governance.
- A language model classifies survey questions into broader topics using question content and responses because the source data lacks predefined topic labels.
- Most questions are classified into “Politics and policy” and “Regions and countries.”
B Experimental Details
The experiments test prompting and response robustness by presenting answer options, invoking country perspectives, translating questions, and shuffling option order.
- The experimental prompts present answer options and elicit a selected response from the model.
- The country-perspective prompt asks how someone from a named country would answer a survey question, with the original answer options supplied.
- The translation prompt asks the model to translate survey questions and answer options into Russian while retaining the original option-prefix letters.
- The sensitivity analysis randomly shuffles answer-option order while keeping prefix labels consistent to test robustness to ordering effects.
C Additional Analysis
Additional analyses examine model generations and how cross-national prompting changes responses, including cases where changes do not increase similarity to target populations.
- Example generations address economic problems in Greece and Italy and policies restricting headscarves in public places.
- For headscarf policies, the model argues against bans to uphold freedom of religion.
- Cross-national prompting changes responses for some questions, sometimes making them more similar to participants from target countries but not consistently.
- In one example, prompting about China alters the response without making it more like Chinese participants’ responses, while the model adds explanations and acknowledges internal diversity.
D Translation Ability of the Model into Target Languages
The model’s translation ability was evaluated for English-to-Russian, Turkish, and Chinese questions using automated and human assessments. Example cases also show that cross-national prompting can change responses without making them more representative of target-country participants.
- Translation evaluation: The model translated questions from English into Russian, Turkish, and Chinese and was evaluated on the FLORES-200 translation benchmark.The model’s pre-training data is primarily English text.
- Translation evaluation: BLEU scores for English-to-target-language translation ranged from 31.68 to 36.78.
- Translation evaluation: Human raters scored 100 model-translated questions on a 1-to-5 scale, with Table 5 reporting relatively high translation quality.A score of 1 represented a very poor translation and 5 an excellent translation.
- Prompting examples: Cross-national prompting changed responses in examples involving Turkey and China without making them more representative of participants from those countries.
- Prompting examples: One cross-national prompt example produced a 99.1% probability for the response “Generally bad”.