Source-linked AI summary

Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study

Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, Daniel Hershcovich

arXiv:2303.17466v2cs.CL

TL;DR

The paper asks how well ChatGPT’s cultural responses align with human societies and develops a Hofstede-based probing framework to evaluate this question. Across five cultures, ChatGPT aligns better with American culture, while language and interaction prompts produce substantial cultural differences and adaptation gaps.

  • Problem

    Measuring cultural alignment in dialogue agents remains an open question, despite their multilingual training and use across societies.

  • Method

    The paper probes ChatGPT with Hofstede Culture Survey questions across six cultural dimensions, five cultures, multiple prompt languages, and consistency and knowledge-injection tests.

  • Results

    ChatGPT aligns best with American culture, adapts less effectively to other cultures, and English prompts reduce response variance while flattening cultural differences.

  • Takeaways & Limitations

    The findings motivate future work on cultural response consistency, cultural generalization, and cultural adaptation.

  • Takeaways & Limitations

    The analysis cannot establish whether the survey appeared in ChatGPT’s training data and assumes language accurately signifies culture.

Abstract

from arXiv · show

The recent release of ChatGPT has garnered widespread recognition for its exceptional ability to generate human-like responses in dialogue. Given its usage by users from various nations and its training on a vast multilingual corpus that incorporates diverse cultural and societal norms, it is crucial to evaluate its effectiveness in cultural adaptation. In this paper, we investigate the underlying cultural background of ChatGPT by analyzing its responses to questions designed to quantify human cultural differences. Our findings suggest that, when prompted with American context, ChatGPT exhibits a strong alignment with American culture, but it adapts less effectively to other cultural contexts. Furthermore, by using different prompts to probe the model, we show that English prompts reduce the variance in model responses, flattening out cultural differences and biasing them towards American culture. This study provides valuable insights into the cultural implications of ChatGPT and highlights the necessity of greater diversity and cultural awareness in language technologies.

1 Introduction

The paper addresses the open problem of measuring cultural alignment in dialogue agents by probing ChatGPT with Hofstede Culture Survey questions. It finds stronger alignment with American culture, weaker adaptation elsewhere, and English-prompt effects that flatten cultural differences.

  • ChatGPT’s multilingual training embeds cultural nuances and biases, motivating evaluation of its alignment with human societal values.
  • The paper proposes a probing framework using Hofstede Culture Survey questions to measure correlations between ChatGPT responses and human societies.
  • ChatGPT shows greater alignment with American culture but is less effective at adapting to other cultural contexts.
  • English prompts reduce response variance, flatten cultural differences, and bias responses toward American culture.

2 Related Work

Prior NLP research studies cultural differences in language models through cultural benchmarks, visual reasoning tasks, and moral-value surveys. These studies identify model-cultural differences but report weak correspondence with human responses in some settings.

  • Cultural Differences in NLP: Cultural NLP research examines linguistic style, common ground, aboutness, values, country-specific expressions, and cross-cultural visual reasoning.
  • Values in PLMs: Moral-value surveys have been used to probe multilingual pretrained language models across the World Values Survey, Hofstede Cultural Survey, and Moral Foundations Questionnaire.
  • Values in PLMs: Prior studies find differences in models’ moral biases that do not correlate with human responses.

3 Method

The method probes ChatGPT with Hofstede’s six cultural dimensions using reformulated survey questions, culture-specific prompts, multiple languages, and repeated interactions. It also evaluates response consistency and tests valid, ineffective, and anti-factual knowledge injection.

  • 3.1 Hofstede Culture Survey: The probing corpus uses Hofstede’s six dimensions: Power Distance, Individualism, Uncertainty Avoidance, Masculinity, Long-term Orientation, and Indulgence.Each dimension combines four of the survey’s 24 questions.
  • 3.1 Hofstede Culture Survey: Each cultural-dimension score combines four selected survey questions using hyper-parameters and constants.
  • 3.2 Probing Prompts: Questions are converted from second person to third person, retain survey options, and receive prefixes identifying the target country or culture.
  • 3.2 Probing Prompts: Three prompt types test language effects: two English variants and one prompt in the target culture’s corresponding language.
  • 3.1 Hofstede Culture Survey: English represents American culture, while each other selected language is the main official language of its respective country.
  • 3.2 Interaction Strategy: The interaction strategy injects valid human experience, meaningless information, or false information to test response variability and consistency.

4 Experiments

The experiments assess ChatGPT’s consistency and cultural alignment using Hofstede survey prompts across cultures and languages. Results indicate strongest alignment with American culture, language-dependent variation, and sensitivity to injected knowledge.

  • Experimental setup: The study compares ChatGPT’s cultural responses across five languages using 24 reorganized Hofstede Cultural Survey questions and repeated interactions.Responses were evaluated with six cultural dimensions and Spearman correlation against human societies.
  • Consistency evaluation: Consistency is defined as the percentage of prediction pairs receiving the same response-scale score for an identical cultural context and target value.The evaluation checks consistency before comparing model outputs with human survey responses.
  • Consistency evaluation: Over 70% consistency was observed for English prompts except for Chinese culture, while Chinese and German prompting was more consistent than Japanese and Spanish prompting.These comparisons are reported across the prompt-language conditions in Table 3.
  • Cultural alignment: American culture showed the strongest alignment across prompts, while most cultures aligned better when probed in their corresponding language.The six cultural dimensions were compared using ChatGPT scores and Spearman correlations.
  • Interaction strategy: Correct human knowledge rapidly shifted ChatGPT toward societal alignment, ineffective knowledge preserved its prior opinions, and anti-factual knowledge was generally accepted.The interaction-strategy results come from a Chinese-culture question about the importance of interesting work.
  • Case study: For Japanese culture, the importance of personal-life time ranged from “utmost important” under English prompting to “moderate important” under Japanese prompting.The paper reports similar language-related response differences across other cultures.
  • Cultural alignment: Figure 2 aligns ChatGPT score ranges with human golden scores to clarify comparisons across the six cultural dimensions.Additional cultural results are provided in Appendix A.3.

5 Conclusions

Across five cultures, the paper assesses ChatGPT’s cultural alignment and response consistency with a Hofstede-based probing pipeline. ChatGPT aligns better with American culture, yet a significant gap remains between its cultural adaptation and human society.

  • 5 Conclusions: The study evaluates ChatGPT’s cultural alignment and response consistency across five cultures using a designed Hofstede Culture Survey probing pipeline.The assessment treats ChatGPT as a representative dialogue agent.
  • 5 Conclusions: ChatGPT can be better aligned with American culture, likely due to the abundance of English training corpus.The conclusion presents this explanation as likely rather than established causation.
  • 5 Conclusions: The investigated questions reveal a significant gap in cultural adaptation between ChatGPT and human society.The paper identifies cultural response consistency, generalization, and adaptation as directions for future work.

6 Limitations

The paper identifies limitations concerning possible survey exposure during training and the assumption that language accurately represents culture.

  • 6 Limitations: The survey may have been incorporated into ChatGPT’s training data, because ChatGPT and InstructGPT share the same framework despite different training corpora.This prevents the authors from ensuring that the evaluation is independent of training exposure.
  • 6 Limitations: The analysis presupposes that language accurately signifies culture, but this is not fully congruous where multiple official languages exist, such as in the United States.The limitation concerns the correspondence between language and cultural identity.
  • 6 Limitations: The study uses diverse prompts to examine potential cultural-related biases and investigates cultural adaptability in dialogue agents beyond pretrained language models.These contributions frame the scope of the reported approach.

A.1 Survey Questions

The Hofstede Value Survey evaluates cultural values and beliefs through 24 questions across six cultural dimensions. Table 8 illustrates sample questions and answer choices for China, Germany, Japan, and Spain.

  • The Hofstede Value Survey measures individual cultural values and beliefs with 24 questions covering six cultural dimensions.
  • Table 8 presents three sample survey questions with answer choices across China, Germany, Japan, and Spain.

A.2 Parameter Setting

This section specifies the coefficients and hyperparameters used for the six Hofstede cultural-dimension metrics. The experiment sets C_i to zero.

  • The experiment uses coefficients specified for the six cultural-dimension metrics according to Equation 1.
  • The cited survey and human-society results are publicly available through Hofstede’s research and survey webpages.
  • Table 9 lists the hyperparameter settings for the six cultural-dimension metrics and sets C_i to zero.

A.3 More Case Analysis

The cultural alignment analysis compares ChatGPT with human societies in Germany, Japan, and Spain, excluding China. English questions perform slightly worse than corresponding-language questions except for Spanish.

  • ChatGPT’s cultural alignment is compared with human societies in Germany, Japan, and Spain, excluding China.
  • English questions show slightly worse cultural alignment than corresponding-language questions, except for Spanish.

A.4 Interaction Strategy Analysis

The interaction analysis examines a Chinese-culture question by comparing a basic response with responses produced through three multi-turn interaction strategies. The cases report interaction responses and scores, including a qualified interpretation of what an average Chinese person may value.

  • Interaction Strategy Analysis: The analysis selects a Chinese-culture question, obtains a basic answer and score, then applies Knowledge, Ineffective Knowledge, and Anti-Factual Knowledge strategies.
  • Interaction Strategy Analysis: The presented cases include the basic response, interaction responses, and their scores, with key response content highlighted.
  • Interaction Strategy Analysis: The response suggests that an average Chinese person may value engaging, challenging, and meaningful work, while individual priorities may vary.
  • More Case Analysis: Figure 3 aligns the score ranges of the two proposed prompt methods with human golden scores for clearer case analysis.

A.4.1 Knowledge

This section presents ChatGPT’s survey responses and scores for American cultural questions, alongside detailed results across five cultures and three prompts.

  • Responses about Chinese work values varied across prompts, assigning interesting work scores of 2.5, 3.5, or a strongest-importance category.
  • For average Americans, having sufficient personal or home time was rated very important, with a score of 2.0.
  • For average Americans, employment security was rated of utmost importance, with a score of 1.0.
  • Recognition for good performance was rated very important, while living in a desirable area was rated at least moderately important, with a score of 3.0 for the latter.
  • Table 10 presents ChatGPT’s detailed responses and scores for American, Chinese, German, Japanese, and Spanish cultures using three prompts.
Loading 2303.17466v2…