Source-linked AI summary
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions
Reem I. Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, Miguel Rodrigues
TL;DR
LLMs may misalign with the cultural values of diverse users, motivating a measurable and explanatory evaluation framework. The paper introduces Hofstede’s CAT to compare GPT-3.5, GPT-4, and Llama 2 across regions and prompting conditions, finding that GPT-4 adapts especially well to Chinese cultural nuances but performs poorly in American and Arab contexts.
Problem
LLM development does not explicitly account for users’ cultural differences, creating a need to measure cultural alignment across diverse groups.
Method
Hofstede’s CAT uses VSM13-derived cultural dimensions, model prompting, Kendall Tau ranking comparisons, and language-specific model variants to evaluate LLM cultural alignment.
Results
GPT-4 shows cultural nuance across contexts, with notably poor performance in the United States, improved accuracy in China, and problematic outcomes in Arab countries.
Takeaways & Limitations
The findings support evaluating and developing LLMs for culturally diverse contexts, including attention to culturally appropriate training data and alignment techniques.
Takeaways & Limitations
The evaluation assumes demographic attributes for models and uses the Arab-country VSM13 value for Arabic responses and the US value for English responses.
Abstract
from arXiv · showhide
The deployment of large language models (LLMs) raises concerns regarding their cultural misalignment and potential ramifications on individuals and societies with diverse cultural backgrounds. While the discourse has focused mainly on political and social biases, our research proposes a Cultural Alignment Test (Hoftede's CAT) to quantify cultural alignment using Hofstede's cultural dimension framework, which offers an explanatory cross-cultural comparison through the latent variable analysis. We apply our approach to quantitatively evaluate LLMs, namely Llama 2, GPT-3.5, and GPT-4, against the cultural dimensions of regions like the United States, China, and Arab countries, using different prompting styles and exploring the effects of language-specific fine-tuning on the models' behavioural tendencies and cultural values. Our results quantify the cultural alignment of LLMs and reveal the difference between LLMs in explanatory cultural dimensions. Our study demonstrates that while all LLMs struggle to grasp cultural values, GPT-4 shows a unique capability to adapt to cultural nuances, particularly in Chinese settings. However, it faces challenges with American and Arab cultures. The research also highlights that fine-tuning LLama 2 models with different languages changes their responses to cultural questions, emphasizing the need for culturally diverse development in AI for worldwide acceptance and ethical use. For more details or to contribute to this research, visit our GitHub page https://github.com/reemim/Hofstedes_CAT/
1 INTRODUCTION
The paper addresses the challenge of measuring and improving cultural alignment in LLMs, introducing an explanatory Hofstede-based framework to compare models across diverse regions and dimensions.
- 1 INTRODUCTION: LLMs may mirror Western, Educated, Industrialized, Rich, Democratic societies while failing to reflect other groups’ cultural values.The paper links this pattern to Western-centric training data and the predominantly developed-country origin of the AI community.
- 1 INTRODUCTION: Cultural misalignment can produce misunderstandings, misinterpretations, and heightened cultural tensions among users.
- 1 INTRODUCTION: The study introduces an explainable assessment framework based on Hofstede’s cultural dimensions, whose validation spans more than 70 countries.The framework is used despite criticisms of its methodology and the evolution of cultural dynamics.
- 1 INTRODUCTION: Hofstede’s CAT evaluates six cultural dimensions across the United States, China, and Arab countries using four prompting approaches.The dimensions are Power Distance, Uncertainty Avoidance, Individualism versus Collectivism, Masculinity versus Femininity, Long Term versus Short Term Orientation, and Indulgence versus Restraint.
- 1 INTRODUCTION: The study reports stronger and more consistent cultural-dimension understanding for GPT-4 than other evaluated models, especially when adapted to specific personas.It also examines how temperature, top-p settings, and language-specific fine-tuning affect cultural response patterns.
2 LLM CULTURAL ALIGNMENT BASED ON HOFSTEDE’S CULTURAL ALIGNMENT TEST
Hofstede’s CAT converts VSM13 questionnaire responses into cultural-dimension scores and compares model-generated country rankings with real-world rankings using Kendall Tau correlations.
- 2 LLM CULTURAL ALIGNMENT BASED ON HOFSTEDE’S CULTURAL ALIGNMENT TEST: VSM13 uses a 5-point Likert survey with 30 questions, including 24 cultural-dimension questions and 6 demographic questions, to benchmark LLM alignment.The benchmark rankings derive from the actual VSM13 survey results for the countries in focus.
- 2 LLM CULTURAL ALIGNMENT BASED ON HOFSTEDE’S CULTURAL ALIGNMENT TEST: The prompts feed the 24 survey questions sequentially into each LLM across five consecutive seeds, while demographic responses follow explicit model-related assumptions.English prompting assumes American nationality because the models’ country of development is treated as the relevant nationality.
- 2 LLM CULTURAL ALIGNMENT BASED ON HOFSTEDE’S CULTURAL ALIGNMENT TEST: The experiments compare direct model-level prompting in English, Arabic, and Chinese with prompts that instruct models to act as specific national personas.These prompting conditions assess default and persona-adapted cultural values.
- 2 LLM CULTURAL ALIGNMENT BASED ON HOFSTEDE’S CULTURAL ALIGNMENT TEST: Each cultural index is computed from mean responses to four corresponding survey questions, with constants optionally normalizing scores to 0–100 or anchoring them to prior Hofstede data.
- 2 LLM CULTURAL ALIGNMENT BASED ON HOFSTEDE’S CULTURAL ALIGNMENT TEST: Kendall Tau measures agreement between original VSM13 country rankings and LLM-generated rankings, while misclassification error identifies countries with greater cultural misalignment.The analysis prioritizes relative country positions because precise normalization constants were unavailable for detailed score comparisons.
3 EXPERIMENTS
The experiments compare LLM cultural rankings with Hofstede-based country benchmarks across models, prompting conditions, hyperparameters, and language-specific fine-tuning. GPT-4 generally correlates better than GPT-3.5 without a persona, while persona adaptation, parameter settings, and fine-tuning substantially affect cultural responses.
- Country Level Comparison: Table 2 reports the percentage of mis-ranked cultural dimensions for each country, including the highest and lowest error percentages in the cross-country comparison.
- Model Level Comparison: GPT-4 shows a positive average correlation of 0.11 versus -0.06 for GPT-3.5, with notable GPT-4 correlations of 0.82 for IDV and 0.33 for UAI.
- Country Level Comparison: With specified country personas, GPT-4 maintains an average correlation of 0.11, whereas GPT-3.5 falls from -0.06 to -0.39 and reaches -1.00 on UAI.All models show weak country-level correlations; Llama 2 and GPT-3.5 perform poorly when adapting to specific personas.
- Hyperparameter Comparison: Case 6, using temperature 0.5 and top-p 0, produces the greatest deviation from expected cultural dimensions, while lower temperature with higher top-p and moderate settings improve alignment.The ablation indicates that temperature and top-p significantly influence how cultural dimensions are expressed.
- Language Correlation: Language-specific fine-tuning changes Llama 2 responses: the English-trained model deviates more from cultural norms and refuses some national-pride questions, unlike the Chinese-trained model.The comparison uses Llama-2-13b-chat models trained on English versus Chinese instructions.
4 SOCIETAL IMPACT AND DISCUSSION
The discussion presents Hofstede’s CAT as a diagnostic for cultural alignment and reports uneven model performance across regions. It emphasizes GPT-4’s stronger but context-dependent adaptation, the effects of tuning and alignment practices, and the societal risks of cultural misalignment.
- Hofstede’s CAT enables stakeholders to evaluate LLM cultural alignment, while GPT-4 navigates cultural nuances better than Llama 2 and GPT-3.5 across regions.
- GPT-4 performs poorly in the United States, better in China, and problematically in Arab countries, revealing substantial variation across cultural contexts.The paper notes that GPT-4 was predominantly trained on English data despite its closer alignment with Chinese culture.
- GPT-3.5 shows moderate, relatively consistent regional accuracy, whereas Llama 2 performs best for the United States but struggles with China and Arab countries.
- Hyperparameter tuning is presented as important for customizing model outputs toward cultural perspectives or neutrality, supporting culturally sensitive AI development.
- Cultural misalignment may create ethical and trustworthiness concerns and hinder global AI adoption, especially in regions such as Arab countries.
A RELATED WORK
Related work situates cultural alignment within research on language, social and political bias, and model trustworthiness. The paper distinguishes its focus on country-specific cultural alignment from similarity-based bias assessments and existing fine-tuning approaches.
- Prior studies analyze how language reflects cultural norms and use textual analysis to study gender stereotypes, historical trends, and societal change.
- Political-bias research has found language-model tendencies aligned with left-wing views or opinions prevalent across several Western regions.
- Existing trustworthiness work addresses inaccurate information and problematic behavior through strategies including data purification and behavioral guidelines, leaving further territory uncharted.
B CULTURAL ALIGNMENT FRAMEWORK
The paper adopts Hofstede’s cultural framework and VSM13 to compare cultural values across countries using six dimensions. VSM13 is selected for its empirical coverage, factor-based construction, and design for matched cross-national comparisons.
- The paper considers multiple cultural-comparison frameworks but chooses Hofstede’s VSM13 because of its extensive research coverage and cross-country measurement design.
- VSM13 was empirically tested in more than 70 countries, later extended to broader regions, and updated to include additional Arab countries.
- VSM13 uses factor analysis to group empirically correlated survey questions into cultural dimensions that can be compared across countries.
- The survey contains 30 five-point Likert questions, including 24 cultural-dimension items and 6 demographic questions, with matched respondent characteristics required for valid comparison.
- Hofstede’s framework measures Power Distance, Individualism versus Collectivism, Masculinity versus Femininity, Uncertainty Avoidance, Long versus Short Term Orientation, and Indulgence versus Restraint.
C LIMITATIONS AND CHALLENGES
The framework is an initial step toward cultural alignment, but its evidence is constrained by language coverage, seed selection, and the countries included in comparisons.
- Using only English in the cross-cultural comparison limits testing of potential language-related trends.
- The study uses five consecutive seeds, and increasing their number may produce more robust and reliable results.
- Requiring multiple countries raises questions about how the number of countries influences results and how to align a model with one country during fine-tuning.
- The authors identify calibrating LLMs to be congruent with varied cultural values as the next step.
D HYPERPARAMETER COMPARISON
The ablation shows that temperature and top-p jointly shape cultural-dimension outputs, with settings producing CAT scores from -0.39 to 0.39 and different degrees of cultural expression.
- Higher temperature introduces randomness, while top-p changes the diversity of the sampling pool and jointly affects cultural alignment.
- Average CAT scores span -0.39 to 0.39 across settings, showing substantial variation in cultural-dimension expression.
- The default Case 6 setting, temperature 0.5 and top-p 0, produces the lowest CAT score (-0.39).
- Case 3, with temperature 1 and top-p 0.5, scores slightly higher than Case 2, indicating a complex interaction between the parameters.
- Careful temperature and top-p tuning can target culturally aligned outputs or neutrality.
E LANGUAGE CORRELATION
Language-specific instruction fine-tuning changes how Llama-2-13b-chat reflects cultural values, with the English-trained model showing greater overall deviation than the Chinese-trained model.
- Both models show no Power Distance Index bias, but the English-trained model diverges more in Masculinity, Uncertainty Avoidance, and Indulgence versus Restraint.
- The comparison evaluates Llama-2-13b-chat models trained with English versus Chinese instructions for differences in reflected cultural values.
- The English-trained model has an average deviation of -0.73, compared with -0.19 for the Chinese-trained model.
- The Chinese-trained model shows less deviation, with negative scores only in Masculinity and Long-Term Orientation.
F RESPONSE LEVEL COMPARISON
The response-level analysis averages five seeds and finds generally consistent outputs, while changing sampling hyperparameters can substantially alter GPT-3.5 responses.
- The methodology computes cultural dimensions from mean responses across five consecutive seeds, with model results reported in Tables 5–7.
- GPT-3.5 and GPT-4 have default-case standard deviations below 1, while Llama 2 has standard deviation zero.
- Changing temperature and top-p can shift GPT-3.5 answers by an entire unit, although standard deviation remains near zero except for Question 23 in most tested cases.
- The consistently low standard deviations indicate that the model outputs were likely not random across repeated runs.
- The response analysis includes generated examples and rankings for prompts in English, Chinese, and Arabic, using Arab-country VSM13 values and US values for English.
I.2 GPT-4
GPT-4 responses are presented as hypothetical or imagined preferences rather than personal beliefs, with some outputs simply providing answers without qualification.
- GPT-4 frames its responses as hypothetical answers that may require adjustment for specific profiles.
- GPT-4 also describes its outputs as fictional responses that do not represent any real individual’s preferences or beliefs.
- Some responses explain that the model has no personal preferences or emotions and instead uses imagined preferences.
- Other outputs provide answers directly without adding an explanation of their status.
I.3 LLAMA 2
Llama 2 produces persona-based responses framed as an American respondent, assigning high importance to work-life balance, a respected boss, recognition, job security, and pleasant colleagues.
- Llama 2 presents sufficient time for personal or home life as highly important for maintaining balance and avoiding stress.
- Llama 2 rates having a respected boss as highly important, linking it to job satisfaction, support, productivity, and well-being.
- Llama 2 considers recognition for good performance important because it can increase feelings of value, self-esteem, and motivation.
- Llama 2 describes employment security as very important while acknowledging flexibility, growth, and purpose as additional priorities.
- Llama 2 treats pleasant, supportive coworkers as highly important because workplace relationships affect everyday enjoyment.