Source-linked AI summary

Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs

Ariba Khan, Stephen Casper, Dylan Hadfield-Menell

arXiv:2503.08688v2cs.CY

TL;DR

Cultural alignment in LLMs remains challenging, and survey-based evaluation methods may not adequately characterize it. This paper investigates cultural dimensions and evaluation design, finding erratic, context-dependent preferences and substantially different outcomes from small methodological changes.

  • Problem

    Cultural alignment in LLMs remains challenging, while existing survey-based methods may not adequately characterize how effectively models align with different cultures.

  • Method

    The paper statistically investigates the interplay between cultural dimensions and uses survey evaluations including a standard 5-point Likert scale.

  • Results

    State-of-the-art LLMs display erratic cultural preferences, with alignment appearing nuanced and context-dependent and outcomes changing substantially after small methodological modifications.

  • Takeaways & Limitations

    Narrow experiments, overly simplistic evaluations, cherry-picking, and confirmation biases can produce incomplete or misleading understandings of LLM cultural alignment.

  • Takeaways & Limitations

    Extrapolability and steerability do not hold in general for state-of-the-art LLMs, although they may hold in specific cases.

Abstract

from arXiv · show

Research on the 'cultural alignment' of Large Language Models (LLMs) has emerged in response to growing interest in understanding representation across diverse stakeholders. Current approaches to evaluating cultural alignment through survey-based assessments that borrow from social science methodologies often overlook systematic robustness checks. Here, we identify and test three assumptions behind current survey-based evaluation methods: (1) Stability: that cultural alignment is a property of LLMs rather than an artifact of evaluation design, (2) Extrapolability: that alignment with one culture on a narrow set of issues predicts alignment with that culture on others, and (3) Steerability: that LLMs can be reliably prompted to represent specific cultural perspectives. Through experiments examining both explicit and implicit preferences of leading LLMs, we find a high level of instability across presentation formats, incoherence between evaluated versus held-out cultural dimensions, and erratic behavior under prompt steering. We show that these inconsistencies can cause the results of an evaluation to be very sensitive to minor variations in methodology. Finally, we demonstrate in a case study on evaluation design that narrow experiments and a selective assessment of evidence can be used to paint an incomplete picture of LLMs' cultural alignment properties. Overall, these results highlight significant limitations of current survey-based approaches to evaluating the cultural alignment of LLMs and highlight a need for systematic robustness checks and red-teaming for evaluation results. Data and code are available at https://huggingface.co/datasets/akhan02/cultural-dimension-cover-letters and https://github.com/ariba-k/llm-cultural-alignment-evaluation, respectively.

1 Introduction

This paper investigates whether survey-based evaluations reliably characterize cultural alignment in LLMs. It tests stability, extrapolability, and steerability, finding substantial inconsistencies and methodological sensitivity.

  • Research questions: The study tests three assumptions: stability, extrapolability, and steerability.These assumptions concern whether alignment is intrinsic, generalizes across issues, and can be induced through prompting.
  • Approach: The experiments combine explicit cultural surveys with implicit preference elicitation through simulated hiring scenarios.The implicit evaluation examines behavior rather than directly asking models about values.
  • Findings: The findings report significant inconsistencies across explicit surveys and implicit preference elicitation.These inconsistencies challenge the reliability of current survey-based evaluation approaches.
  • Findings: Small methodological variations can substantially shift conclusions about LLM cultural alignment.The case study found that results from prior work did not replicate when models could select indifference between alternatives.
  • Implications: Narrow experiments or selective evidence assessment may produce an incomplete picture of LLM cultural alignment.The authors therefore call for critical re-examination of popular survey-based methods.

2 Related Work

Prior work evaluates cultural alignment through discriminative surveys and generative outputs, but evidence indicates that model preferences may not correspond reliably to human cultural values.

  • Cultural alignment: Cultural alignment concerns how closely model behavior reflects common beliefs about what is desirable and proper within a culture.Researchers study whether model behaviors reflect or diverge from different cultural perspectives.
  • Evaluation paradigms: Discriminative assessments ask models to select among predetermined options, whereas generative assessments analyze free-form outputs.Generative evaluations may use single-turn or multi-turn interactions.
  • Benchmarks: VSM and GQA are survey-based resources used to assess model responses across cultural dimensions and contexts.VSM measures six cultural dimensions, while GQA combines World Values Survey and Pew Research questions.
  • Methodological concerns: Prior studies found weak correlations between LLM outputs and established human cultural-value surveys.Other work also found that model preferences can be sensitive to prompting.
  • Methodological concerns: Existing approaches may struggle to capture and contextualize the nuances of LLM cultural preferences.This concern motivates systematic examination of evaluation assumptions.

3 Identifying Key Assumptions

The paper formalizes stability, extrapolability, and steerability as assumptions underlying cultural-alignment evaluations. Prior evidence and the paper’s framing question whether these assumptions hold for LLMs.

  • Core assumptions: The paper identifies three assumptions underlying current cultural-alignment evaluations: stability, extrapolability, and steerability.These assumptions concern consistency across equivalent formats, generalization across dimensions, and reliable persona induction.
  • Stability: Stability assumes cultural alignment remains consistent across semantic-preserving variations in evaluation methodology.Prior studies found that prompt phrasing can produce variations larger than differences between preferences.
  • Extrapolability: Extrapolability assumes alignment on a limited set of cultural issues predicts alignment on unobserved issues.The paper tests whether cultural dimensions reliably predict one another in LLMs and humans.
  • Extrapolability: Prior research reports weak correlations between LLM outputs and established cultural value surveys despite alignment on selected issues.Human cultural dimensions also show partial correlation and country-specific variation.
  • Steerability: Steerability assumes prompting can make LLMs consistently embody coherent cultural perspectives.The paper examines whether this remains possible beyond narrow contexts, even with prompt optimization.
  • Steerability: Prior work reports that cultural prompting can improve alignment for some countries while failing or increasing bias for others.This motivates testing steerability across tasks and models rather than treating persona induction as uniformly reliable.

4 Experimental Setup

The experiments evaluate five LLMs with standardized survey and implicit-preference procedures. They use VSM, GQA, and culturally varied cover-letter comparisons, with normalization, effect-size calculations, and permutation tests.

  • Models and protocol: The study evaluates GPT-4o, Claude 3.5 Sonnet, Gemini 2.0 Flash, Llama 3.1 405B, and Mistral Large 2411.The models span different architectures, releases, and developers.
  • Models and protocol: All experiments use Likert-scale questions, with temperature 0.0, three independent runs, and averaged measurements unless noted otherwise.The protocol is designed to support consistent and reproducible results.

5 Evaluating Key Assumptions

The experiments test whether cultural preferences remain stable across evaluation designs, extrapolate across dimensions, and respond reliably to cultural prompt steering. Across explicit and implicit evaluations, minor methodological changes produce large shifts, extrapolation is dimension-sensitive, and prompted outputs remain unlike human cultural responses.

  • Stability: LLM responses changed significantly when question direction or response format varied, despite no semantic change.The study compared ascending versus descending options and identifier-only versus option-text responses.
  • Stability: Effect sizes from presentation changes frequently exceeded the 0.114 between-country human standard-deviation benchmark.
  • Stability: Comparative versus absolute cover-letter ratings produced different preference distributions, with comparative ratings showing wider variance and more extreme negative values.
  • Stability: Reasoning requirements, Likert scale size, and prompted role each altered expressed cultural preferences across dimensions.Reasoning effects were especially pronounced for Long/Short Term Orientation and Indulgence/Restraint, while role changes were strongest for Long/Short Term Orientation and Uncertainty Avoidance.
  • Extrapolability: Extrapolation from few cultural dimensions was near the random-chance baseline, whereas ARI exceeded 0.8 for humans and LLMs when substantially more dimensions were included.Extrapolation also depended strongly on which dimensions were observed.
  • Steerability: Prompt-steered LLM responses were erratic and unlike human responses, with human-model distance tests returning p=0.0 for both prompting methods.Human responses from different countries clustered more closely than the prompted model outputs.

6 Manipulating LLM Evaluations with Forced Binary Choices: a Case Study

The case study shows that forced binary-choice evaluation can create an apparent hierarchy of country preferences that disappears when models may choose neutrality. Thus, evaluation design can materially change the observed interpretation of GPT-4o’s valuation of human lives.

  • Evaluation design: The case study tested whether excluding a neutral option changes apparent country preferences in forced-choice evaluations.It compared a 5-point Likert scale with neutrality against a forced-choice 4-point scale.
  • Results: The 4-point forced-choice condition produced non-uniform country preferences and a visible preference hierarchy.
  • Results: The 5-point neutral condition selected “No preference” in 100% of country-pair comparisons and gave all 11 countries a normalized score of 0.50.
  • Interpretation: Allowing neutrality led GPT-4o to indicate equal valuation of human lives regardless of nationality, supporting a more nuanced interpretation of earlier apparent preferences.
  • Implications: The findings caution that outputs from specific elicitation paradigms may be constructed responses rather than stable internal preferences.

7 Discussion

The discussion argues that state-of-the-art LLMs display erratic, context-dependent cultural preferences, making broad conclusions from narrow evaluations unreliable. It recommends stronger robustness practices while acknowledging that cultural alignment evaluations may still hold in narrow cases.

  • State-of-the-art LLMs show erratic cultural preferences, with apparent cultural alignment varying substantially across contexts and small methodological changes.The discussion cautions that narrow experiments can produce substantially different outcomes.
  • Narrow evaluations, cherry-picking, and confirmation biases may produce incomplete or misleading understandings of cultural alignment.
  • Stability, extrapolability, and steerability do not hold in general for state-of-the-art LLMs, although they may hold in specific narrow circumstances.The authors do not claim that cultural alignment evaluations are fundamentally invalid.
  • Pre-registration could reduce cherry-picking and p-hacking, while red-teaming could strengthen assessments that are unreliable and sensitive to adversarial modifications.
  • Future work could improve steerability and preference consistency, develop evaluations reflecting contextual complexity, and study deployment effects on power dynamics.

A Prompts

The prompt appendix documents survey variations used in the evaluation, including response ordering, response format, rating scale type, and Likert-scale length.

  • Ascending and descending response-option orderings represent direction variations for otherwise identical survey questions.
  • Identifier-only and full-text response requirements represent response-type variations while keeping question content identical.
  • Comparative scales ask respondents to choose between two cover letters, whereas absolute scales rate individual cover letters.
  • The evaluation compares 4-point, 5-point, and 6-point Likert scales as alternative response-option formats.

[Hiring Manager]

The prompt assigns the model the role of a hiring manager at an organization focused on improvement and progress.

  • The prompt frames the evaluator as a hiring manager.
  • The hiring organization is characterized as focused on improvement and progress.
  • The role framing establishes an organizational evaluation context for the prompt.

[Career Coach]

The prompt assigns the model the role of a career coach helping organizations evaluate candidates and assess their fit for open positions.

  • The prompt frames the evaluator as a career coach.
  • The career coach helps organizations evaluate candidates.
  • The role includes assessing candidates’ potential fit for open positions.

[Job Applicant]

The evaluation frames the model as a job applicant seeking a suitable position and varies the professional perspective used to assess cover letters.

  • The model is instructed to act as a job applicant seeking a position aligned with its skills and career aspirations.
  • Cover-letter evaluation varies the professional context or perspective used to assess the same letters.The variation is intended to examine how role context might influence assessment.

[Without Reasoning]

The evaluation asks models to compare options, including lives from different countries, using rating scales that vary in format and reasoning requirements. These variations include prompts with or without reasoning, neutral or forced-choice scales, and cultural perspective-taking instructions.

  • Some versions require two or three sentences of reasoning before the rating, while others assess preferences without that requirement.The reasoning prompt asks models to explain the factors influencing their preference between cover letters.
  • A cultural perspective-adopting prompt instructs the model to respond from the viewpoint of a typical person from a specified country.The prompt is used in both extrapolation and steerability experiments.
  • Models rate their preference between two options describing observations about the state of the world.
  • The task compares saving the life of a person from country_A with saving the life of a person from country_B.
Loading 2503.08688v2…