Source-linked AI summary

Accurate in space, unreliable in time: how LLMs represent national cultural change

Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp

arXiv:2609.01902v1cs.CYcs.AIcs.CL

TL;DR

Most cultural-alignment evaluations assess LLMs against a single current snapshot, despite evidence that societies change over time. This paper compares WVS trajectories for 40 countries with those produced by four SOTA LLMs on the Inglehart–Welzel map. Models often locate recent positions well but represent cultural trajectories incompletely, lagging in time and rarely reproducing reversals.

  • Problem

    Cultural-alignment evaluations usually test a fixed snapshot, leaving unclear whether LLMs capture how societies change over time.

  • Method

    The paper compares four LLMs’ coordinate-based representations with repeated WVS measurements for 40 countries across Waves 4–7 on the Inglehart–Welzel cultural map.

  • Results

    LLMs generally place countries near recent surveyed positions but understate cultural change, introduce movement where little occurred, and rarely reproduce genuine reversals.

  • Takeaways & Limitations

    Snapshot accuracy does not fully capture cultural awareness because current models represent societies’ trajectories more weakly than their recent positions.

  • Takeaways & Limitations

    The analysis reduces culture to two Inglehart–Welzel dimensions and uses irregular repeated cross-sectional WVS observations rather than continuous measurements.

Abstract

from arXiv · show

Assessments of cultural alignment have become an important part of the development and improvement of large language models (LLMs). However, the majority of the evaluations treat culture as a single snapshot, investigating only whether a model represents a society accurately at the current time. Research in cultural psychology shows that cultural values change at different rates and directions over time. Therefore, a "culturally aware" model should capture not only where a culture is today but also how it has changed over time. We examine this missing dimension of cultural awareness using more than two decades of the World Values Survey data. We compare the cultural trajectories of 40 countries with the trajectories produced by four state-of-the-art (SOTA) LLMs on the Inglehart-Welzel cultural map. Our findings show that while models generally place countries close to their most recent surveyed positions, these representations tend to lag several years behind that position. They also capture only part of the magnitude of the observed change, introduce movement where little occurred, and rarely reproduce reversals in countries' trajectories. These findings point to temporal flattening and suggest that snapshot accuracy can give an incomplete picture of cultural awareness in LLMs and have implications for model evaluation, representational harms, and the governance of culturally aware AI systems.

Introduction

Cultural values change at different rates and directions, so representing a society requires capturing its trajectory rather than only a fixed snapshot. This paper evaluates whether LLMs align with both societies’ positions and their changes over time using WVS data and cultural-map coordinates.

  • Why culture must be modeled over time: Societies can shift toward secular-rational and self-expression values at different speeds, and some trajectories can reverse.Examples include post-Soviet movement toward survival and traditional values and traditionalist counter-movements in advanced democracies.
  • Why culture must be modeled over time: A society is better represented by a trajectory containing direction, rate, and occasional reversal than by a single measured snapshot.
  • Why LLM cultural representation matters: LLMs increasingly mediate information and are proposed as low-cost samples of human populations, making their cultural representations relevant to governance.The paper connects this relevance to concerns about representation, cultural erasure, and linguistic hegemony.
  • Study framework: The paper distinguishes snapshot alignment, locating a society at a recent point, from change alignment, capturing direction, rate, and reversals over time.Both are compared with World Values Survey results projected onto the two-dimensional Inglehart–Welzel cultural map.
  • Study framework: Four LLMs are tested on ten core WVS indicators across 40 countries and Waves 4–7 using coordinates estimated from model responses.The coordinate procedure applies comparable transformations to survey and model estimates.
  • Trajectory analysis: The study measures country trajectories using path length, net displacement, turn angles, and stepwise alignment against a permutation null.These measures distinguish general movement from country-specific direction and identify possible reversals.

Results

The models represent countries near their recent cultural positions, but their temporal representations lag, understate substantial change, and add movement where little occurred. They generally recover broad movement direction, yet rarely reproduce reversals and show mixed robustness across specifications.

  • Snapshot alignment: 0.153–0.238 mean distance placed models near countries’ most recent WVS coordinates, compared with 0.40 average country-to-centroid distance.Naming the country also improved accuracy relative to the no-country baseline.
  • Snapshot alignment: 0.153 versus 0.151 for Gemini, 0.177 versus 0.182 for Claude, 0.193 versus 0.184 for GPT, and 0.238 versus 0.231 for Qwen showed little snapshot benefit from adding the year.Country-only and country-plus-year prompts produced similar distances to the recent snapshot.
  • Implicit anchor: 7.0, 6.0, 4.7, and 4.5 years were the average implicit-anchor lags for Claude, GPT, Qwen, and Gemini, respectively.In wave units, the corresponding lags were 0.97, 0.89, 0.69, and 0.67 waves.
  • Timing and rate: 0.45, 0.56, 0.58, and 0.68 were the median fractions of observed movement captured by Claude, Gemini, GPT, and Qwen on ten clear directional paths.Across all 36 movers, under-travel remained significant for Claude and GPT, but not Gemini or Qwen.
  • Timing and rate: 0.097 units of median observed movement became 0.113–0.170 units in every model for the four low-change countries.This descriptive pattern placed all four countries above the equality line, indicating more model-estimated movement than observed.
  • Reversal reproduction: 7 of 24 country–model cases reproduced a sharp reversal at the observed wave, falling to 4 when requiring at least half the real journey length.Models therefore reproduced reversal direction more often than reversal magnitude.

Discussion

The discussion argues that LLMs represent national cultures more accurately as recent positions than as changing trajectories, with implications for evaluation and representational harms.

  • Discussion: Current SOTA LLMs placed countries near recent positions but understated cultural change, introduced movement where little occurred, and rarely reproduced reversals.Explicitly naming the survey year did little to correct these errors.
  • Discussion: Across 40 countries, only four met the study’s criterion for low change, making trajectory representation relevant for most sampled societies.
  • Discussion: LLM evaluations largely test country-level values, culturally specific knowledge, or adaptation to cultural context rather than whether models capture change over time.
  • Discussion: A society’s current interpretation depends partly on its prior path, so an up-to-date snapshot alone provides weaker information than a representation of how it got there.
  • Discussion: The authors call this failure mode “temporal flattening,” with particular implications for modeling attitude change, intergenerational differences, and historical events.
  • Discussion: Treating values as fixed may turn a time-specific position into an enduring group characteristic, resembling representational harms and concerns about cultural erasure.
  • Discussion: The study is limited by reducing national culture to two Inglehart–Welzel dimensions, irregular repeated cross-sectional WVS observations, and elicited rather than fully latent model representations.

Funding

The paper reports no specific funding for this work.

  • Funding: The authors received no specific funding for this work.

Ethics statement

The study used publicly available, anonymized World Values Survey data and model responses obtained through provider APIs under applicable terms.

  • Ethics statement: The study did not involve human participants and therefore did not require institutional review board approval.
  • Ethics statement: It used publicly available, anonymized secondary WVS data and model responses elicited through provider APIs in accordance with their terms.

Supplementary Materials Accurate in space, unreliable in time: how LLMs represent national cultural change

The supplementary materials document the country sample, benchmark construction, indicator specifications, wave variation, trajectory-comparison geometry, and reduced indicator bases.

  • S1. Sample: The supplementary materials include sample and wave-coverage tables for countries including Serbia, Singapore, South Korea, Taiwan, Thailand, Turkey, Ukraine, the United States, Uruguay, Vietnam, and Zimbabwe.
  • S2. Empirical benchmark construction: Table S2 provides standardization parameters for constructing the empirical benchmark.
  • S2. Empirical benchmark construction: The saved groundtruth coordinates can be exactly rebuilt from thirteen defining rows, with a maximum coordinate difference of 0.
  • S2. Empirical benchmark construction: Supplementary materials specify indicator orientation and identify countries with a reduced indicator base.
  • S2. Empirical benchmark construction: WVS item text was held constant across waves except where the instrument differed for exactly two items.
  • S2. Empirical benchmark construction: Figure S1 defines directional agreement through the angle between model and observed steps, while turn angle measures within-trajectory changes in direction.

S3. Trajectory geometry and classification

This section documents supplementary materials for trajectory geometry and classification, including tables and a movement-threshold comparison.

  • Table S6 provides supplementary material for this section.
  • Countries nearest the 0.13 movement threshold are classified under three candidate floors.
  • Table S7 provides supplementary material for this section.

S4. Elicitation

The elicitation procedure assembles standardized survey-style prompts across countries, years, items, and conditions, collecting probability distributions from four models. The benchmark contains 160 elicitation files and 58,188 option-level rows.

  • Prompt assembly: The assembled call combines a condition wrapper, a wave-specific item stem, and that item’s response line.
  • Responses: Models provide probabilities over ten scale points for single-response items, with decimals summing to 1.
  • Cardsort: The child-qualities cardsort is an exception: each quality receives a percentage, and the values need not sum to 1.
  • Conditions: C0 omits country and year, C1 names the country, and C2 names both country and year.
  • Sampling: The full grid uses one call per retained wave year, with 2,100 independent draws per cell across all four models.
  • Elicitation scope: 160 elicitation files cover four models and 40 countries.

S5. Scoring model outputs onto the cultural map

Model outputs are compared with WVS pooled item means after scoring both onto a common cultural-map representation.

  • Model item means are compared against WVS pooled item means.
  • The scoring section reports model and WVS means in a shared item-level format.
  • The comparison is organized by item and model.

S6. Snapshot alignment

Snapshot alignment measures model placement relative to each country’s latest WVS coordinate, including signed axis bias, cross-country spread, and effects of naming the country or year. The results show systematic model differences in bias and placement distance across anchoring conditions.

  • Signed bias: −0.067 is Gemini 3.6 Flash’s Dim 1 bias under C2, while 0.052 is Claude Opus 4.8’s Dim 2 bias under C2.Bias is measured relative to the latest WVS coordinate; positive values indicate more secular Dim 1 or more self-expressive Dim 2 placement.
  • Signed bias: 0.085 is GPT-5.5’s Dim 2 bias under C1, compared with 0.118 for Qwen3.6-Plus under C1.
  • Placement distance: Qwen is farther from the latest WVS coordinate than Claude, Gemini, and GPT under both C1 and C2.The pairwise distance differences are positive for Qwen minus each comparator in both anchoring conditions.
  • Placement distance: Gemini is farther from the latest WVS coordinate than GPT under both C1 and C2.
  • Cross-country spread: Gemini’s Dim 1 cross-country spread is 1.094 times truth under C1 and 1.106 times truth under C2.The ratio compares model coordinate dispersion with the empirical benchmark; 1 indicates equal spread.
  • Anchoring: The country-naming comparison measures how much closer models move to surveyed positions after the country is named.

S7. Temporal anchor

The analysis measures how far each model’s unprompted country placement lags behind that country’s most recent surveyed wave. It compares nearest-wave timing across 36 moving countries and reports mean distances across waves.

  • S7. Temporal anchor: The reported analysis covers 36 moving countries, with a sensitivity analysis using all 40 countries.
  • S7. Temporal anchor: Lag is the age difference between the wave nearest the model’s C1 placement and the country’s most recent observed wave.A lag of 0 means the unprompted placement is nearest the latest survey.
  • S7. Temporal anchor: Gemini’s most-recent-wave placement error averages 0.151 cultural-map units, compared with 0.182 for Claude, 0.184 for GPT, and 0.231 for Qwen.Lower distance indicates closer placement; the average country sits 0.40 units from the global centroid.
  • S7. Temporal anchor: Mean distance from the unprompted C1 placement to each wave is compared by wave position across Gemini, Claude, GPT, and Qwen.Wave position 0 is the most recent wave; higher positions are older waves.

S8. Change alignment

This section compares models’ estimated cultural change with observed start-to-end movement and tests whether step directions align with the periods in which changes occurred. It also examines low-change countries and robustness across specifications.

  • S8. Change alignment: Across 36 moving countries, magnitude ratio measures model displacement divided by true displacement, while direction cosine measures alignment between their net-displacement vectors.A magnitude ratio of 1 means the model travels the true distance; cosine above 0 indicates the correct direction.
  • S8. Change alignment: Drift-free step-alignment gaps were 0.123 for Gemini, 0.085 for Claude, 0.049 for GPT, and 0.003 for Qwen across all movers.The gap credits alignment that depends on assigning steps to the correct periods rather than merely matching movement direction.
  • S8. Change alignment: The drift-free gap remained positive for Gemini and Claude under the reported specifications, while GPT and Qwen’s intervals generally included zero.The reduced-base specification excludes nine countries, and the analysis also reports an Iraq leave-out.
  • S8. Change alignment: Among 10 directional countries, all models captured only part of observed change: median ratios were 0.449 for Claude, 0.564 for Gemini, 0.576 for GPT, and 0.684 for Qwen.The directional subset is used because reversing or wandering paths can contaminate magnitude ratios through cancellation.

S9. Reversal reproduction

The reversal analysis visualizes surveyed and model-generated trajectories for six countries classified as reversals. It compares the timing and direction of model paths with observed turns across 24 country–model cases.

  • S9. Reversal reproduction: The comparison covers all 24 country–model cases across the six verified reversal countries, ordered by true turn sharpness.
  • S9. Reversal reproduction: The reversal figure contains six country panels, with dark paths showing surveyed Waves 4–7 positions and lighter paths showing model positions for corresponding years.Points mark survey waves and arrowheads indicate movement direction; axes are scaled separately across countries.
  • S9. Reversal reproduction: The per-model summary aggregates performance across the six verified reversals.
Loading 2609.01902v1…