Source-linked AI summary
CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
TL;DR
Cultural evaluations often reduce competence to single-turn factual recall, leaving practical, culturally grounded help-seeking underexamined. CultureConverse introduces a multilingual multi-turn simulation and evaluation harness with hidden cultural constraints, finding strong human alignment and downstream gains from fine-tuning its dataset. The authors position it as a scalable framework for interactive cultural evaluation, while noting that its taxonomy, generation fidelity, and training gains remain limited.
Problem
Existing cultural evaluations emphasize single-turn factual recall, leaving subgroup variation, context maintenance, and actionable multi-turn assistance insufficiently assessed.
Method
CultureConverse simulates scored multi-turn episodes across culturally situated users and assistants, with hidden cultural state, oracle-guided judging, and auditable episode data.
Results
Fine-tuning on 27,860 high-quality samples improves in-domain assistance and transfers to cultural MCQ and safety classification benchmarks; GPT-5 mini achieves the highest assistance quality.
Takeaways & Limitations
Human-aligned automated evaluation and out-of-domain transfer suggest that culturally grounded multi-turn harnesses can scale evaluation beyond cultural trivia.
Takeaways & Limitations
The benchmark is not an exhaustive account of cultural truth: its demographic taxonomy compresses community diversity, generation fidelity varies, and training gains are modest.
Abstract
from arXiv · showhide
Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Each simulated and evaluated episode produces a scored interaction where the assistant assists the user and infers cultural constraints from partial information. The resulting CultureConverse-DS dataset contains 14,610 benchmark (evaluation) episodes and 274,295 oracle-guided (gold-mode) dialogues. In our benchmark evaluation of 18 models, GPT-5 mini achieves the highest assistance quality. Human annotation experiments suggest that our evaluation framework is a sufficient proxy for human judgment. Performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks. We release the harness, both splits, and judge prompts to support interactive evaluation of cultural competency.
1 Introduction
Cultural evaluation often misses subgroup variation, multi-turn context maintenance, and the ability to turn cultural knowledge into practical assistance. CultureConverse addresses these gaps with hidden cultural constraints, multilingual simulated episodes, automated judging, and a released training dataset.
- Motivation: Existing evaluations underrepresent practical help-seeking by focusing on country-level labels, single-turn prompts, and factual MCQs.These formats can obscure within-country subgroup disparities and cannot test context maintenance or actionable assistance.
- Approach: CultureConverse evaluates multi-turn assistance in which the assistant infers cultural constraints from user turns rather than receiving them explicitly.The episode uses a hidden oracle available to the simulator and judge, while the assistant sees only the dialogue history.
- Approach: 10 regions, 58 subgroup identities, and 7 domains are covered by the multilingual simulation harness, grounded in curated Cultural and Taboo Knowledge Bases.The harness targets stateful, culturally constrained interactions across East and Southeast Asia.
- Evaluation: 18 frontier LLMs are benchmarked across 14,610 simulated dialogues, with automated metrics reported as strongly aligned with human consensus.The benchmark evaluates assistance in culturally sensitive, multi-turn scenarios.
- Dataset and contribution: 274,295 guided training trajectories support the released CultureConverse-DS dataset, and fine-tuning on 27,860 high-quality samples improves downstream cultural and safety benchmarks.The released corpus is intended for supervised fine-tuning and evaluation research.
2 Related Work
Prior work documents regional, linguistic, demographic, and conversational weaknesses in LLMs, while existing cultural evaluations often remain static or shallow. CultureConverse combines stateful multi-turn interaction with culturally grounded simulation and human-calibrated judging.
- Multilingual and regional benchmarks: Multilingual benchmarks reveal performance variation across languages and regions, including exposure bias, local commonsense gaps, and weaknesses in lower-resource languages.SEA-HELM and SEACrowd/SEA-LION broaden coverage for underrepresented Southeast Asian languages.
- Positioning of CultureConverse: CultureConverse responds by evaluating stateful, multi-turn assistant interaction and calibrating its LLM judge against diverse human annotators.Its pipeline expands persona-template seeds into culturally grounded dialogue simulations before automated scoring.
- Cultural alignment and bias: Prior studies find Western- or English-centric value patterns and difficulties with local norms, etiquette, and demographic bias under persona prompting.Related work measures these gaps through sociological values, survey replication, morality probes, and culture-sensitive alignment methods.
- Interactive evaluation: Multi-turn evaluations show that deployed assistants can lose context or make premature assumptions, while synthetic personas may introduce demographic artefacts.Interactive frameworks use generative agents, role-play, and multi-turn probes to study conversational social intelligence.
- Automated judging: LLM-as-a-judge methods scale evaluation, but cultural judgments remain value-laden, context-dependent, and sensitive to demographic preferences and prompt instability.Bias research further emphasizes power relations and intersectionality rather than treating demographic labels as neutral.
3 CULTURECONVERSE Construction
CULTURECONVERSE constructs auditable, culturally grounded episodes through a seed-to-evaluation pipeline that progressively reveals constraints during multi-turn interaction. Its design separates assistant, simulator, and judge views while covering diverse regions, identities, domains, and challenge tiers.
- Overview: The evaluation unit is an auditable episode: a scored multi-turn dialogue requiring progressive inference and application of implicit cultural norms.Each episode stores persona information, knowledge-base evidence, simulator directives, and scores linked to its originating seed.
- Pipeline: Episodes follow a five-stage seed → blueprint → shard → simulation → evaluation chain that specifies scenario content, information timing, role visibility, and judging.The stages support explicit analysis of ecological validity, judge–human alignment, and out-of-distribution transfer.
- Information asymmetry: The assistant sees public context and chat history, the simulator gradually surfaces private constraints, and the judge accesses oracle success criteria and safety tripwires.This asymmetric partition evaluates assistants under realistic partial observability while retaining oracle-guided scoring.
- Templates and seeds: 348 personas paired with 420 templates yield 146,160 seeds distributed 50% / 25% / 25% across Standard, Contextual, and Normative tiers.The tiers represent baseline assistance, cross-cultural friction, and taboo or safety-trap detection.
- Blueprints: Blueprints retrieve region- and persona-filtered cultural knowledge plus taboo evidence, then partition populated scenario fields into public, private, and oracle context.The knowledge sources include 180,992 region-tagged Wikipedia passages and 962 curated taboo entries.
- Shards and simulation: Shard plans contain 3–5 ordered phases with user goals, language directives, and transition conditions, enabling progressive disclosure rather than front-loading constraints.Runtime validation checks structure, cultural plausibility, answerability, and challenge sufficiency; failed outputs are regenerated within a bounded retry budget.
- Dataset: 288,905 CULTURECONVERSE-DS episodes span 10 regions, 58 subgroup identities, and 7 domains.The dataset breakdown is reported by split, language mode, region, and challenge configuration.
4 Experimental Setup
The experiments evaluate generated dialogue quality and culturally grounded assistance with oracle-informed metrics across a stratified multilingual benchmark. Human annotation is used to assess whether the automated judge provides a credible aggregate evaluation.
- Evaluation unit: The evaluation unit records a frozen shard, hidden oracle, safety and cultural tripwires, curated knowledge base, and dialogue turns with asymmetric visibility across roles.The assistant sees only the dialogue, while the simulator and judge receive progressively broader scenario information.
- Metrics: TDQ averages Naturalness, Scenario Plausibility, and Cultural Typicality to measure generated dialogue realism and ecological validity.High TDQ is intended to indicate that the benchmark approximates real-world interactions.
- Metrics: 3H averages Helpfulness, Honesty, and Harmlessness against oracle criteria and tripwires, alongside a binary clean rate for episodes without violations.The metrics target task resolution, calibrated accuracy, and safety or cultural sensitivity.
- Benchmark design: 18 target assistants are evaluated on 14,610 paired episodes per assistant, split evenly between global-English and native-language modes across identical frozen shards.The evaluation remains stratified across 10 regions, 58 subgroup identities, and 7 domains, with fixed prompts and GPT-5 mini as simulator and default judge.
5 Experimental Results
Across 14,610 episodes per assistant, the evaluation measures dialogue realism, assistance quality, safety, human alignment, transfer, and subgroup disparities. GPT-5 mini leads mean assistance quality, while fine-tuning on 27,860 high-quality episodes improves assistance and transfers to other benchmarks.
- Evaluation setup: 14,610 episodes per assistant are evaluated with TDQ, 3H, and clean rate.TDQ averages Naturalness, Plausibility, and Typicality; 3H averages Helpfulness, Honesty, and Harmlessness, while clean rate tracks avoidance of hidden safety tripwires and taboos.
- Benchmark results: GPT-5 mini achieves the highest 3H score at 4.28, while GPT-5.4 achieves the highest clean rate at 91.0%.Assistance Quality ranges from 3.25 to 4.28 across evaluated models.
- Benchmark results: TDQ and 3H are correlated but distinct, separating conversational realism from culturally grounded assistance.Some models assist well despite weaker realism, whereas others converse realistically but provide weaker assistance.
- Human-aligned evaluation: 90.1% ±1 agreement and 0.581 MAE show GPT-5 mini’s judge aligns favorably with human consensus.The human–human baseline is 87.5% agreement and 0.727 MAE; the judge is used for aggregate comparisons, not expert adjudication of cultural truth.
- Downstream generalisation: 27,860 perfect-score episodes improve overall assistance quality after LoRA fine-tuning and transfer to 7 cultural MCQ and 10 safety classification datasets.Llama-3.1-8B-IT gains +0.88 pp on MCQ and +2.49 pp on CLS, while SEA-LION-v4-8B-IT gains +0.52 pp and +0.80 pp.
- Performance disparities: GPT-5 mini shows a 0.165 3H spread across identities, with lower scores for historically under-represented groups.The reported gender range is negligible at 0.011 3H.
6 Discussion
The discussion frames multi-turn cultural alignment as more practically relevant than static factual recall, while highlighting simulator realism and representation as unresolved challenges. The benchmark’s value depends on controls for conversational diversity and careful interpretation of demographic results.
- Utility of Multi-turn Cultural Alignment: Multi-turn interactions expose implicit values and pragmatic utility that cultural MCQs and single-turn prompts do not test.The authors connect human alignment and out-of-domain transfer to the utility of interactive cultural evaluation.
- The Challenge of Simulator Realism: Simulator realism remains a challenge because LLM-driven users can collapse into predictable, formulaic prompting patterns.CULTURECONVERSE adds stylistic directives to increase conversational diversity while preserving reproducibility.
- Subgroup Disparities and Cultural Essentialism: Equal weighting across 58 subgroups upweights minoritised communities, but its superiority to proportional representation remains unresolved.Observed demographic variance also reflects sociolinguistic bottlenecks in the underlying base models.
- Subgroup Disparities and Cultural Essentialism: Explicit persona labels create tension between minority visibility and the risk of cultural essentialism.The authors identify avoiding monolithic stereotypes as an open challenge for authentic cultural representation.
7 Conclusion
CULTURECONVERSE and CULTURECONVERSE-DS provide a scalable framework and corpus for culturally grounded, multi-turn assistance, with evaluations aligned to human consensus and downstream generalisation. The authors caution that the benchmark is a structured evaluation artefact rather than an exhaustive or perfect representation of cultural truth or human behaviour.
- CULTURECONVERSE is a scalable simulation and evaluation harness, while CULTURECONVERSE-DS is a benchmark and training corpus for culturally grounded, multi-turn assistance.
- The experiments differentiate models’ abilities to apply cultural nuance in help-seeking scenarios and show strong alignment with human consensus.
- Fine-tuning on CULTURECONVERSE-DS yields measurable downstream generalisation across interactive and static cultural benchmarks.
- The benchmark cannot exhaustively represent cultural truth because demographic taxonomies compress community diversity and generation fidelity varies across identities and languages.
- Generated episodes are targeted evaluation artefacts, not a perfect proxy for human prompting behaviour, and modest training gains indicate proof of feasibility rather than large or consistent improvements.
- Discrete personas risk essentialism, so generated episodes should not support broad sociological claims about real populations.
A Related Benchmark Comparison
CULTURECONVERSE is positioned as a culturally grounded benchmark for multi-turn, help-seeking assistance across diverse Asian regions, identities, languages, and everyday domains. Its scenario design uses broad scaffolds whose cultural specificity emerges from persona and retrieved local knowledge rather than fixed scripts.
- Related benchmark comparison: CULTURECONVERSE addresses a benchmark gap by evaluating multi-turn, culturally grounded assistant help-seeking rather than only regional knowledge or single-turn etiquette.
- Persona universe: The persona universe spans ten East and Southeast Asian regions and combines salient ethnicity, religion, language, and regional background where publicly available.
- Persona universe: Equal identity sampling supports subgroup-level disparity analysis, producing 58 × 3 × 2 = 348 persona configurations across age cohorts and gender labels.
- Scenario universe: The scenario universe covers 7 broad domains and 42 subdomains to represent varied everyday situations without overfitting to manually authored scripts.
- Scenario universe: Broad subdomain descriptions act as flexible scaffolds, with cultural specificity emerging dynamically from persona pairing and retrieved local knowledge.
- Scenario universe: The domains include culinary practices, interpersonal dynamics, ritual and observance, economic and professional conduct, domestic and community life, and travel or personal expression.
KB samples injected into context
The generation pipeline retrieves broad cultural context and verified taboo constraints, then renders them into templates, blueprints, and ordered conversational shards. This separates public user information from hidden oracle state so assistants must infer culturally relevant constraints under partial observability.
- Template: A generic scenario template provides undefined slots and cultural dimensions, enabling diverse situations across demographic permutations without manual authoring for each case.
- Blueprint: The example instantiates a Singaporean host, a Malaysian client, communal-dining constraints, kitchen limitations, and inferred Muslim dietary and prayer considerations.
- Blueprint: The blueprint pairs a template with a persona and retrieved evidence, while separating the public initial message from oracle-only taboos, tripwires, and success criteria.
- Shards: Shard guidance requires brief, progressive user turns and evaluates whether the assistant avoids pork and alcohol, proposes feasible dishes, and accounts for prayer timing and arrival.
- Shards: Blueprints are decomposed into 3–5 ordered conversational shards, each specifying a private goal, cultural signifiers, and pass/fail transition logic.
E.4 Transcript and Evaluation Record Transcript and evaluation excerpt
The transcript and evaluation record captures a complete multi-turn episode alongside execution metadata, judge scores, and generation configuration. The surrounding pipeline uses structured prompts and hidden constraints to produce culturally grounded, multi-hop scenarios under partial information.
- Transcript and evaluation excerpt: A recorded example contains an eight-turn English episode in which all four shards are cleared.The run reports 4 / 4 shards cleared, with shard completion taking 1, 2, 2, and 3 turns.
- Transcript and evaluation record: A final episode record combines the complete dialogue with judge scores and execution metadata for transparent, machine-parseable analysis.These records support filtering high-quality trajectories for downstream fine-tuning.
- Pipeline configuration: The pipeline separates template, blueprint, shard, and user-simulation stages, with GPT-5 mini used across core generation and evaluation roles.The blind benchmark varies the target assistant model while GPT-5 mini remains the simulator, gold-mode assistant, and evaluator.
- Scenario generation: Templates begin as culture-agnostic everyday scenarios with variable slots, subtle cultural nuance, and multi-hop reasoning requirements.The design calls for 3–5 inference steps while avoiding over-constrained scripts and excessive cultural term stuffing.
- Blueprint constraints: Blueprints enforce public, private, and oracle-only separation while encoding harmful-action tripwires and challenge requirements.Initial messages may contain at most one essential niche cultural term, and global-English mode requires Latin script.
K.1 Use of Gold Mode for Training Data
Gold mode generates CultureConverse-DS training trajectories by giving the assistant privileged access to hidden scenario information, addressing missed cultural nuances, unsafe advice, and hallucinations during standard inference. The matched comparison reports improved assistance quality and safety with negligible impact on scenario realism.
- Gold-mode motivation: Gold mode grants the training assistant access to complete hidden information to generate high-fidelity cultural-competence trajectories.The configuration is intended to overcome missed nuances, unsafe advice, and hallucinations when assistants observe only public chat history.
- Gold-mode comparison: Positive ∆values in the matched GPT-5 mini comparison indicate Gold mode improves assistance quality and safety while barely affecting scenario realism.Table 10 summarizes Gold-mode versus blind generation on the matched test split using TDQ for scenario realism.
K.2 Ablations for GPT-5 mini as Judge
The judge ablation compares individual models with three-judge ensembles against human-annotated dialogues. GPT-5 mini is selected as the sole evaluator because ensembles provide only marginal alignment gains while reducing within-one agreement and increasing cost and complexity.
- Individual judge comparison: GPT-5 mini is the strongest individual judge, achieving 0.581 MAE and 90.1% within-one agreement against human consensus.The comparison covers 11 candidate judges and 165 possible three-judge ensembles.
- Slice robustness: GPT-5 mini remains close to the human–human agreement bandwidth across regional and metric slices.Tables 12 and 13 report agreement by evaluation metric and region.
- Evaluation prompts: The evaluation prompts dynamically populate transcript, persona, oracle, and metadata fields for six core metrics.Figures 20–25 provide the metric-specific prompt templates.
M Human Annotation and Judge Alignment
Human annotation evaluates six episode-level qualities on 1–5 Likert scales, while GPT-5 mini is compared with human consensus as an automated judge. The validation supports GPT-5 mini as a sufficiently aligned evaluator, while examples and breakdowns expose safety failures and subgroup disparities.
- Human annotation: Human validation covers 1,250 Gold-mode dialogues rated by 42 annotators across helpfulness, honesty, harmlessness, naturalness, plausibility, and typicality.Each item–metric cell receives three independent human scores.
- Agreement metrics: Within-one agreement measures scores within one Likert point, while MAE measures average absolute score distance and signed bias captures judge strictness.A negative signed bias means the judge scores more strictly than the human mean.
- Failure example: A low-scoring episode receives average 3H = 2.67 because recommending tuak without checking alcohol suitability produces Harmlessness = 1.The same episode scores Helpfulness = 2 and Honesty = 5, illustrating multidimensional evaluation.
- Challenge effects: Normative and contextual trap episodes consistently increase task difficulty, as indicated by negative performance deltas.The benchmark reports these effects in Table 14.
- Disparity analysis: The benchmark analyzes subgroup means and score ranges across model, region, identity, age, gender, language, and challenge type.Composite 3H and TDQ scores are computed as episode-level averages before subgroup aggregation.
O.1 Bootstrap Confidence Intervals for Headline Results
The paper reports bootstrap confidence intervals for headline metrics and finds a small, consistent reduction in 3H for trap-labelled episodes. It also describes the narrow fine-tuning pool used for transfer experiments.
- Bootstrap confidence intervals: 95% non-parametric bootstrap confidence intervals are reported for headline model families and the six metrics underlying Table 4.The intervals use episode-level resampling and are summarized in Table 17.
- Trap-stratum result: 0.04–0.08 points: trap-labelled episodes reduce 3H across evaluated models, with a mean decrease of −0.06.This reduction is reported relative to matched standard items.
- Transfer-training setup: 27,860 trajectories with perfect helpfulness, honesty, and harmlessness scores form the intentionally narrow LoRA SFT training pool.The setup tests transfer to static out-of-domain tasks without exposing their answer formats during training.
P.1 OOD Evaluation Suite
The out-of-domain evaluation suite tests whether interactive cultural fine-tuning transfers to static normative, regional, safety, and localisation tasks. It spans multiple datasets and modalities, while accompanying analyses examine model, language, regional, identity, age, and gender variation.
- Model-level analysis: Figure 33 represents each model-level heatmap cell as an assistant-level mean across all evaluated episodes.The figure compares component and family-average scores.
- Evaluation suite composition: 17 datasets and 29,828 examples comprise the OOD suite: 7 MCQ datasets with 11,988 examples and 10 CLS datasets with 17,840 examples.MCQ tasks target normative and regional knowledge, while CLS tasks evaluate localised safety-related capabilities.
- Language analysis: Figures 34 and 35 show native-minus-English language deltas overall and by region, where positive values indicate better native-language performance.The regional view is explicitly organized by model and region.
- Subgroup and slice analysis: Age and gender 3H variance remains negligible at ≤0.05 points, while region–identity is consistently the widest disparity axis and age and gender approach zero.The subgroup figures also include regional deviations, region–identity disparities, and identity-level means for GPT-5 mini.
- Score distributions: Figure 41 displays GPT-5 mini judge-score distributions separately for English and native-language slots.The figure supports comparison of score distributions across the two language conditions.
- OOD transfer metrics: Table 18 reports per-dataset OOD transfer using accuracy for MCQ tasks and macro-F1 for CLS tasks, with deltas measured in percentage points versus each base model.The comparison follows LoRA SFT on the 27,860-sample perfect-3H CultureConverse-DS subset.