Source-linked AI summary

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha

arXiv:2608.23026v1cs.CLcs.AI

TL;DR

Multilingual LLM audits risk conflating social bias, identity representation, and cross-cultural patterns with surface cues such as names and wording. The paper introduces a human-validated, multi-agent audit that separates these questions and finds systematic language- and task-dependent variation, while translation and name masking reduce source-language recognition. The evidence is bounded to three language-linked contexts and a limited marker sample, so it does not establish cultural authenticity or broad generalizability.

  • Problem

    Systematic evidence is limited on how sociocultural patterns vary across languages and whether apparent cross-cultural differences persist after surface cues are controlled.

  • Method

    The paper uses a human-validated, multi-agent audit with separate lenses for bias representation, identity-linked semantic separation, and cross-cultural patterns across explicit and implicit task conditions.

  • Results

    Bias representation varies across languages and tasks, while cue removal reduces identity-label prediction in English and Chinese more than in French; translation and name masking reduce source-language recognition.

  • Takeaways & Limitations

    Separating contextual relevance from language recognition provides a more falsifiable way to evaluate apparent cultural grounding without equating language with a bounded culture.

  • Takeaways & Limitations

    The cross-cultural analysis covers 3,926 marker-bearing narratives and validates 178 detected markers, so findings describe this corpus rather than cultural authenticity, causality, or broader generalizability.

Abstract

from arXiv · show

Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording can reveal the source language. Treating all these signals as evidence of cultural grounding may obscure potential biases. We present a human-validated, multi-agent audit that separates three questions: whether outputs reproduce social biases, whether identity groups are represented differently, and whether outputs reflect cross-cultural patterns. The study analyzes 89,253 outputs from 12 LLMs in English, French, and Chinese, spanning 18 occupations and three task conditions. We find that bias representation varies systematically across languages and tasks. Removing direct identity cues sharply reduces identity-label prediction in English and Chinese, but has a much smaller effect in French. Across all language-genre settings, the cultural context associated with the source language receives the highest average relevance score, with moderate agreement between automated and human ratings. However, the ability to identify the source language drops substantially after translation and again after masking names. Without these controls, multilingual audits may mistake surface cues for cultural understanding, leading to misleading conclusions about cross-cultural variation and bias. Our audit offers a practical framework for separating such shortcuts from more meaningful cross-cultural patterns.

1 Introduction

Sociocultural bias in multilingual LLM narratives can emerge through implicit roles, causal framing, and language-linked assumptions, while existing evaluations provide limited evidence about how these patterns vary across linguistic settings. The paper addresses this gap with a multi-agent audit spanning bias representation, identity-linked semantic separation, and cross-cultural patterns.

  • Motivation: Implicit narrative bias can shape assumptions about who is suited to particular occupations through repeated role and value patterns rather than explicit declarations.Such patterns include portraying some identities as rational protagonists and others as supportive or emotional laborers.
  • Research gap: Existing LLM bias evaluations often emphasize explicit sentence-level tests, leaving subtler biases in task allocation, persona settings, and nuanced language choices less examined.The paper motivates evaluating bias in more complex, contextualized generation.
  • Research gap: Systematic evidence remains limited on whether sociocultural patterns in generated narratives vary across languages and persist after surface cues are controlled.This gap concerns linguistic settings that encode distinct discursive norms and cultural traditions.
  • Approach: The audit organizes multilingual sociocultural evaluation around bias representation, identity-linked semantic separation, and cross-cultural patterns.These lenses examine socially patterned associations, identity-label recoverability after direct-cue deletion, and context relevance alongside source-language separability.
  • Approach: The study conceptualizes multilingual sociocultural evaluation as a unified, multi-level audit rather than a single test of explicit bias.The framework connects the three analytical levels within one experimental setup and multi-agent workflow.

2 Related Work

Prior LLM bias research has moved from controlled sentence-completion benchmarks toward implicit, contextualized generation, but multilingual audits add the problem of separating cultural patterns from linguistic and surface cues. This motivates measuring both representation and semantic separation while applying surface-cue controls.

  • Task formulations: LLM bias evaluations include explicit sentence-completion benchmarks and implicit tasks where bias is woven into stories or news reports.The two formulations differ in whether they test isolated associations or contextualized framing.
  • Implicit bias: Long-form generation can expose bias through semantic construction, evaluative framing, and task context that isolated metrics may miss.Language structure, grammatical gender, cultural context, and training-data composition are identified as relevant factors.
  • Evaluation approaches: Dynamic and agentic approaches provide infrastructure for diagnosing long-form behavior, extending earlier static bias evaluations.Examples include RUTEd, multi-agent debate frameworks, and interaction simulations with biased AI.
  • Multilingual audit challenge: Multilingual cultural audits face an identification problem because markers can reveal their source language through wording and named entities.The paper therefore treats surface-cue controls as part of evaluating apparent cross-cultural differences.

3 Methodology

The audit evaluates multilingual LLM outputs through three complementary lenses—bias representation, identity-linked semantic separation, and cross-cultural patterns—using coordinated agents and human validation. It compares controlled and contextualized tasks across languages, then applies explicit metrics and surface-cue controls.

  • Audit framework: The framework coordinates extraction, classification, cultural-scoring, and translation agents to support the three-lens audit.Agents identify protagonist labels and cultural markers, score markers across contexts, and prepare non-English markers for surface-cue controls.
  • Audit framework: Human studies validated cultural-marker annotations and protagonist-label extraction, with automated marker-inclusion decisions receiving 80.0% human endorsement.Weighted agreement between automated and human protagonist labels was 91.3% for EN, 82.6% for FR, and 81.6% for ZH.
  • Task and data design: The study uses an Explicit sentence-completion task and Implicit Story and News generation to contrast controlled bias measurement with contextualized generation.The Explicit task selects one gendered label for an occupation, while the Implicit task generates long-form narratives in two genres.
  • Task and data design: The evaluation covers 18 occupations and 12 LLMs across English, French, and Chinese, targeting 50 outputs per available language–task–model–occupation configuration.All models contribute English and Chinese cells, while nine contribute French cells.
  • Analytical measures: Bias representation is quantified with stereotype and balance metrics, aggregated across available occupations and analyzed using Type-III ANOVAs over language, task, interaction, and model.Unidentified labels remain in the denominator but contribute to neither gender numerator.
  • Analytical measures: Identity-linked separation uses multilingual E5-large embeddings and a grouped logistic probe, comparing full-text label recovery with a paired subset after deleting names and direct gender terms.ROC AUC is computed from out-of-fold predictions, with 95% intervals from 2,000 bootstrap resamples of model–occupation groups.
  • Analytical measures: Cross-cultural analysis averages marker relevance within narratives and evaluates source-language separability for original, translated, and named-entity-masked markers.Self-context advantage compares source-linked relevance with the mean relevance of the other two contexts.

4 Results

Across languages and task conditions, LLM narratives show systematic differences in bias representation and identity-linked recoverability. Cultural-context relevance favors source-linked contexts, while translation and name masking reduce source-language separability.

  • Bias Representation: Bias metrics vary by language and task condition, with English separating most clearly by task, French primarily by Balance, and Chinese showing overlapping moderate-to-high Stereotype.English Stereotype ranges from 0.5 to 0.8 in Explicit, 0.1 to 0.4 in Story, and approximately 0 in News; French Story has Balance below −0.4, while Chinese Stereotype ranges from 0.2 to 0.7.
  • Bias Representation: Language–task interactions remain detected across analyses, including estimates conditioned on female or male labels and trial-level analyses.Trial-level analyses detect language–genre and model terms with all p < 0.001.
  • Identity-Linked Semantic Separation: Deleting names and direct gender terms reduces identity-label probe AUC by 0.230–0.365 in English and 0.320–0.342 in Chinese, but only 0.003–0.037 in French.Full-text probe AUCs are 0.992–1.000, serving as a positive control because direct cues remain.
  • Cross-Cultural Patterns: Human raters endorse 80.0% of automated marker-inclusion decisions, while automated–human relevance agreement within two scale points is 0.641.Across 624 comparisons, quadratic-weighted κ is 0.280, MAE is 2.09, and marker-level Spearman correlation is 0.555.
  • Cross-Cultural Patterns: Source-linked contexts receive the highest mean automated relevance in every source-language row, with positive self-context advantages across all six language–genre cells.Mean self-context advantages are 2.42, 4.26, and 2.94 for English, French, and Chinese Story, and 2.64, 4.78, and 3.64 for News.
  • Cross-Cultural Patterns: Source-language separability falls from 0.991 to 0.721 after translation and to 0.558 after named-entity masking.The three-class chance level is 0.333, showing that surface-cue controls substantially reduce—but do not eliminate—separability.

5 Discussion

The discussion shows that sociocultural findings depend on language, elicitation conditions, and the construct being evaluated. Surface cues contribute to apparent cultural grounding, while human validation remains especially important for interpretive judgments.

  • Language and task dependence: The language–task interaction is the largest effect, with stereotype and balance patterns changing across English, French, and Chinese conditions.English Explicit outputs align strongly with stereotypes, English News approaches zero Stereotype with positive Balance, French varies mainly in Balance, and Chinese conditions overlap more.
  • Identity-linked semantic separation: Cue deletion reduces identity-label recoverability in English and Chinese but has a smaller effect in French, where grammatical agreement remains informative.The result indicates that direct-cue controls do not remove all identity-linked information equally across languages.
  • Cross-cultural patterns: Every language–genre cell shows a positive self-context advantage, yet translation and entity masking substantially reduce source-language separability.Names and wording therefore contribute to apparent cultural grounding in the audit.
  • Evaluation and validation: Cultural relevance is more interpretive than structured protagonist labeling: human–human agreement is moderate and automated–human agreement is lower.LLM judges provide consistent coverage but may share model priors, whereas human raters offer a more independent reference at higher cost.
  • Evaluation and validation: Human involvement is essential for defining constructs, exposing legitimate disagreement, and calibrating automated judges before scaling evaluations.The framework positions automated systems as extensions of a validated rubric rather than replacements for human judgment.

6 Conclusion

The audit separates social bias, identity representation, and cross-cultural patterns rather than treating them as interchangeable signals. It finds language- and task-dependent representation, cue-sensitive identity prediction, and reduced source-language recognition after controls.

  • Framework: The audit separates social bias, identity representation, and cross-cultural patterns into distinct analytical questions.This separation is the organizing contribution of the framework.
  • Main findings: Representation varies across languages and tasks, while cue removal reduces identity-label prediction much more in English and Chinese than in French.The reported asymmetry shows that identity-linked recoverability is language-dependent.
  • Main findings: Source-linked contexts receive the highest average relevance scores, but translation and name masking substantially reduce source-language recognition.These controls test whether apparent cross-cultural patterns persist when wording and names are weakened.
  • Implication: Names and wording can be mistaken for cultural understanding, so multilingual audits should evaluate these signals separately.The framework tests whether apparent cross-cultural patterns remain after surface cues are weakened.

Limitations

The study’s scope is limited by its language-linked corpus, uneven occupation coverage, modeling assumptions, and marker-based cross-cultural analysis. These boundaries constrain interpretation of generalizability, causality, and cultural authenticity.

  • Scope: The evidence covers three high-resource language-linked contexts—English, French, and Chinese—rather than bounded cultures.The comparisons are analytical contrasts among model outputs, not definitions of authentic cultural communities.
  • Data coverage: The dataset includes 18 occupations, two implicit genres, unequal model pools, and three protagonist labels, with partial coverage in four aggregate cells.Unidentified labels also affect some News estimates.
  • Methodological boundaries: The embedding probe measures identity-label recoverability rather than causal effects, and cue deletion may leave grammatical or indirect cues, especially in French.Aggregate estimates also depend partly on the ANOVA specification.
  • Cross-cultural analysis: The cross-cultural analysis covers 3,926 marker-bearing narratives, while human validation examines 178 detected markers and does not measure extraction recall.The probe translates individual markers rather than full narratives, so findings describe this corpus rather than cultural authenticity, competence, causality, or broad generalizability.

Ethical Considerations

The paper frames language-linked comparisons as diagnostic analyses of model behavior, not claims about authentic cultures. Its ethical practices include informed consent, participant compensation, and author verification of AI-assisted work.

  • Interpretive safeguards: Language-linked comparisons are analytical contrasts among model outputs, not definitions of authentic English, French, or Chinese culture.The paper avoids attributing observed patterns to inherent community values or prescribing culturally authentic narratives.
  • Interpretive safeguards: The findings are intended as diagnostic signals about model behavior rather than fixed categories of language groups or stereotypes.This limits how the reported sociocultural patterns should be interpreted.
  • Participant protections: Participants in both Prolific studies provided informed consent, and cultural-marker participants were compensated at least GBP 15 per hour.These procedures apply to the human validation studies.
  • AI use: Generative AI assisted with writing and icon editing, while the authors verified all claims, analyses, and citations.The statement describes the paper’s use of AI assistance and subsequent author verification.

B Experiment Materials

The study spans 18 occupations, three task conditions, and English, French, and Chinese outputs from 12 models. Cultural markers are organized into a ten-category taxonomy covering contexts such as places, institutions, names, cuisine, rituals, and arts.

  • Study design: 18 occupations are evaluated across Explicit Task, Story, and News conditions in English, French, and Chinese.The corpus includes outputs from 12 models, with exact model versions and language coverage reported separately.
  • Cultural-marker taxonomy: The classification agent maps retained cultural markers to ten categories, including toponyms, institutions, anthroponyms, cuisine, rituals, religion, mythology, and value orientations.The taxonomy also includes material culture and arts and media.
  • Cultural-marker taxonomy: Examples of toponyms include New York, Silicon Valley, Paris, the Latin Quarter, and Beijing.
  • Cultural-marker taxonomy: Examples of institutions and personal names include MIT, the Université de Paris-Sorbonne, the National Library of France, Ms. Brown, Monsieur Léon, and Grandma Li.
  • Cultural-marker taxonomy: The taxonomy also captures cuisine and rituals, including apple pie, croissant, Sichuan cuisine, Thanksgiving, Noël, and the Mid-Autumn Festival.
  • Cultural-marker taxonomy: Religion, material culture, arts and media, mythology, and value orientations include examples such as church, abacus, Shakespeare, the Little Prince, dragons, individualism, and collective dedication.

C Bias-Representation Robustness

The robustness analysis tests whether bias-related findings persist across model and denominator specifications, statistical formulations, and validated multilingual procedures. It also documents the sampling, ethics, and participant instructions underlying the audit.

  • Statistical robustness: The primary Type-III ANOVAs use all available outputs, while complementary specifications condition probabilities on female or male labels and repeat measures on nine common models.Trial-level binomial GLMs model protagonist labels and label identification with language–genre interactions, model, and occupation terms and cluster-robust covariance.
  • Statistical robustness: The language–task interaction remains significant for both aggregate metrics in every ANOVA specification.
  • Audit procedure: The extraction agent proposes text-grounded cultural markers, while the classification agent decides inclusion and assigns each included marker to a ten-category taxonomy.Markers are excluded when they do not meet the prespecified inclusion rule.
  • Sampling and coverage: Generation targets 50 outputs per occupation–language–task cell, with nine models represented in French and all 12 models represented in English and Chinese.French outputs are unavailable for three listed models.
  • Human validation: Participants received localized consent, comprehension checks, training, and instructions to judge markers and protagonist identity using only text-expressed evidence.Protagonist gender could not be inferred from occupation, name, nationality, or expectation.
  • Human validation: Cultural-marker evaluation produced 624 marker-level comparisons, with bootstrap intervals and within-language permutation baselines used for uncertainty and calibration.Table 3 reports agreement between automated and human cultural-marker relevance scores.

D.2.2 Protagonist-Label Study

The protagonist-label study evaluates human agreement with automated labels across languages and controls identity-probe comparisons through blinded narrative judgments and cue deletion. Its validation design includes stratified sampling, quality filtering, and geographic coverage reporting.

  • Study design: 120 implicit narratives per language are stratified across six Story/News-by-label cells, with automated labels hidden from raters.Participants labeled 12 blinded narratives before completing additional association, gold, attention, and demographic items.
  • Study design: 116 retained raters produced 1,392 formal judgments after consent, comprehension, training, gold-item, attention, completeness, and speed checks.The retained sample included 49 English, 36 French, and 31 Chinese raters.
  • Validation results: Human inter-rater reliability was Fleiss’ κ = 0.72–0.82, while weighted automated–human agreement was κ = 0.72–0.86.
  • Rater coverage: English, French, and Chinese responses span 5, 3, and 6 normalized country groups, respectively.
  • Cue-deletion analysis: The paired cue-deletion analysis samples 100 female- and 100 male-labeled narratives per language–genre cell and removes protagonist names and direct gender terms without placeholders.Full-text and cue-deleted probes use the same folds, with AUC changes bootstrapped at the model–occupation-group level.
Loading 2608.23026v1…