Source-linked AI summary
Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs
Namya Bhatnagar
TL;DR
English-centric safety alignment leaves limited evidence about whether LLM safety generalizes to Indian regional and code-mixed languages, an important gap for multilingual spoken technologies. The paper introduces INCLUDE to evaluate Indian socio-cultural bias across six languages and finds that open-source and closed-source models exhibit opposite cross-lingual patterns.
Problem
Safety filters and fairness checks receive less rigorous coverage in low-resource and code-mixed Indian languages than in English, limiting meaningful algorithmic protection for non-English speakers.
Method
INCLUDE evaluates culturally grounded stereotype prompts across six languages and measures Indian-centric bias across six axes in ten open- and closed-source models.
Results
Open-source models showed lower bias for English than Indian-language prompts, while closed-source embedding models showed the reverse pattern, with English producing stronger stereotype associations.
Takeaways & Limitations
INCLUDE provides a framework for systematically detecting cross-lingual cultural bias and supports more culturally informed evaluation of multilingual AI systems.
Takeaways & Limitations
Replacing the GPT-4o-based translation pipeline with professional human translators would strengthen the benchmark’s cross-lingual validity.
Abstract
from arXiv · showhide
Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken dialogue systems may produce stereotype-reinforcing outputs, bypassing the standard English-focused safety alignments and propagating harmful bias to non-English speaking communities. For spoken language technologies deployed across India's linguistically diverse population, this represents a critical failure mode. To address this cross-lingual gap, we introduce INCLUDE (Indian Cultural Lens for Understanding and Detecting Embedded Biases), a multilingual evaluation benchmark designed to quantify Indian-centric socio-cultural biases. INCLUDE consists of 2,604 prompts spanning six prompt languages: English, Hindi, Bengali, Marathi, Tamil, and Hinglish (Hindi-English code-mix). We evaluate ten open- and closed-source LLMs against this benchmark, analyzing 14,988 bias scores. Our statistical results reveal two key findings. First, Bengali yielded the highest average bias score in open-source models. Second, English demonstrated a notable reversal, producing the lowest bias in open-source models but the highest bias in closed-source models.
I. INTRODUCTION · II. RELATED WORK
The paper frames a cross-lingual safety gap in which English-aligned models can propagate stereotypes in low-resource and code-mixed Indian languages, then introduces INCLUDE to measure this problem across six languages and ten models. Related work shows that existing English and Indian-centric benchmarks leave multilingual and code-mixed safety underexamined.
- I. INTRODUCTION: English-focused safety filters and fairness checks are rarely extended equally to low-resource or code-mixed languages, allowing models that pass English audits to propagate embedded stereotypes.Safety filters include post-training alignment practices such as Constitutional AI.
- I. INTRODUCTION: Voice assistants integrating LLM backends disproportionately expose low-literacy users, children, and speakers of low-resource or code-mixed languages to underlying model biases.Typing in non-Latin scripts such as Devanagari can be cumbersome, motivating reliance on voice interfaces.
- I. INTRODUCTION: INCLUDE evaluates culturally grounded stereotype prompts across six languages: English, Hindi, Bengali, Marathi, Tamil, and Hinglish.The benchmark evaluates six open-source models through log-likelihood scoring and four closed-source embedding models through the Word... method described in the passage.
- I. INTRODUCTION: All five low-resource Indian languages elicit significantly higher bias than English in open-source models, revealing safety-alignment asymmetries in the Indian context.This is presented as the paper’s first direct contribution to spoken language technologies.
- I. INTRODUCTION: Closed-source embedding models show a significant reversal: English prompts elicit the strongest stereotype associations, implicating pretraining corpus composition.These models are used in spoken-language retrieval pipelines.
- II. RELATED WORK: StereoSet and CrowS-Pairs quantify bias effectively in English but are unsuitable for non-Western contexts, while Indian-BhED and IndiBias cover only one or two languages, respectively.Other Indian-centric frameworks evaluate bias across ten Dravidian languages, but the passage does not specify their broader coverage beyond that figure.
- II. RELATED WORK: Hinglish remains particularly neglected because spelling is not standardized and code-switch training data are sparse, creating a persistent challenge for multilingual systems.The passage identifies this structural under-representation as a plausible explanation for lower stereotype activation in open-source models relative to monolingual languages.
- II. RELATED WORK: Prior studies find that safety-aligned models generate unsafe responses more often in low-resource languages and that high-resource-language RLHF yields minimal safety improvement elsewhere.Those studies do not evaluate code-mixed languages, a limitation INCLUDE addresses directly.
III. METHODOLOGY · A. Defining Bias Segments · B. Expert Validation
The methodology defines Indian-centric socio-cultural bias through six axes and 27 identity groups, then validates these segments through iterative review by a 20-person expert panel.
- A. Defining Bias Segments: The benchmark defines six Indian-centric socio-cultural bias axes: Caste, Intersectionality, Region, Religion, Socioeconomic Status (SES), and Spoken Language.These axes were selected based on their documented prevalence in scientific, cultural, and historical literature.
- A. Defining Bias Segments: The six bias axes contain 27 identity groups representing populations commonly targeted by stereotyping within their respective axes.The passage gives Brahmins within the Caste axis as an example of an identity group.
- A. Defining Bias Segments: Identity groups are defined as populations commonly targeted by stereotyping within a specific bias axis.For example, Brahmins are identified within the Caste axis.
- B. Expert Validation: A 20-person Expert Panel of schoolteachers, university faculty, and native speakers reviewed the bias segments for stereotypical accuracy.The panel composition combined educational professionals and native-language expertise.
- B. Expert Validation: The bias segments were revised iteratively in response to panel feedback.The revisions addressed attributes that conflated caste with ethnicity or were considered too archaic for contemporary discourse.
- B. Expert Validation: Attributes conflating caste with ethnicity or judged too archaic for contemporary discourse were removed or reworded.These changes were examples of adjustments made after expert review.
- B. Expert Validation: Anonymous feedback was recorded, and the questionnaire collected no personally identifiable information.The validation process therefore documented feedback without collecting personally identifiable information.
C. Prompt Construction · D. Prompt Translation
The study constructs cue-free prompts tailored to open- and closed-source models, then translates and validates them across five low-resource Indian languages. Translation quality is assessed through weighted agreement, within-1-point consistency, and back-translation checks.
- C. Prompt Construction: Cue-free cloze prompts and target-attribute tuples were designed to elicit stereotyping quantitatively while attributing model preferences to training data and architecture.The prompts were intended to be comprehensive, lexically natural, and free of overt stereotype cues.
- C. Prompt Construction: Closed-source tuples instantiate the target identity group, stereotypical attribute, anti-stereotypical attribute, and prompt language.Distinct templates were required because closed-source models provide limited architecture transparency.
- D. Prompt Translation: The complete prompt set was translated into Hindi, Bengali, Hinglish, Tamil, and Marathi using an automated GPT-4o pipeline.Native bilingual expert-panel annotators independently reviewed translations for stereotype-intensity preservation on a 5-point Likert scale.
- D. Prompt Translation: κ = 0.977 for Hinglish and κ = 0.927 for Tamil indicated almost-perfect inter-annotator agreement on translation quality.Agreement was measured using linearly weighted Cohen’s κ on a five-point rating scale.
- D. Prompt Translation: κ = 0.639 for Hindi, κ = 0.585 for Marathi, and within-1-point agreement rates of 88.9% or higher showed broad annotator alignment.The reported κ interpretations classified Hindi as substantial and Marathi as moderate agreement.
- D. Prompt Translation: Back-translation checks evaluated each closed-source tuple term with binary judgments to ensure non-English embedding-based bias scores reflected language-native semantic representations.Annotators either confirmed each term’s fit or proposed a more contextually accurate alternative.
E. Prompt Evaluation · F. Statistical Analysis
The evaluation used language-specific scoring procedures for open- and closed-source models. Statistical analysis examined repeated-measures bias-score differences across six prompt languages using complementary non-parametric methods.
- E. Prompt Evaluation: Open-source models were evaluated with conditional log-likelihood scoring for each cloze-style prompt.The procedure followed the method described in Section IV.
- E. Prompt Evaluation: Closed-source models were evaluated with WEAT cosine similarity scoring for each target-attribute tuple.The procedure followed the method described in Section V.
- F. Statistical Analysis: 14,988 raw bias score observations were analyzed to assess whether differences across prompt languages were statistically reliable.The analyses accounted for the repeated appearance of the same prompts across six languages.
- F. Statistical Analysis: Standard ANOVA was not used because repeated prompts across languages violate its independence assumption and could inflate Type I error rates.A linear mixed-effects model was therefore introduced instead.
- F. Statistical Analysis: Overall language-group differences were assessed with the Friedman test, a non-parametric repeated-measures alternative.The test reports Kendall’s W (0-1) and a p-value for the null hypothesis of no group differences.
- F. Statistical Analysis: Kendall’s W (0-1) measures group-level consistency, with higher values indicating stronger consistency.The accompanying p-value indicates the probability of observing the result under the null hypothesis.
- F. Statistical Analysis: p < 0.05 throughout indicated statistically significant overall results.This threshold was reported for the Friedman-test analyses.
- F. Statistical Analysis: Wilcoxon signed-rank tests with Holm correction identified the specific language pairs responsible for pairwise differences.These were conducted as post-hoc comparisons following the overall group assessment.
IV. OPEN-SOURCE EVALUATION METRICS
The open-source pipeline evaluates model preferences among stereotype, anti-stereotype, and meaningless completions using cloze-style loss scores. Bias scores compare stereotype and anti-stereotype naturalness, clip negative values, and aggregate into Average Bias Scores across axes, languages, and models.
- Loss-based evaluation: Cloze-style prompts measure each model’s surprise for stereotype, anti-stereotype, and meaningless candidate completions using loss scores.A low loss indicates that the model finds a completion natural.
- Bias-score aggregation: For each prompt language and target bias axis, the pipeline computes bias scores across 216 individual result files.The files are organized by model, prompt language, and target bias axis.
- Bias-score aggregation: Bias scores interpret δi > 0 as stereotype preference and δi < 0 as anti-stereotype preference.Negative values are clipped to zero before aggregation so anti-bias does not cancel stereotype preference.
- Bias-score aggregation: Average Bias Scores average six per-axis scores per model, then average those results across all six open-source models for each prompt language.This produces a language-level ABS from per-axis and per-model aggregates.
- Bias interpretation: ABS magnitude is categorized as < 50.0 negligible bias, 0.50-0.99 mild bias, 1.00-1.49 moderate bias, or ≥1.50 strong bias.These thresholds define the interpretation of the aggregated score.
V. CLOSED-SOURCE EVALUATION METRICS
Closed-source models are evaluated with WEAT cosine-similarity scores rather than internal loss scores. Interpretation relies on established significance thresholds, while WEAT and open-source loss-based scores are compared only through relative structure, not absolute magnitude.
- WEAT-based measurement: WEAT measures bias by computing cosine similarity between target and attribute embedding vectors.Words are represented as vectors in high-dimensional space, with similar meanings clustering and opposing meanings diverging.
- WEAT-based measurement: Score ≈−1 indicates opposing vectors and an anti-stereotypical association.Targets such as Dalit are paired with stereotypical and anti-stereotypical attributes.
- Significance thresholds: < 0.01 is negligible and not reported; ≥0.05 indicates statistically significant bias; ≥0.10 indicates extreme bias.These significance thresholds follow established conventions.
- Cross-metric comparison: WEAT scores and open-source loss-based scores use different numerical scales and cannot be directly compared by absolute magnitude.Only relative structure is compared across cohorts, with each metric interpreted in its architecture’s native idiom.
VI. RESULTS … 2) Statistical Validation: Language Effects:
Across 13,716 open-source-model bias scores, prompt language significantly affected elicited bias, with Bengali highest on average and English uniquely lower than every other language. Hinglish statistically grouped with Indian languages rather than English, while Bengali was the only significant internal difference among those languages.
- 1) Average Bias Scores by Prompt Language:: 1.46 ABS was the highest average bias score, observed for Bengali across six prompt languages in open-source models.The remaining averages were Hindi 1.41, English 1.40, Marathi 1.39, Tamil 1.38, and Hinglish 1.30.
- 1) Average Bias Scores by Prompt Language:: All six prompt languages fell within the moderate bias range of 1.00-1.49.This range follows the ABS interpretation thresholds established in Section IV.
- A. Open-Source Models: Generative Bias Patterns: 0.38 SD for Bengali indicated greater variability in its bias scores than its average alone captured.The passage describes Bengali’s standard deviation as elevated, suggesting dispersed bias responses.
- 2) Statistical Validation: Language Effects:: 13,716 bias scores supported a significant overall language effect in open-source models, with Friedman W = 0.074 and p < 0.001.The test established reliable variation in bias across prompt languages.
- 2) Statistical Validation: Language Effects:: English was the only language significantly different from every other language, eliciting lower bias than Bengali (padj = 0.002), Hindi and Marathi (padj < 0.001), and Hinglish and Tamil (padj = 0.001).Pairwise Wilcoxon tests used Holm correction.
- 2) Statistical Validation: Language Effects:: Hinglish differed significantly only from English (padj = 0.001), while its comparisons with Bengali, Hindi, Marathi, and Tamil were all non-significant (padj = 1.000).Despite sharing English’s ABS threshold, Hinglish’s statistical distribution aligned with the low-resource Indian-language group.
- 2) Statistical Validation: Language Effects:: Bengali differed significantly from Hindi (padj = 0.040), the only significant internal difference among the Indian languages.All remaining Indian-language pairs were statistically indistinguishable, with padj = 1.000 for all pairs.
3) The Hinglish Exception: · B. Closed-Source Models: Embedding Bias Patterns
Hinglish preserves bias patterns associated with low-resource Indian languages, as both syntactic structure and Indian-cultural referents remain influential. For open-source models, language-level ABS is computed by averaging six bias scores per language.
- 3) The Hinglish Exception:: Syntactic structure retains bias patterns characteristic of low-resource Indian languages.
- 3) The Hinglish Exception:: Indian-cultural referents also retain bias patterns associated with low-resource Indian languages.
- 3) The Hinglish Exception:: The retained bias patterns are characteristic of low-resource Indian languages.
- 3) The Hinglish Exception:: Six bias scores are reported per language for each open-source model.
- 3) The Hinglish Exception:: The six language-specific bias scores are averaged to produce the ABS.
- 3) The Hinglish Exception:: ABS is defined for the respective prompt language.
1) Language-Level Variation: … 4) Effect Size Evidence:
Closed-source embedding models showed significant cross-lingual variation, with English producing the highest bias and reversing the open-source pattern. This reversal was linked to distinct mechanisms and supported by a practically meaningful effect size.
- 1) Language-Level Variation:: W = 0.334, p = 0.005 confirmed significant overall language-level variation across 1,296 closed-source embedding bias scores.The language effect was stronger than in open-source models, where W = 0.074.
- 1) Language-Level Variation:: Closed-source models showed a markedly stronger language effect than open-source models, with W = 0.334 versus W = 0.074.This indicates that prompt language was more decisive for closed-source embedding bias than open-source generative bias.
- 2) The English Reversal:: English embedding-based bias scores were significantly higher than Indian prompt languages, especially Bengali, with padj = 0.029.This directly reversed the open-source pattern, in which Indian-language prompts produced the greatest bias.
- 2) The English Reversal:: English elicited the highest bias scores among the prompt languages in closed-source models.The remaining language pairs showed no significant differences, with padj ≥0.328 for all comparisons.
- 3) Two Distinct Mechanisms of Cross-Lingual Bias:: Open-source gaps reflected English-focused post-training safety alignments, whereas closed-source gaps reflected English-dominant pretraining corpora encoding stronger stereotype associations.The two model cohorts therefore exhibited fundamentally different sources of cross-lingual bias.
- 4) Effect Size Evidence:: r = 0.495 was the mean rank-biserial correlation for English targets across all four closed-source models.The medium-to-large effect indicated that the English reversal was practically meaningful beyond statistical significance.
C. Summary of Cross-Lingual Bias Findings · VII. CONCLUSION AND FUTURE WORK
The findings show that cross-lingual bias varies by prompt language and model type, reflecting interactions between English-centric alignment, pretraining composition, and latent cultural priors. INCLUDE provides a culturally informed benchmark for identifying these biases and motivates broader multilingual and multimodal evaluation.
- C. Summary of Cross-Lingual Bias Findings: In open-source models, English prompts suppress bias relative to all Indian-language prompts, while Bengali elicits the highest bias.Bengali and Hindi are statistically distinguishable from English and each other.
- C. Summary of Cross-Lingual Bias Findings: Hinglish produces a hybrid outcome with the lowest ABS while remaining statistically grouped with other prompt-language results.
- C. Summary of Cross-Lingual Bias Findings: In closed-source embedding models, English targets show stronger stereotype associations than Bengali, Hindi, and Tamil.The reversal points to pretraining corpus composition as an additional source of cross-lingual bias manifestation.
- C. Summary of Cross-Lingual Bias Findings: Cross-lingual bias varies by prompt language, especially between English and low-resource Indian languages.
- VII. CONCLUSION AND FUTURE WORK: The multilingual safety gap is attributed to Western-dominated pretraining corpora and English-centric alignment practices.Bias emerges through interactions among representation density, mitigation strength, and latent-prior activation under uncertainty.
- VII. CONCLUSION AND FUTURE WORK: INCLUDE offers a comprehensive benchmark for systematically identifying and mitigating harmful biases affecting marginalized communities.It provides a practical framework for detecting bias across languages and cultural contexts and supports more equitable voice assistants.
- VII. CONCLUSION AND FUTURE WORK: Future work should extend bias evaluation to multimodal and image-generation systems beyond text-only settings.Benchmark development should also incorporate geographically and culturally nuanced frameworks spanning global contexts beyond India.
- VII. CONCLUSION AND FUTURE WORK: Additional studies should investigate bias in other low-resource and code-mixed languages within the Indian linguistic landscape.