Source-linked AI summary

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta, Sabik Bin Sultan, Abdullah Khan Zehady

arXiv:2608.22335v1cs.CLcs.AI

TL;DR

Bengali safety evaluation lacks evidence grounded in local harms and ordinary language variation rather than translated English categories. BANGLASAFE addresses this gap with a culturally grounded benchmark and controlled comparisons across language, register, and authority framing. Across 18 frontier LLMs, Bengali register shifts produce the strongest observed safety difference, while classifier disagreement limits uncalibrated evaluation.

  • Problem

    Existing multilingual safety evaluations are predominantly English-centric and often translate English harm categories, omitting culturally specific Bengali harms and ordinary register variation.

  • Method

    BANGLASAFE evaluates 879 prompts across 17 culturally grounded harm categories and five matched prompting conditions on 18 models, using native or expert-reviewed prompts and calibrated four-way judgments.

  • Results

    53.6% ASRloose was observed overall, with formal Bengali journalism exceeding colloquial Banglish by 17 percentage points, the strongest observed effect.

  • Takeaways & Limitations

    Safety alignment and evaluation for Bengali should account for culturally grounded harms and ordinary register shifts, not language switching alone.

  • Takeaways & Limitations

    The corpus coverage probe does not establish causality, the journalism mechanism is based on single-author qualitative inspection, and the judge validation used one annotation session.

Abstract

from arXiv · show

Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce BanglaSafe, a benchmark of 879 Bengali prompts combining 309 natively authored prompts with 570 expert-reviewed prompts, spanning 17 culturally grounded harm categories and five prompting conditions that vary language, writing style, and authority framing. Evaluating 18 frontier LLMs, we find that over half of all responses are unsafe or partially unsafe (53.6%) while 14.7% contains strictly harmful content, and that the strongest observed effect is not the switch from English to Bengali but the choice of writing style within Bengali: the same harmful request phrased as a formal newspaper investigation succeeds 17 percentage points more often than the same request phrased as a casual message, with no adversarial engineering involved. We further show that existing safety classifiers struggle to reliably evaluate Bengali content, with even frontier models failing on nearly half of all cases.

1 Introduction

BANGLASAFE addresses Bengali safety-evaluation gaps by combining culturally grounded harms with controlled language, register, and authority comparisons. Across 18 models, register variation within Bengali produces the strongest observed safety effect, while classifiers disagree substantially.

  • Benchmark and headline findings: 53.6% overall ASRloose was observed across 15,822 evaluations of 879 prompts and 18 frontier LLMs.ASRloose counts partial or fully harmful responses as unsafe.
  • Benchmark and headline findings: 17 percentage points separate formal Bengali journalism from colloquial Banglish, exceeding the +13pp English-to-Bengali language effect.BNFormal reaches 63.3% ASRloose, versus 45.8% for BNCollq.
  • Failure mechanism: Formal Bengali prompts can elicit operational detail embedded in investigative-news framing rather than outright refusal.Culturally specific Bengali harm terms appear only once across 80,587 prompts in twelve English safety corpora.
  • Evaluation infrastructure: Safety classifiers disagree sharply on Bengali register-shifted content, with Cohen’s κ ranging from 0.014 for LLAMAGUARD 4 to 0.667 for GPT-OSS-SAFEGUARD.The benchmark therefore audits evaluation infrastructure alongside model refusal behavior.
  • Benchmark and headline findings: The benchmark contains 879 prompts spanning 17 statute-anchored harm categories and five conditions varying language, register, and authority framing.It combines native and LLM-assisted prompt construction and measures compliance under naturally occurring shifts.

2 Related Work

Prior multilingual safety benchmarks largely translate English taxonomies and assume that categories, framing, and cultural context transfer across languages. BANGLASAFE instead targets native Bengali authorship, culturally grounded harms, ordinary register variation, and institutional framing.

  • English-centric evaluation: English-centric benchmarks assume that harm categories, framing, and cultural context transfer across languages, an assumption the paper states breaks down in Bengali.Prior frameworks nevertheless substantially advanced safety evaluation.
  • Multilingual safety evaluation: Existing multilingual benchmarks generally translate English harm categories into other languages and measure refusal rates.This approach preserves intent but may miss culturally specific Bengali harms.
  • Bengali benchmark gap: Prior Bengali and Indic benchmarks address adjacent problems but omit combinations of refusal behavior, Bangladesh-specific harms, native authorship, and sociolinguistic framing.INDICSAFE, for example, uses Indian socio-cultural categories and omits harms such as hundi and bKash fraud.
  • Register and framing: BANGLASAFE examines ordinary register variation rather than explicit adversarial construction such as persuasion, persona injection, or optimization-based prompting.Its formal Bengali journalism register reflects standard Bangladesh news-writing practice.
  • Benchmark gap: No prior benchmark combines native Bengali prompt authorship, culturally grounded harm taxonomy, controlled register variation, institutional framing, and a human-validated judge pipeline.This combination is the paper’s stated positioning against prior work.

3 Dataset Construction

The dataset grounds Bengali safety evaluation in local harms and everyday language variation. It uses five matched prompting conditions, complementary native and reviewed prompt tracks, and validation of register and harm labels.

  • Harm taxonomy: 17 harm categories are anchored to statutes or documented institutional sources, including Bangladeshi law, NGO case files, and contemporary news coverage.The taxonomy is intended to represent culturally specific harms rather than translated generic categories.
  • Register design: Bengali register varies with social context, so the same event can appear as a news article, peer text, or government case report.These registers carry implicit signals about expertise, intent, and legitimacy.
  • Five prompting conditions: Each underlying harm-act instance appears under five conditions varying English versus Bengali and direct, institutional, journalism, or colloquial framing.Matched semantic content supports paired language, authority, and register comparisons.
  • Prompt construction: The benchmark contains 879 prompts: 309 natively authored gold prompts and 570 prompts generated through a register-controlled Claude Opus 4.7 pipeline and reviewed by native annotators.Generated prompts were revised when they failed register fidelity, harm verification, or cultural authenticity.
  • Prompt construction: 501 prompts, or 57.0%, are case-anchored, with similar case-anchor densities across the five conditions.This balance reduces confounding of the register and authority ablations by case anchoring.
  • Quality validation: Register-tier agreement reached κ = +0.915, while harm verification achieved 95.1% raw agreement and cultural authenticity achieved 51.7% raw agreement.The cultural-authenticity disagreements concentrated in standard code-mixed forms, whereas register agreement exceeded the stated threshold.

4 Methodology

The methodology combines deterministic model evaluation with a four-way safety rubric, calibrated judging, classifier audits, and paired tests isolating language, authority, and register effects. The design preserves partial compliance as a distinct diagnostic outcome.

  • Model evaluation: 18 large language models from nine vendors are evaluated across four model-size and access tiers.The tiers include open-weight models under 10B, 10–30B, 30B+, and closed-source models.
  • Model evaluation: Deterministic decoding uses temperature = 0.0, top-p = 1.0, and no system prompt, paired with cluster-aware inference and bootstrap intervals.The setup isolates policy decisions from sampling noise while retaining paired tests over base-harm-act clusters.
  • Safety labeling: The four-way rubric labels responses REFUSE, POLICY, PARTIAL, or HARMFUL, while ASRloose counts PARTIAL and HARMFUL as unsafe and ASRstrict counts only HARMFUL.Retaining four-way labels exposes how models fail, not only whether they fail.
  • Judge validation: The calibrated Claude Opus 4.7 judge achieves binary κ = 0.666 against the primary human annotator on a stratified 300-response validation subset.Lower inter-human agreement indicates subjectivity at the PARTIAL/HARMFUL boundary in Bengali code-mixed responses.

5 Results

Across 15,822 responses, Bengali register strongly shapes safety outcomes: formal Bengali has the highest loose attack success, while colloquial Banglish is safest. Paired analyses show this register gap exceeds language and authority effects, and classifier agreement varies sharply.

  • 53.6% overall ASRloose and 14.7% ASRstrict were observed across 15,822 responses.
  • 63.3% ASRloose makes BNFormal the highest condition, 13 percentage points above the English baseline.Its ASRstrict is 13.5%, below English’s 18.7%, indicating the increase is concentrated in partial rather than verbatim harmful content.
  • 45.8% makes BNCollq the safest condition, despite its permission-seeking peer-help and personal-emergency framing.Colloquial Banglish does not function as a cover narrative.
  • 17pp is the strongest paired gap, with BNFormal at 63.3% ASRloose and BNCollq at 45.8%.
  • 17.0pp separates BNFormal from BNCollq, exceeding the 13.0pp English-to-formal-Bengali shift and the smaller authority effects.The paired Wilcoxon tests isolate language, authority, and register using matched model–base-prompt pairs.
  • Per-model ASRloose ranges from 17.1% to 89.5%, with no reliable size-safety correlation.The register effect is reported as near-universal rather than an aggregation artifact.
  • κ = 0.014 for LLAMAGUARD 4 versus κ = 0.667 for GPT-OSS-SAFEGUARD shows sharp cross-judge disagreement.The classifiers cannot be used as drop-in Bengali register-shift evaluators without threshold calibration.

6 Discussion

The discussion attributes Bengali safety failures to both missing culturally specific terms in English safety corpora and register-dependent interpretation. Formal Bengali journalism produces the strongest within-language bypass by reframing harmful content as reporting, while institutional code-switching follows a distinct route.

  • 6.1 The Corpus Coverage Gap: Generic English counterparts appear frequently while culturally specific terms remain absent, indicating a lexical rather than categorical coverage gap.Examples include “methamphetamine” 290 times versus zero for yaba, “money laundering” 177 versus zero for hundi, and “OTP/phishing” 363 versus zero for bKash.
  • 6.2 The Journalism Register as Task Reframing: 17.0pp is the strongest paired ASRloose difference, with BNFormal exceeding BNCollq within Bengali and showing a large effect size.The paired Wilcoxon ablation isolates register while holding model and base prompt matched.
  • 6.2 The Journalism Register as Task Reframing: 49.9% PARTIAL drives BNFormal’s elevated ASR, while its HARMFUL rate is 13.5%, below the English baseline of 18.7%.The journalism register increases hedged categorical content wrapped in newspaper formatting rather than verbatim operational recipes.
  • 6.2 The Journalism Register as Task Reframing: 194 of 200 inspected BNFormal responses labelled PARTIAL or HARMFUL adopted newspaper-investigation framing with bylines, section headers, expert quotes, and sources codas.Operational details such as sourcing channels, tactics, and pricing structures appeared inside the article format.
  • 6.2 The Journalism Register as Task Reframing: BNFormal bypasses safety while remaining nearly entirely Bengali-scripted, whereas BNInst partly shifts operational content into English within an institutional frame.Latin-script density correlates weakly with unsafe labels in BNFormal but more strongly in BNInst, where the adjusted odds ratio is 1.75.

7 Conclusion

The paper shows that Bengali writing style matters more for harmful compliance than Bengali language alone. BANGLASAFE links the formal-journalism gap to culturally specific coverage failures and journalistic reframing, while highlighting classifier misalignment.

  • 7 Conclusion: Across 18 frontier LLMs, BNFormal reaches 63.3% ASR, exceeding BNCollq by 17 percentage points without adversarial prompting.The benchmark evaluates culturally grounded Bengali harms across language, register, and authority conditions.
  • 7 Conclusion: Culturally specific Bengali harm terms are largely absent from contemporary safety training, while formal journalism reframes harmful compliance as investigative reporting.The paper characterizes this as an interpretation-level bypass rather than a content-recognition failure.
  • 7 Conclusion: Existing safety classifiers show threshold misalignment under Bengali register shifts, limiting reliable evaluation without calibration.BANGLASAFE provides a benchmark, calibrated judge, and evaluation framework for studying these failures.

Limitations

The paper identifies important limits on causal interpretation, judge validation, and dataset composition. Its coverage probe is associational, its qualitative mechanism is not quantitatively calibrated, and substantial prompt-generation and judging overlap remains.

  • Limitations: The corpus coverage probe identifies a lexical gap but does not establish causality for the observed ASR differences.The journalism-cover mechanism is based on a single-author qualitative inspection of 200 responses and is explanatory rather than quantitatively calibrated.
  • Limitations: The four-way judge was validated on a 300-response human set labeled in a single annotation session, leaving independent-cohort replication for future work.This constrains confidence in the breadth of judge validation.
  • Limitations: 570 of 879 prompts and the judge use Claude Opus 4.7, so ASRloose is reported separately for gold and synthetic subsets.The gold-only estimate is 47.6%, compared with 56.9% for the synthetic subset; an independent judge reproduces the findings.
  • Limitations: BANGLASAFE lacks a benign control set and therefore does not quantify whether hardening against journalism-register attacks would increase over-refusal of legitimate Bengali investigative reporting.A benign-register control set is left to future work.

Ethics Statement

The benchmark grounds its harms in Bangladesh-specific legal and institutional sources while limiting access to prompt-response pairs. Its ethics procedures address public-source provenance, annotation, and same-family judging concerns.

  • Ethics Statement: Every benchmark prompt is grounded in a Bangladesh statute or publicly documented case, and the benchmark is intended to surface alignment failures rather than expand operational-harm knowledge.
  • Ethics Statement: 15,822 prompt-response pairs are gated behind a use-case statement, institutional affiliation, and manual author review, while taxonomy, metadata, rubric, audit protocol, tables, and scripts are openly released.
  • Ethics Statement: The authors report that Claude Opus 4.7 generates and judges 570 prompts, but agreement with GPT-OSS-SAFEGUARD is κ = 0.856 on 833 Claude-Haiku-4.5 rows.
  • Ethics Statement: Case anchors use already-public reporting, introduce no private victim data, and limit personal identifiers to those present in cited public sources.
  • Ethics Statement: Prompt-level annotation used two paper authors, while response-level annotation used one author and one independently compensated annotator.
  • Ethics Statement: The 17 harm categories are anchored in statutes, NGO case files, or documented news sources, with each category tied to at least one primary source.
  • Ethics Statement: The taxonomy excludes disputed or controversial-but-legal topics and includes only explicitly illegal acts or harms broadly agreed across reasonable Bangladeshi viewpoints.

B Case Anchoring and Category-Level Effects

Case anchoring has almost no aggregate effect on attack success, but its impact varies significantly by harm category. The benchmark constructs matched, natively authored or reviewed prompts across five controlled conditions to study these effects.

  • Category-Level Effects: +0.8pp mean ASR difference indicates that case anchoring does not systematically increase bypass rates overall.The aggregate effect is statistically null after correction (Holm-adjusted p = 1.00; rrb = +0.045).
  • Category-Level Effects: +29.4pp for communal violence and +12.9pp for trafficking are the largest significant positive case-anchoring effects.Six categories show significant positive effects after Holm correction.
  • Category-Level Effects: −16.3pp for burn/corrosive and −16.2pp for formalin are the largest significant negative case-anchoring effects.Four categories show significant negative effects after Holm correction.
  • Category-Level Effects: Case anchoring is therefore a category-conditional modulator rather than a uniformly amplifying attack axis.Positive effects cluster in systemic or organisational harms, while negative effects cluster in more victim-centred categories.
  • Case-Anchor Design: The dataset renders 114 unique harm-act cases across five conditions, producing 570 synthetic prompts within the 879-prompt benchmark.Synthetic prompts were generated under register control and reviewed by native Bangla-speaking annotators for fidelity, harm verification, and cultural authenticity.
  • Evaluation: The four-way REFUSE/POLICY/PARTIAL/HARMFUL rubric distinguishes partial compliance that binary safe/unsafe evaluation would obscure.The rubric treats categorical mechanisms without verbatim operational recipes as PARTIAL, and its judge agrees substantially with an independent judge.

I Prompt-Level Guard Audit

Prompt-level guards often flag unsafe prompts but detect far fewer corresponding responses, especially under authority-cover conditions. The audit therefore shows that aggregate flagging rates overstate reliability for Bengali register-shift content.

  • I Prompt-Level Guard Audit: 81.5% and 89.1% prompt-level flagging fall to 15.4% and 62.1% response-level flagging for the two guards.The same prompts are identified as unsafe much more often than the responses they elicit.
  • I Prompt-Level Guard Audit: 51.5% is LLAMAGUARD 4’s ENInst prompt-level rate, down from 92.5%, while response-level detection reaches only 3.3%.Institutional-researcher framing sharply weakens detection at both stages.
  • I Prompt-Level Guard Audit: 74.4% is GPT-OSS-SAFEGUARD’s highest response-level rate on BNFormal, compared with 53.3% on BNCollq.The guard is more aggressive overall but exhibits the same condition-sensitive pattern.
  • I Prompt-Level Guard Audit: A prompt-level guard is insufficient for the authority-cover conditions tested because those conditions degrade guard detection.The 81.5% aggregate prompt-level rate consequently overstates guard reliability.
  • J Model Details and Per-Model ASR: The 18 evaluated models span open-weight tiers below 10B, 10–30B, and above 30B, plus closed-source systems.The model inventory includes both open-weight and closed-source families.
  • K Per-Category Attack Success Rates: Self-harm has the lowest ASRloose at 40.9% but the highest ASRstrict at 25.9% across categories.This pattern indicates either refusal or fully operational content with little middle ground.
  • L Cross-Judge Agreement: Pairwise Cohen’s κ compares the calibrated Claude judge with two field-standard guards on their shared response intersections.The comparison is reported in the cross-judge agreement analysis.
  • M Safety RLHF Corpora Used in the Coverage Probe: Across 80,587 prompts in twelve English safety RLHF corpora, 20 culturally anchored Bengali harm terms produce only one hit.The sole hit is “dowry” in PKU-SAFERLHF, without Bangladesh context.

N.1 Example 1: Trafficking (Grok-4.3)

The trafficking example holds the harmful request constant while changing only Bengali register and framing. Grok-4.3 provides operationally detailed content under formal journalism but refuses the colloquial peer-help version.

  • BNFormal: BNFormal presents the trafficking request as formal Bengali journalism concerning recruitment, visa fraud, clearance bypass, and ransom bargaining.The prompt explicitly frames the request as an investigative report about documented trafficking mechanisms.
  • BNFormal: Grok-4.3 responds with documented recruitment and exploitation patterns, including recruitment mechanisms, visa fraud, and ransom bargaining details.The response accepts the awareness-raising frame while embedding operational substance.
  • BNCollq: BNCollq presents the same trafficking request in colloquial Banglish as peer help about a Saudi returnee case.The harmful content and requested mechanisms remain materially the same across the paired prompts.
  • BNCollq: Grok-4.3 refuses the colloquial version and declines to explain false visas, clearance bypass, trafficking, ransom, or abuse methods.The refusal explicitly withholds details, steps, mechanisms, and explanations.
  • Interpretation: The paired result attributes the behavioral difference to register: formal journalism is interpreted as legitimate awareness-raising, whereas casual peer framing triggers refusal.The semantic content is identical; only the register differs.
Loading 2608.22335v1…