Source-linked AI summary
'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection
Fawzia Zehra, Kara-Isitt, Sonal Khosla, Stephen Swift
TL;DR
Urdu is largely absent from mainstream LLM safety evaluation and WOAH proceedings despite its 246 million speakers. The paper audits this gap and tests five LLMs across six datasets and script conditions. It reports uneven cross-script safety assurance, with label instability of 15.9%–31.6% and Missed-in-Urdu rates of 2.4%–9.9%.
Problem
Urdu remains largely absent from mainstream LLM safety evaluation and WOAH proceedings despite 246 million speakers and exposure to online harm.
Method
The paper audits 205 ALW/WOAH papers and evaluates five LLMs across six datasets spanning Nastaliq Urdu, Roman Urdu, English, and code-switched Urdu-English.
Results
Label instability ranges from 15.9% to 31.6% and Missed-in-Urdu rates from 2.4% to 9.9%, while the audit confirms zero dedicated Urdu papers across nine editions.
Takeaways & Limitations
Current LLMs provide uneven safety assurance across Urdu script varieties, with open-weight models performing substantially worse than frontier models on both measures.
Takeaways & Limitations
The same model performs translation and classification, so results cannot distinguish translation quality from classification consistency.
Abstract
from arXiv · showhide
Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings. To investigate whether this absence has measurable consequences for content moderation reliability, five large language models, GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5, and Llama-3.1, were tested across six datasets spanning Nastaliq Urdu, Roman Urdu, English, and code-switched Urdu-English. Across the five Urdu-script datasets, label instability between original-script and English-translation classification ranged from 15.9% (Gemini 2.5 Flash) to 31.6% (Qwen-2.5), with a 'Missed-in-Urdu' rate, content flagged as harmful in English translation but passed as normal in the original script, ranging from 2.4% to 9.9% (median 4.3%). A complete enumeration of all 205 papers across nine ALW/WOAH editions via the ACL Anthology API confirms zero dedicated Urdu papers across the entire period. Results indicate that current LLMs provide uneven safety assurance across Urdu's script varieties, with smaller open-weight models showing substantially higher instability and missed-harm rates than frontier closed models.
1 Introduction
Urdu is a major but persistently neglected language in online-harm research and mainstream LLM safety evaluation. The paper audits this gap, evaluates cross-script classification consistency, identifies missing resources, and tests whether model differences are systematic.
- Research gap: 246 million Urdu speakers remain largely absent from mainstream LLM safety evaluation and WOAH proceedings.A complete audit of 205 papers across nine ALW/WOAH editions found no work substantively engaging with Urdu data.
- Research questions: The paper asks whether WOAH contains sufficient Urdu online-abuse research and whether LLM classifications vary across Nastaliq, Roman Urdu, English, and code-switched inputs.It also examines missing harm dimensions and whether model differences reflect systematic label shifts rather than random variation.
- Contributions: The study audits ALW/WOAH proceedings from 2017–2025 and evaluates five LLMs across six datasets and multiple script conditions.The evaluation measures label instability and a Missed-in-Urdu rate, while statistical tests validate observed model differences.
- Contributions: The paper identifies a structural resource gap: no dedicated Nastaliq Urdu-English code-switched hate-speech dataset exists.It also proposes statistical validation of cross-model differences as part of its contributions.
2 Background and Related Work
Urdu matters for online-safety research because a large, digitally active population faces documented online harm while its scripts and datasets remain underrepresented. Existing multilingual LLM evaluations and WOAH work leave Urdu outside substantial coverage.
- Script diversity: Nastaliq and Roman Urdu are distinct writing systems, while Urdu-English code-switching introduces additional variation in social-media text.Nastaliq is a right-to-left Perso-Arabic calligraphic style; Roman Urdu is a Latin-script transliteration.
- Why Urdu matters: Urdu combines 246 million speakers, substantial social-media activity, and documented exposure to hate speech and threatening content.Pakistan-focused reports identify religious minorities, political opponents, and journalists among targets of online hate.
- Representation gap: Urdu and its Nastaliq social-media variants are underrepresented in large-scale pretraining corpora relative to the speaker population.This underrepresentation coexists with a documented concentration of online harm in Urdu’s primary national context, Pakistan.
- Related work: An evaluation spanning eight non-English languages covered five top-ten languages but left Urdu, Mandarin, Bengali, and Russian unevaluated.The same related work found that prompt design affects hate-speech detection and that languages benefit from different prompting strategies.
- Related work: Prior work found translating inputs to English before classification outperformed prompting in the original language, providing a direct precedent for comparing original and translated inputs.Hausa research offers a structural analogue within WOAH and frames low-resource-language gaps as systemic rather than isolated.
3 Methodology
The methodology evaluates five LLMs on six datasets under controlled script and translation conditions, using a unified three-class taxonomy and divergence-based consistency metrics. Paired statistical tests assess whether observed differences are systematic.
- Datasets: Six datasets spanning three scripts are evaluated, with a Roman Urdu-English emotion corpus serving as a proxy because no dedicated Nastaliq Urdu-English hate-speech dataset exists.The proxy’s absence is treated as a finding in its own right.
- Datasets: All datasets are mapped to Hate, Offensive, and Normal, requiring conversion from binary or fine-grained source label spaces.Table 1 reports source labels normalized to H/O/N.
- Experimental conditions: Four conditions isolate script, transliteration, and code-switching while holding semantic content constant across models.C1 uses original Nastaliq Urdu, C2 classifies the same model’s English translation, and the proxy condition evaluates mixed-script robustness.
- Models: Five LLMs perform zero-shot classification with the same standardized prompt and no in-context examples.The models are GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5-7B-Instruct, and Llama-3.1-8B Instant.
- Metrics: Label instability is the proportion of instances whose classifications differ between C1 and C2.Missed-in-Urdu is the proportion harmful under C2 but Normal under C1, representing harm passed in the original Urdu script.
- Statistical testing: Paired binary outcomes use McNemar’s test, three-class outcomes use Stuart-Maxwell, and multiple comparisons receive Bonferroni correction.These tests account for paired classifications of the same text across models or conditions.
4 Results
Results cover WOAH/ALW Urdu coverage, cross-script classification consistency, dataset gaps, and statistical reliability. The study finds no dedicated Urdu paper across the reviewed proceedings, measurable cross-script label divergence, missing Nastaliq code-switched resources, and statistically tested model differences.
- Coverage of Urdu in WOAH Literature: 205 papers across nine ALW/WOAH editions were enumerated via the ACL Anthology API to assess Urdu coverage.The search used titles and abstracts with controlled language keywords and represents the indexed record.
- Coverage of Urdu in WOAH Literature: Urdu received zero dedicated papers across the entire reviewed period, despite being identified as urgently requiring technological support in 2022.Several less widely spoken European languages had multiple dedicated papers, while Bengali was also absent.
- Cross-Script Classification Consistency: 0.0% instability on the control dataset was observed for all five models by construction, attributing remaining instability to language and script rather than a pipeline issue.The control condition uses an English gold-standard dataset.
- Cross-Script Classification Consistency: 18.0% median label instability and 4.3% median Missed-in-Urdu were observed across models and datasets when equivalent content was presented in Nastaliq Urdu versus English translation.Every tested model assigned different moderation labels under the two presentation conditions.
- Missing Harm Dimensions in Urdu Resources: No dedicated Nastaliq Urdu-English code-switched hate-speech dataset exists, so the Roman Urdu-English RU-EN Emotion corpus serves as the closest proxy.The proxy uses emotion rather than harm labels, limiting direct measurement of this input variety.
- Missing Harm Dimensions in Urdu Resources: 1,216 instances were labelled Hate by original annotators but Normal by all five models under C1, while possible annotation bias remains unreviewed at scale.The paper leaves qualitative review of whether these cases reflect annotation bias, model failure, or both for future work.
- Statistical Significance: McNemar, chi-square, and Stuart-Maxwell tests with Bonferroni correction were used to assess instability, Missed-in-Urdu, and three-class label shifts.The tests used α = 0.005 across ten pairwise comparisons.
5 Conclusion
Across five Urdu-script datasets, LLM safety labels vary between original Urdu and English translation, while Urdu remains absent from dedicated WOAH research. The paper identifies this uneven coverage and cross-script inconsistency as measurable safety gaps.
- 15.9%–31.6% label instability and 2.4%–9.9% Missed-in-Urdu rates were observed across five Urdu-script datasets and five models.
- Open-weight models performed substantially worse than frontier models on both label instability and Missed-in-Urdu rates.
- Zero dedicated Urdu papers appeared across nine ALW/WOAH editions, despite Urdu’s large speaker population and an explicit 2022 invitation.
- The paper calls for Nastaliq Urdu-inclusive, English code-switched datasets and Urdu-inclusive safety benchmarks.
Limitations
The evaluation has several scope and interpretation constraints, including incomplete sampling, proxy labels, unreviewed disagreements, and dependence on provider APIs.
- N = 853–854 for one dataset because an interrupted run reached 85% of its target.
- The same model translated inputs and classified translations, so results cannot distinguish translation quality from classification consistency.
- Binary source datasets lack an Offensive class, which may inflate apparent instability.
- The RU-EN Emotion corpus proxies code-switched hate speech, with emotion labels mapped to harm labels as an approximation.
- The ACL Anthology audit searched titles and abstracts, and API-based evaluation may not remain reproducible as providers version their models.
Ethical Considerations and use of AI
The study reports responsible handling of harmful content and separates AI assistance for formatting and figures from the authors’ research responsibilities.
- Datasets were public or registration-accessible, no personally identifying information was stored or reported, and harmful content was analysed only for research.
- Claude Sonnet 4.6 assisted with LaTeX formatting and figure generation, while the authors retained responsibility for design, analysis, interpretation, and conclusions.
A Prompt Templates
The experiment used standardized moderation and translation prompts across models, comparing original Nastaliq inputs with model-generated English translations.
- Two prompt templates were applied identically across all five models and datasets.
- The classification prompt required one label—HATE, OFFENSIVE, or NORMAL—with explicit operational definitions.
- The user classification turn placed each post inside a labeled Post field.
- The C2 translation prompt instructed the same model to preserve tone, intensity, and meaning without sanitizing the Urdu text.
- Responses lacking a recognized label were recorded as REFUSED and excluded from instability calculations.
B Additional Experimental Details
The study documents dataset coverage, exclusions, and statistical tests used to assess cross-script label shifts and pairwise model differences. Figures and tables summarize dataset/model timing, label flows, instability controls, cleaned counts, and significance tests.
- Data cleaning: Error rows from API failures and refusal rows from failed C2 translations were excluded from all analyses.The exclusions covered 13 error rows and 29 refusal rows in total.
- Experimental materials: The release timeline places all datasets between 2021 and 2024 and all models between mid-2024 and early 2025.Model colours are consistent across the paper’s figures.
- Label-flow analysis: Figure 6 aggregates C1-to-C2 label flows across five Urdu-script datasets and shows more cross-flow for open-weight models, especially from Hate and Offensive into Normal.Cross-flow width represents the degree of label instability.
- Controls: HateXplain shows 0.0% label instability across all five models, providing a pipeline control for language- and script-related instability.The remaining five datasets are the relevant comparison set in the figure.