Source-linked AI summary
LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review
Valdini Douglace Lemofouet, Blessing Ngozi Uzor, Paula Chikaodinaka Anyanwu, Danielle Blanche Kapsa, Sukairaj Hafiz Imam, P Sam Sahil, Abigail Oppong, Tassallah Abdullahi, Clemencia Siro, Idris Abdulmumin, Seid Muhie Yimam, Shamsuddeen Hassan Muhammad
TL;DR
Safety alignment evidence remains limited and uneven for low-resource and multilingual languages. This systematic review analyzes 50 studies using PRISMA 2020 and finds that alignment approaches and benchmarks generalize poorly across linguistic and cultural contexts.
Problem
Safety alignment methods, training data, and evaluation remain predominantly English-centered, limiting evidence about their effectiveness across low-resource linguistic and cultural contexts.
Method
The paper conducts a PRISMA 2020 systematic literature review of 50 studies identified through Semantic Scholar, arXiv, and OpenAlex.
Results
Safety alignment approaches generalize poorly across diverse linguistic and cultural contexts, while translated benchmarks often miss culturally grounded harms and multilingual adaptation can reduce low-resource safety performance.
Takeaways & Limitations
Progress requires native-language and culturally grounded benchmarks alongside alignment methods designed for multilingual and low-resource settings.
Takeaways & Limitations
Benchmark resources for African languages remain limited compared with those available for other multilingual regions.
Abstract
from arXiv · showhide
Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual settings than in high-resource languages. In this paper, we conduct a Systematic Literature Review (SLR) of LLM safety alignment in low-resource languages by adopting the PRISMA 2020 methodology. Out of roughly 1,500 papers identified from Semantic Scholar, arXiv, and OpenAlex, 50 relevant studies have been selected and analyzed. Our review is organized around four themes: safety alignment methods, multilingual safety risks, evaluation benchmarks, and cross-lingual transferability. We further propose a taxonomy of safety alignment approaches based on three adaptation mechanisms: data adaptation, objective optimization, and mechanistic alignment. Across literature, translated English benchmarks fail to sufficiently represent culturally rooted harms, and multilingual models are more vulnerable to cross-lingual jailbreaks, code-switching attacks, and safety degradation in underrepresented languages. These failures are driven by several key factors, including uneven multilingual pre-training coverage, insufficient native-language preference data, poor transfer of safety representations, and a lack of culturally aware evaluation frameworks. The review also notes that many low-resource languages, especially African languages, have fewer safety benchmarks available than other multilingual regions. Overall, the results reveal a persistent multilingual safety gap, and suggest that future progress will require culturally grounded benchmarks, participatory data collection, balanced multilingual pre-training, and scalable multilingual alignment methods.
1 Introduction
LLM safety alignment remains weaker in low-resource languages because methods and evaluations are concentrated in high-resource settings, especially English. This review synthesizes multilingual safety risks, alignment methods, benchmarks, and cross-lingual transfer using PRISMA 2020.
- Most alignment techniques and evaluation frameworks target English and depend on large annotated datasets, benchmarks, and human-feedback pipelines unavailable for many underrepresented languages.These resource constraints limit the applicability of established safety methods beyond high-resource languages.
- English safety alignment does not reliably transfer across languages, allowing models that reject harmful English prompts to comply with equivalent prompts in low-resource languages.Imbalanced multilingual pre-training coverage is identified as a factor related to these transfer failures.
- The review applies the PRISMA 2020 framework to synthesize research on safety alignment methods, multilingual safety risks, evaluation benchmarks, and cross-lingual transferability.It also identifies key challenges, methodological limitations, research trends, and future directions for safer and more inclusive multilingual models.
- The review examines how alignment methods adapt to low-resource languages, what multilingual safety risks and cultural harms emerge, and which benchmarks and evaluation frameworks exist.Its research questions also address factors affecting cross-lingual transfer of safety alignment.
2 Related Work
Prior work identifies a persistent multilingual safety gap but lacks formal, Africa-focused synthesis. This review addresses these limitations through a PRISMA-based review dedicated to African and low-resource languages.
- Existing reviews: Yong et al. (2025b) reviewed nearly 300 publications and found a persistent, growing language gap in LLM safety research.Banerjee et al. (2026) synthesized findings for Global South languages and reported sharply weaker safety guardrails on low-resource and code-mixed inputs.
- Limitations: Existing studies do not follow a formal SLR methodology with explicit research questions, inclusion and exclusion criteria, and a PRISMA flow.
- Limitations: Prior work subsumes African languages under generic low-resource categories and does not structure findings into actionable dimensions such as methods, risks, and benchmarks.
- Limitations: Existing studies do not synthesize mechanistic explanations for cross-lingual transfer failures.
- Novelty: This work presents the first PRISMA-based systematic literature review dedicated exclusively to LLM safety alignment in African and low-resource languages.
3 Methodology
The study used a systematic literature review based on PRISMA 2020 guidelines. Articles were collected from Semantic Scholar, arXiv, and OpenAlex using queries spanning technology, safety, and linguistic scope.
- Review protocol: The review followed the PRISMA 2020 guidelines, with its screening process documented in a flow diagram.The PRISMA flow diagram is presented in Fig. 1.
- Search strategy: Relevant articles were collected from Semantic Scholar, arXiv, and OpenAlex.
- Search strategy: Search queries covered technology, safety, and linguistic scope through terms including LLM, alignment, jailbreak, multilingual, and African languages.The linguistic-scope terms also included low-resource languages, cross-lingual, and code-switching.
4 Safety Alignment Methods for Low-Resource Languages
Safety alignment methods for low-resource languages are organized around data adaptation, objective optimization, and mechanistic alignment. Evidence highlights culturally grounded data, cross-lingual consistency objectives, and direct parameter or layer modification as complementary strategies.
- Data adaptation: Culturally grounded data pipelines use local-language generation, safety filtering, native-speaker validation, and parameter-efficient fine-tuning to improve alignment beyond translated English data.CultureGuard was trained on 386k samples across nine languages, including Hindi and Thai, and achieved zero-shot transfer to unseen languages.
- Method taxonomy: The review taxonomy groups methods by three transfer mechanisms: data adaptation, objective optimization, and mechanistic alignment.Data adaptation improves multilingual resources; objective optimization transfers safety behavior across languages; mechanistic alignment modifies internal model representations.
- Cross-lingual transfer: Cross-lingual transfer remains sensitive to filtering: unfiltered refusal distillation increased jailbreak success rates by up to 16.6 points, whereas filtering mitigated the degradation.The finding shows that transferred refusal behavior can harm safety when teacher responses contain ambiguous boundary refusals.
- Mechanistic alignment: Mechanistic approaches bypass extensive data collection or full retraining by editing safety-critical weights or transplanting safety-relevant transformer layers from high-resource models.Sparse Weight Editing projects harmful low-resource representations into safety subspaces, while layer transplantation improves MultiJail safety without reducing general-benchmark performance.
5 Safety Risks in Multilingual and Low resource Language Contexts
Safety vulnerabilities increase in low-resource and underrepresented languages, extending from translated prompts and multilingual jailbreaks to adaptation-induced degradation and code-switching attacks. These risks also include culturally contextualized harms, stereotypes, and biases that conventional benchmarks often miss.
- Cross-lingual attacks: 79% jailbreak success on GPT-4 followed translated harmful prompts in low-resource languages, while high-resource languages remained below 15%.This disparity indicates that safety alignment does not generalize evenly across linguistic space.
- Cross-lingual attacks: Low-resource languages have about 3 times the likelihood of encountering harmful content as high-resource languages, while instruction following also becomes less reliable.The reported risks include both unintentional and intentional multilingual jailbreak scenarios.
- Adaptation risks: Benign fine-tuning on new or synthetic languages can degrade safety alignment by interfering with previously established safety constraints.For African language adaptation pipelines, introducing new linguistic domains without safety recalibration may reduce robustness.
- Code-switching attacks: Code-switching bypass rates reached 67.23% on GPT-3.5 and 40.34% on GPT-4, with effects depending on linguistic distance and prompt structure.The risk is especially relevant where code-switching is a common communicative norm, including African contexts.
- Cultural harms: Culturally contextualized harms include unsafe Hausa outputs, meaning changes during translation, stereotype propagation, and culturally dependent biases missed by conventional benchmarks.Reported examples include toxic product recommendations, culturally inappropriate content, and hate-speech amplification.
6 Safety Evaluation Benchmarks
Safety benchmark development is shifting from translated English datasets toward native-language, culturally grounded, and multilingual adversarial evaluation. However, African languages remain substantially under-represented in available benchmark resources.
- Multilingual benchmark expansion: XSafety introduced 28,000 annotated instances covering 14 safety issues across 10 languages, while POLYGUARDPROMPTS combines multilingual human-LLM interactions with verified machine translations.These benchmarks represent the broader move toward more culturally diverse multilingual evaluation settings.
- Toxicity and policy evaluation: PolygloToxicityPrompts evaluates toxicity using 425K naturally occurring prompts spanning 17 languages, while ML-Bench provides policy-grounded evaluation across 14 languages using regional regulations.Together, these benchmarks target toxicity and policy-sensitive multilingual safety.
- Culturally localized evaluation: SEALSBench, SEA-SafeguardBench, SGToxicGuard, Qorgau, IndicSafe, and IndicJR extend evaluation toward regional socio-cultural risks, native-language prompts, adversarial behavior, toxicity, and culturally grounded harms.The cited benchmarks cover Southeast Asian, Singaporean, Kazakh-Russian, and South Asian settings.
- Regional coverage gaps: African-language benchmark resources remain limited: LSR evaluates refusal degradation in Yoruba, Hausa, Igbo, and Igala, Ubuntu-Guard introduces an African policy-based benchmark, and Uhura studies truthfulness and safety constraints.The review identifies African languages as substantially under-represented relative to other multilingual benchmark ecosystems.
- Overall trend: Overall, benchmark development is shifting toward native-language evaluation, culturally grounded harms, and multilingual adversarial testing, while African languages remain substantially under-represented.This synthesis captures the section’s overall result and limitation.
7 Cross-Lingual Transferability of Safety Alignment
Safety alignment is not language-agnostic: behaviors learned in high-resource languages often degrade in low-resource languages because safety-relevant representations and supervision are unevenly distributed. Recent work seeks to improve transfer through representation-level edits, mechanistic targeting, multilingual clustering, and cross-language consistency objectives.
- Transfer limitations: Safety alignment learned in high-resource languages often degrades in low-resource languages because safety-relevant knowledge remains concentrated in high-resource multilingual representations.The literature repeatedly identifies uneven multilingual pre-training coverage as the primary bottleneck behind cross-lingual transfer degradation.
- Transfer limitations: Morphologically rich languages and underrepresented scripts face fragmented tokenization and weaker semantic representations, reducing the reliability of safety reasoning and refusal behavior.These problems are especially visible in African and other low-resource languages, where pre-training data and alignment supervision are scarce.
- Representation and mechanistic transfer: Representation-level approaches improve transfer by editing lightweight parameters, replacing safety-critical transformer layers, or selectively targeting shared multilingual safety neurons.Mechanistic findings also indicate a shared directional structure in refusal behavior across safety-aligned languages.
- Scalable transfer: Efficient multilingual transfer can use carefully selected language subsets with representation clustering or multilingual consistency objectives that enforce alignment agreement across languages.These approaches aim to improve transfer without requiring extensive target-language supervision.
8 Discussion
Safety-alignment methodologies remain English-centric across data, alignment, and evaluation, producing a persistent multilingual gap. African languages are especially underrepresented, while translation-based benchmarks overlook culturally specific harms and can distort safety annotations.
- Safety-alignment development remains predominantly based on English-language data, spanning pre-training, alignment procedures, and evaluation techniques.
- RQ1 and RQ4: Parameter-efficient fine-tuning, data synthesis, and cross-lingual transfer remain poorly validated for African languages, with marginal gains when pre-training coverage is insufficient.
- RQ2: Cross-lingual jailbreak transfer, code-switching vulnerabilities, and safety degradation after benign fine-tuning are reported but seldom evaluated.
- RQ2: English-centric benchmarks neglect culturally specific risks, while annotation discrepancies and culturally biased benchmarks further weaken evaluation measures.
- RQ3: African languages remain underrepresented in emerging multilingual benchmarks, and translation-based dataset construction can manipulate meanings and mislead safety annotations.
- Inadequate pre-training coverage, weak representation, limited alignment capabilities, and poor benchmarks form a reinforcing cycle requiring simultaneous progress across all fronts.
9 Conclusion
The review synthesized 50 studies and concludes that safety alignment remains largely designed for high-resource languages, limiting generalization across diverse linguistic and cultural contexts. It also finds that translated benchmarks often fail to capture culturally grounded harms.
- Evidence base: 50 studies were synthesized on LLM safety alignment in low-resource language settings.The review examined research focused specifically on low-resource language contexts.
- Core finding: Current safety alignment approaches remain largely designed for high-resource languages and generalize poorly across diverse linguistic and cultural contexts.This conclusion identifies a persistent limitation in the applicability of existing alignment methods.
- Evaluation gap: Existing benchmarks rely heavily on translated datasets that often fail to capture culturally grounded harms.Translation-based evaluation is insufficient for representing harms rooted in local cultural contexts.
A Appendix … A.3 Further Analysis
The appendix documents the review’s database-search queries, LLM-assisted screening procedure, eligibility criteria, and further analysis of multilingual safety disparities. The analysis shows that English dominates research coverage, while low-resource and code-switched settings exhibit substantially greater safety risk and utility loss.
- A.1 Search Strings and Criteria: The appendix describes the database queries used to retrieve articles for the systematic literature review.These queries were applied across different databases.
- A.2 LLM Usage: LLM Screening Prompt: The screening workflow submitted each candidate record once to Claude Sonnet without cross-record context and parsed the JSON output programmatically.The model’s decision and justification were appended to the screening spreadsheet.
- A.2 LLM Usage: LLM Screening Prompt: The screening prompt assessed relevance to LLM safety alignment for low-resource and African languages using four research questions.The questions covered alignment methods, multilingual safety risks and cultural harms, datasets and benchmarks, and cross-lingual transfer.
- A.2 LLM Usage: LLM Screening Prompt: Eligibility required post-2020 English-language empirical studies focused primarily on LLM safety and covering or applying to low-resource languages.The exclusion criteria removed English-only studies, non-LLM work, peripheral safety topics, studies without empirical results, and inaccessible papers.
- A.2 LLM Usage: LLM Screening Prompt: The screening output recorded addressed research questions, inclusion and exclusion decisions, keep-or-remove status, confidence, and a one-sentence justification.The response format required valid JSON.
- A.3 Further Analysis: English dominates study coverage across benchmarks, alignment methods, and adversarial attacks, whereas low-resource languages receive minimal representation.The appendix links this coverage imbalance to disparities in multilingual safety alignment.
- A.3 Further Analysis: More than seven times higher cross-lingual jailbreak failure rates occur in low-resource languages than in high-resource baselines.This result is reported for Figure 3(A).
- A.3 Further Analysis: Low-resource and code-switched settings combine elevated harmful outputs with substantial utility loss, unlike high-resource languages’ lower-risk, lower-degradation cluster.The contrast is described as a separation in the risk–utility landscape.