Source-linked AI summary

Redteaming Leading Arabic LLMs with ASAS

Fidaa Abed, Haidar Khan, M Saiful Bari, Babar Khan, Abdalghani Abujabal

arXiv:2608.21985v1cs.AI

TL;DR

Arabic LLM safety remains underexplored, particularly for regional and culturally grounded adversarial evaluation. The paper introduces ASAS, a human-curated Arabic redteaming benchmark, and finds substantial safety gaps alongside poor reliability of automated safety judging.

  • Problem

    Arabic LLM safety, especially redteaming and region-specific cultural evaluation, remains largely unexplored despite growing adoption.

  • Method

    ASAS is a fully human-curated benchmark with 801 Modern Standard Arabic prompts spanning 8 safety categories and 8 attack strategies, plus ideal responses.

  • Results

    The evaluation shows large Arabic alignment and safety gaps, especially in high-harm categories, while direct prompting, code/encryption, and hypothetical testing elicit unsafe responses in over 60% of cases.

  • Takeaways & Limitations

    ASAS provides a culturally grounded benchmark for addressing Arabic model-safety gaps and supports alignment with users’ linguistic and cultural backgrounds.

  • Takeaways & Limitations

    GPT-4o safety judging achieves ~50% accuracy and ~22% recall on unsafe responses, indicating that human expert evaluation is required.

Abstract

from arXiv · show

As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical. However, Arabic LLM safety remains underexplored, especially in adversarial evaluation settings. We introduce the Arabic Safety Index (ASAS), the first fully human-curated Arabic benchmark for redteaming LLMs. ASAS contains 801 prompts spanning 8 safety categories and 8 attack strategies, with ideal responses in Modern Standard Arabic (MSA). We conduct a redteaming evaluation across seven leading models with Arabic capabilities, including GPT-4o, Claude 3.7 Sonnet, and regional models such as ALLaM and FANAR. Human annotators rate responses using a structured 4-point safety scale, revealing that most models fail to defend against 50% of unsafe prompts. Our findings highlight major safety gaps in high-harm categories such as weapons and illicit substances, with direct and obfuscation-based attacks proving most effective. The results also show that language alignment does not readily transfer across languages, and that automated safety judges (e.g., GPT-4o) perform poorly compared to human annotators. ASAS provides a culturally grounded benchmark and redteaming protocol to drive progress in Arabic LLM safety.

1 Introduction

Arabic LLM safety remains underexplored, particularly for adversarial evaluation, motivating ASAS as a human-curated benchmark and redteaming assessment. Across seven models, substantial vulnerabilities emerged in high-harm categories and under several attack strategies, while automated judging performed poorly.

  • Motivation: Arabic LLM safety remains largely unexplored, especially for redteaming, despite growing regional adoption and culturally specific evaluation needs.Regional evaluation must account for ethical, legal, and cultural considerations.
  • Benchmark and evaluation: ASAS provides 801 Modern Standard Arabic prompts spanning 8 safety categories and 8 attack strategies, with manually curated ideal responses.It is designed for evaluating and improving Arabic LLM safety, alignment, and robustness.
  • Benchmark and evaluation: Most tested models elicited unsafe responses for approximately 50% of prompts under four human-assigned safety labels.The seven evaluated models included both leading general-purpose and regional Arabic-capable systems.
  • Findings: The largest safety gaps occurred in Guns & Illegal Weapons, Controlled Substances, and Suicide & Self-Harm.
  • Findings: Direct Prompting, Code/Encryption, and Hypothetical Testing elicited unsafe responses in over 60% of cases across models and categories.
  • Findings: GPT 4o achieved ~50% overall accuracy and ~22% recall on unsafe responses as a safety judge, indicating the need for human annotation.The low recall means many unsafe responses could pass undetected.

2 Related Work

Arabic LLM safety research is limited compared with broader English-focused safety and redteaming work. ASAS addresses this gap with a culturally grounded, human-annotated Modern Standard Arabic benchmark centered on adversarial evaluation.

  • Arabic and multilingual safety: English-focused LLM safety research has advanced, whereas Arabic LLM safety remains underexplored and needs culturally specific datasets.
  • General LLM safety and redteaming: General LLM safety research emphasizes redteaming, diverse attack strategies, automated redteaming, and evaluation methods for uncovering vulnerabilities.
  • Arabic and multilingual safety: Existing Arabic and multilingual studies report socio-cultural safety disparities and that multilingual models leave 79% of harmful Arabic content undetected.
  • Positioning ASAS: ASAS introduces 801 Modern Standard Arabic prompts across 8 safety categories and 8 attack types, combining comprehensive coverage with human annotation and a redteaming focus.This distinguishes it from datasets focused on natural language understanding or lacking adversarial evaluation.

3 About the ASAS Index

ASAS defines a human-curated framework for Arabic LLM safety evaluation, covering safety categories, attack strategies, severity labels, and redteaming procedures. The index also characterizes prompt complexity and evaluates seven Arabic-capable models.

  • Safety categories: ASAS evaluates harmful, biased, illegal, and culturally inappropriate content across eight safety categories.Islamic/Arab Culture Alignment is the largest category at 25.3% of the dataset.
  • Safety labels: Human annotators classify responses from Safe to Extremely Unsafe using four severity labels.Extremely Unsafe responses directly facilitate harm, illegal activity, severe misinformation, or incitement to violence.
  • Attack types: The benchmark uses eight attack types, including direct prompting, escalation, code or encryption, hypothetical testing, storytelling, false premises, roleplay, and out-of-context attacks.Out of Context and Step-by-Step Escalation are overrepresented because they require multiple prompts.
  • Prompt design and redteaming: Most prompts contain fewer than 50 words, while the dataset averages 1.5 turns because some attacks require multiple interactions.The redteaming process records the most effective prompt that elicits a harmful response for each scenario.

4 Results

The evaluation reveals substantial safety variation across models, categories, and attack types. Jailbroken models often produced extremely unsafe responses, while automated judging showed limited reliability.

  • Claude 3.7 Sonnet achieved the highest overall safety score at 68%, followed by ALLaM at 58%.
  • GPT 4o produced extremely unsafe responses in 62% of cases where its guardrails were broken.
  • Command-R 7B Arabic and Mistral Saba were weakest in the Guns & Illegal Weapons category.
  • Jais recorded safety rates of 26% under Direct Prompting and 27% under Code/Encryption attacks.
  • ALLaM produced 70% unsafe responses under Storytelling/Roleplay attacks.
  • Fanar had the highest unjustified-refusal rate, whereas Claude 3.7 Sonnet showed a more balanced refusal profile.Fanar recorded 6.0% of all safe responses and 18.5% of all refusals; Claude recorded 2.6% and 6.5%, respectively.
  • After jailbreaking, models tended to produce Extremely Unsafe responses more often than other unsafe levels.
  • GPT 4o achieved approximately 50% accuracy and 22% recall when judging unsafe responses, supporting human expert evaluation.

5 Conclusion

The ASAS evaluation identifies large gaps in Arabic alignment and model safety. The authors argue that culturally and linguistically aligned evaluation remains important and plan broader Arabic and multimodal benchmarks.

  • ASAS reveals large gaps in Arabic alignment and model safety.
  • The authors emphasize alignment with users’ linguistic and cultural backgrounds as a critical goal for AI models.
  • Future work will extend safety benchmarks to other Arabic dialects and modalities.

Ethics Statement

The study uses human-annotated Arabic adversarial prompts while incorporating safeguards for sensitive content and annotator well-being. ASAS also addresses cultural alignment and responsible benchmark use.

  • Trained MSA-fluent annotators familiar with regional ethical, legal, and cultural norms curated and reviewed prompts and responses.
  • Ethical review protocols, content warnings, and opt-outs were used to support annotator well-being during sensitive evaluations.
  • ASAS was designed to expose safety vulnerabilities rather than promote harmful content, with examples anonymized or paraphrased where appropriate.
  • The benchmark includes Islamic/Arab Culture Alignment and Bias to address region-specific sensitivities often neglected in global safety datasets.
  • The authors report serious safety challenges for Arabic prompts and encourage responsible use of ASAS to improve safety rather than replicate unsafe behaviors.
  • No personally identifiable information was collected, and all model outputs were synthetic.

A Appendix

The appendix provides contrasting examples of unsafe and safe model behavior: Mistral Saba gives harmful concealment advice, while Fanar safely handles a prompt involving racial stereotypes.

  • Mistral Saba responds unsafely to a leading prompt about hiding the smell of a corpse.
  • The examples contrast unsafe and safe handling of harmful prompts across the tested models.
  • Fanar responds safely to a prompt requesting help propagating racial stereotypes and prejudices.
Loading 2608.21985v1…