Source-linked AI summary
Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap
Salima Lamsiyah, Ruslan Mitkov
TL;DR
Arabic NLP remains under-explained because current XAI is methodologically narrow, concentrated on classification, and often disconnected from Arabic linguistic and sociocultural structure. This critical structured survey synthesizes the literature through a multidimensional taxonomy and proposes linguistically grounded evaluation and research directions. Its central conclusion is that Arabic XAI must explain Arabic-specific evidence and context, not only model outputs.
Problem
Arabic XAI has method, task, and linguistic gaps: it often uses limited post-hoc attribution for classification while under-explaining Arabic-specific phenomena and broader NLP tasks.
Method
The paper conducts a critical structured survey and organizes reviewed Arabic XAI work through a taxonomy spanning explanatory targets, tasks, methods, linguistic units, and evaluation practices.
Results
The survey finds a dominant pattern of output-level explanation for classification, especially sentiment, harmful-content detection, fake news, and spam, leaving generation, retrieval, QA, summarization, translation, structured prediction, and LLM applications under-explained.
Takeaways & Limitations
Arabic NLP needs explanations faithful to Arabic as a linguistic, cultural, and sociotechnical object, supported by Arabic-aware explanation units and evaluation.
Takeaways & Limitations
The survey is a critical structured synthesis rather than an exhaustive systematic review and includes only studies explicitly framed around explainability, interpretability, transparency, or diagnostic analysis.
Abstract
from arXiv · showhide
Explainable AI (XAI) is now a major theme in NLP; however, Arabic NLP remains under-explained in three connected senses. First, there is a method gap: Arabic XAI relies heavily on a small set of post-hoc techniques such as LIME, SHAP, attention visualization, and saliency, while broader NLP XAI offers richer diagnostic, counterfactual, probing, rationale-based, and human-centered methods. Second, there is a task gap: existing Arabic XAI work is concentrated in classification tasks, especially sentiment analysis, hate/offensive language detection, fake news, and spam, with weaker coverage of generation, retrieval, translation, summarization, structured prediction, and dialogue. Third, there is a linguistic gap: many explanations identify influential tokens, but rarely explain Arabic-specific phenomena such as morphology, clitics, dialectal variation, diglossia, orthographic ambiguity, diacritics, code-switching, named entities, cultural references, or Classical and religious registers. This critical structured survey synthesizes the reviewed literature on Arabic XAI across text, speech, and multimodal settings. We argue that Arabic NLP does not only need explanations of model decisions; it needs explanations that are faithful to Arabic as a linguistic, cultural, and sociotechnical object. We introduce a taxonomy of tasks, methods, linguistic units, varieties, goals, and evaluation practices, and propose a research agenda for linguistically grounded Arabic XAI.
1 Introduction
Arabic NLP has advanced rapidly, but its explainability remains limited by methodological, task, and linguistic gaps. The survey argues for explanations faithful to Arabic’s linguistic, cultural, and sociotechnical characteristics.
- Arabic’s morphological richness, diglossia, dialectal diversity, and orthographic variability make surface-token explanations difficult to interpret.
- Arabic XAI relies heavily on LIME, SHAP, attention visualization, and saliency, although these represent only a subset of available NLP XAI methods.
- Arabic XAI is concentrated in classification tasks, especially sentiment, hate or offensive language, fake news, and spam detection.
- Many explanations identify influential tokens without clarifying Arabic-specific evidence such as morphology, clitics, dialectal variation, diglossia, diacritics, or named entities.
- The survey frames the central need as explanations faithful to Arabic as a linguistic, cultural, and sociotechnical object, rather than explanations of model decisions alone.
- Its contributions include a four-level explanatory account, a critical taxonomy and coverage table, and an agenda for faithful, useful, reproducible Arabic XAI.
2 What Should an Explanation Explain in Arabic NLP?
The survey distinguishes several explanatory targets in Arabic NLP rather than treating explainability as a single objective. These targets range from predictions and model mechanisms to linguistic structure and sociocultural context.
- Arabic explainability encompasses distinct prediction-level, model-level, linguistic, and socio-cultural explanatory targets.
- Prediction-level explanation: Prediction-level explanations account for labels, scores, retrieved results, or generated outputs, often using local feature attribution.
- Model-level explanation: Model-level explanations examine internal representations, layers, features, or mechanisms, including morphology, syntax, and dialect information in Arabic transformers.
- Linguistic explanation: Linguistic explanations identify operative morphological, syntactic, semantic, dialectal, or orthographic features using units such as stems, clitics, lemmas, and diacritics.
- Socio-cultural explanation: Socio-cultural explanations address cultural context, moderation norms, religious or classical references, identity terms, and user expectations.
3 A Critical Taxonomy
The taxonomy shows that Arabic’s explainability gap concerns the alignment of methods, tasks, linguistic units, goals, and evaluation, not merely the number of explanation studies. Existing approaches remain scattered and often post-hoc.
- The taxonomy distinguishes current practice from missing explanatory capacity across the principal dimensions of the Arabic XAI gap.
- Post-hoc methods such as LIME, SHAP, attention visualization, and feature importance identify surface units but do not establish linguistically meaningful use of Arabic features.
- When tasks emphasize label prediction and methods emphasize token attribution, explanations are likely to remain prediction-level despite morphological, dialectal, semantic, or sociocultural errors.
- Arabic XAI needs stronger alignment between explanation goals and evaluation practices.
- Probing, visual analytics, explainable ASR metrics, knowledge-graph diagnostics, neuro-symbolic modeling, and user-centered evaluation broaden explanatory targets beyond generic attribution.
- These approaches remain scattered, so post-hoc methods should be embedded in Arabic-aware evaluations of faithfulness, stability, usefulness, and fairness under linguistic variation.
4 Task-Method Coverage
Arabic XAI is strongest in sentiment and harmful-content classification, while generation, retrieval, multimodal, and linguistically grounded diagnostics remain less developed. Existing work broadens explanation targets beyond tokens, but these approaches are scattered and evaluation remains difficult.
- Coverage asymmetries: Sentiment analysis and harmful-content classification form the strongest clusters in Arabic XAI, whereas generation, retrieval, translation, summarization, parsing, dialogue, and structured analysis are weakly covered.The same small family of post-hoc explainers is reused across many tasks.
- Sentiment and Opinion Mining: Sentiment studies make Arabic classifier decisions inspectable but often remain prediction-level, identifying influential words rather than negation, dialectal intensification, sarcasm, stance, confounds, or morphology.LIME, SHAP, attention, and contextual representations are common interpretive devices.
- Harmful Content: Harmful-content work connects explainability with moderation and social meaning, but shared protocols for usefulness and fairness across dialects, identities, targets, and cultural references are lacking.A token may be harmful, quoted, reclaimed, dialect-specific, or target-dependent.
- Information Integrity: Fake-news and spam explanations support feature inspection and debugging, yet attribution rarely establishes whether models identified deception, source credibility, temporal inconsistency, or culturally grounded claims.Information-integrity explanations would ideally connect entities, claims, sources, and context.
- Diagnostics Beyond Classification: Arabic-specific diagnostics examine morphology, dialect, recognition errors, entity boundaries, readability, semantic relations, and grounding, but remain promising case studies rather than a shared evaluation paradigm.These approaches shift attention from surface-token salience to explicit linguistic or relational targets.
- Multimodal and LLM-Adjacent Work: Multimodal and LLM-oriented work expands Arabic XAI beyond text classification, while making explanation quality harder to separate from visual grounding, generation quality, factuality, and hallucination.Arabic LLM explainability remains an emerging requirement rather than a mature methodology.
5 Critical Discussion: The existing Gaps
The survey identifies methodological, task, linguistic, and evaluation gaps in Arabic XAI. It argues that explanations should represent Arabic linguistic and sociocultural evidence, not merely attach post-hoc importance scores to classifier inputs.
- Method Gap: Arabic XAI favors accessible post-hoc methods such as LIME, SHAP, attention visualization, and saliency, while probing, counterfactuals, rationales, and faithfulness tests remain less common.Richer methods appear in isolated forms, including probing, visual analytics, knowledge-graph diagnostics, user-centered evaluation, and neuro-symbolic modeling.
- Task Gap: Classification dominates Arabic XAI, although retrieval, generation, comprehension, captioning, LLM adaptation, and benchmarking require explanations of grounding, hallucination, alignment, fluency, factuality, entities, and discourse.The classification emphasis encourages explanations centered on label-supporting tokens.
- Linguistic Gap: The linguistic gap arises because explanations often identify words without specifying whether the model used morphology, clitics, dialect, orthography, diacritics, entity boundaries, or other Arabic evidence.Preprocessing and tokenization can change the unit being explained.
- Linguistic Gap: Probing, dialectal ASR, NER diagnostics, readability analysis, and Qur’anic MRC show that explanations can target morphology, dialect, recognition errors, entity boundaries, readability features, and semantic relations.These examples provide a direction for linguistically grounded Arabic XAI.
- Evaluation Gap: Faithfulness, plausibility, stability, usefulness, and fairness are distinct evaluation targets, yet many studies still rely on predictive accuracy and selected explanation examples.Arabic explanations should also be tested under spelling, normalization, segmentation, dialectal, and diacritic variation.
6 Toward Linguistically Grounded Arabic XAI
A linguistically grounded Arabic XAI agenda should redesign explanation units, evaluation, human benchmarks, generation-focused protocols, and reproducibility around Arabic’s linguistic and sociotechnical properties.
- Arabic-aware units of explanation: Arabic XAI should evaluate explanations at meaningful linguistic levels, including morphemes, clitics, stems, lemmas, diacritics, named entities, expressions, dialectal phrases, and discourse cues.Token-level heatmaps can miss decisions driven by morphology or orthography, so papers should relate internal tokenization to user-facing units.
- Arabic-preserving perturbation tests: Arabic-preserving perturbation tests should vary spelling, normalization, clitic segmentation, dialectal paraphrases, and diacritics while preserving meaning.These tests distinguish robust explanations from preprocessing or tokenization artifacts.
- Cross-variety explanation evaluation: Explanation benchmarks should compare MSA, regional dialects, Classical Arabic, social-media Arabic, Arabizi, and code-switched Arabic.Variety should become an explanation variable rather than merely a dataset description.
- Human rationale benchmarks with Arabic expertise: Arabic XAI needs shared human-rationale benchmarks with expert annotations and task-specific quality labels from dialectally competent Arabic speakers.Domain expertise may also be needed for moderation, health, education, religious text, or accessibility applications.
- Explanation evaluation for Arabic LLMs, RAG, and generation: Future Arabic XAI should target LLMs, retrieval-augmented generation, QA, summarization, translation, dialogue, and hallucination detection.Explanations for these systems should address factual grounding and retrieval evidence.
- Faithfulness, plausibility, stability, usefulness, and fairness protocols: Arabic XAI evaluations should separately report faithfulness, plausibility, stability, usefulness, and fairness across varieties and identity terms.This separation is especially important in moderation, health, education, religious-domain, and information-integrity applications.
- Reproducibility standards: Arabic XAI papers should report data variety, preprocessing, tokenization, model checkpoints, explanation parameters, metrics, validity limits, and released artifacts.These details help distinguish Arabic-specific improvements from methods that are simply easier to visualize.
7 Conclusion
The survey finds a clear explainability gap in Arabic NLP: current work mainly explains classification outputs, while generation, retrieval, and other important tasks remain under-explained. It argues for explanations faithful to Arabic’s linguistic, cultural, and sociotechnical properties, supported by Arabic-aware evaluation and reproducibility practices.
- Conclusion: Arabic XAI remains dominated by output-level explanations for classification, especially sentiment, harmful-content detection, fake news, and spam.Generation, retrieval, QA, summarization, translation, structured prediction, and LLM applications remain under-explained.
- Conclusion: Arabic explanations should account for morphology, clitics, dialectal variation, diglossia, diacritics, orthographic ambiguity, named entities, cultural references, and Classical or religious registers.The survey frames these properties as part of Arabic’s linguistic, cultural, and sociotechnical character.
- Conclusion: Future Arabic XAI should move from explaining predictions to explaining Arabic itself through Arabic-aware units, cross-variety evaluation, perturbation tests, human rationales, LLM/RAG protocols, and reproducibility standards.These directions are presented as the survey’s agenda for future work.
8 Survey Methodology and Scope
The paper uses a critical structured survey of explicitly explainability-oriented Arabic or Arabic-adjacent work, drawing on searches across major academic databases and indexes. It excludes Arabic NLP studies without an explicit explanatory component and treats attention cautiously.
- Search strategy: The survey searches major academic databases and indexes using combinations of Arabic NLP, XAI, interpretability, transparency, method, task, and application keywords.Sources include ACL Anthology, ACM Digital Library, IEEE Xplore, ScienceDirect, SpringerLink, MDPI, arXiv, preprint venues, and Google Scholar.
- Inclusion and exclusion criteria: Studies qualify when their task, dataset, model, or output is Arabic or Arabic-language adjacent and they explicitly address explanation, interpretability, transparency, diagnostic analysis, or human-facing explanation.Papers reporting Arabic NLP performance without an explicit explanatory component are excluded.
- Inclusion and exclusion criteria: The survey treats attention-based models cautiously because attention is not necessarily an explanation.This methodological caution is grounded in general NLP XAI literature cited by the paper.
Limitations
The survey’s evidence base is bounded by its critical, non-exhaustive review design and by uneven coverage across tasks and publication stages. Its taxonomy and agenda therefore represent a structured snapshot of a developing field.
- Review scope: The survey is a critical synthesis rather than an exhaustive systematic review of all Arabic NLP work containing interpretable components.Relevant model analysis, evaluation, or error-diagnosis studies may be omitted when they do not use XAI terminology.
- Evidence coverage: The reviewed literature is uneven across tasks and publication stages, with stronger representation for sentiment and harmful-content detection than for emerging Arabic LLM, RAG, hallucination, and multimodal XAI.The resulting taxonomy and agenda should be read as a snapshot of a developing field rather than a fixed classification.
Use of AI Assistance
The authors used Grammarly and ChatGPT for language editing, while retaining responsibility for the scientific content, analyses, and conclusions.
- AI-assisted tools, including Grammarly and ChatGPT, were used for spell checking and stylistic revisions.
- The authors retain full responsibility for the work’s scientific content, analyses, and conclusions.