Source-linked AI summary
Beyond Cultural Knowledge: Evaluating Arabic Cultural Appropriateness of Large Language Models
Enes Altinisik, Hamdy Mubarak, Masoomali Fatehkia, Husrev_Taha_Sencar Husrev Taha Sencar
TL;DR
Existing cultural evaluations often test what language models know rather than how they behave in culturally sensitive, open-ended interactions. AraBehave evaluates Arabic cultural appropriateness through native-speaker judgments and finds that leading models can achieve similar scores while failing on different dimensions: normative stance versus grounded cultural accuracy.
Problem
Cultural evaluations largely test knowledge rather than culturally appropriate behavior in open-ended recommendations, opinions, and guidance for Arabic-speaking users.
Method
AraBehave uses 1,623 culturally grounded Arabic prompts, 29,214 native-speaker judgments from five Arab regions, and a scoring model evaluated across six language models.
Results
The leading general-purpose and Arabic-centric models score similarly—Gemini 3.1 3.84 versus Fanar 2.0 3.83—but fail on opposite axes, while cultural instruction raises Gemini 3.1 to 4.57.
Takeaways & Limitations
Cultural appropriateness is a multidimensional capability: stance is prompt-sensitive, whereas grounded accuracy tracks scale and Arabic alignment data and disappears under culture-neutral tuning.
Takeaways & Limitations
AraBehave targets broadly shared Arab norms rather than intra-cultural variation, uses subjective judgments from three annotators per pair, and its scoring model is trained on six models.
Abstract
from arXiv · showhide
Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test what a model knows rather than how it behaves when giving open-ended recommendations, opinions, and guidance. We introduce AraBehave: 1,623 culturally grounded, open-ended Arabic prompts with 29,214 cultural-appropriateness judgments from native speakers across several Arab regions, plus a scoring model whose predictions correlate strongly with human judgments on unseen systems (Pearson r=0.74). Evaluating three Arabic-centric and three frontier LLMs, we find that cultural appropriateness is not a single capability but decomposes into two largely independent components: normative stance and grounded cultural accuracy. The best general-purpose and best Arabic-centric models score identically (3.84 vs. 3.83 of 5) yet almost never fail for the same reason: general-purpose models exhibit strong factual grounding but a culturally inappropriate normative stance, being penalized for secular framing and false balance on culturally settled matters (28--33% of their low-score rationales), while the best Arabic-centric model adopts the expected stance but is penalized for fabricated hadith and misquoted verses (29%). Stance is cheap and fragile: one sentence of cultural instruction lifts Gemini to 4.57, above every Arabic-specialized model. Conversely, a generic ``answer clearly and objectively'' prompt costs Allam-7B 0.68 points, while asking the same questions in English lowers scores for every model but one. Grounding instead tracks scale and Arabic alignment data, and disappears when culturally aware instruction tuning is replaced by a culture-neutral corpus. General safety benchmarks see none of this: they saturate above 89 while cultural scores span 2.71-3.84. We will release the benchmark, annotations, and the scoring model.
Introduction
Existing cultural evaluations largely test factual knowledge rather than open-ended behavior, leaving culturally appropriate Arabic interaction undermeasured. AraBehave addresses this gap and shows that models can achieve similar scores while failing through different cultural mechanisms.
- Introduction: AraBehave targets culturally grounded Arabic recommendations, opinions, and guidance rather than factual cultural knowledge.Its benchmark contains 1,623 curated prompts rated by native Arabic speakers with written justifications.
- Introduction: Cultural appropriateness decomposes into normative stance and grounded cultural accuracy, so a single score can conceal the needed model fix.General-purpose models are typically penalized for secular or both-sides framing, whereas Arabic-centric models are penalized for scriptural and factual errors.
- Introduction: Gemini 3.1 and Fanar 2.0 are statistically tied at 3.84 and 3.83, yet their low-score rationales are nearly disjoint.This contrast demonstrates that equal aggregate performance can reflect opposite failure profiles.
- Introduction: A single cultural instruction lifts Gemini 3.1 to 4.57, while generic objectivity prompting lowers Allam-7B by 0.68 points.The findings characterize stance as highly sensitive to configuration and prompting.
- Introduction: General safety benchmarks saturate at 89.4–98.7 while cultural scores span 2.71–3.84, indicating that safety scores do not capture these cultural differences.The paper also reports a negative correlation between the two evaluation types.
Arabic Cultural Benchmarks
Arabic benchmarks cover safety, cultural knowledge, commonsense, entities, values, and linguistic competence, but most reduce evaluation to objective targets. AraBehave extends this landscape by measuring graded human judgments of open-ended culturally appropriate behavior.
- Arabic Cultural Benchmarks: Existing Arabic benchmarks evaluate complementary dimensions including culturally grounded safety, cultural knowledge, commonsense, entities, values, and dialectal competence.Examples include Arabic Safeguard Benchmark, AraSafe, AraTrust, ArabCulture, PALM, CAMeL, AraDiCE, JAWAHER, and DialectalArabicMMLU.
- Arabic Cultural Benchmarks: Most prior evaluations use a correct answer, safe response, or automated-evaluator match as the objective target.These formulations fit factual knowledge and policy compliance better than interactions where multiple responses may be appropriate to varying degrees.
- Arabic Cultural Benchmarks: AraBehave advances open-ended evaluation by using multiple native Arabic speakers, Likert scores, and mandatory rationales instead of relying on an automated judge.The rationales support separating stance from grounding failures and reveal measurable, irreducible disagreement.
AraBehave
AraBehave is a 1,623-item benchmark of culturally grounded Arabic behavioral scenarios, assembled from user logs, tester-seeded samples, and expert curation. Native Arabic speakers rate six models across eight thematic categories using 5-point judgments and written justifications.
- AraBehave: AraBehave combines user log analysis, tester-seeded samples, and expert curation to construct culturally sensitive Arabic behavioral prompts.The final dataset retains prompts matching the paper’s definition of value-laden scenarios rather than general safety violations.
- AraBehave: The final benchmark contains 1,623 prompts organized into eight categories spanning religion, gender, family, honor, prohibited conduct, health, politics, and identity.Categories were defined inductively from the three sources and validated against prior Arabic cultural NLP work.
- AraBehave: Table 1 reports mean human cultural-appropriateness scores by model and thematic category, with higher values indicating greater cultural alignment.The table covers the full benchmark and includes the number of prompts per category.
- AraBehave: Six models—three Arabic-centric and three general-purpose—produce 9,738 prompt-response pairs for evaluation.The model set includes Fanar 2.0, Allam-7B, Jais-70B, GPT-5, Gemini 3.1, and Qwen3-30B.
- AraBehave: Each response receives a 1–5 cultural-appropriateness rating from three native Arabic speakers, with a mandatory written justification.The procedure yields 29,214 annotations across annotators from five Arab regions, and mean ratings become pair labels.
AraBehave Scoring Model
The AraBehave scoring model is a regression-based evaluator trained on human annotations to predict 1–5 cultural-sensitivity scores. Leave-one-model-out evaluation tests whether it generalizes to systems whose response styles were absent from training.
- AraBehave Scoring Model: The scoring model continues training FanarGuard with a single cultural-sensitivity head that predicts a 1–5 score from AraBehave annotations.It uses mean human ratings as targets and trains with mean squared error loss.
- AraBehave Scoring Model: Leave-one-model-out training holds out all 1,623 responses from one model and averages performance across six held-out models.This split measures generalization to unseen model styles, while a separate random 75/10/15 split provides an additional evaluation setup.
Annotation Results
Across 1,623 prompts, Gemini 3.1 and Fanar 2.0 achieved statistically comparable overall scores, but their low-score rationales revealed different weaknesses in normative stance versus grounded accuracy. Human judgments showed moderate agreement, while irreducible disagreement concentrated on value-laden topics.
- Overall Results: Gemini 3.1 scored 3.84 and Fanar 2.0 scored 3.83, statistically comparable overall and ahead of the other evaluated models.GPT-5 scored 3.55, Allam-7B 3.47, Jais-70B 3.06, and Qwen3-30B 2.71.
- Failure Analysis: 28% and 33% of Gemini and GPT-5 low-score rationales cited secular or Western framing, whereas Fanar’s rationales cited doctrinal or factual errors in 29%.Gemini and GPT-5 rarely lost points for factual inaccuracy (6%), while Fanar’s framing was cited in only 6%; Qwen combined accuracy problems with weak stance.
- Failure Analysis: Cultural appropriateness decomposed into normative stance and grounded accuracy, so aggregate scores concealed which correction each model required.The failure analysis was coded from annotators’ written justifications rather than model responses or an automated judge.
- Per-Category Results: On Vice & Prohibited, Fanar scored 3.77 versus 2.49 for GPT-5 and 1.96 for Qwen, while Fanar led value-laden categories and Gemini led knowledge-oriented ones.Vice & Prohibited and Politics & Sect were hardest for all models; response length was uncorrelated with score (r = −0.01).
- Inter-Annotator Agreement: Annotators agreed within one scale point on 74.1% of pairs, with ICC 0.58, weighted κ 0.54, and mean absolute pairwise difference 1.00.Each annotator correlated with the leave-one-out mean at r = 0.66 and ρ = 0.65.
- Inter-Annotator Agreement: Every one of the 50 highest-disagreement pairs contained both a rating of 1 and 5, with coherent rationales often reflecting incompatible cultural criteria.Disagreement concentrated in Politics (α = 0.470) and Gender (α = 0.458), motivating graded judgments with rationales instead of a single gold label.
Analyses
The analyses show that cultural alignment separates into prompt-sensitive stance and scale- and training-sensitive grounding, while the scoring model generalizes to unseen systems. Prompt language and generic instructions can materially alter scores, whereas general safety scores do not track cultural appropriateness.
- Scoring Model Performance: Pearson 0.73: a Gemma-3-4B scoring-model head predicts human scores well, preserves model rankings exactly, and supports scoring-model-based analyses beyond the annotated systems.Leave-one-model-out performance also indicates generalization to unseen models; the main residual gap is driven by low-rated, out-of-distribution Qwen3-30B.
- Stance Is Cheap to Add and Easy to Lose: 3.92 to 4.57: a Cultural prompt raises Gemini 3.1 above every Arabic-specialized model, while Allam-7B falls 3.47 to 2.79 under a Generic prompt.Prompting changes the leaderboard because stance is deployment-configuration dependent; it does not help GPT-5 on Vice & Prohibited, where willingness to state a ruling remains limited.
- Cultural Alignment Is Language-Conditioned: Every model except Qwen3-30B loses cultural appropriateness under English or Chinese prompts; Allam-7B drops 3.47 to 2.97 in English and 2.50 in Chinese.Gemini 3.1 also falls from 3.92 to 3.26 and 3.07, showing that the effect is not confined to Arabic-specialized systems.
- Grounding Scales with Model Size: 1.98 to 3.77: Qwen3.5 cultural appropriateness rises monotonically with scale, while Gemma4 increases from 3.13 to 3.86; the largest models match Fanar 2.0’s 3.86.The result holds within both families under a fixed recipe across roughly two orders of magnitude in parameters; the 35B mixture-of-experts model matches the 27B dense model, consistent with its smaller active budget.
- Cultural Alignment Must Be Trained, Not Pretrained: 0.66 vs. 0.28 and 0.16: replacing culturally aware alignment with UltraChat lowers appropriateness most for Fanar-1, and Fanar-1 then falls below Gemma2-9B’s instruction-tuned variant.This supports explicit culturally informed alignment data, rather than Arabic pretraining alone, as the main source of the specialized advantage.
- Safety Evaluation: 89.4 to 98.7 vs. 2.71 to 3.84: general-safety scores are compressed near the top while cultural appropriateness remains widely separated, with no significant positive correlation.Allam-7B and Qwen3-30B have the two highest safety scores but opposite cultural scores, whereas Fanar 2.0 has the second-highest cultural score at the lowest safety score.
Conclusion
AraBehave shows that cultural appropriateness is not a single competence: models can score similarly while failing on different axes. Stance is configuration-dependent, grounding depends on scale and Arabic alignment data, and safety evaluation does not reveal these cultural differences.
- Gemini 3.1 and Fanar 2.0 score similarly, yet fail on opposite axes: normative stance versus grounded cultural accuracy.The paper recommends reporting these components separately because a single score conceals which fix a model needs.
- A single cultural instruction raises stance-sensitive scores, whereas generic instructions or changing prompt language can reduce them.Grounding instead follows scale and explicit Arabic alignment data and disappears under culture-neutral tuning.
- Cultural alignment is configuration-dependent and is not captured by general safety evaluation.The deployed system prompt should therefore be treated as part of a model’s cultural alignment.
Limitations
AraBehave’s limitations concern cultural scope, subjective disagreement, evaluator generalization, cross-lingual scoring, and possible Western-centric phrasing in generated data.
- The benchmark targets broadly shared Arab norms rather than variation across Gulf, Levantine, and North African contexts, and reflects values at collection time.
- Three annotators per pair can produce mean labels that mask minority disagreement on contested items.
- The scoring model is trained on six models and may not calibrate to future systems.
- Cross-lingual scoring uses back-translation that is better validated for Arabic-English than for Chinese, while LLM-generated prompts may contain Western-centric phrasings.
D Scoring Model Training Details
The scoring model adapts FanarGuard into a single-output cultural-sensitivity regressor trained on AraBehave annotations. Continued FanarGuard training outperforms alternative initialization and transfers across held-out models.
- Scoring Model Architecture and Training: Continuing FanarGuard training on AraBehave annotations yields the adopted single-output model for predicting 1–5 cultural-sensitivity scores.The harmlessness head is removed, and training uses mean human ratings with Mean Squared Error loss.
- Scoring Model Design Choices: MAE 0.537 and Pearson 0.812 for continued FanarGuard training outperform Gemma-3-4B initialization, with MAE 0.654 and Pearson 0.728.
- Scoring Model Design Choices: MAE 0.537 versus 0.840 and Pearson 0.812 versus 0.508 show that continued annotation training sharply improves calibration over FanarGuard off-the-shelf.
- Generalization and Bias Checks: Leave-one-model-out evaluation holds out all 1,623 responses from one model and trains on the other five, with consistent performance across held-out models.
- Generalization and Bias Checks: The scoring model slightly overestimates mean human scores but preserves model rankings and overall trends.
F Why Models Succeed or Fail: A Comment-Grounded Analysis
Annotator rationales show that model quality depends on stance, grounding, depth, and context-sensitive refusal rather than answer length alone. Frontier and Arabic-centric models exhibit distinct strengths and failure patterns.
- The correlation between log answer length and score is negligible at r = −0.01; Fanar 2.0 scores 3.83 with much shorter answers than Qwen3-30B, which scores 2.71.Annotators reward substantive grounding rather than verbosity.
- Low-score rationales reveal opposite failure modes: frontier models are penalized for values and Arabic-centric or smaller models for accuracy.
- Gemini and GPT-5 earn high scores through comprehensive, integrated responses, while Fanar and Jais face a brevity ceiling despite cultural reliability.About 53% of their high-score rationales praise depth and coverage, whereas roughly 24% of Fanar and Jais rationales cite brevity or missing evidence.
- High scores require a clear Islamic stance, authentic Qur’an/Sunnah evidence, practical halal alternatives, and a firm but non-hostile tone.Gemini earns 5s when it adopts this frame with depth, indicating framing rather than model provenance drives the score.
- Refusal has no systematic score effect: refusal-flagged rationales average 3.44 versus 3.41 otherwise, because context determines whether refusal is appropriate.
- Gemini is comprehensive but polarizing, Fanar 2.0 is consistently reliable but brief, GPT-5 has the most value-based failures, and Qwen3-30B is dominated by hallucination and fabricated citations.
G Data Source Analysis and GPT-5 Contamination
AraBehave prompts were sourced through three stages, with GPT-5 generating only tester-seeded samples and otherwise serving as an eliminator. No contamination pattern appears in GPT-5’s scores, and annotation timing shows no quality concern.
- Data sources: GPT-5 only filtered prompts in Stages 1 and 3, while it transformed human-written cultural contrasts into adversarial scenarios in Stage 2.
- Contamination check: 3.36: GPT-5 scored lowest on the Stage 2 prompts it helped generate, providing no evidence of systematic contamination.The study compares scores by prompt source because GPT-5 generated Stage 2 prompts; its Stage 2 score was below both other source groups.
- Annotation quality: Annotation times were not implausibly short, and assigned ratings showed no correlation with time spent on responses.
I Analysis of Annotator Disagreement
Annotator disagreement on difficult AraBehave items is substantial, coherent, and often reflects incompatible standards rather than carelessness. Mean aggregation is reasonable for most items but can conceal bimodal disagreement on contested, value-laden prompts.
- Disagreement overview: Every one of the 50 highest-disagreement pairs contains ratings of both 1 and 5, while almost all annotators provide coherent rationales.Unsupported ratings are treated as annotator error; the remaining disagreements reflect different criteria for cultural appropriateness.
- Criterion-level disagreement: Annotators diverge over whether cultural appropriateness requires explicit Arab–Islamic stances or neutral presentation, including disagreements about refusal, religious scrutiny, and practical framing.The cited examples show refusals, jurisprudential errors, and alcohol-related advice receiving opposite judgments under different standards.
- Criterion-level disagreement: Contested doctrine and fiction expose irreducible pluralism: knowledgeable annotators may disagree about theology or whether creative content endorses the values it depicts.Examples include divine creation of evil and a fictional inheritance split, where rationales remain internally coherent but apply different criteria.
- Rater effects: Rater means range from 4.0–4.3 for lenient annotators to 1.0–2.1 for strict annotators, whose longer evidence-citing rationales suggest a calibration opportunity.Per-annotator calibration could reduce the stable leniency/strictness effect, but not legitimate criterion-level disagreement.
- Rater effects: Mean-of-three aggregation is reasonable for most items but masks bimodal splits on the contested prompts AraBehave is designed to surface.
J Per-Category Results
Per-category analyses show that cultural appropriateness is especially sensitive to prompting language and cultural stance, while culturally aligned instruction tuning and scale support performance. Replacing culturally aware tuning with a neutral corpus lowers scores across categories, especially on norm-dependent content.
- Stance Is Cheap to Add and Easy to Lose: Cultural prompting benefits general-purpose models more than Arabic-specialized models across categories, except Vice & Prohibited, where gains are minimal or negative.Average gains are +0.92 vs. +0.45 on Honor & Reputation and +0.83 vs. +0.47 on Health & Body; GPT-5 falls from 3.60 to 3.31 on Vice & Prohibited.
- Cultural Alignment Is Language-Conditioned: Arabic-to-English/Chinese score drops are steepest on Vice & Prohibited for every model, showing that norm-dependent alignment is tightly coupled to the Arabic language cue.Fanar 2.0 falls from 3.85 in Arabic to 2.68 in English and 2.53 in Chinese.
- Grounding Scales with Model Size: The evaluated Qwen3.5 and Gemma4 families span roughly two orders of magnitude in parameter count while covering dense and sparse architectures.
- Cultural Alignment Must Be Trained, Not Pretrained: UltraChat tuning lowers cultural appropriateness in every category for all three base models, with the largest decline on Vice & Prohibited.For example, Fanar-1 falls from 3.01 to 2.03 on Vice & Prohibited; the Arabic-specialized model loses the most, tying its advantage to explicit alignment.
- Translation Quality for Cross-Lingual Scoring: Machine translation largely preserves the cultural-safety signal for Arabic–English scoring, but Arabic–Chinese results warrant caution because that translation direction is less studied.Translation may slightly affect surface quality, including poetic quality, without materially affecting the targeted cultural-safety content.