Source-linked AI summary

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg

arXiv:2608.29990v1cs.CLcs.AIcs.CV

TL;DR

MSA-centric Arabic benchmarks leave dialectal and culturally grounded competence insufficiently measured. The paper introduces an expert-authored Saudi benchmark with a two-phase rubric methodology and finds broadly low, tightly clustered performance across four models, alongside distinctive error patterns. It releases the prompts, ground truths, and scored rubrics for reproducible evaluation.

  • Problem

    MSA-dominated Arabic benchmarks provide limited evidence about models’ competence in everyday dialectal and culturally grounded Saudi Arabic.

  • Method

    The paper evaluates 31 expert-authored Saudi prompts using shared ground-truth-derived positive criteria and model-specific penalties across four systems.

  • Results

    42.7%–53.1% macro-average scores across the four systems show a narrow performance band, with no model exceeding 55%.

  • Takeaways & Limitations

    Dialect-aware evaluation requires reference-guided human rubrics that capture register, nuance, and culturally grounded meaning rather than relying on MSA-centric or purely automatic metrics.

  • Takeaways & Limitations

    The study uses a focused 31-prompt Saudi Arabic set, so aggregate percentages are indicative signatures rather than precise population estimates.

Abstract

from arXiv · show

Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is consequential: everyday Arabic is largely dialectal, and dialect encodes social meaning that MSA-centric evaluation cannot capture. We present a rubric-based benchmark for the Saudi dialect, comprising 31 expert-authored prompts spanning idiomatic, pragmatic, lexical, and culturally-embedded phenomena, each paired with an expert-established ground truth. Our methodology separates evaluation into a model-agnostic phase, in which atomic, MECE positive criteria are derived solely from the ground truth, and a model-specific phase, in which four state-of-the-art systems -- Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3 -- are scored against those criteria and penalised for errors they actively introduce. Across 124 model-prompt evaluations we catalogue 466 error instances under a nine-category taxonomy. The four systems cluster within a narrow macro-average band (42.7%-53.1%), with no model exceeding 55% and every model recording at least one negative-scoring prompt, confirming that Saudi dialectal competence remains broadly unsolved. Notably, Ambiguous Framing is the dominant failure mode (37.3% of errors) while outright Hallucination accounts for only 11.2%, indicating that models fail less by stating falsehoods than by distorting register and flattening pragmatic nuance. We further observe a consistency-versus-ceiling trade-off and model-distinctive error signatures. We release the full prompt set, ground truths, and scored rubrics to support reproducible dialectal evaluation.

A Reference-Guided, Expert-Authored Evaluation of Four State-of-the-Art

The paper benchmarks Saudi dialect and cultural competence, addressing the dominance of MSA-focused evaluation with expert-authored, ground-truth-linked prompts and reproducible rubrics.

  • Benchmark and evaluation scope: The benchmark addresses a gap left by Arabic evaluations dominated by Modern Standard Arabic rather than everyday dialectal communication.Dialect carries culturally load-bearing meanings that MSA evaluation may miss.
  • Benchmark and evaluation scope: 31 expert-authored prompts target Saudi dialect and cultural knowledge requiring local, lived experience.Each prompt is paired with SME-validated ground truth.
  • Rubric methodology: The evaluation uses a model-agnostic positive-rubric phase followed by model-specific negative criteria for errors actively introduced.Positive criteria are derived from ground truth and reused across systems.
  • Comparative evaluation: Four state-of-the-art models are compared using a nine-category taxonomy covering 466 discrete error instances.The evaluated systems are Claude Opus 5, Gemini 3.7, GPT-5.6, and Kimi K3.

2 Background and Motivation

The background motivates dialect-aware evaluation because MSA-centric benchmarks overlook socially situated Arabic, while the benchmark is designed around difficult, locally grounded Saudi prompts.

  • Related evaluation work: Capability-specific benchmarks are valuable because broad aggregate testing can miss failure conditions in particular deployment domains.Related work applies this approach to long-context modelling, document reading, clinical text, and other settings.
  • Background and motivation: Arabic evaluation has gravitated toward MSA because annotated resources are more abundant, creating a blind spot for spoken regional varieties.Formal-language performance cannot be assumed to transfer to socially situated language use.
  • Background and motivation: Dialectal competence includes culturally transmitted knowledge such as proverbs, etiquette, conflict-resolution norms, and material-culture references.These phenomena are difficult to acquire from web text at scale.
  • Evaluation design rationale: Rubric decomposition replaces single-reference matching for open-ended pragmatic explanations with independently marked, interpretable criteria.Atomicity and MECE coverage make the decomposition rigorous and repeatable.
  • Benchmark construction: The 31 prompts are native-authored, dialect-written, conversational, and designed to require local knowledge while supporting layered answers.Prompts are classified by dominant linguistic or cultural scope and screened for discriminativeness.

4 Evaluation Methodology

The methodology fixes a shared evaluation standard before model outputs are observed, then scores each response against atomic positive criteria and model-specific penalties.

  • 4.1 Two-phase workflow: Phase A derives a reusable prompt-plus-positive-rubric pair from SME ground truth without consulting model outputs.Criteria are atomic, dimension-tagged, weighted, and finalized before evaluation.
  • 4.1 Two-phase workflow: Phase B scores each model response against shared positive criteria and adds penalties for errors that the model actively introduces.Negative criteria are model-specific because systems make different mistakes.
  • 4.2 Criterion design: Atomicity assigns one testable property to each criterion, while MECE coverage prevents overlap and covers the complete ground truth.Together they support faithful, non-double-counted scores.
  • 4.3 Criterion fields: Positive and negative criteria record dimensions, objectivity, explicitness, weights, score types, and maximum scores.The fields distinguish correctness requirements from judgment-based qualities such as register naturalness.
  • 4.4 Scoring model: A prompt’s maximum score depends only on positive criteria, while its model score combines earned positive points with active-error penalties.Penalties are not bounded by positive points, so misleading responses can receive negative scores.
  • 4.5 Evaluation controls: Outputs are collected under uniform conditions with memory and conversation history disabled, one prompt per session, and no re-prompting.The first response is preserved verbatim with its original formatting.

5 Failure Taxonomy

The benchmark uses a nine-category taxonomy to classify failed criteria and introduced errors, with Ambiguous Framing requiring particular interpretive scrutiny.

  • Taxonomy construction: The final taxonomy contains nine categories, expanded from an initial six as recurring annotation behaviours emerged.Formatting and Language Mixing were added during annotation, while Instruction Violation became a first-class category.
  • Interpretive considerations: Ambiguous Framing is intentionally broad and captures register and nuance distortions central to dialectal competence.Its breadth makes it informative but especially dependent on evaluator judgment and inter-rater scrutiny.
  • Interpretive considerations: Language Mixing is reserved for inappropriate switches of language or variety, while Ambiguous Framing covers distortions within the correct variety.The boundary is an explicit modelling choice that future work may refine.

6 Quantitative Performance

Across models, Saudi-dialect performance remains low and broadly difficult, while aggregate rankings are robust but conceal substantial differences in reliability, prompt difficulty, and error severity.

  • Leaderboard: 42.7%–53.1% macro-average scores place all four systems in a narrow band, with no model exceeding 55% and every model producing at least one negative-scoring prompt.The leaderboard uses macro-average over per-prompt percentages; negative scores occur when active errors outweigh earned points.
  • Consistency versus peak performance: 61.5% is GPT-5.6’s highest median and 88.9% its highest single-prompt score, but its SD is 33.7 and its worst prompt reaches −44.8%.GPT-5.6 therefore has the highest ceiling alongside the widest dispersion and most severe single failure.
  • Consistency versus peak performance: Gemini 3.7 is the most consistent system at SD 16.6 with only one negative prompt, whereas Claude Opus 5 has four negative prompts and SD 29.2.These distributions establish a consistency–ceiling trade-off relevant to deployments that cannot tolerate occasional severe failures.
  • Macro versus micro aggregation: Ranking is invariant between macro- and micro-averages, with gaps of at most ∼1.6 points, indicating that standings are not artifacts of a few high-weight prompts.GPT-5.6 has the largest divergence because it performed well on several high-max prompts, including SAU-25.
  • Item difficulty: The hardest prompts concentrate on highly localized lexis and metaphor, while common greetings and high-frequency expressions are easiest across models.Figure 4 represents universally difficult prompts as vertical cool bands and broadly stronger models as warmer rows.
  • Score–error relationship: Error frequency and severity are distinct: high-penalty errors can produce negative scores, so total error counts do not necessarily determine leaderboard rank.Claude Opus 5 has the most error instances and lowest macro-average, while GPT-5.6 has fewer but more consequential errors on high-max prompts.

7 Error Analysis

Across 124 evaluations, the models produced 466 errors dominated by Ambiguous Framing, with distinct normalized error profiles despite broadly similar overall loads.

  • Aggregate error load: 466 discrete errors were recorded across 124 model-prompt evaluations.The benchmark covers 31 prompts evaluated on four models.
  • Aggregate error load: GPT-5.6 recorded the fewest errors at 103, while Claude Opus 5 recorded the most at 132.Gemini 3.7 and Kimi K3 recorded 117 and 114 errors, respectively.
  • Dominant failure mode: 37.3% of all errors were Ambiguous Framing, followed by Omission at 18.0% and Irrelevant Addition at 13.5%.Hallucination ranked fourth at 11.2%, indicating that nuance and completeness failures were more frequent than factual fabrication.
  • Cross-model comparison: Ambiguous Framing led every model, while Gemini concentrated 46.2% of its errors in that category.GPT-5.6 had the fewest hallucinations but the highest formatting load; Opus 5 had the most irrelevant additions.
  • Normalized error signatures: Normalized profiles isolate behavioral tendencies from total error load and reveal distinct signatures across the models.Gemini is framing-heavy, GPT-5.6 emphasizes omission and formatting, Opus 5 emphasizes irrelevant addition, and Kimi combines framing with elaboration.

8 Discussion

The discussion argues that fluent outputs can still mishandle Saudi register and pragmatic force, making error profiles and rubric design important for deployment decisions.

  • Fluency masks incompetence: Ambiguous Framing exceeds Hallucination because models often distort register and pragmatic force without fabricating facts.A grammatically correct MSA gloss for a Saudi idiom can therefore still be an evaluation failure.
  • Fluency masks incompetence: Reference-guided human rubrics expose dialectal errors that MSA-centric and automatic fluency metrics cannot detect.The methodology targets culturally load-bearing omissions, register mismatches, and other subtle failures.
  • Deployment relevance: Model error signatures should be matched to product use cases rather than selected solely by leaderboard position.Framing-heavy, over-elaborating, and omission-heavy profiles create different risks in high-context, terse-response, and factual-reliability settings.
  • Limitations: The largest error category is also the most judgement-dependent, making inter-rater reliability measurement a priority.Atomic criteria and item-level justifications constrain scorer discretion but do not remove the need for formal reliability analysis.
  • Threats to validity: The study’s main external-validity threat is specificity to Saudi Arabic, although the authors report that the methodology is variety-agnostic and transfers to other Arabic dialects.The prompt set is also focused rather than statistically broad, so aggregate percentages represent indicative signatures rather than population estimates.

9 Conclusion

The paper concludes that Saudi dialectal and cultural competence requires rubric-based, reference-guided evaluation, supported by released materials for reproducible extension.

  • Conclusion: The benchmark fixes shared ground-truth-derived positive criteria before model outputs are observed and adds model-specific penalties for introduced errors.This two-phase design supports both fair cross-model comparison and diagnostic failure analysis.
  • Conclusion: Across four systems, Ambiguous Framing was the dominant failure mode, followed by omission and unsolicited over-elaboration.The conclusion characterizes these as distortions of register, nuance, completeness, and response scope rather than primarily fabrication.
  • Future work: The authors recommend inter-rater reliability statistics, a larger prompt set, broader dialect coverage, and targeted mitigation studies.These directions address category subjectivity, limited scale, and the benchmark’s current Saudi-focused scope.

10 Reproducibility and Ethics Statement

The benchmark uses native Saudi expert authorship, controlled single-turn collection, and released scoring materials to support respectful, reproducible evaluation.

  • Reproducibility: Native Saudi subject-matter experts authored the prompts and validated their difficulty before rubric construction.This process grounds the benchmark in expert linguistic and cultural judgment.
  • Reproducibility: Model outputs were collected with memory disabled in single-turn sessions and captured verbatim.The protocol supports independent replication and rescoring.
  • Ethics and release: The released materials include prompts, ground truths, full scored rubrics, model-specific negative criteria, and justifications.Sensitive cultural prompts were framed respectfully.

A Saudi Dialect Benchmark Prompts

The benchmark appendix presents the complete Saudi dialect prompt set, comprising 31 expert-authored prompts reproduced verbatim. Table A1 organizes the full prompt collection across the appendix pages.

  • 31 expert-authored prompts make up the complete Saudi dialect benchmark set.The appendix states that the prompts are reproduced verbatim.
  • Table A1 is the complete Saudi dialect benchmark prompt set.
  • The prompt list continues across multiple appendix pages.The appendix explicitly marks continuation from the previous and next pages.
Loading 2608.29990v1…