Source-linked AI summary
YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models
Mahmoud Reda, Salam Khalifa, Reham Marzouk, Nizar Habash
TL;DR
Arabic LLM evaluations have provided limited evidence about systematic morphological control from explicit lexical and feature specifications. YallaMorph addresses this gap with a large-scale benchmark spanning Arabic word classes, cliticization, invalid configurations, and controlled sampling, and evaluates multilingual and Arabic-oriented LLMs. The results show that generation remains difficult, especially for cliticized, unseen, zero-frequency, and rare forms, with performance suggesting reliance on exposure frequency and memorization.
Problem
Existing Arabic LLM evaluations mainly target downstream tasks or broad generation quality rather than controlled generation from explicit lexical and morphological specifications.
Method
YallaMorph benchmarks controlled Arabic morphological generation across verbs, nouns, adjectives, cliticized forms, and invalid configurations using over 600K sampled forms balanced across linguistic and frequency dimensions.
Results
Controlled Arabic morphological generation remains challenging, especially in diacritized and cliticized settings, with substantial performance drops for zero-frequency and rare morphospecification forms.
Takeaways & Limitations
The results suggest that current LLMs rely strongly on distributional exposure and memorization rather than robust morphological generalization.
Takeaways & Limitations
The benchmark focuses on Modern Standard Arabic in CamelMorph MSA, uses guided rather than exhaustive sampling, and is constrained by the underlying resource.
Abstract
from arXiv · showhide
Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly test controlled morphological generation from explicit lexical and feature-based input. We introduce YallaMorph, a large-scale benchmark for Arabic morphological generation covering verbs, nouns, adjectives, their cliticized forms, and invalid configurations. We evaluate multilingual and Arabic-oriented LLMs under diacritized and undiacritized settings over 600K benchmark entries. Results show that Arabic morphological generation remains difficult, especially for cliticized, unseen, and morphologically rare forms.
1 Introduction
YallaMorph addresses the gap between fluent Arabic generation and systematic morphological control by benchmarking generation from explicit lexical and feature specifications. It evaluates a large, linguistically varied benchmark and finds persistent difficulty, particularly for cliticized, unseen, and rare forms.
- Arabic morphology challenges LLMs because templatic, concatenative, morphosyntactic, and cliticization features create many productive forms.
- Existing Arabic LLM evaluations mainly assess downstream tasks or broad generation quality, limiting insight into controlled generation from lexical and morphological specifications.
- YallaMorph maps a lemma, part of speech, gloss, and target feature bundle to an Arabic surface form or an invalid-configuration indication.
- The benchmark covers verbs, nouns, adjectives, cliticized forms, and invalid configurations across over 600K evaluation instances.
- Evaluations span proprietary, multilingual, and Arabic-oriented instruction-tuned LLMs in diacritized and undiacritized settings.
- The benchmark introduces a linguistically motivated sampling framework and analyzes model performance across linguistic and distributional dimensions.
2 Related Work
Arabic morphological generation builds on established inflection benchmarks and explicit morphological generators, while recent LLM evaluations have only partially addressed Arabic. YallaMorph is positioned within this gap by using a comprehensive Arabic resource for controlled benchmark construction and generation validation.
- SIGMORPHON shared tasks established morphological inflection as a benchmark mapping lemmas and morphosyntactic feature bundles to inflected forms.
- Arabic morphological analyzers and generators have long modeled root-and-pattern structure through finite-state, templatic, and lexical-resource approaches.
- CamelMorph provides over 100K MSA lemmas with rich morphological features and broad paradigm coverage.
- The paper uses CamelMorph’s lexicon and database to construct its controlled benchmark and its CAMeL Tools generation engine to generate and validate forms.
- Recent LLM morphology evaluations include multilingual inflection studies and Arabic-focused work, but Arabic coverage remains limited in scope.
3 Linguistic Background & Terminology
Arabic morphology combines large inflectional paradigms with productive cliticization and interactions between templatic and concatenative morphology. The paper defines its core units through root, pattern, stem, lemma, and inflectional examples involving broken plurals and clitic attachment.
- Linguistic Background: Arabic morphology produces many surface forms through large inflectional paradigms and productive cliticization involving conjunctions, prepositions, articles, and pronouns.
- Linguistic Background: Arabic forms combine templatic root-and-pattern morphology with concatenative affixes or clitics, often accompanied by orthographic and morphophonological alternations.
- Terminology: A lemma abstracts over a lexical item’s inflectional forms, while a root encodes core lexical meaning and a pattern specifies vocalic and templatic structure.
- Terminology: A stem is formed by combining roots and patterns before affixation, separating internal templatic derivation from later morphological attachment.
- Examples: The committee example illustrates broken-plural formation through a pattern change while preserving the root, followed by further variation from enclitic attachment.
- Examples: Clitic combinations can trigger orthographic assimilation, as illustrated by the form involving the preposition li and the plural noun.
4 YallaMorph Benchmark Design
YallaMorph is designed as a large-scale, linguistically balanced benchmark for controlled Arabic morphological generation across major open word classes, inflectional phenomena, cliticization, and invalid configurations. It uses guided, dimension-aware sampling rather than exhaustive enumeration to maintain coverage while controlling benchmark size.
- Benchmark scope: YallaMorph contains over 600K sampled forms spanning verbs, nouns, adjectives, inflectional morphology, and cliticization phenomena.The benchmark is organized into complementary baseword and cliticization subsets.
- Benchmark scope: The benchmark covers five core verbal configurations and separates baseword inflection from interactions between stems and attached clitics.The verbal configurations are active and passive perfective, active and passive imperfective, and command.
- Sampling dimensions: Sampling balances root class, lemma frequency, paradigm completeness, and stem complexity through categorical labels and their joint strata.Lemma frequencies are grouped into low, medium, and high bands, while overlapping root and stem-complexity classifications are resolved with fixed priority hierarchies.
- Sampling dimensions: Paradigm completeness distinguishes masculine-only, feminine-only, and full paradigms, enabling null references for unsupported gender–number combinations.Verbal and adjectival paradigms are always full, whereas nouns may belong to any of the three categories; around 13% of entries have null gold references.
- Sampling procedure: Fixed per-stratum sampling selects up to 10 lemmas from each non-empty stratum, including all lemmas in smaller strata, to preserve coverage and limit over-representation.Cliticized forms are sampled separately from the baseword subset because proclitic and enclitic combinations create substantial expansion.
- Task formulation: The task maps a lemma, part of speech, gloss, and morphological specification to valid inflected forms or an explicit null output for invalid configurations.Outputs may contain one valid form, multiple valid realizations, or a null reference; multiple-reference cases comprise 3.2% of entries.
5 Evaluation
YallaMorph evaluates controlled Arabic morphological generation across models, prompts, orthographic settings, linguistic dimensions, and frequency distributions. Results show strong overall model differences and persistent difficulty with cliticized, rare, unseen, and invalid configurations.
- Overall Performance: GPT substantially outperforms all other models across Any Match Accuracy and F1 in every evaluation setting, followed by Gemini.Arabic-oriented open models perform considerably worse overall, although Arabic prompts help ALLaM and Jais-8B on several zero-shot metrics.
- Detailed Results: GPT shows substantial baseword–cliticization differences, especially under diacritized evaluation, while differences are smaller for undiacritized and normalized outputs.Adjectives and active perfective verbs perform best for basewords, whereas nouns lead cliticized diacritized accuracy.
- Surface-Form Frequency: 12.5% accuracy is achieved on Null-reference configurations, which comprise 12.9% of the benchmark and are the most difficult surface-form category.Performance decreases as surface forms become rarer, although attested forms remain considerably easier.
- Specification Frequency: GPT performance decreases steadily from frequent to rare morphological specifications, with completely unseen specifications especially difficult under diacritized evaluation.The results indicate sensitivity to both lexical exposure and the frequency of underlying morphosyntactic patterns.
- Error Analysis: Cliticization mainly increases the difficulty of realizing the host word correctly rather than producing direct clitic errors.In cliticized errors, 39% involved diacritization-only differences and 39% involved incorrect stems, affixes, or both.
- Error Analysis: Jais-8B returned an unrelated default “book” form in 30,085 instances, while Jais-70B produced romanized forms in 3,473 Arabic-prompt outputs.These recurring behaviors were identified through inspection of the complete test set.
6 Conclusion and Future Work
YallaMorph provides a broad benchmark for controlled Arabic morphological generation and shows that current LLMs struggle particularly with diacritized, cliticized, rare, and zero-frequency forms. Future work will extend the benchmark to additional Arabic varieties and contextual generation.
- Conclusion: YallaMorph covers verbs, nouns, adjectives, cliticized forms, and invalid configurations using controlled sampling across frequency, root class, paradigm completeness, and stem complexity.The benchmark contains over 600K carefully sampled entries.
- Conclusion: Current LLMs remain challenged by Arabic morphological generation, especially in diacritized and cliticized settings, with substantial drops for zero-frequency forms and rare morphological specifications.The findings suggest reliance on distributional exposure and memorization rather than robust generalization across the full morphological space.
- Future Work: Future work will extend YallaMorph to additional Arabic varieties and contextual generation settings, with finer-grained analysis of agreement, stem selection, diacritization, and cliticization errors.
Limitations
The benchmark is bounded by Modern Standard Arabic, resource coverage, guided sampling, and evaluation choices that may affect reproducibility and generalizability.
- YallaMorph focuses on Modern Standard Arabic in CamelMorph MSA, excluding the full diversity of Arabic dialects and mixed-register usage.
- The benchmark is constrained by the coverage, analyses, and generation decisions of its underlying morphological resource.
- Guided sampling includes many forms but does not exhaustively enumerate the full morphological space.
- Lemma-frequency estimates were produced without final BAREC-10M annotations and may differ from frequencies derived from those annotations.
- Prompt-based results may vary with alternative prompting strategies, decoding settings, or model versions, while exact match may miss acceptable orthographic variation.
Ethics Statement
The benchmark uses existing linguistic resources and selected, restricted feature and clitic configurations to evaluate Arabic morphology while acknowledging Arabic’s linguistic diversity.
- YallaMorph is derived from existing linguistic resources and contains no private, personal, or user-generated sensitive information.
- The benchmark targets Modern Standard Arabic and should not support broad claims about Arabic language competence across dialects and registers.
- Core morphological features are POS-dependent: verbs use aspect, person, gender, number, voice, and mood, while nouns and adjectives use gender, number, case, and state.
- Clitics are represented through POS-sensitive proclitic and enclitic positional dimensions, with selected values summarized in the benchmark tables.
- The clitic space is restricted to preserve informative variation while keeping evaluation feasible, rather than enumerating every combinatorially possible configuration.
- Restrictions depend on baseword features, including case and state for nominals and adjectives and verbal group and mood for verbs.
B Model Configuration and Time & Cost Analysis
The evaluation fixes model-generation settings and analyzes runtime and monetary cost, while using selected clitic configurations with compatibility restrictions.
- Model Configuration: All models use temperature 0 and a maximum output length of 256 tokens; open-weight models run locally on NVIDIA A100 and V100 GPUs.
- Model Configuration: Selected clitic values cover representative attachment behaviors across verbs, nouns, and adjectives, but do not exhaust the full inventory.
- Time & Cost Analysis: GPT-5.4 achieves the strongest overall performance and incurs the highest reported monetary cost, totaling $1,996.60.
- Time & Cost Analysis: Open-weight models have no API usage cost, but local computational and infrastructure costs are not included.
- Time & Cost Analysis: GPT-5.4 records the lowest aggregate runtime among evaluated LLMs at 8,104 minutes, while CAMeL Tools reference generation requires 233 minutes without API cost.
- Model Configuration: Clitic configurations are restricted by morphological features, including case and state for nominal and adjectival forms and verbal group and mood for verbs.
C Precision and Recall Results
The section reports micro-averaged Precision and Recall scores corresponding to the F1 results presented elsewhere.
- Table 16 reports micro-averaged Precision and Recall scores corresponding to the F1 results in Table 6.
D Output Cardinality Analysis
Output cardinality analysis distinguishes no-answer, exactly-one, and multiple-answer outputs across prompting settings. Most models predominantly produce one answer, while ALLaM and Jais-8B produce multiple answers more often, especially zero-shot; this analysis concerns output behavior rather than correctness.
- Most models predominantly produce exactly one answer, whereas ALLaM and Jais-8B generate multiple answers more frequently, particularly zero-shot.
- A feature bundle may correspond to no valid form, exactly one valid form, or multiple valid forms, so output cardinality reflects the task design as well as model behavior.
- Table 17 categorizes outputs as No Answer, Exactly One, or Multiple and includes a reference distribution for gold answers.Percentages may not sum exactly to 100% because of rounding.