Source-linked AI summary
Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning
Mena Attia, Mona Diab, Thamar Solorio
TL;DR
Whether cultural knowledge and figurative-language understanding support one another in LLMs remains untested. Through controlled Arabic fine-tuning experiments, the study finds limited, inconsistent, model-dependent transfer, with poetry improving idiom interpretation by +2.33% (p = 0.021).
Problem
Whether cultural knowledge and figurative-language understanding support one another in large language models remains untested.
Method
The study uses controlled Arabic fine-tuning and evaluation experiments to test bidirectional transfer between cultural and figurative-language understanding.
Results
Transfer is limited, inconsistent, and model-dependent; poetry fine-tuning improves idiom interpretation by +2.33% (p = 0.021), while cultural fine-tuning lowers proverb accuracy in Arabic-centric models.
Takeaways & Limitations
The relationship between cultural and figurative knowledge is not straightforwardly captured through fine-tuning alone.
Takeaways & Limitations
The cross-domain transfer directions use different datasets, benchmarks, and task formulations, limiting direct comparison and the generality of cross-domain claims.
Abstract
from arXiv · showhide
Figurative language is deeply culturally embedded; fluent use requires not just linguistic competence but cultural immersion. We ask whether LLMs can learn this link: does fine-tuning on cultural data improve figurative language understanding, and vice versa? We conduct a systematic study across four models (ALLaM-7B, Fanar-1-9B, Qwen3-8B, Llama-3.1-8B) and six Arabic datasets spanning cultural commonsense, proverbs, and poetry across diverse dialects and regions. Fine-tuning on poetry improves idiom comprehension (+2.33%, p<0.05), a gain our ArabicMMLU control does not reproduce, indicating that it stems from figurative content rather than Arabic language adaptation and pointing to a sensitivity to non-literal meaning that transfers across figurative types. Cultural fine-tuning, by contrast, lowers proverb-interpretation accuracy in both Arabic-centric models. Transfer between the two domains is otherwise indistinguishable from noise, with Arabic models frequently regressing after fine-tuning, suggesting prior saturation of relevant knowledge, while multilingual models show greater adaptation headroom. Error analysis further reveals that fine-tuning reinforces experiential cultural knowledge while destabilizing historically grounded factual knowledge. Our findings suggest that the relationship between culture and figurative language, though conceptually natural, is not straightforwardly captured through fine-tuning alone.
1 Introduction
The section motivates studying cultural knowledge and figurative competence as coupled capacities and introduces controlled Arabic fine-tuning experiments to test bidirectional transfer. It presents transfer as limited, uneven, and model-dependent, with poetry fine-tuning the only reliable positive source of figurative transfer.
- Motivation: Figurative interpretation draws on culture-specific values, norms, historical experience, and community familiarity beyond linguistic competence alone.Cognitive-linguistic and figurative-language theories document cross-cultural variation in how figurative meanings are conventionalized and interpreted.
- Research gap: Cultural knowledge and figurative-language understanding are studied separately in NLP, leaving their potential mutual support in large language models untested.Cultural benchmarks assess factual, commonsense, social, and regional knowledge, whereas figurative benchmarks assess proverbs, idioms, poetry, and metaphor.
- Approach: The study uses controlled Arabic fine-tuning experiments to test whether cultural supervision improves idiom and proverb interpretation and whether figurative supervision improves culturally grounded question answering.The framework examines both directions of cross-domain transfer and also investigates transfer across figurative forms.
- Contributions: Fine-tuning effects are limited, unevenly distributed, and dependent on model family, dialect, topic, and knowledge type.The contribution includes model-, dialect-, and topic-level analyses of cross-domain transfer.
- Main finding: Poetry fine-tuning is identified as the only reliable source of positive figurative transfer, while overall transfer remains small, inconsistent, and highly model-dependent.These findings indicate that the relationship between cultural and figurative knowledge is not straightforwardly captured through fine-tuning.
2 Related Work
Prior work largely treats figurative competence and cultural knowledge as separate adaptation targets, while Arabic benchmarks include figurative language only peripherally. This study instead examines cross-domain transfer between cultural and figurative supervision, including transfer from poetry to proverbs and idioms.
- Benchmark coverage: Existing figurative-language benchmarks span English and multiple languages but provide thinner multilingual coverage than English resources.Fig-QA frames figurative interpretation as multiple-choice question answering, FLUTE adds human-written idiom explanations, and MABL evaluates eight languages.
- Benchmark coverage: Figurative language remains peripheral in Arabic and global benchmarks, motivating dedicated evaluation.ArabCulture includes five idiom samples per country, while Palm treats proverbs as one of 20 topics without in-depth evaluation.
- Adaptation approaches: Prior adaptation work uses task-specific fine-tuning, structured prompting, cultural fine-tuning, or culture-specific adapters to improve implicit or culturally grounded understanding.Arabic cultural adaptation includes the PalmX shared task, while lightweight alignment has shown cultural-commonsense transfer across Arab countries.
- Cross-domain transfer: This study tests whether supervision in one domain transfers to the other and whether poetry transfers across figurative forms to proverbs and idioms.Unlike prior work, which targets cultural knowledge or evaluates figurative competence independently, the study directly investigates cross-domain transfer.
3 Experimental Setup
The study uses a controlled, bidirectional fine-tuning framework to test transfer between cultural knowledge and figurative-language understanding in Arabic LLMs. It compares cultural, figurative, and general-Arabic control training across multiple datasets and model families.
- Experimental design: The framework tests bidirectional transfer by measuring how cultural fine-tuning affects idiom and proverb understanding, and how figurative-language fine-tuning affects the same benchmarks.A general-Arabic condition tests whether observed transfer reflects cultural content rather than language adaptation.
- Datasets: Training uses cultural supervision from ArabCulture and Palm, figurative supervision from FannOrFlop and Jawaher, and evaluates culture on AraDiCE and figurative understanding on Jawaher and Kinayat.ArabicMMLU provides the language-adaptation control.
- Data design: Each condition uses 1,000 training samples, except Jawaher, which uses 800; proverb-topic items are removed from Palm to prevent figurative supervision in cultural training.Full-corpus training for FannOrFlop and ArabCulture is reported as an ablation.
- Models and evaluation: Evaluation uses identical lm-eval procedures for base and fine-tuned models, scoring multiple-choice items by answer-option log-likelihood and reporting accuracy averaged across three runs.Generative datasets use expressions as inputs and explanations or interpretations as targets.
4 Results
Cross-task transfer between cultural knowledge and figurative-language understanding is generally small, model-dependent, and statistically indistinguishable from zero. The sole reliable aggregate gain is poetry fine-tuning on idiom interpretation, while cultural fine-tuning does not reliably improve figurative benchmarks and can produce supported regressions.
- Overall transfer: +2.33% is the only statistically reliable pooled effect, produced by poetry fine-tuning on Kinayat idiom interpretation.Its 95% confidence interval excludes zero, and the effect is driven primarily by LLaMA-3.1-8B, whose largest single-seed gain is +10.00% (p = 0.008).
- Model-level variation: −3.70% and −3.03% are the supported pooled regressions on Jawaher, affecting ALLaM-7B and Fanar-1-9B after Palm fine-tuning.Both regressions are consistent across all three seeds; significant individual seeds reach −5.56% (p = 0.013) and −3.54% (p = 0.039), respectively.
- Cultural fine-tuning: Cultural fine-tuning does not reliably improve Jawaher or Kinayat, with ArabCulture and Palm producing pooled effects whose confidence intervals span zero.ArabCulture yields −0.08% on Jawaher and +0.78% on Kinayat, while Palm yields −0.46% on Jawaher and +1.22% on Kinayat.
- Figurative-language fine-tuning: −0.19% and +0.00% are the pooled AraDiCE effects of proverb and poetry fine-tuning, respectively, with no detectable improvement.The effects have p = 0.81 and p = 0.99, and model-level variation includes a cautious ALLaM-7B regression of −1.85% under poetry fine-tuning.
- Control fine-tuning: −2.39% and −0.84% are the ArabicMMLU control effects on Kinayat and Jawaher, respectively, contrasting with poetry’s positive Kinayat effect.The control moves in the opposite direction from poetry fine-tuning on Kinayat, a divergence of nearly five percentage points, while all four models improve on AraDiCE by +2.41%.
- Model-level variation: Average effects across fine-tuning conditions range from −0.46% to +2.33%, and Arabic-centric models benefit less than the multilingual models.ALLaM-7B and Fanar-1-9B record negative deltas in most combinations, whereas Qwen3-8B and LLaMA-3.1-8B often show positive point estimates.
5 Error Analysis
Figurative fine-tuning produces nearly balanced item-level improvements and regressions overall, but its effects vary across models, countries, and topics. Where predictions shift, fine-tuning favors experiential cultural knowledge over historically grounded or domain-specific factual knowledge.
- Aggregate Effects: 185 improvements and 189 regressions yield a net degradation of −4 and negligible average change of −0.11%.These aggregate counts motivate item-level analysis rather than a claim of meaningful overall gain or loss; category patterns are descriptive and untested individually.
- Model-Level Trends: ALLaM shows the largest degradation, particularly under poetry (−1.87%), although that direction is better supported than its magnitude.The −1.87% estimate rests on one significant seed (−5.00%, p = 0.022) alongside +1.11% and −1.67% in the other seeds; Fanar and Qwen show small positive estimates, while LLaMA is mixed.
- Geographic Variation: Jordan has the strongest net improvement (+32), whereas Syria (−18) and Egypt (−12) show the largest geographic degradations.Country counts aggregate improvement and regression instances across models and fine-tuning datasets, so one question can contribute multiple counts.
- Topic-Level Analysis: Food/Cuisine (+15) and Traditional Games (+13) show the strongest topic gains, while Traditional Clothing (−14), History/Civilization (−8), and Religion (−8) show the largest losses.History/Civilization and Religion effects rely on few items and should be interpreted cautiously; improved questions cluster around popular cultural practices and celebrations.
- Knowledge-Type Sensitivity: Figurative fine-tuning tends to strengthen culturally embedded, experiential knowledge while destabilizing historically grounded or domain-specific factual knowledge.This topical sensitivity appears beyond model or dialect differences, including shifts involving traditional clothing and infrastructural knowledge.
6 Conclusion
Across four models and six Arabic datasets, fine-tuning transfer between cultural knowledge and figurative language is limited, inconsistent, and model-dependent. The reliable improvement occurs within the figurative domain: poetry fine-tuning improves idiom interpretation, while cultural fine-tuning yields small, unreliable pooled gains and supported regressions.
- 6 Conclusion: Transfer across cultural and figurative domains is limited, inconsistent, and highly model-dependent across four models and six Arabic datasets.The datasets span diverse Arabic dialects and regions.
- 6 Conclusion: +0.78% for ArabCulture and +1.22% for Palm on Kinayat are small and unreliable pooled gains from cultural fine-tuning.Per-model effects are polarized in sign.
- 6 Conclusion: Palm fine-tuning produces the only statistically supported regressions in the study, lowering accuracy for both Arabic-centric models.Arabic-centric models start from higher baselines and account for both supported regressions.
- 6 Conclusion: +2.33% poetry fine-tuning gain on idiom interpretation is statistically supported (p = 0.021), unlike the cross-domain cultural-transfer effects.The reliable effect occurs within the figurative domain rather than across the cultural–figurative divide.
Limitations · A Prompt Templates · B Fine-tuning Setup
The paper notes limitations in adaptation-method coverage and cross-domain comparability, while documenting standardized prompt templates and an efficient LoRA fine-tuning setup. The setup freezes base parameters and uses fixed low-rank updates, training controls, and validation monitoring.
- Limitations: The fine-tuning experiments use only LoRA because computational constraints prevented testing full fine-tuning or alternative adaptation methods.The authors note that other methods may yield different results.
- Limitations: Hyperparameter search was limited, leaving learning rate, rank, and training duration insufficiently explored.The authors indicate that more careful tuning could affect the results.
- Limitations: The two cross-domain transfer directions are not directly comparable because they use different source datasets, target benchmarks, and task formulations.Differences in magnitude may reflect dataset difficulty, domain specificity, or evaluation sensitivity rather than directional asymmetry.
- A Prompt Templates: The MCQ understanding prompt presents a question with candidate answers and elicits a single-letter selection.The corresponding AraDiCE template uses the same structure with three candidate options.
- A Prompt Templates: Fine-tuning prompts cover FannOrFlop poetry explanation, Jawaher proverb explanation, and ArabCulture open-ended completion.The prompt templates are shown in Figures 5, 6, and 7.
- A Prompt Templates: The Palm instruction-following prompt uses no task-specific preamble, with the instruction field carrying the full context.This template is shown in Figure 8.
- B Fine-tuning Setup: LoRA is applied to attention and feed-forward projection layers with rank r = 4, scaling factor α = 8, and dropout 0.1.The adaptation configuration is described in the fine-tuning setup.
- B Fine-tuning Setup: Training runs for 3 epochs at a learning rate of 5e −5 with batch size 1, validation monitoring, frozen base parameters, and a single NVIDIA RTX 5000 Ada Generation GPU.The setup is intended to enable efficient adaptation while keeping the base model parameters frozen.
C Additional Results · D Ablations
The additional results assess stability across random seeds and visualize how fine-tuning on several datasets changes performance relative to base models. These analyses use accuracy scores (↑) and per-dataset performance differences defined as fine-tuned accuracy − base accuracy.
- C Additional Results: Table 7 evaluates Jawaher, Kinayat, and AraDiCE across three random-seed runs for base and subset-fine-tuned models.Multiple runs assess whether observed trends are stable under random initialization.
- D Ablations: No separate ablation result is specified in the supplied passages for section D Ablations.The provided evidence describes additional-results tables, heatmaps, and a baseline comparison only.
- C Additional Results: Figures 9–12 visualize per-dataset performance differences for models fine-tuned on ArabCulture, FannOrFlop, Jawaher, and Palm.Each difference is defined as fine-tuned accuracy − base accuracy.
- C Additional Results: Figure 9 reports performance differences across datasets after fine-tuning on ArabCulture.The plotted measure is diff = finetuned accuracy - base accuracy.
- C Additional Results: Figure 10 reports performance differences across datasets after fine-tuning on FannOrFlop.The plotted measure is diff = fine-tuned accuracy - base accuracy.
- C Additional Results: Figure 11 reports performance differences across datasets after fine-tuning on Jawaher.The plotted measure is diff = fine-tuned accuracy - base accuracy.
- C Additional Results: Figure 12 reports performance differences across datasets after fine-tuning on Palm.The plotted measure is diff = fine-tuned accuracy - base accuracy.
- C Additional Results: The additional analysis includes a cross-model performance-difference view and an ArabicMMLU baseline comparison.Figure 13 defines the comparison as diff = fine-tuned accuracy - base accuracy.
D.1 Increasing the size of the fine-tuning data · D.2 Changing LoRA hyperparameters
Increasing fine-tuning data produced model- and dataset-dependent changes, with full FannOrFlop training helping some models but harming others. Alternative LoRA hyperparameters generally failed to improve ALLaM, apart from a few isolated gains.
- D.1 Increasing the size of the fine-tuning data: Full-dataset fine-tuning on FannOrFlop and ArabCulture produced results that varied across models and evaluation sets.The comparison used 6,984 FannOrFlop samples and 3,482 ArabCulture samples instead of the default 1,000-sample subsets.
- D.1 Increasing the size of the fine-tuning data: ALLaM and Qwen consistently benefited from full FannOrFlop training across Jawaher, Kinayat, and AraDiCE.The passage reports positive performance differences across all three benchmarks.
- D.1 Increasing the size of the fine-tuning data: Fanar declined across all FannOrFlop benchmarks, with its largest loss on Jawaher at −3.37%.LLaMA showed its largest FannOrFlop drop on Kinayat at −8.44%.
- D.1 Increasing the size of the fine-tuning data: For ArabCulture, ALLaM improved modestly across benchmarks, while Fanar gained +1.48% on AraDiCE and Qwen and LLaMA declined.The supplied passage ends while describing LLaMA’s declines, so no additional value is inferred.
- D.2 Changing LoRA hyperparameters: Alternative LoRA settings largely failed to improve ALLaM relative to the default configuration.Table 9 compares different rank, alpha, and learning-rate settings on Jawaher, Kinayat, and AraDiCE.
- D.2 Changing LoRA hyperparameters: The configuration r = 16, α = 32, and lr=1e-5 improved Jawaher by 1.18% and AraDiCE by 2.04% on Palm Subset.These were the only notable gains identified for the alternative LoRA settings.
E Confidence Intervals and Significance Testing. · F Error Analysis · F.1 Using Figurative Language to Improve Cultural Understanding
The paper quantifies uncertainty with bootstrap confidence intervals and exact tests, finding that only pooled poetry fine-tuning on Kinayat idioms shows a significant effect. Error analysis then examines the selective, asymmetric transfer of figurative-language fine-tuning across cultural knowledge categories.
- E Confidence Intervals and Significance Testing.: 95% confidence intervals were computed for every reported accuracy difference using 20,000-item paired-bootstrap resamples.The evaluation sets contain 198 Jawaher, 150 Kinayat, and 180 AraDiCE items.
- E Confidence Intervals and Significance Testing.: When fewer than ten predictions differed, the exact test was treated as authoritative because percentile-bootstrap distributions were too sparse.Across most conditions fewer than twenty predictions changed, and several Fanar and Qwen seeds had fewer than ten changes.
- E Confidence Intervals and Significance Testing.: Cluster bootstrap resampled items jointly across models and seeds, avoiding an independence assumption for repeated measurements of the same item.The clustering correction remains valid under arbitrary dependence among measurements of each resampled item.
- E Confidence Intervals and Significance Testing.: The reported intervals were generally wide and straddled zero, reflecting that most conditions changed fewer than twenty predictions.Tables 19–21 report per-seed intervals alongside improvement and regression counts for Kinayat, Jawaher, and AraDiCE.
- E Confidence Intervals and Significance Testing.: +2.33, p = 0.021 was the only significant pooled effect: poetry fine-tuning on Kinayat idioms.Every other pooled effect was within noise of zero, including all three AraDiCE conditions; pooled AraDiCE estimates were essentially null.
- E Confidence Intervals and Significance Testing.: Only five individual runs reached p < 0.05 among 72 per-seed tests, close to the number expected under the global null.The paper therefore treats per-cell results as descriptive and bases substantive claims on pooled estimates.
- F Error Analysis: Error analysis identifies which cultural knowledge categories improve or degrade after figurative-language fine-tuning.The analysis aggregates AraDiCE questions across models fine-tuned on proverbs and poetry to characterize transfer in this direction.
- F.1 Using Figurative Language to Improve Cultural Understanding: The transfer from figurative language to cultural understanding is asymmetric and selective rather than uniformly beneficial.The analysis provides a granular view of categories most receptive to transfer and those most prone to degradation.
F.2 Using Culture and Poetry to Improve Figurative Language Understanding
Transfer across figurative-language fine-tuning is selective and geographically uneven across proverb dialects. Cultural and poetry fine-tuning produce persistent gains for some varieties, regressions for others, and instability at the individual-expression level.
- Item-Level Analysis: Individual expressions show selective and unstable transfer, with some idioms and proverbs improving in certain models or seeds while regressing in others.Persistent dialect-level gains and losses suggest that evaluation-set properties may matter more than the choice of fine-tuning data.
- Dialect-Level Analysis: −11 for Algeria, −10 for Sudan, and −8 for Libya mark the sharpest proverb regressions under cultural fine-tuning, while Mauritanian (+20) and Yemeni (+14) benefit most.Iraqi and Tunisian varieties also improve by +5. These patterns do not align straightforwardly with representation in the cultural fine-tuning data.
- Dialect-Level Analysis: −8 for Sudanese, −8 for Qatari, and −5 for Algerian varieties show the strongest proverb regressions under poetry fine-tuning, while Yemeni (+8) and Mauritanian (+8) benefit most.Jordanian varieties also improve by +7, and the geographic split largely persists across fine-tuning conditions.
- Dialect-Level Analysis: Qatari proverbs diverge between conditions, becoming a clear regression category under poetry fine-tuning because one proverb accounts for 6 regressions.The same variety was only mildly negative under cultural fine-tuning.
G Use of AI Assistants · ¯ ÈAªË@ · àA¿ ñË ð Q J.ºË@ ÐC¿ ©ÖÞ @
The manuscript used LLM-based writing assistants for rephrasing and editing, while authors retained responsibility for research decisions and final wording. The accompanying analyses catalogued unstable and frequently changed idioms and proverbs across fine-tuning conditions and dialect varieties.
- G Use of AI Assistants: LLM-based writing assistants supported manuscript writing and editing through rephrasing and refinement suggestions.The authors made all research design, analysis, interpretation, and final wording decisions.
- G Use of AI Assistants: 15 idioms were unstable under poetry fine-tuning, improving in some seeds but regressing in others across all models.The table reports counts for a representative subset of these unstable idioms.
- G Use of AI Assistants: 7 proverbs were unstable under poetry fine-tuning, with improvements in some seeds and regressions in others.The table identifies these unstable proverbs under the poetry condition.
- ¯ ÈAªË@: Fine-tuning effects on idioms were broken down for cultural data and poetry data using accuracy changes, prediction counts, confidence intervals, and exact McNemar tests.The analysis covers Kinayat idioms across ArabCulture, Palm, and FannOrFlop.
- ¯ ÈAªË@: Fine-tuning effects on proverbs were similarly evaluated across cultural and poetry conditions using accuracy changes, prediction counts, confidence intervals, and exact McNemar tests.The analysis concerns Jawaher proverbs and distinguishes improved from worsened individual predictions.
- ¯ ÈAªË@: Figurative fine-tuning effects on AraDiCE-Culture were assessed for Jawaher and FannOrFlop using accuracy changes, prediction counts, confidence intervals, and exact McNemar tests.The reported change is from base to fine-tuned accuracy.
- àA¿ ñË ð Q J.ºË@ ÐC¿ ©ÖÞ @: The paper lists the most frequently improved and regressed idioms and proverbs across cultural and poetry fine-tuning, including counts and dialect varieties.These summaries cover cultural idioms, poetry idioms, cultural proverbs, and poetry proverbs.