Source-linked AI summary

Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close

Rania Elbadry, Ahmed Heakl, Saeed Almheiri, Fan Zhang, Muhra AlMahri, Xueqing Peng, Mohsinul Kabir, Shuyao Wang, Yi Han, Saadeldine Eletter, Duzhen Zhang, Preslav Nakov, Yuxia Wang, Fajri Koto, Zhuohan Xie

arXiv:2609.00999v1cs.CLcs.AIcs.LG

TL;DR

In normative pluralism, questions can have valid answers under different frameworks, requiring models to select a framework and answer correctly within it. The paper evaluates this problem in Islamic finance with a four-choice taxonomy and finds that cultural cues can steer models toward a framework while exposing incorrect within-framework answers.

  • Problem

    Questions about Islamic finance can admit valid answers under different normative frameworks, making both framework selection and within-framework correctness relevant.

  • Method

    The paper introduces normative pluralism and evaluates it with a bilingual Islamic-finance benchmark and four-choice taxonomy separating framework selection from within-framework correctness.

  • Results

    Up to 66% of responses are factually wrong within the Islamic framework selected, while cultural signals shift framework selection across twelve models, two languages, and fifty demographic signals.

  • Takeaways & Limitations

    Separating framework selection from within-framework correctness exposes the stereotype trap, which two-choice evaluations can miss when cultural cues redirect models toward the Islamic framework.

  • Takeaways & Limitations

    The findings may not generalize beyond Islamic finance in Arabic and English, and the four-choice format remains vulnerable to answer-position and wording effects.

Abstract

from arXiv · show

When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-framework correctness. This separation reveals the stereotype trap: a cultural cue steers a model toward one framework, but the model selects an incorrect answer within that framework. Across twelve models, two languages, and fifty demographic signals, cultural cues change framework selection and reveal substantial differences in accuracy, especially among non-frontier models. Under the strongest signal, large open-weight models select the Islamic framework 97% of the time. A two-choice evaluation would report near-perfect alignment, although 57--66% of those selections are incorrect. These findings motivate, but do not directly test, the competence-conditioned routing hypothesis: models may favor frameworks where they are more accurate, while cultural cues may expose framework-specific competence gaps.

1 Introduction

Normative pluralism arises when questions admit valid answers under multiple frameworks, making context-sensitive framework selection and within-framework correctness distinct evaluation problems. The paper addresses this gap with a four-choice taxonomy that exposes culturally steered but factually wrong framework selections.

  • Normative pluralism describes questions with valid answers under multiple frameworks, where the appropriate response depends on context rather than a single ground truth.
  • Standard bias benchmarks and preference-only instruments cannot distinguish competent framework selection from stereotyped selection followed by a factually wrong answer.
  • Up to 66% of responses selected the Islamic framework but were factually wrong within it.
  • The four-choice taxonomy crosses framework selection with within-framework correctness using CI, CW, II, and IW answer cells.
  • Across twelve models, two languages, and fifty demographic signals, cultural cues redirect framework selection while within-framework competence varies across model tiers.

2 Related Work

Prior cultural-bias, Islamic-knowledge, financial, and steering benchmarks leave framework selection and within-framework correctness insufficiently separated. This paper extends preference measurement to framework routing in an expert Islamic-finance domain.

  • Existing stereotype benchmarks assume one correct answer, whereas normative framework selection can involve two legitimate answers.
  • Preference-only instruments measure whether a model selects Framework A or B but not whether its selected framework answer is factually correct.
  • Islamic knowledge benchmarks evaluate jurisprudence and scripture but do not isolate finance or test routing across frameworks.
  • Financial bias surveys show demographic signals shift recommendations without controlling for framework possession, while financial NLP benchmarks assume the applicable framework is fixed.
  • Prior steering studies document performance costs and cultural-direction effects but do not compute cross-model relationships between activation and within-framework accuracy.

3 Methods

The benchmark converts expert Islamic-finance material into bilingual, framework-neutral questions with validated Islamic, Western, and distractor answers. Its design combines bilateral and control sets with coherent cultural-signal injections for four-cell evaluation.

  • Benchmark construction: The benchmark contains 304 bilateral questions, 64 Western-anchor controls, and 41 Islamic-anchor controls for separating framework selection from within-framework competence.
  • Benchmark construction: Four stages transform expert-validated SAHM sources through stem neutralization, Western-answer generation, distractor generation, and signal injection.
  • Corpus filtering: Two experts classify 811 samples, yielding 430 candidate-bilateral and 41 Islamic-only questions after excluding 340 non-advisory samples.
  • Stem neutralization: Neutralization preserves financial substance while removing framework-specific terminology and produces parallel English translations validated by experts.
  • Answer construction: Western answers are grounded in primary regulatory text and validated by three financial experts against source documents.
  • Answer validation: Experts confirm substantively different CI and CW products for surviving items with κ = 0.82.
  • Signal injection: Fifty cultural signals span nine families and are injected as short prefixes, with only coherent pairings retained after model classification and expert validation.

4 Results

Across twelve models, two languages, and fifty cultural signals, cues redirected framework selection, but correctness within the selected Islamic framework varied sharply by model capability. The results distinguish cultural activation from competent rule application, exposing a direction-specific stereotype trap.

  • Evaluation setup: Models were evaluated across twelve systems, two languages, 304 bilateral questions, and Western- and Islamic-anchor controls using five reported metrics.The bilateral taxonomy records correct and incorrect Islamic and Western responses, while controls isolate framework-specific competence and general financial competence.
  • Framework Sensitivity: Baseline Islamic routing was higher for recognisable Islamic product names but near zero for topics defined by institutional rules.Only Opus defaulted Islamic at p_islamic = 0.84; a Shariah-compliance request lifted every topic above p_islamic = 0.77.
  • Framework Sensitivity: The strongest Islamic cue, KEYWORD_SHARIA, increased Islamic routing by +0.66, while the strongest Western cue shifted it by only −0.11.Identity-vocabulary signals reliably activated Islamic framing, whereas structurally embedded roles without identity vocabulary did not.
  • Cultural signals: Theophoric names produced the strongest name shift at +0.13, exceeding Muhammad at +0.11, while Arab cultural names produced no measurable shift.Behavioral cues such as Ramadan or Zakat reached two-thirds of the lift from an explicit Muslim declaration, with Ramadan/Zakat at +0.30 versus +0.49.
  • Within-Framework Competence: Frontier models kept Islamic Fake Rate IFR between 0.02 and 0.11, whereas open-weight models ranged from 0.33 to 0.79 under high-activation signals.The worst frontier IFR, 0.114, was three times lower than the best non-frontier IFR, 0.333; Qwen-14B reached IFR 0.66.
  • Within-Framework Competence: Under the Shariah keyword, non-frontier IFR reached 0.58–0.68 while same-item WFR remained 0.10–0.27, and Islamic-anchor IFR remained 0.52–0.58.Across eight signals spanning a 16× activation range, non-frontier IFR stayed at 0.52–0.68 while WFR stayed at 0.21–0.29.

5 Analysis

Cultural cues route models toward Islamic finance, but routing and competence are largely separate: the cue selects the frame while model tier determines correctness within it. Location cues also trigger a coarse geographic rule that misses important regulatory distinctions.

  • Across a 16× activation range, non-frontier IFR stays at 0.52–0.68 while WFR stays at 0.21–0.29.The trap remains comparatively stable as signal activation changes.
  • Tier explains 71.6% of IFR variance, whereas signal family explains 2.6%, indicating that the cue picks the frame while tier determines its contents.
  • Switching from English to Arabic shifts frontier routing by +0.10 to +0.54 while changing frontier IFR by at most 0.03.For non-frontier models, IFR instead moves with routing by +0.03 to +0.17.
  • The Shariah keyword increases IFR on six of seven clusters, but Arabic F_SARF is the exception, with ∆IFR ≈ −0.16 across Arabic-centric models.The exception is governed by AAOIFI Standard No. 1, the most codified rule in the corpus.
  • Models order jurisdictions broadly by Islamic-finance prevalence but cannot distinguish statutory single-system from dual-system regulatory regimes.Explicit regulatory framing adds no information beyond the geographic prior in the Gulf context.
  • In Gulf-anchored conflicts, location overrides explicit identity, applying Islamic routing even where both frameworks are legally available.The rule is correct only in the two statutory single-system jurisdictions identified in the passage.

6 A Single Gate Underlies the Trap

Mechanistic analyses identify a model-controlled gate that commits to a framework in the back half of the network, while cues mainly determine whether their framing survives it. Competence within the selected frame is separate: several models suppress a correct answer, whereas Gemma-3 models lack an early correct lead.

  • A Single Gate Underlies the Trap: Across 40 model-cue cells, patching localizes framework commitment to one back-half gate at proportional depth 0.55–0.84.Patching at or after the gate recovers the Western answer in 59–100% of eligible items.
  • A Single Gate Underlies the Trap: Within models, five cues commit at nearly identical depths, but across models the same cue lands up to ∆=0.42 apart.The gate is therefore cue-invariant but architecture-specific.
  • Cue Survival: KEYWORD_SHARIA raises routing by ∆p_islamic = +0.66 versus +0.13 for a name because its framing survives the gate.Late-layer p_islamic exceeds other cues by +0.19, while weaker inferential cues collapse toward Western framing.
  • Competence Within the Frame: In six of eight models, the CI–II margin is positive early and turns negative at the gate; the two Gemma-3 models never show a positive margin.For most models, the trap is deletion of a correct answer rather than its absence.
  • Competence Within the Frame: Gate depth predicts stereotype rate with r = −0.78, while baseline routing does not with r = −0.01.These results support routing and competence as orthogonal axes.
  • Directional Asymmetry: Islamic traps outnumber Western traps 1,084 to 122, an 8.9× asymmetry favoring CW→II over CI→IW.Arabic-finance specialists occupy the favorable extreme, with lower asymmetry and deeper gates.

7 Conclusion

The paper frames cultural-bias evaluation as normative pluralism and separates framework selection from correctness within the selected framework. This decomposition exposes stereotype traps, in which cultural cues shift models toward Islamic finance but nine of twelve models choose incorrect answers there.

  • The benchmark formalizes normative pluralism as a cultural-bias evaluation setting with a four-choice taxonomy separating framework selection from within-framework correctness.
  • Nine of twelve models select incorrect options within the Islamic framework after cultural cues shift them toward it.

Limitations

The study’s scope is limited to Islamic finance in Arabic and English, with mechanistic analysis restricted to open-weight models and a constrained multiple-choice evaluation design.

  • Findings may not generalize beyond Islamic finance in Arabic and English to other multi-framework domains.
  • Mechanistic results cover only open-weight models and limited cues, so they do not establish that the internal patterns hold for other signal families or closed frontier models.
  • The four-choice format remains vulnerable to answer-position bias and format instability, and untested answer orders and prompt paraphrases leave residual effects unresolved.
  • Only 11 of 23 pre-registered demographic conditions were implemented, limiting coverage of the intended signal space.
  • The evaluations reflect a single model-release snapshot and cannot capture changes introduced by later versions or updates.

Ethical Considerations

The evaluation separates framework selection from correctness and shows that cultural cues can expose serious within-framework competence gaps. The authors caution that identity and location cues do not establish which framework is appropriate for a user.

  • Risks and model behaviour: 57–66% of large open models’ Islamic-framework selections under the strongest signal are incorrect within that framework.The benchmark treats an Islamic-framed incorrect option as an error rather than successful alignment.
  • Risks and model behaviour: Incorrect options can appear authoritative because they preserve standard numbers, citations, and tone while introducing materially wrong financial rules.The paper gives examples involving a fabricated deposit requirement and liability transfer before possession.
  • Interpretive scope: The evaluation measures model behaviour, not which framework any user should receive, and gives equal credit to correct answers under either framework.Error rates are computed within whichever framework the model selects.
  • Interpretive scope: Identity cues function as proxies for framework selection even though they alone do not determine applicable rules or user preferences.Arabic Christian names and explicit Christian declarations can increase Islamic framing, while Gulf locations often dominate conflicting cues.
  • Interpretive scope: Location prompts identify cue effects rather than demonstrated knowledge of national banking mandates or regulators.The paper interprets Tehran and Riyadh findings as location-based routing, not evidence of distinguishing Iran’s or Saudi Arabia’s regulatory systems.
  • Interpretive scope: The benchmark covers products with different procedures under two frameworks, not rules assigning different entitlements to different people.It therefore does not determine when framework-specific personalisation is appropriate.

A.1 Data Availability

The benchmark and its annotation pipeline are released with bilingual datasets, documented rubrics, and interfaces for reviewing neutralisation and answer quality. Construction uses expert classification, coherence filtering, and domain-specific verification.

  • Released resources: The released materials include a bilateral four-cell dataset, a full 50-signal evaluation grid, and a neutralisation audit set.The bilateral base set contains 304 Arabic-English items with CI/CW/II/IW cells.
  • Dataset construction: Two Islamic-finance experts classify source samples as non-advisory, Islamic-only, or candidate-bilateral.The rubric excludes incoherent demographic framing and distinguishes constructs without Western counterparts from questions admitting substantively different answers.
  • Annotation pipeline: Neutralisation review removes framework leakage while preserving product type and customer situation, and bilingual review checks semantic equivalence without added or missing facts.Terms such as “profit-sharing” and “lease-to-own” are identified as possible framework hints.
  • Annotation pipeline: Domain-matched experts verify distractor errors, semantic substance, and plausibility, while Western answers are checked for accuracy and source entailment.Generated answers are evaluated against the full source document; the Islamic answer serves only as a structural reference.
  • Signal evaluation: After coherence filtering, each question retains an average of 38.6 coherent signals per language.Name and religion signals retain more than 90% of cells, whereas conflict and occupation signals retain about 60%.

H Additional Validity Checks

Additional checks reject tokenisation and prompt-length effects as explanations for framework shifts and confirm that distractors discriminate competence. A major limitation remains residual position confounding from one deterministic shuffle per item.

  • Mechanism checks: Tokenisation is rejected as the mechanism: correlations between signal subword count and Δp_islamic are +0.31 in English and +0.16 in Arabic, opposite the fragmentation hypothesis.The analysis uses six open-weight tokenisers.
  • Mechanism checks: Prompt-length attraction is ruled out: the length-matched placeholder differs from ZERO_SIGNAL by 0.024 in mean absolute p_islamic deviation without systematic direction.Fourteen cells favour Western framing and ten favour Islamic framing.
  • Robustness checks: The coherence-classification model preserves the signal hierarchy, with per-signal values within 0.02 of the panel mean.This recomputation uses that model’s evaluation rows alone.
  • Position-bias limitation: Residual position confounding cannot be excluded because each item used only one deterministically shuffled answer order.Position sensitivity peaks at 52.5 percentage points for Llama-8B in Arabic, and the authors plan all 24 letter assignments per item.
  • Distractor validation: Distractors discriminate competence: point-biserial correlations are −0.157 for II items and −0.337 for IW items, with 73% and 95% respectively below −0.1.Higher-scoring models systematically avoid these distractors.

J Trajectory Threshold Sensitivity

Trajectory categories remain directionally stable across stricter and more lenient thresholds, with frontier models showing cleaner trajectories than large models. The Religion=Islam pattern is uniformly tier-dependent across product clusters.

  • Threshold sensitivity: The 30:0 versus 0:26 directional contrast persists under strict, default, and lenient trajectory thresholds.Threshold changes alter absolute counts but not the directional contrast.
  • Threshold sensitivity: Frontier models record zero TRAP cells except one borderline lenient-threshold cell and the highest per-tier CLEAN_LIFT count.Large models record zero CLEAN_LIFT cells and the highest per-tier TRAP count.
  • Continuous trap measure: Between-tier variance in the continuous trap coefficient τ exceeds within-tier variance by a factor of 10.5.This supports structural uniformity in τ beyond the discrete trajectory categories.
  • Cross-cluster stability: Cluster-level signal orderings are highly consistent overall, with mean pairwise Spearman correlation 0.96, but insurance is an outlier at 0.76.Six of seven cluster diagonals excluding self have correlations of at least 0.89.
  • Religion=Islam confusion: Frontier models show negative Christian-cue shifts across all seven clusters, whereas large models show positive shifts across all seven.The ranges are −0.072 to −0.190 for frontier and +0.143 to +0.366 for large models.

N Trap Localisation: M1 Experiment Details

The experiment localizes when cultural cues flip models from correct Western answers to incorrect Islamic answers, using activation patching and logit-lens analyses across 40 model–cue cells. A representative keyword slice illustrates the resulting trajectories, while the scope-wide claims use all 40 cells.

  • Experiment scope: 40 model–cue cells combine eight open-weight models with five cultural cues, using trap-flip items that shift baseline CW answers to cue-conditioned II answers.The five cues are keyword, religion, occupation, location, and name.
  • Method 1: Activation patching: Activation patching replaces the cue-conditioned last-token residual with its baseline counterpart at each decoder layer to test recovery of the original Western answer.Sweeping the replacement layer identifies the gate at which the cue-conditioned behavior changes.
  • Method 1: Logit lens: Logit lens decodes every layer’s residual into the four-choice distribution {pCI, pCW, pII, pIW} to track falling pCW and rising pII.The analysis compares baseline and cue-conditioned forward passes layer by layer.
  • Representative slice: The representative KEYWORD_SHARIA slice covers three models across Large, Midsize, and Arabic-Centric groups, with gate depths of 0.67–0.84 and post-gate recovery of 0.59–0.96.These rows illustrate model-level trajectories; cross-model and cross-cue claims rely on all 40 cells.
Loading 2609.00999v1…