Source-linked AI summary
DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs
Bhuvan Arora, Devesh Saraogi, Sravya Varada, Dhruv Kumar
TL;DR
Cultural-bias benchmarks often assume a single correct answer, making default preferences difficult to measure when multiple culturally grounded responses are valid. DiSCo introduces a distribution-first forced-choice framework with a C0–C3 context gradient and evaluates six LLMs on a 304-item, 12-culture benchmark. Models show concentrated cultural priors, prompt steering widens the high- versus low-resource selection gap, and fact injection barely changes distributions.
Problem
Existing evaluations often assume a correct answer or target distribution, limiting characterization of default cultural preference priors when multiple culturally grounded responses are valid.
Method
DiSCo evaluates culturally grounded forced-choice items across C0–C3 to separate default priors from contextual adaptation, using DiSCo-Bench and six LLMs.
Results
Prompt-based steering widens the high- versus low-resource selection gap, while fact injection produces negligible distributional disruption across models.
Takeaways & Limitations
Cultural preference priors persist across model families despite explicit cultural prompting and added cultural facts, exposing a gap between surface personalization and deeper cultural adaptation.
Takeaways & Limitations
The dataset is English-only and operationalizes culture through country or region labels, which the authors identify as limitations for future work.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Existing cultural benchmarks evaluate accuracy against a single "correct" answer, making it difficult to characterise an LLM's cultural preference prior when multiple culturally grounded responses are all valid; they also conflate default preferences with context-driven adaptation. We propose DiSCo, a distribution-first forced-choice evaluation framework that isolates default cultural priors and tests steerability via a four-level context gradient (C0--C3). Using DiSCo-Bench (304 items) derived from BLEnD spanning 12 cultures, we evaluate six diverse instruction-tuned LLMs. Default priors are heavily concentrated, with UK and US together absorbing approximately 35\% of all selections despite representing only 2 of 12 cultures. Most critically, prompt-based steering consistently widens the selection gap between high- and low-resource cultures, and injecting explicit cultural facts produces negligible distributional disruption, confirming that cultural preference bias cannot be resolved through prompt-based personalisation alone.
1. Introduction
DiSCo addresses gaps in cultural-bias evaluation by measuring default selection distributions when multiple culturally grounded answers are valid and separating priors from contextual adaptation. It evaluates a graded context framework and finds concentrated, persistent cultural preferences across models.
- Cultural preference bias concerns systematic overselection of some cultures’ options and underselection of others, independent of instruction or contextual signals.
- Existing evaluations often assume a correct answer or target distribution, obscuring default cultural priors when multiple responses are culturally valid.
- DiSCo uses a four-level context escalation from no cultural signal through location cues, locally appropriate instructions, and untargeted cultural fact injection.
- The benchmark spans 12 cultures and tests six instruction-tuned LLMs with option-order and C3 context-order rotations.
- All models show concentrated UK/US-dominant priors; C1 compliance is 0.43–0.57, C2 stickiness is 0.34–0.50, and JSD remains ≤0.018.
2. Methodology
The methodology constructs a balanced forced-choice benchmark and evaluates cultural preference through a C0–C3 context gradient with distributional, steerability, parity, and disruption metrics. Rotations and exposure normalization address multiple-choice and representation artefacts.
- DiSCo-Bench presents four culturally grounded lifestyle options that are all valid, allowing preference distributions to be measured rather than accuracy.
- The dataset covers 304 questions and 12 cultural regions, with approximately uniform culture frequency across option slots.
- The C0–C3 protocol progresses from default prior measurement to location hints, explicit local-appropriateness intent, and simultaneous per-option fact injection.
- C3 independently rotates option and context-fact order in a 4 × 4 = 16 combination design to measure and control primacy bias.
- CSD is the exposure-normalised selection probability per culture; KL, Gini, compliance, Signal Lift, PSI, SPD, and JSD quantify concentration, steering, parity, and distributional shift.
- High PSI indicates persistence of the default prior under explicit location and locally appropriate instructions, while near-zero JSD indicates fact injection leaves the prior unperturbed.
3. Experimental Setup
Six instruction-tuned LLMs were evaluated zero-shot through the OpenRouter API under deterministic, JSON-constrained output conditions.
- Six instruction-tuned LLMs were evaluated zero-shot with temperature = 0 and constrained JSON output.
4. Results & Discussion
Across six models, cultural preferences remain concentrated, unevenly steerable, and resistant to factual prompt grounding. Steering improves compliance but disproportionately benefits high-resource cultures, widening the equity gap.
- C0: Default Cultural Prior: All six models show concentrated default cultural priors, with UK and US accounting for approximately 35% of C0 selections despite representing 2 of 12 cultures.KL values range from 0.124 to 0.197, with DeepSeek and LLaMA strongest and Gemma and Qwen weakest.
- C1–C2: Steering: C1 location cues raise compliance above the 0.25 random baseline, while explicit locally appropriate directives improve it further at C2.CR ranges from 0.43–0.57 at C1 and 0.50–0.66 at C2.
- C1–C2: Cross-Culture Equity: The weakest location-cue signal occurs for cultures that resist maximum prompting most strongly, producing a compounding disadvantage for underrepresented cultures.Northern Nigeria, Ethiopia, and Assam have low Signal Lift, while high-resource cultures such as the UK receive moderate lift partly from existing priors.
- C2: Prior Stickiness: C2 improves compliance for every model but leaves substantial prior stickiness, especially for Northern Nigeria, Ethiopia, and Assam.DeepSeek and LLaMA still resist approximately 34% of the time, while Northern Nigeria reaches PSI = 0.592.
- C3: Fact Injection: Fact injection produces negligible distributional disruption, while primacy effects remain above random but content-driven culture consistency remains dominant.JSD ranges from 0.007 to 0.018, and C3 shows PPR of 0.28–0.42 with CSC of 56–69%.
- Cross-Condition Pattern: Across conditions, Northern Nigeria, Ethiopia, and Assam consistently combine the lowest default selection, weakest Signal Lift, and highest Prior Stickiness Index.This pattern is consistent across all six model families and is presented as a structural property associated with training-data composition.
- Equity Implications: Prompt-based steering widens the high- versus low-resource equity gap for every model, with SPD reductions ranging from −0.037 to −0.085.Steering is more effective toward high-resource cultures than low-resource cultures.
5. Conclusion
Across six model families, DiSCo finds concentrated cultural priors, limited and unequal prompt steerability, and negligible disruption from explicit cultural facts. These results position cultural preference bias as a distributional equity problem requiring interventions beyond prompt engineering.
- All models favour UK and US options by default, with KL = 0.12–0.20 and Gini = 0.33–0.40.
- Location cues produce only moderate steering, while maximum identity-based steering leaves 34–50% residual stickiness.
- Prompt-based steering widens the equity gap between high- and low-resource cultures, with SPD increasing at C2 for every model.
- Explicit cultural fact injection produces negligible distributional disruption, with JSD < 0.02 bits across all models.
- Future work should test category, language, retrieval, calibration, and fine-tuning effects on cultural preference bias.
Impact Statement
The paper frames cultural preference bias as consequential for globally deployed systems because culturally mismatched suggestions can reduce trust and amplify harms. It argues that prompt-based personalisation is insufficient and motivates training-level and broader evaluation interventions.
- The released materials are English-only and operationalise culture through country or region labels, leaving these as stated limitations.
- Existing work covers stereotypes, cultural knowledge, values, persona assignment, and cultural steering, but these approaches address different constructs.
- The study measures benign lifestyle preferences as distributions over culturally tagged choices rather than stereotype violations or toxicity scores.
- The dataset pipeline generalises BLEnD questions, removes geographic anchors, and filters rows containing dummy-culture placeholders.
- After filtering, DiSCo Dataset contains 150,816 rows covering 304 unique questions, with 89 original questions dropped entirely.
B.2. Evaluation Benchmark Curation (Extended)
DiSCo-Bench is curated from a larger BLEnD-derived dataset into a balanced 304-item benchmark with four culturally grounded options per item. Its prompt suite preserves the question and options while escalating cultural context from none to location, intent, and fact injection.
- Dataset construction: Each DiSCo Dataset row contains a country-neutral question, four equally valid culturally grounded options, hidden culture metadata, and four per-option cultural facts.
- Dataset construction: The benchmark filters to 12 regions covering all 304 questions, yielding 43,870 eligible rows before sampling.
- Dataset construction: One row per question is sampled to maximise balanced representation across the 12 cultures and option slots.
- Dataset construction: Culture frequencies across 1,216 option slots remain near uniform, supporting the 1/12 ≈0.083 baseline.
- Prompt design: Core fields include the neutral question, fixed instructions, four options, hidden metadata, and context fields aligned to the options.
- Prompt design: C0–C3 hold the question, instructions, options, and output constraint constant while adding progressively stronger cultural context.
- Evaluation design: The evaluation uses four option-order rotations and 12,160 base runs per model to control letter-position effects and compare conditions.
- Metrics: KL divergence measures deviation from a uniform reference, while JSD is preferred for comparing full distributions without KL’s asymmetry or zero-probability issue.
C. Extended C0 Results
C0 results show substantial and consistent inequality in default cultural selections across all six models. UK and US dominate the distribution, while several cultures receive persistently low selection shares.
- Gini values range from 0.327 to 0.397 across models, confirming substantial inequality in default cultural selection distributions.
- UK and US together account for approximately 35% of selections, while the bottom half of cultures receive less than 30%.
- UK and US consistently dominate the C0 distribution across models, while Ethiopia and Northern Nigeria are persistently underselected.
- All models exceed the Gini baseline of 0 for perfectly equal selection.
- All six Lorenz curves substantially depart from the diagonal representing perfectly uniform selection.
D. Extended C1/C2 Results
C1 location cues carry genuine cultural signal, while C2 intent directives improve compliance but leave culture-dependent steering differences. Signal lift is stronger for high-resource cultures, whereas underrepresented cultures show greater resistance under maximum steering.
- C2 consistently improves compliance over C1, although no model approaches perfect compliance.
- High-resource cultures show moderate signal lift, while Northern Nigeria and Ethiopia show near-zero lift across models.
- Underrepresented cultures, including Northern Nigeria, Ethiopia, and Assam, have the highest resistance to identity-based steering at C2.
- C2 compliance gains are culture-dependent: UK and US gain less from the directive, while some underrepresented cultures gain more from a lower absolute base.
E. Extended C3 Results
Explicit cultural fact injection barely changes cultural selection distributions, while maximum steering widens the gap between high- and low-resource cultures. C3 evaluations also reveal position-sensitive selection effects that require controlled rotations.
- All C3-to-C0 Jensen-Shannon Divergence values remain below 0.02 bits, indicating less than 2% mean absolute distributional shift.
- C2 increases Statistical Parity Difference for every model relative to C0, widening rather than narrowing the selection gap.
- The C3 cultural selection distribution closely mirrors C0 across cultures and models.
- Primacy bias leads models to overselect the option whose fact appears in slot 1; GPT-5.4 Nano has PPR = 0.42 and Claude Haiku 4.5 has PPR = 0.28.
- Independent option and context rotations create 16 combinations per scenario, separating cultural content effects from fact-position effects.
- The CSC random baseline is approximately 1.6%, because four independent rotations must select the same culture by chance.
F.4. Entropy Analysis
Rotation controls largely remove letter-position preference, but context position still contributes measurable variance in C3 selections. High entropy therefore coexists with primacy bias rather than ruling it out.
- Normalised selection entropy remains high across models and rotations, ranging from 0.92 to 1.00.
- Most models achieve 56–69% Culture Selection Consistency, far above the approximately 1.6% random baseline but below perfect consistency.
- GPT-5.4 Nano has the strongest primacy bias, with PPR = 0.42 and the most pronounced entropy drop.
- Four option-order rotations keep CR, JSD, and CSD estimates stable, indicating that no single option rotation drives the findings.
- Residual PPR increases at some context rotations reflect primacy bias rather than option-order instability.
- After four option-order rotations, letter probabilities remain close to the 0.25 uniform baseline, with residual deviations such as Qwen P(B) = 0.277 and LLaMA P(B) = 0.283.
- GPT-5.4 Nano combines comparatively low default concentration with the highest prior stickiness and JSD, showing that flatter priors need not be more steerable.
- Explicit cultural facts produce JSD values below 0.02 bits, leaving the default selection distribution almost intact.
H.3. Primacy Bias Methodological Implications
Primacy bias complicates interpretation of fact-injection effects by allowing context position to mimic cultural preference. The paper therefore treats context-order control as a methodological requirement for C3-style evaluations.
- Some apparent selection shifts after fact injection may reflect context-position effects rather than genuine content processing.
- Uncontrolled sequential fact injection can confound position effects with cultural preference effects.
- The proposed 4 × 4 rotation design provides a tractable correction for multi-fact injection evaluations.
- Primacy severity varies by model, with PPR ranging from 0.28 to 0.42, so position sensitivity should be characterized independently.
- Higher stickiness for Northern Nigeria, Ethiopia, and Assam aligns with their underrepresentation in the default C0 distribution.