Source-linked AI summary

When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice

Muhammad Salar Khan, Hamza Umer, Hasan Mahmud, Sandra Rothenberg

arXiv:2608.16909v1cs.CYcs.AIcs.CL

TL;DR

Religious bias in LLM-generated financial advice remains underexamined. Using a mixed-methods analysis of simulated advisory outputs, the paper finds that bias operates both structurally and through linguistic framing, with unbiased responses in fewer than one in five cases.

  • Problem

    The paper examines how LLMs handle religion in financial advice, addressing limited evidence about religious bias in this expanding application.

  • Method

    The study applies Critical Algorithm Studies to interpret religious framings in AI-generated financial advisory messages as reflections of sociocultural hierarchies shaping model training.

  • Results

    Religious bias operates through both structural patterns and linguistic framing, with explicit religious framing dominating and unbiased responses occurring in fewer than one in five cases.

  • Takeaways & Limitations

    The findings show that bias manifests in both recommendation substance and framing, highlighting the need to assess religious language as well as financial advice content.

  • Takeaways & Limitations

    The study’s limitations motivate extending research beyond text outputs to multimodal advisory interfaces such as voice.

Abstract

from arXiv · show

Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok) using 432 simulated advisor-client interactions spanning 16 religious identity pairings (Christian, Muslim, Hindu, and non-religious) and three core household financial decisions: stock investment, house purchase, and life insurance. Combining regression and reflexive thematic analyses, we identify structural biases across models and decision contexts and the discursive mechanisms through which they are linguistically enacted. Unbiased advice appeared in only 12-18% of cases. Gemini consistently produced more bias than Grok, while ChatGPT's outputs were statistically comparable to Grok's. Religiously symmetric advisor-client pairings almost always triggered explicit religious framing, and non-religious clients often received advisor-centered religious appeals. Qualitative findings show that bias is linguistically manifested through religious anchoring, uneven cultural signaling, and tone modulation, varying by model and financial scenario. Stock investment prompts produced more financially technical responses, whereas life insurance advice triggered stronger religious language. The study develops a dual-dimensional framework linking structural bias rooted in model training and design with discursive bias expressed through language, advancing understanding of algorithmic bias in LLM-generated financial advice. It also shows that such advice adapts linguistically to identity cues, revealing a managerial dilemma between personalization and neutrality. Finally, it highlights implications for businesses, financial institutions, and regulators seeking to ensure neutrality, cultural sensitivity, and trust in AI-mediated advice.

1. Introduction

The paper examines religious bias in LLM-generated financial advice as both a structural problem rooted in model systems and a discursive problem expressed through framing, tone, and cultural references. It shows that personalization can conflict with neutrality, creating ethical, managerial, and governance challenges for financial advice.

  • Research problem and gap: Religious bias in AI-generated financial guidance remains critically underexplored despite the expanding use of LLMs in financial advice and other financial tasks.The gap is especially important because personalization can introduce bias in morally and culturally sensitive contexts such as religion.
  • Conceptual framework: The paper distinguishes structural bias, rooted in training data and model design, from framing bias, expressed through tone, moral language, and cultural references.The framing-bias concept differs from decision bias because it concerns the style rather than the substance of advice.
  • Contributions and implications: The study contributes a mixed-methods, comparative approach integrating critical algorithm studies, sociolinguistics, and moral economy perspectives across LLMs, decisions, and identity pairings.It frames religious bias as a personalization–neutrality tension and highlights safeguards to keep AI-mediated financial advice culturally sensitive yet unbiased.
  • Empirical findings: 12–18%: unbiased outputs occurred in only 12–18% of cases, while Gemini showed significantly higher bias than Grok and ChatGPT’s bias rates were statistically comparable to Grok’s.In-group advisor–client pairings almost always elicited explicit religious framing, and non-religious clients frequently received advice framed through the advisor’s faith.
  • Discursive mechanisms: Bias was manifested qualitatively through religious framing, uneven cultural signaling, tone modulation, and shifts between moral and technical justifications depending on client identity.These findings motivate examining variation across religious identities, financial decision types, and LLMs, as well as the linguistic mechanisms of bias.

2. Related Literature

Prior research shows that financial AI and LLMs can reproduce structural and discursive biases, while religious bias remains comparatively underexamined. Behavioral-finance research also establishes that religion shapes financial judgment, motivating systematic study of how LLM personalization intersects with neutrality across core financial decisions.

  • Algorithmic bias in finance: Financial AI systems reproduce and amplify inequities across gender, racial, and socioeconomic lines, including when marginalized identities intersect.Prior work links these disparities to historical biases in data and to credit algorithms’ treatment of multiple identities.
  • LLM bias and personalization: LLM bias arises from skewed training data, subjective fine-tuning, and deployment feedback loops, affecting political, gender, racial, and cultural outputs.Unlike numerical predictive systems, LLMs generate personalized textual advice that can alter substantive outcomes and reshape tone or moral language.
  • Religious bias in LLMs: Religious bias in LLMs remains scant and emerging, with studies documenting Abrahamic overreliance, asymmetric treatment of Islam and Hinduism, and Christian dominance in simulated debates.Other research reports religious stereotyping, including associations between Islam and violence or terrorism.
  • Structural and discursive bias: Religious bias operates structurally through data and fine-tuning asymmetries and discursively through communicated moral and cultural meanings in supposedly neutral contexts.This distinction frames religion as a measurable domain of LLM bias comparable to race and gender.
  • Religion and financial decision-making: Religion shapes financial judgment through value framing, ethical orientation, and community norms across investing, life insurance, and housing.The literature links religiosity to saving, risk-taking, ethical evaluations, asset selection, and trust, while shared norms may alienate clients from different faith backgrounds.
  • Research gap and contribution: The study addresses a research gap by examining whether and how LLMs reproduce, modify, or amplify religious framing across stock investment, life insurance, and housing purchase.This extension focuses on how AI-mediated personalization intersects with fiduciary neutrality in applied financial communication.

3. Theoretical Framework

The framework combines Critical Algorithm Studies with sociolinguistic and moral economy perspectives to explain religious bias as both structurally generated and discursively enacted in AI-mediated financial advice. It conceptualizes bias across upstream socio-technical asymmetries and downstream linguistic, moral, and cultural framing.

  • Critical Algorithm Studies: Critical Algorithm Studies treats algorithms as socio-technical systems shaped by cultural norms, institutional priorities, historical power relations, data, and model design.This perspective interprets bias as structural asymmetry rather than an isolated technical error.
  • Discursive framework: The framework extends CAS with a sociolinguistic–moral economy bridge explaining how bias is performed through tone, politeness, cultural signaling, and moral framing.Sociolinguistics addresses identity performance and relational stance, while moral economy explains the values, obligations, and power relations underlying linguistic choices.
  • Moral economy: Uneven religious framing can privilege some moral economies over others and differentially affect financial-advice uptake, trust, and perceived appropriateness.Examples include Islamic finance’s prohibition of interest and Christian stewardship framing when such references are adopted unevenly.
  • Dual-dimensional framework: The resulting dual-dimensional model links structural bias from upstream training, design, and deployment asymmetries with discursive bias enacted downstream through language and moral framing.Structural bias shapes the conditions under which patterns arise, while discursive bias appears in tone, moral framing, and cultural signaling.
  • Research design: The framework informs a comparative, multi-model, multi-scenario mixed-methods design that examines bias prevalence, linguistic mechanisms, and variation across identities, decisions, and models.CAS motivates analysis of output-level socio-technical asymmetries, while the sociolinguistic–moral economy bridge guides qualitative analysis of how bias is expressed and justified.

4. Methodology

The study uses a controlled mixed-methods design to test religious bias in LLM-generated financial advice across identity pairings, household financial scenarios, and three major models. Standardized prompts, repeated executions, and isolation procedures support comparisons of systematic, model-specific, temporal, and regional patterns.

  • Research focus: Religious bias is defined as systematic differences in religious language, concepts, or references that alter recommendations or linguistic and moral tone.The analysis primarily focuses on framing bias rather than decision bias.
  • Financial scenarios: The study examines global-stock investment, house purchase, and life-insurance decisions to compare technical and morally salient household financial contexts.These scenarios represent core household portfolio components and enable scrutiny across differing contextual demands.
  • Experimental design: Replacing advisor and client identity cues with Christian, Muslim, or Hindu labels creates 16 possible advisor-client combinations for each scenario.The prompts retain the baseline wording while modifying only the designated identity portions.
  • Models and data collection: Three researchers independently tested ChatGPT, Gemini, and Grok with 48 prompts per model, producing 144 prompts per researcher and 432 outputs overall.The same standardized prompts were executed across models to compare bias presence, temporal variation, and relative magnitude.
  • Controls and variation: Controls reduced carryover from prior interactions by running the religion-free prompt first, using incognito windows for ChatGPT and Grok, and disabling Gemini app activity.Researchers also varied execution timing and location, including May 2025 runs from the US and Japan and a later US run.

5. Quantitative Study

Across models and financial decisions, religious bias was usually explicit and varied systematically by model, task context, and advisor–client identity pairing. Regression results likewise identify model and interaction structures as important predictors of explicit bias, while financial decision type was generally not significant.

  • Model-level patterns: 15% of ChatGPT, 13% of Gemini, and 18% of Grok outcomes were religiously unbiased, while ChatGPT and Grok showed about 73% and 72% explicit bias, respectively.Implicit bias ranged from 4% to about 12% depending on the LLM model.
  • Decision-level patterns: 18% of stock-investment outcomes were unbiased, compared with 12.5% for life-insurance advice, indicating that bias magnitude varied by task context.The study interprets task context as conditioning whether religious framings become salient.
  • Advisor–client identity patterns: 89% of same-religion religious pairings, excluding baseline–baseline interactions, produced explicit interaction-based religious bias.This occurred in 24 of 27 scenarios, with three exceptions, including Hindu–Hindu stock-investment prompts.
  • Advisor–client identity patterns: 78% of interactions between religious advisors and religiously unspecified clients produced advisor-based bias, including 20 explicit and one implicit case.The pattern suggests that outputs often assume the client shares the advisor’s religion.

6. Qualitative Study

The qualitative study used reflexive thematic analysis to explain how religion, identity, and financial context shape the tone, content, and cultural framing of LLM-generated financial advice. It extends structural quantitative findings by identifying the linguistic mechanisms through which bias is performed.

  • Analytical approach: Reflexive thematic analysis examined how religion, identity, and financial context influenced the tone, content, and cultural framing of financial advice.The analysis complemented regression models by examining linguistic expression, not only bias frequency.
  • Theoretical contribution: The qualitative themes operationalized the study’s two-dimensional framework by linking structural patterns to discursive mechanisms such as moral anchoring, cultural signaling, and tone or politeness.The analysis revealed how religious worldviews were encoded or performed in outputs in ways not visible through quantitative models alone.
  • Analytical approach: 432 outputs were read closely and coded across three interpretive dimensions: tone, content, and phrasing.The coding tracked formal versus casual tone, religious references, culturally specific phrasing, technical terminology, emotional register, and moral or doctrinal content.
  • Analytical approach: AI-assisted coding combined inductive–deductive analysis with human review and adjudication of codes.ChatGPT generated semantic and latent codes through iterative reading, while deductive cues drew on literature about religious framing, politeness, and algorithmic bias.

Theme 1: Religious Framing as Moral Anchor (n = 374)

Religious framing was the dominant theme, especially in advisor–client pairs involving Muslim or Christian identities. Out-group pairings elicited stronger discursive moral anchoring through religious economic language.

  • Religious Framing as Moral Anchor: Religious framing was particularly salient in advisor–client pairs involving Muslim or Christian identities.The theme centered financial advice on religious identity and moral responsibility.
  • Religious Framing as Moral Anchor: Phrases such as “wise stewardship,” “as part of our faith,” and “guided by Sharia” anchored financial responsibility in moral or religious terms.A cited example framed future planning as fulfilling duty to family and community.
  • Religious Framing as Moral Anchor: The framing reflects religious economic norms embedded in LLM training corpora.The passage links religious language in financial advice to norms represented in model training data.
  • Religious Framing as Moral Anchor: Religious framing was more common in out-group pairings, possibly reflecting algorithmic overcompensation or identity signaling.This structural configuration triggered stronger discursive moral anchoring.

Theme 2: Cultural Signaling: Surface Depth, Uneven Spread (n = 232)

LLMs used culturally specific religious language unevenly, with Islamic terminology appearing more frequently and fluently than Hindu equivalents. Christian expressions also appeared in baseline messages, reflecting dominant cultural norms in Anglophone training data.

  • Theme 2: Cultural Signaling: Surface Depth, Uneven Spread: Cultural signaling uses lexical or symbolic references to religious identity without advancing explicit moral or theological reasoning.This framing captures surface cultural cues rather than substantive religious argument.
  • Theme 2: Cultural Signaling: Surface Depth, Uneven Spread: Islamic finance concepts such as Sharia and halal appeared more frequently and fluently than Hindu equivalents such as Grihastha or Lakshmi.Examples include references to Sharia-compliant investments and the Grihastha stage of life.
  • Theme 2: Cultural Signaling: Surface Depth, Uneven Spread: Uneven training-data exposure, particularly to Islamic finance discourse, yielded uneven discursive cultural signaling across traditions.The finding links structural exposure in training data to differences in how cultural cues are expressed.
  • Theme 2: Cultural Signaling: Surface Depth, Uneven Spread: Christian expressions appeared frequently even in baseline messages, reflecting the dominance of Christian cultural norms in Anglophone training data.The passage attributes this pattern to the cultural composition of the training data.

Theme 3: Technical Framing and Rationalized Advice (n = 320) · Theme 4: Communication Style: Deference vs. Assertion (n = 242)

The paper finds that LLMs shift between technical, objective financial reasoning and moral or emotional framing depending on the financial context and identity pairing. It also shows systematic tone differences, with deference directed toward Hindu and Muslim clients and greater assertion toward Christian and baseline clients.

  • Theme 3: Technical Framing and Rationalized Advice (n = 320): Technical Framing emphasized objective financial logic, especially in stock investment scenarios or when religion was absent.Examples included diversification, risk mitigation, and tax planning.
  • Theme 3: Technical Framing and Rationalized Advice (n = 320): “Diversifying internationally can help mitigate country-specific risk” illustrates the technical framing used in financial advice.
  • Theme 3: Technical Framing and Rationalized Advice (n = 320): Technical Framing occurred more often in out-group interactions, where LLMs appeared to avoid culturally specific commitments.
  • Theme 3: Technical Framing and Rationalized Advice (n = 320): Technical Framing dropped sharply in insurance-related outputs, where moral and emotional framings were privileged in family-protection contexts.
  • Theme 4: Communication Style: Deference vs. Assertion (n = 242): Tone varied systematically across identities: messages to Hindu and Muslim clients were more formal or deferential, while messages to Christian or baseline clients were more assertive.The contrasting formulations included “you may wish to consider…” and “you should…”.
  • Theme 4: Communication Style: Deference vs. Assertion (n = 242): These patterns suggest that LLMs learn tone differentiation tied to cultural scripts and that structural identity pairing organizes deference versus assertion.

Theme 5: Neutral

Neutral outputs were exceptionally rare: only three messages, all from Gemini, contained no explicit religious, cultural, or technical framing. The broader findings show that structural patterns by pairing, model, and scenario shape the discursive mechanisms through which bias is expressed.

  • Theme 5: Neutral: Only three messages, all from Gemini, contained no explicit religious, cultural, or technical framing; these outputs were minimal and professional.Their rarity underscores the extent to which LLM outputs are structured through learned patterns.
  • Structural distributions: Out-group pairings triggered more of every theme, especially Religious Framing and Communication Style, producing more overtly polite and morally framed responses.This amplification may reflect performative caution and over-signaled respect when models navigate perceived cultural difference.
  • Model differences: Grok produced the highest proportion of Communication Style codes (n = 134), Gemini exhibited the strongest Cultural Signaling (n = 104), and ChatGPT showed a balanced thematic distribution.ChatGPT’s comparable frequencies were Religious Framing (n = 120), Technical Framing (n = 104), and Communication Style (n = 107).
  • Discursive mechanisms: In-group interactions tended to feature direct religious anchoring, whereas out-group interactions more often carried deferential tone and cultural signaling.These forms indicate a shift toward cautious accommodation across identity pairings.
  • Scenario effects: Life insurance advice was the most moralized and religiously framed, stock advice remained primarily technical, and house-related scenarios occupied an intermediate mix.Together, quantitative and qualitative results show structural regularities predicting when bias occurs and discourse specifying how it is realized.

7. Discussion: findings and implications

The study identifies religious bias in LLM-generated financial advice as both structural and discursive, with explicit religious framing in most of 432 outputs and unbiased responses in only 12–18% of cases. Bias varied across models, scenarios, and identity cues, creating a managerial tension between personalization and neutrality in high-stakes financial advising.

  • Main findings: 12–18% of 432 outputs were unbiased, while explicit religious framing dominated, providing systematic evidence of religious bias in LLM-generated financial advice.The findings extend algorithmic bias research into the high-stakes financial advisory domain, where neutrality is expected.
  • Main findings: Gemini produced higher bias than Grok, while ChatGPT’s outputs were comparable to Grok’s.Regression results also showed higher bias for Gemini relative to Grok and no significant difference for ChatGPT.
  • Identity cues and discourse: In-group pairings almost invariably elicited explicit religious framing, whereas neutral advisor–client prompts almost never produced religious content.Models adapted to identity cues through moral anchoring, cultural signaling, and religious nudging, including advisor-centered appeals to neutral clients.
  • Scenario effects: Stock investment advice was less biased and more technical, whereas life insurance advice elicited stronger religious and cultural framing.Life insurance’s associations with mortality, family protection, and intergenerational responsibility likely made religious and moral language more readily invoked.
  • Discursive mechanisms: Religious framing operated through moral anchoring, uneven cultural signaling, and tone modulation, including stronger Islamic-finance fluency and more deferential phrasing toward Muslim and Hindu clients.The same financial recommendation could be linguistically reframed according to perceived identity through shifts in tone, morality, or cultural signaling.
  • Implications: The findings frame AI advising as a tension between personalization and neutrality, requiring personalization to enhance inclusivity without compromising fairness, objectivity, or trust.Excessive religious personalization can eliminate secular options, undermine perceived objectivity and fiduciary integrity, and create an adaptive-versus-over-personalized behavioral continuum.

8. Limitations and Future Research

The study’s conclusions are bounded by its models, religious coverage, interpretive coding, and lack of downstream outcome measures. Future research should broaden identities and model comparisons, examine multimodal advice, and connect bias to client behavior and consequences.

  • Limitations: Findings are bounded by the examined models and timeframes, as evolving training data and fine-tuning strategies may change bias patterns.The study compared three major LLMs across multiple advisor–client religious pairings and decision scenarios.
  • Limitations: The analysis covered Christianity, Islam, and Hinduism, providing breadth but omitting other significant traditions and limiting generalizability.These traditions were selected for their global prevalence.
  • Limitations: Interpretive coding judgments remain a limitation despite independent replications and cross-checking; alternative frameworks or more coders could add nuance and reliability.The study identified religious-bias presence and expression but did not measure how clients interpret, trust, or act on biased advice.
  • Limitations: The study documents religious framing’s prevalence and expression without evaluating whether it benefits or harms outcomes, leaving appropriateness, accuracy, and effectiveness unresolved.Downstream behavioral consequences remain unmeasured.
  • Future Research: Future research should expand religious identities, compare free and paid tiers across model generations, study multimodal interfaces, and link bias to trust, decision quality, and recommendation uptake.These directions would assess the economic, ethical, and managerial consequences of bias in real-world advisory interactions.
  • Future Research: Linking biased advice to behavioral outcomes could move research from detecting bias toward assessing its real-world economic, ethical, and managerial consequences.Candidate outcomes include client trust, decision quality, and recommendation uptake.

9. Conclusion

The study establishes religious bias as a systematic feature of LLM-generated financial advice, operating through both recommendation substance and linguistic framing. It highlights the managerial and regulatory challenge of preserving personalization while ensuring fairness, neutrality, cultural respect, and fiduciary trust.

  • Core findings: 432 outputs showed explicit religious framing dominated, with unbiased responses occurring in fewer than one in five cases.Bias appeared in both the substance of recommendations and their framing.
  • Model and scenario differences: Gemini exhibited consistently higher levels of bias than Grok, while ChatGPT’s outputs were broadly comparable to Grok’s.Bias was most pronounced in in-group advisor–client interactions and descriptively in life insurance advice.
  • Dual-dimensional bias: Bias operated structurally through model- and scenario-level regularities and discursively through moral anchoring, cultural signaling, and tone modulation.The dual-dimensional account frames algorithmic advice as both a technical and communicative act, making fairness depend on how advice is framed as well as what it recommends.
  • Managerial and regulatory implications: AI advisory systems must balance personalization that enhances engagement with neutrality that preserves fairness and fiduciary integrity.Financial firms should ensure religious and cultural neutrality in client-facing applications, while regulators should include religion as a fairness criterion in AI audits.
  • Limitations and future research: The study documents the prevalence and expression of religious framing but does not assess whether it benefits clients when aligned with their preferences.Future research should examine how religious or moral framings affect user trust and decision quality, distinguishing meaningful personalization from inappropriate over-personalization.
Loading 2608.16909v1…