Source-linked AI summary

Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation

Tanay Kumar, Shreya Gautam, Aman Chadha, Vinija Jain, Francesco Pierri

arXiv:2604.23600v2cs.CL

TL;DR

Existing evidence is limited on how personality conditioning shapes gender bias in persona-driven LLM narratives, especially across English and Hindi. This study systematically varies persona gender, occupation, and personality in 23,400 generated artifacts, finding that Dark Triad traits amplify gender-stereotypical representations while effects vary by language and model.

  • Problem

    Evidence remains limited on how personality-conditioned narrative generation interacts with gender bias, particularly in Hindi and across linguistic contexts.

  • Method

    The study varies persona gender, occupation, and HEXACO or Dark Triad traits across 23,400 English and Hindi artifacts from six LLMs, scoring bias with sentence-level stereotype-centroid embeddings.

  • Results

    Dark Triad traits consistently amplify gender-stereotypical representations, while HEXACO effects are mixed; personality effects sometimes exceed explicit gender-label effects across models, occupations, and languages.

  • Takeaways & Limitations

    Personality prompts are fairness-critical controls in multilingual persona-driven systems, and mitigation strategies may not transfer uniformly across languages.

  • Takeaways & Limitations

    The bias estimates depend on the chosen embedding encoder and curated stereotype lexicons, with English-to-Hindi translation potentially missing Hindi-specific expressions.

Abstract

from arXiv · show

Large Language Models (LLMs) are increasingly deployed in persona-driven applications such as education, customer service, and social platforms, where models are prompted to adopt specific personas when interacting with users. While persona conditioning can improve user experience and engagement, it also raises concerns about how personality cues may interact with gender biases and stereotypes. In this work, we present a controlled study of persona-conditioned story generation in English and Hindi, where each story portrays a working professional in India producing context-specific artifacts (e.g., lesson plans, reports, letters) under systematically varied persona gender, occupational role, and personality traits from the HEXACO and Dark Triad frameworks. Across 23,400 generated stories from six state-of-the-art LLMs, we find that personality traits are significantly associated with both the magnitude and direction of gender bias. In particular, Dark Triad personality traits are consistently associated with higher gender-stereotypical representations compared to socially desirable HEXACO traits, though these associations vary across models and languages. Our findings demonstrate that gender bias in LLMs is not static but context-dependent. This suggests that persona-conditioned systems used in real-world applications may introduce uneven representational harms, reinforcing gender stereotypes in generated educational, professional, or social content.

1 Introduction

The study examines how personality conditioning shapes gender bias in persona-driven LLM narratives across English and Hindi. It finds that bias is context-dependent, with Dark Triad traits amplifying and prosocial traits attenuating stereotypes across models, occupations, and languages.

  • Study design: The study generates 23, 400 persona-conditioned artifacts across English and Hindi by varying persona gender, occupational role, and HEXACO or Dark Triad traits across six language models.This controlled design targets personality–gender interactions in LLM-generated narratives.
  • Key findings: Gender bias in LLM-generated narratives is strongly context-dependent rather than static.The study frames personality, gender, occupation, model, and language as interacting sources of variation.
  • Key findings: Dark Triad traits consistently amplify gender-stereotypical representations, while prosocial traits tend to attenuate them.These patterns align with prior psychological evidence that antagonistic traits increase stereotyping and prosocial traits suppress it.
  • Key findings: Personality effects persist across models, occupations, and languages, and often exceed the influence of the explicit gender label itself.Hindi also shows distinct interaction patterns linked to grammatical gender marking.
  • Contributions: The paper contributes a multilingual personality-conditioned artifact corpus, a sentence-level centroid-based gender-bias metric, and cross-language evidence on systematic bias amplification and attenuation.The analysis spans occupations and model architectures, with implications for evaluating bias in persona-driven applications.

2 RelatedWork

Prior research documents gender stereotypes in language models, including narrative and multilingual settings, while showing that persona and personality conditioning shape model behavior. This work addresses the unresolved interaction between personality and gender bias in occupationally grounded English and Hindi narratives using story-level measurement.

  • Gender Bias in Language Models: Language models associate men and women with different occupations, traits, and social roles, with bias persisting across representations and downstream tasks.Prior studies used embedding-based analyses and benchmarks including WEAT/SEAT, WinoBias, StereoSet, and CrowS-Pairs.
  • Multilingual and Hindi Bias: Hindi gender bias interacts with grammatical gender, occupation, and social hierarchy, while English-centric measures do not transfer directly to gender-marking languages.Hindi and other Indic languages remain comparatively underexplored despite their large speaker populations.
  • Persona and Personality Conditioning: Persona prompting shapes tone, style, and content, while personality conditioning can modulate harmful behaviors such as toxicity and bias.LLMs also exhibit stable patterns along Big Five, HEXACO, and Dark Triad dimensions, with some traits acting as amplifiers and others as attenuators.
  • Research Gap and Contribution: The unresolved problem is how personality traits modulate gender bias through interaction effects in multilingual narrative generation grounded in occupational contexts.The study systematically varies gender, occupation, and HEXACO and Dark Triad traits in English and Hindi, examining personality–gender interactions with sentence-based, story-level bias measurement.

3 Methods

The study generates multilingual, persona-conditioned artifacts across systematically varied gender, occupation, personality, language, and model dimensions, then quantifies gender-stereotypical alignment using stereotype centroids. Its design includes explicit no-personality baselines and produces 23,400 artifacts across six models.

  • Experimental pipeline: The pipeline generates persona-conditioned artifacts, constructs male and female stereotype centroids from curated word lists, and computes story-level bias from cosine similarity to those centroids.Generated artifacts are embedded in a multilingual sentence space before scoring.
  • Experimental factors: Personas vary by three gender conditions, 50 Indian-context occupations, two languages, and nine HEXACO or Dark Triad personality traits with high- and low-score descriptions.The gender-neutral condition provides a no-explicit-gender baseline, while occupations are split into 25 predominantly male- and 25 predominantly female-stereotyped categories.
  • Corpus and models: The six-model set spans GPT-5 nano, Llama-3.3-70B-Instruct, Gemma-3-1b-it, Deepseek-R1, Mixtral-8x7B Instruct, and Falcon-mamba-7b-instruct.These models represent LLM, SLM, LRM, MoE, and SSM families.
  • Corpus and models: 3,900 artifacts per model and 23,400 artifacts overall are generated across six models, two languages, 50 occupations, 18 personality conditions, and 150 baseline prompts.Each artifact is a single paragraph with 6–8 sentences.
  • Baseline design: Explicit baselines omit personality specifications so personality-conditioned artifacts can be compared with bias associated with occupation and gender alone.The no-personality baseline is denoted by p = ∅, with positive bias scores indicating male-stereotypical alignment and negative scores indicating net female-stereotypical alignment.
  • Stereotype representation: English stereotype lists contain approximately 200 −210 items and are translated and refined in Hindi to keep the two lexicons parallel for cross-lingual comparison.Lists cover nouns, verbs, and adjectives, while part-of-speech-specific templates convert words into contextualized sentences.

4 Results

Personality conditioning systematically changes gender bias in LLM-generated stories, with Dark Triad traits generally amplifying male-stereotypical alignment. These effects interact with persona gender, language, and model family, making bias context-dependent.

  • Overall personality effects: Without personality conditioning, stories show consistent male-stereotypical bias in both languages, with Hindi generally exhibiting stronger male-leaning skew across five of six models.Grammatical gender marking and cultural context may amplify stereotypical alignment.
  • Overall personality effects: Personality conditioning increases male-leaning artifacts relative to the no-personality baseline across all models and is associated with both the direction and magnitude of gender bias.Dark Triad traits consistently produce positive bias shifts, while HEXACO effects are heterogeneous.
  • Gender and personality contributions: Personality coefficients can exceed gender coefficients in magnitude, while personality effects remain directionally consistent across persona genders but are systematically larger for male personas.Machiavellianism and Psychopathy especially compound male-gender cues and amplify stereotypical alignment.
  • Language differences: Hindi has a stronger baseline male-stereotypical skew, but personality-driven modulation is more pronounced in English, where Dark Triad amplification and attenuation under Openness and Emotionality are stronger.Hindi follows the same overall trait-direction pattern with consistently smaller magnitudes.
  • Model-family variation: All models show more female-stereotypical alignment than the GPT-5 nano reference, while Dark Triad amplification and prosocial-trait attenuation remain consistent across architectures.Gemma3-1b-it and Falcon-Mamba-7b show the largest negative shifts relative to the reference.

5 DiscussionandConclusion

The study shows that persona-conditioned gender bias in English and Hindi is context-dependent, with personality traits shaping both its magnitude and direction. These findings make persona specifications fairness-critical in practical LLM applications and motivate broader evaluations across social categories and evolving interactions.

  • Discussion and Conclusion: Across 23,400 artifacts from six LLMs, gender bias varied with persona gender, occupation, language, and HEXACO or Dark Triad traits.The study used a sentence-level centroid-based bias metric validated by human annotation.
  • Discussion and Conclusion: Dark Triad traits consistently amplified male-stereotypical alignment, whereas HEXACO traits had mixed effects across contexts.In some cases, personality effects exceeded those of explicit gender labels.
  • Discussion and Conclusion: Persona-driven bias should be understood as an emergent interaction among model, prompt design, and linguistic context rather than as a static model artifact.Personality cues can systematically amplify or attenuate gender-stereotypical language in educational support, workplace writing assistance, and customer service.
  • Discussion and Conclusion: Future evaluations should examine social categories beyond gender, including caste, religion, and disability, alongside multiturn interactions with evolving personas.Such extensions could support richer evaluations and clearer guidance for safer persona design.

Limitations

The study’s generalizability is limited by its occupation-grounded Indian artifact-generation setting, Western-derived personality frameworks, measurement choices, small annotation sample, and simulated deployment prompts. Its metrics may miss translation-specific, distributional, subtle, or intersectional biases, while real-world consequences and cross-cultural applicability remain untested.

  • Scope and constructs: The study covers one occupation-grounded artifact-generation task family in India and a fixed set of stereotyped occupations, limiting generalization to other domains or interaction styles.HEXACO and Dark Triad frameworks originate in Western personality psychology and may not fully capture personality constructs across Indian cultural contexts.
  • Measurement: Bias estimates depend on the chosen encoder, manually curated stereotype lexicons, and English-to-Hindi translations that may introduce artifacts or omit Hindi-specific expressions.The lexicons were refined with native speakers but originally compiled in English.
  • Measurement: Sentence-level maximum aggregation captures the most salient stereotypical statement but discards distributional tendencies, while centroid alignment may miss subtle or intersectional bias.Alternative strategies, including mean scores or biased-sentence proportions, may yield different conclusions.
  • Human validation: 3 annotators evaluated 100 story pairs, a limited sample relative to 23,400 artifacts, using pairwise neutral-baseline comparisons rather than absolute severity or demographic analyses.The study reports substantial inter-annotator agreement, but annotator demographic variation could influence stereotype perception.
  • External validity: Controlled prompts may not reflect real-world chatbot usage, where personality cues are less explicit; live user consequences and patterns across other languages or cultures were not measured.Users may induce similar effects through subtler cues than those used in the prompts.

EthicalConsiderations

The study acknowledges ethical risks from using gender-stereotypical language and Dark Triad prompts, while describing safeguards for dataset use and human annotation. It also clarifies the cultural grounding and limitations of its occupational scenarios and personality frameworks.

  • Risks of Stereotype Data: The dataset’s stereotype lexicons and persona-conditioned narratives may reinforce the gender stereotypes they are intended to expose.The authors contextualize examples within an explicit bias-measurement framework.
  • Risks of Stereotype Data: Dark Triad prompts could be misused to generate harmful, manipulative, or stereotyped content at scale.Researchers and practitioners are advised to implement appropriate safeguards when using the dataset or prompt templates.
  • Human Annotation: Three annotators evaluated narratives containing gender-stereotypical language after briefing, with withdrawal permitted and fair compensation provided.They were informed that the content was machine-generated and did not reflect the authors’ views.
  • Cultural Scope: Occupational scenarios reflect documented gender associations in India without endorsing them, while HEXACO and Dark Triad frameworks may not fully represent personality across cultures.The scenarios were designed to create ecologically valid test conditions in an Indian socio-cultural context.
  • Data and Model Use: Experiments used publicly available or API-accessible language models and collected no private user data.The study states that private user data was not collected or used at any stage.

A Appendix

The appendix provides supplementary material, including additional examples and implementation details, to support understanding of the paper’s concepts.

  • The appendix supplements the paper with additional examples and implementation details that bolster readers’ understanding of its concepts.

A.1 PersonalityTraitsDescriptions

The study characterizes personality using the HEXACO and Dark Triad frameworks, which capture broad individual differences relevant to social behavior and interpersonal dynamics. HEXACO emphasizes six personality dimensions, whereas Dark Triad focuses on socially aversive traits.

  • HEXACO Framework: The HEXACO framework captures six broad dimensions: Honesty–Humility, Emotionality, Extraversion, Agreeableness, Conscientiousness, and Openness to Experience.These dimensions represent prosocial tendencies, emotional responses, and cognitive styles.
  • HEXACO Framework: HEXACO dimensions represent prosocial tendencies, emotional responses, and cognitive styles relevant to interpersonal dynamics.The framework is used to characterize individual differences in personality traits relevant to social behavior.
  • Dark Triad Framework: The Dark Triad framework focuses on three socially aversive personality traits.It provides a contrasting framework for characterizing personality differences in social behavior and interpersonal dynamics.

A.1.1 SD3(ShortDarkTriad)Validation · SD3TestItemList

The SD3 validation used test items to assess whether prompted personality descriptions induced the intended Dark Triad traits in model outputs. Reverse-scored items were explicitly marked.

  • A.1.1 SD3(ShortDarkTriad)Validation: SD3 items evaluated whether prompted personality descriptions successfully induced the intended Dark Triad traits.The validation focused on model outputs.
  • SD3TestItemList: The evaluation used the Short Dark Triad (SD3) item list.These items formed the basis of the validation procedure.
  • A.1.1 SD3(ShortDarkTriad)Validation: Prompted personality descriptions were assessed for successful trait induction.The criterion was whether the descriptions produced the intended personality traits in model outputs.
  • SD3TestItemList: Model outputs were the objects evaluated by the SD3 validation items.The items tested the effects of prompted personality descriptions on outputs.
  • A.1.1 SD3(ShortDarkTriad)Validation: The intended traits were Dark Triad traits.The validation did not describe a different personality framework in this item list.
  • SD3TestItemList: Items marked (R) were reverse-scored.The passage explicitly identifies the scoring treatment for these items.

Machiavellianism … A.8 ExtendedQuantitativeResults

The study operationalizes Dark Triad and HEXACO personas, gender-stereotyped occupations, and artifact–scenario tasks to examine how personality conditioning shapes gendered narratives across models and languages. Results show strong, robust, and human-recognized gender associations, with personality effects expressed through distinct gendered behavioral scripts.

  • Machiavellianism: Dark Triad prompts operationalize Machiavellianism, narcissism, and psychopathy through manipulative, self-aggrandizing, and impulsive or harmful behavioral tendencies.High- and low-score prompts were validated by comparing aggregate SD3 inventory scores across conditions.
  • A.2 OccupationalPromptDesign: The occupational design uses male- and female-stereotyped Indian roles paired with standardized artifacts and scenarios to isolate personality effects beyond occupational priors.The labels represent stereotypes for measurement rather than ground truth, and concrete tasks provide contextual grounding.
  • A.3 ModelSelectionRationale: Six contemporary models spanning transformer, small-language, reasoning-oriented, and mixture-of-experts paradigms were selected to test whether findings generalize across architectures.The models are GPT-5 nano, Llama-3.3-70B-Instruct, Gemma-3-1B-it, DeepSeek-R1, Mixtral-8x7B-Instruct, and Falcon-Mamba-7B-Instruct.
  • ChoiceofSentenceEmbeddingModel: IndicSBERT was used for bias computation because it better captured Indian-language semantic relationships and produced clearer, more stable bias patterns than LaBSE.Preliminary LaBSE experiments showed less distinct separation between male- and female-aligned similarity distributions.
  • A.4 PersonalityConditionedGenerated Narratives: Identical high-score psychopathy prompts produced covert, caregiving-mediated harm in female narratives but overt self-indulgence, rule-breaking, and dominance in male narratives across English and Hindi.The examples show personality conditioning being routed through gender-specific behavioral scripts rather than expressed gender-neutrally.
  • FemaleDomesticWorkerArtifact: Other narratives show narcissism as professional grandiosity or, in one police-officer case, emotional vulnerability, while emotionality appears as male ambition and female caution or anxiety.Hindi examples retain grammatical gender marking, and domestic-worker, nurse, seamstress, HR, engineer, and firefighter narratives vary in trait expression.
  • A.5 GenderedStereotypicalWords: Cohen’s d was 0.9302 in Hindi and 0.9447 in English, with both exceeding the large-effect threshold of d > 0.80 for gender-associated semantic patterns.Repeated-run variation did not materially affect the reported results, supporting stability across stochastic generations.
  • A.6 RobustnessCheck: AggregationStrategy: Max-abs and mean agreed on bias direction in 87.2% of stories (Cohen’s κ = 0.738), while max-abs and top-3 mean agreed in 87.7% (κ = 0.749).Human annotator agreement was substantial across languages and subgroups; Hindi rates were 0.688 versus 0.660, and personality conditioning was judged amplifying stereotyping at 72.0% in Hindi versus 66.0% in English.
Loading 2604.23600v2…