Source-linked AI summary

One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization

Franziska Weeber, Vera Neplenbroek, Jan Batzner, Sebastian Padó

arXiv:2601.18572v3cs.CL

TL;DR

Existing persona-based evaluations often rely on a single sociodemographic cue, despite differences in prompt sensitivity and real-world cue availability. The paper compares six cues across seven LLMs and four task types, finding that highly correlated cues can still yield different persona disparities. It therefore recommends evaluating multiple cues, prompt variations, and tasks, while noting scope limits around intersectionality, English-only settings, and proxy names.

  • Problem

    Prior work commonly uses one persona cue, leaving it unclear whether cue variation changes measured bias and personalization findings.

  • Method

    The paper benchmarks six persona cues across ten personas, three sociodemographic variables, seven LLMs, and multiple writing and advice tasks.

  • Results

    Persona cues are highly correlated overall, but produce distinct disparities across personas; explicit user-prompt mentions yield 20/24 significant differences versus 1/24 for system-prompt names.

  • Takeaways & Limitations

    Robust claims about personalization bias should examine multiple persona cues, prompt variations, and evaluation tasks rather than rely on one cue.

  • Takeaways & Limitations

    The study tests only isolated sociodemographic variables, five wording and value variations, and English examples, making intersectional and broader evaluations computationally difficult.

Abstract

from arXiv · show

Personalization of LLMs by sociodemographic subgroup often improves user experience, but can also introduce or amplify biases and unfair outcomes across groups. Prior work has employed so-called personas, sociodemographic user attributes conveyed to a model, to study bias in LLMs by relying on a single cue to prompt a persona, such as user names or explicit attribute mentions. This disregards LLM sensitivity to prompt variation and the rarity of some cues in real interactions (external validity). We compare six commonly used persona cues across seven open and proprietary LLMs on four writing and advice tasks. While cues are overall highly correlated, they produce substantial variance in responses across personas that can change findings on persona-induced differences and bias. We therefore caution against claims based on single persona cues, especially when they are overly explicit and have low external validity.

1 Introduction

LLM personalization by sociodemographic subgroup can improve helpfulness but can also produce unequal responses. This paper examines whether bias findings remain robust when persona cues, personas, tasks, and models vary.

  • Personalized LLM responses can differ across sociodemographic groups even for high-stakes questions where the persona should not matter.Prior work reports disparities in college recommendations, hiring decisions, and health advice.
  • Prior studies typically introduce personas with a single cue, such as conversation history, names, or explicit mentions.These cues differ in how naturally they occur in real user–LLM interactions.
  • Because LLMs are sensitive to prompt formulation, it is unclear whether different persona cues produce the same personalization and bias conclusions.The paper frames this uncertainty across personas, cues, evaluation tasks, and models.
  • The study compares six persona cues across multiple personas, datasets, and LLMs using a flexible evaluation benchmark.Its contributions include analyzing how cue choice affects bias findings and measuring correlations across cues and models.
  • Average responses are strongly correlated across cues and LLMs, but individual cues produce distinct disparities across personas.The paper reports stronger personalization bias for highly explicit, potentially unnatural cues than for more implicit cues.

2 Related Work

Prior personalization research uses personas and sociodemographic cues, but the external validity and cross-cue consistency of these evaluations remain underexplored. This paper addresses how cue choice affects measured LLM differences across personas.

  • Persona research has progressed from qualitative user representations to computationally generated personas, but accurate representation of target populations is rarely evaluated.This motivates attention to how personas are operationalized in personalization studies.
  • External validity concerns whether findings generalize beyond study conditions to real-world contexts, and weakly operationalized personas can limit that generalization.The related work recommends comparing natural operationalizations and examining their correlations.
  • Prompt formulation can affect stereotyping and alignment with human survey responses when models simulate users from sociodemographic groups.Prior work also finds that country-perspective prompts can yield population-like responses and that explicit indicators can shift social judgments.
  • LLM bias and personalization evaluations use indirect cues such as names, conversation histories, dialects, and explicit profiles, but their effects have rarely been compared at scale.Existing studies include small comparisons of synthetic profiles with real-world users.
  • The paper studies how cue choice changes answers across personas and which cues best reflect real-world LLM use.It also relates its focus to concurrent work finding stronger effects for dialect cues and explicit mentions than for names or conversation histories.

3 Methodology

The benchmark varies personas, cue formats, evaluation tasks, and models to measure how sociodemographic cues affect LLM outputs. It combines implicit and explicit cues with writing, advice, misinformation, and judgment tasks.

  • The benchmark is flexible, allowing researchers to add, remove, or modify personas, cues, datasets, and models.Figure 2 presents the benchmark’s four dimensions and experimental components.
  • Personas: Ten personas cover gender, race/ethnicity, and age, with each persona defined by only one sociodemographic marker.Age is evaluated in three bins: 18-34, 35-54, and 55+.
  • Persona cues: Six cues represent different external-validity levels: system- and user-prompt names, system- and user-prompt explicit mentions, and human- or LLM-generated conversation histories.Human-written conversation histories are treated as most externally valid because they naturally occur after the first conversation turn.
  • Persona cues: Names are selected from North Carolina voter-registration data, while explicit mentions insert demographic attributes through prompt templates.Names are treated as demographic proxies when 90% of people in the data share the persona-defining attribute.
  • Evaluation tasks: The evaluation covers writing and advice requests plus medical misinformation and judgment tasks, including SBB, MMMD, AITA, and IssueBench.Persona cues precede the question, and IssueBench evaluates political writing assistance.

4 Experimental Setup

The experiment combines questions with persona cues and tasks, samples multiple responses, and evaluates output-based metrics across models. It then measures cue similarity and persona disparities with correlation and significance tests.

  • Sampling: Each question is paired with six persona cues or no personalization, producing three sampled responses per prompt at temperature 1.0.The full experiment yields 10,530,786 model responses.
  • Evaluation: Closed-ended tasks exclude invalid responses, which average 0.98% of outputs.The evaluation retains responses containing one of the two permitted answer options.
  • Metrics: Task metrics include accuracy for AITA and MMMD, mean numerical answers for selected SBB and IB outputs, and additional subset-specific measures.For AITA, accuracy reflects agreement with the Reddit verdict; for MMMD, it reflects agreement with the ground-truth answer.
  • Cue comparison: Spearman correlations assess similarities between persona cues or models after aggregating sampled responses and cue variations.The correlation analysis compares result metrics rather than raw generated text.
  • Statistical analysis: Persona disparities are evaluated with per-persona means and standard deviations, one-way ANOVA, and Tukey-Kramer post-hoc tests.The tests determine whether result metrics differ across personas and identify which groups differ when they do.

5 Results

Across models, persona cues are strongly correlated overall, but their effects vary substantially by evaluation task, cue, and sociodemographic attribute. These differences can change observed disparities, making diverse cues and tasks important for evaluating personalization bias.

  • Correlation Between Persona Cues: Persona-cue correlations are very strong overall (.91 ≤ρ ≤ .96), with names and explicit mentions more closely aligned than conversation-history cues.All correlations are significantly different from 1.
  • Correlation Between Persona Cues: Model correlations remain substantial (.68 ≤ρ ≤ .88) but are lower than persona-cue correlations, with stronger similarity within model families.The analysis reports no clear model-size effect.
  • Disparities in Results across Personas: Persona cues produce different disparities: explicit user-prompt mentions yield significant persona differences in 20/24 combinations, versus 1/24 for system-prompt names.Human-written conversation histories yield significant differences in 9 combinations, and other cues usually disagree with those patterns.
  • Disparities in Results across Personas: Evaluation tasks show high variance, with AITA producing significant differences in 11/18 combinations, SBB medical in 10/18, and MMMD in 6/18.The results also vary by question type and evaluation format.
  • Disparities in Results across Personas: Age, gender, and race/ethnicity disparities depend on the cue: older personas received up to $5,000 higher salary recommendations, while non-binary personas had over 5 percentage points lower AITA accuracy than male personas.Race/ethnicity effects also reversed between conversation-history and explicit-mention cues on AITA and IB.

6 Conclusion

The study finds that persona cues are highly correlated overall but can produce different persona disparities and conclusions about bias. It recommends evaluating multiple cues, prompt variations, and tasks rather than relying on one cue.

  • Although persona cues are highly correlated, they produce distinct disparities across personas and can change conclusions about personalization and bias.
  • Human-written conversation histories can yield different statistical conclusions across personas than names or explicit mentions, especially explicit user-prompt cues.
  • Highly explicit cues with low external validity may overestimate personalization effects relative to an uncued baseline.
  • Robust personalization-bias claims require multiple persona cues, prompt variations, and evaluation tasks because no single cue captures the complexity of LLM interactions.

Limitations

The study is limited by its restricted persona values and variables, isolated rather than intersectional analysis, English-only cue and value variations, and constrained task coverage. It also uses potentially unrepresentative name proxies, assumes persona-invariant evaluation data in some cases, and cannot control all confounding factors.

  • The study tests limited, largely U.S.-centered values across only three sociodemographic variables, omitting other relevant personalization factors.
  • Personas are evaluated in isolation, so potential intersectional effects are not examined; broader combinations are computationally costly.
  • Only five wording variations, five values per persona, and English-language data are included, limiting coverage of possible cue and persona variations.
  • Names come from one North Carolina voter file and may be culturally ambiguous, selection-biased, or unreliable proxies for gender, especially for unisex names.
  • The evaluation includes only one open-ended task with few examples because LLM-as-a-judge evaluation is complex.
  • The study assumes evaluation datasets should not differ across personas, although that assumption may fail for domains such as medical advice.
  • Uncontrolled factors include socioeconomic status, dialect, and conversation-history content, while conflicting cues may underestimate personalization differences.

Ethical Considerations

The study uses public data and controlled persona construction to evaluate sociodemographic effects without human participants or new data collection. It controls non-target demographic variables and varies cue wording to improve comparability.

  • The study uses publicly available, research-licensed data and involves no human participants or new data collection.
  • Names are selected from voter-registration data using a 90% threshold for gender–ethnicity associations, with unisex names for non-binary personas.
  • Five names are sampled for each demographic attribute while controlling the other two demographic variables.
  • The evaluation uses five paraphrases for name and explicit-mention cues in system and user prompts to account for wording variation.
  • When varying one sociodemographic variable, the study keeps the other two constant, although non-binary reference combinations are limited by PRISM data availability.
  • The evaluation includes balanced closed-ended datasets with 500 examples each and an open-ended IssueBench sample of 166 examples.

C Experimental Details

The experiments compare seven instruction-tuned models across six persona cues, multiple datasets, personas, and response samples. The resulting evaluation contains 10,530,786 model responses, with a low average invalid-response rate.

  • Seven instruction-tuned models are evaluated, including six open-source models from the Llama, Gemma, and Qwen families and ChatGPT.
  • The setup produces 10,530,786 model responses across models, questions, persona cues, sampled responses, personas, and cue indicators.
  • The experiments use gpt-4o-mini-2024-07-18 at temperature 1 and single GPUs selected according to open-source model size.
  • Generation takes from a few hours for some dataset–cue combinations to up to 4 days for IssueBench with writing-style cues.
  • The average invalid-response rate is 0.98%, ranging from 0.08% on SBB to 3.49% on IB.

E IssueBench Evaluation

IssueBench converts open-ended model responses into stance scores using a separate LLM judge and a five-level scale plus refusal. The reported correlation analysis finds nearly identical persona-cue patterns across personas.

  • IssueBench responses are converted into numeric stance values through stance detection for comparisons across demographics, cues, and models.
  • Llama-3.3-70B is used as the stance-detection judge with the best-performing prompt template from prior work.
  • The stance scale distinguishes pro and con orientations from 1 to 5, with refusal as an additional label.
  • The prompt template instructs the judge to label the generated text using the stance scale or refusal, based only on the text and instructions.
  • Figures 6–8 show almost no differences in persona-cue correlation patterns across the ten personas.

G Correlations by Model

Across models, explicit mentions and names are more strongly correlated with each other than with conversation histories, although some larger models respond more inconsistently to cues. Most aggregate patterns persist across model–dataset combinations, with notable exceptions.

  • G Correlations by Model: For each model, explicit mentions and names correlate more strongly with each other than with conversation histories.
  • G Correlations by Model: Larger Llama and Gemma models appear more inconsistent across persona cues, with no observed difference between open-source models and GPT-4o-mini.
  • G Correlations by Model: Most model- and dataset-level combinations preserve the aggregate correlation patterns across persona cues.
  • G Correlations by Model: Some exceptions include lower cue correlations on MMMD, exceptionally low name–conversation-history correlation for Llama-3.1-70B on AITA, and a Gemma-3-27B IB case where names correlate most with conversation histories.

H Result Metrics per Dataset and Personas

The paper reports result metrics across datasets, persona cues, and personas, with figures separating model-specific and aggregated patterns. Persona differences are tested through mean comparisons, with significance evaluated using the Tukey-Kramer test.

  • Result metrics: Figures 11–32 report average result metrics for each dataset subset and persona cue, both by model and aggregated across models.The figures cover accuracy, average answers, salary, and model stance across the evaluated datasets and demographic dimensions.
  • Cross-model patterns: Although result-metric magnitudes differ across models, the reported patterns tend to remain the same.
  • Statistical comparisons: Tables 5–7 report mean differences between persona pairs and whether those differences are significant at α = 0.01 using the Tukey-Kramer test.
  • Interpretation caveat: Effect sizes should be compared only within the same dataset because the outcome variable has different meanings across datasets.
  • Persona dimensions and datasets: The evaluations compare persona-conditioned outcomes across age, gender, and race/ethnicity personas on writing, advice, bias, and salary-related datasets.The listed figures include AITA and MMMD accuracy, SBB domain outcomes and salary, and IB model stance.
Loading 2601.18572v3…