Source-linked AI summary
LLM Generated Persona is a Promise with a Catch
Ang Li, Haozhe Chen, Hongseok Namkoong, Tianyi Peng
TL;DR
Traditional persona data collection is expensive, privacy-constrained, and incomplete, motivating scalable LLM-generated alternatives. The paper systematizes persona generation and evaluates it through large-scale experiments and bias analysis. It finds substantial bias and deviations from real-world outcomes, motivating a more rigorous science of persona generation and careful application.
Problem
Realistic persona data is difficult to collect at scale because of cost, logistics, privacy constraints, and limited coverage of multidimensional subjective attributes.
Method
The paper systematizes persona-generation approaches, conducts large-scale evaluations, analyzes generated persona profiles, and outlines methodological directions for rigorous generation.
Results
Current persona-generation practices produce substantial bias, and increasing LLM-generated persona content can drive simulations farther from real-world outcomes.
Takeaways & Limitations
Reliable persona simulation requires principled methodologies, empirical foundations, benchmarks, interdisciplinary support, and careful attention to risks and biases.
Takeaways & Limitations
None of the existing methods evaluated reliably produces realistic societal-level opinions.
Abstract
from arXiv · showhide
The use of large language models (LLMs) to simulate human behavior has gained significant attention, particularly through personas that approximate individual characteristics. Persona-based simulations hold promise for transforming disciplines that rely on population-level feedback, including social science, economic analysis, marketing research, and business operations. Traditional methods to collect realistic persona data face significant challenges. They are prohibitively expensive and logistically challenging due to privacy constraints, and often fail to capture multi-dimensional attributes, particularly subjective qualities. Consequently, synthetic persona generation with LLMs offers a scalable, cost-effective alternative. However, current approaches rely on ad hoc and heuristic generation techniques that do not guarantee methodological rigor or simulation precision, resulting in systematic biases in downstream tasks. Through extensive large-scale experiments including presidential election forecasts and general opinion surveys of the U.S. population, we reveal that these biases can lead to significant deviations from real-world outcomes. Our findings underscore the need to develop a rigorous science of persona generation and outline the methodological innovations, organizational and institutional support, and empirical foundations required to enhance the reliability and scalability of LLM-driven persona simulations. To support further research and development in this area, we have open-sourced approximately one million generated personas, available for public access and analysis at https://huggingface.co/datasets/Tianyi-Lab/Personas.
1 Introduction
LLM-generated personas offer scalable representations for population-level simulation, but current generation practices introduce biases that can undermine realism and downstream validity. The paper systematizes persona generation, evaluates its biases, analyzes their sources, and proposes foundations for a more rigorous field.
- Motivation: LLM-driven simulations use demographic, psychographic, and behavioral personas to approximate real-world populations and simulate individual decisions at scale.These synthetic agents can support applications across social science, politics, economics, marketing, psychology, entertainment, and business.
- Motivation: Traditional persona data collection is costly, time-consuming, privacy-constrained, and often misses joint distributions and subjective attributes.Large real-world datasets may provide marginal demographics without capturing combinations of characteristics, lifestyle preferences, or nuanced beliefs.
- Research Gap: Current scalable persona-generation methods are significantly biased and nonrepresentative, creating serious concerns for opinion simulation and decision-making.The paper highlights a 2024 presidential-election simulation as an example of these deviations.
- Contributions: The paper systematizes existing approaches, evaluates current practices at scale, analyzes bias sources, and outlines methodological and institutional directions for rigorous persona generation.It also releases all generated personas from the experiments to support community research.
- Research Gap: The paper studies persona-generation bias itself, showing that generated profiles can contain substantial unexamined bias even when simulation outputs appear unbiased.This shifts attention from traditional LLM simulation bias toward fairness and representativeness in the generation step.
2 Persona-Driven Opinion Simulation
The paper constructs increasingly detailed persona types and uses them in LLM-based opinion simulations. Its design compares structured and freeform generation while testing whether additional LLM-generated content improves diversity or instead amplifies bias.
- Experimental Setup: The study implements commonly used persona-generation and simulation strategies to establish a foundation for large-scale experiments.It constructs three persona types: Meta, Tabular, and Descriptive Personas.
- Meta Personas: Meta Personas use census-based demographic information, but marginal distributions can produce incongruous combinations because joint probabilities are not captured.These personas contain no LLM-generated content and provide a foundation for more varied persona construction.
- Tabular Personas: Tabular Personas extend Meta Personas with structured objective and subjective attributes generated by LLMs, using predefined categories for measurable fields and open-ended responses for less standardized ones.Objective attributes include income, education, occupation, and industry; subjective attributes include political affiliation and leisure preferences.
- Descriptive Personas: Descriptive Personas prompt LLMs to generate detailed freeform, narrative-style profiles conditioned on Meta Personas.Their flexibility requires careful prompting and post-processing to preserve plausibility and coherence.
- Simulation Design: Increasing LLM-generated persona content is intended to broaden opinion distributions but can exacerbate simulation bias across domains.The simulation samples 1,000 personas per state, generates 1,000 augmented personas for each type, and evaluates survey responses under prompt variants.
- Simulation Design: The simulation treats LLMs as impartial world models and asks them to select survey responses that reflect each persona’s attributes.The pipeline provides high-level neutrality instructions, a persona, and a multiple-choice question.
3 Can Personas Simulate the Society?
The experiments evaluate persona-based simulations of elections and public opinion, finding that increasing LLM-generated persona content systematically reduces alignment with real-world outcomes and shifts simulated perspectives leftward. These patterns extend across models and domains, raising concerns about using such personas for research and business decisions.
- Experiments: Experiments cover U.S. elections from 2016, 2020, and 2024, plus opinion surveys spanning diverse social, economic, and personal topics.The study compares persona types with varying amounts of LLM-generated content against real-world data where available.
- 3.1 A Case Study with US Presidential Election Voting: Increasing LLM-generated persona attributes makes election simulations increasingly deviate from reality, with Meta Personas closest and Generative Personas most divergent.In one Llama 3.1 70B setup, the simulated electorate progressively shifted toward left-leaning stances.
- 3.1 A Case Study with US Presidential Election Voting: Historical election exposure does not remove the bias: 2016 and 2020 simulations show a similar leftward drift.The result runs contrary to the expectation that training-data exposure would improve electoral distributions.
- Quantitative evaluation: The alignment metric is 1 −W(ˆp, p), where higher scores indicate simulated support rates closer to actual voting rates; cross-simulation tests whether bias depends on the model.Figure 5 reports alignment across simulation and persona-generation models.
- 3.2 Exploring Bias Across Domains: Across broader opinion scenarios, more LLM-generated content shifts perspectives from traditional to progressive, including preferences for expensive eco-friendly cars, liberal arts, and La La Land.These qualitative examples lack ground-truth responses for alignment measurement.
- Implications: The experiments are not exhaustive, but they indicate that persona-model differences can influence societal and individual decision-making scenarios.The paper warns that such preferences could mislead business decisions about products and entertainment markets.
4 A Closer Look into Persona Profiles
The paper examines how LLMs construct persona profiles as they receive more freedom to generate details. More detailed personas become more subjective and positive, with language patterns emphasizing optimism, social connection, achievement, and stability.
- Persona characteristics: The analysis asks what characteristics and biases emerge when LLMs receive greater freedom to generate persona details.This follows the finding that increasing generated content biases simulation outcomes.
- Sentiment analysis: TextBlob sentiment analysis finds that subjectivity increases as personas contain more generated details.The analysis compares personas with different levels of LLM-generated detail.
- Sentiment analysis: More generated details also produce more positive sentiment, with descriptive personas showing significantly higher sentiment polarity.This pattern is reported quantitatively in the sentiment analysis.
- Language patterns: Word clouds reinforce the positive framing through terms such as “love,” “proud,” “family,” “community,” “education,” and “work.”The vocabulary suggests optimistic outlooks, strong social connections, and emphasis on achievement and stability.
- Interpretation: The paper links positive sentiment and emotional characterization in personas to skewed responses, especially on social issues where emotional reasoning may matter more.This provides a proposed explanation for the simulation biases observed earlier.
5 Paths Forward: Toward a Scientific Approach to Persona
The paper identifies unreliable societal-level opinions as a central limitation and proposes methodological, data, and collaboration priorities for scientific persona generation. These priorities include identifying essential persona information, calibrating generated populations, building open benchmarks, and coordinating interdisciplinary research.
- Current persona generation methods do not reliably produce realistic societal-level opinions, motivating a more principled scientific approach.
- Identifying essential information needed in a persona: Essential persona information must be identified and represented according to what drives realistic simulation outcomes, rather than merely listing attributes.
- Calibrating LLM-generated personas towards real population: Generated personas should be calibrated to target populations by recovering realistic joint attribute distributions from fragmented data sources.
- Open-source benchmark and datasets: An open-source benchmark could evaluate generation methods, support training, and provide diverse population-level personas for direct simulation use.
- Interdisciplinary research and broad community collaborations: Interdisciplinary collaboration could support rapid, cost-effective human experiments across academic fields and industry applications such as software and web design.
6 Related Work
Prior work spans traditional synthetic populations, LLM-based simulations, and investigations of bias in role-play and persona assignment. However, existing studies have largely established feasibility or examined specific biases rather than providing rigorous evaluation of persona simulations.
- Persona generation: Traditional persona generation uses census and survey data, while agent-based models sample attributes to address statistical validity, scalability, and diversity.
- LLM Driven Simulation: LLM-driven simulation research covers digital twins, role-playing, personality simulation, and societal-scale interactions.
- LLM Driven Simulation: Prior LLM simulation studies primarily demonstrate feasibility rather than conducting rigorous evaluations.
- LLM Bias: Related bias research reports harms from persona assignment and risks of biased or harmful outputs in role-play settings.
- LLM Bias: A separate multi-step reasoning framework addresses simulation biases and adapts to evolving political contexts for greater accuracy.
A.1 Language Model Selection
The study evaluates six open-source language models selected to vary in alignment strategy and geographic training origin. This design supports analysis of whether model refinement and training location relate to biases or responses.
- The evaluation includes six open-source models: Athene 70B, Llama 3.1 8B, Llama 3.1 70B, Mistral-8x7B Instruct V0.1, Nemetron 70B, and Qwen 2.5B.
- Alignment strategy: Models differ in alignment strategy, including instruction tuning alone versus additional Reinforcement Learning from Human Feedback refinement.
- Geographic diversity: The selection includes models developed primarily outside the United States to examine whether training origin influences biases or responses.
A.2 A Special LLM
Most evaluated language models show left-leaning political bias, while Yi-34B Chat is a notable right-leaning exception. The contrast emphasizes careful model selection and alignment for persona simulation.
- Most language models exhibit significant bias toward politically left-leaning perspectives.
- Yi-34B Chat demonstrates substantial bias toward right-leaning political viewpoints, contrasting with the broader model pattern.
- The observed exception reinforces the importance of carefully selecting and aligning models for persona generation and simulation tasks.
A.3 Feedback Collection
The approach aggregates selected-choice counts across the simulated population instead of computing token-level log probabilities for each choice.
- Selected-choice counts are aggregated across the simulated population.This replaces token-level log-probability computation for each choice.
- The method leverages modern LLMs’ instruction-following capabilities.
- Feedback collection operates at the population level rather than the individual token-probability level.
B Persona Generation and Simulation Prompts
The appendix specifies prompts and templates for generating objective, subjective, and descriptive personas, then simulating opinions from those personas. It combines constrained demographic categories with richer subjective and descriptive attributes.
- B.1 Persona Generation System Prompts: The Persona Generation System Prompts instruct the model to create specific, realistic, and diverse personas from demographic information in a comprehensive JSON template.
- B.2 Objective Tabular Persona: Objective tabular personas are generated by filling a final template consistently with all metadata while selecting values from prescribed ranges and categories.The template includes demographic, employment, income, household, birthplace, veteran, disability, and insurance fields.
- B.2 Objective Tabular Persona: The objective persona schema organizes occupations into broad categories and detailed subcategories.
- B.3 Subjective Tabular Persona: Subjective tabular personas extend the objective schema with ideology, political views, and Big Five personality scores.The instructions distinguish categorical features from fields requiring concise, objective descriptions.
- B.4 Descriptive Persona: Descriptive personas are designed to be detailed, diverse, vivid, and three-dimensional while remaining consistent with the supplied metadata.The instructions require specific values within ranges, varied perspectives and backgrounds, and committed details for each individual.
- B.5 Opinion Simulation Prompts: Opinion simulation prompts ask the system to generate realistic opinions from a given persona about a specific topic.