Source-linked AI summary
Large language models that replace human participants can harmfully misportray and flatten identity groups
Angelina Wang, Jamie Morgenstern, John P. Dickerson
TL;DR
The paper asks whether LLMs can replace human participants while preserving the influence of demographic identity. It analyzes four LLMs against human identity-group portrayals and finds misportrayal and flattening, while identity prompts can also essentialize identities. The authors urge caution about replacement and report inference-time alternatives that reduce, but do not remove, these harms.
Problem
Replacing human participants with identity-prompted LLMs risks failing to represent demographic groups’ differing perspectives and lived experiences.
Method
The paper compares four LLMs with human in-group portrayals and out-group imitations across demographic identities, question types, and multiple measurements.
Results
Across four LLMs, responses showed misportrayal toward out-group imitations and flatter representations than humans, including GPT-4 and 3.5 covering only 3 of 5 multiple-choice possibilities.
Takeaways & Limitations
LLM replacement should be approached cautiously when demographic identity is relevant, while alternatives such as identity-coded names may mitigate harms in supplementation settings.
Takeaways & Limitations
Identity prompting can essentialize identities by legitimizing them as rigid and innate, and the harms of replacement vary with the reason for prompting and the use setting.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasing in capability and popularity, propelling their application in new domains -- including as replacements for human participants in computational social science, user testing, annotation tasks, and more. In many settings, researchers seek to distribute their surveys to a sample of participants that are representative of the underlying human population of interest. This means in order to be a suitable replacement, LLMs will need to be able to capture the influence of positionality (i.e., relevance of social identities like gender and race). However, we show that there are two inherent limitations in the way current LLMs are trained that prevent this. We argue analytically for why LLMs are likely to both misportray and flatten the representations of demographic groups, then empirically show this on 4 LLMs through a series of human studies with 3200 participants across 16 demographic identities. We also discuss a third limitation about how identity prompts can essentialize identities. Throughout, we connect each limitation to a pernicious history of epistemic injustice against the value of lived experiences that explains why replacement is harmful for marginalized demographic groups. Overall, we urge caution in use cases where LLMs are intended to replace human participants whose identities are relevant to the task at hand. At the same time, in cases where the benefits of LLM replacement are determined to outweigh the harms (e.g., the goal is to supplement rather than fully replace, engaging human participants may cause them harm), we provide inference-time techniques that we empirically demonstrate do reduce, but do not remove, these harms.
Preliminaries
The study examines when demographic identity prompts may be used, how identity-related questions are categorized, and how four LLMs are evaluated against human identity-group portrayals. It uses 16 identities across five demographic axes, free-response questions, and multiple complementary measurements.
- Study design: The analysis covers race, gender, intersectional identity, age, and disability, totaling 16 identities across five demographic axes.Participants were recruited on Prolific and compensated $12 per hour.
- Prompting rationales: The paper distinguishes four prompting rationales: contingent, relevant, subjective, and coverage.Coverage concerns increasing response coverage, while the first three concern identity’s relationship to meaning, truth, or subjectivity.
- Prompting rationales: Coverage is analyzed separately because it presupposes that at least one of the contingent, relevant, or subjective rationales applies.The nine questions include one contingent, two relevant, three subjective, and three coverage questions.
- Measurement: The experiments use free-response questions, 100 responses per demographic group and generation source, and both embedding-based and categorical measurements.Multiple measurements are used to assess robustness and expose possible contradictions between measures.
- Measurement: GPT-4 results compare LLM responses with out-group imitations and in-group portrayals using t-statistics across question types and identity axes.Statistical significance is marked at p < .05, with figures distinguishing significant and nonsignificant comparisons.
LLMs can misportray groups as out-group imitations
Across four LLMs, demographic personas often resemble out-group imitations more than in-group portrayals, especially for non-binary people and people with impaired vision. Identity-coded names can improve alignment for Black personas, but the improvement is incomplete and less consistent for White personas.
- Misportrayal evidence: Many GPT-4 personas are more similar to out-group imitations than in-group representations, especially for White men, women, non-binary people, Generation Z, and people with impaired vision.The comparison uses nearest-neighbor distances between SBERT embeddings and two-sided Welch’s t-tests.
- Misportrayal evidence: Across four LLMs, R1-Contingent misportrayal was significant for White people in 23/24, non-binary people in 16/24, and people with impaired vision in 18/24 measurement-model comparisons.These counts summarize six metrics across the four models.
- Misportrayal evidence: For R2-Relevant, misportrayal appeared for non-binary people in 32/48, people with impaired vision in 27/48, Generation Z in 27/48, and women in 26/48 comparisons.White-person misportrayal was less frequent in this setting, appearing in 15/48 comparisons.
- Misportrayal evidence: R3-Subjective showed no misportrayal effects because demographic identities and personas generated minimal differences in the more constrained annotation tasks.The paper reports weaker effects for this question type than for contingent and relevant questions.
- Reason for harm: Non-binary people and people with impaired vision were misportrayed across R1-Contingent and R2-Relevant for all four LLMs, affecting groups that are historically excluded and underrepresented.The paper connects this pattern to the history of speaking for marginalized groups and the erasure of lived perspectives.
- Alternative: Identity-Coded Names: Across four LLMs, identity-coded names made Black men’s and Black women’s responses more, though not fully, aligned with in-group representations on R1-Contingent and R2-Relevant.The improvement had exceptions, including Black men on Llama-2, and was less evident for White men and White women.
LLMs flatten groups and portray them one-dimensionally
LLMs generate responses that are less diverse than human in-group responses, flattening demographic groups into narrower and sometimes one-dimensional portrayals. Increasing temperature can produce more varied text, but only at settings that introduce incoherence and still fail to match human diversity across most measures.
- LLMs across all four models and nearly all diversity measures generated flatter responses than human participants.GPT-4 and GPT-3.5 typically covered only 3 of 5 multiple-choice possibilities in 100 responses per scenario.
- This flattening can erase within-group heterogeneity, especially for marginalized groups historically portrayed as one-dimensional.The paper links this harm to failures to recognize intersectional differences and varied experiences within demographic groups.
- For non-binary identities, LLM responses often uniformly described pronoun-related difficulty while missing variation in pronoun use and gender identity.Human in-group responses included people using combinations such as he/him and they/them.
- At temperature 1.4, GPT-4 reached human diversity only on unique n-grams, while its outputs became incoherent and remained less diverse on three semantic measures.The analysis tested temperatures 1.0, 1.2, and 1.4 on the intersectional demographic axis.
- Prompting and temperature tuning may increase response heterogeneity but are unlikely to fully match the range of human experiences.
Alternatives to demographic personas for increasing coverage
Identity prompts are not necessary to increase response coverage, because non-demographic alternatives can achieve comparable or greater coverage. Using such alternatives can also avoid essentializing demographic identities as rigid and innate.
- R4-Coverage uses identity prompting to increase the quantity of semantically distinct responses for simulations, harm anticipation, and user-study exploration.
- The study compares sensitive demographic prompts with personality types, behavioral personas, political leanings, astrology signs, and generic prompts.
- Alternative prompts can achieve coverage as high as or higher than sensitive demographic prompts across the evaluated metrics.The figure evaluates three response-coverage metrics across three R4-Coverage questions for GPT-4.
- No model required sensitive demographic attributes to achieve the highest response coverage.Random personas produced the highest coverage on all models except Wizard Vicuna Uncensored, where astrology and Myers-Briggs performed well; generic prompting usually had the lowest coverage.
- When alternatives exist, they may reduce the harm of essentializing identities as rigid and innate and amplifying perceived differences between groups.
- GPT-4 and Llama-2 produced stereotyped identity-coded greetings and expressions for Black women compared with White men.Examples included “Hey girl!” and “Oh, girl” for Black women versus “Hey buddy” and “Hey mate” for White men.
Discussion
The paper identifies persistent harms in identity-prompted LLMs and offers mitigation strategies for use cases where supplementing humans may still be worthwhile. It also emphasizes that harms depend on the reason for prompting and that the analysis covers only a limited set of groups.
- Persistent limitations: Two critical limitations and one further consideration are expected to persist under current online-text training and likelihood-based losses.The authors argue that newer models trained in the same way may not easily resolve these problems.
- Mitigation: Inference-time alternatives can mitigate, but not eliminate, harms when LLMs supplement humans, costs are prohibitive, or participation risks distress.The paper discusses identity-coded names and temperature manipulation as mitigation strategies.
- Use-case boundaries: The reason for prompting mediates harm: R1-Contingent uses carry higher normative consequences, while some R3-Subjective and R4-Coverage uses may be more permissible or justifiable.R1 treats social location as determining meaning and truth; R2 and R3 treat it as bearing on meaning and truth.
- Use-case boundaries: The discussion distinguishes whether LLM replacement can be done from whether it should be done, including autonomy-related concerns about replacing human voices.The paper frames this as a distinction between capability and normative justification.
- Scope: The study covers 16 demographic groups in America, while warning that online-text training may underserve people without Internet access and cultures centered on oral traditions.The authors connect this scope boundary to the risk of erasing marginalized and offline voices.
Methods
The study compares identity-prompted LLM responses with human responses across four question rationales, five demographic axes, and free-response and discretized analyses. Multiple metrics and bootstrap confidence intervals are used to reduce dependence on any single measurement choice.
- Participants and models: The four-model analysis includes two accessible open-source 7-billion-parameter models and two closed-source models, with human data withheld because it is sensitive and personal.LLM-generated data and code are released at the cited OSF location.
- Questions and rationale: The study defines four prompting rationales: contingent, relevant, subjective, and coverage, using nine questions distributed across them.R1 has one question, R2 has two, R3 has three, and R4 has three.
- Response representation: The analysis uses free responses alongside five-point categorical discretizations, comparing LLM generations with human participant responses.Human participants map their own responses to the scale, while GPT-3.5 classifies LLM responses.
- Statistical analysis: The study uses 95% bootstrap confidence intervals and multiple metrics to reduce dependence on statistical artifacts or any single measurement choice.For Fig. 4, questions are treated as separate clusters during bootstrapping.
- Measures: Misportrayal is measured with n-gram Jaccard and SBERT distances, while flattening uses n-gram uniqueness, SBERT diversity statistics, and multiple-choice uniqueness.These measures compare response similarity or diversity across generated and human response sets.
A Results Across all 4 LLMs
Across the four LLMs, the results include multiple-choice analyses and show model-specific differences in political-response inflation. Wizard Vicuna Uncensored inflates groups as more liberal, whereas GPT-3.5 and GPT-4 show little such inflation relative to in-group participants.
- Supplementary results: The supplementary results present all four LLMs and correspond to the main-text figures for misportrayal, flattening, and related analyses.The appendix includes results that were not accommodated in the main text.
- Multiple-choice results: Across R1, R2, and R3, Wizard Vicuna Uncensored overinflates all groups as more liberal than the other LLMs.The result differs from prior intuitions that alignment creates politically liberal biases.
- Multiple-choice results: GPT-3.5 and GPT-4 show little liberal inflation compared with in-group members in the multiple-choice responses.These results appear in the analysis corresponding to R1, R2, and R3.
B Establishing Premises
The premise checks establish that identity prompts change LLM responses and that human in-group and out-group responses can differ. These differences provide the basis for evaluating whether LLM portrayals align more closely with in-group or out-group perspectives.
- Premises: The analysis first tests whether demographic identity prompts change LLM responses and whether human in-group and out-group participants respond differently.These premises are necessary for interpreting the paper’s misportrayal analysis.
- LLM response differences: Across all four LLMs, responses change according to the demographic identity used in the prompt.The observed identity-conditioned difference is reported in Fig. 13.
- LLM response differences: The identity-prompted difference can exceed the difference represented by out-group human imitations, especially outside R3-Subjective questions.The paper reports much less identity-conditioned difference for R3-Subjective questions.
- Human response differences: Differences between human in-group and out-group responses vary by identity and questioning rationale, with the smallest differences in R3-Subjective.The paper notes stronger differences for identities such as Black people or Black women, while many cases are not statistically significant.
- Human response differences: The premise checks support comparing LLM portrayals with in-group and out-group human responses, while acknowledging that baseline closeness may limit the size of some comparisons.This comparison is central to assessing misportrayal.
C Prompt Details
The study uses identity and topic prompts to elicit LLM responses across four identity-prompting rationales, demographic axes, and question types.
- Prompt construction: Each prompt combines a demographic identity with a topic, using model-specific input formats for GPT-3.5, GPT-4, Llama-2-Chat, and Wizard-Vicuna-Uncensored.The identity prompt is placed in the system input for GPT-3.5 and GPT-4 and formatted differently for the other models.
- Prompt construction: The identity prompts cover race, gender, intersectional identities, age, and disability for R1-Contingent, R2-Relevant, and R3-Subjective analyses.The listed identities include 16 demographic identities, such as Black, Asian, and White people; men, women, and non-binary people; and age groups.
- Response collection: Responses generally request one paragraph of 4–5 sentences, while R3-Subjective prompts use shorter task-specific response instructions.The study converts free responses to five-point Likert choices using GPT-3.5 three-shot examples, with humans selecting corresponding answers themselves.
- Study design: The questions include lived experience, healthcare, gun regulation, immigration, criminal justice, climate change, toxicity, and positive reframing.For R2-Relevant, demographic axes are paired with political topics selected using empirical maximum entropy.
D Prompt Phrasing Robustness
The authors test whether findings depend on prompt wording by comparing four identity-prompt variants and examining response embeddings.
- Evaluation: The robustness analysis compares bag-of-words n-gram and SBERT embeddings across prompt formulations.Similar embeddings across formulations would indicate that results are not driven solely by prompt artifacts.
- Prompt variants: Four prompt variants modify whether the model is told to be an identity, speak exactly like one, or remember that identity is only one part of identity.The variants are evaluated for Black women answering where they like to vacation.
- Results: Heavy overlap appears across prompt variants for most models, while GPT-3.5 shows differences between the first two and the latter two formulations.Llama-2 shows a smaller version of this pattern, and qualitative inspection finds no notable differences.
- Interpretation: The authors use Prompt 3 because it is the simplest and most likely to represent actual use cases, while cautioning that wording may matter for GPT-3.5.The robustness study supports limited concern that the broader findings are artifacts of the selected prompt wording.
E Noise in Human and LLM Generations
The authors identify and address noise from model refusals, identity markers, and possible human use of LLMs in survey responses.
- Data limitations: The collected human and LLM response datasets likely contain residual noise despite cleaning efforts.The paper treats this remaining noise as a known characteristic of the datasets rather than claiming complete removal.
- LLM-generated responses: Fewer than 5% of responses for each model are refusals, which the authors rerun to estimate an upper bound on perspective representation.The refusals occur when alignment prevents a model from answering potentially harmful questions.
- Data cleaning: Identity markers such as “As a woman” are removed before analysis so textual labels do not create apparent demographic differences.Cleaning is more difficult for behavioral personas whose characteristics recur throughout responses.
- Human-generated responses: The authors do not remove suspected LLM-generated human responses because available heuristics were insufficient for reliable filtering.They estimate that labeled scenarios contain fewer than 10% apparent LLM-assisted responses, compared with a prior estimated prevalence of 30%.
F Related Work
The paper builds on prior studies of demographic prompting, stereotypes, identity simulation, and subjective annotation while broadening the set of tasks and harms considered.
- Demographic prompting: Prior work finds that demographic prompting shifts LLM political opinions toward human groups without fully aligning them.This paper extends that premise to more questions and specific hypotheses tied to histories of harm.
- Stereotypes: Related research reports stereotype amplification when LLMs describe themselves under demographic prompts.The present work differs in task range and in the types of problems it studies.
- Identity simulation: A framework for identity simulation measures individuation and exaggeration, with individuation closest to one premise examined here.The paper distinguishes its exaggeration analysis from that framework’s measure.
- Subjective annotation: Studies of subjective annotation use multiple-choice LLM responses and emphasize label accuracy rather than the social harms examined here.Those studies map most directly to the paper’s R3-Relevant category.
- Synthesis: Together, these studies provide complementary, harm-specific analyses that the paper combines to cover a broader range of limitations in identity-prompted LLMs.The authors frame this collective coverage as strengthening precision about where harms stem from.
G Human Participant Demographics
The paper documents human-participant demographics and compares LLM responses with human in-group and out-group representations across identity-related measures. These supplementary figures examine similarity, diversity, coverage, prompt phrasing, and response differences.
- Human participant demographics: The human-participant table records self-reported gender, race, and age categories for 100 participants in each study.Participants could select multiple genders or races and could decline to disclose identity.
- Representation comparisons: LLM responses are compared with human in-group and out-group representations across six similarity metrics and multiple sets of reasons.Figures 7 and 8 assess demographic prompts and identity-coded names or explicit identity labels.
- Group diversity: Across all question types and demographic groups, LLM responses are less diverse than human in-group responses.Figure 9 reports four diversity metrics with 95% confidence bars across 100 responses per demographic group.
- Temperature analysis: Varying temperature does not solve flatness for Wizard Vicuna Uncensored; at 1.8, responses become incoherent.Although unique n-gram diversity exceeds the human value at that setting, no other semantic metric reaches human diversity.
- Response coverage: Alternative prompts achieve response coverage as high as or higher than prompts containing sensitive demographic attributes.Figure 11 compares no-identity, sensitive-attribute, and alternative prompts across three coverage metrics and three questions.
- Response differences and robustness: Prompted demographic identities produce different LLM answers, while comparisons of human groups and prompt variations examine the sources of those differences.Figures 13–15 respectively compare response differences, human in-group versus out-group differences, and embeddings across four prompt phrasings.