Source-linked AI summary
From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction
Chaemin Jang, Junsik Min, Jaewoo Choi, Donggyu Lee, Haiin Lee, Junyoung Park, Namhee Kim, Hyunwoo Kim, Jungwon Kim, Juho Kim, Nuri Kim, Jihee Kim
TL;DR
The paper asks whether LLM-based deliberation can represent population opinion patterns and produce conclusions shaped by peer interaction. Using census-grounded Korean personas and national-survey benchmarks, it finds that these properties can separate: agents generate reasoned exchanges and move positions, but representation and interaction-driven change remain unreliable.
Problem
LLM deliberation requires evidence that persona agents represent population opinion patterns and that peer exchange shapes what agents consider and conclude.
Method
The study benchmarks census-grounded Korean personas on real policy questions against national surveys and evaluates deliberation with controlled interaction conditions.
Results
Persona agents often produce concentrated, demographically misaligned responses, while deliberations generate reasoned exchanges and stance movement that sealed-monologue controls largely reproduce.
Takeaways & Limitations
Population simulation, argument surfacing, and interaction-driven opinion change should be treated as distinct uses with different validation requirements.
Takeaways & Limitations
Argument-surfacing validity remains untested because the available benchmarks measure opinions rather than arguments or perspectives.
Abstract
from arXiv · showhide
Multi-agent LLM deliberation has been explored as a scalable way to simulate public deliberation. For such simulations to be informative, persona agents should reflect population opinion patterns and interaction should shape their conclusions. We evaluate whether LLM-based deliberation can meet these two conditions using census-grounded Korean personas debating real policy questions benchmarked against national surveys. Persona agents do not reliably reproduce population opinion patterns: responses are often far more concentrated and frequently reverse demographic differences in the human data. Deliberations nonetheless produce reasoned, reciprocal, and varied arguments alongside substantial stance movement. Yet much of this movement does not require peer exchange: sealed-monologue agents change position at similar rates and reach nearly the same final balance as full debates, while groups initialized with very different positions often converge to similar endpoints. Anchoring population-informed starting positions, meanwhile, sharply suppresses updating. Thus, population representation, argument generation, and interaction-driven opinion change do not necessarily go together. The simulations readily surface arguments on both sides, though whether they capture the diversity of human perspectives remains untested, leaving open a promising role for argument surfacing even as population simulation requires further validation.
1 Introduction
The paper asks whether LLM deliberation can represent population opinions and whether peer interaction shapes agents’ conclusions. It evaluates these conditions and distinguishes population simulation from argument surfacing.
- The study evaluates LLM deliberation through population opinion representation and interaction-driven conclusion formation.
- Persona agents often produce concentrated responses and incorrect demographic differences, while sealed-monologue groups reach nearly the same final composition as full debates.
- The results separate population alignment, visible deliberative features, and peer-interaction-driven stance change as distinct properties.
- It benchmarks persona-conditioned responses against demographic patterns in national surveys and tests whether mismatches persist across specifications, formats, populations, and models.
- Sealed-monologue controls test whether stance changes depend on hearing and responding to other agents or arise without peer exchange.
2 Related work
Prior work studies LLMs as simulated publics, persona representation, simulated debates, and deliberation quality. This paper extends those lines by testing population alignment and interaction in census-grounded Korean deliberation.
- Earlier studies use persona-conditioned LLMs to simulate social behavior, survey distributions, mediation, and multi-agent deliberation.
- Research on population representation reports demographic misalignment and cautions that response artifacts, format sensitivity, and question variants affect survey-derived alignment.
- Studies of simulated debates identify persona fragility, while debates on verifiable tasks differ from deliberation because deliberation has no answer key.
- Deliberation quality can be assessed through balanced exchange, discourse-quality dimensions, and the diversity of arguments given for each side.
3 Setup
The study constructs balanced Korean persona pools and evaluates eight two-position policy questions against national survey benchmarks. Six personas then deliberate across three rounds using a cumulative shared transcript.
- Figure 1 first elicits baseline responses, then has six personas deliberate for three rounds over a cumulative transcript.Each agent reads the full log, speaks once, and appends its turn in every round.
- The Korean benchmark covers environmental and low-birthrate policies using eight two-position questions from two national surveys.Four questions come from the KEI survey and four from the PCASPP survey.
- Source questions with more than two response options are reduced to two substantively contrasting positions, with original-format analyses provided separately.
- The primary pool contains 640 Korean personas jointly balanced across sex, age, education, and region.Each persona retains demographic attributes, occupation, household information, and a background narrative.
- The study also constructs a parallel U.S. pool of 256 personas and benchmarks five matched policy questions against Pew surveys.
4 Personas as survey respondents
Persona responses fail to reproduce human demographic opinion patterns despite survey controls and persona conditioning. Adding demographic attributes reduces average gaps, but the mismatch persists across tested setups, populations, and models.
- 29 percentage points is the mean absolute gap between persona responses and human survey benchmarks, with group-level means ranging from 24 to 34 points.
- On five of eight questions, at least one persona group is near 0 or 100 while corresponding human shares are less concentrated.
- Persona responses match the survey direction of demographic differences in only 34 of 110 comparisons.
- On the two divisive questions, personas match the human ordering of demographic groups in only 9 of 28 comparisons.For education–care, personas in their twenties have an A-share of 45%, versus 61% among respondents in their twenties.
- Adding four demographic attributes reduces the mean absolute gap from 43 to 30 points, while the full narrative profile yields a similar 29-point gap.The demographic attributes account for nearly all of the tested improvement in average alignment.
- The mismatch persists across alternative response formats, original survey procedures, a U.S. persona pool, and four models including a frontier model.
5 Personas as deliberating agents
The simulated discussions show substantial stance movement and increasingly varied, reasoned exchange, but final outcomes are weakly tied to starting composition and much movement occurs without peer interaction. Carrying assigned positions into the transcript strongly anchors agents and suppresses updating.
- Between 36 and 53% of agents change position over three rounds, while rooms converge to a majority on five of six questions.
- Agents produce well-justified, respectful, reciprocal exchanges, increasing their distinct reasons from 2.6 in the first turn to 4.9 by debate end.
- Rooms starting with all-A, balanced, or all-B compositions end within about two agents of one another on the three tested questions.
- Final room composition does not consistently track the population benchmark, personas’ starting responses, or the model’s no-persona response.
- Sealed-monologue rooms differ from full debates by at most one agent at the end, with a mean difference near zero.
- When assigned positions are restated or argued, stance movement falls to 0–3% in all but three cells and never exceeds 11%.
- Removing agents’ own turns from stance-survey context restores 61% movement, whereas removing same-side peers’ turns leaves movement at 1%.
- The findings distinguish argument generation, population representation, and interaction-driven opinion change as separate properties of LLM deliberation.
6 Discussion and Future Directions
The findings caution against treating LLM deliberation as population simulation, while leaving open a narrower role in surfacing arguments and perspectives for human inspection. That role remains unvalidated and raises challenges of scope, mechanism, and presentation.
- Implications for the use of LLM deliberation: LLM deliberation should not be used as a population simulation without evidence that agents represent populations and respond to one another.The agents do not reliably reproduce baseline opinion patterns, and stance movement can arise without peer interaction.
- Implications for the use of LLM deliberation: Argument surfacing remains a potentially useful alternative because agents generate varied, both-sided arguments and expand their argument repertoires.Whether these arguments reflect the diversity of human perspectives remains unvalidated.
- Limitations and open questions: Validation must separately assess population representation, interaction-driven updating, and the diversity of human perspectives.The study’s available benchmarks measure opinions rather than arguments or perspectives.
- Limitations and open questions: A broader scope boundary remains: deliberation dynamics are tested with one model and two policy domains in Korea, while discourse scores rely on an LLM judge without human validation.Future work should test whether the patterns persist across models, populations, policy domains, and deliberation designs.
- Limitations and open questions: Substantial stance movement without peer exchange remains mechanistically unresolved, with self-consistency, self-persuasion, context accumulation, and prompt-induced dynamics still possible.More targeted interventions on agents’ own discourse histories are needed to distinguish these explanations.
- Limitations and open questions: Making argument-surfacing useful to readers requires representing long interaction histories and multiple possible deliberative trajectories.The paper identifies overview, traceability, and comparability as design requirements for visualization.
A Question wordings and full by-group results
The appendix provides the full Korean policy-question wording and demographic comparisons, showing that persona responses can diverge sharply from population distributions and demographic opinion patterns across instruments.
- Question wordings and full by-group results: The appendix reports eight Korean policy questions and complete persona-versus-population comparisons across demographic groups.The questions cover environmental, climate, work–family, education–care, economic-support, and housing topics.
- Question wordings and full by-group results: Persona and population shares differ across both saturated and divisive policy questions, including environmental, climate, education–care, work–family, economic-support, and housing items.The tables report Position 1 shares by demographic group for each question.
- Instrument comparison: All four language-model response instruments fail to recover population opinion patterns, each in a different way.The instrument comparison uses mean absolute gaps from the population A-share averaged over 88 demographic groups.
- Instrument comparison: The appendix identifies presentational bias in the signed bipolar scale, where the scale sign changes which position is labeled +2.This factorial isolates option order, answer letter, scale sign, and scale direction one factor at a time.
C Asking the personas the human questionnaire
Asking personas the source questionnaire verbatim and applying the human scoring rule does not improve population fit. Instead, it reveals that personas often collapse onto a few options rather than merely choosing the wrong side of a binary mapping.
- Method: The study gives personas the original questionnaire wording, options, and order, then applies the survey’s published recoding rule.Human responses are recomputed from respondent-level microdata for direct comparison.
- Results: 29 to 46 points: the mean absolute population gap rises under the original format on seven defined items, increasing on six.Reversing option order changes little, with 46 against 44, so ordering does not explain the result.
- Results: Personas concentrate on two options even when human respondents distribute across every option, including climate strategy and housing.Within demographic cells, same-pair choices occur 50.8 to 92.9 percent of the time for personas versus 3.7 to 7.6 percent for humans.
- Results: On housing, none of 640 personas selects either benchmark option in its top two, while 97.7 percent select dedicated housing supply.Because the personas are not in the same option region as humans, no comparable A-share exists.
- Conclusion: Giving personas the human questionnaire and scoring rule makes the discrepancy larger and shows collapse onto a small set of options.The representation failure is therefore attributed to persona-conditioned model responses rather than answer elicitation or conversion.
D Robustness across models and populations
The survey-stage failure persists across models and populations. All four tested models over-polarize relative to the population, and a parallel U.S. persona pool shows a similar gap.
- Robustness across models and populations: 29, 28, 33, and 43 points: GPT-4.1-mini, GPT-5.5, Llama-3.3-70B, and Qwen-2.5-72B show these mean absolute population gaps across eight issues.The frontier model is no more moderate than the small model, and none of the four tracks the population.
- Robustness across models and populations: 13 to 94 points: the four models’ A-shares span this range on individual questions, with 54 points on average.The frontier model is the outlier most often, indicating model-dependent answers.
- Robustness across models and populations: 94 to 100 percent: U.S. personas saturate at these levels on three of five matched Pew items.The U.S. pool reproduces the same concentration pattern in a second language and population.
- Robustness across models and populations: 28 points: the U.S. pool’s mean population gap matches the Korean pool’s 29 points.This cross-population comparison supports the persistence of the survey-stage failure beyond the Korean setting.
E Deliberation details
The deliberation evaluation combines controls for response noise and presentation format with judge-based measures of discourse quality and argument diversity. Results indicate that discourse quality and argument repertoire can remain high even when apparent deliberative change is absent, while several interaction variations leave movement and convergence largely unchanged.
- Measurement controls: 5 to 10% of answers flip on repeated no-debate surveys at temperature 0, rising to 15 to 25% at temperature 0.7.Debate-induced movement is reported only when it exceeds this noise floor.
- Benchmarking: The evaluation benchmarks persona responses against demographic-group survey patterns using multiple response formats and model comparisons.The supplied tables cover original survey formats, within-cell dispersion, cross-model natural-choice surveys, and U.S. persona comparisons.
- Measurement controls: The judge-based rubric scores justification, common-good orientation, respect, reciprocity, and distinct reasons for each position.Round-over-round trends are treated as more reliable than absolute levels because judge labels were not hand-validated.
- Results: Every frozen protocol matches or exceeds the live room on all four Discourse Quality Index dimensions and produces a larger, more two-sided argument repertoire despite almost no mind changes.The comparison uses the same gpt-4.1 judge and rubric across the live and frozen conditions.
- Results: Across 294 debates, changing demographic disclosure or opening-side order leaves movement and convergence unchanged.These tests examine whether structural interaction levers alter the final balance.
G Injection: representation is preserved only by freezing
The injection experiment preserves population-informed representation by assigning a balanced three-to-three starting split based on demographic-cell probabilities. That representation is maintained because agents stop updating, and the decomposition indicates that carrying one’s own stated position forward is the anchoring mechanism.
- Injection design: The injected protocol assigns side A to the three agents with the highest demographic-cell probabilities and side B to the remainder, producing a three-to-three split.The assignment matches which cells lean toward each side, not the population’s overall share.
- Benchmarking: The supplied U.S. persona table reports first-position shares against published Pew survey benchmarks rather than the Korean injection experiment.It is included as a comparison table in the broader evaluation materials.
- Injection design: The study uses six rooms per question across eight questions, with balanced natural rooms as the control.The injected rooms preserve the assigned balanced composition through the debate protocol.
- Results: 2% of injected agents change side, while balanced natural rooms move about twenty times as often and converge to the model’s pole.The injected distribution is preserved on every question, but the preservation reflects suppressed updating rather than persuasion.
- Mechanism: Recording an assigned side without having agents argue it produces 50% movement, whereas stating and carrying the agent’s own side freezes updating at 1%.The decomposition therefore isolates the agent’s own stated position as the anchoring component.
H Prompt templates and compute
The implementation uses persona prompts, private attitude surveys, and speech-turn templates across controlled deliberation protocols. Experiments rely on hosted-model API calls, with structural variations including side cues, committed openings, and injected starting positions.
- Prompt templates: The experiments run in Korean, with persona attributes and debate transcripts inserted into prompt templates; private surveys use temperature 0 and speech turns use temperature 1.0.Survey responses are not written back into the transcript.
- Prompt templates: The persona system prompt specifies a South Korean citizen profile including age, sex, region, education, occupation, household, and narrative background.Agents are instructed to answer only in the specified JSON format.
- Prompt templates: Private attitude surveys ask agents to choose one of two policy positions after each round, with option order randomized per persona and fixed across rounds.Before the debate, the transcript block is omitted; after each round, it is inserted.
- Prompt templates: Natural speech prompts expose agents to the issue, two positions, the discussion transcript, and the speaker identity without restating the agent’s round-0 answer.A first speaker instead receives the placeholder text that no one has spoken yet.
- Protocol variants: The controlled protocols add either a side cue in round 1 or a committed round-0 opening whose reasons enter the transcript read by round 1.These variants test whether explicitly restating or arguing the initial side changes later updating.
- Compute: The reported experiments comprise roughly 106,000 hosted-model API calls, including about 68,000 survey responses, 32,000 deliberation-room calls, and 5,400 judge calls.Each experiment completes in minutes to about an hour using 12 to 16 parallel requests.