Source-linked AI summary
ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models
Zhen Wang, Yuqi Ren, Yuehan Cui, Hongxiang Wang, Jianxiang Peng, Zhaoxia Zhang, Bingkun Zhu, Tongxuan Zhang, Dezhi Tong, Deyi Xiong
TL;DR
Individual value simulation is limited by fragmented survey-to-prompt profiling and static evaluation that misses value expression in dynamic interactions. ExpertIVS reconstructs WVS responses with 14 sociological expert agents and evaluates simulated individuals through multi-agent debate. Across 480 individuals from 12 countries, it reports high-fidelity value reproduction and improved value generalization, while retaining limitations in linguistic style fidelity and safety-alignment compatibility.
Problem
Existing methods fragment survey responses into profiles, while static evaluation underexplores value consistency during dynamic deliberation and behavior.
Method
ExpertIVS uses 14 sociological expert agents to reconstruct WVS responses into structured individual profiles and evaluates them through multi-agent debate.
Results
90.78% value restoration fidelity across 12 countries and a 5.3% leave-one-out accuracy improvement over baseline demonstrate high-fidelity reproduction and improved value generalization.
Takeaways & Limitations
ExpertIVS supports dynamically evaluated, interpretable simulation of individual value systems with strong consistency across complex interactions.
Takeaways & Limitations
The model aligns with value logic but struggles to reproduce high-fidelity linguistic style, and safety alignment may constrain simulation fidelity on controversial topics.
Abstract
from arXiv · showhide
Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts, which suffer from semantic fragmentation, failing to capture the internal coherence of human value systems. The value systems of LLMs are typically assessed using static multiple-choice questions, which fail to evaluate the value orientation in real-world dialogue interactions. To address these issues, we propose ExpertIVS, a framework employing 14 Sociological Expert Agents to interpret World Values Survey (WVS) responses through structured professional perspectives, rather than direct responses concatenation. These expert agents perform deep semantic reconstruction to generate robust and internally consistent individual profiles. To evaluate the consistency between LLMs and individual value systems during dynamic interactions, we further introduce a multi-agent debate mechanism. Extensive experiments across 480 individuals from 12 countries demonstrate that ExpertIVS achieves 90.78% value restoration fidelity and significantly outperforms baselines in value generalization (+5.3%). Moreover, ExpertIVS exhibits strong personality discriminability and behavioral consistency, enabling a shift from mere response concatenation to genuine sociological role-playing.
1 Introduction
Existing individual simulations often fail to represent coherent value systems or evaluate them in dynamic interaction. ExpertIVS addresses these gaps by reconstructing WVS responses through sociological experts and testing value consistency through debate.
- Individual value systems remain difficult to model from real-world data, despite the importance of values for explaining divergent behavior in shared contexts.
- Flat survey fragments can miss dependencies among value dimensions, weakening generalization beyond observed questionnaire items.
- Static survey retaking underexplores value consistency because values also appear in deliberation and behavior under conflict.
- ExpertIVS uses 14 value-dimension-specific sociological expert agents to reconstruct fragmented WVS responses into structured, coherent individual profiles.The profiles are intended to improve generalization and realism in simulated individuals.
- ExpertIVS evaluates simulated individuals through multi-agent roundtable debate using Value Alignment, Style Simulation, and Persona Distinctiveness.The mechanism assesses consistency with real individuals during high-density ideological debates.
- 5.3% improvement in leave-one-out accuracy over baseline demonstrates stronger value generalization across value dimensions.Experiments covered 480 real individuals from 12 countries.
2. We design a multi-agent debate mechanism
ExpertIVS introduces three metrics to evaluate whether simulated individuals maintain consistent values during dynamic opinion interactions, and reports superiority in value restoration and generalization.
- Three metrics—Value Alignment, Style Simulation, and Persona Distinctiveness—quantify value consistency during dynamic opinion interactions.
- ExpertIVS outperforms other methods in value restoration fidelity and value generalization capability across extensive experiments.The experiments also analyze how demographic attributes affect individual value simulation.
2 Related Work
Related work spans individual, social, and society-scale simulation, while highlighting limitations in persona construction and static evaluation. ExpertIVS organizes expert-based profiling and debate-based assessment to address these gaps.
- LLM-driven simulations range from individual and scenario modeling to society-scale interaction, motivating a focused treatment of individual value simulation.
- Individual Simulation: Prior individual simulation uses role prompts, sociodemographic personas, WVS reasoning, interviews, narratives, and story-world grounding.
- Individual Simulation: ExpertIVS is structured as expert construction, profile generation, individual simulation, and multi-agent debate evaluation.
- Individual Simulation: Persona-based methods may propagate stereotypes or exhibit stereotype and deviation biases, while debate may inherit systematic biases from base models.
- Social Simulation: Social simulation research studies network diffusion, emergent sandbox behavior, social-contract order, and experiments with more than 10k agents.
- Social Simulation: Few studies quantitatively assess the robustness and consistency of deep-seated values during high-intensity social deliberation.
3 ExpertIVS
ExpertIVS reconstructs individual value systems from WVS responses through sociological experts, then simulates profiles and evaluates them in dynamic multi-agent debates.
- Framework Overview: ExpertIVS uses four stages: expert construction, profile generation, individual simulation, and multi-agent debate evaluation.The framework creates 14 sociological expert agents, generates respondent profiles, instantiates individual agents through in-context learning, and assesses value simulation in dynamic interactions.
- Multi-Agent Debate Evaluation: The debate evaluation uses eight simulated agents and a Base LLM in an opening statement followed by 40 rounds of free debate.After each round, agents update topic agreement scores; the mechanism evaluates alignment, style, and persona distinctiveness.
- Expert Construction: Fourteen experts separately summarize demographic information and 13 WVS value dimensions.One expert focuses on demographics, while the remaining 13 correspond to specific value dimensions.
- Profile Generation: Expert-generated narrative segments aggregate into a complete cognitive profile rather than a concatenation of raw survey responses.Each dimension-specific expert performs semantic compression and value summarization, producing logically consistent value-belief descriptions that are combined into the individual profile.
- Individual Simulation: The comprehensive profile is injected as an LLM system prompt to guide personalized responses grounded in individual value logic and demographic background.Given a context, response generation is conditioned on the individual profile.
- Multi-Agent Debate Evaluation: Persona Distinctiveness Index measures deviation from the Base LLM, while Value Alignment and Style Simulation assess value fidelity and demographic style alignment.Higher PDI values indicate a more independent persona, and VAS and SSS are averaged over five independent runs.
4 Experiments
Experiments evaluate ExpertIVS on value restoration, generalization, demographic effects, and cross-value uncertainty across individuals from 12 countries. ExpertIVS achieves high restoration fidelity and improves held-out value prediction over baselines, with performance varying by value dimension and country.
- Experimental Setup: Experiments cover value restoration, leave-one-out generalization, demographic ablations, and multi-agent debate-related evaluation across 480 respondents from 12 countries.The respondents were drawn from the World Values Survey, with 40 individuals per country.
- Experimental Setup: The baseline converts WVS questionnaires into natural-language value statements, while Simple Concatenation and Anthology provide additional profiling comparisons.VOA accuracy evaluates whether predicted Likert-scale responses match the ground-truth tendency range.
- Value Restoration Fidelity: 90.78% average VOA accuracy measures ExpertIVS value restoration across 12 countries, with country scores ranging from 88.08% in Russia to 93.56% in China.Cross-country population variance is 2.09, with SD = 1.45; the results report no evident culture-specific bias.
- Value Restoration Fidelity: Restoration is strongest for corruption (98.85%), security (98.67%), religious (98.07%), and migration (97.35%), but lower for economy (78.49%) and political cultural (82.72%).The paper relates higher fidelity to clear semantic anchors and lower fidelity to competing priorities requiring fine-grained distinctions.
- Value Generalization Capability: 57.4% average VOA accuracy gives ExpertIVS a 5.3% improvement over the Baseline in leave-one-out value generalization, while also achieving the best overall performance against Simple Concatenation and Anthology.The test masks one value dimension while retaining the other 12 dimensions and demographic attributes.
- Value Generalization Capability: VOA accuracy is negatively correlated with entropy (Pearson r = −0.58, p < 0.001), while gains vary across countries and value dimensions.Britain, the USA, and Germany show improvements of +13.7%, +11.1%, and +9.0%, respectively; religious values, corruption, and migration gain +8.7%, +7.8%, and +7.7%.
- Impact of Demographics: Removing demographics changes restoration accuracy from 90.78% to 90.54% but reduces generalization accuracy from 57.4% to 54.8%, indicating different roles for demographics when direct value evidence is present or absent.Value narratives provide dense signals for restoration, whereas demographics serve as sociological priors during uncertain inference.
- Impact of Demographics: Value narratives provide limited information for reconstructing precise demographics: average VOA accuracy ranges from 39.5% in Kenya to 59.2% in the UK, with education and age especially difficult to predict.Country is easier to recover than education or age, reaching 100% in Germany while education in India reaches 7.5% and age in Kenya 10.0%.
5 Discussion
Discussion examines how ExpertIVS shapes behavior in dynamic debate and how its profiles support more interpretable value reasoning. Agents exhibit culturally differentiated responses and coherent arguments rather than shallow value labels.
- Dynamic Behavioral Evaluation: Multi-agent debate evaluates simulated individuals through Value Alignment, Style Simulation, and Persona Distinctiveness metrics.The analysis also examines scoring trajectories and case-study reasoning across agents with diverse cultural backgrounds.
- Dynamic Behavioral Evaluation: Immigrant-background agents show higher acceptance of “newcomers” than native agents, while the Base LLM consistently shows the highest acceptance.USA_native and Russia_native agents remain below 50 across 40 rounds, maintaining cautious and skeptical stances.
- Interpretability of Value Reasoning: ExpertIVS anchors value preferences to individual profiles and produces logically coherent arguments, unlike baseline outputs that primarily list binary labels or generic value tags.The Japan_immigrant case links concerns about public safety and employment availability to the agent’s profile.
- Interpretability of Value Reasoning: Semantic distillation by sociological experts captures the intrinsic logic of individual value systems, enabling interpretable value reasoning in the case study.This contrasts with rigid and shallow simulation from superficial value labels.
6 Conclusion
ExpertIVS summarizes individual value systems through sociological expert agents and evaluates consistency during dynamic interactions with multi-agent debate. The framework is presented as enabling high-fidelity value reproduction and interpretable reasoning, while retaining important limitations in linguistic-style fidelity and safety alignment.
- ExpertIVS summarizes individual value systems through sociological expert agents.The framework is described as using expert agents to structure individual value systems.
- Multi-agent debate evaluates value consistency during dynamic interactions.The debate mechanism is designed to assess individual agents in interaction rather than only through static responses.
- The framework demonstrates high-fidelity value reproduction, improved value generalization, and interpretable reasoning over individual value systems.The conclusion reports experiments across 480 respondents from 12 countries and highlights leave-one-out generalization and debate-based interpretability.
- The study remains limited by a value–voice gap and tension between persona fidelity and safety alignment.Simulated individuals may default to standardized written expressions, while safety mechanisms can constrain high-fidelity simulation for aggressively aligned models or fringe groups.
E Results Without Demographics
Without demographic fields, the experiments report value restoration and leave-one-out value generalization, alongside a multi-agent debate procedure for evaluating consistency in dynamic interactions.
- Results Without Demographics: The demographic ablation removes all demographic fields from the reported value restoration and generalization experiments.Table 7 reports value restoration fidelity, while Table 8 reports leave-one-out value generalization; AVG averages 13 value dimensions and Average averages 12 countries.
- Multi-Agent Debate: The debate setup uses nine participants and one moderator discussing a single topic.The participants include eight simulated individuals and one unconditioned base model, while the moderator coordinates the process without contributing a stance.
- Stage 1: Opening Statements: Stage 1 collects independent opening statements, moderator summaries, and 0–100 agreement scores before free debate.Opening statements follow a fixed order, are initially kept independent, and are followed by participant agreement scoring.
- Stage 2: Free Debate: Stage 2 updates agreement scores after selected utterances and repeats the process for 40 debate rounds.The moderator selects speakers based on the ongoing discussion, and participants update scores using the current utterance and accumulated debate record.
H Detailed Experimental Settings in Multi-Agent Debate
This section reports the detailed experimental settings used for the Multi-Agent Debate evaluation.
- The paper provides detailed experimental settings for its Multi-Agent Debate procedure.
H.1 Agents and Model Assignment
The debate instantiates eight profile-conditioned individuals, one unconditioned base agent, and a separate moderator using Gemini 2.5-flash. Agents share accumulated public utterances while logging dialogue and agreement scores.
- Agents and Model Assignment: Gemini 2.5-flash is used to instantiate nine debate agents plus a separate moderator.The nine agents comprise eight profile-conditioned individuals and one unconditioned base agent.
- Agents and Model Assignment: Participant system prompts inject individual profiles, whereas the base agent receives no profile injection.The moderator uses a dedicated prompt enforcing neutrality, turn-taking, and faithful summarization.
- Agents and Model Assignment: The eight simulated participants represent native and immigrant individuals from the USA, Germany, Russia, and Japan.The listed profiles include one native and one immigrant respondent for each of the four countries.
- Shared Information and Logging: The shared public information pool contains accumulated utterances that all agents receive at the beginning of each speaking round.The experiment logs chronological utterances, agreement scores, topic IDs, and individual IDs.
I Evaluation Prompt
The evaluation prompt asks expert evaluators to judge whether simulated agents restore a target persona in dynamic discourse rather than remain generic assistants. The appendix also lists per-individual VAS and SSS evaluations across debate topics and evaluator models.
- I Evaluation Prompt: The prompt is used with Gemini 3 Pro and ChatGPT 5.2 Thinking as evaluators.
- I Evaluation Prompt: The appendix labels the evaluation prompt as VAS and SSS and includes ExpertIVS, IndieValue, and Simple Anthology entries.
- I Evaluation Prompt: It assigns the evaluator an expert psychologist and sociologist role specializing in profiling and psycholinguistic analysis.
- I Evaluation Prompt: The task is to evaluate the simulated agent’s restoration degree using its target profile and discourse speeches.
- I Evaluation Prompt: The evaluation determines whether the agent becomes the profiled person or remains a generic AI assistant.
- I Evaluation Prompt: Table 10 presents the evaluation prompt for assessing persona restoration in dynamic discourse.
- Individual VAS SSS VAS SSS VAS SSS VAS SSS: Table 11 reports per-individual VAS and SSS scores for the Accept Newcomers debate topic evaluated by Gemini 3 Pro.
- Individual VAS SSS VAS SSS VAS SSS VAS SSS: Table 12 reports the same Accept Newcomers per-individual measures evaluated by ChatGPT 5.2 Thinking.
K Human Evaluation Details
The study supplements LLM-based evaluation with anonymized human assessment and reports tables covering individual debate evaluations and averaged human scores.
- K Human Evaluation Details: Three evaluators with government, sociology, and IT backgrounds conducted the human evaluation for the Accept Newcomers setting.
- K Human Evaluation Details: Each human evaluator received USD 70 in shopping vouchers as compensation.
- K Human Evaluation Details: The method and IndieValueCatalog baseline were anonymized as Method A and Method B to avoid evaluation bias.
- K Human Evaluation Details: Table 13 lists per-individual VAS and SSS evaluations for Local or Universal assessed by Gemini 3 Pro.
- K Human Evaluation Details: Table 14 lists corresponding Local or Universal per-individual evaluations assessed by ChatGPT 5.2 Thinking.
- K Human Evaluation Details: Table 15 reports human evaluation scores for the Accept Newcomers setting averaged across three evaluators.