Source-linked AI summary
SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users
Xinnong Zhang, Jiayu Lin, Xinyi Mou, Shiyue Yang, Xiawei Liu, Libo Sun, Hanjia Lyu, Yihang Yang, Weihong Qi, Yue Chen, Guanying Li, Ling Yan, Yao Hu, Siming Chen, Yu Wang, Xuanjing Huang, Jiebo Luo, Shiping Tang, Libo Wu, Baohua Zhou, Zhongyu Wei
TL;DR
Existing social simulations face alignment challenges involving real-world environments, target users, interaction mechanisms, and behavioral patterns. SocioVerse addresses these challenges with an LLM-agent-driven world model, four alignment components, and a 10-million-person real-world user pool. Across political, news, and economic simulations, the framework reflects large-scale population dynamics, while remaining gaps appear in complex situations requiring contextual knowledge.
Problem
Existing social simulation methods face unresolved alignment challenges between simulated environments and real-world environments, users, interactions, and behavioral patterns.
Method
SocioVerse combines four modular alignment components with a 10-million-person real-world user pool to drive LLM-agent social simulations.
Results
Across three real-world scenarios, SocioVerse demonstrates large-scale social simulation, with demographic distributions and user histories significantly improving simulation accuracy.
Takeaways & Limitations
The results support standardized, diverse, and representative social simulations, while showing that LLMs can simulate human responses in complex social contexts.
Takeaways & Limitations
LLMs underperform in complex situations requiring contextual knowledge, and the current framework implements only part of its proposed architecture.
Abstract
from arXiv · showhide
Social simulation is transforming traditional social science research by modeling human behavior through interactions between virtual individuals and their environments. With recent advances in large language models (LLMs), this approach has shown growing potential in capturing individual differences and predicting group behaviors. However, existing methods face alignment challenges related to the environment, target users, interaction mechanisms, and behavioral patterns. To this end, we introduce SocioVerse, an LLM-agent-driven world model for social simulation. Our framework features four powerful alignment components and a user pool of 10 million real individuals. To validate its effectiveness, we conducted large-scale simulation experiments across three distinct domains: politics, news, and economics. Results demonstrate that SocioVerse can reflect large-scale population dynamics while ensuring diversity, credibility, and representativeness through standardized procedures and minimal manual adjustments.
1 Introduction
SocioVerse addresses four alignment challenges in social simulation: keeping environments current, matching target users, standardizing interactions, and reproducing group behavior. It combines modular engines with a 10-million-person real-world user pool and evaluates the framework across three scenarios.
- Traditional social research methods face high costs, limited sample sizes, and ethical concerns, motivating alternative approaches to studying human behavior.
- Existing social simulations struggle to align simulated environments, agents, interaction mechanisms, and behavioral patterns with the real world.
- SocioVerse uses modular Social Environment, User Engine, Scenario Engine, and Behavior Engine components to address these alignment questions.
- A 10 million-person real-world user pool supports diverse, large-scale simulations through sampling strategies that extract target groups for customized tasks.
- The framework is evaluated through presidential election prediction, breaking news feedback, and national economic survey simulations compared with real-world situations.
2 Methods
SocioVerse implements a four-part pipeline that aligns social context, users, interaction structures, and behavior. Its engines combine updated information, real-world user data, scenario templates, and agent-based or LLM-based behavior models.
- Social Environment: The Social Environment aligns simulations with real-world conditions by supplying social structure, time-sensitive dynamics, and personalized context.
- User Engine: The User Engine retrieves and describes target users using a large social-media-derived pool and extensive user labels.
- User Engine: The 10-million-user pool aggregates filtered social-media footprints, while demographic labels are refined through LLM annotation, human review, and trained classifiers.
- Scenario Engine: The Scenario Engine maps task formulations to questionnaire, interview, behavior-experiment, and social-media-interaction templates with different interaction structures.
- Behavior Engine: The Behavior Engine combines user history, scenario interaction mechanisms, and social context to predict individual behavior using traditional agent-based models and LLMs.
3 Implementation of Specific Scenarios
The framework is instantiated in political, news, and economic simulations using standardized pipelines and domain-specific target groups, questionnaires, and evaluation metrics. These scenarios compare simulated responses with elections, user attitudes, or official economic statistics.
- Three scenarios cover U.S. presidential election prediction, breaking news feedback, and a national economic survey of China.
- Presidential Election Prediction: The election simulation models demographic and ideological diversity using Census Bureau and ANES data, then evaluates state outcomes with accuracy rate and RMSE.
- Breaking News Feedback: The breaking-news task samples technology-interested Rednote users while preventing pre-news context leakage, then compares attitudes with a ChatGPT-discussion ground-truth set.
- Breaking News Feedback: Breaking-news feedback is evaluated using NRMSE for point-wise answers and KL-divergence for six-dimensional answer distributions.
- National Economic Survey: The economic survey samples nationwide agents by regional population and income distributions, evaluating spending values with NRMSE and eight-item distributions with KL-divergence.
4 Results
Across three scenarios, SocioVerse evaluates politics, breaking news, and economics against real-world outcomes using standardized scenario settings. Results show strong alignment overall, while user distributions, real-world knowledge, model choice, and task-specific biases affect precision.
- Overall Evaluation: The evaluation covers presidential election prediction, breaking news feedback, and a national economic survey, with overall and selected subset results reported.Table 2 identifies the three scenario configurations, while Table 3 reports their overall results and relevant subsets.
- Presidential Election Prediction: Over 90% of state voting results are typically predicted correctly under the winner-takes-all rule, indicating high-precision macroscopic reproduction of election outcomes.GPT-4o-mini and Qwen2.5-72b show competitive Acc and RMSE performance, including for battleground states.
- Breaking News Feedback: GPT-4o and Qwen2.5-72b align best with real-world breaking-news reactions on KL-Div and NRMSE, respectively.The evaluation measures consistency between simulated reactions and real-world users’ attitudes.
- National Economic Survey: All models closely align with real-world economic statistics, with Llama3-70b performing significantly better overall and developed-region subsets outperforming overall results.Spending habits are reported as accurately reproduced, especially in developed regions.
- Overall Findings: The framework supports diverse and accurate massive simulations through a standard pipeline with minimal human-expert changes, but underlying LLM choice affects precision across scenarios.Breaking-news simulations also show generally more disagreeable simulated responses than real responses, indicating a potential public-opinion bias risk.
- Ablation Analysis: Prior demographic distributions improve election accuracy, while real-world user posts improve fine-grained performance, especially for Llama3-70b in Acc and all models in RMSE.The ablation compares prior demographic distributions with random distributions and evaluates the contribution of users’ past social-media posts.
5 Discussion
SocioVerse shows that LLM-based social simulation can model responses across three real-world scenarios, while revealing important limits in complex, context-dependent situations and opportunities for further development.
- Demographic distributions and users’ historical experiences significantly improved simulation accuracy.The findings support a large, demographically rich user pool complemented by multidimensional user tagging for modeling group-specific behaviors.
- LLMs produced broadly similar simulations of human attitudes and ideologies under consistent measurement protocols.GPT-4o-mini showed notable inconsistencies, indicating that model-specific preferences or biases remain influential.
- LLMs perform well in simple daily scenarios but underperform in complex situations requiring contextual knowledge.This pattern underscores the need to align model behavior with real-world experiences and social contexts.
- The current version implements only part of the framework, leaving potential to enhance simulation accuracy and quality.Future directions include refining module collaboration, incorporating up-to-date social environments, expanding scenario formats, adapting models to complex user groups, and adding autonomous analysis planning.
- SocioVerse could provide social scientists with a cost-effective, scalable platform for social experiments with minimal setup.The framework is positioned for analyzing theories and exploring large-scale social impacts and long-term changes in virtual societies.
A.1 Content Data Extraction
The content extraction procedure collects only post-related content from social media platforms to avoid violating privacy policies.
- Only post-related content is extracted from all social media platforms to avoid violating privacy policies.The data collected from each platform is listed in Table 6.
A.2 Abnormal Data Filtering
The filtering procedure removes abnormal users based on textual repetition, while annotation resources are sampled and varied across data sources to reduce costs.
- Users whose textual repetition ratio surpasses 0.3 are considered likely robots or advertisers and filtered.The ratio is calculated from all textual content associated with the same user.
- A subset of the user pool is sampled and annotated with multiple powerful LLMs to save costs.The LLMs used vary by source, including X and Rednote users.
B.2 Human Evaluation
Human annotators verify LLM-generated demographic annotations, with each record checked by at least two annotators and consistency reported across models.
- Seven professional human annotators re-annotate demographic factors without seeing LLM labels.Their results are used to verify the LLM annotations.
- Every datum is verified by at least two human annotators.Table 7 reports consistency between human judgments and different LLMs.
B.3 Classifier Training
The classifier pipeline constructs a training dataset using majority-voted LLM labels and platform-specific language models, then reports demographic-classifier performance on a test set.
- Majority-voted labels from multiple LLMs form the demographic-classifier training dataset.LongFormer is used for X data, while Bert-base-chinese is used for Rednote data.
- LongFormer processes X data, whereas Bert-base-chinese processes Rednote data.
- Table 9 reports demographic-classifier performance for each demographic factor on the test set.
B.4 Overall Distribution of the User Pool
The demographic classifiers annotate the user pool to produce overall demographic distributions, with Figure 5 presenting distributions for X and Rednote users.
- Demographic classifiers annotate all users in the user pool to obtain overall demographic distributions.
- The annotation pipeline uses GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro for X users.
- For other demographics not included in prior distributions, only sampled users are annotated using majority votes from LLMs.
- Figure 5 presents demographic distributions for the X and Rednote user pools.
C.1 Iterative Proportional Fitting
The section describes procedures for matching simulated populations to target distributions, including IPF, direct sampling, regional income mixtures, and simulation questionnaires.
- Iterative Proportional Fitting: IPF iteratively adjusts a two-way table so estimated marginals approach targeted row and column totals.The iterations stop when estimated marginals are sufficiently close to real marginals or stabilize.
- Iterative Proportional Fitting: 888 of 918 presidential-election marginals fall within 5% of actual values despite frequent IPF nonconvergence.The resulting estimated joint distribution and marginals are then used for massive simulation.
- Identical Distribution Sampling: Identical distribution sampling draws agents from an available joint distribution of demographic features.For breaking-news feedback, the ground-truth demographics provide the joint distribution needed for direct sampling.
- National Economic Survey: Regional income priors combine a log-normal distribution for lower and middle incomes with a Pareto distribution for high-income groups.The mixture typically assigns 90% of the population to the log-normal component and 10% to the Pareto component.
- Simulation Questionnaires: The questionnaire materials cover presidential voting and policy preferences, ChatGPT cognition and perceived risks, and household spending and financial pressure.The economic questions include food, clothing, housing, services, transportation, education, entertainment, healthcare, and insurance.