Source-linked AI summary

User Behavior Simulation with Large Language Model based Agents

Lei Wang, Jingsen Zhang, Hao Yang, Zhiyuan Chen, Jiakai Tang, Zeyu Zhang, Xu Chen, Yankai Lin, Ruihua Song, Wayne Xin Zhao, Jun Xu, Zhicheng Dou, Jun Wang, Ji-Rong Wen

arXiv:2306.02552v3cs.IRcs.AI

TL;DR

Reliable user behavior simulation is difficult because human decisions are intricate and real behavioral data are costly or constrained to obtain. The paper develops an LLM-based agent framework with a sandbox environment, and reports simulated behaviors close to those of real humans while reproducing information cocoons and conformity behaviors.

  • Problem

    Reliable user behavior data are fundamental to human-centered applications, but real data are difficult to acquire and existing simulators face limitations from simplified decision mechanisms and real-data dependence.

  • Method

    The paper designs an LLM-based agent framework with profile, memory, and action modules, supported by an intervenable and resettable sandbox for interactive behavior simulation.

  • Results

    The simulated behaviors are reported as very similar to real humans’, and the simulator reproduces information cocoons and user conformity behaviors.

  • Takeaways & Limitations

    The simulator provides a paradigm for studying user behavior and social phenomena in recommender systems and social networks.

  • Takeaways & Limitations

    The round-by-round execution discretizes time and restricts actions between rounds, while prompts may not be robust across different LLMs.

Abstract

from arXiv · show

Simulating high quality user behavior data has always been a fundamental problem in human-centered applications, where the major difficulty originates from the intricate mechanism of human decision process. Recently, substantial evidences have suggested that by learning huge amounts of web knowledge, large language models (LLMs) can achieve human-like intelligence. We believe these models can provide significant opportunities to more believable user behavior simulation. To inspire such direction, we propose an LLM-based agent framework and design a sandbox environment to simulate real user behaviors. Based on extensive experiments, we find that the simulated behaviors of our method are very close to the ones of real humans. Concerning potential applications, we simulate and study two social phenomenons including (1) information cocoons and (2) user conformity behaviors. This research provides novel simulation paradigms for human-centered applications.

Introduction

The paper addresses the difficulty of obtaining reliable human behavior data and the limitations of existing simulators that rely on real datasets and simplified decision models. It proposes an LLM-based agent framework and sandbox, finding simulated behaviors close to human behavior while reproducing information cocoons and conformity.

  • Real user data are often prohibitively expensive or ethically difficult to acquire because of commercial confidentiality and privacy concerns.
  • Existing simulators use simple decision functions that do not capture the intricate mechanisms of human cognition, potentially reducing behavioral reliability.
  • Traditional methods depend on real-world datasets, creating a “chicken and egg” problem and limiting simulation to patterns found in known datasets.
  • LLMs provide a basis for believable simulation because many user behaviors are expressed in language and behavior corpora from the web have been learned into these models.
  • The proposed framework combines profile, memory, and action modules with an intervenable, resettable sandbox, and its simulated behaviors are reported as close to real humans.

Results

RecAgent combines LLM-based profiles, memories, and actions in an intervenable sandbox, producing believable behaviors and supporting studies of information cocoons and conformity.

  • Simulator framework: RecAgent models each user with profile, memory, and action modules in a sandbox where agents interact with recommenders and one another.The simulator supports external intervention and participation by real humans.
  • Believability evaluation: 68% average improvement over the best baseline and an 8% gap below Real Human results indicate highly believable recommendation behavior.The comparison included Embedding, RecSim, and real-human selections across recommendation settings.
  • Believability evaluation: Most chatting and broadcasting believability ratings exceeded 4 out of 5, but all fell below 4 after 15 rounds.The authors speculate that accumulated memory burdens the LLM’s attention during longer simulations.
  • Memory mechanism: RecAgent’s memory outputs were close to human judgments, while removing short-term, long-term, or reflection components reduced informativeness or relevance.About 40% favored RecAgent for short-term summarization, 1.7% below non-expert humans; reflection support exceeded the comparison human result by about 3.3%.
  • Social phenomena: After about five rounds, recommendation entropy decreased by 8.5%, reproducing an information cocoon; randomness and added diverse friends improved entropy, especially together.The Rec-Strategy was more effective than the Soc-Strategy in the reported intervention, while combining them produced further improvement.
  • Social phenomena: Conformity emerged as initially dispersed movie scores concentrated at 6 and 7, with agents having more friends more likely to change scores.Agents exchanged opinions through private chats and broadcasts before rescoring the movie each round.

Discussion

The paper presents LLM-based agents as a new paradigm for simulating user behaviors, reporting similarity to real humans and applications to social phenomena. It also identifies temporal, adaptability, and prompt-robustness limitations while envisioning broader use across human-centered AI.

  • Discussion: The simulator extends LLMs with profile, memory, and action modules to model user behaviors.The framework supports simulation-based studies of information cocoons and user conformity behaviors.
  • Discussion: Simulated behaviors are reported to be very similar to those of real humans across extensive experiments.
  • Discussion: Discretized round-by-round execution restricts users from acting between rounds, reducing flexibility relative to real-world scenarios.
  • Discussion: The simulator lacks specific LLM fine-tuning for recommendation problems, and behavior prompts may not transfer robustly across different LLMs.The paper gives ChatGPT and GPT-4 as an example of models potentially requiring distinct prompts.
  • Discussion: A flexible interface could allow the simulator to improve with LLM development and extend beyond recommender systems and social networks.The authors position RecAgent as an example for other subjective simulation problems in human-centered AI.

Methods

The simulator uses LLM-based agents with profile, memory, and action modules, supported by a cognitive-neuroscience-inspired memory system and controllable human-agent interaction.

  • Agent framework: Each agent combines profile, memory, and action modules to simulate user behaviors.Profiles encode backgrounds and preferences; memory supports dynamic behavior; the action module produces specific behaviors.
  • Profiling module: Profiles encode user identity, demographics, traits, career, interests, and behavior features such as Watcher, Explorer, Critic, Chatter, and Poster.The behavior features describe feedback, search, criticism, conversation, and posting tendencies.
  • Memory module: The memory system comprises sensory, short-term, and long-term memory modeled on human memory mechanisms.Sensory memory processes observations, short-term memory mediates transfer, and long-term memory stores information reusable in similar or unseen situations.
  • Memory module: Sensory memory compresses observations and assigns importance and timestamps before producing records for later memory operations.Observations are represented as natural-language events, and more important information is more likely to be recalled.
  • Memory module: Repeatedly encountered similar short-term observations are enhanced and transformed into long-term memories after K enhancements.Similarity is computed from embeddings, and transformed records are summarized into higher-level insights.
  • Memory module: Memory reading combines top-N relevant long-term records with all short-term memories to capture general and recent preferences.Long-term records are retrieved using the current observation as a query.
  • Human-agent collaboration: Humans can act as agents and correct erroneous or hallucinatory virtual behaviors, creating trade-offs between accuracy and cost.Human-agent collaboration supports intermediate states between fully real and fully virtual execution.
  • System intervention: The simulation can be globally interrupted, questioned, paused, and modified to study emergency events or counterfactual behaviors.Users can interview agents or change factors such as profiles before resuming execution.

An example of the first step in sensory memory

The sensory-memory example compresses a dialogue about shared movie interests into a single natural-language observation emphasizing the participants’ preferences and reasons.

  • Prompt: An LLM prompt compresses a multi-turn movie dialogue into one independent sentence while emphasizing movie interests and reasons.The prompt also instructs the output to use third person when names appear and omit explicit profiles.
  • Compressed observation: The compressed output represents David Smith and David Miller as sharing enthusiasm for mind-bending movies and planning a movie night.It names films including Interstellar, Inception, The Matrix, Blade Runner, and The Prestige.

An example of the insight generation process in short-term memory

The insight-generation example transforms detailed movie-related memories into a higher-level characterization of the user.

  • Input memories: The stored memories describe David Miller’s interest in mind-bending movies and his search for recommendations.The memories include his enjoyment of Interstellar and Inception and his desire for further discussion and recommendations.
  • Insight generation: An LLM prompt asks for a one-sentence character insight that differs substantially from the original memories’ content and structure.The prompt uses both memory and observation inputs to generate an abstract characterization.
  • Generated insight: The generated insight characterizes David Miller as curious and open-minded because he actively seeks recommendations and discussions about mind-bending movies.The output abstracts specific movie interactions into a character-level description.

Efficiency analysis

The efficiency analysis examines how simulation time changes with agent count, API-key count, rounds, and behavior type, with costs rising as workloads accumulate.

  • Experimental design: The analysis studies time costs as the numbers of agents, API keys, and epochs change, and across different agent behaviors.For agent count, the experiment observes one simulator round with 1 to 500 agents while fixing API keys at 1.
  • Agent scaling: 220s for 10 agents and 1.9 hours for 100 agents show that per-round time increases sharply with agent count under one API key.All agents take actions in this experiment, whereas fewer active agents may reduce practical time cost.
  • Parallel invocation: Additional API keys lower time cost, indicating that parallel API-key invocation improves simulator efficiency.The reported results fluctuate substantially and have high variance, possibly because of unstable network speeds.
  • Round scaling: Time cost rises with an increasing acceleration rate as the number of rounds grows, possibly because accumulated information requires longer processing.This explanation is presented as a possibility rather than a confirmed mechanism.
  • Behavior costs: Friend chatting is the most expensive behavior because it requires generating more complex content.The simulator costs about 0.25 dollars per round for 10 agents using ChatGPT.

Examples of system intervention

The simulator supports interventions that alter agents’ profiles and can change their conversational behavior and recommendations accordingly. It also models long-tail activity distributions in real-world recommendation datasets.

  • Experimental setup: The intervention experiment compares an intervention branch with an original branch after five rounds of simulation.
  • Profile intervention: After changing David Smith’s preferences from sci-fi to family-friendly movies, his dialogue shifts toward family-friendly and romantic movie interests.
  • Active interviewing: Agents change their movie recommendations when their preferences are altered, and explain those recommendations during active interviews.
  • Activity modeling: By varying α, ppxq effectively models the long-tail activity distributions of MovieLens, Amazon-Beauty, Book-Crossing, and Steam.

Prompt Examples for different Agent Behaviors

RecAgent prompts combine an agent’s profile summary, reaction to observations, and instructions specific to the intended action.

  • RecAgent prompts contain a personal-profile summary, a reaction to the given observation, and action-specific instructions.

Summary

The summary component extracts profile information relevant to the agent’s current observation. The surrounding examples illustrate this process within the agent framework.

  • Summary extracts and condenses information from a user’s profile that is relevant to the current observation.
  • A profile-summary prompt asks the agent to identify details relevant to a conversation between David Miller and David Smith.
  • The resulting summary combines demographic traits, movie interests, system-related activities, standards, and a friendship relation.

Reaction

The reaction framework integrates summaries, memories, observations, and other context to produce individual or interpersonal agent behavior. Its examples show reactions formatted for one agent and for interactions between two agents.

  • The shared reaction prompt integrates summary, memory, observation, and related information for individual actions and two-agent dialogues.
  • An individual reaction example combines an agent’s summary with recent social-media information, recommender-system activity, and recent observations.
  • The framework distinguishes reactions between two agents by including summaries and recent movie-related information for both participants.

Action

The section specifies the actions agents can perform in recommender and social-media environments, together with prompts and formatted outputs for simulated behavior. Agents can choose among browsing, searching, watching, conversing, posting, or doing nothing.

  • Action specification: The section presents action prompts with examples of inputs and outputs for the agent environment.A separate recommender-action heading introduces the detailed interaction examples, including David Smith receiving recommendations and selecting Son of Flubber.
  • Available actions: Agents can enter the recommender system, enter social media, or do nothing.Entering the recommender system enables movie recommendations, watching, or searching; entering social media enables chatting or publishing posts.
  • Recommender actions: A recommender interaction shows David Miller searching for Interstellar and receiving five movie results.The example output then selects a search action for Inception rather than buying from the returned list.
  • Recommender actions: In the recommender system, agents choose one movie, view the next page, search for an item, or leave.The prompts specify formatted commands for buying one listed movie or searching for a title.
  • Social-media actions: The framework models social interaction through simulated dialogue between named agents, with one agent initiating the conversation.The dialogue prompt restricts agents from discussing movies they have not watched or heard about.
  • Social-media actions: Agents may publish one-line posts about recently watched recommender-system movies, such as David Miller recommending Inception.The posting instructions prohibit discussion of movies the agent has not watched or heard about.
Loading 2306.02552v3…