Source-linked AI summary
Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model
Ross Williams, Niyousha Hosseinichimeh
TL;DR
The paper asks how sensitive generative-agent behavior is to prompt wording, contextual changes, and persona names. Using a generative-agent epidemic model and mobility-curve comparisons, it finds that nonsynonymous and contextual changes affect outcomes, while close synonyms and persona names generally do not. The study presents this as an initial framework for prompt sensitivity analysis, while noting that its tested prompt space is limited.
Problem
Generative-agent research lacks a unified approach to prompt sensitivity analysis despite the large space of linguistic expressions that can influence behavioral responses.
Method
The study modifies semantic wording, agent contexts, and persona names in a generative-agent epidemic model, then compares mobility curves with baseline results using multilinear regression.
Results
Synonymous phrasing and persona-name changes produced no significant deviations, whereas nonsynonymous phrasing and contextual shifts changed the model’s mobility outcomes.
Takeaways & Limitations
The findings provide an initial basis for prompt sensitivity analysis and practical guidance for crafting prompts in generative agent-based models.
Takeaways & Limitations
The study cannot cover all prompt modifications and tests only one very close synonymous pair, so more distant synonyms could still change agent behavior.
Abstract
from arXiv · showhide
As generative AI gains traction, researchers are investigating its potential to serve as proxies for humans. From undergoing cognitive psychology experiments to experiencing an epidemic, generative agents, agents powered by generative AI models, produce realistic human behavior when prompted. This study explores the sensitivity of these generative agents' behavior to prompt modifications and varied persona names of the agents. To assess this sensitivity, we use a generative agent epidemic model, wherein each agent is prompted daily on whether it wants to isolate or commingle with other agents. We found that using synonymous prompts results in negligible changes to the model's outcomes. However, minor variations in prompts, as well as contextual changes, do influence the model's results. Lastly, our data indicates that different persona names assigned to generative agents, specifically those imbued with personas, do not significantly impact epidemic outcomes.
1. Introduction
Generative agents are increasingly used to simulate human behavior, but their responses may depend on how prompts are worded and contextualized. This paper uses an epidemic model to examine semantic, contextual, and persona-name sensitivity.
- Generative agents: Generative agents use large language models to simulate realistic human behavior in dynamic social-science models.These models receive prompts and generate responses that guide agent behavior over time.
- Research gap: Existing generative-agent models lack robust sensitivity analysis for prompt wording despite the vast space of linguistic expressions.Prompt sensitivity supplements traditional parameter sensitivity because different expressions can convey the same meaning.
- Research gap: The epidemic model provides a comparatively stable prompt and a measurable mobility-curve outcome for testing prompt sensitivity.Mobility curves relate current mobility to the percentage of new active cases in the previous time step.
- Study focus: The study tests whether synonymous wording, nonsynonymous wording, contextual shifts, and persona-name substitutions alter generative-agent behavior.The analyses compare agent responses and epidemic outcomes across these prompt or name modifications.
- Study focus: Synonymous prompts produced no statistically significant deviations, whereas nonsynonymous phrasing and contextual shifts did; persona names produced no significant deviations.These findings are reported from the study’s experiments using the generative-agent epidemic model.
2. Materials and Methods
The study modifies the epidemic model’s prompts and agent names across semantic, contextual, and persona-name experiments. It compares resulting agent mobility with the baseline using the model’s daily isolation decisions and mobility curves.
- Epidemic model: Each time step, every agent receives a prompt containing static personal information and dynamic health information before deciding whether to stay home.The model uses self and societal health feedback to elicit decisions about isolation or commingling.
- Sensitivity analyses: The study conducts semantic variation, contextual shift, and persona-name sensitivity analyses.These analyses respectively alter wording, revise the agent’s scenario, or substitute assigned names.
- Sensitivity analyses: Experiment #3 changes the work-and-savings context so agents already have enough savings but still go to work to earn money.This modification tests whether a different environmental circumstance changes behavior.
- Sensitivity analyses: Experiments #4 and #5 assign alternative name sets to test sex-bias and individual-name effects.Experiment #4 uses the 100 most popular female American names; experiment #5 assigns Jose to 50 agents and Maria to 50.
3. Results
The results show that close synonymous substitutions and persona-name changes leave epidemic outcomes statistically unchanged, while nonsynonymous wording and contextual shifts alter mobility relative to baseline.
- Statistical comparisons: Experiments #1, #4, and #5 were not statistically significant relative to baseline.These correspond to the synonymous semantic variation and persona-name sensitivity modifications.
- Semantic variation: Experiment #2 produced a different mobility outcome: agents who “learns” about the epidemic were more risk-averse than baseline agents.The “learns” run exhibited reduced mobility compared with the “knows” run.
- Contextual shift: Experiment #3 produced a different mobility curve from baseline after changing the agents’ financial context.Agents with enough savings displayed less risk aversion and were more inclined to work than agents who had to support themselves.
- Persona names: Persona-name experiments #4 and #5 did not significantly alter the model’s results.The tested name substitutions therefore showed no statistically significant epidemic-outcome deviation from baseline.
4. Discussion
The study finds that prompt changes can alter generative-agent epidemic outcomes, although synonymous wording and persona names generally leave results unchanged. It presents prompt sensitivity analysis as an initial framework for assessing model robustness while noting substantial coverage and computational limits.
- Semantic variation: Synonymous phrases produced no discernible change in mobility, whereas nonsynonymous semantic variations reduced mobility in the “learns” run relative to the “knows” run.The comparison used “knows about” versus “is aware of” for synonymous wording and contrasted the “learns” and “knows” runs for nonsynonymous variation.
- Contextual shift: Contextual shifts produced departures from baseline behavior and, contrary to the initial hypothesis, resulted in more mobile agents.The authors hypothesize that the model may assume voluntary workers enjoy going to work, but state that this explanation is uncertain.
- Persona name sensitivity: Persona names did not statistically influence model results when coupled with personality traits, indicating no major effect on simulated epidemic outcomes.This finding supports the authors’ assertion that persona names do not substantially affect the model’s epidemic outcomes.
- Contribution: The paper introduces a first step toward prompt sensitivity analysis for generative agent-based modeling, addressing a field that lacks unified prompting and sensitivity-analysis methods.The authors distinguish prompt sensitivity analysis from traditional parameter sensitivity analysis and identify prompt-space exploration as a methodological challenge.
- Contribution: The epidemic model’s general behavior responding to cases remains robust despite prompt-dependent result changes, potentially increasing epidemiological modelers’ confidence in the base model.The paper frames this conclusion as a contribution to epidemiology through robustness analysis of the generative agent epidemic model.
- Limitations: Prompt sensitivity analysis cannot cover all potential prompt modifications, while generative agent-based models are resource-intensive to run.The authors note that even close synonym pairs may not represent more distant alternatives and report that 100 agents across 50 time steps could take over two hours and cost approximately $2 per run as of June 2023.