Source-linked AI summary
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
Preethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde, Sameer Singh, Seraphina Goldfarb-Tarrant
TL;DR
Agentic benchmarks increasingly use LLM-simulated users, but whether they robustly and fairly represent real users remains unclear. Through a diverse user study on τ-Bench retail tasks, this paper finds that simulated-user evaluations miscalibrate agent performance and represent populations unevenly. The findings expose risks of misrepresenting agent capabilities across diverse users and motivate more robust, human-validated evaluation.
Problem
Prior agentic benchmarks largely overlook validation against real human interactions and variation across user populations, leaving the robustness, validity, and fairness of simulation uncertain.
Method
The paper conducts a user study across the United States, India, Kenya, and Nigeria, comparing simulated and human users on difficulty-balanced τ-Bench retail tasks.
Results
Simulated-user evaluations lack robustness, misestimate agent performance across task difficulty, show demographic disparities, and introduce conversational artifacts relative to human users.
Takeaways & Limitations
Current user-simulation practices can misrepresent agent capabilities across diverse populations and obscure real-world deployment challenges.
Takeaways & Limitations
The evaluation is conducted entirely in English and focuses on a single retail domain, limiting conclusions about multilingual settings and other domains.
Abstract
from arXiv · showhide
Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. Through a user study with participants across the United States, India, Kenya, and Nigeria, we investigate whether LLM-simulated users serve as reliable proxies for real human users in evaluating agents on τ-Bench retail tasks. We find that user simulation lacks robustness, with agent success rates varying up to 9 percentage points across different user LLMs. Furthermore, evaluations using simulated users exhibit systematic miscalibration, underestimating agent performance on challenging tasks and overestimating it on moderately difficult ones. African American Vernacular English (AAVE) speakers experience consistently worse success rates and calibration errors than Standard American English (SAE) speakers, with disparities compounding significantly with age. We also find simulated users to be a differentially effective proxy for different populations, performing worst for AAVE and Indian English speakers. Additionally, simulated users introduce conversational artifacts and surface different failure patterns than human users. These findings demonstrate that current evaluation practices risk misrepresenting agent capabilities across diverse user populations and may obscure real-world deployment challenges.
1 Introduction
Agentic benchmarks increasingly use LLM-simulated users to scale evaluation of multi-turn interactions, but their robustness, validity, and fairness remain uncertain. This study compares simulated and human users across diverse populations and finds systematic miscalibration, demographic disparities, and artificial conversational patterns.
- Motivation: LLM-simulated users enable scalable evaluation of sustained, context-aware agent interactions beyond static single-turn benchmarks.These benchmarks assess natural conversation, policy adherence, instruction following, and tool use across multiple turns.
- Research Questions: User simulation raises unresolved questions about consistency across user LLMs, validity as a proxy for real people, and fairness across populations.Without validation against actual users, simulated interactions may misrepresent agent capabilities through miscalibration.
- Study Scope: The study directly evaluates user simulation with participants from the United States, India, Kenya, and Nigeria.It uses τ-Bench retail tasks as a case study for comparing simulated and human users.
- Findings: User simulation lacks robustness and systematically misestimates agent performance across task difficulty levels.It underestimates success on the most challenging tasks and overestimates outcomes on moderately difficult scenarios.
- Findings: Simulated users show demographic biases, perform particularly poorly as proxies for AAVE speakers, and introduce heightened question-asking and politeness.These patterns challenge user simulation as a stand-alone evaluation paradigm.
2 Related Work
Prior work studies simulated-user conversational realism, fidelity, and demographic skews in language technologies, but this paper examines whether simulation reliably represents diverse human users in agentic evaluations.
- User Simulation: Prior user-simulation studies examine conversational characteristics, behavioral realism, human-rating alignment, and replication of human intermediate steps.Existing work spans math tutoring, daily planning, and online shopping interactions.
- Positioning: This work extends prior research by studying demographic representation and proxy validity in agentic evaluation rather than conversational realism alone.Its focus is whether simulated users reflect actual users across diverse populations.
- Demographic Skews: Research on demographic skews shows that model and dataset judgments vary systematically across race, gender, political affiliation, and other demographic axes.These differences include perceptions of safety, offensiveness, toxicity, and politeness.
- Demographic Skews: Sociodemographic prompting has produced mixed results, sometimes failing to improve alignment consistently or relying on harmful stereotypes.The cited literature motivates examining demographic representation directly rather than assuming prompting resolves it.
3 Methodology
The study adapts τ-Bench retail tasks, recruits diverse English-proficient participants, balances task difficulty, and compares agent outcomes with human and simulated users using success rate and calibration metrics.
- Benchmark: τ-Bench evaluates interactive customer-service tasks requiring tool calls, policy adherence, information gathering, and realistic database changes.Tasks involve collaboration between an agent and a user receiving specific objectives.
- Benchmark: Task success requires the correct final database state and responses containing all information specified in the instructions.The benchmark applies automated substring matching against ground-truth annotations.
- Task Sampling: 18 retail tasks were selected by difficulty from six GPT-4o-based success-rate levels, with three tasks per level.Difficulty levels range from 0/5 to 5/5 successful runs, or 0% to 100%.
- User Study: Participants from the United States, India, Kenya, and Nigeria completed randomized English-language tasks while the agent model remained GPT-4o.The study recruited primarily through Prolific, using snowball sampling in Nigeria.
- User Study: US participants were stratified by SAE versus AAVE dialect and by age, while participants from other countries were recruited only from ages 18–34.The study targeted approximately 40 participants per age × dialect/country group, except the AAVE 55+ group with 22 participants.
- Evaluation Metrics: Success rate is the percentage of tasks completed with all required actions and requested information correctly delivered.The same automated evaluation procedure is applied to human and simulated interactions, with averages computed across difficulty levels.
- Evaluation Metrics: ECEHuman–LLM measures weighted average absolute deviation between simulated-user and human-user success rates across difficulty levels.Lower values indicate better calibration, and perfect calibration equals 0.
4 Results
The results show that user-simulated evaluations are sensitive to the simulation model, miscalibrated against human users, and uneven across dialect, age, and country groups. Simulated and human interactions also produce different conversational behaviors and failure patterns.
- 4.1 Robustness: Nearly 9 percentage points separate success rates when only the user LLM changes, indicating limited robustness.With GPT-4o fixed as the agent, Sonnet 3.7 and Sonnet 4.5 produce this difference; the study recommends reporting multiple user models.
- 4.2 Validity: 45.2% success with US participants and ECEHuman–LLM = 15.1 show substantial miscalibration even in the expected strongest-alignment setting.The calibration gap reaches ECEHuman–LLM = 25.9 across the 1st and 4th difficulty bins.
- 4.3.1 Dialect and Age (United States): 50.6% success and ECEHuman–LLM = 11.7 for SAE contrast with 39.4% and 20.3 for AAVE, producing an 11.2-point performance decrease and 8.6-point ECE increase.Dialect disparities in performance also grow with age, while calibration patterns differ across age groups.
- 4.3.2 Countries: 41.0%-49.2% success rates across countries are less dispersed than within-US dialect and age differences, and cross-country differences are not statistically significant.The cross-country analysis is restricted to participants aged 18–34.
- 4.3.2 Countries: ECEHuman–LLM = 18.9 is reported for AAVE and Indian participants, compared with 13.0 for SAE, while simulated users overestimate performance on easy tasks.The findings indicate that simulated users are poorer proxies for some populations and can underestimate deployment difficulty across diverse users.
- 4.4 Interaction Patterns: Simulated users ask questions in 18.8% of user turns and use politeness indicators in 39.2%, versus 9.8% and 19.9% for human users.Differences are larger between simulated and human users than among human users.
- 4.4.2 Errors: 32.2% argument errors and 31.4% output errors occur with simulated users, while output errors are 12.2%-23.6% for human users.Simulated-user interactions involve fewer missing and extra actions but more output errors.
- 4.4.2 Errors: Agents account for 48.9% of failures in simulated conversations versus 24.5% in human conversations, indicating distinct failure patterns.Human interactions show more user-side errors, whereas simulated interactions place greater burden on agent execution.
5 Discussion and Conclusion
The paper finds that LLM-simulated users can misestimate agent performance and obscure disparities across real user populations. These artifacts suggest that benchmark robustness does not necessarily generalize to users whose communication styles are underrepresented by simulation.
- LLM-simulated users may misestimate agent performance for actual users and obscure demographic disparities.
- Heightened politeness and question-asking in simulated interactions may not generalize across diverse user groups.
Limitations
The study’s conclusions are bounded by its English-only evaluation, single retail domain, and use of one fixed agent. These choices isolate user-simulation effects but limit generalization across languages, domains, and agents.
- English-only evaluation limits assessment of LLM-simulated users in multilingual settings.The authors note that language can influence user and agent behavior, capabilities, and interaction norms.
- The study evaluates only retail customer-service scenarios from τ-Bench, so simulation quality and agent-performance patterns may differ in other domains.Longer interactions and greater cultural or stylistic variation could make differences more apparent elsewhere.
- A single fixed agent, GPT-4o, enables isolation of user-simulation effects but leaves variation across agents unexamined.The authors identify cross-agent calibration gaps and performance disparities as important for understanding robustness, validity, and fairness.
Ethics Statement
Participants received standardized compensation, informed-consent information, and the option to withdraw without penalty. The study was classified as minimal risk because it involved simulated retail interactions without sensitive personal-data collection, while the authors note risks from demographic bias in simulated-user evaluation.
- Participants received standardized compensation regardless of country and could withdraw at any point without penalty.They were also given an overview of study procedures upfront and provided informed consent by choosing to participate.
- The study was classified as minimal risk because it used simulated retail customer-service interactions without collecting sensitive personal information.
- Adopting biased simulated users as a standard evaluation practice could systematically underserve certain demographic groups.The authors frame this as a broader societal implication of unreliable proxy behavior and demographic bias.
Acnkowledgements
The authors acknowledge contributors who piloted the user study and provided feedback, as well as the UCI NLP community. The work was conducted primarily during an internship and received support from HPI and an NSF CAREER award.
- Several contributors piloted the user study and provided feedback.
- The authors thank UCI NLP members for helpful discussions and comments.
- The work was conducted primarily during Preethi’s Cohere internship and supported by HPI and NSF CAREER award IIS-2046873.
A.1 Design Choices
The study holds the agent model constant to isolate effects of user variation, while recognizing that its conclusions do not cover differences across agent capabilities or benchmarks.
- Agent model: GPT-4o is used as the agent model throughout, isolating the effect of user variation while holding agent capability constant.The design attributes observed performance differences to user simulation rather than differences in agent performance.
- Agent model: The authors state that the core questions about robustness, validity, and fairness are agent-agnostic.
- Scope boundaries: The study cannot assess whether these issues vary across agents with different capabilities.The authors identify this as a limitation and call for comparisons involving weaker and stronger agents.
- Scope boundaries: The analysis uses τ-Bench as a case study, so benchmarks with different tasks, metrics, or prompting strategies may show different degrees of these issues.A comprehensive investigation across multiple benchmarks remains future work.
A.2 Dialect Screening
The study screens participants by dialect and evaluates human–simulation differences using repeated-measures models, controlled task instructions, and targeted prompting interventions. These design choices reveal that behavioral and demographic prompting can change task difficulty, success, and calibration patterns.
- Dialect screening: US participants are retained as White/SAE or Black/AAVE based on self-identified race and primary English dialect.Screening distinguishes mostly Standard American English from mostly African American Vernacular English, while excluding mixed or uncertain responses.
- Statistical analysis: Generalized Estimating Equations model binary task success while accounting for four repeated tasks per participant and clustering observations by participant ID.Models include demographic, experience, usage, and task-difficulty covariates, with separate age-stratified analyses for dialect effects.
- Task setup: Task instructions remove names and behavioral cues, replacing them with anonymized IDs to avoid identity- and behavior-based bias.Human participants interact through a Streamlit chat interface and are instructed to complete all requests while behaving naturally.
- Behavioral intervention: 11 of 18 tasks shift difficulty bins after limiting simulated-user politeness, while overall success decreases from 50.0% to 46.7%.The intervention uses GPT-4o as both agent and user LLMs and recomputes task bins from the updated distribution.
- Behavioral intervention: Calibration improves for AAVE, Indian, and Kenyan participants but worsens for SAE and Nigerian participants after reduced-politeness prompting.The comparison holds human-study results fixed while recomputing task difficulty bins from the revised simulated-user results.
- Demographic intervention: Country prompting yields success rates of 55.6 ± 6.1% for India, 51.1 ± 6.5% for Kenya, and 47.8 ± 4.4% for Nigeria across five runs.Calibration changes are mixed: Indian participants improve by −7.9, Nigerian participants change by 0.5, and Kenyan participants worsen by 8.6.