Source-linked AI summary
MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
Ashutosh Hathidara, Julien Yu, Vaishali Senthil, Sebastian Schreiber, Anil Babu Ankisettipalli
TL;DR
User proxies offer scalable human-interaction data and evaluation, but naive prompting can produce unrealistic utterances, creating a need for principled human-likeness measurement. MirrorBench benchmarks proxy utterances against human users using complementary lexical and LLM-judge metrics while decoupling evaluation from task success. Across four datasets, it finds systematic proxy–human gaps, realism–diversity tensions, and sensitivity of judge scores to the chosen judge model.
Problem
User proxies are increasingly used for conversational evaluation and data generation, but naive prompting can produce unrealistic user behavior, motivating independent measurement of human-likeness.
Method
MirrorBench evaluates context-conditioned user-utterance human-likeness with MATTR, HD-D, Yule’s K, GTEval, Pairwise Indistinguishability, Rubric-and-Reason, and calibration controls.
Results
Across four public datasets, MirrorBench finds systematic gaps between proxies and human users, a regime-dependent realism–diversity tension, and judge-sensitive scores or orderings.
Takeaways & Limitations
Human-likeness benchmarking should use complementary metric families, calibration controls, and transparent multi-judge reporting when comparing user proxies.
Takeaways & Limitations
The benchmark is limited to four English-centric datasets, fixed assistant configurations with limited randomization, and possible bias from judge models and intent summarization.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning data. However, naive "act-as-a-user" prompting often yields verbose, unrealistic utterances, motivating principled evaluation of *user proxy agents*. We present **MirrorBench**, a reproducible and extensible benchmarking framework that evaluates user proxies solely on their ability to produce human-like user utterances across diverse conversational regimes, explicitly decoupled from downstream task success. **MirrorBench** combines three lexical-diversity metrics (**MATTR**, **Yule's~$K$**, and **HD-D**) with three LLM-judge-based metrics (**GTEval**, **Pairwise Indistinguishability**, and **Rubric-and-Reason**), and contextualizes judge scores using Human-Human and Proxy-Proxy calibration controls. Across four public datasets, **MirrorBench** yields variance-aware comparisons and reveals systematic gaps between user proxies and real human users. The framework is open sourced at https://github.com/SAP/mirrorbench and includes a command-line interface for running and managing user-proxy benchmarking experiments.
1 Introduction
MirrorBench addresses the need for scalable, realistic user-proxy evaluation by measuring human-likeness independently of downstream task success. It combines complementary lexical-diversity and behavioral-realism metrics across conversational settings, revealing systematic proxy–human gaps and judge sensitivity.
- Motivation: User proxy agents are LLMs prompted to emulate specified human personas for scalable automated testing and post-training data generation.They are increasingly used for regression testing, tool-use and grounding evaluation, synthetic dialogue generation, and stress testing.
- Motivation: Naive “act-as-a-user” prompting often produces verbose, overly cooperative utterances that diverge from real user behavior.This motivates measuring proxy human-likeness separately from downstream task success.
- Benchmark framing: MirrorBench evaluates whether proxy utterances match real human turns in the same task context, rather than measuring an unconstrained notion of general human behavior.The framework treats human-likeness as an utterance-level property conditioned on conversational regime.
- Benchmark design: MirrorBench combines human-anchored lexical-diversity statistics—MATTR, HD-D, and Yule’s K—with LLM-judge metrics for behavioral realism.The judge metrics include GTEval, Pairwise Indistinguishability, and Rubric-and-Reason, evaluated over four open conversational corpora.
- Findings: Across datasets, strong LLM proxies show systematic gaps from real users, including a recurring realism–diversity tension that varies by conversational regime.Judge-based realism can remain high while lexical or distributional characteristics under-match humans in clarification-centric settings.
- Findings: Judge choice can shift absolute scores and sometimes fine-grained proxy orderings, supporting transparent multi-judge reporting and calibration controls.MirrorBench uses human–human and proxy–proxy controls to contextualize judge scores.
2 Related Work
Prior work has used simulators and LLM judges for dialogue and agent evaluation, but generally does not isolate whether the simulated user itself is human-like. MirrorBench builds on these strands while targeting user-role realism directly.
- User simulation and user proxies: Early user simulators were typically goal-driven or rule-based, whereas LLM-based proxies support more open-ended conversational variability.LLM-based proxies also introduce verbosity, role drift, and inconsistent intent adherence.
- User simulation and user proxies: Broader agent frameworks evaluate agents in environments but do not isolate the human-likeness of the user role.This leaves the realism of simulated users distinct from overall agent performance.
- User-proxy evaluation: SimulatorArena evaluates whether profile-conditioned simulators can replace human judges for ranking assistants, while SimUSER measures behavioral alignment in recommender-system interactions.These approaches treat simulators primarily as downstream evaluation tools rather than objects of human-likeness assessment.
- LLMs as judges: LLM-as-a-judge methods provide scalable dialogue assessment through ratings or comparisons, but most setups focus on assistant outputs rather than the user role.MirrorBench repurposes judge-based evaluation to assess user-proxy realism.
3 MirrorBench
MirrorBench evaluates user-proxy human-likeness through controlled synthetic rollouts against reference conversations, combining calibrated judge-based realism and human-anchored lexical-diversity metrics. Its protocol standardizes datasets, goals, user-side outputs, aggregation, and robustness-aware comparison.
- Benchmark Protocol: MirrorBench synthesizes proxy–assistant dialogues from diverse reference datasets, using goals and explicit role separation before scoring the resulting transcripts.Goals are taken from annotations when available or generated from reference dialogues otherwise; the assistant is conditioned on reference history to preserve the same conversational trajectory.
- Benchmark Protocol: The evaluation compares proxy-generated and reference dialogues under identical context while scoring only proxy user behavior, tone, and style.Assistant response content is excluded from the human-likeness target.
- Aggregation & Reporting: MirrorBench reports per-metric means, standard deviations, and two-sided 95% confidence intervals for comparisons across proxies, datasets, and seeds.The evaluation record also includes per-sample values, model identifiers, random seeds, telemetry, and metric reliability diagnostics.
- Datasets & Tasks: The benchmark uses 795 conversations from QULAC, ClariQ, OASST1, and ChatbotArena, with stratified sampling and dataset-specific coverage criteria.Conversations are normalized to alternating English user–assistant turns with at least two turns before stratification.
4 Experiments & Results
Across four conversational datasets, MirrorBench compares user proxies with lexical-diversity and judge-based realism metrics, finding systematic, regime-dependent gaps from human users. Results also show judge sensitivity, generally stable rankings under assistant and seed changes, and a need to evaluate both realism and diversity.
- Experimental setup: MirrorBench evaluates user-proxy human-likeness across ChatbotArena, ClariQ, OASST1, and QULAC using MATTR, Yule’s K, HD-D, GTEval, PI, and RNR.The experiments also assess judge reliability, proxy behavior reproduction, and computational cost, latency, and throughput.
- Proxy comparison: Gemini-2.5-Pro and Claude-4-Sonnet are the most human-like by judge criteria across datasets, with GPT-4o generally competitive but behind.On ClariQ and QULAC, the two leading proxies approach the Human–Human ceiling under RNR, while PI win margins are positive.
- Proxy comparison: Lexical diversity depends strongly on dataset regime: proxies exceed human diversity on ClariQ, fall below it on QULAC, and show smaller deviations on ChatbotArena and OASST1.Gemini-2.5-Pro and GPT-4o provide the strongest overall diversity alignment, while Claude-4-Sonnet, GPT-5, and GPT-OSS-120B often differ more from human baselines.
- Proxy comparison: High judge realism does not guarantee human-level diversity, showing that judge-based realism and lexical diversity capture partially decoupled facets of proxy quality.Claude-4-Sonnet and Gemini-2.5-Pro lead on judge metrics but undershoot diversity on QULAC, whereas GPT-4o shows more stable diversity with moderate judge-realism gains.
- Judge sensitivity: Judge choice substantially changes realism scores: GTEval spans approximately 0.45–0.81, PI varies from near-zero or negative to clearly positive deltas, and RNR ranges approximately 0.79–0.98.The reported sensitivity ordering is PI > GTEval > RNR when the assistant and proxy are fixed to GPT-4o.
- Judge sensitivity: GTEval aligns strongly with human judgments while PI shows moderate correlation, with GTEval ρ ranging from 0.607 to 0.697 and PI ρ from 0.532 to 0.671.All reported correlations have p < 0.001, supporting consistent judge–human alignment across evaluated conditions.
- Robustness: Assistant swaps usually produce only slight score changes and preserve proxy ordering, with largely stable ranks and one GTEval crossover.Multi-seed runs likewise show nearly flat metric trajectories and consistent proxy ordering.
- Robustness: Across 12 additional multi-seed conditions, standard deviation remains ≤0.016 and proxy rankings are stable in 11 of 12 cases.The sole exception is a near-tie on OASST1 with Δ = 0.004.
5 Conclusion & Future Work
MirrorBench benchmarks whether user-proxy utterances resemble real human users across conversational tasks, independently of downstream task success. Its analyses expose systematic proxy–human gaps and realism–diversity trade-offs, while defining scope boundaries for future evaluation.
- MirrorBench evaluates user-proxy human-likeness across four public datasets using judge-based realism and human-anchored lexical diversity metrics.
- The benchmark is explicitly independent of downstream task success and intended to support systematic comparison of user proxies.
- MirrorBench reveals systematic gaps between proxies and human users, including a realism–diversity trade-off that varies by dataset regime.
- Judge-based metrics may reflect model-family bias and prompt sensitivity despite Human–Human and Proxy–Proxy calibration controls.
- The current evaluation is limited to four English-centric datasets, primarily uses a fixed assistant configuration, and evaluates fixed reference trajectories.
A Extended Quantitative Results
Extended analyses test whether MirrorBench’s judge correlations and multi-seed stability generalize across datasets and proxy models. They report consistent human alignment and low GTEval variability across the evaluated conditions.
- GTEval Spearman correlations range from 0.607 to 0.697 across datasets and proxy models, while PI correlations range from 0.532 to 0.671.All reported correlations have p < 0.001.
- Human–judge alignment remains consistent across ChatbotArena, OASST1, and ClariQ and across Gemini-2.5-Pro, Claude-4-Sonnet, and GPT-4o proxies.
- GTEval standard deviation is ≤0.016 in all 12 multi-seed conditions across the four datasets.The extension uses five seeds for ChatbotArena and three for each remaining dataset.
B Goal-Conditioning Ablation: Qualitative Examples
Qualitative examples show that full goal conditioning preserves the user’s concrete objective, whereas topic-only conditioning produces fluent but generic related questions. This distinction explains why PI can remain high while GTEval falls.
- B Goal-Conditioning Ablation: Qualitative Examples: Goal_topic produces fluent, human-sounding turns but can pursue a plausible related question instead of the user’s actual intent.
- B Goal-Conditioning Ablation: Qualitative Examples: The divergence between high PI and lower GTEval under goal_topic separates surface-level fluency from adherence to the concrete conversational objective.
- B Goal-Conditioning Ablation: Qualitative Examples: Goal_full keeps the proxy on the user’s specific request, while goal_topic deflects to a generic related question.
C.1 Architecture
MirrorBench uses a six-layer architecture that separates infrastructure from evaluation logic. Lower layers provide execution and data services, while upper layers host pluggable components, task logic, interfaces, and reporting.
- MirrorBench is organized as a six-layer stack separating infrastructure from evaluation logic.
- The lowest layer manages execution and data, including execution backends and persistence.
- Execution backends accept decomposed task units and support synchronous or asynchronous evaluation jobs through extensible interfaces.
- Persistence stores runs, units, episodes, execution duration, metric values, telemetry, and aggregate statistics in SQLite.
- Layer 2 provides backend-agnostic data models, registries, and configuration-management components for downstream modules.
- Configuration and manifest management decomposes jobs into self-contained executable units for parallelization across asynchronous and remote backends.
C.1.3 Orchestration.
The orchestration layer coordinates evaluation planning, execution, observability, and persistence through modular components. Standardized interfaces, registries, caching, and unified model and agent adapters support extensible and reproducible runs.
- Five orchestration modules manage component lifecycles, execution planning, and runtime operations.
- The core component layer provides pluggable evaluation logic through standardized interfaces and registry-based registration.
- Model clients unify provider SDKs, capture token, latency, and cost telemetry, and support transparent response caching.
- User-proxy adapters normalize message formats, tool usage, and execution models into a standardized AgentSession interface.
- Datasets are loaded from diverse sources into standardized episode objects with splitting, sampling, and reference-statistics metadata.
- Metrics combine lexical-diversity and LLM-as-judge measurements while declaring compatibility requirements such as task type and reference-statistics needs.
C.1.5 Task Drivers.
Task drivers define how proxy-assistant conversations are synthesized and how resulting episodes are evaluated. MirrorBench supports both single-turn utterance evaluation and multi-turn conversation-level assessment through user-facing APIs and CLI tools.
- Task drivers bridge high-level evaluation specifications and low-level execution by orchestrating proxy-assistant interactions and metric engagement.
- The default single-turn driver generates one user-proxy response from episode context for prompt-completion and turn-level utterance evaluation.
- The Mirror Conversation driver sequences user-proxy responses within episodes for conversation-level human-likeness evaluation across extended interactions.
- The Runner Facade exposes high-level Python methods for planning, execution, and result retrieval while abstracting orchestration complexity.
- Metrics are computed per episode and aggregated within each evaluation unit with confidence intervals.
- The CLI supports planning, dry runs, execution, reporting, run management, and cache operations.
C.2 Execution Flow
MirrorBench turns a declarative configuration into a replayable execution plan, synthesizes proxy-assistant conversations episode by episode, computes metrics, and persists aggregated results. The flow fixes component assignments at planning time and supports multiple execution backends, telemetry, caching, and reproducible reporting.
- A configuration specifies proxies, datasets, metrics, assistant, backend settings, observability, concurrency, and seeds for repeated-trial variance estimation.
- The Planner validates component compatibility, enumerates independent proxy-dataset-metric-seed units, and persists a replayable manifest.
- The Run Controller selects synchronous, asynchronous, or distributed execution and stores task-driver outputs and metric values in a run-specific SQLite database.
- Each unit enumerates immutable reference-dialogue episodes and delegates them to an episode executor for conversation synthesis and metric computation.
- Task drivers alternate user-proxy and assistant invocations across the reference dialogue length to construct synthetic histories.
- Episode artifacts bind synthetic transcripts to references, persist synthetic history for metric reuse, and retain auxiliary outputs for inspection and reproducibility.
- Unit-level aggregation computes means, standard deviations, and 95% confidence intervals, while RunSummary consolidates statistics, telemetry, and execution metadata.
- Caching reduces repeated LLM calls, and the plan manifest enables exact replays of component instantiation and invocation.
C.3 Implementation Details of Metrics
MirrorBench standardizes metric computation, aggregation, judge prompting, reproducibility artifacts, and run infrastructure. Lexical metrics use human-anchored normalization, judge metrics use repeated cached evaluations and calibration, and the framework records operational details for inspection and reporting.
- Lexical metrics tokenize concatenated user utterances with the GPT-4o tokenizer, while judge responses use content-based caching with metric-specific namespaces.
- Judge metrics run multiple independent samples per episode, average their scores, and retain individual samples and reasoning traces for audit.
- Aggregated metrics report means, standard deviations, and 95% confidence intervals using Student’s t-distribution, with optional bootstrap intervals.
- Lexical scores are converted to per-episode human-baseline z-scores, whereas calibrated judge results separately compute Human-Human and Proxy-Proxy statistics.
- Episodes lacking required human references are excluded from GTEval, Pairwise Indistinguishability, and z-score computation, with exclusions logged.
- Lexical estimates require at least five tokens by default; shorter utterances can be flagged as unstable or excluded to limit outlier effects.
- Evaluations fix LLM temperature at 0 and max_tokens at 2048, with exponential-backoff retries enabled by default.
- The framework records latency, generation time, token counts, cost, hierarchical traces, and structured logs across runs, units, and episodes.