Source-linked AI summary
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
Yuhang Zhu, Mingxuan Du, Benfeng Xu, Jie Gao, Lingyun Yu, Hongtao Xie
TL;DR
Existing RPA benchmarks evaluate continuations from fixed histories with user-independent rubrics, limiting assessment of independently constructed multi-turn role-playing and individual satisfaction. PALATE uses per-user simulators, free multi-turn interactions, and personalized rubrics; it finds that capability advantages differ across tracks and users, rather than yielding one dominant ranking.
Problem
Existing benchmarks condition RPA evaluation on fixed dialogue histories and user-independent rubrics that need not reflect individual satisfaction.
Method
PALATE trains per-user simulators for free multi-turn interaction and evaluates resulting user–RPA trajectories with personalized, generic, and whole-session rubrics.
Results
Across 16 candidates, capability advantages differ across generic turn quality, long-horizon sessions, and per-user satisfaction, with no single candidate winning all users.
Takeaways & Limitations
PALATE produces interactive evaluation profiles that reveal cross-track capability mismatches and per-user differences instead of a single user-independent leaderboard.
Takeaways & Limitations
PALATE currently covers five extensively annotated users and lacks an end-to-end human-ranking reference spanning every candidate in the main table.
Abstract
from arXiv · showhide
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.
1 Introduction
PALATE addresses two limitations of fixed-history, user-independent evaluation: inherited dialogue quality can distort RPA scores, and static rubrics may misalign with individual user satisfaction. It instead evaluates free multi-turn user–RPA trajectories with person-aligned simulators and personalized rubrics, characterizing specific pairs across multiple quality dimensions.
- Limitations: 0.21 points: high-quality character-side histories increased candidates’ mean overall continuation score on a five-point scale relative to original histories.The experiment used five real conversations longer than 20 turns and four candidate RPAs, with all continuations scored by the same generic rubric.
- Limitations: Fixed-history benchmarks conflate an RPA’s capability with the quality of externally supplied dialogue histories.High-quality borrowed histories inflate continuation scores, whereas degraded histories deflate them.
- Limitations: Static user-independent rubrics may misjudge satisfaction because interaction preferences differ across users.A gradual relationship arc may suit slowburn preferences but frustrate users seeking rapid conflict.
- PALATE: PALATE trains dedicated simulators from real dialogue histories and lets them freely interact with candidate RPAs from the character’s opening.This design makes each candidate help construct its evaluated trajectory rather than inherit another system’s character-side history.
- PALATE: PALATE evaluates specific user–RPA pairs using personalized experience, generic turn quality, and whole-session quality.The benchmark was applied to 16 candidates and is accompanied by released multi-turn conversations with user-experience annotations.
2 Related Work
Related work has advanced role-playing agents from persona-grounded dialogue toward character-specific modeling and interactive evaluation. Recent personalized evaluation and user simulation motivate PALATE’s shift from task-conditioned to person-conditioned assessment.
- Modern RPAs build on persona-grounded dialogue through character-specific data, synthetic personas, instruction tuning, self-alignment, multi-character portrayal, and human-like reasoning.
- Existing benchmarks evaluate character knowledge, persona maintenance, voice, dialogue consistency, psychological traits, social and emotional capabilities, plot progression, dialogue points, trajectories, and multi-agent simulation.
- Personalized evaluation links participant profiles to live feedback, exposes variation hidden by aggregate rankings, and derives user-conditioned criteria from interaction histories.
- PALATE extends adaptive evaluation from task-conditioned criteria to person-conditioned evaluation.
- User simulation has evolved from agenda-based and corpus-trained task-oriented methods toward tool-agent and interactive evaluation, with LLM simulators learning natural user turns or conditioning on inferred profiles.
3 PALATE Benchmark
PALATE evaluates role-playing agents through person-aligned, free-form multi-turn interactions rather than borrowed histories. It combines per-user simulators, frozen character panels, personalized rubrics, and three-track scoring for user experience, turn quality, and whole-session quality.
- Benchmark pipeline: PALATE comprises user-behavior modeling, free interaction, person-aligned evaluation, and four operational steps from simulator training through three-track trajectory scoring.The four steps are simulator training and validation, frozen free interaction, personalized rubric construction, and scoring.
- Data and splitting: 300 bilingual structured character cards form the released pool, while 10 pre-frozen cards constitute the main evaluation panel.The benchmark also includes distinctive IP anchors, and the released pool supports future alternative panels.
- Data and splitting: 5 users contribute 5,133 annotated user turns, with sessions split into training and evaluation sets for simulator learning and validation.The main cohort provides approximately 1,000 turns per user, and volunteers choose characters themselves.
- Free interaction: Each simulator interacts freely with every candidate RPA on the 10-card panel, producing 5 × 10 × C × 2 trajectories, with C = 16 in the paper.Each simulator–character–candidate cell is independently repeated twice, and trajectories are constructed by alternating simulator and RPA generations.
- Person-aligned evaluation: Personalized rubrics convert each user’s annotated training history into frozen, interpretable experience criteria, while shared tracks preserve generic turn and whole-session comparisons.The rubric identifies rewards, penalties, applicability conditions, counterexamples, and subsequent-reaction semantics under a common schema.
4 Evaluation Design and Results
PALATE evaluates 16 candidates through personalized, generic, and session-level tracks on candidate-constructed multi-turn trajectories, revealing capability mismatches rather than a single dominant ranking. Personalized rubrics with user reactions improve agreement with human judgments, while results remain highly stable across repetitions and judges.
- Evaluation design: Table 2 evaluates 16 candidates across five personalized-user tracks, generic turn quality, and whole-session quality using GPT-5.5 judgments.Each user–candidate cell has two independent rollouts; Personalized and Generic score 40 frozen valid nonterminal decision points, while Session scores 20 complete trajectories.
- Evaluation results: GPT-5.4 leads generic turn quality, Claude Sonnet 4.6 leads Session, and different users favor Qwen3-Max, DeepSeek V4 Pro, Claude Opus 4.8, or Claude Sonnet 4.6.No candidate wins all users, and Qwen3-Max ranks markedly lower for U4 despite matching U1 best.
- Evaluation results: GLM-5.1 reaches the leading tier among open-weight models, whereas MiniMax M2-her and CoSER-Llama-3.1-70B remain lower despite role-play-oriented training or evaluation.Model provenance and training orientation therefore do not determine interactive performance across characters and users.
- Robustness: 0.959 is the minimum Spearman correlation across independent model-rank repetitions, while cross-family rejudging preserves wide-margin Personalized and Generic ordering.Session rankings and small top-end gaps are more judge-sensitive.
- Rubric validation: 0.613 is the best macro agreement for Personalized with reaction, versus 0.551 without reaction; Generic changes from 0.480 to 0.507, and MiniMax-aligned obtains 0.467.The full setting improves the equal-user signal for every user, although U2 favors Generic, and chance is 0.500.
5 Conclusion
PALATE reframes role-playing evaluation from isolated, fixed-history assessment to user–RPA pair evaluation that combines general quality with user-perspective satisfaction. Across 16 candidates, its interactive profiles expose cross-track capability mismatches and per-user differences rather than producing a single leaderboard.
- Motivation: Fixed-history evaluation confounds an RPA’s capability with externally supplied dialogue history, while user-independent scoring obscures individual differences in satisfaction.These limitations motivate changing the basic unit of evaluation.
- Evaluation unit: PALATE decomposes each person into a behavior policy learned from real dialogue and an individual utility supervised by experience annotations.This supports evaluation centered on the user–RPA pair.
- Evaluation unit: PALATE shifts the basic unit from an isolated RPA to a user–RPA pair, adding a user-perspective satisfaction reference while retaining general quality evaluation.The framework therefore evaluates both generic quality and personalized experience.
- Findings: Across 16 candidates, advantages on the three evaluation tracks do not coincide, and the five users do not share a single best candidate.PALATE consequently reports interactive evaluation profiles that locate cross-track mismatches and per-user differences.
6 Limitations
PALATE currently covers five extensively annotated users, and its validation relies on disjoint held-out sessions rather than an end-to-end human-ranking reference spanning every main-table candidate. This limitation reflects the cost of collecting long conversations, turn-level satisfaction labels, and repeated evaluations of every candidate.
- Coverage: PALATE currently covers five extensively annotated users.The benchmark’s current user coverage is limited to five users.
- Human-ranking validation: Cost constraints prevent an end-to-end human-ranking reference spanning every candidate in the main table.Collecting long conversations and turn-level satisfaction labels, then repeatedly asking the same people to evaluate every candidate, is costly.
- Held-out validation: Validation instead uses disjoint held-out sessions to assess user-simulator behavioral fidelity and agreement between personalized scoring and human judgments.These held-out experiments do not provide a human-ranking reference covering every candidate in the main evaluation.
7 Ethics Statement
The study collected real role-playing conversations and turn-level experience annotations under informed consent, compensation, and explicit authorization for public release of de-identified data and per-user LoRA weights.
- 7 Ethics Statement: Participants provided informed consent, received compensation, and authorized public release of de-identified research data and per-user LoRA weights trained from their data.The data collection covered real role-playing conversations and turn-level experience annotations.
A Experimental Details · A.1 Controlled Fixed-History Experiment
The controlled fixed-history experiment rewrites five long platform conversations into matched HQ and degraded histories, then evaluates candidate responses independently across frozen depths. Results show that inherited history chiefly affects what candidates generate, with substantial candidate-specific sensitivity and ranking changes.
- A.1 Controlled Fixed-History Experiment: Five real conversations with more than 20 user–character rounds were selected, fixing user turns, character actions, relationships, commitments, and plot events.Each original character-side history was rewritten into two realizations while preserving plot facts and continuation compatibility.
- A.1 Controlled Fixed-History Experiment: HQ histories used specific, natural, persona-grounded language, whereas degraded histories rendered the same facts more generically, repetitively, and with less detail.Neither rewrite could add or remove plot facts, create contradictions, refuse, break character explicitly, or introduce unsafe content.
- A.1 Controlled Fixed-History Experiment: 900 replies were scored with the Generic rubric across one shared Original arm and four candidates evaluated on HQ and degraded histories.The Original arm contributed 5×20 = 100 replies; each candidate contributed 100 replies per rewritten arm, with candidate outputs not written back.
- A.1 Controlled Fixed-History Experiment: DeepSeek V4 Pro was most sensitive to inherited history at 0.57, while GPT-5.1 had the smallest gap at 0.10 and remained strong in both arms.Claude Sonnet 4.6 ranked above Gemini 2.5 Pro under HQ history but below it under degraded history.
- A.1 Controlled Fixed-History Experiment: The crossed analysis froze continuations from both histories in all 400 candidate–conversation–depth cells and swapped only the history shown to the judge.This separates generation changes from direct judge responses to visible history.
- A.1 Controlled Fixed-History Experiment: +0.49 was the aggregate continuation main effect, compared with −0.16 for the visible-history main effect.The visible-history effect was negative for every candidate, while the larger continuation effect indicates that inherited history chiefly changes candidate generation.
A.2 User-Simulator Training and Selection · A.3 Repeat Stability and Cross-Judge Rejudging · B Dataset and Character Panel
The paper trains per-user simulators as next-action predictors, selects a lightweight 35B-A3B Instruct + LoRA configuration for evaluation, and assesses repeat stability through frozen cross-judge rejudging. The supplied passages do not provide substantive evidence about Dataset and Character Panel.
- A.2 User-Simulator Training and Selection: User simulators predict each human user message from the preceding RPA message, with naturally ended sessions represented by a [QUIT] target.Training and inference share the same frozen task instruction, and models receive no names, demographic attributes, or handwritten profiles.
- A.2 User-Simulator Training and Selection: The training ablation compares Qwen3.5-4B and Qwen3.5-35B-A3B across Base and Instruct initialization, no adaptation, per-user LoRA, and full-parameter tuning.Comparisons use the same frozen decision points for users U3 and U4.
- A.2 User-Simulator Training and Selection: Full tuning has the highest point estimate in the small ablation, supporting user-specific training.The passage does not report the corresponding numerical point estimate.
- A.2 User-Simulator Training and Selection: The main experiment selects 35B-A3B Instruct + LoRA because its fidelity is near the chance-indistinguishable region while requiring only a lightweight adapter per user.This choice avoids the substantially higher cost of training, storing, and deploying a full model per user.
- A.2 User-Simulator Training and Selection: Full-parameter updating also poses greater individual-data overfitting risk when each user has fewer than 1,000 examples.The passage presents this as an additional reason to prefer lightweight per-user adaptation.
- A.3 Repeat Stability and Cross-Judge Rejudging: Cross-judge rejudging freezes 4 Personalized points, 4 Generic points, and 2 Session trajectories from every candidate–user–repetition cell, yielding 1,600 inputs for identical-input evaluation.The two additional judges receive the same frozen inputs; the supplied passages identify agreement tables but provide no values.
B.1 Human Data and Splits … D Complete Prompts and Output Protocols
PALATE’s evaluation is built from frozen, session-isolated human data, a constrained 300-card character pool, and three complementary scoring tracks. Its prompts operationalize generic quality, whole-session capability, user simulation, personalized satisfaction, and human-agreement testing through fixed protocols.
- B.1 Human Data and Splits: 1,233 held-out user turns are tested, including 1,206 with valid replies, satisfaction labels, and adjacent reactions, while splits remain isolated by complete session.Training satisfaction labels may inform rubric construction, but held-out satisfaction is used only for final agreement calculation.
- B.2 The 300 Character Cards: The 300-card pool contains 240 original cards and 60 distinctive IP anchors with bilingual metadata, newly authored wording, and manual review for user-agency violations.Cards include name, introduction, description, opening, and a structured candidate-side persona.
- B.3 Frozen Ten-Card Panel: The frozen ten-card panel requires 8 original and 2 IP cards, balanced gender, ten archetypes, complete bilingual fields, self-consistent openings, and preserved user agency.Because panel cards may appear in training conversations, evaluation measures new free trajectories and cross-candidate reuse rather than character-level holdout generalization.
- C.1 Generic Turn-Level Rubric: Generic scoring uses 12 shared dimensions plus an independent overall score, with GPT-5.4 leading overall while candidates specialize in different turn-level strengths.GPT-5.4 leads coherence, character knowledge, persona behavior, voice, response fit, intent response, and action continuity; DeepSeek V4 Pro leads lexical richness, local narrative contribution, and interaction appeal; Qwen3-Max leads fluency.
- C.2 Whole-Session Rubric: Whole-session scoring uses 11 shared dimensions and an independent overall, with Claude Sonnet 4.6 leading accumulated trajectory quality rather than polished individual-turn quality.Claude Sonnet 4.6 is strongest on emotional and relationship development, pacing, memory, and recovery; GPT-5.4 leads persona and motivation stability, instruction cooperation, and cross-turn logic; DeepSeek V4 Pro leads plot progression and structural diversity.
- C.3 Personalized-Rubric Construction: Personalized-rubric construction uses 50 stratified training turns and the full rating distribution to infer each user’s rating prior, four drivers, reaction semantics, contrastive rules, anchors, caps, uncertainty, and scoring policy.The evaluator is explicitly person-specific rather than a summary of general role-playing quality.
- C.4 Personalized Scoring Interface: All users share the same personalized-scoring interface, while only the frozen rubric differs and the judge outputs normalized probabilities p1 through p5.The no-reaction ablation removes only the next user message and scores from preceding context plus the candidate reply.
- C.5 Human-Agreement Protocol: Human agreement is computed over 1,206 turns and 20,477 within-session pairs, aggregated within users and macro-averaged, with the MiniMax-aligned baseline reconstructed rather than officially reproduced.Restricting comparisons to one session avoids confounding character, phase, or rating-prior differences.
E Release, Ethics, and Reproducibility
The planned release supports reproducible PALATE evaluation through de-identified data, frozen character and rubric resources, simulator configurations, trajectory indices, and scoring code. Ethical safeguards include informed consent, compensation, privacy review, restricted simulator use, and separate review of IP-anchor cards.
- Reproducibility: The release package includes de-identified dialogues and satisfaction labels, 300 bilingual character cards, five frozen rubric files, simulator configurations, and scoring resources.It also includes frozen splits, checksums, construction and scoring templates, authorized per-user LoRAs, indices for 1,600 trajectories, and aggregation code.
- Ethics: Participants gave informed consent and received compensation, while direct identifiers, timestamps, unnecessary metadata, and personal information in free text will be removed or reviewed.These measures apply before release of the collected materials.
- Ethics: Per-user simulators are authorized only for research evaluation, excluding impersonation, third-party contact, and user-facing deployment.IP-anchor cards will undergo separate source and redistribution review before release.