Source-linked AI summary
Towards Customized Multimodal Role-Play
Chao Tang, Jianzong Wu, Qingyu Shi, Ye Tian, Aixi Zhang, Hao Jiang, Jiangning Zhang, Yunhai Tong
TL;DR
Existing multimodal models do not adequately customize persona, dialogue style, and visual identity together for consistent character interaction. The paper introduces CMRP, RoleScape-20, and UniCharacter’s two-stage training framework, which outperforms prior approaches and improves multimodal alignment and role-play quality. The authors conclude that this combination supports coherent, personalized virtual characters, while noting important limits in video and long-turn interaction.
Problem
Existing systems and multimodal models do not jointly customize persona, dialogue style, and visual identity for persona-driven interaction.
Method
The paper introduces CMRP and RoleScape-20, then adapts a unified multimodal model with Unified-SFT and Character-GRPO using character-specific multimodal data.
Results
UniCharacter outperforms competitive baselines in role consistency, dialogue authenticity, image fidelity, and cross-modal alignment.
Takeaways & Limitations
CMRP with unified modeling supports coherent multimodal role-play by aligning persona, dialogue style, and visual identity.
Takeaways & Limitations
The current task is limited to text and images and single-turn scenarios, leaving video consistency and long-turn stability untested.
Abstract
from arXiv · showhide
Unified multimodal understanding and generation models enable richer human-AI interaction. Yet jointly customizing a character's persona, dialogue style, and visual identity while maintaining output consistency across modalities remains largely unexplored. To mitigate this gap, we introduce a new task, Customized Multimodal Role-Play (CMRP). We construct the RoleScape-20 dataset comprising 20 characters, including training and evaluation data that cover persona, stylistic descriptions, visual/expressive cues, and text-image interactions. Building on a unified model, we devise UniCharacter, a two-stage training framework containing Unified Supervised Finetuning (Unified-SFT) and character-specific group relative policy optimization (Character-GRPO). Given only 10 images plus corresponding interaction examples, the model acquires the target character and exhibits coherent persona, style, and visual identity in both generated text and images. This process takes about 100 GPU hours. Experiments on the RoleScape-20 dataset show that the proposed method substantially outperforms prior approaches. Ablation studies further validate the effectiveness of our cross-modal consistency design and few-shot customization strategy. We argue that CMRP, coupled with unified modeling, provides a basis for next-generation characterful and immersive interactive agents.
1. Introduction
The paper introduces CMRP to customize a character’s persona, dialogue style, and visual identity jointly, addressing the limitations of single-modality systems. It proposes UniCharacter and reports stronger performance across role consistency, dialogue quality, image fidelity, and cross-modal alignment.
- Existing systems typically customize either how a character speaks or how it looks, rather than both simultaneously.
- CMRP adapts a general-purpose multimodal model using a textual profile, reference images, and example dialogues to generate in-character responses and appearance-consistent images.
- RoleScape-20 contains 20 characters with profiles, 5–15 reference images, 150–250 role-playing dialogues, and fine-grained multimodal annotations.
- UniCharacter outperforms competitive baselines in role consistency, dialogue authenticity, image fidelity, and cross-modal alignment.
- UniCharacter uses Unified-SFT followed by Character-GRPO to support coherent multimodal role-play and reduce T2I overfitting while preserving text-image consistency.
2. Related Work
Prior customization methods generally operate in one modality or support limited multimodal interaction. The paper positions CMRP and UniCharacter as a joint text-image customization setting with richer character modeling.
- Training-based and tuning-based customization methods represent two broad approaches to adapting models to user-provided roles.
- Existing customization methods are limited to text or images during inference, making joint-output interactions difficult.
- CMRP requires joint text and image generation from user inputs, and UniCharacter is presented as a tuning-based method for this setting.
- Unified multimodal models integrate understanding and generation, but coherent personalization for complex interactive scenarios remains challenging.
3. The RoleScape-20 Dataset
RoleScape-20 defines multimodal role-play around a character’s profile, reference images, and reference dialogues, requiring coordinated text and image outputs. Its construction pipeline produces dialogue, T2I, knowledge-QA, and VQA data for unified modeling.
- Problem Formulation: CMRP represents a character with a textual profile, core reference images, and reference dialogues capturing personality, visual identity, and speaking style.
- Problem Formulation: Given a user query, the model must generate text following the character’s personality and style alongside a contextually relevant image depicting its visual features.
- Problem Formulation: The multimodal response is modeled as a joint distribution whose chain-rule decomposition generates text before a conditional image.
- Dataset Construction: RoleScape-20 contains 20 diverse characters spanning real-world figures, anime and game characters, and animals.
- Dataset Construction: Compared with text-only, personalized multimodal, and image-customization datasets, RoleScape-20 combines visual modality, in-character dialogue, reasoning, generation instructions, knowledge QA, and VQA.
- Dataset Construction: The construction pipeline transforms character materials into expanded dialogues, multimodal role-play and T2I pairs, knowledge QA, and visual QA data.
4. Method
UniCharacter combines unified supervised fine-tuning with Character-GRPO to support multimodal role-play while improving T2I diversity and character consistency.
- Unified-SFT: Unified-SFT frames training as multitask learning across role-play chatting, thinking, VQA, and Knowledge QA.These tasks target in-character responses, image-generation reasoning, image-based questions, and character knowledge.
- Unified-SFT: The text-generation tasks maximize target-text likelihood using cross-entropy losses, while T2I generation uses a Rectified Flow MSE objective.The T2I branch generates images from text conditions using a noise-to-clean residual objective.
- Character-GRPO: Character-GRPO addresses visual overfitting by generating multiple samples per character-specific prompt instead of relying on one ground-truth image.The sampling process expands character-image mappings without requiring ground-truth images.
- Character-GRPO: Its reward combines text-image alignment, diversity, and training-set similarity penalties to guide character-consistent and varied image generation.Alignment uses CLIP and VQA rewards, while diversity uses perceptual variation and similarity penalties.
- Character-GRPO: The composite reward uses α = 0.45, β = 0.3, γ = 0.1, and δ = 0.15 as default hyperparameters.The coefficients weight the reward components in the final sample-level signal.
5. Experiment
Experiments evaluate UniCharacter against personalized multimodal, image-generation, and vision-language baselines across role-play, generation, comprehension, and ablation settings. The reported results show stronger overall performance, with qualitative evidence of improved identity and character consistency.
- Experiment Setup: Experiments compare UniCharacter with UniCTokens, DreamBooth-equivalent customization, and Qwen2.5-VL-based role-playing baselines.Evaluation covers personalized unified modeling, T2I generation, and personalized understanding and role-play.
- Experiment Setup: The evaluation spans text-based role play, T2I, multimodal role play, Knowledge QA, and VQA using judge scores, CLIP-I, CLIP-T, DINO, and accuracy.Text role play is scored for Memorization, Personality, and Diversity by an LLM-as-Judge methodology.
- Quantitative Results: UniCharacter outperforms leading baselines across image generation, text-based role play, Knowledge QA, and VQA, and surpasses UniCTokens across most evaluated tasks.The authors report stronger role embodiment, generative capability, comprehension, and unified modeling capacity.
- Qualitative Results: Qualitative comparisons show better image quality, text-image alignment, identity fidelity, and character-matched responses than the compared baselines.The reported role-play example contrasts concise, sarcastic Chandler responses with Qwen2.5-VL’s long-winded, out-of-character responses.
- Qualitative Results: Qualitative results show consistent character identity and personality, controllable character states, and reduced overfitting during user-directed interactions.The authors report consistency across the characters shown in Figure 5 and alignment between responses and personality traits.
- Ablation Studies: GRPO improves image quality and training-set similarity metrics on T2I and multimodal role-play tasks compared with omitting the GRPO stage.Reward ablations report that alignment improves image quality, while diversity reduces similarity to the training set.
6. Conclusion
The paper introduces CMRP and UniCharacter for coherent multimodal character customization, while identifying limits in video generation, long-turn dialogue, real-time deployment, safety, and user-in-the-loop customization.
- UniCharacter jointly models role-play chatting, thinking processes, knowledge QA, VQA, and T2I generation while aligning persona, dialogue style, and visual identity.
- The two-stage framework mitigates few-shot overfitting, improving image diversity and generalization.
- Ablation studies support the importance of the reward design and Character-GRPO stage for role-play quality and multimodal alignment.
- Current CMRP is limited to text and images and single-turn scenarios, leaving customized video, temporal consistency, long-turn stability, and role-drifting prevention for future work.
Impact Statement
The supplied passages describe supporting materials and architectural preliminaries for the unified multimodal framework, including an introduction video and GRPO optimization details.
- The paper provides a 5-minute introduction video to help readers quickly grasp its primary idea.
- The architecture uses a unified multimodal model capable of simultaneous image-text understanding and generation.
- GRPO estimates advantages by normalizing rewards across trajectories sampled from the same prompt.
- The GRPO objective combines a clipped surrogate loss with a KL-divergence penalty for policy stability, although the paper sets β = 0.
C. More Qualitative Results
The paper places its extensive qualitative results in a separate anonymous local HTML page supplied with the supplementary materials.
- Extensive qualitative results across various tasks are presented on a separate anonymous local HTML project page.
- Readers are directed to the supplementary file’s “index.html” page for these results.
- The separate page is used because of the large amount of qualitative material.
D. More Ablation Studies
The ablation studies examine Character-GRPO, training-data composition, inference-time thinking, and user preferences across multimodal role-play tasks. Results indicate trade-offs between textual performance, image quality, diversity, and training-set similarity.
- Training stage: Models trained with Character-GRPO achieve superior text-image alignment and image diversity compared with models without GRPO.
- Training data composition: Adding Extension Dialogues improves Memorization, Personality, and Diversity by reducing repetitive role-play responses.
- Training data composition: Adding Thinking Process data preserves strong textual metrics and improves image-quality metrics beyond the Original setting.
- Inference strategy: Inference-time thinking has different effects depending on training: it does not improve Unified-SFT models, whereas its effect is evaluated separately for Unified-SFT plus Character-GRPO models.
- User study: UniCharacter outperformed DreamBooth, Qwen2.5-VL, and UniCTokens across T2I generation, multimodal role-play, and text role-play in the user study.
- Dataset coverage: RoleScape-20 includes 9 human, 4 animal, and 7 anime characters.
F.1. Data Collection
RoleScape-20 combines diverse character sources with profiles, dialogues, images, multimodal annotations, and task-specific QA data to support role-play evaluation and training.
- Character and source collection: RoleScape-20 covers human, animal, and anime characters collected from diverse online, film, television, game, and anime sources.
- Character and source collection: Character resources include profiles, example dialogues, reference images, and manually reviewed dialogue expansions when original samples are insufficient.Approximately 10 example dialogues are expanded to around 200 dialogues using Qwen3, with a subset used for diversity testing.
- Multimodal annotation: Each image-dialogue pair receives a reasoning process and standardized generation instruction for multimodal role-play generation.GPT-4o is used to annotate the “Thinking Process” and “Generation Instruction.”
- Evaluation data: The dataset includes specialized personality and memorization test questions designed to evaluate character-specific capabilities.Approximately 20 test questions are annotated for each capability.
- Evaluation data: Knowledge QA and VQA receive separate training and test sets with character- and image-specific question samples.Knowledge QA includes approximately 100 training samples and 10 test questions per character; VQA includes approximately 20 training samples and 5 test questions per character image.
G.1. More Implementation Details
Implementation combines task-balanced Unified-SFT with Character-GRPO image sampling and multi-objective rewards, while text role-play is assessed through LLM-based scoring.
- Unified-SFT: Unified-SFT samples text-to-image and image-understanding tasks at a 200:1 ratio while ensuring both task types appear in every batch.Optimization uses AdamW with a learning rate of 2e-5 and no warmup.
- Character-GRPO: Character-GRPO samples G = 8 images per prompt and uses conservative optimization with learning rate 1e-5, batch size 6, and β = 0.Training stability relies on strict clipping rather than a KL-divergence penalty.
- Reward design: The reward signal combines aesthetic quality, text alignment, identity preservation, and perceptual diversity objectives.CLIP similarity, VQA consistency, LPIPS diversity, and DINO-based trainset similarity penalties implement these objectives.
- Evaluation: Text-based role play is evaluated with an LLM-as-Judge on Personality, Memorization, and Diversity metrics.Personality and Memorization average per-question scores, whereas Diversity jointly evaluates 20 user-input–response pairs.
- Evaluation: All three text role-play scores use a 1-to-7 scale, with evaluation prompts specified for Personality, Memorization, and Diversity.