Source-linked AI summary
HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence
Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian
TL;DR
Existing task-specific resources do not provide a shared, aligned basis for multimodal human-centered understanding and generation. HUG-VIS builds such a benchmark from synchronized multimodal performances and evaluates four tasks under a common zero-shot protocol. The results identify consistent task-specific weaknesses and show that evaluation and difficulty patterns must be analyzed across models, emotions, metrics, and tasks.
Problem
Existing resources do not jointly provide emotion annotations, body-motion information, alpha supervision, and synchronized audio–video–text, limiting reliable cross-task comparisons.
Method
HUG-VIS uses 8,400 performances from 30 actors on an identical 280-assignment grid, packaging synchronized video, audio, text, and alpha mattes for unified four-task zero-shot evaluation.
Results
Across four tasks, linguistic content dominates emotion recognition, moving boundaries remain the main matting obstacle, generation metrics and human judgments diverge at top rankings, and difficulty varies jointly by emotion, capability, and criterion.
Takeaways & Limitations
A shared actor-by-assignment foundation enables empirical analysis of trade-offs and relationships among human-centered understanding and generation tasks.
Takeaways & Limitations
Evaluation itself remains a bottleneck because no established cross-task metrics exist, motivating a newly designed suite built from task-specific measures and statistical analysis.
Abstract
from arXiv · showhide
Visual intelligence seeks to perceive, interpret, and synthesize the visual world and is central to modern computer vision. Human-centered visual intelligence is especially demanding because it studies people as expressive, socially situated subjects whose meaning is rarely conveyed by appearance alone. It couples vision with audio and language across four representative tasks: human emotion recognition, human video generation, human voice cloning, and human video matting. Yet existing resources remain task-specific, providing modalities and annotations for individual problems rather than a shared foundation coordinating understanding and generation. This limits multimodal signal use and broader research. We address this gap with HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence. It contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alpha mattes. We evaluate diverse open- and closed-source models across the four tasks under a unified zero-shot protocol using automatic metrics, criterion-specific mean opinion scores, and multiple cross-task analyses. Results show that (i) linguistic content dominates current emotion recognition, while purely visual affect recognition is weakest; (ii) in video generation and voice cloning, automatic metrics and human judgment agree overall but differ in their top rankings, requiring joint reporting; (iii) boundary fidelity under motion is the main remaining obstacle for human matting; and (iv) task difficulty varies across emotions, models, and metrics, with notable cross-task correlations. The dataset and results are available at https://github.com/GML-MMGroup/HUG-VIS.
1 Introduction
Existing human-centered resources isolate modalities, annotations, and tasks, preventing aligned comparisons across understanding and generation. HUG-VIS addresses this gap with a unified condition-aligned benchmark and reports four-task zero-shot findings.
- Motivation: Existing datasets provide task-specific modalities and annotations, leaving shared multimodal evaluation across human-centered tasks unavailable.Emotion, video generation, voice cloning, and matting resources differ in identities, behaviors, and acquisition conditions.
- Benchmark: Each clip bundles synchronized RGB video, noise-suppressed audio, an assigned prompt, its transcript, and a sequence-level alpha matte.This packaging supports the four benchmarked understanding and generation tasks within one condition-aligned resource.
- Evaluation: The authors evaluate diverse open- and closed-source systems across all four tasks under a unified zero-shot protocol and conduct cross-task analyses.The evaluation combines automatic metrics, criterion-specific human judgments, and analyses of relationships among task outcomes.
- Findings: Linguistic content dominates current emotion recognition, while purely visual affect recognition remains the weakest setting.In generation tasks, automatic metrics and human judgments agree overall but diverge in top rankings; matting chiefly struggles with moving boundaries.
- Findings: Cross-task analysis finds that task difficulty emerges jointly from emotions, models, and metrics rather than being intrinsic to a task.The reported analyses include correlations between recognition and matting errors and consistency between voice-quality predictors.
2 Related Work
Human-centered visual intelligence has expanded toward multimodal understanding and controllable generation, but its data resources remain fragmented. HUG-VIS places these capabilities on a shared actor-by-assignment foundation to support comparable and cross-task analysis.
- Field Scope: Human-centered visual intelligence connects appearance, identity, motion, speech, and affect in people as expressive and socially situated subjects.The field studies both understanding and generation of human-centered content across these linked signals.
- Field Scope: Modeling has shifted toward broadly pretrained, instruction-tuned systems, while signals have expanded from vision alone toward joint multimodal use.These developments broaden the field along modeling and signal axes.
- Fragmentation: Existing resources remain fragmented because they target individual capabilities or modality families and capture different identities under different conditions.Affective datasets, generation datasets, and matting datasets supply task-specific coverage rather than a common experimental basis.
- Fragmentation: Cross-benchmark comparisons conflate model capability with population, behavior, and acquisition confounds.This prevents reliable establishment of relationships among capabilities, including whether task difficulty is consistent or errors are correlated.
- HUG-VIS Response: HUG-VIS uses identical emotion–action–prompt performances from the same actor grid with synchronized video, audio, text, and alpha mattes.The shared structure is intended to place understanding and generation on a common footing while preserving task-specific inputs, outputs, and metrics.
3 The Dataset
HUG-VIS uses a controlled actor-by-assignment design to align multimodal assets across 8,400 performances. Its standardized protocol and synchronized releases support consistent evaluation and condition-aligned comparisons.
- Design Principle: HUG-VIS crosses 30 actors with 280 identical emotion–action–prompt assignments, producing 8,400 performances aligned across actors and conditions.Each evaluation unit contains synchronized RGB video, audio, assigned prompt, transcript, and alpha matte.
- Performance Design: The grid covers seven conditions, four action templates per emotion, and ten textual prompts per action.The seven conditions comprise six basic emotions plus Neutral, yielding 40 assignments per emotion and 280 overall.
- Recording Protocol: Each clip follows a rest–expression–rest protocol under fixed frontal green-screen capture with seated half-body framing and synchronized audio.Video is recorded at 1920 × 1080 and 240 FPS, then released at 30 FPS.
- Dataset Composition: The released dataset evenly distributes 8,400 clips across seven emotions, with 1,200 clips per condition and synchronized RGB, audio, text, transcript, and alpha assets.Each clip is packaged as a self-contained evaluation unit usable for understanding and generation tasks.
- Dataset Comparison: Compared with representative resources, HUG-VIS is the only listed dataset satisfying all five criteria: emotion annotation, half-body capture, alpha supervision, A-V-T alignment, and condition-aligned cross-task analysis.Its shared actor, source, and condition identifiers enable analyses on matched identities and instructed conditions.
4 Benchmark Protocols and Metrics
HUG-VIS evaluates four human-centered tasks with task-specific metrics and subjective scores under a common zero-shot protocol. It also introduces cross-task analyses for comparing emotion, source-clip, and motion difficulty.
- Evaluation Protocol: All four tasks use a zero-shot protocol that excludes benchmark samples from training, fine-tuning, calibration, and model selection.Emotion recognition predicts seven-class labels from image, video, audio, text, or modality combinations.
- Automatic Metrics: Task metrics match each output and reference: accuracy for emotion recognition, identity and synchronization measures for video generation, speech quality and speaker similarity for voice cloning, and alpha errors for matting.Matting reports spatial MAD, MSE, gradient, connectivity, and temporal dtSSD errors.
- Human Evaluation: Subjective studies report criterion-specific MOS on a 1–5 scale for identity, emotion naturalness, synchronization, motion naturalness, and voice similarity.The assessed criteria differ by generation regime and task.
- Cross-Task Analysis: The four cross-task analyses characterize difficulty by instructed emotion, source clip, and video motion, providing comparisons beyond single-task metrics.They are intended to diagnose where models break down and compare methods along shared axes.
- Emotion Difficulty: Emotion difficulty averages model results within each metric, orients values so larger scores indicate greater difficulty, and min-max normalizes conditions from easiest to hardest.The resulting profiles are visualized in Fig. 8.
- Metric Agreement: Metric agreement compares normalized emotion-difficulty profiles using Spearman rank correlation, where positive, negative, and near-zero values indicate corresponding correlation patterns.The agreement matrix is reported in Fig. 9.
- Source-Level Difficulty: Source-level generation difficulty normalizes metric values across matched clips, averages metrics and models within audio-driven and vision-driven regimes, and fits their distributions.The analysis uses a shared coordinate system for the two regimes.
- Motion Difficulty: Motion difficulty combines percentile-ranked frame-to-frame alpha variation and normalized foreground displacement, then relates it to averaged matting dtSSD error.The resulting analysis characterizes motion challenges and their relationship with matting performance.
5 Experimental Results
Across four tasks, HUG-VIS exposes modality-specific weaknesses and evaluation disagreements: language dominates emotion recognition, human judgments refine generative rankings, and motion-sensitive boundaries remain difficult for matting.
- Multimodal Emotion Recognition: 83.74% versus 27.77% shows Qwen2.5-Omni-7B gains substantially from adding text to video, while adding audio contributes only 0.05 point.Other models show the same pattern: text adds 30.51% or 28.60%, whereas further audio adds 0.74% or 0.18%.
- Human Video Generation: Automatic and subjective evaluations distribute leadership differently in audio-driven generation: Ditto reaches CSIM 0.904, while LatentSync leads Sync-C at 5.43 and subjective lip synchronization at 4.47.The qualitative comparison also identifies Sonic and LatentSync as producing the most stable identities and expression dynamics.
- Human Video Generation: No single vision-driven system dominates all automatic criteria: X-NeMo leads Head Animation LPIPS, PSNR, and FID, while AniPortrait leads CSIM and SSIM; Animate-X and Wan2.2 split Body Animation leadership.Qualitative results show differences in trajectory timing, gesture completion, appearance stability, and localized hand structure.
- Voice Cloning: Subjective voice-cloning leaders align more closely with speaker preservation than with reference-free predictors: IndexTTS2 leads open-source criteria, while Eleven Multilingual v2 leads closed-source criteria.OpenAudio S1 and Inworld TTS-1.5 lead UTMOS or DNSMOS but not subjective evaluation.
- Human Video Matting: 2.30 MAD, 0.82 MSE, 2.04 dtSSD, 10.43 gradient error, and 4.31 connectivity error make BiRefNet best on every matting criterion.Residual errors concentrate on fine hand boundaries, transient motion contours, and foreground leakage, making boundary fidelity under motion the primary challenge.
- Cross-Task Analysis: Difficulty varies by emotion and metric: Afraid is hardest for MER-RE, Sad for several generation metrics, Angry for voice-quality predictors, and Happy for VM-MAD.Metric agreement is selective, including strong alignment between VC-UTMOS and VC-DNSMOS at ρ = 0.82 and identical ordering across two video-generation regimes at ρ = 0.99.
- Cross-Task Analysis: Cross-task analyses establish connections unavailable through single-task evaluation, linking difficulty and errors across understanding and generation tasks.The shared actor and source support comparisons across tasks, emotions, models, and metrics.
6 Discussion
The discussion identifies an affective asymmetry between understanding and generation, an evaluation bottleneck, and relational task difficulty. It argues for stronger visual and cross-modal modeling together with integrated evaluation and cross-task analysis.
- 6.1 Insights: Linguistic content is strongly predictive for emotion understanding, while visual evidence is weakest; generation instead requires affect to be synthesized through pixels and motion.This creates a representational gap between reading affect and producing it.
- 6.1 Insights: Automatic and qualitative evaluations agree on overall trends in video generation and voice cloning but diverge in their top rankings.The discussion therefore identifies evaluation itself as a bottleneck and motivates refining task-specific and cross-task metrics.
- 6.1 Insights: Cross-task analyses show that difficulty is relational rather than intrinsic, arising jointly from the emotion, model, and evaluation criterion.These analyses are presented as guidance for unified multimodal understanding and generation models.
- 6.1 Insights: The proposed directions emphasize visual affect representations, fusion mechanisms, joint automatic metrics and MOS, extended cross-task measures, and multi-task architectures.These directions target the documented asymmetry and evaluation bottleneck.
7 Conclusion
HUG-VIS places human-centered understanding and generation on a shared actor-by-assignment foundation, enabling aligned multimodal evaluation. Its analyses identify modality, evaluation, matting, and task-difficulty patterns across the benchmark.
- Conclusion: HUG-VIS unifies human-centered understanding and generation tasks on a shared actor-by-assignment grid.The dataset includes 8,400 seated half-body performances from 30 actors across 280 identical assignments, with synchronized video, audio, text, and alpha mattes.
- Conclusion: Linguistic content dominates present-day emotion recognition, while purely visual affect recognition remains the weakest setting.
- Conclusion: Automatic metrics and human judgments agree in overall trend but diverge at the top in both generation tasks, requiring joint reporting.
- Conclusion: Boundary fidelity under motion is the principal remaining obstacle for human matting.
- Conclusion: Task difficulty emerges jointly from the emotion, capability, and evaluation criterion rather than being intrinsic to a condition.
Statements and Declarations
The study reports participant consent for research capture, authorized use of identifiable likenesses, voices, and behaviors, and publication of identifiable examples for covered research purposes.
- Statements and Declarations: Actors provided written informed consent for research capture and authorized use of identifiable likeness, voice, and performed behavior.
- Statements and Declarations: Written authorizations allow publication and demonstration of identifiable examples for covered research purposes.
- Statements and Declarations: The dataset’s synchronized audio, video, text, and alpha assets are released through the project.