Source-linked AI summary

CelebV-HQ: A Large-Scale Video Facial Attributes Dataset

Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, Chen Change Loy

arXiv:2207.12393v1cs.CV

TL;DR

Face research has lacked a large-scale video dataset combining high quality, diversity, and rich facial-attribute annotations. CelebV-HQ addresses this gap with manually labeled face videos, analyzes their distributions and temporal properties, and evaluates them on generation and editing. The dataset contains 35,666 clips involving 15,653 identities, and experiments demonstrate its effectiveness and potential.

  • Problem

    Face research lacks a video dataset combining large scale, high quality, diversity, and rich facial-attribute annotations.

  • Method

    CelebV-HQ constructs and manually annotates face videos, then analyzes their distributions and evaluates them on video generation and facial-attribute editing.

  • Results

    35,666 clips involving 15,653 identities and 83 attributes are analyzed, while experiments demonstrate CelebV-HQ’s effectiveness and potential.

  • Takeaways & Limitations

    CelebV-HQ provides a foundation for future face-video research and may benefit neural rendering and face-analysis studies.

  • Takeaways & Limitations

    Video alignment remains challenging because keypoint alignment can be jittery, while unaligned faces can degrade generation quality.

Abstract

from arXiv · show

Large-scale datasets have played indispensable roles in the recent success of face generation/editing and significantly facilitated the advances of emerging research fields. However, the academic community still lacks a video dataset with diverse facial attribute annotations, which is crucial for the research on face-related videos. In this work, we propose a large-scale, high-quality, and diverse video dataset with rich facial attribute annotations, named the High-Quality Celebrity Video Dataset (CelebV-HQ). CelebV-HQ contains 35,666 video clips with the resolution of 512x512 at least, involving 15,653 identities. All clips are labeled manually with 83 facial attributes, covering appearance, action, and emotion. We conduct a comprehensive analysis in terms of age, ethnicity, brightness stability, motion smoothness, head pose diversity, and data quality to demonstrate the diversity and temporal coherence of CelebV-HQ. Besides, its versatility and potential are validated on two representative tasks, i.e., unconditional video generation and video facial attribute editing. Furthermore, we envision the future potential of CelebV-HQ, as well as the new opportunities and challenges it would bring to related research directions. Data, code, and models are publicly available. Project page: https://celebv-hq.github.io.

1 Introduction

CelebV-HQ addresses the lack of large-scale, high-quality face video data with diverse facial-attribute annotations. The dataset’s analyses and experiments demonstrate its diversity, video quality, and utility for generation and editing.

  • Current facial datasets mainly provide either static images with attributes or videos with insufficient scale and diversity.
  • Its construction combines multilingual, diverse Internet queries with automatic preprocessing to filter and organize high-quality face clips.The queries cover 11 languages, 8,376 entities, and 3,717 actions before filtering.
  • CelebV-HQ contains 35,666 in-the-wild clips at least 512 × 512, covering 15,653 identities and 83 manually labeled attributes.The attributes comprise 40 appearance, 35 action, and 8 emotion categories.
  • Statistical analysis evaluates age, ethnicity, temporal quality, brightness variation, head-pose distribution, motion smoothness, and data quality.Comparisons with CelebA, CelebA-HQ, and VoxCeleb2 support the reported distribution and video-quality analysis.
  • Experiments use unconditional video GANs and temporal-constrained image-to-image baselines for video generation and facial-attribute editing.Action-specific subsets successfully generate their corresponding actions.
  • The dataset is presented as useful beyond the evaluated tasks, including neural rendering and face analysis.

2 Related Work

Existing face datasets span image and video modalities, but they do not jointly provide large-scale, diverse videos with rich facial-attribute annotations. CelebV-HQ is positioned to fill this gap for face-video research.

  • Face-video editing methods commonly operate on StyleGAN latent spaces but often train on images because large-scale, high-quality video datasets are lacking.This creates a temporal-consistency problem for video-based editing.
  • Image datasets support facial recognition and attribute analysis, while video datasets support speaker, talking-face, reenactment, and emotion research.
  • CelebA and LFWA provide 40 facial-attribute annotations, while CelebA-HQ raises image resolution to 1024×1024.
  • Existing video datasets include five long CelebV videos, audiovisual VoxCeleb datasets, and MEAD with 60 actors across seven view directions.
  • These datasets either contain attribute-labeled images or unlabeled videos with insufficient diversity.

3 CelebV-HQ Construction

CelebV-HQ is constructed through collection, preprocessing, and manual annotation designed to preserve scale, quality, diversity, and temporal coherence. Its annotation scheme separates appearance, action, and emotion attributes.

  • The construction pipeline comprises data collection, data preprocessing, and data annotation, targeting large-scale, high-quality, diverse clips reflecting real-world distributions.
  • Data Collection: Multilingual queries covering celebrity names, movie trailers, street interviews, and vlogs retrieve diverse Internet videos while limiting duplicate identities.
  • Data Pre-processing: Automatic preprocessing detects 98 facial landmarks, filters faces smaller than 450 pixels, and splits clips when adjacent frames do not match in motion or identity.
  • Data Pre-processing: 512×512 is selected as the normalized resolution to avoid significant upsampling while keeping clips at a common training resolution.The reported source-resolution distribution is 0.6% at 450²∼512², 76.6% at 512²∼1024², and 22.7% above 1024².
  • Data Annotation: Annotation factors face videos into appearance, action, and emotion, representing time-invariant attributes, sequence-related actions, and high-level mental states.
  • Data Annotation: The final scheme contains 40 appearance, 35 action, and 8 emotion attributes, with multi-label appearance and action classes but single-label emotions.

4 Statistics

CelebV-HQ combines large scale, high resolution, rich facial annotations, natural attribute distributions, and diverse temporal content. Comparisons with image and video datasets indicate broad facial coverage, higher quality, and varied head movement.

  • Dataset comparison: CelebV-HQ provides richer information than image datasets by combining appearance annotations with action and emotion annotations in video.Its resolution is more than twice that of CelebA, with scale comparable to CelebA-HQ.
  • Attribute distribution: Attribute distributions are diverse and natural, with balanced hair colors, varied appearance and action attributes, and emotion proportions reflecting real-world data rather than laboratory control.The appearance distribution includes both common and long-tail attributes, while neutral, happiness, and sadness are the most represented emotions.
  • Image-dataset comparison: CelebV-HQ has smoother age distribution, an ethnicity distribution close to CelebA-HQ, and more even representation of Latino Hispanic, Asian, Middle-eastern, and African groups.Its face-shape ratio distribution is also more uniform than CelebA-HQ’s.
  • Temporal and geometric diversity: CelebV-HQ contains more diverse head-pose and face-shape distributions than CelebA-HQ, covering both stable and substantially moving videos.The dataset includes videos with less than 20° of movement and videos ranging from 75° to 100°.
  • Video-dataset comparison: Compared with VoxCeleb2, CelebV-HQ offers higher image and video quality, more stable brightness, broader head-pose distributions, and richer facial action variation.The analysis evaluates quality with BRISQUE and VSFA, brightness variance, head pose, and Action Units; VoxCeleb2 is mainly composed of talking videos.

5 Evaluation

CelebV-HQ is evaluated with established baselines for unconditional video generation and video facial attribute editing. The experiments show temporal consistency, action-conditioned generation, and improved video editing consistency while preserving comparable image quality.

  • Unconditional Video Generation: Four unconditional video GANs are evaluated on CelebV-HQ using full-data and action-specific subset settings, with FVD and FID measuring video and image quality.The models are VideoGPT, MoCoGAN-HD, DIGAN, and StyleGAN-V.
  • Unconditional Video Generation: CelebV-HQ-trained MoCoGAN-HD, DIGAN, and StyleGAN-V generate temporally consistent videos, while action-specific subsets produce the corresponding facial actions.The benchmark compares these models across FaceForensics, Vox, MEAD, and CelebV-HQ.
  • Unconditional Video Generation: The benchmark reports similar model rankings across datasets, and models obtain good FVD/FID metrics compared with Vox at similar data size.The authors interpret this as evidence that CelebV-HQ supports more diverse and higher-quality generated results.
  • Video Facial Attribute Editing: Video facial editing evaluates StarGAN-v2 and MUNIT after adding a vanilla temporal constraint based on optical flow.The comparison uses the Gender appearance attribute and contrasts modified video versions with the original image models.
  • Video Facial Attribute Editing: The “Video” versions outperform the “Original” versions on FVD in all cases and achieve comparable FID, indicating improved temporal consistency without comparable image-quality loss.Qualitative results show hair-shape inconsistencies and jittering color blocks for the original image models.

6 Discussion

The discussion identifies future opportunities for video facial editing, coherent video alignment, and broader uses of CelebV-HQ. It also highlights temporal modeling and annotation-rich learning as open directions.

  • Empirical Insights: The growing prevalence of short videos motivates transferring face editing from static images to videos as an emerging research direction.The discussion cites applications such as TikTok and Snapchat while noting that current applications are mainly image-based.
  • Empirical Insights: An effective video alignment strategy must retain temporal information while aligning faces, because jittery alignment can harm coherence and misalignment can degrade generation quality.The authors suggest that jointly preserving temporal information and aligning faces may improve temporal consistency.
  • Future Work: CelebV-HQ may support video generation, editing, reenactment, face swapping, text-to-video generation, and neural-rendering tasks that rely on dataset scale and quality.The discussion also identifies attribute-based synthesis and disentanglement as uses of its facial annotations.
  • Future Work: Temporal modeling remains important because some video-generation methods emphasize frame-level quality while neglecting temporal information.The paper presents smooth and realistic video generation as requiring further investigation.
  • Future Work: CelebV-HQ could enable image tasks such as attribute recognition to be transferred to videos through learned spatio-temporal representations.

7 Conclusion

The paper introduces CelebV-HQ as a large-scale, diverse, high-quality video dataset with extensive facial annotations. Statistical analyses and two video tasks demonstrate its dataset properties and practical potential.

  • The dataset analysis demonstrates diversity across age, ethnicity, brightness, motion smoothness, pose diversity, and data quality.
  • Unconditional video generation and video facial attribute editing experiments demonstrate CelebV-HQ’s effectiveness and future potential.
  • The authors position CelebV-HQ as a source of new opportunities and challenges for related research directions and plan continued evolution of its scale, quality, and annotations.

A Data Pre-processing

CelebV-HQ uses detection, tracking, geometric cropping, and size normalization to preprocess videos. Additional analyses compare its attribute, duration, and action-unit distributions with established datasets.

  • Pre-processing Pipeline: Each frame is processed with bounding-box detection, identities are tracked, and bounding-box sequences are converted into minimum bounding rectangles.
  • Pre-processing Pipeline: Rectangles smaller than 512×512 are expanded to that size before videos are cropped using the resulting rectangles.
  • Dataset Comparisons: CelebV-HQ has a similar appearance-attribute distribution to CelebA-HQ, with most attributes close and no significant distribution deviation.
  • Dataset Comparisons: CelebV-HQ clips are shorter than Vox clips and all remain under 20 seconds because videos are truncated to limit attribute changes and support consistency and annotation accuracy.
  • Action-Unit Analysis: CelebV-HQ action units are smoother than VoxCeleb2 and more evenly distributed across different action-unit values.

C Additional Experiments

The additional experiments define a controlled evaluation protocol for video generation and editing, using fixed real and generated test sets and FID/FVD metrics. The reported editing comparison favors video-trained models on FVD while maintaining comparable FID.

  • FID and FVD assess image and video quality for video generation and editing models.The evaluation uses FID5 for image quality and FVD6 for video quality.
  • 2048 randomly selected videos form the real test set, with 2048 generated videos for unconditional generation and editing evaluations.Editing produces one fake result for each real test video.
  • The “Video” version achieves lower FVD scores and comparable FID performance than “Original”.The table marks lower values as better with “↓”.
  • 8192 images are used for FID testing by sampling 4 frames from each of the 2048 real and fake videos.

C.2 Additional Video Facial Attribute Editing Results

Additional editing experiments test Brown Hair with StarGAN-v2 and Eyeglasses with MUNIT. Adding temporal regularization improves realism and coherence, enabled by CelebV-HQ’s annotations and facial dynamics.

  • Additional Video Facial Attribute Editing Results: Brown Hair is edited with StarGAN-v2, while Eyeglasses is edited with MUNIT.
  • Additional Video Facial Attribute Editing Results: Adding a temporal regularization term improves StarGAN-v2 and MUNIT results in realism and coherence.
  • Additional Video Facial Attribute Editing Results: Table A1 reports quantitative results for two video facial editing baselines, comparing “Video” and “Original” versions.The “Video” version achieves lower FVD and comparable FID performance.
  • Additional Video Facial Attribute Editing Results: CelebV-HQ enables temporal regularization through its rich annotations and facial dynamics.

C.3 Experiment on labeled Vox

The labeled Vox experiment compares models trained on CelebV-HQ with algorithmically labeled Vox data. Models trained on CelebV-HQ perform better, indicating that automatic labeling is not a suitable substitute in this experiment.

  • Experiment on labeled Vox: Models trained on CelebV-HQ yield better performance than models using algorithmically labeled Vox data.
  • Experiment on labeled Vox: The experiment concludes that algorithmically labeling an existing dataset is not a suitable substitute for CelebV-HQ.
  • Experiment on labeled Vox: The complete attribute list is reported in Table A3.
Loading 2207.12393v1…