Source-linked AI summary
ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analysis
Yinan He, Bei Gan, Siyu Chen, Yichun Zhou, Guojun Yin, Luchuan Song, Lu Sheng, Jing Shao, Ziwei Liu
TL;DR
Existing face forgery datasets provide limited diversity or coarse-grained analysis, motivating a broader benchmark. ForgeryNet constructs a mega-scale, unified dataset spanning image and video classification plus spatial and temporal localization, and reports studies that help understand facial forgery in real-world scenarios.
Problem
Existing face forgery datasets have limited diversity or support only coarse-grained analysis, despite the need to determine whether and where media are manipulated.
Method
ForgeryNet constructs a mega-scale benchmark with image- and video-level data, 15 forgery approaches, diverse perturbations, and unified annotations across four analysis tasks.
Results
The benchmark’s extensive task results help researchers better understand facial forgery toward real-world scenarios.
Takeaways & Limitations
ForgeryNet provides a broader basis for studying image and video classification together with spatial and temporal forgery localization.
Abstract
from arXiv · showhide
The rapid progress of photorealistic synthesis techniques has reached at a critical point where the boundary between real and manipulated images starts to blur. Thus, benchmarking and advancing digital forgery analysis have become a pressing issue. However, existing face forgery datasets either have limited diversity or only support coarse-grained analysis. To counter this emerging threat, we construct the ForgeryNet dataset, an extremely large face forgery dataset with unified annotations in image- and video-level data across four tasks: 1) Image Forgery Classification, including two-way (real / fake), three-way (real / fake with identity-replaced forgery approaches / fake with identity-remained forgery approaches), and n-way (real and 15 respective forgery approaches) classification. 2) Spatial Forgery Localization, which segments the manipulated area of fake images compared to their corresponding source real images. 3) Video Forgery Classification, which re-defines the video-level forgery classification with manipulated frames in random positions. This task is important because attackers in real world are free to manipulate any target frame. and 4) Temporal Forgery Localization, to localize the temporal segments which are manipulated. ForgeryNet is by far the largest publicly available deep face forgery dataset in terms of data-scale (2.9 million images, 221,247 videos), manipulations (7 image-level approaches, 8 video-level approaches), perturbations (36 independent and more mixed perturbations) and annotations (6.3 million classification labels, 2.9 million manipulated area annotations and 221,247 temporal forgery segment labels). We perform extensive benchmarking and studies of existing face forensics methods and obtain several valuable observations.
1. Introduction
ForgeryNet addresses limited dataset diversity and coarse-grained evaluation by providing a mega-scale, comprehensively annotated benchmark for real-world face forgery analysis.
- Existing datasets often exceed 99% accuracy because their limited scale and diversity make benchmark performance saturate.They also commonly support classification without locating manipulated image regions or untrimmed-video segments.
- ForgeryNet combines image- and video-level data across four tasks for real-world digital forgery analysis.The tasks cover image classification, spatial localization, video classification, and temporal localization.
- ForgeryNet uses wild original data collected with diversity in angle, expression, identity, lighting, and scenario.
- The paper defines face forgery as learning-based modification of identity, expressions, or attributes in an image or video.This definition distinguishes face forgery from Cheap-Fakes and the narrower identity-swapping use of DeepFakes.
- ForgeryNet includes 15 learning-based forgery approaches spanning encoder-decoder, generative-adversarial, graphics-formation, and RNN/LSTM models.
- ForgeryNet applies 36 perturbations, including optical distortion, multiplicative noise, random compression, and blur, to model re-rendering challenges.
2. Related Works
Earlier face forgery datasets progressed from small, narrow collections to larger video datasets, but continued to lack manipulation diversity and comprehensive task annotations; ForgeryNet broadens both.
- First-generation datasets: Early datasets such as DF-TIMIT, UADFV, SwapMe, and FaceSwap used small collections centered mainly on face swapping.DF-TIMIT generated 640 videos, UADFV contained 98 videos, and SwapMe plus FaceSwap produced 2,010 images.
- ForgeryNet comparison: ForgeryNet compares datasets across scale, diversity, media level, manipulation approaches, and post-processing perturbations.Its data include image- and video-level samples, 15 manipulation approaches within four categories, and 36 perturbation types from four distortion kinds.
- Data and forgery diversity: ForgeryNet’s original-data examples come from four face datasets, while its sampled forgeries cover identity-remained and identity-replaced categories.
- Second-generation datasets: Second-generation datasets improved scale and quality but still lacked diversity in forgery approaches and task annotations.Examples include Google DeepFake Detection, Celeb-DF, and FaceForensics++, which primarily provide manipulated videos and classification-oriented data.
- Third-generation datasets: Third-generation datasets reached tens of thousands of videos and tens of millions of frames, yet practical localization remained necessary beyond classification.
3. ForgeryNet Construction
ForgeryNet is a large, diverse face-forgery dataset built with 15 manipulation approaches, realistic perturbations, and unified image- and video-level annotations spanning classification and localization tasks.
- Dataset scale and diversity: ForgeryNet provides over 2.9M still images, more than 220k video clips, 15 manipulation approaches, over 36 mix-perturbations, and 9.4M annotations.Seven approaches generate images only, while eight also generate videos.
- Forgery pipeline: The forgery pipeline prepares target and conditional-source representations, generates a forged target face, re-renders it into the full image, and applies perturbations.Conditional sources can include images, sequences, sketches, parsing masks, audio, labels, or noise; model variants include encoder-decoder, GAN, Pix2Pix, RNN/LSTM, and graphics formation architectures.
- Dataset scale and diversity: The dataset draws original data from CREMA-D, RAVDESS, VoxCeleb2, and AVSpeech to diversify identities, angles, expressions, and scenarios.Source resolutions range from 240p to 1080p, with face yaw angles from −90 to 90 degrees represented.
- Forgery categories: Forgery approaches are divided into identity-remained and identity-replaced categories, with reenactment, editing, transfer, swap, and stacked manipulations represented.Identity-remained approaches alter identity-agnostic or external attributes, whereas identity-replaced approaches preserve the source identity after transferring or swapping content.
- Forgery annotations: ForgeryNet supplies three image-label granularities and untrimmed videos that splice real and manipulated frames, enabling temporal segment localization.Image classification includes two-way, three-way, and 16-way labels; video data includes locations of manipulated segments.
- Forgery annotations: Spatial annotations localize manipulated face regions using pre-perturbation forgery distributions, while perturbations can modify the entire image beyond the face.The dataset also includes examples of real images, forgery images, and corresponding spatial annotations.
4. ForgeryNet Settings
ForgeryNet evaluates image and video forensics through controlled train-validation-test splits, intra- and cross-forgery protocols, and task-specific classification and localization metrics.
- Dataset preparation: Image- and video-level data are split into training, validation, and test subsets at a ratio close to 7:1:2, with matched identities and roughly balanced real-to-fake ratios.Forgery categories and distributions are illustrated for both sets.
- Classification protocols: Protocol 1 trains on all real and fake training data and evaluates on validation data across two-way, three-way, and n-way classification variants.The n-way setting uses n=9 for video classification, while image labels support the stated classification variants.
- Classification protocols: Protocol 2 tests cross-forgery generalization by training on one manipulation type and testing on others, using binary classification.The manipulation type may be general, such as identity-replaced, or specific, such as ATVG-Net.
- Evaluation metrics: Spatial localization is evaluated with two Intersection over Union variants and L1 distance using images paired with forgery masks.Video classification uses the same metrics as image classification, while temporal localization follows ActivityNet-style evaluation with segment boundaries and confidence values.
5. Image Forgery Analysis Benchmark
The image benchmark evaluates classification and spatial localization across increasingly fine-grained forgery settings. Results show greater classification difficulty with more categories, limited cross-forgery generalization, and stronger spatial localization from HRNet.
- Image Forgery Classification: Training on one forgery family generalizes poorly to unseen identity-replaced or identity-remained approaches.The cross-forgery protocol reports significant performance drops in both directions.
- Image Forgery Classification: Classification becomes more difficult as the number of categories increases, although mAP indicates higher discrimination ability.The benchmark reports accuracy, mAP, and AUC for three-way and 16-way settings.
- Image Forgery Classification: More auxiliary classification information slightly boosts F3-Net compared with training using only binary labels.The comparison concerns mapping multi-class training back to binary classification.
- Image Forgery Classification: Training on ATVG-Net, StyleGAN2, or BlendFace gives the best average cross-forgery generalization, while DiscoFaceGAN is most generalizable and SC-FEGAN hardest to generalize to.Approaches with stronger similarity tend to produce better cross-forgery performance, and same-meta-category approaches usually correlate more highly.
- Spatial Forgery Localization: HRNet outperforms the other spatial localization methods, exceeding them by more than 10% on IoUdiff at threshold 0.01.A slight beard change remains difficult to detect, and a real image is misjudged as manipulated in one example.
6. Video Forgery Analysis Benchmark
The video benchmark evaluates classification and temporal localization using video backbones and boundary-aware proposal methods. Video-based temporal localization substantially outperforms frame-based localization, with BMN achieving approximately 87 average AP.
- Video Forgery Classification: Video-level classification generally achieves higher accuracy and AUC than image-level evaluation.SlowFast obtains the best classification performance, while the smaller X3D-M also gives satisfying results.
- Video Forgery Classification: Cross-forgery video classification shows substantial performance drops when models trained on one identity category are tested on unseen approaches from the other category.The drops are reported as more significant than their image-level counterparts.
- Temporal Forgery Localization: Video-based temporal localization methods significantly outperform the frame-based method, demonstrating the importance of boundary-aware networks.The evaluated video-based models use BSN or BMN on X3D-M and SlowFast features.
- Temporal Forgery Localization: BMN outperforms BSN by large margins and achieves approximately 87 average AP on the validation set.The benchmark reports AP, AR, and mAP for temporal localization.
7. Conclusion
ForgeryNet is presented as a mega-scale benchmark for image- and video-level face forgery analysis. Its broader sources, manipulation variety, re-rendering processes, and annotations support classification and spatial and temporal localization studies.
- Conclusion: ForgeryNet provides a mega-scale benchmark covering image and video classification, spatial localization, and temporal localization.The paper positions these applications as tools for understanding facial forgery in real-world scenarios.
- Conclusion: Compared with existing datasets, ForgeryNet offers greater variety and more comprehensive coverage of wild sources, manipulation approaches, re-rendering processes, and annotations.The authors invite further forgery approaches and analysis on the dataset.
A. Original Data Collection
ForgeryNet uses diverse face datasets as original data to build a wild and varied forgery benchmark. The source material spans identities, demographics, emotions, expressions, actions, and other facial conditions.
- Original Data Collection: The original data are selected from four face datasets to diversify identities, angles, expressions, and actions.This design contrasts with prior datasets built from narrower briefing or television scenarios.
- Original Data Collection: CREMA-D contributes 7,442 clips from 91 actors aged 20 to 74, with varied ethnicities and six emotions.The passage identifies 48 male and 43 female actors.
- Original Data Collection: RAVDESS contributes 7,356 files from 24 professional actors covering eight emotions and two lexically matched statements.The files include both video footage and sound tracks.
B. Original Data Preprocessing
ForgeryNet preprocesses diverse source media, identifies target faces and attributes, applies 15 forgery approaches, and adds realistic re-rendering perturbations.
- Source data: 43,941 VoxCeleb2 and 43,584 AVSpeech videos longer than six seconds are randomly selected, then generally truncated to 6–10 seconds.
- Target preparation: Face tracking groups detections into identity-consistent face tubes and selects a target face for manipulation.
- Target preparation: Attribute labels unavailable in source data are predicted with Slim-CNN for conditional facial-attribute manipulation.
- Forgery generation: 15 forgery approaches span encoder-decoder, vanilla GAN, Pix2Pix, RNN/LSTM, and graphics-formation architectures.
- Re-rendering: Re-rendering aligns faces, performs color matching and Poisson blending, and applies motion blur, compression, or superresolution to video sequences.
- Perturbations: Approximately 98% of data receive 2–4 randomly selected mixed perturbations, while 1% receive a single perturbation and 1% remain unchanged.
F. ForgeryNet Annotation
ForgeryNet provides unified image- and video-level annotations for classification and localization, with metrics and benchmark models spanning these tasks.
- Image annotation: Image classification supports balanced two-way, three-way, and 16-way labels distinguishing real images and forgery categories.
- Image annotation: Spatial localization derives a pixelwise difference map between each forgery image and its corresponding real image, using the pre-perturbation map as ground truth.
- Video annotation: Untrimmed videos splice 1–4 manipulated segments into corresponding real videos, with each segment containing at least 9 frames.
- Video annotation: Video annotations include three classification label types, fragment labels, and temporal localization targets for manipulated segments.
- Evaluation: Classification uses balanced Accuracy, binary AUC, or multiclass mAP, while spatial localization uses IoU variants and L1 distance.
- Benchmark models: The benchmark includes 11 image-classification methods and three representative spatial-localization models, including Xception-based baselines and HRNet.
H.3. Implementation Details
The implementation benchmarks image, video, and localization models with task-specific losses, pretrained backbones, augmentation, and temporal proposal procedures.
- Training: Classification uses cross-entropy, while localization adds either soft-target binary cross-entropy or MSE segmentation loss selected by validation.
- Training: All image models use ImageNet pretraining and synchronous SGD with batch size 128 over 100k iterations.
- Evaluation: Image and video classification use the same metrics, whereas temporal localization evaluates AP at tIoU thresholds, average AP, and Average Recall@K.
- Video models: Video classification compares TSM, SlowFast, Slow-only, and X3D-M, while proposal methods pair SlowFast or X3D-M features with BSN or BMN.
- Temporal localization: Frame-based temporal localization thresholds Xception frame scores at 0.25 and groups compatible predictions across tolerance values {1, 3, 5, 7}.
- Temporal localization: Temporal proposal systems retain the top 10 predictions per video by confidence score.
I.4. More Experiments
Additional experiments examine augmentation, temporal shuffling, and temporal localization, showing limited dependence on continuous temporal flow and accurate boundary proposals in an example.
- Ablation studies: Video classification is less affected by augmentation than image classification, while temporal disruptions have considerable but not very major performance impact.
- Ablation studies: Weak temporal shuffling slightly boosts AUC relative to no shuffling in the reported experiment.
- Ablation studies: The shuffling result implies that video models may exploit information beyond continuous temporal flow.
- Temporal localization: In the Figure 11 example, BMN accurately localizes all endpoints of two manipulated segments, with similar predictions suppressed by SoftNMS.