Source-linked AI summary
DF26: We Cannot Tell Fake From Real Anymore
Severyn Shykula, Andrii Yermakov, Ivan Samarskyi, Dmytro Mishkin, Jan Cech, Anastasiia Mishchuk
TL;DR
Modern text-to-video and image-to-video systems can produce realistic full-scene public-speaking videos that existing deepfake evaluations may not represent. DF26 builds a controlled benchmark of matched real and synthetic clips across seven generators, finding that both humans and state-of-the-art detectors struggle near chance on modern generated content.
Problem
Existing deepfake benchmarks do not adequately represent fully synthetic public-speaking videos from modern text-to-video and image-to-video generators.
Method
DF26 constructs matched real and synthetic public-speaking clips across three scenarios, seven generators, and controlled prompts, metadata, and generation settings.
Results
State-of-the-art detectors perform poorly on DF26, with several near random chance and substantial variation across generators and generation settings.
Takeaways & Limitations
DF26 motivates evaluation protocols that explicitly test robustness to modern generator distribution shifts and include commercial systems.
Takeaways & Limitations
DF26 is limited to visual modality and does not evaluate audio realism, speech quality, lip-sync consistency, or audio-visual synchronization.
Abstract
from arXiv · showhide
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.
1. Introduction
DF26 addresses the difficulty of detecting fully generated videos in single-person public-speaking scenarios, where existing benchmarks and human judgments provide limited evidence of robustness. It introduces a controlled benchmark spanning modern generators and reports near-chance human performance on its synthetic videos.
- Motivation: DF26 targets photorealistic, fully generated videos in single-person public-speaking settings that legacy face-manipulation benchmarks do not represent well.The benchmark covers direct-to-camera/casual videos, official footage, and studio interviews.
- Motivation: Existing detectors often fail to generalize from legacy benchmarks to unseen modern generators, motivating controlled evaluation of generator distribution shifts.Prior benchmarks cover face manipulation, speech-driven synthesis, or avatar-based generation without isolating full-scene T2V and I2V generation.
- Benchmark design: The benchmark pairs each real video with synthetic clips generated from aligned semantic context while controlling scenario, source, prompt, and generation configuration.This design varies the generator while preserving matched public-speaking content.
- Benchmark: DF26 contains 271 real and 2,420 generated clips across three public-speaking scenarios and seven modern video generators.Open-source systems are evaluated in supported text-to-video and image-to-video modes, while commercial systems use text-to-video.
- Findings: Human accuracy on DF26 deepfakes is 52.6%, near chance, compared with 74.5% on CelebDF++ and 69.8% on DeepSpeak v2.Accuracy on real videos was similar across the datasets.
2. Related work
Related benchmarks have expanded beyond face manipulation, but many still lack controlled metadata or full-scene coverage for modern text-to-video and image-to-video systems. DF26 addresses this gap with a metadata-rich benchmark focused on single-person public-speaking videos.
- Existing benchmarks: Legacy datasets primarily evaluate face swaps, facial reenactment, and localized facial artifacts rather than complete scenes synthesized by modern video generators.These datasets remain useful for controlled training and historical comparison.
- Existing benchmarks: In-the-wild benchmarks improve realism but often lack generator metadata, reproducible protocols, and controlled source splits needed to isolate detector failure factors.Uncontrolled variation can involve generator shift, source mismatch, compression, or scene content.
- DF26 positioning: DF26 targets full-scene single-person public-speaking videos generated by modern text-to-video and image-to-video systems, unlike talking-head benchmarks focused on facial animation.It also narrows the setting to a misinformation-relevant public-speaking context.
- DF26 positioning: DF26 provides generator, scenario, prompt, and generation-mode metadata to enable systematic cross-generator evaluation of detector generalization.This controlled design distinguishes it from broad open-domain and in-the-wild datasets.
3. DF26 Design
DF26 is designed as a controlled, evaluation-primary benchmark for visual detection of fully synthetic public-speaking videos. It standardizes clip properties and metadata while comparing matched real videos with outputs from multiple generators and source datasets.
- Dataset design: DF26 normalizes duration and resolution and filters overlays, multiple visible people, missing faces, and severe technical artifacts to reduce trivial detector cues.The controls focus evaluation on visual evidence of synthetic generation.
- Evaluation design: DF26 evaluates pretrained or externally trained detectors on held-out data without training or fine-tuning on the benchmark in the main protocol.Train-on-DF26 experiments are treated as separate diagnostic ablations.
- Dataset composition: The dataset contains 271 real clips across three public-speaking scenarios, each standardized as a 5-second 1280×720 MP4 with prompts, keyframes, and provenance metadata.The same real videos are used for per-generator comparisons.
- Raw sources and provenance: Real videos are sourced from OpenVid-1M, TalkingCelebs, and MAVOS-DD, with automatic filtering, semantic classification, and manual selection.OpenVid-1M primarily supports Direct-to-Camera/Casual and Studio Interview scenarios, while TalkingCelebs supports Official Statement clips.
4. Dataset creation pipeline
The DF26 creation pipeline progressively filters and curates real videos, assigns semantic scenarios, generates matched prompts, and produces controlled synthetic clips. It also supports detector analysis across the full benchmark and separate I2V/T2V subsets.
- Pipeline overview: The pipeline applies objective filtering, semantic classification, manual curation, prompt generation, synthetic generation, and quality control in sequence.This multi-stage process converts noisy raw pools into controlled benchmark samples.
- Filtering and normalization: Metadata filtering enforces 1280×720 or 1920×1080 resolution and a minimum duration of 5.0 seconds before normalization to 5 seconds at 1280×720.These constraints establish consistent input conditions for the benchmark.
- Content checks: OCR rejects videos containing text in any of ten evenly spaced frames, while face detection retains clips with no frame containing multiple faces and at least 80% of frames containing exactly one face.The checks limit shortcut cues and enforce the single-speaker setting.
- Manual curation: Manual curation verifies the single speaking person, target scenario, absence of overlays, technical requirements, visual quality, and lack of severe artifacts.Within scenarios, reviewers also encouraged diversity in speakers, backgrounds, framing, lighting, and recording style.
- Prompt generation: A detailed prompt derived from four uniformly sampled frames describes observable scene and appearance properties, reducing content mismatch between real and generated samples.The prompt controls person appearance, background, lighting, framing, camera style, and speaking context without source-dataset identifiers.
- Dataset distribution and evaluation: The real-video distribution is organized by scenario and source dataset, while detector precision–recall curves are reported for the full benchmark and separate OS I2V and OS T2V subsets.These views support comparisons across generation settings.
5. Deepfake generation
DF26 generates matched synthetic counterparts using seven recent systems across text-to-video and image-to-video settings, while controlling source, scenario, prompts, and generation configuration.
- Seven recent video AI systems generate counterparts spanning commercial and open-source models, with text-to-video and image-to-video modes where supported.
- The generation process uses documented checkpoints and inference settings for open-source models, including approximately 440 GPU-hours for 1,626 videos.
- Commercial models were accessed through Higgsfield in March 2026 and excluded from Official Statement generations because of platform restrictions.
- Commercial outputs were screened to contain no visible watermarks, logos, platform overlays, or other export markers.
- Table 5 compares cross-dataset video-level AUROC and EER for detectors trained on FaceForensics++, including performance on CDFv3 and DF26.
6. Benchmark
DF26 evaluates state-of-the-art detectors and human observers under cross-dataset and cross-generator shifts. Detector and human performance often approaches chance, although retraining on modern data improves commercial-generator detection.
- 6.1. Frame-based and temporal deepfake detectors: 48.2 and 61.6 AUROC on DF26 contrast with 94.3 and 92.3 on CDFv3 for temporal detectors, while GenD-PE reaches 69.7, making DF26 substantially harder.
- 6.2. Per-generator analysis of detectors: PwTF-DVD achieves 92.9 AUROC on Wan 2.2 T2V but performs no better than chance on HunyuanVideo 1.5, exposing strong generator-specific variation.
- 6.3. Retraining state-of-the-art detector on DF26: Retraining GenD-PE on HunyuanVideo 1.5 yields at least 93.1 AUROC on Grok Imagine 1.0, Veo 3.1, and Wan 2.6 samples.
- 6.4. Open-source vs. commercial generators and scenario analysis: Both temporal detectors perform substantially better on open-source than closed-source generators, but the analysis is diagnostic rather than causal.
- 6.4. Open-source vs. commercial generators and scenario analysis: DFD-FCG remains close to chance across all three scenarios, while PwTF-DVD varies by scenario and performs best on official statements.
- 6.5. Human study: Human fake-video accuracy is 52.6% on DF26 versus 74.5% on CDFv3 and 69.8% on DSv2, while real-video accuracy remains similar across datasets.
7. Dataset Limitations
DF26 has two stated limitations: its dataset is relatively small, and it evaluates only visual evidence rather than multimodal realism.
- Dataset scale: DF26 contains 2,691 videos, which is small relative to large-scale benchmarks.The authors consider scaling it through paid actors, similarly to DSv2.
- Limited to visual modality: DF26 does not evaluate audio realism, speech quality, lip-sync consistency, or audio-visual synchronization.The authors identify multimodal extension as future work.
8. Release policy
DF26 is distributed through controlled access, with use conditions that distinguish evaluation from permitted training and restrict the closed-source subset.
- Access requirements: DF26 access requires applicants to provide affiliation, research purpose, intended use, and agreement to the dataset terms.The default use is held-out evaluation of pretrained or externally trained detectors.
- Use restrictions: Training and fine-tuning are permitted only on the open-source subset under the predefined cross-generator protocol.The closed-source subset is restricted to evaluation-only use.
- Use restrictions: The closed-source subset may not be used for training commercial models, fine-tuning, distillation, or other model-improvement procedures.
9. Conclusion
DF26 evaluates vision-based detectors on modern generated public-speaking videos and finds substantial difficulty for both machines and humans. The results expose generator-specific failures and motivate stronger evaluation protocols.
- Benchmark scope: DF26 covers four closed-source text-to-video and three open-source image-to-video and text-to-video setups in public-speaking scenes.
- Detector evaluation: Several state-of-the-art detectors perform near random chance on DF26, with substantial variation across generators.Temporal detectors that score highly on CelebDF++ perform poorly on DF26.
- Detector evaluation: Per-generator analysis shows that aggregate benchmark scores can hide severe failures, including performance approaching or falling below chance on some generators.
- Human study: DF26 contains deepfakes that people find harder to distinguish from real samples than those in CelebDF++ or DeepSpeak v2.
- Implications: The benchmark motivates better evaluation protocols and more robust detection mechanisms to mitigate disinformation.
S1. Additional results of the human study
The additional human-study materials report accuracy and response-time analyses across DF26, CDFv3, and DSv2, including generator-level results and controlled testing procedures. LTX 2.3 distilled I2V produced the lowest human accuracy, while several supplementary figures document the study setup and response distributions.
- Additional results: The additional results cover accuracy distributions for fake and real subsets across DF26, CDFv3, and DSv2, plus generator-level results in Table S1.
- Generator-level results: 25.2% human accuracy was recorded for LTX 2.3 distilled I2V, comparable to the 24.0% rate of labeling DF26 real videos as fake.The comparison is matched because these I2V clips were conditioned on the first frames of the same real videos.
- Response times: Correct answers had a 6.5-second median response time, compared with 8.7 seconds for incorrect answers.
- Study materials: Supplementary figures show the response-time distribution by dataset, participant instructions, the task interface, and last-frame comparisons for real and generated videos.
- Study procedure: Participants viewed silent clips with up to 10 replays, while the interface remained identical and clip order randomized across datasets.
S3. Additional samples from DF26
Additional DF26 samples pair real videos with generated counterparts from all seven generators, using I2V and T2V where supported. Bootstrapped analyses report pooled, per-generator, and per-scenario detector metrics, with uncertainty patterns supporting generator-shift effects.
- Additional samples from DF26: Figure S6 pairs two additional real videos with generated counterparts from all seven generators in supported I2V and T2V modes.Frames are taken from each clip’s final frame, where conditioning-frame drift is largest.
- Temporal detector results: Bootstrapped 95% confidence intervals quantify uncertainty for two temporal detectors using 1,000 video-resampling iterations.Pooled, per-generator, and per-scenario metrics are reported in Tables S2–S4.
- Temporal detector results: Per-generator intervals for PwTF-DVD separate Wan 2.2 T2V from HunyuanVideo 1.5 in both modes, indicating the spread is not sampling noise.The cited comparison concerns interval non-overlap across those generator conditions.
- Temporal detector results: Per-scenario intervals overlap substantially for DFD-FCG, consistent with failures being driven by generator shift rather than public-speaking scenario.