Source-linked AI summary
Deepfake Detection by Human Crowds, Machines, and Machine-informed Crowds
Matthew Groh, Ziv Epstein, Chaz Firestone, Rosalind Picard
TL;DR
Deepfakes create a societal need to distinguish authentic from manipulated video, but the relative strengths of humans and machines remain uncertain. Across two online studies, the paper compares ordinary observers, crowds, and a leading computer-vision model while testing model assistance, video features, and face-processing interventions. Humans and the model are similarly accurate overall but make different errors; model feedback can improve human accuracy yet mislead them, while disrupting face processing harms humans more than the model.
Problem
Deepfakes undermine video’s evidentiary value, motivating evidence about how ordinary people, crowds, and automated systems compare at detecting manipulated media.
Method
Two online experiments compare human judgments with a leading computer-vision model, test optional model feedback, examine video-level features, and randomize emotion and face-processing interventions.
Results
Humans and the leading model are similarly accurate overall but make different errors; model feedback improves accuracy overall while inaccurate predictions often mislead participants.
Takeaways & Limitations
Face-processing interventions reduce human accuracy, whereas only inversion changes model performance, supporting a role for specialized face processing in human detection.
Takeaways & Limitations
Results generalize to ordinary people detecting a limited video sample, with only two political videos containing lip-syncing manipulations and minimal-context clips that may not represent persuasive deepfakes.
Abstract
from arXiv · showhide
The recent emergence of machine-manipulated media raises an important societal question: how can we know if a video that we watch is real or fake? In two online studies with 15,016 participants, we present authentic videos and deepfakes and ask participants to identify which is which. We compare the performance of ordinary human observers against the leading computer vision deepfake detection model and find them similarly accurate while making different kinds of mistakes. Together, participants with access to the model's prediction are more accurate than either alone, but inaccurate model predictions often decrease participants' accuracy. To probe the relative strengths and weaknesses of humans and machines as detectors of deepfakes, we examine human and machine performance across video-level features, and we evaluate the impact of pre-registered randomized interventions on deepfake detection. We find that manipulations designed to disrupt visual processing of faces hinder human participants' performance while mostly not affecting the model's performance, suggesting a role for specialized cognitive capacities in explaining human deepfake detection performance.
Introduction
Deepfakes undermine video’s traditional evidentiary value, motivating comparisons among ordinary observers, crowds, and automated detectors. The paper uses online experiments to test human and machine detection, collaboration, and the role of face processing.
- Motivation: Deepfakes manipulate faces to swap identities or alter speech, challenging video as evidence that an event occurred.The paper frames this as a societal challenge because visual media can no longer be treated as an unquestioned authenticity standard.
- Data and benchmark: The DFDC dataset contains 23,654 original videos and 104,500 corresponding deepfakes generated with seven visual synthesis techniques.The videos show unknown actors making uncontroversial statements in nondescript locations, minimizing auditory and contextual cues.
- Data and benchmark: The competition’s leading computer-vision model achieved 65% accuracy on 4,000 balanced holdout videos after outperforming 2,115 other teams’ models.The model detects fakes from static-frame face localization and feature encoding, making it the paper’s machine comparison benchmark.
- Research questions: The paper asks whether individuals, averaged crowds, and the leading model differ in accuracy, and whether model predictions help or hinder human judgments.It also examines video-level variation, emotional priming, and specialized face processing as possible sources of performance differences.
- Study design: The study tests whether humans’ specialized face-processing abilities contribute to detecting synthetic facial manipulations.Prior perceptual research motivates this mechanism, including reduced face-recognition accuracy for inverted or misaligned images.
- Study design: Two randomized online experiments compare paired real-versus-deepfake choices with single-video confidence judgments, including optional model feedback and processing interventions.The interventions target incidental emotion and face processing through inversion, misalignment, and occlusion.
Results
Across two experiments, human performance varied by task and video, while collective responses often matched or exceeded the leading model. Model-informed participants generally improved, but confidently incorrect model predictions could reduce accuracy; face-processing obstructions selectively impaired human detection.
- Individual versus machine: 82% of participants who saw at least 10 video pairs outperformed the model’s 65% holdout accuracy in Experiment 1.Half of the 56 pairs were identified correctly by more than 83% of participants, while 12 pairs fell below 65%.
- Individual versus machine: 66% of recruited participants and 69% of non-recruited participants correctly identified sampled holdout videos, versus 80% for the leading model.Among non-recruited participants who saw at least 10 videos, accuracy reached 72%.
- Crowd wisdom: The crowd mean reached 76% accuracy for recruited participants and 80% for non-recruited participants, rising to 86% among non-recruited participants who saw at least 10 videos.The leading model accurately identified 80% of videos, placing collective human performance on par with the model.
- Machine-informed crowds: Updating responses after seeing the model’s prediction increased recruited participants’ accurate identification from 66% to 73% of observations.Participants updated confidence in 24% of trials, crossing the 50% accuracy threshold in 12%.
- Machine-informed crowds: 2% and 8% model probabilities for Kim Jong-un and Vladimir Putin deepfakes were confidently incorrect, reducing participant accuracy after updating.Accuracy dropped from 56% to 34% for Kim Jong-un and from 70% to 55% for Vladimir Putin.
Discussion
Ordinary people and the leading model detect deepfakes with comparable overall accuracy, but their strengths differ across video contexts. Human performance depends on face processing and can be improved or impaired by model advice depending on its correctness.
- Human–machine comparison: 82% of participants outperformed the leading model in Experiment 1, while 13%–37% did so in the more challenging single-video Experiment 2.Aggregated Experiment 2 responses were as accurate as the model’s prediction.
- Human–machine comparison: Participants and the model made different errors: the model struggled with stylistically different political-leader videos, whereas humans generalized across those contexts.The model confidently misclassified the Kim Jong-un and Vladimir Putin deepfakes, while most participants identified them correctly.
- Video features: The model performed slightly better on grainy, blurry, or very dark videos, while participants were more robust when two actors appeared.Both groups performed similarly on standard-quality videos and detected flickering faces well; both were less accurate with dark-skinned actors, significantly so for participants.
- Interventions: Anger elicitation reduced participants’ accuracy for real videos by 5.2 percentage points, although it did not significantly affect overall accuracy.The treatment underperformed control participants on real-video identification in the preregistered follow-up analysis.
- Interventions: Inversion, misalignment, and partial occlusion decreased human accuracy, whereas only inversion changed the model’s performance.These findings support a role for specialized face processing in human deepfake detection; the model’s robustness to some obstructions may reflect overfitting.
Limitations
The study’s conclusions are bounded by its video sample, participant population, comparison model, and experimental treatment results. The authors also caution that incidental-emotion findings are not firm.
- Scope: The evaluation covered 167 videos, included few lip-sync manipulations, recruited ordinary rather than expert detectors, and compared participants with the leading 2020 model.These constraints limit generalization to expert performance, newer models, and broader deepfake types.
- Ecological validity: The videos primarily showed unknown people making non-controversial statements in nondescript settings, which may not represent persuasive or realistic deepfakes.Only four political-leader videos began to address more realistic examples, and the authors call for larger samples.
- Ecological validity: The experiments used a 50% deepfake base rate, unlike misinformation’s much lower prevalence in contemporary media ecosystems.Future work should test detection without foreknowledge of the base rate and within social-media contexts containing contextual information.
- Emotion evidence: The two experiments produced different emotion findings, so the authors draw no firm conclusions about incidental emotions in deepfake detection.Experiment 1 found no significant effect, while Experiment 2’s significant result was near the cutoff for authentic videos and nonsignificant for deepfakes.
Implications
The findings support combining human crowd wisdom with machine predictions, while recognizing that their strengths vary by video type. They also point toward face-sensitive and context-aware detection systems rather than visual classification alone.
- Human–machine complementarity: Humans and the leading model are similarly accurate overall but excel on different video classes.Participants perform better on political and attention-check videos, whereas the model performs slightly better on blurry, grainy, and very dark videos.
- Human–AI systems: Decision-support tools should carefully weight human and model predictions because incorrect model outputs can make people less accurate.The authors suggest explainable AI and video-subtype information rather than presenting only a model likelihood.
- Perceptual mechanisms: Specialized face processing helps humans detect deepfakes, so manipulations that disrupt this processing and models that learn similar processing deserve attention.
- Perceptual mechanisms: Visual cues remain useful, but authenticity judgments also involve context, world knowledge, critical reasoning, and belief updating.
Methods
The study was hosted on Detect Fakes, an online website used to run both experiments. Supplementary materials provide the interface details and the remaining methods.
- Study platform: The experiments were hosted on the Detect Fakes website, with the user interface shown in Supporting Information Figure S4.The authors state that the rest of the methods are described in the Supplementary Information section.
Supporting Information
The supporting materials document the experiments’ sampling, randomized video presentation, interventions, analysis models, participant handling, and supplementary tables. Together, they specify how human deepfake judgments and treatment effects were measured.
- Experiment 1: The pilot selected 56 video pairs from the DFDC training data, including difficult cases identified through model-confidence criteria.
- Experiment 1: Participants viewed one authentic and one manipulated version side-by-side, guessed which contained the deepfake, and received feedback after each choice.
- Experiment 1: Participants were randomly assigned video pairs without replacement, allowing repeated judgments across the 56-pair sample.
- Experiment 1: The pilot cross-randomized inversion, emotion-elicitation, and reflection interventions across participants and video trials.
- Analysis: Experiment 1 estimated treatment effects with video fixed effects and participant-clustered robust errors using a preregistered regression model.
- Experiment 2: Experiment 2 randomly sampled 50 videos from the DFDC holdout set, balanced between 25 deepfakes and 25 authentic videos, with some model-targeting distractions.
- Experiment 2: Participants completed an attention check, then viewed one randomly selected video assigned to control or holistic-processing obstruction conditions.