Source-linked AI summary

How well are open sourced AI-generated image detection models out-of-the-box: A comprehensive benchmark study

Simiao Ren, Yuchen Zhou, Xingyu Shen, Kidus Zewde, Tommy Duong, George Huang, Hatsanai, Tiangratanakul, Tsang, Ng, En Wei, Jiayu Xue

arXiv:2602.07814v1cs.CVcs.AI

TL;DR

Reliable AI-generated image detection is difficult to assess because existing benchmarks largely emphasize fine-tuned models rather than the pretrained detectors practitioners deploy. This paper conducts a comprehensive zero-shot benchmark across diverse detectors, datasets, and generators, finding unstable rankings, large performance variation, strong training-data effects, and severe weakness on modern commercial generators. The results support threat-specific detector selection rather than reliance on a universal benchmark winner.

  • Problem

    Existing benchmarks predominantly evaluate fine-tuned detectors, leaving out-of-the-box performance—the common deployment setting—insufficiently understood.

  • Method

    The study evaluates pretrained detectors across diverse datasets and generators using a comprehensive zero-shot benchmark with standardized statistical analysis.

  • Results

    Detector rankings are unstable and performance varies widely: Community-Forensics reaches 75.0% mean accuracy, while the weakest detector reaches 37.5%.

  • Takeaways & Limitations

    Detector choice should be matched to the specific threat landscape rather than based on published benchmark performance alone.

  • Takeaways & Limitations

    The evaluation is a snapshot as of mid-2024 and covers only detectors with publicly available pretrained weights.

Abstract

from arXiv · show

As AI-generated images proliferate across digital platforms, reliable detection methods have become critical for combating misinformation and maintaining content authenticity. While numerous deepfake detection methods have been proposed, existing benchmarks predominantly evaluate fine-tuned models, leaving a critical gap in understanding out-of-the-box performance -- the most common deployment scenario for practitioners. We present the first comprehensive zero-shot evaluation of 16 state-of-the-art detection methods, comprising 23 pretrained detector variants (due to multiple released versions of certain detectors), across 12 diverse datasets, comprising 2.6~million image samples spanning 291 unique generators including modern diffusion models. Our systematic analysis reveals striking findings: (1)~no universal winner exists, with detector rankings exhibiting substantial instability (Spearman~$ρ$: 0.01 -- 0.87 across dataset pairs); (2)~a 37~percentage-point performance gap separates the best detector (75.0\% mean accuracy) from the worst (37.5\%); (3)~training data alignment critically impacts generalization, causing up to 20--60\% performance variance within architecturally identical detector families; (4)~modern commercial generators (Flux~Dev, Firefly~v4, Midjourney~v7) defeat most detectors, achieving only 18--30\% average accuracy; and (5)~we identify three systematic failure patterns affecting cross-dataset generalization. Statistical analysis confirms significant performance differences between detectors (Friedman test: $χ^2$=121.01, $p<10^{-16}$, Kendall~$W$=0.524). Our findings challenge the ``one-size-fits-all'' detector paradigm and provide actionable deployment guidelines, demonstrating that practitioners must carefully select detectors based on their specific threat landscape rather than relying on published benchmark performance.

1 Introduction

The paper addresses the deployment-relevant gap in zero-shot AI-generated image detection, where practitioners use pretrained detectors without target-specific retraining. Its benchmark finds strong context dependence, large performance differences, and substantial training-data effects.

  • Motivation: Out-of-the-box detector performance remains largely unexplored despite practitioners commonly deploying pretrained models without fine-tuning.Existing benchmarks mainly evaluate models trained or fine-tuned on target datasets.
  • Research Questions: The study investigates overall performance, ranking stability, generalization factors, and systematic failure patterns across diverse datasets.
  • Key Findings: Spearman correlations from 0.01 to 0.87 and statistically significant Friedman results show that detector rankings are highly dataset-dependent.The Friedman test reports χ2=121.01, p < 10^-16, and Kendall W=0.524.
  • Key Findings: 75.0% mean accuracy for Community-Forensics versus 37.5% for AIGCDetectBenchmark CNNSpot creates a 37 percentage-point performance gap.Even top-tier detectors show 20–30% standard deviation across evaluation scenarios.
  • Key Findings: Identical detector architectures vary by 20–60% with different training data, while Flux Dev, Firefly v4, and Midjourney v7 reach only 18–30% average detection accuracy.These findings connect detector generalization to training-data alignment and emerging generator capabilities.
  • Contributions: The paper contributes a comprehensive zero-shot benchmark, statistical generalization analysis, and a three-pattern failure taxonomy intended to support detector selection and future development.The benchmark is framed as a bridge between research evaluation and practical deployment.

2 Related Work

Prior work established detector taxonomies, standardized comparisons, and cross-dataset generalization concerns, but systematic zero-shot evaluation remains limited. Existing evidence shows that performance often degrades under domain shift and that training data and real-world conditions matter substantially.

  • Detection Methods: Deepfake detection methods span CNN, transformer, frequency-domain, and ensemble or hybrid architectural families.
  • Research Gap: Existing methods often achieve strong in-dataset performance, but their zero-shot capabilities remain largely unexplored in systematic benchmarks.
  • Benchmark Datasets: Modern datasets increasingly include diffusion and commercial generators, although rapidly evolving generation techniques continue to outpace evaluation datasets.
  • Generalization: Prior studies report accuracy drops from above 95% to below 50% across datasets and find no model consistently exceeding 80% AUC across datasets.
  • Generalization: Reported barriers include dataset-specific artifacts, GAN-to-diffusion training mismatch, fairness issues, and domain shift, while proposed remedies include robust alignment, adversarial training, augmentation, fusion, and ensembles.
  • Prior Benchmarks: Earlier benchmarks provide standardized protocols and broad trained-model comparisons, but commonly rely on fine-tuning or retraining rather than pretrained detectors used as-is.

3 Experimental Setup

The experimental setup compares pretrained detectors and diverse datasets under a strictly standardized zero-shot protocol. It combines balanced evaluation, fixed thresholds, threshold-independent AUC, and complementary statistical analyses.

  • Detector Selection: The benchmark evaluates 16 detection methods and 23 independently treated pretrained detector instances selected for public weights, architectural diversity, recency, and established benchmark performance.
  • Detector Selection: The detector pool covers CNN, transformer, frequency-domain, and ensemble or hybrid families, enabling architectural and training-strategy comparisons.
  • Dataset Selection: The datasets span GAN, diffusion, and commercial API generators across 2021–2025, with approximately 2.6 million images and 46 distinct generators in MNW fake.
  • Evaluation Protocol: Evaluation uses uniformly subsampled and balanced splits, standard detector-specific normalization, fixed 0.5 thresholds, and AUC alongside accuracy.
  • Statistical Analysis: The analysis applies Friedman testing, Spearman rank correlation, and coefficient of variation to quantify significance, ranking stability, and performance variability.The coefficient of variation is defined as CV = σ/µ.

4 Results

Across 23 pretrained detector variants and 12 datasets, zero-shot performance is highly heterogeneous and dataset-dependent: detector rankings shift substantially, dataset difficulty varies widely, and accuracy and AUC generally agree.

  • Overall performance: 75.0% mean accuracy separates the best detector from the worst detector’s 37.5% mean, demonstrating large zero-shot performance heterogeneity.Community-Forensics achieves 75.0% mean accuracy, while AIGCDetectBenchmark CNNSpot reaches 37.5%.
  • Overall performance: χ2=121.01 and Kendall’s W=0.524 indicate statistically significant and substantial differences among detectors.The Friedman test reports p = 1.85×10^-16 with df=18.
  • Ranking stability: No detector ranks first across all datasets: Community-Forensics leads eight datasets, while PatchCraft and DRCT variants lead elsewhere.The best detector on one dataset can fall substantially on another, contradicting a universally best method.
  • Ranking stability: 0.01–0.87 Spearman rank correlations across dataset pairs show that datasets probe partially distinct detector capabilities.The median rank correlation is 0.52; GenImage and MNW fake have low correlation at approximately 0.23.
  • Dataset difficulty: 80.0% mean accuracy on GenImage contrasts with 37.6% on community forensics test, establishing a broad dataset-difficulty hierarchy.MNW fake and Nano-banana also remain difficult, at 41.7% and 44.6% mean accuracy.
  • Metric consistency: 0.82 Pearson correlation between AUC and accuracy indicates strong overall agreement, although some pairs combine high AUC with low accuracy.These outliers indicate ranking ability can coexist with poorly calibrated fixed thresholds.

5 Analysis

The analysis links zero-shot generalization primarily to training-data alignment, generator coverage, and ensemble diversity rather than architecture alone. Recent commercial and diffusion generators remain especially difficult, while detector performance declines for newer generators.

  • Training-data alignment: 99.7% accuracy for AIDE GenImage on GenImage illustrates direct alignment, whereas AIDE variants diverge sharply on diffusion-heavy datasets.The AIDE variants share ResNet-50 architectures but use different training data.
  • Training-data alignment: 20–60% performance variance within identical detector architectures is attributed to training-data alignment with the test generators.The effect can exceed variance between architectures and becomes extreme under severe training–test mismatch.
  • Generator difficulty: Flux Dev, Firefly v4, and Midjourney v7 average only 18–30% detection accuracy, whereas older generators remain substantially more detectable.Flux Dev reaches 21%, Firefly v4 18%, and Midjourney v7 24% mean accuracy.
  • Generator difficulty: Mean detection accuracy declines from approximately 79% for 2020–2021 generators to around 38% for 2024 models.A small subset of detectors still performs near-perfectly on some recent generators, but many fail almost entirely.
  • Generalization factors: Community-Forensics combines diverse training sources and five models, achieving 78.0% mean accuracy and rank standard deviation 1.27.Its broader coverage and ensemble aggregation are identified as factors supporting robust zero-shot generalization.
  • Architecture and features: Architecture groups show substantial within-group variation, indicating that training data and ensemble strategies matter more than any single architecture.Reported mean ranges are 51–72% for DRCT, 41–65% for frequency methods, and 67.5% for PatchCraft.

6 Discussion

The discussion argues that zero-shot detector performance is highly context-dependent: training data, generator age, dataset difficulty, and evaluation choices all shape generalization. It recommends broader benchmarks, adaptive and generator-agnostic methods, and practical trade-offs between accuracy, robustness, interpretability, and compute.

  • Implications: Ranking instability and training-data effects challenge the assumption of a universal best detector.Cross-dataset Spearman ρ ranges from 0.01–0.87, while training-data alignment produces 20–60% performance variance.
  • Implications: Training data explains more performance variance than architecture within the AIDE and DRCT families.The discussion favors diverse, carefully curated training datasets over incremental architectural refinements.
  • Implications: 78.0% mean accuracy is achieved by top ensemble methods, compared with 37–72% for single models.The authors attribute the advantage to complementary patterns that improve robustness to distribution shift.
  • Implications: Mean detection accuracy declines from roughly 79% for 2020 generators to approximately 38% for 2024 generators.The decline is uneven: a small subset of detectors remains highly accurate on some recent models, while many fail substantially.
  • Implications: Benchmark difficulty varies sharply, with GenImage averaging approximately 80% and the community forensics test approximately 38%.The authors argue that broader, more challenging datasets better represent deployment risks.
  • Limitations: The evaluation is limited by its mid-2024 snapshot, public-detector focus, image-only scope, fixed threshold, incomplete detector coverage, and partial generator coverage.These boundaries affect currency, proprietary-system generalization, video applicability, operating-point optimization, dataset comparability, and coverage of adversarial or commercial models.
  • Future Work: Future research should pursue rapid adaptation, continual learning, generator-agnostic features, multimodal detection, adversarial robustness, explainability, and resource efficiency.These directions target evolving generators, changing artifacts, broader media types, adaptive attacks, user trust, and deployment cost.

7 Conclusion

The paper presents a large-scale zero-shot benchmark showing that detector performance varies substantially across datasets, training data, and modern generators. It concludes that real-world deployment requires representative validation, context-specific selection, and complementary defenses rather than reliance on published benchmark results alone.

  • Conclusion: The benchmark evaluates 16 state-of-the-art detectors across 12 datasets, 2.6 million samples, and 291 unique generators.It is presented as the first comprehensive zero-shot evaluation of AI-generated image detectors.
  • Conclusion: Detector rankings are unstable, performance differs by 37 percentage points between best and worst detectors, and training data explains 20–60% variance.The reported ranking correlations range from Spearman ρ 0.01–0.87 across datasets.
  • Conclusion: Flux Dev, Firefly v4, and Midjourney v7 defeat most detectors, reaching only 18–30% accuracy.The conclusion identifies three systematic failure patterns affecting cross-dataset generalization.
  • Deployment: Community-Forensics reaches 75.0% mean accuracy but only 35–42% on newest commercial generators.This contrast motivates sustained research and multilayered defense strategies.
  • Deployment: Practitioners should validate detectors on representative target data, use ensembles when feasible, and temper expectations for state-of-the-art generators.The paper frames technical detection as one component alongside provenance, platform policies, and media literacy.

A.1 Complete Performance Matrix

The complete performance matrix reports accuracy values for 21 detectors or variants across 12 datasets.

  • Performance Matrix: The complete matrix contains all accuracy values for 21×12 detector–dataset combinations.It provides the tabulated basis for comparing detector performance across datasets.

A.2 Statistical Test Details

The study uses Friedman testing, Kendall’s W, and pairwise Spearman correlations to assess detector differences and ranking consistency across datasets.

  • Friedman Test: Friedman testing compares detector performance across datasets using ranks aggregated over the evaluation set.The analysis defines N as the number of datasets, k as the number of detectors, and Rj as detector j’s summed ranks.
  • Kendall’s W: Kendall’s W measures concordance in detector rankings across datasets, with the reported effect indicating substantial agreement despite ranking fluctuations.
  • Spearman Rank Correlation: Spearman’s ρ is computed for each dataset pair from detector rank differences, with the resulting correlations visualized as a matrix.For each pair, di is the rank difference for detector i and n is the number of detectors evaluated on both datasets.

A.3 Per-Dataset Rankings

The paper reports complete detector performance and ranking tables across datasets, alongside supplementary visualizations, implementation details, and data or code-release information.

  • Resources: The study identifies publicly available datasets and states that detector weights and evaluation code will be released.
  • Supplementary Analysis: The supplementary materials provide per-generator heatmaps, ROC curves, confusion matrices, training-data composition analysis, and temporal generalization curves.
  • Performance Matrix: Table 4 reports accuracy for all 19 detectors across all 12 datasets, highlighting the best performance per dataset.Missing values indicate detector–dataset combinations that were not evaluated.
  • Per-Dataset Rankings: Table 5 lists detector rankings from 1=best to 21=worst across datasets and reports rank standard deviation as a stability measure.
Loading 2602.07814v1…