Source-linked AI summary

Micro-Defects Expose Macro-Fakes: Detecting AI-Generated Images via Local Distributional Shifts

Boxuan Zhang, Jianing Zhu, Qifan Wang, Jiang Liu, Ruixiang Tang

arXiv:2605.09296v1cs.CVcs.AIcs.LG

TL;DR

AI-generated images are increasingly difficult to distinguish from natural images because modern generators leave sparse, localized forensic deviations that global representations may miss. MDMF addresses this with a patch-based PFS forensic space and MMD-based distributional detection, and experiments report strong, stable performance across diverse benchmarks. The analysis also identifies an optimal patch granularity rather than unbounded gains from increasing patch count.

  • Problem

    Modern generators leave sparse, localized artifacts, while global detectors can be dominated by semantic content rather than localized forensic deviations.

  • Method

    MDMF uses learnable PFS to reparameterize semantic patch embeddings into a forensic space, then applies MMD to quantify distributional discrepancies against reference real images.

  • Results

    Across widely used benchmarks, MDMF consistently achieves strong and stable detection performance and generalizes across diverse generative architectures and training paradigms.

  • Takeaways & Limitations

    Patch-wise distributional modeling captures subtle localized forensic signals and aggregates them into robust macro-level detection signals across emerging and conventional generative models.

  • Takeaways & Limitations

    Patch-wise advantages are not unbounded: finite-sample estimation and patch-resolution effects produce an optimal patch granularity.

Abstract

from arXiv · show

Recent generative models can produce images that appear highly realistic, raising challenges in distinguishing real and AI-generated images. Yet existing detectors based on pre-trained feature extractors tend to over-rely on global semantics, limiting sensitivity to the critical micro-defects. In this work, we propose Micro-Defects expose Macro-Fakes (MDMF), a local distribution-aware detection framework that amplifies micro-scale statistical irregularities into macro-level distributional discrepancies. To avoid localized forensic cues being diluted by plain aggregation, we introduce a learnable Patch Forensic Signature that projects semantic patch embeddings into a compact forensic latent space. We then use Maximum Mean Discrepancy (MMD) to quantify distributional discrepancies between generated and real images. Our theory-grounded analysis shows that patch-wise modeling yields provably larger discrepancies when localized forensic signals are present in generated images, enabling more reliable separation from real images. Extensive experiments demonstrate that MDMF consistently outperforms baseline detectors across multiple benchmarks, validating its general effectiveness. Project page: https://zbox1005.github.io/MDMF-project/

1 Introduction

MDMF addresses the difficulty of detecting realistic AI-generated images by modeling localized forensic evidence rather than relying on semantic-dominant global representations. It introduces PFS and MMD to amplify subtle statistical irregularities into image-level distributional discrepancies, with theory and experiments supporting reliable separation.

  • Modern generative models produce realistic images, making reliable separation of AI-generated and natural images increasingly challenging and important.
  • Existing image-level detectors can over-rely on global semantics, reducing sensitivity to sparse, localized forensic traces.
  • MDMF models images as collections of localized evidence and uses distributional discrepancies to amplify micro-scale statistical irregularities into image-level signals.
  • PFS reparameterizes semantic patch embeddings into a forensic space that deemphasizes semantic content while preserving and amplifying generation-induced statistical irregularities.
  • MMD quantifies discrepancies between patch-level PFS representations of test images and reference real images, while theoretical analysis establishes provable separation.
  • Across diverse benchmarks, MDMF achieves strong and stable detection performance and shows robustness to diverse generative architectures and training paradigms.

2 Micro-Defects Expose Macro-Fakes.

MDMF detects AI-generated images by converting localized patch-level forensic evidence into distributional discrepancies, rather than relying on a single globally pooled representation. It combines a learnable Patch Forensic Signature with MMD, with theory showing positive separation when generated images contain localized defects.

  • AI-generated image detection asks whether a test image comes from the real-image distribution P or a generator-induced alternative distribution Q.
  • Global representations can be dominated by semantic content, biasing detection away from sparse, localized forensic deviations.
  • The Patch Forensic Signature maps semantic patch embeddings into a compact forensic space that deemphasizes semantic variation and amplifies generation-induced statistical deviations.
  • MDMF compares distributions of PFS patch signatures between real and generated images using MMD, avoiding dilution from plain image-level pooling.
  • Training and detection pipelines form PFS representations for real and generated images and use a learned projection and kernel-based discrepancy to produce detection scores.
  • Theoretical analysis shows that localized defects produce a positive PFS separation and larger generated-image MMD scores than real-image fluctuations when the separation dominates.

3 Experiments

Experiments evaluate MDMF across diverse benchmarks, architectures, perturbations, and aggregation choices. Results show strong generalization, robustness, and benefits from combining patch-level forensic representations with distributional aggregation.

  • Datasets: MDMF is evaluated on ImageNet, LSUN-Bedroom, GenImage, WildRF, LDMFakeDetect, and an OpenSora video-frame case study.The evaluation spans standard image benchmarks and generated video content.
  • Main Results: MDMF consistently performs strongly across nine ImageNet generators spanning diffusion, GAN, and transformer models.The results indicate robust generalization across diverse generative mechanisms, including recent diffusion models.
  • Ablation: PFS without MMD remains competitive, while MMD improves performance when combined with PFS but not with global pooling.This supports PFS as the representation that preserves forensic evidence for MMD-based amplification.
  • Further Analysis: Performance varies non-monotonically with patch size, remaining above baselines while favoring intermediate granularity over overly coarse or fine partitions.W = 56 lacks spatial resolution, whereas W = 16 weakens forensic evidence within each patch.
  • Robustness: MDMF outperforms F-ConV across DINOv2 backbones and retains higher AUROC under JPEG, blur, and noise perturbations.At the most severe levels, AUROC gaps are +3.0, +8.5, and +9.9 for JPEG, blur, and noise, respectively.
  • Qualitative Analysis: Grad-CAM shows localized responses on generated images and diffuse activations on real images, unlike global pooling’s semantically driven patterns.The visualization is consistent with MDMF capturing localized generation-induced irregularities.

4 Conclusion

The paper presents MDMF as a distributional detector built from localized visual evidence. Its PFS representation and MMD aggregation are supported by theory and experiments across multiple benchmarks.

  • Conclusion: MDMF models images as collections of localized evidence rather than single global feature vectors.PFS suppresses semantic invariances while amplifying generative artifacts, and MMD aggregates the resulting patch evidence.
  • Conclusion: Theoretical analysis establishes advantages for PFS and separation between real and generated images, while experiments demonstrate effectiveness and generalization.

Reproducibility Statement

The paper documents datasets, assumptions, implementation resources, and proof-related modeling conditions. The theoretical analysis relies on sub-Gaussian embeddings, weak spatial dependence, and regularity of the PFS mapping.

  • Data and Resources: Experiments use public ImageNet, LSUN-Bedroom, GenImage, WildRF, and LDMFakeDetect benchmarks plus OpenSora-generated video frames.
  • Assumption: The detector follows a training-based setting in which training uses a designated dataset and evaluation tests generalization across generators and benchmarks.
  • Reproducibility: Source code, training and evaluation scripts, and applicable pretrained checkpoints are included in the supplementary materials.
  • Theory Assumptions: The proofs assume sub-Gaussian patch embeddings, weak spatial dependence across patches, and second-order regularity of the PFS mapping.The dependence assumption is used when relating patchwise and global pooling.
  • Theory Assumptions: Under exponential mixing, the effective patch count scales linearly with the number of patches up to constants.The result is expressed as Keff = Θ(K).

A.3 Proof of Proposition 2.5

The proof shows that patch-wise PFS modeling amplifies second-order defect shifts relative to global pooling. It also establishes that finite-sample noise and defect dilution produce an optimal finite patch granularity.

  • Proposition 2.5: Weak spatial dependence is summarized through an effective patch count, with independent patches giving Keff = K and exponential mixing yielding Keff = Θ(K).
  • Proposition 2.5: Patch-wise PFS modeling amplifies the same second-order defect signature relative to global pooling by approximately Keff.Under exponential mixing, this advantage is linear in K up to constants.
  • Optimal Granularity: Finite-sample estimation and defect-power dilution imply that the signal-to-noise ratio has a finite maximizer rather than improving indefinitely with patch count.The patch advantage eventually saturates beyond an optimal granularity.
  • Optimal Granularity: When g(K) = cK^-η with η > 0 under exponential mixing, the signal-to-noise ratio is eventually non-increasing after a finite K⋆.
  • MMD Analysis: The Gaussian PFS-space surrogate is used for a closed-form MMD expression, while the positivity and monotonicity conclusions extend beyond the exact Gaussian case.

A.6 Proof of Theorem 2.7

The proof derives high-probability empirical MMD bounds for real and generated test images, then combines them to establish empirical separation when the population discrepancy is sufficiently large.

  • A.6 Proof of Theorem 2.7: The proof handles real and generated test-image distributions as separate cases before completing the empirical ordering argument.The real case uses identical distributions, whereas the generated case uses the population discrepancy between P and Q.
  • A.6 Proof of Theorem 2.7: A union bound makes both deviation inequalities hold simultaneously with probability at least 1−2δ.This joint event supports comparison of the empirical statistic across real and generated cases.
  • A.6 Proof of Theorem 2.7: If the population MMD gap is positive, the separation condition is satisfied once reference and test sample sizes are sufficiently large.The condition uses MMD2(P, Q; kω) > 0 and increasing sample sizes M and N.

C.1.1 Details of Image Benchmarks

The evaluation uses image and video benchmarks spanning standard datasets, diverse generators, social-media distortions, and zero-shot cross-generator testing, with specified preprocessing and detector protocols.

  • Image benchmarks: ImageNet and LSUN-Bedroom use 256 × 256 images from DGM-Eval, with LSUN inputs randomly cropped to 224 × 224.The benchmark generators include both diffusion and GAN families.
  • Image benchmarks: GenImage covers real ImageNet images and heterogeneous sources including Midjourney, Stable Diffusion, ADM, GLIDE, Wukong, VQDM, and BigGAN.Image resolutions vary across GenImage subsets.
  • Image benchmarks: WildRF contains real and AI-generated images collected from Reddit, X, and Facebook after authentic-content and AI-generated-content hashtag filtering.The benchmark is designed to reflect in-the-wild social-media data.
  • Image benchmarks: LDMFakeDetect trains detectors on Stable Diffusion v1.4 and evaluates zero-shot across nine modern generators.The benchmark includes engines such as Midjourney, FLUX, Kandinsky, Playground, and Würstchen.
  • Video benchmarks: OpenSora and MSR-VTT each contribute 3,275 sampled videos, with ten frames per video and 224 × 224 random-cropped inputs.OpenSora supplies generated videos, while MSR-VTT supplies natural video data.
  • Evaluation protocol: Main experiments use DINOv2 ViT-L/14 patch embeddings, W=32 pooling, learned PFS parameters, and jointly optimized kernel bandwidth.Training uses random crops and flips, while testing uses center crops.

D.1 Results on Additional Benchmarks

Additional benchmark results show strong cross-dataset, cross-generator, perturbation, encoder, and granularity performance, while ablations identify PFS-plus-MMD as the strongest configuration.

  • LSUN-Bedroom: MDMF achieves the best average AUROC on LSUN-Bedroom while maintaining highly competitive average AP across diverse generators.The result covers diffusion-based models and GAN variants.
  • GenImage: MDMF achieves the best average accuracy on GenImage across proprietary, diffusion, and GAN sources.The result is reported under heterogeneous generative sources.
  • WildRF: MDMF attains the highest mean ACC and AP across WildRF’s three social platforms, outperforming LaDeDa on both metrics.The localized cues remain discriminative after social-media compression and processing.
  • LDMFakeDetect: MDMF achieves the best average AUROC and AP on LDMFakeDetect when trained on SD v1.4 and evaluated zero-shot on remaining generators.This supports cross-generator evaluation across nine modern diffusion engines.
  • Core ablation: Replacing global pooling with PFS-based patch evidence improves detection across generators, while coupling PFS with MMD produces the best performance.Simple mean, max, and top-k PFS aggregations already outperform global baselines.
  • Patch granularity: Intermediate patch granularity, such as W=32, offers the best overall trade-off between overly fine and overly coarse partitions.Fine partitions increase variability, whereas coarse partitions can miss sparse localized artifacts.
  • Robustness and encoders: MDMF consistently improves over F-ConV across DINOv2 encoder variants and retains higher AUROC and AP under JPEG, blur, and noise perturbations.The perturbation comparison uses clean reference images and corrupted test inputs.

D.8 Full Results for the Comparison with Patch-Level Hard Voting

The hard-voting comparison isolates aggregation by sharing the detector backbone and patch evidence, then contrasts threshold-sensitive binary patch decisions with continuous MMD aggregation.

  • Comparison setup: Hard voting independently classifies patches and combines binary decisions into an image-level fake-patch ratio.The comparison shares model components with MDMF except for aggregation.
  • Threshold dependence: Hard voting requires both a per-patch cutoff θpatch and an image-level threshold τ, whereas MDMF uses only τ.MDMF aggregates PFS evidence continuously through the MMD score.
  • Results: MDMF outperforms hard voting on every generator under every tested θpatch in both AUROC and AP.At θpatch=0.20, MDMF’s per-generator AUROC advantages include +1.42 on ADM and +3.02 on DiT-XL/2.
  • Results: Voting AUROC peaks at 94.43 for θpatch=0.20 and declines to 86.70 at 0.03 and 89.99 at 0.30.The single-peaked profile indicates sensitivity to the patch threshold.
  • Results: MDMF’s collective AUROC advantage is +2.23 at the voting peak and grows to +10.48 at θpatch=0.30.The comparison combines performance differences with the additional burden of tuning a generator-dependent patch threshold.

D.9 Failure Case Analysis: Borderline Real Images

MDMF can flag genuinely real images when photographic or stylistic conditions create local distributional shifts resembling generated-image artifacts. Its attention remains localized and interpretable, while comparisons across generators show a contrast with global pooling’s semantic focus.

  • Failure cases: Four genuinely real ImageNet images receive high MDMF fake-side scores, forming representative borderline cases.The cases include compressed foliage, an impressionist painting, a soft-focus bear photograph, and a black-and-white film photograph.
  • Why these images are borderline: Compression, painterly texture, defocus, monochrome film, and related processing shift local statistics away from MDMF’s clean color-photograph reference distribution.These conditions replace or suppress expected high-frequency or color information without making the images synthetic.
  • Attention and interpretation: MDMF highlights the specific regions whose local statistics deviate most, including a branch corner, painted face, fur boundary, and textureless bus body.The per-patch grid traces each misclassification to a concrete image region.
  • Attention and interpretation: These errors are design-consistent: MDMF detects unusual local distributions, regardless of whether their cause is generation or real-world photographic processing.The same operating principle that identifies synthetic artifacts also responds to non-synthetic post-processing artifacts.
  • Cross-generator qualitative comparison: Across ADM, ADMG, LDM, and DiT visualizations, global pooling mainly attends to semantic regions, whereas MDMF shows greater sensitivity to localized artifacts.The global baseline exhibits similar attention patterns on real and generated samples, while MDMF produces more localized responses.

F Limitations and Discussion

MDMF has two practical boundaries: inference depends on a real-image reference set, and severe perturbations still reduce performance for all evaluated detectors. The authors also frame reliable detection as relevant to limiting harms from synthetic visual content.

  • Practical considerations: MDMF requires a small reference set of real images at inference, so strict standalone deployment without test-time real images is outside its scope.Performance remains essentially stable from 1k to 10k references, but the operational dependency remains.
  • Robustness boundary: Severe perturbations, such as Gaussian blur with σ=5, cause a non-trivial performance drop for MDMF and all evaluated detectors.Real images with strong compression or denoising artifacts can have PFS distributions resembling generated samples.
  • Broader relevance: More reliable detection could help platforms, fact-checkers, and users flag synthetic content and support the integrity of online discourse.The paper connects this goal to risks including visual disinformation, identity impersonation, and erosion of trust.
Loading 2605.09296v1…