Source-linked AI summary

SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

Junxiao Yang, Minghao Zhang, Xiaoce Wang, Haoran Liu, Shiyao Cui, Hongning Wang, Minlie Huang

arXiv:2606.03348v1cs.CVcs.AI

TL;DR

Synthetic credibility is a visual misinformation threat in which generated artifacts use realistic embedded text, authoritative formats, and circulation traces to appear evidential. The paper introduces SYNCRED-BENCH and FP450 to evaluate detection across diverse synthetic artifacts and real-image negatives. Existing systems and human annotators remain unreliable, especially under low false-positive-rate constraints.

  • Problem

    Realistic embedded text and coherent credible layouts create a visual misinformation threat that prior forgery and multimodal misinformation settings do not fully capture.

  • Method

    The paper constructs SYNCRED-BENCH, a benchmark of 600 claim-bearing AI-generated images spanning six credible-form categories and seven circulation styles, with FP450 real-image negatives.

  • Results

    15 MLLMs achieve only 10.5% average TPR under a 5% FPR constraint, while open-source AIGC detectors remain below 5% average TPR and commercial APIs reach 57.6% average accuracy.

  • Takeaways & Limitations

    Synthetic credibility remains difficult for current MLLMs, AIGC detectors, and humans, motivating detectors that reason beyond superficial credibility cues.

  • Takeaways & Limitations

    The benchmark focuses on English and Chinese images and may not transfer directly to other languages, regions, scripts, or institutional styles.

Abstract

from arXiv · show

Recent generative models can now produce visual artifacts with realistic embedded text and layouts, creating a new misinformation threat: synthetic credibility. We introduce SYNCRED-Bench, a benchmark of 600 AI-generated misinformation images balanced across six credible-form categories and seven fine-grained circulation styles, together with FP450, a real-image negative set for measuring false positives. Extensive evaluation shows that existing systems remain unreliable: under a 5% false-positive-rate constraint, 15 MLLMs achieve only 10.5% true positive rate (TPR), open-source AIGC detectors achieve less than 5%, and commercial APIs reach 57.6%. Human annotators also struggled to identify synthetic credibility, reaching only 63% TPR. These findings establish synthetic credibility as a severe and underexplored visual misinformation challenge, and provide a benchmark for developing detectors that reason beyond superficial credibility cues.

1 Introduction

Generative models can create self-contained visual artifacts that combine authoritative formats with realistic circulation traces, a threat termed synthetic credibility. SYNCRED-BENCH formalizes this challenge, and evaluations show that models, detectors, and humans remain unreliable.

  • Threat definition: Synthetic credibility describes generated artifacts that gain persuasiveness by imitating authoritative formats and traces of authentic circulation.Credible form includes notices and news layouts, while credible circulation includes scans, screen photographs, and online compression.
  • Benchmark: SYNCRED-BENCH contains 600 balanced AI-generated misinformation images spanning six credible-form categories and seven fine-grained circulation styles.FP450 adds 450 matched real images for measuring false positives.
  • Evaluation: 15 MLLMs achieve only 10.5% average TPR under a 5% false-positive-rate constraint, while their average positive-class accuracy is 31.2%.These results indicate that general multimodal systems remain unreliable for credibility-bearing forgeries.
  • Evaluation: Open-source AIGC detectors achieve less than 5% average TPR at 5% FPR, whereas commercial APIs reach 57.6% average accuracy.The evaluation covers specialized detectors as well as multimodal language models.
  • Evaluation: Human annotators reach 63.0% TPR with 27.0% FPR, showing that synthetic credibility challenges both automated systems and people.The paper presents this difficulty as a severe visual misinformation challenge.

2 Related Work

Prior research covers multimodal misinformation, synthetic-image detection, and document forensics, but these lines of work address different assumptions from synthetic credibility. The paper positions SYNCRED-BENCH as targeting fully generated, credibility-bearing artifacts rather than only mismatches or localized edits.

  • Multimodal misinformation detection: Multimodal misinformation detection commonly tests whether images and text jointly support claims, using textual, visual, social, contextual, retrieval, and world-knowledge cues.Later work also studies image-caption mismatch, out-of-context use, mixed-source distortions, and synthetic data.
  • Synthetic-image detection: Synthetic-image detection distinguishes real from generated images through fingerprints, reconstruction, frequency features, or foundation-model representations.Existing benchmarks also evaluate cross-generator and degradation robustness.
  • Visual-text and document forensics: Document-forensics methods typically detect localized edits in otherwise genuine materials, including scene text, documents, receipts, and identity records.These assumptions differ from artifacts generated as complete visual objects.
  • Positioning: SYNCRED-BENCH instead studies AI-generated claim-bearing artifacts that imitate credibility-bearing visual conventions.Its focus is the synthetic origin of a credibility-rich image rather than claim verification or localized tampering.

3 SYNCRED-BENCH Construction

SYNCRED-BENCH constructs a balanced dataset of synthetic claim-bearing artifacts along credible-form and credible-circulation axes, paired with real images for false-positive evaluation. Samples are generated from structured claim items, filtered for fidelity and consistency, and organized to preserve realistic form–style pairings.

  • Dataset: SYNCRED-BENCH contains 600 AI-generated images balanced across six credible-form categories, with 100 samples per category.FP450 provides 450 real-image negatives for false-positive evaluation.
  • Task scope: A synthetic credibility artifact must be generated from text, contain an interpretable embedded claim or record, and imitate at least one credibility-bearing visual convention.Examples include official notices, platform screenshots, certificates, receipts, and dashboards.
  • Credible form: The benchmark organizes positive samples by credible form, covering Media Layout, Institutional Notice, Platform Interface, Credential Record, Analytical Display, and Assessment Material.Each category includes multiple subtypes to avoid collapsing the benchmark into a small set of templates.
  • Credible circulation: Credible circulation models how artifacts appear to have been accessed, reproduced, photographed, or reshared, including scans, phone photos, screenshots, crops, and compressed reposts.These traces imply a plausible acquisition history.
  • Credible circulation: Seven fine-grained circulation styles span Native Rendering, Scanned Copy, Camera Copy, Fax Copy, Screen Photograph, Cropped View, and Online Compression.The distribution covers all styles while preserving realistic form–style pairings rather than enforcing equality across every cell.
  • Generation: Each positive sample begins with a self-contained claim item specifying claim, form category, subtype, language, and circulation style.The generated image is intended to communicate its core information without captions, webpages, or metadata.
  • Generation: Generation applies claim rendering, circulation styling, and rendering constraints for readable text, coherent layouts, and realistic typography.Invalid candidates are discarded and generation may be retried up to three times before each category reaches 100 valid samples.
  • Quality control: Five authors verify candidates for content fidelity, form consistency, circulation consistency, and benchmark validity after construction-based label initialization.FP450 is collected from matched real visual domains and manually edited to balance underrepresented styles.

4 Experiments

Experiments evaluate MLLM judges and dedicated AIGC detectors using accuracy and low-FPR TPR, revealing weak and uneven detection of synthetic credibility. Error analyses show that credible layouts, acquisition cues, and circulation styles often lead models to treat synthetic artifacts as authentic.

  • 4.1 Experimental Setup: The evaluation combines MLLM judges with dedicated AIGC detectors and reports positive-class accuracy, FPR, and TPR at constrained operating points.FP450 supplies real images for measuring false positives, while threshold sweeping selects TPR at an FPR not exceeding the specified budget.
  • 4.2 Main Results: 15 MLLMs average 31.2% positive-class accuracy and 10.5% TPR(5), showing poor detection under a 5% false-positive constraint.Claude Opus 4.6 is the strongest usable MLLM, reaching 69.5% accuracy with 5.0% FPR.
  • 4.2 Main Results: Dedicated AIGC detectors average 48.3% accuracy and 30.3% TPR(5), but calibration varies sharply across systems.Hive AI reaches 75.2% accuracy with 0.9% FPR, whereas AI-vs-Real reaches 78.7% accuracy with 69.1% FPR.
  • 4.3 Threshold sweeping reveals weak low-FPR operating points.: At 5% FPR, Claude Opus 4.6 reaches about 69.5% TPR and Sonnet 4.6 reaches 55.3%, while most other MLLMs remain below 30%.Several models improve mainly at looser FPR budgets, indicating poorly separated detection signals from real-image scores.
  • 4.4 Credible layout is frequently converted into authenticity evidence.: Structured layouts and UI/screenshot conventions appear in 66.1% and 61.1% of false-negative rationales, followed by typography at 46.9% and camera perspective/lighting at 46.0%.Per-model rationales vary, but judges commonly reinterpret coherent layout, content, and acquisition cues as authenticity evidence.
  • 4.5 Circulation styles can systematically distort authenticity judgments.: Camera Copy, Screen Capture, and Cropped View reduce detection relative to Native Rendering by 6.5, 11.1, and 10.3 percentage points.Scanned Copy and Fax-like Copy instead increase average detection by 4.1 and 9.4 percentage points.

5 Conclusion

The paper frames synthetic credibility as an emerging visual misinformation threat and introduces SYNCRED-BENCH to study it. Results indicate that current detectors remain unreliable, motivating systems that distinguish visual plausibility from provenance.

  • Synthetic credibility is a visual misinformation threat in which AI-generated artifacts imitate authoritative formats and realistic circulation traces.
  • The benchmark evaluates this threat using 600 claim-bearing AI-generated images spanning six credible-form categories and seven circulation styles, plus FP450 real images for false-positive evaluation.
  • Current MLLMs and AIGC detectors remain unreliable, particularly under low false-positive-rate constraints.
  • Future systems should combine provenance verification, watermarking, broader detector training data, and alignment that separates visual plausibility from image provenance.

Limitations

The benchmark provides initial evidence rather than exhaustive coverage of synthetic credibility. Its scope is constrained by language and cultural coverage, generator diversity, validation design, dataset scale, detection-only evaluation, and changing model behavior.

  • SYNCRED-BENCH is an initial rather than exhaustive benchmark, omitting additional formats, languages, cultural conventions, interfaces, and domain-specific credibility cues.
  • The dataset focuses on English and Chinese claim-bearing images, so conclusions may not transfer directly to different languages, regions, or institutional styles.
  • The synthetic set uses a single text-to-image pipeline, potentially introducing generator-specific artifacts, prompt regularities, and rendering biases.
  • Human verification checks construction-based labels but does not establish whether ordinary users perceive artifacts as credible, persuasive, or harmful.
  • The evaluation covers 600 generated images and 450 real-image false-positive samples, so limited scale may introduce randomness and benchmarking artifacts.
  • The benchmark measures synthetic-image detection rather than full claim veracity or provenance verification.
  • Results may change with model updates, policy changes, prompt formatting, preprocessing, and threshold calibration, while model rationales are not direct causal evidence of internal decisions.

Ethical Considerations

The work treats SYNCRED-BENCH as a dual-use research resource because credibility-bearing artifacts could support detection research while enabling deception. It therefore recommends controlled access and safeguards against misuse and false accusations.

  • Credibility-bearing artifacts can support detection research but may also enable misinformation, impersonation, fraud, or deception.
  • The authors recommend research-only access, identity and intent verification, and explicit restrictions and safety warnings for released materials.
  • Detector outputs should not be the sole basis for accusing a source of fabrication because authentic images can be mislabeled as AI-generated.
  • Real-world use should combine predictions with provenance signals, watermark or content-credential checks, source verification, human review, and claim-level fact-checking.
  • The real-image set raises copyright, attribution, privacy, and data-governance concerns requiring licensing, privacy protection, source documentation, and removal mechanisms.

A.1 Detailed Dataset composition.

The dataset combines balanced credible-form categories with broad but plausibility-constrained circulation styles, bilingual claims, matched real negatives, metadata for auditing, and controlled release. Construction also uses manual edits where web data underrepresents realistic circulation styles.

  • Dataset composition: The positive split contains 600 AI-generated images, balanced across six artifact categories with 100 images per category.
  • Dataset composition: The circulation axis covers how artifacts are rendered, reproduced, photographed, cropped, degraded, or reshared, without forcing uniformity across form–style cells.
  • Dataset composition: Reproduced Style unifies scanned, camera, and fax-like copies, while Circulated Content unifies cropped views and online circulation.
  • Dataset composition: Claims appear in English and Chinese across education, administration, finance, healthcare, commerce, platforms, rankings, and analytical reports.
  • FP450 Negative Set: FP450 is a matched real-image negative set covering the same broad visual space for false-positive evaluation.
  • FP450 Negative Set: Because web corpora underrepresent screen captures, camera copies, cropped views, and compressed styles, the authors manually edited images to emulate consistent presentation styles.
  • Metadata and evaluation: Each generated image has metadata for construction, stratified reporting, and error analysis, while detectors receive only the image.
  • Metadata and evaluation: The release provides documentation, taxonomies, schemas, splits, code, aggregate statistics, and predictions publicly, while restricting images and prompts to research use.

C Details of Human Detection Experiment

The human detection experiment used balanced samples from SynCred-Bench and FP450, with five independent judgments per question. Annotators achieved limited performance in identifying synthetic credibility.

  • 20 university students from varied academic backgrounds performed the annotation task.The group included 4 Ph.D. students and 16 undergraduates.
  • 100 SynCred-Bench samples and 100 FP450 samples were selected through category-balanced sampling.The two sets were combined and randomly shuffled into annotation questions.
  • Each question was independently evaluated by five participants.
  • 63% average accuracy was achieved by human annotators, while the best individual annotator reached no more than 80% overall accuracy.The annotations also exhibited a relatively high false positive rate.
  • The experiment reports anonymous subject-level annotation performance in Table 8.

E Models Used in Our Experiments

The experiments evaluated closed-source and open-source multimodal models alongside closed-source and open-source AIGC detectors. The listed systems span commercial APIs, named model families, and publicly available detector checkpoints.

  • Models: Closed-source models included GPT-5.4, GPT-4o, Grok-4.3, Claude Opus 4.6, Claude Sonnet 4.6, and Gemini variants.
  • Models: Open-source models included Qwen, Pixtral Large, and Llama-3.2-11B-Vision variants.
  • Detectors: Closed-source detectors included Sightengine, AI or Not, and Hive AI.
  • Detectors: Open-source detectors included AI-vs-Real, AI-vs-Human, and Deepfake-vs-Real checkpoints.

F Data Examples

The data examples cover false-negative MLLM judgments and diverse artifact and circulation styles across AI-generated and real document images. The examples emphasize how plausible layouts, text, interfaces, and capture artifacts can resemble authentic imagery.

  • Dataset examples: Figures 10 and 11 show circulation-style and artifact-type examples, while Figures 12 and 13 compare AI-generated and real images across provenance and capture styles.The examples include scanned copies, native renderings, compressed layouts, screen photographs, cropped notices, and scanned or camera-captured records.
  • MLLM failure examples: A GPT-5.4 judge labeled a synthetic Douban movie-page screenshot real with 0.98 confidence.Its rationale cited consistent UI layout, crisp text, coherent metadata, and absent visual artifacts.
  • MLLM failure examples: AI-generated screenshots, documents, presentation scenes, and broadcast-style images were repeatedly misclassified as real by MLLM judges.Rationales relied on plausible layouts, coherent typography, realistic camera artifacts, and familiar interface conventions.
  • MLLM failure examples: A Claude Sonnet 4.6 judge labeled a synthetic conference-room presentation image real with 0.97 confidence.The rationale treated projector glare, audience silhouettes, blur, and room lighting as evidence of a genuine photograph.
  • MLLM failure examples: A Grok-4.3 judge labeled a synthetic exam PDF screenshot real with 0.9 confidence.The rationale emphasized sharp text, consistent browser and PDF controls, and resemblance to Cambridge A-level physics papers.
Loading 2606.03348v1…