Source-linked AI summary
One Prompt Is Enough: Watermark Laundering Through Foundation Image Models
Jidong Yang, Qi Li, Wei Zong, Yang-Wai Chow, Willy Susilo, Huaike Yu, Chunpeng Wang, Suo Gao
TL;DR
Invisible watermark evaluations often omit prompt-conditioned reconstruction by public foundation image models, which can preserve content while undermining payload recovery. The paper formalizes and evaluates this threat with a joint payload–fidelity profile across models and watermark schemes, finding strong disruption in OpenAI models and high-fidelity vulnerability for DwtDct in Nano Banana 2. It concludes that foundation-model reconstruction is a distinct operational attack interface and a missing robustness condition.
Problem
Invisible watermark robustness evaluations typically focus on predefined perturbations and lack evidence about single-prompt reconstruction through public foundation image models.
Method
The paper defines watermark laundering and evaluates black-box prompt-conditioned reconstruction using BER together with separate visual, semantic, and image-quality fidelity measures.
Results
OpenAI models produce the strongest payload disruption across evaluated schemes, while Nano Banana 2 leaves DwtDct vulnerable under high-fidelity reconstruction.
Takeaways & Limitations
Foundation-model reconstruction should be treated as a distinct operational attack interface and a missing robustness condition in invisible watermark evaluation.
Takeaways & Limitations
The evidence is limited to three watermark schemes, six editing models, 100 images, one prompt family, and a conventional baseline suite without an additional diffusion reconstruction method.
Abstract
from arXiv · showhide
Invisible watermarks are typically evaluated against predefined perturbations such as compression, blur, noise, cropping, and denoising. Public foundation image models expose a distinct threat: an attacker can submit a watermarked image with a single reconstruction prompt and obtain a visually faithful output from which the invisible watermark can no longer be decoded reliably. We formalize this failure mode as watermark laundering and evaluate it using a joint payload-fidelity profile that combines bit error rate (BER) with visual and semantic preservation. Across six OpenAI and Google image editing models, three representative watermarking schemes, and 1,800 reconstructed outputs, we identify two complementary laundering regimes: OpenAI models produce the strongest payload disruption across the evaluated schemes, whereas Nano Banana 2 shows that DwtDct remains vulnerable under high-fidelity reconstruction. Prompt ablations show that no single removal-oriented instruction is necessary for payload disruption, indicating that the effect is primarily induced by the reconstruction pathway rather than by explicit attack wording. Comparisons with conventional attacks further show that prompt-conditioned reconstruction constitutes a distinct operational attack interface. These findings motivate foundation-model reconstruction as a missing robustness condition in invisible watermark evaluation.
1 Introduction
The paper defines watermark laundering as prompt-conditioned reconstruction that preserves visible or semantic content while making an original invisible payload unreliable. It evaluates this black-box threat and argues that foundation-model reconstruction should be included in watermark robustness testing.
- 1 Introduction: Watermark laundering preserves usable visual or semantic content while weakening the original watermark’s recoverability.The output must remain usable; payload disruption caused by severe image destruction does not qualify.
- 1 Introduction: The threat model uses a public foundation image editor accessed through one natural-language reconstruction prompt, without decoder, key, model, gradient, or feedback access.This distinguishes the setting from attacks requiring a local or controllable generative pipeline.
- 1 Introduction: The evaluation spans six public foundation image models, three watermarking schemes, and 1,800 reconstructed outputs.The study uses a stratified set of 100 images and examines complementary high-disruption and high-fidelity regimes.
- 1 Introduction: OpenAI models produce the strongest payload disruption, while Nano Banana 2 shows that DwtDct remains vulnerable during high-fidelity reconstruction.The results identify complementary laundering regimes rather than a single universal model behavior.
- 1 Introduction: The paper establishes payload–fidelity evaluation as a missing robustness condition for invisible watermark assessment rather than proposing a dedicated state-of-the-art removal attack.The framework jointly considers payload disruption and preservation of usable content.
- 1 Introduction: Prompt ablations find that explicit hidden-information-removal wording is unnecessary, while high-frequency residual analysis only partially explains payload disruption.Reconstruction, restoration, or artifact-removal framing can still weaken invisible evidence.
2 Threat Model and Theoretical Framework
The paper models watermark laundering as black-box prompt reconstruction that disrupts victim-payload recovery while preserving visual or semantic content. It formalizes this threat with information-bottleneck assumptions and a joint payload–fidelity evaluation profile.
- Watermark laundering occurs when reconstruction makes the victim decoder unreliable while the output remains visually faithful and semantically related to the input.
- The joint profile reports raw BER alongside fidelity to the watermarked input, fidelity to the clean image, semantic preservation, and no-reference quality.
- BER = 0.5 represents random binary recovery, while values above 0.5 are not treated as stronger disruption.
- The threat requires a single natural-language instruction, newly rendered pixels, and simultaneous payload disruption and content preservation.
- The study avoids a universal thresholded laundering-success rate because application-specific thresholds would obscure the BER–fidelity tradeoff.
- The behavioral bottleneck assumes scene-relevant information is retained while payload information conditional on visible content is small, without decoder-informed adaptation.
- Under uniform independent payload bits, vanishing information about each bit drives optimal recovery toward random decoding, whereas empirical results use raw BER from fixed victim decoders.
- Prompt wording need not explicitly request hidden-information removal when payload loss occurs before rendering, although wording can affect whether output fidelity remains sufficient.
3 Experimental Setup
The experiments evaluate single-prompt black-box reconstruction across six public image models and three watermarking schemes using a stratified 1,800-call protocol. Prompt ablations and separated fidelity metrics assess payload disruption without conflating source similarity, semantic preservation, and image quality.
- The model set includes GPT Image 1, GPT Image 1.5, GPT Image 2, Nano Banana, Nano Banana Pro, and Nano Banana 2.
- The victim schemes are DwtDct, DwtDctSvd, and RivaGAN, with clean references resized to 1024 × 1024 before embedding and editing.
- The factorial grid contains 3 × 6 = 18 watermark–model cells and 100 outputs per cell, totaling 1,800 primary API calls.
- The structured prompt constrains objects, layout, canvas size, luminance, color, regional geometry, visible content, hidden information, and output validation in one reconstruction request.
- Prompt ablations on DwtDct use Nano Banana 2 and GPT Image 2 while separately removing content, luminance/color, appearance/geometry, and hidden-information clauses.
- Reference, semantic, and no-reference metrics remain separate because image-quality improvements can coincide with greater source deviation, and BER is not interpreted alone.
4 Results
Across six foundation image models and three watermarking schemes, reconstruction produces complementary laundering regimes: OpenAI models maximize payload disruption, while Nano Banana 2 combines high fidelity with substantial DwtDct disruption. Prompt ablations and conventional-attack comparisons show that this threat is tied to the reconstruction interface rather than a single removal instruction.
- Joint BER–Fidelity Profile: GPT Image 1 reaches mean BER 0.4808, closest to random binary recovery, while GPT Image 2 retains substantial disruption with BER values of 0.4266, 0.4517, and 0.3620 for DwtDct, DwtDctSvd, and RivaGAN.These results come from 18 watermark–model configurations with N = 100 outputs per configuration.
- Joint BER–Fidelity Profile: Nano Banana 2 reaches mean PSNR 30.27 and mean SSIM 0.879 against the watermarked input, while DwtDct remains disrupted at BER 0.4110.This supports high-fidelity laundering within the evaluated setting because disruption persists without obvious image destruction.
- Joint BER–Fidelity Profile: Later model versions improve source fidelity while mean BER moves farther from 0.5, indicating that laundering is not monotonically stronger across model generations.Mean BER changes from 0.4808 to 0.4655 to 0.4134 across GPT Image versions and from 0.3282 to 0.2941 to 0.2667 across Nano Banana versions; DwtDct remains approximately 0.41 across Google models.
- Prompt Sensitivity of Watermark Laundering: Removing any tested prompt module keeps Nano Banana 2 DwtDct BER within 0.0023 of 0.4110, while fidelity changes more substantially.Removing appearance and geometry constraints lowers PSNR by 0.302, and the minimal prompt lowers PSNR by 2.369 on Nano Banana 2.
- Comparison with Conventional Attacks: The comparison separates foundation-model reconstruction from fixed local attacks: Nano Banana 2 provides higher-fidelity laundering, whereas Gaussian blur and BM3D preserve the clean reference more closely but are less consistently disruptive.The comparison is intended to distinguish the public reconstruction interface from fixed local operators and does not exhaust regeneration methods.
- High-Frequency Residual Analysis: Across 1,800 matched pairs, HFR correlates positively with raw BER, but weak within-model results limit it to one contributor rather than a complete explanation of decoder errors.The aggregate correlations are Pearson r = 0.3158 and Spearman ρ = 0.3607; HFR does not identify payload-carrying coefficients or decoder invariances.
5 Security Implications for Watermark Laundering
Watermark laundering creates a provenance-security mismatch: permitted reconstruction prompts can disrupt victim payloads while preserving usable content, and new provider markers do not restore the victim watermark or replace provenance.
- A visually benign recreate, restore, or clean instruction can repurpose permitted editing into disruption of the victim payload.
- Phrase blocking alone is insufficient because explicit hidden-information-removal wording is not required for the observed disruption.
- Provider policy compliance at the wording level does not guarantee provenance preservation when reconstruction discards the victim carrier.
- C2PA or Content Credentials attached to reconstructed outputs do not recover the victim payload or constitute provenance substitution.
- Defenses should bind authorization and provenance across transformations and test outputs for victim-signal loss rather than relying only on prompt moderation or new provider marks.
6 Discussion and Conclusion
The study finds that single-prompt reconstruction is a practical laundering interface within the evaluated setting, but its evidence is bounded by the tested schemes, models, prompts, baselines, and temporal snapshot. It therefore calls for reconstruction-based robustness evaluation without claiming universal provenance substitution or coverage of generative watermarks embedded within models.
- Single-prompt reconstruction can disrupt invisible victim payloads while returning usable outputs in the evaluated setting.
- OpenAI models produce the strongest disruption across schemes, while Nano Banana 2 shows DwtDct vulnerability at substantially higher fidelity.
- Future evaluations should pair foundation-model reconstruction with classical perturbation, denoising, and generative reconstruction baselines while reporting BER with visual and semantic fidelity.
- The evidence covers three watermark schemes, six editing models, 100 images, one prompt family, and conventional baselines without an additional diffusion reconstruction method.
- Formal confidence intervals and effect-size tests are absent, and the reported values form a temporal snapshot because model versions and interfaces evolve.
- The conclusion does not claim coverage of generative watermarks embedded within models or universal provenance substitution.