Source-linked AI summary
BadWorld: Adversarial Attacks on World Models
Linghui Shen, Mingyue Cui, Xingyi Yang
TL;DR
Visual world models may be unstable under small input perturbations, especially when future controls are unknown. BadWorld introduces label-free velocity attacks with trajectory-adaptive optimization, revealing that imperceptible perturbations can severely corrupt rollouts and pose safety risks while offering a privacy application.
Problem
The robustness of world-model dynamics under small input perturbations, including unknown future controls, remains insufficiently assessed for safety-critical applications.
Method
BadWorld combines a self-supervised velocity attack with trajectory-adaptive bi-level optimization to learn imperceptible, control-agnostic perturbations without future videos or action knowledge.
Results
Across representative models with continuous and discrete controls, imperceptible adversarial images severely degraded rollouts through incomplete denoising, structural collapse, semantic drift, and inconsistent control.
Takeaways & Limitations
The findings expose safety risks for deploying visual world models and suggest imperceptible perturbations as a practical mechanism for protecting images from unauthorized interactive generation.
Abstract
from arXiv · showhide
Visual world models (VWMs) synthesize interactive, action-conditioned rollouts from a single context image. However, it remains an open question how robust these models are to adversarial perturbations. Standard adversarial attacks fail to assess this vulnerability because attackers lack ground-truth future videos and cannot predict subsequent user controls. We introduce BadWorld, a label-free adversarial framework tailored for autoregressive VWMs that systematically overcomes both constraints. First, to bypass the need for future supervision, we propose a self-supervised velocity attack that directly disrupts the early denoising dynamics of the model. Second, to ensure the attack generalizes across unpredictable user actions, we formulate a trajectory-adaptive bi-level optimization that actively mines hard control sequences to forge control-agnostic perturbations. Evaluated on representative VWMs with continuous and discrete controls, BadWorld exposes severe structural fragility. Visually indistinguishable adversarial images reliably trigger catastrophic degradation in future rollouts, leading to incomplete denoising, structural collapse, and control inconsistency. These findings reveal critical risks for deploying VWMs in safety-critical systems while highlighting a practical mechanism for privacy protection.
1 Introduction
BadWorld frames adversarial perturbations as a label-free stress test for the temporal robustness of interactive visual world models. It addresses missing future supervision and unknown future controls through self-supervised velocity attacks and trajectory-adaptive optimization, exposing fragility in representative models.
- Motivation: Visual world models synthesize action-conditioned future videos from a single context image, supporting interactive games, robotics, and autonomous driving.Their growing rollout coherence motivates testing whether learned dynamics remain stable under small input perturbations, especially in safety-critical applications.
- Challenges: Existing generative-model attacks assume fixed conditions, reference outputs, or localized degradation, unlike interactive autoregressive world models.World-model adversaries observe only one context image and must handle changing future camera paths, navigation commands, or discrete actions.
- Framework: BadWorld learns an imperceptible perturbation that drives future rollouts toward out-of-distribution behaviors without paired future videos or knowledge of future actions.The framework operates from a pretrained world model and a single clean context image.
- Self-supervised velocity attack: BadWorld attacks predicted velocity using the model’s denoising dynamics, early-denoising approximation, and a context-based history proxy to avoid unavailable ground truth.This self-supervised objective requires neither future videos nor action annotations.
- Trajectory-adaptive optimization: Trajectory-adaptive bi-level optimization searches for hard trajectories while updating perturbations against them, improving effectiveness across continuous camera and discrete action controls.The resulting control-agnostic adversarial image is harder to bypass by changing the action sequence, and the findings reveal robustness risks alongside a potential privacy application.
2 Related Work
Related work traces VWMs to interactive video simulation and frames adversarial generative-model research around vulnerability assessment and privacy protection. Recent video and world-model attacks broaden the threat surface to temporal dynamics, physical conditions, and automated attack search.
- Interactive world models: Diffusion and flow-matching advances have shifted video generation toward interactive simulation, while VWMs model physics and action- and history-conditioned transitions.Autoregressive adaptation of pretrained models is a prominent trend for viewpoint-aware rollouts.
- Adversarial generative modeling: Adversarial generative-model research typically targets structural vulnerability assessment and privacy protection through imperceptible perturbations or training-data poisoning.The cited passage describes early image-generation attacks and poisoning against unauthorized Text-to-Image personalization.
- Video and world-model attacks: Video-generation attacks exploit temporal dynamics through prompt backdoors, jailbreaking, adversarial trajectories, physical-condition perturbations, and automated attack search.These paradigms extend adversarial research beyond image-based generation to video and world-model settings.
3 Background
Autoregressive world models generate video chunks sequentially from context, history, controls, and prompts, but their rollout structure can amplify perturbations. BadWorld therefore targets label-free denoising dynamics and adapts perturbations to unknown future controls.
- Autoregressive World Models: Autoregressive world models decompose videos into K chunks generated from context, prior chunks, controls, and optional prompts.The control τ_i typically encodes camera motions or actions, while prior chunks are supplied through a sliding history window.
- Autoregressive World Models: Diffusion- or flow-matching-based chunk generators produce sequential self-rollouts, whose temporal dependencies can amplify small perturbations through error accumulation.Flow Matching interpolates target chunk latents between data and Gaussian noise and trains a velocity network to predict transport direction.
- Adversarial Objective: An ideal attack seeks a small bounded context-image perturbation that maximizes the difference between adversarial and clean rollouts under the same controls.Large rollout differences can manifest as visual artifacts, temporal inconsistency, motion collapse, or semantic drift.
- Attack Challenges: The attack lacks future-video supervision, so the rollout-level objective must be label-free and self-supervised.Only the context image is observed, with no predefined correct future rollout.
- Attack Challenges: Unknown and user-dependent future controls require perturbations that remain effective across diverse trajectories rather than a single fixed control sequence.BadWorld addresses this through trajectory-adaptive optimization that searches for hard control trajectories and optimizes against them.
4 Methodology
BADWORLD combines a label-free velocity attack with trajectory-adaptive bi-level optimization to disrupt autoregressive visual world models without future videos or fixed control sequences. It uses early-denoising and history-proxy approximations, hard-trajectory mining, and perturbation updates to produce control-robust adversarial contexts.
- Overview: BADWORLD has two components: a label-free velocity objective and trajectory-adaptive rollout optimization for effectiveness across control signals.The velocity objective addresses missing future supervision, while trajectory optimization mines hard control trajectories.
- Label-free velocity attack: The attack optimizes velocity predictions because corrupted denoising directions disrupt local generation and compound across autoregressive steps.Direct rollout discrepancy is unavailable without future videos and would require computationally prohibitive backpropagation through the entire sequence.
- Label-free velocity attack: Velocity-Max and Velocity-Min respectively amplify or suppress velocity magnitude, causing aggressive distortions or retained noise and incomplete structures.These localized denoising failures compound over long-horizon autoregressive rollouts.
- Label-free realization: Early-denoising queries and a repeated encoded-context history proxy enable velocity optimization without ground-truth future videos or explicit annotations.At early denoising timesteps, the interpolated latent is approximated by Gaussian noise; the proxy repeats E(x + δ) to the required history length.
- Trajectory-adaptive optimization: Trajectory-adaptive bi-level optimization alternates hard-trajectory mining with perturbation updates, yielding adversarial contexts that induce unstable rollouts across diverse controls.The outer loop finds trajectories resisting the current attack, while the inner loop updates δ using PGD; CMA-ES searches coherent camera-motion patterns without gradients.
5 Experiments
Experiments on Astra and Matrix-Game-2.0 show that Velocity-Min is the strongest attack objective, with Matrix-Game-2.0 broadly vulnerable and Astra more selective. Ablations and hard-sample tests show that early denoising, context-based history simulation, and trajectory-adaptive optimization improve attack effectiveness.
- Experimental setup: Experiments use 100 images per model, three camera or action sequences, and perturbation budget η = 0.05 to compare attack objectives.Astra uses landscape frames, while Matrix-Game-2.0 uses GTA gameplay frames; attacks are evaluated with VBench, VBench++, CLIP, and MEt3R.
- Attack-objective comparison: Velocity-Min is the strongest objective: it is selectively effective on Astra but consistently degrades quality across both benchmarks, while Matrix-Game-2.0 is vulnerable to more objectives.On Matrix-Game-2.0, Velocity-Min reduces background consistency to 0.845, imaging quality to 0.513, and reaches MEt3R of 0.264.
- Ablations: Increasing η from 0.03 to 0.10 consistently worsens video-quality metrics and raises MEt3R, but also makes perturbations more perceptible, motivating η = 0.05.The default budget balances attack potency with imperceptibility.
- Ablations: Without early-denoising timestep selection, attack efficacy collapses: Aesthetic Quality rises to 0.514 and MEt3R falls to 0.167.The result supports targeting initial structural formation to neutralize generative priors.
- Bi-level attack: Trajectory-adaptive bi-level optimization provides moderate gains on the full Astra benchmark and more strongly degrades the 10 most resilient samples across seven trajectories.On hard samples, it reduces Aesthetic Quality to 0.425 and produces more pronounced noise and geometric distortion than the baseline.
6 Conclusion
BadWorld is an adversarial framework for autoregressive visual world models that addresses missing future-video supervision and unpredictable future controls. It exposes critical safety risks for deploying VWMs.
- BadWorld is an adversarial framework designed for autoregressive visual world models (VWMs).
- A self-supervised velocity attack bypasses the need for ground-truth future videos by disrupting early denoising dynamics.
- Trajectory-adaptive bi-level optimization handles unpredictable future controls by mining hard control sequences to forge control-agnostic perturbations.
- These findings expose critical safety risks for visual world model deployment.
Supplementary Material · A Explanation of the Velocity Magnitude Objective
The supplement expands BadWorld with objective explanations, robustness evaluations, implementation details, and additional experiments. Section A shows that velocity magnitude directly controls visual quality and that reducing it most effectively disrupts autoregressive video coherence.
- Supplementary Material: The supplementary material covers objective explanations, robustness evaluations, implementation specifics, and additional experimental results.It organizes these additions across Sections A–D.
- A Explanation of the Velocity Magnitude Objective: In flow-matching frameworks, the denoising network predicts a velocity field specifying latent-state direction and rate of change.The supplement isolates velocity magnitude by scaling the predicted velocity during inference with a positive scalar s.
- A Explanation of the Velocity Magnitude Objective: Scaling velocity preserves prediction direction while changing only its magnitude before the scheduler updates the latent state.The update uses the scaled velocity s · ˆvθ at each denoising step.
- A Explanation of the Velocity Magnitude Objective: Velocity magnitude directly determines visual quality: s < 1 produces gray, blurry, structurally disordered outputs, while s > 1 causes saturation and pixel distortion.Under-scaling yields under-denoised outputs; over-scaling pushes latents toward extreme values and can introduce overshooting artifacts.
- A Explanation of the Velocity Magnitude Objective: At equivalent scaling intensity, reducing velocity magnitude disrupts video coherence more severely than increasing it.Weakened updates prevent stable-manifold reachability, allowing denoising errors to accumulate through autoregressive generation.
- A Explanation of the Velocity Magnitude Objective: Velocity-Min is superior to Velocity-Max for triggering rapid structural collapse in adversarial attacks.The finding follows the observed asymmetry between weakened and amplified velocity updates.
B Robustness and Transferability … C.1.4 Covariance Matrix Adaptation Evolution Strategy (CMA-ES)[23]
BadWorld remains effective under image preprocessing and transfers across models, while its trajectory search combines compact camera parameterization, stochastic sampling, adaptive hard-trajectory mining, and CMA-ES optimization. These procedures target physically plausible, control-generalizable trajectories and distill high-velocity candidates for adversarial optimization.
- B Robustness and Transferability: Across 10 context images and three camera trajectories, the method is evaluated for robustness and transferability.The evaluation follows the setup in Sec. C.2.
- B Robustness and Transferability: Under light Gaussian noise (σ = 0.5), Velocity-Min maintains an MEt3R score of 0.255, while JPEG compression and TV denoising only partially mitigate the attack.The passage also reports significantly low imaging quality under light Gaussian noise.
- B Robustness and Transferability: Cross-model evaluation between Matrix-Game-2.0 and Astra shows clear black-box transferability, with video quality metrics declining relative to the clean baseline.Adversarial context images are resized to match the target model’s input configuration before inference.
- C.1.1 Camera Trajectory Formulation: Camera controls use the triplet c_t = (ψ_t, f_t, s_t) for yaw, forward displacement, and lateral shift, mapped uniquely to a relative pose matrix.For T = 8 frames, the continuous trajectory is vectorized as τ ∈ R^3T and constrained to remain physically plausible and in-distribution.
- C.1.2 Stochastic Trajectory Sampling via Random Walk: Baseline self-supervised PGD samples a fresh trajectory at every step, using a uniformly drawn initial control followed by a temporally coherent Gaussian random walk.The resulting trajectories are clipped to satisfy variation and range constraints, supporting generalizability across camera motions.
- C.1.3 Trajectory-Adaptive Bi-Level Execution: The trajectory-adaptive bi-level attack searches for high-loss trajectories, evaluates candidates by Monte Carlo approximation, and maintains history-length-specific hard trajectory pools.The default is M = 1; random-walk sampling remains the standard choice for short horizons, while extended horizons query mined hard trajectories.
- C.1.4 Covariance Matrix Adaptation Evolution Strategy (CMA-ES)[23]: CMA-ES replaces prohibitive and unstable exact gradients with derivative-free evolutionary search over the continuous trajectory space.It samples λ candidates, projects them into the feasible region, updates the mean from the top µ candidates, and adapts covariance to temporal correlations in hard camera motions.
- C.1.4 Covariance Matrix Adaptation Evolution Strategy (CMA-ES)[23]: CMA-ES successively updates trajectories toward higher velocity norm and distills globally top-ranked candidates into the hard trajectory pool P_hard.The pool is then used by the inner PGD optimization.
C.2 Additional Experimental Details
The experiments adapt training, conditioning, inference, and evaluation to settings without paired future supervision while preserving diverse autoregressive rollouts. Astra and Matrix-Game 2.0 use model-specific controls and length-aware video-quality and consistency assessments.
- Training: Training restricts optimization to early denoising timesteps 950–1000, where target frames can be approximated as pure noise, and duplicates the context frame to simulate historical context.These adjustments address the absence of paired ground-truth videos and viewpoint annotations in realistic attack scenarios.
- Conditioning: Astra uses 30 GPT-4o-mini prompt variants with randomly sampled prompts and unique camera poses, whereas Matrix-Game 2.0 samples discrete actions without text prompts.Each prompt is semantically aligned with the context frames, and sampling occurs per attack or latent frame as specified.
- Evaluations: Astra is evaluated with standard VBench metrics, while Matrix-Game 2.0 uses VBench-Long and shorter segmented clips to assess video quality and contextual consistency.The length-aware strategy is intended to make evaluation more rigorous and reliable for longer-form Matrix-Game-2 videos.
D Extension Experiments on Matrix-Game-2 Variants
On Matrix-Game-2 Universal, BADWORLD uses a specified perturbation and inference configuration, and the Velocity-Min objective substantially degrades quantitative quality while inducing structural collapse in rollouts.
- Evaluation Setup: BADWORLD trains with a 0.05 attack budget, 0.002 step size, and 700 optimization steps under the Matrix-Game-2 Universal evaluation setup.Inference uses a local attention window of 6, 3 latent frames per chunk, 3 sampling steps, and 27 generated latent frames per input.
- Quantitative Results: 0.652 to 0.515 in Imaging Quality under the Velocity-Min objective demonstrates substantial degradation on Matrix-Game-2 Universal.The reported values compare the unperturbed and attacked conditions, respectively.
- Quantitative Results: 0.513 to 0.418 in Aesthetic Quality further confirms that Velocity-Min substantially degrades Matrix-Game-2 Universal rollouts.The reported values compare the unperturbed and attacked conditions, respectively.
- Qualitative Results: Qualitative results show that the adversarial perturbation induces structural collapse in Matrix-Game-2 Universal rollouts.The findings support BADWORLD’s generalization across different model configurations.
E Additional Qualitative Results
This section presents additional qualitative results for Astra and Matrix-Game-2.0, including visualizations of attack performance and behavior under different camera trajectories.
- Additional qualitative results: Additional qualitative results are provided for Astra and Matrix-Game-2.0.The section explicitly covers both visual world models.
- Astra: Figure 9 presents qualitative results on Astra.The figure is specifically labeled as qualitative results for Astra.
- Attack performance: Figures 10 and 11 show bi-level attack performance, including performance under different camera trajectories.Figure 10 concerns bi-level attack performance, while Figure 11 concerns different camera trajectories.
- Matrix-Game-2.0: Figure 12 presents qualitative results on Matrix-Game-2.0.The figure is specifically labeled as qualitative results for Matrix-Game-2.0.