Source-linked AI summary
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram Đorđević, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
TL;DR
Current video generators can produce plausible continuations, but existing evaluation does not establish that they recover the distribution of valid futures under the same observation and action. This paper formalizes probabilistic alignment, introduces PAWBench and PAWEval, and finds that eleven systems do not consistently match reference probabilities while covering valid outcomes; it also tests language, sampling, and fine-tuning as interventions.
Problem
Existing evaluations assess generated videos individually, leaving whether repeated rollouts recover the correct distribution of possible futures under fixed observations and actions unresolved.
Method
The paper formalizes probabilistic alignment and evaluates repeated video rollouts with PAWBench and PAWEval, while probing language, noise sampling, and fine-tuning interventions.
Results
No model consistently achieves accurate outcome probabilities, broad valid-future coverage, and reliable performance across scenes among eleven current video generators.
Takeaways & Limitations
Plausible individual videos do not imply probabilistically aligned world modeling, motivating distribution-level evaluation of video generators.
Takeaways & Limitations
PAWBench evaluates stochastic futures through terminal outcomes and controlled, visually parseable scenarios rather than full trajectories or interactive embodied environments.
Abstract
from arXiv · showhide
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.
1 INTRODUCTION
Video generators are increasingly treated as world models, but plausible individual rollouts do not establish that they represent the distribution of valid futures. PAWBench addresses this gap by evaluating repeated rollouts and probing whether language, sampling, or training can reshape predictive distributions.
- Motivation: Probabilistic alignment requires modeling both which futures are physically possible and how likely each is under a fixed observation and action.This matters for interaction and planning because decisions depend on possible consequences and their relative likelihoods.
- Evaluation gap: Existing evaluations judge videos independently for visual quality, coherence, alignment, and plausibility, leaving the induced distribution over futures untested.A model can therefore produce plausible individual videos while misrepresenting repeated-outcome behavior.
- Findings: Eleven current video generators fail to consistently achieve accurate probabilities, broad valid-future coverage, and reliable performance across scenes.More rollouts support the main sampling budget, and PAWEval agrees with human judgments on videos with clear terminal outcomes.
- Approach: PAWBench contains 50 scenarios across eight mechanism groups, while PAWEval converts repeated rollouts into empirical outcome distributions for calibration and coverage tests.The benchmark fixes the initial observation and action across rollouts and evaluates probability-mass alignment and valid-support recovery.
- Interventions: The paper probes language guidance, initial-noise sampling, and fine-tuning as distinct ways to affect model predictive distributions.The interventions separate inference-time steering and exploration from changes to the learned distribution.
2 PROBABILISTICALLY ALIGNED WORLD MODELING
The paper defines a world model as a conditional distribution over future trajectories and distinguishes support alignment from probability-mass alignment. It reduces trajectory outcomes to a finite terminal-outcome distribution for evaluation.
- World-model representation: A video world model M is represented by the conditional trajectory distribution P_M(τ | x, a), whose support and probabilities describe possible futures.Here x is the initial observation, a is the action, and τ denotes future trajectories.
- Support alignment: Support alignment asks whether the model realizes the distinct valid outcomes available under the same condition.This criterion detects collapse onto only a subset of possible futures.
- Probability-mass alignment: Probability-mass alignment additionally requires covered outcomes to occur in the correct proportions.A model may cover every outcome yet still assign incorrect probabilities.
- Outcome abstraction: Mapping each trajectory τ to its terminal outcome g(τ) induces an outcome distribution that abstracts away trajectory differences leading to the same result.This finite outcome representation makes distributional comparison tractable.
3 PAWBENCH
PAWBench measures probabilistic alignment through repeated rollouts under fixed observations and actions, separating calibrated probability matching from valid-support recovery. PAWEval maps readable videos to terminal outcomes and scores the resulting empirical distributions.
- Benchmark design: Each benchmark item fixes a source image and action prompt, specifies valid terminal outcomes, and repeats the same conditioning pair across model queries.The resulting videos are mapped to outcomes and aggregated into an empirical distribution.
- Scenario suites: PAWBench includes 50 scenarios across eight mechanism groups, divided into 25 PAW-Calibration and 25 PAW-Coverage scenarios.Calibration uses analytically specified or symmetry-derived probabilities; Coverage uses enumerated outcomes without reliably specified probabilities.
- Alignment regimes: PAW-Calibration tests probability-mass alignment, whereas PAW-Coverage tests recovery of valid outcomes when their relative probabilities are unavailable.Reporting them separately distinguishes biased probability allocation from missing outcomes.
- Scenario curation: Scenarios require visible physical stochasticity, one atomic judgeable intervention, and a finite set of visually identifiable terminal outcomes.These requirements make repeated-rollout statistics interpretable.
- PAWEval: PAWEval uses an outcome rubric to map each readable, in-schema rollout to a terminal label or an outcome-readout failure.A rubric-based judge processes each rollout before outcome labels are normalized into an empirical distribution.
- Metrics: Calibration uses total variation distance against the reference distribution, while Coverage measures the fraction of valid outcomes observed.Both metrics use readable, in-schema outcomes, with the readout gate determining whether a scene is scoreable.
4 EVALUATION ON PAWBENCH
PAWBench evaluates eleven video generators on probability alignment, valid-future coverage, and response to causal versus non-causal interventions. Across these tests, models show complementary distributional failures, incomplete causal responsiveness, and calibration gaps that cannot be explained by rollout count or outcome-label disagreement.
- Evaluation setup: PAWBench evaluates eleven current video generators using conditional TVD for probability alignment and valid-support recovery for outcome coverage.Table 1 defines calibration through conditional TVD and coverage through valid-support recovery, with averages computed over passing scenes and SPR measuring scene-level reliability.
- Benchmark results: No model consistently combines accurate probabilities, broad valid-future coverage, and reliable performance across scenes.Cosmos 3 Super I2V has the lowest Calibration TVD but 80.0% Calibration scene-pass rate, while LTX-2.3 has the highest Coverage average over 72.0% passing scenes.
- Benchmark results: Broad recovery of valid futures remains distinct from correct probability allocation, so PAW-Coverage and PAW-Calibration expose complementary failures.Models with broad support are not necessarily those matching reference probabilities, and either support recovery or calibration can fail independently.
- Benchmark results: Models underreact or respond incorrectly to physically causal interventions while shifting probability mass in response to non-causal cues.The pencil tilt changes the physical transition, whereas Galton-board text leaves it unchanged; nevertheless, distractor text can redirect model probability mass.
- Robustness of the evaluation protocol: Increasing rollout budgets improves coverage but leaves calibration largely unchanged, and finite sampling does not explain the observed calibration gap.With K = 50, observed TVD averages 31.2 across generators, while 99% of matched reference-distribution simulations remain below 9.22.
- Robustness of the evaluation protocol: PAWEval agrees with decisive human terminal-outcome labels on 722 of 888 videos, or 81.3%, so label disagreement alone cannot explain the benchmark gap.The comparison uses seven independent human judgments per video and the same scene-specific outcome space.
5 TOWARD PROBABILISTICALLY ALIGNED WORLD MODELING
This section examines whether language, noise sampling, and training can reshape video generators’ distributions over possible futures. Language and noise alter requests or exploration, while training changes outcome frequencies but does not reliably produce scene-conditioned distributions.
- 5.1 PROMPT ENGINEERING: Language-based control has two failure points: controllers select misaligned future distributions, and generators often fail to realize supplied targets.Oracle PE improves Calibration TVD and Coverage for every generator, yet generators realize only 37.6–58.1% of requested outcomes.
- 5.1 PROMPT ENGINEERING: Direct VLM sampling is already misaligned before video generation, while passing its selections to generators improves SPR but inconsistently affects calibration and coverage.GLM-5V Turbo records Calibration TVD 34.8, and Gemini 3.5 Flash records Coverage 46.6%.
- 5.2 INITIAL NOISE SAMPLING: C2C couples K = 50 initial-noise samples while preserving each sample’s standard Gaussian marginal, holding the prompt, generator, and rollout budget fixed.The design isolates whether repulsive noise coupling changes finite-budget exploration.
- 5.2 INITIAL NOISE SAMPLING: Across passing scenes, C2C lowers mean Calibration TVD and raises mean Coverage for all three generators, but does not consistently raise SPR.The gains broaden exploration of the model’s existing possibilities rather than changing its learned distribution.
- 5.3 UPDATING MODEL PARAMETERS: Increasing left-falling examples in training raises generated left-fall frequency in both pencil scenes, with a nonlinear rather than one-to-one relationship.None of the five adapted models matches both the upright 50/50 and left-leaning 100/0 references; global frequency shifts help one scene while worsening the other.
- 5.3 UPDATING MODEL PARAMETERS: Training-data composition provides only coarse distributional control because the same global adjustment moves different physical scenes in the same direction.Probabilistic alignment requires distributions to change with each scene’s initial physical state under a fixed action.
6 RELATED WORK
Related work covers world models, video-generation and physics-focused benchmarks, and calibration and distributional evaluation. These lines of work motivate evaluating action-conditioned visual prediction and probabilistic behavior.
- Video Generators as World Models: World models support planning by predicting action consequences, spanning classical model-based reinforcement learning, latent dynamics, and decision-centric models.Recent video systems pursue this role through controllable visual rollouts and interactive world models.
- Video-generation and world-model benchmarks: Video-generation benchmarks assess visual quality, temporal consistency, and text-video alignment, while physics-focused benchmarks probe physical plausibility and action following.The cited benchmark families broaden evaluation beyond visual appearance alone.
- Calibration and coverage under fixed actions: Calibration and distributional evaluation distinguish accuracy, fidelity, probability assignment, and support recovery in generative systems.Recent work also shows that confidence estimation is feasible and important for generative-video uncertainty.
7 LIMITATIONS
PAWBench makes distributional comparison tractable through terminal outcomes, finite-rollout estimates, and controlled visual scenarios. These choices limit coverage of trajectories, evaluation efficiency, and real-world interaction.
- Scope and measurement: Terminal-outcome evaluation enables tractable distributional comparison but does not fully capture trajectory-level dynamics or intermediate physical processes.The limitation concerns what PAWBench measures, not whether terminal outcomes are valid scenario targets.
- Evaluation budget: Finite rollout budgets make estimates less reliable, while larger budgets increase evaluation cost and do not correct biased model distributions.More samples reveal the induced distribution more reliably but cannot fix distributional bias by themselves.
- Scope and future extensions: Controlled, visually parseable scenarios isolate stochastic future modeling but leave longer-horizon, interactive, and embodied environments for future study.Future work also targets richer state trajectories, efficient rollout estimators, interactive scaling, and calibrated training objectives.
8 CONCLUSION
The paper concludes that plausible or controllable rollouts do not establish probabilistically aligned world modeling. PAWBench evaluates this gap through repeated rollouts and shows that current interventions do not reliably recover scene-conditioned future distributions.
- Conclusion: Across eleven video generators, no model consistently matches reference probabilities while recovering valid futures, and the gap is not explained by finite sampling or human disagreement on clear outcomes.Language requests individual futures, coupled noise broadens finite-budget exploration, and fine-tuning shifts frequencies, but none reliably recovers scene-conditioned distributions.
- Benchmark setup: PAWBench fixes each scenario’s initial observation, atomic action, valid terminal outcomes, readout criteria, and, when defensible, a reference distribution q before model evaluation.This setup makes repeated-rollout distributional tests comparable across scenarios.
- Benchmark setup: PAW-Calibration derives reference distributions from physical or structural evidence, while PAW-Coverage handles scenarios whose valid outcomes are enumerable but whose relative probabilities are not reliably specified.Reference distributions are not inferred from model outputs or assumed uniform by default.
A.1.2 PAW-COVERAGE TASKS
PAW-Coverage tests whether repeated rollouts recover the set of valid terminal outcomes when their relative probabilities are unspecified. The benchmark uses curated, readable scenarios, outcome-level readout, and separate trustworthiness diagnostics before computing conditional coverage.
- PAW-Coverage covers scenarios with enumerable valid outcomes but no justified relative probabilities, testing support recovery without assuming equal frequencies.
- PAWBench contains 50 scenarios evenly divided between PAW-Calibration and PAW-Coverage across eight stochastic mechanism groups.
- Each scenario fixes a source image, one action prompt, distinguishable terminal outcomes, and outcome-readout criteria so repeated variation reflects the model’s future distribution.
- PAWEval assigns readable in-schema rollouts terminal labels and sends unreadable or out-of-schema endpoints to the shared readout-failure label ⊥.
- Trustworthiness audits separately record action execution, physical process, and object continuity without changing terminal labels or PAWBench scores.
- A scene passes the outcome-readout gate when at least 20 of 50 rollouts are readable, and conditional coverage uses only readable outcomes.
- Higher coverage means recovering more valid terminal outcomes, with primary comparisons fixed at K = 50 and sensitivity examined at K = 100.
B.4 FINITE-SAMPLE EFFECTS ON CALIBRATION
The matched finite-sample analysis shows that calibration errors substantially exceed the TVD expected from limited readable samples. Paired controls then test whether model distributions respond appropriately to causal changes and remain stable under non-causal changes.
- Matched baselines preserve each model’s passing scenes and readable sample counts while replacing generated outcomes with samples from the reference distribution.
- 31.2 observed TVD exceeds 8.33 TVD from matched reference samples, whose simulated averages remain below 9.22 at the 99th percentile.Each generator’s observed TVD also exceeds its own matched baseline’s 99th percentile.
- Table 8 shows the finite-sample conclusion remains under pooled aggregation and a stricter nreadout ≥30 requirement.The checks do not alter benchmark scores or model rankings.
- The paired-control diagnostic compares base and causal or non-causal variants across eleven video generators and five vision-language models using K = 50 samples.
- Causal variants change the physical transition and reference distribution, whereas non-causal variants preserve both and should leave the outcome distribution stable.
- The non-causal controls include irrelevant appearance, distractor, sign, dispenser, and coin changes while preserving the underlying transition or chance process.
C HUMAN STUDY AND PAWEVAL ALIGNMENT
The human study evaluates whether PAWEval agrees with human terminal-outcome labels on a bounded comparable subset, while language and sampling probes examine ways to alter predictive distributions. Language can steer requested outcomes, but repeated predictions and video responses remain imperfectly aligned.
- Human study: The agreement result characterizes the comparable subset and supports a bounded conclusion about clear terminal-outcome labels, not PAWBench calibration, coverage, or model rankings.
- Human study: PAWEval and humans provide comparable labels on 888 of 1,500 videos after parallel eligibility filters.Humans provide decisive labels on 1,128 videos, while PAWEval provides in-schema labels on 1,024.
- Human study: 81.3% exact agreement means PAWEval matches the decisive human label on 722 of 888 comparable videos.Agreement rises from 58.7% with three or four matching votes to 91.7% with six or seven.
- Language probes: The language probes use VLM predictions without outcome labels or probabilities and target-conditioned video prompts that explicitly request a named outcome.
- Language probes: Oracle PE lowers mean Calibration TVD and raises mean Coverage for all four tested generators, whereas PE raises Calibration TVD for all four and improves Coverage for only two.
- Language probes: Only 37.6–58.1% of rollouts reach the requested outcome, showing that language changes individual rollouts without reliably aligning their aggregate distribution.
- Sampling probe: Coupled noise preserves each rollout’s standard-Gaussian marginal and changes gallery dependence, probing finite-budget exploration rather than the learned one-rollout conditional law.
D.3 TRAINING-DISTRIBUTION INTERVENTION
The training intervention varies only the left/right composition of adaptation data to test whether relative outcome frequency shifts the learned distribution. Interior mixtures increase left-fall frequency in both pencil scenes, but do not establish exact ratio recovery or scene-conditioned probability learning.
- Five 2,000-example datasets use left/right ratios of 0/100, 20/80, 50/50, 80/20, and 100/0 while repeating source videos uniformly within direction.
- Each mixture adapts the same Wan2.2 I2V-A14B base model with rank-32 LoRA, 832×480 clips, 49 frames, learning rate 10^-4, and 2,000 optimizer updates per expert.
- Base and five LoRA conditions use matched K = 50 seeds across an upright pencil scene with a 50/50 reference and a left-leaning scene with a 100/0 reference.
- Across the three interior mixtures, conditional left-fall frequency increases in both scenes.
- The experiment shows training composition shifts mass among supported futures but establishes neither exact ratio recovery nor scene-conditioned probability learning.Both scenes respond similarly despite their different reference distributions.
- Qualitative case cards illustrate successful trials and failure modes but do not enter any reported aggregate metric.