Source-linked AI summary
Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information
Bingqi Huang, Bingchuan Wei, Yingkai Cai, Zhaokui Wang
TL;DR
Camera viewpoint changes are a major robustness challenge for WAMs because their denoising targets mix view-invariant actions and state variables with view-dependent future scenes. The paper introduces SCVC, which selectively enforces cross-view consistency on invariant outputs and evaluates held-out viewpoints with controlled paired-view splits, improving extrapolation while preserving in-distribution competence.
Problem
WAM consistency training must handle denoising targets whose future-scene coordinates change with viewpoint while actions, future proprioception, and value remain invariant.
Method
SCVC renders same-state cross-view pairs, uses shared noise draws, and applies consistency only to the invariant output block while retaining per-view supervision for the future scene.
Results
+12.2 points in closed-loop success on held-out orbital extrapolation viewpoints over the matched control, with replication across two further camera axes and preserved in-distribution performance.
Takeaways & Limitations
Selective output consistency improves WAM robustness beyond the training camera envelope without requiring camera information, view synthesis, or deployment-interface changes.
Takeaways & Limitations
Closed-loop WAM evidence is simulated, real-robot evidence remains open, and consistency is supervised only on demonstration states rather than failure or recovery states.
Abstract
from arXiv · showhide
World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, on which output coordinates should a consistency loss be imposed? The WAM denoising target mixes view-covariant coordinates, namely the predicted future scene, with view-invariant coordinates, namely the action chunk, future proprioception, and value. We show that consistency applied to the covariant block is provably harmful, shrinking legitimate view-specific content to a fraction $1/(1+4λ)$ of its true value, and we verify this shrinkage law in controlled experiments. Selective cross-view consistency (SCVC) therefore constrains only the invariant block, requires no camera labels, extrinsics, depth, or view synthesis at training or test time, and leaves the deployment interface unchanged. We introduce a carve-and-hold-out evaluation protocol on the LIBERO-Plus camera track that separates a distribution-matched ceiling from genuine interpolation and extrapolation to held-out viewpoints, with a matched pair-trained control isolating the effect of the consistency term from pair exposure. On held-out orbital viewpoints beyond the training envelope, SCVC improves closed-loop success over the matched control by 12.2 points (95% CI [7.4, 17.0]; +15.5, CI [11.7, 19.4], under an independent second seed) -- an effect two further camera axes replicate -- while interpolation within the envelope shows no gain in either seed (-1.2 and -4.3 points) and in-distribution competence is preserved (-0.6, -0.2). We also report a cross-backbone audit showing that published camera-robustness numbers are confounded by wrist-camera pose stability.
I. INTRODUCTION
The paper identifies viewpoint change as a difficult WAM robustness axis and argues that consistency must respect the denoising target’s mixed transformation laws. SCVC constrains only invariant outputs and evaluates held-out viewpoints with a protocol separating pair exposure from objective effects.
- Camera perturbations cause some of the largest robustness drops across WAM and VLA families, motivating focused evaluation of viewpoint generalization.
- Wrist-camera stability confounds published camera-robustness measurements, so the study trains and evaluates scene-only policies.A wrist-equipped policy can retain substantial performance when the scene camera is blacked out, unlike when its wrist camera is removed.
- Consistency on the covariant future-scene block shrinks legitimate view-specific content to exactly 1/(1+4λ) of its true value, motivating selective application to invariant outputs.The invariant block contains actions, future proprioception, and value; the future scene remains supervised per view.
- The carve-and-hold-out protocol covers each camera axis except deliberately held-out bands and partitions tasks into in-distribution, interpolation, and extrapolation buckets.A matched control trained on identical carved pairs with zero consistency weight isolates the objective from pair exposure.
- +12.2 points on held-out extrapolation-regime orbital viewpoints, with +15.5 under a second seed, while interpolation shows no gain and in-distribution success is preserved.The pattern replicates on two further camera axes.
II. RELATED WORK
Related work spans world action models, viewpoint-robust visuomotor policies, and consistency regularization. The paper distinguishes SCVC by selectively constraining structured outputs rather than representations, pixels, explicit camera inputs, or fully invariant action outputs.
- World action models post-train video-generation backbones to jointly predict future observations and robot actions, with robustness pursued through architecture, representation, diagnosis, and benchmarks.
- ReViWo learns view-invariant representations while keeping its decoder view-dependent, whereas SCVC selectively constrains model outputs by their transformation laws.
- Viewpoint-robust policy methods include test-time view synthesis, camera conditioning, augmentation, reinforcement smoothness, and analytic equivariance, each differing from SCVC’s data-based output selectivity.
- Consistency-regularization theory commonly assumes the entire label is invariant, but WAM targets violate that assumption on the future-scene block.The paper adds noise-draw matching for denoising pairs, which deterministic predictors do not require.
- An action-equivalent pair consists of same-state observations under nominal and perturbed cameras sharing the action chunk, future proprioception, and value label.The WAM target separates these invariant coordinates from the camera-dependent future scene.
IV. SELECTIVE CROSS-VIEW CONSISTENCY
SCVC applies cross-view consistency only to invariant denoising coordinates while retaining per-view supervision for the covariant future scene. Shared noise draws and a ramped consistency weight implement this selective objective without changing the deployment interface.
- The denoiser receives paired views under one shared noise level and realization, while the branches differ in scene conditioning and covariant target content.
- The consistency loss is projected onto invariant frames, comparing the two branch predictions only for actions, future proprioception, and value.
- The covariant future-scene frame receives only its per-view supervised loss, preserving freedom to predict each camera’s future.
- The λ=0 configuration is a camera-augmented control, and no camera parameters, extrinsics, depth, or synthesized views enter training or testing.
B. Why selectivity is necessary
The consistency objective separates cross-view means from view-residuals: it is unbiased on invariant coordinates but shrinks legitimate view-specific targets when applied to covariant coordinates.
- Why selectivity is necessary: The mean-residual decomposition leaves the supervised anchor controlling the cross-view mean while scaling only the view-residual penalty.On invariant coordinates, this creates unbiased pressure toward agreement.
- Why selectivity is necessary: For covariant targets, consistency shrinks the predicted cross-view difference to 1/(1+4λ) of its supervised value.The pointwise minimizer preserves the mean but attenuates the view-specific residual.
- Why selectivity is necessary: Consistency on the covariant scene block pulls future predictions toward a ghosted average, with per-view bias approaching half the true view difference as λ →∞.This makes the ordinary VLA consistency recipe unsuitable for the mixed-coordinate WAM target.
- Why selectivity is necessary: Independent noise draws add a variance penalty even when the denoiser is exactly view-invariant.At D0 ≡ Dp ≡ D, the added term is 2 tr Varn(D) > 0 when predictions depend on noisy input.
- Why selectivity is necessary: A sound cross-view comparator matches state, coordinates, and noise, differing only in camera.The controlled figure illustrates the resulting difference: selective consistency preserves view-specific futures while wrong-coordinate consistency ghosts them.
C. Controlled verification of the shrinkage law
A two-scale controlled study tests whether stochastic training reproduces the analytically predicted shrinkage law for selective consistency.
- Controlled verification of the shrinkage law: An exact simulation with independent per-pair minimizers reproduces the predicted shrinkage ratio to machine precision.A shared-trunk denoiser is then trained to convergence on synthetic invariant and covariant target blocks to test the law under shared-network optimization.
A. Why the protocol is needed
The evaluation protocol is designed to distinguish genuine held-out-view generalization from performance explained by camera coverage or matched training distributions.
- Why the protocol is needed: Nominally trained baselines cannot provide a zero-shot comparison because consistency methods must train on perturbed viewpoints.Near-nominal training can also fail on severe views simply because those viewpoints lack coverage.
- Why the protocol is needed: The carve-and-hold-out protocol trains across each axis’s severity range while reserving bands for in-distribution, interpolation, and extrapolation evaluation.LIBERO-Plus spans orbital azimuth, elevation, dolly scale, and endpoint reorientation.
- Why the protocol is needed: The integrated training manifest contains 583,648 same-state pairs and excludes held-out camera values from every training row.The severe-azimuth training range retains 66,437 pairs, addressing the stated coverage failure mode.
- Why the protocol is needed: The paired control uses the identical carved manifest with λCV = 0, isolating the consistency objective from exposure to cross-view pairs.The analysis plan pre-specifies azimuth interpolation and extrapolation as primary endpoints with paired-task bootstrap confidence intervals.
VI. EXPERIMENTS
The experiments audit what camera robustness measures before evaluating scene-only policies, showing that wrist-camera stability can dominate published camera-track scores.
- What camera-track numbers measure: the wrist audit: Masking the scene camera reduces released Cosmos success from 80.5% to 66.6%, whereas masking the wrist reduces it to 0.2%.The audit uses 1,200 rollouts per condition with Wilson 95% confidence intervals.
- What camera-track numbers measure: the wrist audit: Wrist-equipped policies lose 19.0 points under camera perturbation versus 67.2 points for the scene-only reference.Their in-distribution success differs by about five points, while camera-track performance differs by more than fifty.
- What camera-track numbers measure: the wrist audit: On OpenVLA-OFT, removing the wrist drops camera score from 56.4 to 10.4.Masking the scene view costs 14 aggregate points in the wrist-equipped policy, with effects varying by task category.
- What camera-track numbers measure: the wrist audit: All mechanism experiments therefore use scene-only policies to measure scene-view invariance without the wrist-stability confound.The scene-only models continue post-training from the released Cosmos LIBERO checkpoint under a common training configuration.
C. Motivating observation: the dream survives where the action fails
The scene-only reference reveals that viewpoint changes can destroy closed-loop control even when predicted-scene fidelity remains intact, motivating an action-side repair. Held-out evaluation then shows SCVC’s benefit is concentrated beyond the training viewpoint envelope, not within it.
- 93.2% in-distribution success falls to 26.0% across the full camera track, with the largest axis-specific drop under in-place reorientation at 61.3%.The corresponding drops are 7.5% for dolly and 21.4% for orbital viewpoint change.
- 53% to 3%: orbital action success collapses across severity levels while relative excess-FID remains between −0.07 and +0.16.Dolly instead shows fidelity and control degrading together, separating visual prediction fidelity from semantic control failure on the orbital axis.
- +12.20 points: SCVC improves extrapolation to held-out orbital viewpoints beyond the training envelope, with 95% CI [+7.40, +17.00].Interpolation is unchanged at 85.1 vs. 86.3, a paired difference of −1.22 points with CI [−4.29, +1.84].
- +4.24 and +8.76 points: dolly-scale and elevation extrapolation replicate the positive held-out-viewpoint pattern, while interpolation counterparts and C3 are null.The three-axis result is therefore selective: gains occur beyond the training envelope rather than uniformly across camera perturbations.
- −0.60 points: full foursuite in-distribution competence is preserved within the pre-registered tolerance.A second seed likewise reports ID preservation of −0.2.
- +15.50 points: an independent second seed reproduces the extrapolation hierarchy, with CI [+11.70, +19.40].The second seed also reports dolly +4.44, elevation +4.84, and interpolation −4.29.
VII. DISCUSSION
The matched comparison attributes SCVC’s extrapolation gain to the consistency term rather than exposure to perturbed camera pairs. Its benefits are bounded: they are strongest where viewpoint coverage runs out and can reverse on grasp-precision and compound-shift cases.
- The matched control shares the checkpoint, carved manifest, initialization, budget, and noise schedule with SCVC, differing only in λCV.Because both models saw identical perturbed cameras, the extrapolation gain is attributed to the consistency term alone.
- Null within the training envelope and CI-separated gains beyond it match the predicted advantage of per-state cross-view equivalence over marginal coverage.Interpolation begins around 85–93%, whereas extrapolation begins at 55–67%, leaving more room for improvement beyond coverage.
- Shuffled-pair training removes the benefit, while controlled shrinkage-law experiments reproduce the wrong-coordinate failure predicted by the algebra.Together these checks distinguish state-matched consistency from generic regularization and support selective coordinate choice.
- 40.8 points: spatial performance improves while object performance regresses by 11.6 points within the primary extrapolation bucket.Under a compound extreme azimuth-and-elevation shift, the object suite regresses by 40.4 points, showing that the trade can become negative.
- The method addresses viewpoint-disorientation failures but leaves fine grasp geometry—and therefore some manipulation-precision failures—outside its reach.The near-null C3 result is consistent with its small ±10° held-out span and limited remaining invariance to recover.
VIII. LIMITATIONS
The study’s evidence is bounded by simulated closed-loop evaluation, demonstration-state supervision, and limited interpolation and elevation coverage, while real-robot WAM evidence remains open.
- The method requires same-state cross-view pairs, which are cheap in simulation but require synchronized multi-camera rigs or paired real-world datasets.Extending the analysis to approximately matched pairs is future work.
- Real-robot evidence for the WAM instantiation remains open despite related flow-VLA transfer to a real robot.
- Closed-loop evidence is simulated, and robustness claims do not cover failure or recovery states because rollout data lacks recoverable simulator states.
- The elevation axis contributes only one extrapolation point, while interpolation buckets contain 46–49 tasks and show high baseline success.Checkpoint variation can exceed the paired sampling interval in these small, high-performing cells.
- The scene-only protocol isolates scene-view invariance but excludes wrist-camera pose perturbation.