Source-linked AI summary

On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models

Yang Yu

arXiv:2608.22197v1cs.LG

TL;DR

The paper asks whether future prediction gives world-action models stronger control than direct behavior cloning under the same observational demonstrations. It compares their policy classes and population targets with action-conditioned world-model optimization, finding equivalence for unrestricted imitation but a distinct information and decision role for specified-action prediction.

  • Problem

    It is unclear whether future prediction strengthens control capability beyond direct behavior cloning when both are trained from the same observational demonstrations.

  • Method

    The paper compares direct behavior cloning, imitation-trained world-action policies, and action-conditioned world-model policies across expressivity, population targets, identification, and deployment.

  • Results

    Direct and world-action imitation have the same unrestricted external policy class and recover the observational behavior policy under realizability, exact optimization, and distribution-preserving deployment, while action-conditioned models support specified-action consequence comparison.

  • Takeaways & Limitations

    Future prediction changes imitation’s learning factorization, whereas policy optimization requires predicting consequences of specified actions and may require additional information when effects are not observationally identified.

  • Takeaways & Limitations

    The exact results assume realizable model classes, exact population optimization, common deployment information, and distribution-preserving deployment; finite-sample learning is not bounded.

Abstract

from arXiv · show

World-action models predict a future outcome and then infer an associated action. Although this factorization can improve representation learning and data efficiency, it is unclear whether it provides stronger control capability than direct behavior cloning when both are trained from the same observational demonstrations. We compare a direct behavior-cloning policy, an imitation-trained world-action policy, and a policy optimized with an action-conditioned world model. At the controller-class level, every world-action policy can be flattened into a direct stochastic policy with the same closed-loop trajectory distribution. At the population level, under realizability, exact optimization, common deployment information, and distribution-preserving deployment, direct behavior cloning and world-action imitation both recover the observational behavior policy. Thus, future prediction changes the learning factorization but not the unrestricted external policy class or ideal imitation target. Action-conditioned world-model learning differs by predicting outcomes under specified actions and comparing them through a control objective. We characterize the irreducible action-specific prediction error of future models that do not condition on the candidate action, identify conditions under which a world-action joint can recover an interventional forward model, and show that observational demonstrations do not identify action effects in general. Finally, we construct an environment family in which every observational learner has positive worst-case regret, whereas one informative intervention permits zero regret. The key distinction is therefore between predicting futures associated with observed behavior and predicting consequences of specified actions for policy optimization.

1 Introduction

The paper asks whether future prediction gives world-action models stronger control than direct behavior cloning when both learn from the same observational demonstrations. It separates architecture, population targets, identification, and deployment, concluding that imitation factorization differs from action-conditioned world-model optimization.

  • 1 Introduction: World-action models predict future observations or states before inferring actions, whereas direct policies map histories directly to actions.Action-conditioned world models instead predict consequences under specified candidate actions for comparison and optimization.
  • 1 Introduction: The paper studies external expressivity, population targets, information and identification, while treating finite-data and optimization effects as outside its main exact results.Its exact results assume realizability and exact population optimization; the approximate result is a sensitivity statement, not a finite-sample bound.
  • 1 Introduction: Every unrestricted world-action controller can be flattened into a direct stochastic policy with the same closed-loop trajectory distribution, and every direct policy has a degenerate world-action form.Thus, the controller classes have the same external control capability when both range over unrestricted stochastic kernels.
  • 1 Introduction: Under realizability, exact optimization, common observational data, and distribution-preserving deployment, direct behavior cloning and world-action imitation recover the observational behavior policy.Behavior cloning targets pµ(A | H), while standard future-then-inverse deployment marginalizes the observational joint rather than comparing specified actions.
  • 1 Introduction: Action-conditioned prediction separates candidate-action consequences from observational future prediction, whose irreducible error is positive when alternative actions induce different outcome distributions.Under consistency, exchangeability, and positivity, an exact world-action joint can identify P(Y | H, do(A)); observational demonstrations alone need not identify unsupported effects.
  • 1 Introduction: In a two-action environment family, every observational population learner has positive worst-case regret, whereas one informative action intervention permits zero regret.The paper therefore distinguishes predicting futures associated with observed behavior from predicting consequences of specified actions for policy optimization.

2 Related Work

Related work distinguishes observational behavior reproduction from action-conditioned prediction and model-based control. The paper positions its contribution as isolating how representation, causal identification, deployment, and finite-sample issues differ across these uses of future prediction.

  • 2 Related Work: Behavior cloning estimates a demonstrator’s conditional action distribution from observational trajectories, while prior work studies finite-sample error, optimization, and distribution shift.The paper’s question is whether an observational future variable changes the population action target or external policy class before those errors arise.
  • 2 Related Work: World-action architectures use future prediction as an intermediate representation for action inference and may exploit video pretraining and temporal structure under behavior-reproduction objectives.This use differs from action-conditioned futures used for policy optimization.
  • 2 Related Work: World models support planning, policy optimization, value expansion, and imagined experience, but predictive accuracy need not coincide with decision-relevant accuracy.Policy-conditioned and long-horizon models examine how predictive structure interacts with control.
  • 2 Related Work: An observational conditional is not automatically an interventional conditional; equality requires consistency, conditional exchangeability, and positivity.This distinction underlies work on invariant action effects, causal world-model factorizations, counterfactual environments, and model-predictive control.
  • 2 Related Work: Offline model-based reinforcement learning faces a support problem when optimized policies select actions whose consequences are poorly represented in the dataset.This connects dataset coverage to the reliability of action-conditioned model-based optimization.

3 Framework and Scope

The framework formalizes sequential environments, histories, trajectories, policies, population objectives, and future labels, then separates controller expressivity from learner information. The comparison assumes common deployment information and observational data, realizability, exact optimization, and distribution-preserving deployment.

  • Sequential environment and trajectories: Controllers map histories to actions in a finite-horizon partially observed process whose trajectories depend on observations, actions, and controlled transitions.The available history may include instructions, goals, task identifiers, observations, and past actions.
  • Demonstrations and population objectives: Future labels are recorded trajectory variables such as future observations, states, latent representations, or video chunks, available during training but not action selection.Population objectives use the underlying demonstration distribution rather than a finite empirical sample.
  • Control-capability and learner classes: A control-capability class records externally distinguishable closed-loop behaviors, not sample efficiency, optimization difficulty, parameter count, or task-specific utility.Controller classes describe representable policies, whereas learner classes describe information available for selecting among them.
  • Scope of the comparison: Observational learners receive the true observational distribution and utility, whereas interventional learners additionally receive outcomes generated under specified action interventions.The environment family, utility, and candidate policy class are shared; only the information available for policy selection differs.
  • Scope of the comparison: The main equivalence comparison removes finite-sample, approximation, and optimization error through realizability, exact population optimization, and common deployment conditions.Behavior sufficiency is additionally required when the observational action conditional is interpreted as the demonstrator’s policy generating the full trajectory distribution.
  • Learning paradigms: Direct behavior cloning models pµ(A | H), while world-action learning factors a joint future-action model into a future predictor and inverse-action predictor.The world-action future predictor does not condition on the current candidate action, and distribution-preserving deployment samples or marginalizes its learned future distribution.

4 Observational Equivalence of Direct and World-Action Policies

The paper compares direct and world-action policies through complete closed-loop trajectory distributions and population optima. Under ideal observational conditions, world-action flattening gives no unrestricted external capability advantage, and both imitation objectives recover the same action target.

  • Control equivalence: Control equivalence is evaluated through complete trajectory distributions, so zero control distance prevents any bounded trajectory-level return, success, safety, or verification statistic from distinguishing policies.Equality under only one benchmark reward would be insufficient for general control equivalence.
  • Architecture-level equivalence: Every distributional world-action controller induces a direct stochastic policy with exactly the same closed-loop trajectory distribution, and conversely direct policies have degenerate world-action representations.Thus unrestricted direct and world-action controller classes are externally control-equivalent.
  • Population action targets: Under realizability and exact population optimization on the same observational process, direct behavior cloning and world-action maximum likelihood recover the same observational action conditional.Sampling a behavior-distribution future and then its posterior behavior action recovers pµ(a | h).
  • Deployment conditions: The equivalence requires distribution-preserving deployment; MAP future selection, verifier reranking, or other joint modifications can change the resulting action marginal.The theorem is therefore about the learned stochastic deployment rule, not arbitrary decoding procedures.
  • Practical factorization effects: Future-conditioned action prediction can have lower training uncertainty by an amount Iµ(A; Y | H), but deployment must predict the unavailable future, trading action uncertainty against future-prediction error.These practical differences can affect learning or representation quality without changing the ideal imitation target.

5 Interventional Identification and Decision Separation

This section distinguishes predicting outcomes associated with observed behavior from predicting consequences of specified actions. It shows that action effects are generally not identified observationally, while identified world-action joints can support planning and interventions can provide strict information gains.

  • Action-specific prediction: An action-unconditioned future predictor has irreducible error whenever candidate actions induce different outcome distributions, whereas an exact action-conditioned model has zero risk.The lower bound combines conditional action–outcome dependence with divergence between the behavior-mixture outcome distribution and the predictor.
  • Identification: A world-action joint can recover the observational forward conditional under exact fitting and positive action support, but causal interpretation additionally requires identification assumptions.The relevant requirements include recovering the observational joint, positivity, and causal identification.
  • Planning: With consistency, conditional exchangeability, and positivity, planning over the recovered action-conditioned outcomes maximizes true interventional expected utility.The same joint can instead be deployed for behavior reproduction, so deployment and causal identification determine its decision role.
  • Observational non-identification: Interventional action effects are not identified by observational pµ(H, A, Y ) in general, both without positivity and when history omits a variable affecting action selection and outcomes.Action coverage and causal sufficiency address distinct identification failures.
  • Decision separation: The model-based optimization corollary assumes relevant action effects are already known or identified, whereas interventions can create a strict information advantage.Exact optimization then achieves the best true value within the candidate policy class.
  • Value of interventions: There exists an environment family where every observational learner has worst-case regret at least 1/4, while one informative intervention identifies the environment and enables zero regret.The observational distribution is identical across environments, so the separation comes from information about action effects rather than unrestricted policy expressivity.

6 Conclusion

The paper separates observational imitation from interventional control: direct behavior cloning and imitation-trained world-action models can be equivalent externally and at the ideal population target, while action-conditioned control requires explicit candidate-action evaluation and may need interventions.

  • Direct behavior cloning and imitation-trained world-action policies have the same external control-capability class under unrestricted stochastic kernels.Every distributional world-action controller flattens to a direct stochastic action kernel with the same closed-loop trajectory distribution.
  • Under realizability, exact optimization, matched deployment information, and distribution-preserving deployment, both imitation approaches recover the observational behavior policy.
  • World-action modeling may improve representations, pretraining, parameter sharing, or finite-sample performance without establishing interventional reasoning.
  • Action-conditioned world-model control specifies candidate actions, evaluates their consequences, and compares their utilities rather than decoding actions from observed-behavior futures.
  • Observational demonstrations do not generally identify action effects; interventions, exploratory coverage, or valid causal assumptions are required when decisions depend on unidentified effects.
  • The conclusions are bounded by class-level, population-level assumptions, matched deployment information, behavior sufficiency, distribution-preserving deployment, and valid identification conditions.The paper does not provide sample-complexity or optimization guarantees for large neural models.

A.2 Proof of Theorem 4.4

The proof constructs a degenerate world-action representation for every direct stochastic policy and uses population objective decomposition to establish equivalence under realizability.

  • Every direct stochastic policy has a degenerate world-action representation using a fixed future outcome and the policy as the inverse component.
  • Under realizability and exact population optimization, the KL term can be minimized to zero, recovering the observational conditional action distribution almost everywhere.
  • The world-action population objective decomposes into future entropy, conditional action entropy, and an inverse-model KL divergence.
  • The proof then uses behavior sufficiency and induction over time to extend equality of action kernels to closed-loop trajectory distributions.

A.4 Proof of Theorem 4.8

The proof establishes the interventional identification result by relating observational conditionals to a joint distribution, applying marginalization and total-variation bounds, and using sequential coupling.

  • The proof defines the joint distribution as the observational future conditional multiplied by the inverse action conditional.
  • Marginalization contracts total variation, and the triangle inequality decomposes the resulting discrepancy into two component bounds.
  • Sequential maximal coupling bounds trajectory disagreement by coupling actions at each common history and using the shared environment kernel after equal actions.
  • The remaining conclusion follows from Theorem 4.3.

A.6 Derivation of Equation (57)

The derivation uses conditional exchangeability to equate potential-outcome and observed-outcome conditionals, yielding the observational conditional for the specified action.

  • Conditional exchangeability equates P(Y(a)=y | H=h) with P(Y(a)=y | H=h, A=a).
  • Consistency then equates P(Y(a)=y | H=h, A=a) with P(Y=y | H=h, A=a).
  • Together, these steps yield T_a(y | h) = pµ(y | h, a).

A.7 Proof of Theorem 5.4

The proof constructs observationally indistinguishable environments showing that observational learners can fail to identify action effects. An informative intervention can resolve the ambiguity and eliminate regret.

  • The proof gives separate constructions for support failure and hidden confounding.
  • Hidden confounding: With positive observational action support, hidden confounding can still make environments observationally identical while their interventional outcomes differ.
  • Every observational learner receives the same observational distribution and utility in both environments and therefore produces the same stochastic action distribution.
  • Support failure: In the support-failure construction, exact behavior-cloning and world-action population optima choose A = 0 with probability one, yielding worst-case regret 1/2.
  • Support failure: An intervention selecting A = 1 in M+ and A = 0 in M− identifies the environment and achieves zero regret.

B.1 Point decoding

Point decoding can preserve some marginal behavior properties while failing to preserve the full action distribution or demonstration trajectory distribution. Restricted direct policy classes and out-of-support deployment create additional boundaries for equivalence and performance.

  • Point decoding: A distributional world-action model can reproduce the behavior policy, whereas a MAP-future decoder may produce a different deterministic controller.
  • Point decoding: Sampling one future and applying its conditional mean preserves equality in expectation but not necessarily equality of the full action distribution.
  • Matching pµ(A | H) need not reproduce the demonstration trajectory distribution when behavior sufficiency fails because dependence on an omitted variable can be lost.
  • Including the omitted variable Z in the history restores the information used by the demonstrator.
  • A restricted direct policy class reproduces the world-action class only when it is closed under the required marginalization.
  • Finite architectures may gain representational or computational advantages from world-action factorization, while optimized policies can still depend on extrapolation outside behavior support.
Loading 2608.22197v1…