Source-linked AI summary
Platonic Representation Hypothesis on World Models
Wenhow Li, Chengwei MA, Hui Xiong, Ying-Cong Chen, Lei Zhang
TL;DR
Whether world models develop Platonic-like shared representations under action-conditioned prediction remains largely unexplored. This paper varies heterogeneous visual priors in DINO-WM and evaluates geometric alignment and model stitching, finding convergence toward shared latent structure and functional compatibility across several high-quality models.
Problem
Whether predictive consistency drives shared latent structure in action-conditioned world-model transitions remains largely unexplored.
Method
The paper varies heterogeneous visual encoders in DINO-WM and evaluates latent geometry and functional compatibility through anchored comparisons and model stitching.
Results
Heterogeneous predictors progressively converge toward geometrically similar latent structures, while simple mappings preserve planning performance across several high-quality model pairs.
Takeaways & Limitations
Predictive world-model training can encourage heterogeneous visual priors to develop shared, transition-compatible latent structure.
Takeaways & Limitations
The conclusions are limited to frozen-encoder DINO-WM models and five simulated control tasks, and the probes do not prove recovery of objective physical laws in general.
Abstract
from arXiv · showhide
World models have demonstrated significant potential for perceiving and simulating complex environments. Despite their strong performance, the fundamental nature of their learned representations remains poorly understood. In this paper, we investigate the Platonic Representation Hypothesis within this domain by proposing the Predictive Consistency Assumption: we posit that the optimization of a shared state transition objective acts as a selective pressure that encourages heterogeneous models to converge toward a shared latent structure. Through systematic experiments with the DINO World Model (DINO-WM), in which we vary visual encoders to create heterogeneous models, we find that capable world models evolve toward geometrically similar internal structures. Moreover, via model stitching, we show that the internal features of one world model can be mapped to another with limited performance degradation, providing evidence of functional compatibility. Our findings suggest that the pursuit of predictive consistency can promote shared, transition-compatible latent structure across world models.
1 Introduction
This paper investigates whether capable world models trained on the same environment converge toward a unified latent representation. Using heterogeneous visual encoders in DINO-WM, it finds geometric alignment shaped by predictive training and architectural biases.
- Motivation: The Platonic Representation Hypothesis proposes that neural networks’ internal representation spaces increasingly align toward a shared statistical model of reality as they scale.The paper applies this hypothesis to world models, whose representations are learned while predicting future environmental states.
- Hypothesis: The paper hypothesizes that sufficiently capable world models learning the same environment converge toward a unified Platonic representation.This hypothesis follows from world models’ objective of predicting future states from current states under the environment’s governing dynamics.
- Method: The study tests this hypothesis with DINO-WM by varying vision encoders to construct heterogeneous world models and analyzing their predictors’ internal representations.The predictors model visual dynamics without reconstructing the visual world and receive diverse visual feature priors.
- Findings: Geometric metrics show that predictors initialized with heterogeneous ViT priors, including DINOv2, SigLIP, and MAE, progressively converge toward a shared latent structure.The reported alignment is shaped by architectural inductive biases, including a persistent topological gap between transformer-based models and the
- Contribution: The findings indicate that predictive training can encourage heterogeneous visual priors to form a shared latent transition structure in embodied prediction systems.The paper presents this convergence as an empirical setting for studying Platonic-like alignment in world models.
2 Related Work
Related work frames world models as increasingly abstract, latent-dynamics systems and motivates testing whether their shared predictive capabilities reflect structural convergence in latent spaces. This study extends the Platonic Representation Hypothesis to action-conditioned state transitions using DINO-WM, Mutual k-NN, and model stitching.
- World Models: World modeling has progressed from recurrent controllers to robust latent dynamics frameworks and generative video foundation models that learn causal interactions as general-purpose physical simulators.These developments increasingly emphasize modeling dynamics and abstract states rather than reconstructing pixels.
- Study Setting: DINO-WM bridges self-supervised visual priors with dynamics learning and serves as the testbed for analyzing representational convergence in world models.Its use follows the apparent functional convergence across diverse approaches to modeling temporal evolution.
- Platonic Representation Hypothesis: Convergent learning shows that networks with different initializations can develop functionally similar features, motivating the Platonic Representation Hypothesis.PRH posits that models converge toward a shared statistical representation of reality as they scale.
- Representation Analysis: Relative Representation frameworks propose that models may differ in absolute coordinates while preserving internal geometry, motivating Mutual k-NN as an alignment metric.Model stitching provides a complementary test of whether latent spaces are functionally interchangeable.
- Research Gap: Although PRH has been studied in static vision and language modalities, its extension to action-conditioned state transitions remains largely unexplored.This work investigates whether predictive consistency drives structural convergence in world-model latent dynamics.
3 Method: Modeling Action-Conditioned State Transitions toward a Unified Latent Representation
This section formulates world modeling as learning action-conditioned state transitions and proposes the Predictive Consistency Assumption, which predicts convergence toward geometrically similar internal latent structures despite differing visual encoders. It tests this hypothesis with DINO-WM using heterogeneous vision priors and evaluates convergence through geometric analysis and functional model stitching.
- Formalization: A World Model approximates environment dynamics by mapping high-dimensional observations into visual features and predicting future encoded observations from feature-action histories.The predictor is a state transition operator with internal layer-wise latent representations, trained against future features from a fixed vision encoder.
- Formalization: The Predictive Consistency Assumption posits that predictors trained for accurate transitions in the same environment converge toward a shared geometrically similar latent structure despite differing input features.State-transition optimization is proposed to act as selective pressure for representational alignment across predictor representation spaces.
- Experimental Framework: DINO-WM isolates vision-prior effects by varying the fixed encoder while holding the predictor architecture constant.The four encoders are DINOv2, SigLIP, MAE, and ResNet, representing distinct inductive biases and including a standard convolutional baseline.
- Analytical Protocols: The study measures representation convergence through geometric m-kNN analysis of predictor hidden-state neighborhoods on paired validation trajectories.m-kNN is computed per predictor layer and reported as a layer average; higher values indicate stronger preservation of latent topology.
- Analytical Protocols: Functional representation stitching splices layers from trained world models using lightweight MLP adapters between frozen components and evaluates downstream planning success with SSR.High SSR indicates exchangeable transition factors rather than merely similar local neighborhoods.
4 Experiments: Evaluating the Platonic Convergence of World Transitions
Experiments test whether heterogeneous world models develop shared latent geometry under a common transition objective. Across geometric analyses and model stitching, convergence and functional compatibility are strongest among ViT-based models, while ResNet remains separated.
- Experimental Setup: DINO-WM predictors use fixed heterogeneous encoders spanning DINOv2 variants, SigLIP, MAE, and ResNet18 across five benchmark environments.The environments are PointMaze, Wall, PushT, Rope, and Granular.
- Experimental Setup: DINOv2-S provides the strongest and most stable epoch-10 zero-shot planning performance and serves as an operational anchor for geometric comparisons.Candidate models are compared at epochs 1, 3, 5, 7, and 10 using anchored m-kNN; the anchor is not treated as a true Platonic ideal.
- Geometric Convergence: DINOv2 scaling variants show high early geometric alignment, whereas SigLIP and MAE exhibit the largest progressive gains toward the anchor’s latent geometry.The DINO family begins in a high-similarity regime, while SigLIP and MAE move toward shared, transition-compatible structure from more disparate initial representations.
- Geometric Convergence: ResNet18 remains near a low similarity baseline, indicating that the transition objective does not bridge the observed convolutional-transformer architectural gap.The broader ViT-vs-ResNet ordering remains visible under alternative anchors and later checkpoints, although similarity scores are not monotonic with training.
- Alignment Analysis: Frozen encoders already share non-trivial similarity, but predictor transformations often reduce early alignment, and later cross-model alignment cannot generally be explained as encoder recovery.In Rope and Granular, predictors become less encoder-like by epoch 10 even as cross-model predictor alignment improves.
- Functional Compatibility: Stitching strong planners often recovers high relative success with lightweight MLP alignment, supporting exchangeable planning-relevant transition factors across models.PointMaze retains high stitched success for many pairings; in Wall, DINOv2→SigLIP improves SSR from 0.30/0.38 at k = 1 to 0.62 at k = 3, while MAE directionality is environment-dependent.
5 Conclusion
The study concludes that shared predictive objectives can drive heterogeneous world models toward geometrically similar and transition-compatible latent structures. These conclusions are limited to frozen-encoder DINO-WM models and five simulated control tasks, while the probes do not establish a universal Platonic ideal or physical-law recovery.
- Conclusion: Shared state-transition objectives can align heterogeneous vision priors toward geometrically similar local latent topologies.The study frames predictive consistency as a driver of shared latent structure despite disparate feature spaces.
- Conclusion: Transformer-based priors such as SigLIP and MAE show the clearest trajectory toward the designated Platonic anchor during training.The conclusion identifies this convergence pattern as especially pronounced for these transformer-based priors.
- Limitations: The conclusions are limited to frozen-encoder DINO-WM models and five simulated control tasks.The DINOv2-S anchor is an operational performance reference rather than a true Platonic ideal, and m-kNN/stitching probe geometry and compatibility without proving general physical-law recovery.
- Conclusion: Predictive world-model training provides a concrete setting for studying Platonic-like alignment in embodied prediction systems.The proposed interpretation centers on shared latent transition structure emerging across heterogeneous visual priors.
A Experiment Setting · A.1 Predictor Architecture Details · A.2 Planning Details
The appendix specifies the DINO-WM predictor architecture and latent-space planning procedure used in the experiments. The predictor processes spatial patch features with causal temporal attention, while MPC with CEM optimizes action sequences against goal latents.
- A Experiment Setting: The experiment uses a DINO-WM-style predictor that processes spatial patch features with an embedding dimension matching the vision encoder.The predictor is implemented as a customized Vision Transformer.
- A.1 Predictor Architecture Details: The predictor transformer has 6 layers, 16 attention heads, and an MLP hidden dimension of 2048.With DINOv2-S, whose embedding dimension is 384, this configuration has approximately 19 million trainable parameters.
- A.1 Predictor Architecture Details: The K-dimensional action vector is projected by a separate MLP and concatenated to every visual patch embedding at the corresponding time step.This provides action conditioning for the predictor.
- A.1 Predictor Architecture Details: A causal attention mask lets tokens attend across previous time steps while allowing bidirectional attention among patches within the same time step.This preserves temporal causality while enabling global spatial reasoning within the predicted frame.
- A.2 Planning Details: Planning follows Model Predictive Control with the Cross-Entropy Method optimizing action sequences a_t:T−1 to minimize a planning cost C.The implementation follows the MPC scheme described in Zhou et al. [2024].
- A.2 Planning Details: The planning cost is the Mean Squared Error between the final predicted latent state ˆz_T and the goal latent state z_g.Predicted states follow ˆz_t = ψ(ˆz_t−1, a_t−1), while the goal latent is z_g = ϕ(o_g).
- A.2 Planning Details: The planner uses a population size of N = 50 action sequences and 30 optimization iterations per step.The horizon T typically ranges from 5 to 25 environment-dependent steps, after which the first k actions are executed before replanning.
B Metric and Stitching Details · B.1 Mutual kNN alignment · B.2 Representation stitching protocol
The paper measures representation alignment using layer-averaged mutual kNN over predictor hidden states from paired trajectories, then tests transition-level compatibility through lightweight-adapter stitching and downstream planning success. These procedures distinguish local geometric similarity from transferable, exchangeable planning-relevant transition factors.
- B.1 Mutual kNN alignment: m-kNN is computed on predictor hidden states rather than frozen vision-encoder outputs, using paired representations from identical held-out trajectories.For a predictor with L layers, the reported score is averaged across layers.
- B.1 Mutual kNN alignment: For each layer and sample, the method identifies the k nearest neighbors in both representation spaces and computes their mutual-neighborhood agreement.The feature vectors are denoted xA_i and xB_i, with neighborhoods Nk(xA_i) and Nk(xB_i).
- B.1 Mutual kNN alignment: A higher m-kNN score indicates stronger preservation of local topology, meaning the models place the same world states in more similar neighborhoods.The metric evaluates geometric alignment, not transition-information transfer.
- B.2 Representation stitching protocol: Stitching tests compatibility by placing early predictor layers from one model before later layers from another model.The splice is performed between two trained world-model predictor stacks.
- B.2 Representation stitching protocol: Only lightweight MLP adapters are trained at the interface while the remaining model parameters are frozen.The interface uses a two-layer MLP adapter A, and a two-layer projection P maps the stitched output into the target feature space.
- B.2 Representation stitching protocol: The stitched model maps an input visual feature z to a predicted future visual feature z* and is evaluated using downstream planning success (SSR).The protocol also includes a lightweight adapter and projection between feature spaces.
- B.2 Representation stitching protocol: High SSR under lightweight adaptation indicates that the predictors expose exchangeable planning-relevant transition factors.Directionality tests swap the front-end and back-end order while retaining the same adapter/projection design.
C Anchor Sensitivity of the Alignment Pattern
The alignment pattern is not dependent on DINOv2-S as the reference anchor. Although changing anchors affects absolute m-kNN values and some local rankings, the broader grouping remains stable, with ViT-based models more aligned than ResNet.
- Anchor selection: The main analysis uses DINOv2-S at epoch 10 as the reference anchor because of its strong planning performance.The study tests whether this performance-motivated choice creates an artifact in the observed alignment pattern.
- Anchor selection: The anchored m-kNN analysis is repeated with SigLIP@10 and MAE@10 across all five environments.Self-anchor comparisons are removed to avoid trivial values.
- Robustness of alignment: Changing the anchor alters absolute m-kNN scale and local rankings, but preserves the broader alignment structure.DINOv2 variants, SigLIP, and MAE remain substantially more aligned with one another than the ResNet baseline.
D Persistence under Longer Training
Longer-training analyses show that the main cross-model alignment ordering persists beyond epoch 10, supporting stability rather than an early-training artifact. However, alignment does not monotonically increase across models and environments.
- Longer-training stability: Later checkpoints around epochs 20, 30, and 50 preserve the main cross-model ordering observed at epoch 10.For two PushT runs ending at epoch 49, the latest available checkpoint serves as the late-training proxy.
- Longer-training stability: Alignment can slightly increase, decrease, or plateau depending on the model and environment, rejecting a monotonic convergence story.The stability results therefore support persistent relative structure without uniform improvement during training.
- Longer-training stability: DINOv2-S remains highest, SigLIP and MAE form a middle tier, and ResNet remains least aligned in Wall, Rope, and Granular.Figure 9 evaluates later checkpoints using DINOv2-S@10 as the reference; “late” denotes epoch 50 for most runs and epoch 49 for two PushT runs.
E Spectral kNN Embedding Overlays of Predictor Features … F.3 Representation geometry controls whether latent MSE matches task distance
The appendix uses spectral kNN overlays and theoretical analysis to connect predictor-feature geometry with planning behavior. It argues that rollout accuracy and bounded geometric distortion determine when latent-space MSE is a meaningful task objective.
- E Spectral kNN Embedding Overlays of Predictor Features: Spectral kNN overlays compare predictor layer3 geometries across encoders and epochs using DINO at epoch 10 as the aligned reference.The procedure builds cosine-similarity kNN graphs with K = 10, computes 2D Laplacian Eigenmaps, and applies Procrustes alignment.
- E Spectral kNN Embedding Overlays of Predictor Features: DINOv2, SigLIP, and MAE show relatively clear geometric similarity in PointMaze, Wall, and Push-T, whereas ResNet differs in a task-dependent manner.In Wall and Push-T, where ResNet performs worse, its embeddings are more dispersed and overlap less with transformer-based priors.
- E Spectral kNN Embedding Overlays of Predictor Features: Granular exhibits clustered, noisier convergence, while Rope has more spread-out point clouds, making the overlays an auxiliary diagnostic rather than proof of representational equivalence.The visual patterns are described as consistent with quantitative results and with stronger planning performance often accompanying more similar distributions.
- F Representation Geometry and Planning-Objective Mismatch: The loss–planning decoupling shows that DINO achieves substantially better CEM planning than MAE despite an evaluation loss roughly two orders of magnitude larger.The appendix attributes this mismatch to latent-space MSE depending on rollout error and distortion of the latent representation.
- F.1 CEM objective in latent observation space: CEM evaluates candidate action sequences by minimizing a per-dimension averaged latent-observation MSE objective at the final planning horizon.The objective uses predicted final-step latent observations relative to a latent goal observation.
- F.2 Rollout error induces bounded planning-objective error: Under bounded latent observations, rollout errors induce a bounded mismatch between the predicted-latent and real-latent planning objectives.The bound follows by expanding the difference of squared distances and applying norm inequalities and the triangle inequality.
- F.3 Representation geometry controls whether latent MSE matches task distance: Latent MSE matches task-relevant distance when the representation satisfies local bounded distortion over the visual and proprioceptive pairs encountered during planning.The task objective is defined using distances on true final visual observations and proprioceptive states relative to the goal.
F.4 Decomposition: planning mismatch = rollout error + geometry distortion
Planning mismatch decomposes into rollout error and geometry distortion. Thus, accurate rollouts alone may not ensure effective planning when distorted latent geometry makes Euclidean distances unreliable for task-relevant ranking.
- Decomposition: The standard decomposition separates planning mismatch into geometry distortion and rollout error.The decomposition follows by combining Lemma 1 and Lemma 2.
- Decomposition: Final-step latent prediction errors ev(π), ep(π) control rollout error, while distortion factors (κ, κ) control geometry distortion.These controls follow from Lemma 1 and Lemma 2, respectively.
- Planning implications: Planning can fail from geometry distortion even with accurate rollouts, because Euclidean distances may poorly proxy task-relevant discrepancies and degrade CEM ranking.This explains why planning performance can decouple from loss-based measures.
- Planning implications: Effective planning requires shared transition-critical factors to preserve task-relevant neighborhood and distance structure for goal-conditioned ranking.The paper identifies geometry distortion as a plausible contributor to observed loss-planning decoupling.
- Planning implications: MAE can achieve lower reconstruction loss than DINO yet plan worse when its latent geometry fails to preserve task-relevant structure.This illustrates that reconstruction loss need not predict planning quality.