Source-linked AI summary

Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control

Alain Bensoussan, Minh-Nhat Phung, Minh-Binh Tran

arXiv:2608.07265v2math.OCcs.LG

TL;DR

The paper addresses when action-conditioned latent prediction can preserve the state distinctions required for deterministic control rather than merely fit one-step data. It introduces an encoder-only local–global metric hinge and proves finite-sample certificates for approximate minimizers under explicit geometric, coverage, approximation, regularity, and optimization assumptions. The resulting certificates support uniform latent dynamics and deterministic planning transfer, while experiments show improved planning for geometry-aware objectives.

  • Problem

    The central gap is that forward prediction alone does not guarantee a non-collapsed, metrically faithful representation or a latent transition uniformly compatible with controlled dynamics.

  • Method

    The paper combines observation-based prediction with an encoder-only local–global metric hinge, an explicit one-sided regularization regime, and finite-sample analysis under observable-geometry and regularity hypotheses.

  • Results

    Every approximate empirical minimizer in the certified regime is co-Lipschitz and uniformly approximately semiconjugate, with explicit error separation and downstream planning-transfer guarantees.

  • Takeaways & Limitations

    Geometry-aware objectives restore latent resolution and strongly improve planning relative to the residual-only reference, while the local–global hinge supplies the direct deterministic certificate.

  • Takeaways & Limitations

    The guarantees require observable-state metric information during training, and the analytic regularization threshold is a class-uniform sufficient condition rather than the smallest empirically useful weight.

Abstract

from arXiv · show

We establish a finite-sample learning-to-control theory for geometrically supervised latent models of nonlinear deterministic systems. Geometric supervision is used only during training: simulator state, proprioception, or state estimates with independently validated metric and directional error bounds supply observable-state distances and tangent directions, while deployment remains observation- and action-conditioned. We introduce an encoder-only local--global metric hinge that enforces directional resolution and separated-state discrimination. Under regular observable-factor, coverage, finite-capacity approximation, and uniform $C^{1,1}$ hypotheses, a computable one-sided regularization regime has a strong selection property: with high probability, every approximate empirical minimizer is simultaneously pointwise co-Lipschitz and uniformly approximately semiconjugate to the controlled dynamics. Approximation, sampling, and optimization errors remain explicit and separate. Norm-constrained tensor-product B-spline classes constructively realize the approximation hypotheses, and the interpolation exponent converting mean residual control into a uniform bound is sharp. A modular deterministic corollary transfers the learned certificates to trajectory, finite-horizon cost, learned-cost-head, and optimizer guarantees, while a validated finite-net result enables sharper model-specific certification. Controlled experiments isolate collapse and folding, quantify the analytic certificate's reserve, and demonstrate the control benefit of restored metric resolution. The principal contribution is a complete finite-sample implication from approximate empirical optimization to metric faithfulness, uniform controlled dynamics, and reliable planning for the same learned model.

1. Introduction

The paper asks when observation-based latent prediction becomes metrically faithful enough for deterministic control, addressing collapse and folding that prediction loss alone cannot exclude. It introduces a geometrically supervised local–global hinge and proves that, under explicit assumptions and regularization, approximate empirical minimizers yield certified representations, dynamics, and planning transfers.

  • Motivation: Prediction loss alone permits collapsed or folded representations, so small one-step error does not ensure state distinctions needed for worst-case deterministic planning.The constant encoder can achieve zero prediction loss, while global variance or covariance does not imply pointwise metric resolution.
  • Method: The encoder-only local–global metric hinge penalizes directional collapse and insufficient separation between states at prescribed observable distances.Its one-sided penalty vanishes after local and global margins are reached, without requiring a decoder, inverse model, or reconstruction objective.
  • Planning transfer: The certified representation and dynamics constants feed a modular deterministic transfer to trajectories, finite-horizon costs, learned cost heads, and approximate optimizers.The framework also provides an explicit regularization half-line and a sharper finite-net, model-specific a posteriori certificate.
  • Theory: Under observable-factor, coverage, approximation, regularity, and optimization hypotheses, every approximate empirical minimizer is pointwise co-Lipschitz and uniformly approximately semiconjugate with high probability.The theorem keeps approximation, statistical, and training-optimization errors explicit and separate.
  • Construction and evidence: Norm-constrained tensor-product B-splines constructively satisfy the finite-capacity hypotheses, while controlled experiments isolate collapse and folding and show improved planning after metric resolution is restored.The experiments compare geometry-aware objectives with a residual-only reference and quantify the analytic certificate’s reserve.

2. Relation to coercivity, metric embeddings, abstraction, and planning transfer

The paper frames latent-control learning as a coercivity problem: prediction residuals become useful only when paired with pointwise geometric stability. Its theorem connects metric faithfulness and uniform controlled dynamics to deterministic planning transfer.

  • Coercivity and stability: Prediction residuals require a stability mechanism before becoming scientifically meaningful for control.The paper identifies this as a recurring coercivity theme in inverse problems and statistical learning.
  • Metric embeddings: The local hinge controls directional singular values, while the separated-pair term resolves distant folds.Together they propagate sampled lower-margin control to a continuum co-Lipschitz certificate under C1,1 regularity and coverage.
  • Abstraction and semiconjugacy: The observable factor isolates recoverable control state and receives both a lower metric certificate and a uniform deterministic controlled-semiconjugacy estimate.This connects observable abstraction with system-level behavioral closeness.
  • Planning transfer: The representation theorem supplies the two constants needed for planning transfer: η for latent dynamics mismatch and c∗ for compatible latent-cost Lipschitz bounds.The deterministic transfer is formulated in sup norm for optimization over action sequences.
  • Finite-sample implication: The paper’s objective-level result forces pointwise non-collapse and uniform semiconjugacy for all approximate minimizers in the certified regime.These certificates feed directly into deterministic optimizer transfer.

3. Deterministic controlled systems and the observable quotient

The paper replaces an unavailable full-state representation with a finite-dimensional observable quotient that preserves distinctions recoverable from controlled observation histories. Under explicit regularity and metric assumptions, the quotient supports Euclidean analysis and identifiable observation coordinates.

  • Observable quotient: When H is non-injective, observation-based representations can recover only information contained in controlled observation histories.The paper therefore uses an observable quotient as the correct state object.
  • Observable quotient: Observational equivalence identifies states whose future observation sequences coincide under every prescribed action sequence.The controlled composition notation makes this equivalence explicit.
  • Regular factor structure: Assumption A1 requires a finite-dimensional Euclidean observable factor, with dynamics, observations, and costs factoring through it.The descended observation map is bi-Lipschitz onto its image and the relevant extensions are C1,1.
  • Deployment interface: The theorem’s deployed encoder uses the current observation through the identifiable factor, not controlled histories.Histories motivate quotient construction but are not inputs to the deployed encoder.
  • Metric scope: All lower co-Lipschitz claims use the Euclidean observable-factor metric, which need not equal the quotient chain metric.Translating bounds to the chain metric requires an additional bi-Lipschitz comparison.
  • Coverage: Coverage is imposed through lower Ahlfors regularity of the training measure on the state-action space.The condition supplies a lower bound µo(Br(s,a)) ≥ m0 r^(do+da) over the prescribed radius range.

4. Action-conditioned latent world models

The action-conditioned latent world model combines an observation encoder with a latent transition, while the analysis constrains their regularity and residual uniformly. Metric non-collapse additionally enables compatible latent costs and existence of regularized empirical minimizers.

  • Model definition: An action-conditioned latent world model consists of an encoder Φ and transition map F, with only Φ restricted to H(S) entering the objective.The deployed model is observation- and action-conditioned through the induced latent state.
  • Metric non-collapse: A positive local directional margin requires Dψ(s) to be injective, while lower co-Lipschitzness yields injectivity and compatible latent costs.The induced total latent cost evolves under the learned transition from z0 = ψ(s0).
  • Regularity assumptions: Metric results require latent dimension dz ≥ ds and specified ambient C1,1 representatives for admissible encoders.The induced map ψ = Φ ◦ H provides the equivalent Euclidean-factor formulation.
  • Latent costs: Metric non-collapse permits exact compatibility extensions for Lipschitz stage and terminal costs on the encoded state set.The construction uses the lower co-Lipschitz estimate and McShane extension arguments.
  • Prediction residual: The JEPA residual is R[Φ,F](s,a) = ψ(G(s,a)) − F(ψ(s),a), and uniform Lipschitz structure bounds it on the state-action domain.This residual measures disagreement between encoding after physical evolution and latent evolution after encoding.

5. Geometric obstructions for pure forward prediction and covariance/spectral spread

Pure forward prediction and global spread penalties do not ensure the pointwise geometry required for deterministic planning. The local–global metric objective addresses this by targeting directional collapse and separated-state overlap directly.

  • Pure prediction: Pure forward prediction admits collapsed zero-loss minimizers, so prediction alone cannot imply non-collapse.A constant encoder and transition can achieve zero residual while retaining no positive lower metric constant.
  • Global spread: Positive latent variance or covariance does not prevent folding or loss of local metric resolution.Global spread can coexist with non-injective representations.
  • Counterexample: Covariance or spectral spread alone cannot imply the desired pointwise lower co-Lipschitz property without additional structural assumptions.These penalties rule out total collapse only in a global second-moment sense.
  • Counterexample: A deterministic example has zero residual and positive variance while ψ(s) = ψ(−s), proving non-injectivity.The example uses S = [−1,1], identity dynamics, and ψ(s) = Cs^2.
  • Geometric remedy: The local–global objective supplies directional and separated-pair information for establishing a uniform co-Lipschitz certificate.Its local and global terms target the distinct geometric failures left unconstrained by prediction and spread penalties.

6. The metric hinge regularizer

The paper introduces a soft, encoder-only local–global metric hinge that promotes directional resolution and separation of well-separated states using geometrically supervised training data.

  • The metric regularizer directly promotes infinitesimal directional resolution and separation of well-separated states without requiring a decoder or inverse model.
  • Geometric supervision supplies observable-state distances and tangent directions during training while deployment remains observation–action conditioned.
  • The local and global penalties are averaged hinge losses encoding directional and separated-pair metric conditions.
  • Metric samples comprise independent state-action, unit-direction, and separated-pair families under lower Ahlfors regularity assumptions.
  • The empirical objective combines observation-based prediction loss with the metric regularizer weighted by λ.

7. Finite-sample geometric coercivity and controlled semiconjugacy

Under explicit geometry, approximation, coverage, regularity, and optimization hypotheses, the theory converts approximate empirical optimization into uniform metric and controlled-dynamics certificates.

  • A smooth bi-Lipschitz latent realization and C1,1 representatives provide the faithful comparator required by the finite-capacity argument.
  • Norm-constrained tensor-product cubic B-spline classes satisfy the approximation hypothesis with capacity-independent Lipschitz and C1,1 budgets.
  • With probability at least 1 −δ, every εtrain-approximate empirical minimizer is pointwise co-Lipschitz and uniformly approximately semiconjugate to the controlled dynamics.
  • The theorem separates approximation, statistical, and optimization errors through the bound ηW,n,δ(εtrain) := ΘM(β(W) + 2εstat,W (n, δ) + εtrain).
  • The admissible regularization strengths form a nonempty computable half-line, while a validated finite-net proposition supplies a typically sharper model-specific certificate.
  • The 1/(q + 2) interpolation exponent converting mean residual control into a uniform estimate is sharp under the stated regularity.

8. Deterministic planning as a downstream corollary

The representation and dynamics certificates feed a modular deterministic planning transfer that bounds trajectory, cost, learned-head, and optimizer errors over finite horizons.

  • Theorem 7.11 supplies representation and dynamics certificates that the planning corollary transfers into deterministic trajectory, cost, and optimizer guarantees.
  • Compatible learned cost heads add explicit stage-cost and terminal-cost discrepancies to the optimizer-transfer bound.
  • For every fixed horizon T and every ξplan-optimal latent action sequence, Lipschitz dynamics and costs yield finite-horizon planning bounds.
  • When LF = 1, the horizon factors are C(T, 1) = T and D(T, 1) = T(T −1)/2; for LF > 1, the worst-case envelope grows geometrically.
  • Uniform residual control is required because a localized high-error region can be visited by a deterministic optimizer despite a small population mean.

9. Numerical evidence

Controlled experiments test collapse, folding, finite-capacity threshold behavior, and nonlinear latent-MPC performance under aligned protocols. Geometry-aware training restores lower metric resolution and is associated with more reliable control than residual-only prediction.

  • Folding stress test: The combined local–global hinge is injective in all 12 plain-initialization stress-test runs.
  • Folding stress test: Under strongly symmetry-biased even pretraining, upper control reduces non-injective outcomes from 9 to 3 out of 12.
  • Finite-capacity threshold diagnostic: The finite-capacity spline proxy shows an intermediate regime where sampled local and global margins strengthen while uniform evaluation residual remains controlled.
  • Nonlinear latent-MPC benchmark: Pure prediction attains the smallest evaluated residual, bηeval = 0.01, while geometry-aware objectives raise the lower metric constant.
  • Nonlinear latent-MPC benchmark: At T = 5 and T = 20, pure prediction’s median paired closed-loop cost differences are 1.41 and 1.55, respectively, between 4.03 and 25.79 times regularized-arm medians.
  • Nonlinear latent-MPC benchmark: The geometry/residual ratio correlates with worst-case paired cost difference at 0.62, whereas evaluated residual alone correlates at −0.34.

10. Discussion: scope, certification, and outlook

The framework uses certified geometric supervision during training while preserving observation- and action-conditioned deployment. Its analytic threshold is conservative, and experiments motivate sharper model-specific certification.

  • Scope: Geometric supervision supplies observable-state distances and tangent directions with validated errors, while deployment remains observation- and action-conditioned.Pixel-based geometry estimators fit the framework only after their errors are absorbed into positive effective margins.
  • Certification: The analytic regularization threshold is a class-uniform sufficient condition rather than the smallest empirically useful weight.A posteriori finite-net certification can provide sharper model-specific guarantees.
  • Outlook: Experiments report that several regularized objectives restore non-collapse and achieve competitive control, while the local–global hinge aligns exactly with the pointwise theorem.The analytic certificate's reserve motivates class-specific or finite-net certification.

11. Conclusion

The paper connects geometric regularization to reliable latent control through certificates for representation and dynamics. A modular planning result then transfers those certificates to trajectory, cost, learned-cost-head, and optimizer guarantees.

  • Conclusion: The local–global hinge supplies a pointwise stability mechanism linking empirical prediction to deterministic planning.The principal theorem establishes this link under explicit assumptions and a certified regime.
  • Conclusion: The principal theorem makes every approximate empirical minimizer co-Lipschitz and uniformly approximately semiconjugate to the controlled dynamics.These are simultaneous representation and transition certificates.
  • Conclusion: The modular planning result transfers the certificates to trajectory, cost, learned-cost-head, and optimizer bounds.This extends the theorem from latent-model properties to finite-horizon planning guarantees.
  • Conclusion: Constructive approximation theory, sharp interpolation, finite-net certification, and controlled experiments complete the learning-to-control chain.The results distinguish analytic certification from finite-model evidence while supporting reliable finite-horizon planning.
  • Conclusion: The paper acknowledges AI assistance for language editing, code review, and manuscript checking, with the authors retaining responsibility for the content.

S1. Exact realizers and detailed finite-capacity spline construction

The supplement constructs exact non-collapsed realizers and verifies quantitative local inverse regularity. These constructions provide the concrete smooth and finite-capacity ingredients needed for the theorem.

  • Exact realizers: The supplement constructs an exact non-collapsed latent realization with encoder, transition, and zero residual.The realization satisfies Φ0(H(s)) = ψ0(s), F0(ψ0(s), a) = ψ0(G(s, a)), and R[Φ0, F0] ≡0.
  • Exact realizers: A smooth affine bi-Lipschitz embedding, range enforcement, coordinatewise extension, and metric projection produce globally admissible representatives.The construction keeps the encoder and transition within the compact latent range.
  • Exact realizers: The exact representatives satisfy the controlled realization identity, with the residual vanishing identically.This supplies the realizability baseline for finite-capacity approximation.
  • Inverse regularity: Quantitative co-Lipschitzness yields injectivity, uniformly invertible derivatives, and local C1,1 inverse charts with uniform radii and bounds.The local estimate is 1/2|s −s′| ≤ |Fs0(s) −Fs0(s′)| ≤ 3/2|s −s′|.

0 A0 is invertible and Ps0 is the Moore–Penrose

The detailed construction verifies finite-capacity C1,1 approximation using range-constrained tensor-product splines and establishes a nonempty margin-clearing regime. These results connect exact realizers to empirical certification.

  • Finite-capacity approximation: Finite-capacity classes are required to approximate exact representatives in C1 while maintaining Lipschitz and C1,1 budgets independent of capacity.The approximation error δ(W) tends to zero as capacity W increases.
  • Spline construction: Range-constrained tensor-product cubic B-splines provide constructive approximation with mesh-dependent C1 error and uniform C1,1 control.For coefficient count W_h ≍ h^-d, the construction gives a concrete architecture-level realization of the approximation assumption.
  • Spline construction: The spline classes enforce latent-range constraints through a smooth projection that remains the identity near the target image.This preserves the approximation on the relevant compact set while controlling the class regularity.
  • Margin-clearing regime: The margin-clearing comparator bridges realizability and empirical optimization by providing a positive separated-state margin.Any α in (0, α0(ρ)) is valid when the comparator's co-Lipschitz separation is positive.
  • Margin-clearing regime: The resulting parameter regime enforces directional resolution and separated-state discrimination through derivative and pairwise lower bounds.The constraints require ∥DψW(s)[v]∥Z ≥ κ and ∥ψW(s) − ψW(s′)∥Z ≥ α on their specified sets.

S2. Architecture-specific covering estimates

The supplementary analysis verifies that norm-constrained tensor-product cubic B-spline encoder and transition classes satisfy the finite-parametric assumptions needed for covering estimates. Uniform spline stability, range enforcement, and parameter-Lipschitz bounds yield coefficient radii and complexity growth independent of mesh width where required.

  • Finite-parametric verification: The concrete spline dictionaries verify the finite-parametric deviation hypothesis for the encoder and transition classes.The result connects the abstract deviation corollary to the norm-constrained spline architecture.
  • Architecture: The encoder and transition dictionaries use cubic tensor-product B-splines on coordinate boxes, with basis counts scaling as h^-d for the relevant input dimensions.The encoder and transition parameter dimensions are determined by their respective mesh widths and domain dimensions.
  • Spline stability: Fixed-degree uniform tensor-product B-splines have uniformly bounded local overlap and mesh-independent coefficient stability constants.The local proof reduces possible cell patterns to finitely many reference configurations and uses local linear independence.
  • Range enforcement: The range-enforcing C1,1 map preserves uniform value and derivative parameter bounds because its derivatives are bounded on a fixed compact set.Composition changes constants through bounds on the range map, while retaining the displayed scaling on the relevant domains.
  • Covering estimates: The combined spline coefficient set lies in a cube with radius independent of capacity, and the resulting covering quantities grow at most polynomially in capacity.This supplies the radius choice RW = R in the finite-parametric deviation corollary.

S3. Secondary deterministic-control diagnostics

The deterministic diagnostics explain why uniform residual control complements metric faithfulness: rollout errors amplify with transition dynamics, mean residuals can hide harmful spikes, and concentrated errors can alter planning outcomes. The supplementary results also establish sharp interpolation limits and identify how metric and dynamical certificates transfer to control guarantees.

  • Rollout amplification: Uniform residual errors accumulate according to the rollout factor: they remain bounded for LF < 1, grow linearly for LF = 1, and grow geometrically for LF > 1.For η = 10^-2, Table S1 reports the corresponding cumulative factors across horizons.
  • Interpolation sharpness: The interpolation exponent 1/(q + 2) converting Lipschitz L2 control into uniform control is sharp under Lipschitz-only regularity.A truncated-cone family attains the scaling that rules out any uniformly larger exponent.
  • Mean versus uniform error: An L2-mean residual can be arbitrarily small while the supremum residual remains fixed, especially as the spike support becomes concentrated in higher dimensions.The spike construction shows why average prediction error does not provide the supremum control required by the simulation lemma.
  • Adversarial planning: Approximately 60% more mean excess closed-loop cost is incurred by Model A despite nearly equal L2 one-step errors, because its worst-case error is concentrated near the planning decision boundary.The adversarial-planner diagnostic compares the two surrogate models under receding-horizon control.
  • Certificate transfer: The metric constant c* controls latent-to-state resolution and cost regularity, while the semiconjugacy modulus η controls dynamical error accumulated over the planning horizon.Together these quantities support cost transfer, optimizer bounds, state recovery, Lyapunov estimates, and latent reachable-set approximations.
  • Scope boundary: A state-space pullback of latent reachable-set guarantees requires an additional decoder or projection assumption.The stated theorem directly provides latent-space Hausdorff control, not an unconditional state-space version.

S4. Numerical protocols, summarized records, and supporting-material availability

The experiments show that geometric objectives improve metric faithfulness and planning relative to pure prediction, while supporting diagnostics clarify folding, conditioning, and optimization effects.

  • Folding and covariance: Zero prediction residual and full-rank latent covariance can coexist with non-injective representations, so covariance alone does not certify metric faithfulness.The constructed folding family remains non-injective despite positive-definite covariance and cleared minimum-eigenvalue penalties.
  • Ordinary initialization: The local–global hinge reaches injective profiles in 0 of 12 non-injective outcomes, with median bcmin = 0.90, bcmax = 1.31, and bκgeom = 1.48.The reported diagnostics combine lower and upper geometry with the geometric conditioning measure.
  • Objective comparisons: Pure prediction records 10 non-injective outcomes and median bcmin = 0.25, whereas geometry-aware objectives restore nontrivial metric resolution and improve planning.The pure-prediction arm also has the smallest evaluated residual, separating residual accuracy from metric faithfulness.
  • Upper-distortion diagnostics: Adding the upper directional term reduces expansive outcomes from 9 to 3 and lowers bcmax from 20.78 to 2.55.The recovered runs retain favorable conditioning, with bκgeom = 1.78 versus 1.47.
  • Finite-capacity calibration: At n = 2048 and λ = 0.03, the dense-grid local defect is 0, lower chord ratio is 0.68, and evaluation-set residual is 0.02.For large λ, restart choice strongly affects the attained lower chord ratio and residual, giving a numerical interpretation of εtrain.
  • Planning benchmark: In the planning benchmark, pure prediction has median bcmin = 1.75×10^-4 and bηeval = 0.01, with paired cost differences of 1.41 at T = 5 and 1.55 at T = 20.These differences are between 4.03 and 25.79 times the corresponding medians of the regularized arms.
Loading 2608.07265v2…