Source-linked AI summary

How Wrong Can a Good Predictor Be? Diverging Updates with Vanishing Predictive KL

Qifu Wen, Shuaijun Liu, Zihan Zhou, Xi Zeng, Ningxin Su

arXiv:2609.11132v1cs.LGstat.ML

TL;DR

The paper asks whether a large mismatch between Bayesian update maps must cause predictive failure. It compares exact Bayesian mixing with a deterministic radial filter in a fixed finite-state Gaussian HMM and proves that internal separation can diverge while predictive KL vanishes, including in stationary expectation. The result is a fixed-K counterexample that identifies decoder sensitivity and loss-weighted separating states as missing links between internal gaps and task cost.

  • Problem

    The paper asks when a strict computational mismatch in a fixed-state sequential model forces worse prediction, beyond distinctions among representability, learnability, and task cost.

  • Method

    The paper compares exact Bayesian mixing and an explicit deterministic radial filter on the same K−1 belief coordinates in a stationary symmetric Gaussian HMM.

  • Results

    For every fixed finite K, internal separation can diverge while predictive KL vanishes at a common witness and in stationary expectation as switching becomes rare.

  • Takeaways & Limitations

    An unbounded internal update gap does not by itself certify predictive failure because decoder sensitivity and the contribution of separating states to expected loss matter.

  • Takeaways & Limitations

    The construction fixes finite K before q→0+, provides no rate uniform in growing K, and does not establish a universal compression criterion or learned-model transfer.

Abstract

from arXiv · show

Accurate posterior prediction need not require accurate approximation of Bayesian updates. We prove that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM. Exact Bayesian mixing and an explicit deterministic radial filter act on the same $K-1$ belief coordinates. As $q\to0^+$, their separation in centered logits in the worst case grows at least linearly in the natural confidence scale $L_K(q)$, while their categorical $D_{\mathrm{KL}}(\mathrm{exact}\|\mathrm{radial})$ vanishes at the same explicit witness. Along stationary HMM trajectories, the expected terminal KL between filtered posteriors also converges to zero at $H(q)=\lceil-\log(q)/c\rceil+1$. Typical blocks without switches drive both filters into a common confidence cone, where softmax curvature suppresses their disagreement; a single Gaussian maximal event controls adaptive noise. A sweep with equally spaced Gaussians over $K\in\{2,4,8\}$ illustrates the opposing trends, and binary controls at long horizons compare saturating and nonsaturating recurrences. The result isolates two missing links between internal update gaps and predictive cost: the contribution of separating states to expected loss and decoder sensitivity. Thus even an unbounded internal update gap does not by itself certify predictive failure. The construction is fixed in $K$ and does not provide a universal criterion for when compression is harmless or characterize when internal gaps must incur task loss.

1 Introduction

The paper studies when a fixed-state computational mismatch incurs predictive loss, distinguishing internal representability from distribution-weighted task cost. It constructs a sequential counterexample where update maps diverge while predictive KL vanishes.

  • 1 Introduction: Predictive cost depends on decoder sensitivity and loss-weighted separating states under the evaluation path law, not pointwise representability alone.The paper separates representability, learnability, and task cost as distinct questions.
  • 1 Introduction: For every fixed finite K, exact Bayesian mixing and a radial filter can have diverging centered-logit separation while categorical predictive KL vanishes at the same witness.The divergence grows with the confidence scale L_K(q).
  • 1 Introduction: Typical no-switch blocks place both filters in a common confidence cone, while a maximal Gaussian event controls adaptive noise and softmax curvature suppresses disagreement.A uniform KL envelope handles the remaining paths.
  • 1 Introduction: A finite-scale sweep over K ∈{2, 4, 8} illustrates separated centered states alongside shrinking predictive KL, while binary controls compare saturation behavior.The sweep is illustrative rather than a universal asymptotic claim.
  • 1 Introduction: Under stationary symmetric Gaussian-HMM trajectories, expected terminal predictive KL also converges to zero as switching becomes rare.The result uses separate recurrent filters driven by shared observations.

2 Related Work

Related work distinguishes recurrence expressivity, learnability, task-aware compression, approximate filtering, stability, and detection delay. The paper positions its contribution as an explicit sequential counterexample linking internal mismatch to distribution-weighted predictive cost.

  • 2 Related Work: Expressivity results identify computations that specified architectures cannot represent, while learnability results show that representable functions can still be difficult to learn.Neither issue alone determines predictive cost under a task distribution.
  • 2 Related Work: Predictive cost additionally depends on decoder sensitivity and the loss-weighted contribution of separating states, with unbounded KL not controlled by vanishing event probability alone.This motivates an explicit sequential construction with a proved path-law loss.
  • 2 Related Work: Task-aware compression studies when discarded information can be harmless, whereas this paper fixes the state dimension and update family and lets their centered-logit distance diverge.The paper does not optimize a representation or bit rate.
  • 2 Related Work: Approximate Bayesian filtering and short-memory benchmarks target different maintained objects or predictive quantities than this theorem’s categorical filtered posteriors.The paper maintains K −1 belief coordinates and studies an explicit approximation operation.
  • 2 Related Work: Filter stability studies merging filters under the correct model, while robustness theory perturbs kernels or controls policy costs; the present recurrence is a different comparison.The related frameworks address different assumptions and objectives.
  • 2 Related Work: The binary affine control is adjacent to quickest-change-detection methods, but the paper uses it as a diagnostic for its recurrence comparison.The related detection literature emphasizes delay criteria and optimality.
  • 2 Related Work: State-space models motivate the question because deterministic-state architectures expose recurrence and tracking constraints, but the paper does not prove that any named architecture implements the construction.The result concerns explicit maps rather than an architecture class.

3 Problem Setup: Gaussian HMM with a Fixed Finite State and Radial Filter

The setup uses a stationary symmetric K-state Gaussian HMM and compares exact Bayesian transition mixing with a deterministic radial map on the same centered belief coordinates. Both coupled filters receive identical observations and produce categorical posterior outputs.

  • 3 Problem Setup: Gaussian HMM with a Fixed Finite State and Radial Filter: The latent process has K ≥2 states, pairwise distinct Gaussian means, symmetric switching probability q, and stationary uniform initialization.The transition stays in the current state with probability 1−q and switches uniformly otherwise.
  • 3 Problem Setup: Gaussian HMM with a Fixed Finite State and Radial Filter: Observations are Gaussian conditional on the latent state, and logits are centered so common shifts do not affect categorical predictions.The centered representation removes the all-ones component from belief logits.
  • 3 Problem Setup: Gaussian HMM with a Fixed Finite State and Radial Filter: Exact transition mixing is defined on centered logits, while the radial filter replaces the update geometry with a norm-dependent radial map.The radial map has continuous value R_q(0)=0 and scale α_K(q) at the origin.
  • 3 Problem Setup: Gaussian HMM with a Fixed Finite State and Radial Filter: Both filters start from zero centered logits and process the same observations through their separate recurrences.Their terminal categorical outputs are compared as approximations to the filtered posterior.
  • 3 Problem Setup: Gaussian HMM with a Fixed Finite State and Radial Filter: The theorem fixes 0<c<d_min and evaluates the stationary recurrence at H(q)=⌈(−log q)/c⌉+1.Constants may depend on fixed K, the full mean vector, σ, and c.

4 Main Result and Proof Mechanism

For fixed finite K, exact Bayesian mixing and the radial filter can separate without predictive failure: their worst-case centered-logit gap diverges, while decoded predictions agree and stationary expected KL converges. The proof separates worst-case map geometry from trajectory-level distributional behavior.

  • Main results: For every fixed finite K, the update maps diverge at a moving witness while their decoded predictions agree there.The centered-separation theorem concerns the maps in the worst case, whereas the matched decoder result concerns categorical KL.
  • Main results: At horizon H(q)=⌈(−log q)/c⌉+1, expected predictive KL converges to zero under the stationary HMM path law.This is a distributional result along trajectories, not merely a claim about one hand-chosen input.
  • Proof mechanism: A no-switch block typically drives both filters into a shared confidence cone, while one Gaussian maximal event controls adaptive noise.Softmax curvature then suppresses the effect of their internal disagreement, and a uniform KL envelope handles remaining paths.
  • Scope: The results are fixed-K statements and are not uniform in growing K, nearly colliding means, bounded inputs, or entire architecture classes.The numerical evidence and exclusions are distinguished from the first three mathematical results in the summary table.

5 Evidence at Finite Scale for Fixed K

A finite-scale Gaussian experiment illustrates the theorem’s opposing trends for fixed K: internal state separation increases while categorical predictive KL decreases. The measurements are illustrative rather than proofs of asymptotic behavior.

  • Finite-scale sweep: Across K∈{2,4,8}, sampled centered state distance increases while categorical KL decreases at finite scale.The experiment uses equally spaced, centered means and evaluates multiple confidence scales with paired stationary paths.
  • Interpretation: The curves do not establish the theorem or a uniform onset in K.Appendix F supplies the estimands, coupling, uncertainty calculation, and summary statistics.

6 Binary Specialization and Geometry in the Finite Regime

The binary specialization reduces centered logits to one scalar and contrasts exact Bayesian mixing with a saturating tanh-style radial witness. Its finite-regime diagnostics show why internal geometry and decoded KL can move in opposite directions.

  • Binary reduction: In the binary case, centered logits collapse to one scalar, with exact mixing represented by Fq(h) and the radial witness by Gq(h).This slice visualizes finite-regime geometry and isolates why an affine map cannot represent exact mixing.
  • Exact-map geometry: The exact mixing map saturates, satisfying |Fq(h)|<L(q), and contracts logit perturbations globally by at most 1−2q.The contraction factor is attained at the midpoint.
  • Finite diagnostics: Figure 2 contrasts growing centered distance with falling categorical KL for K=2,4,8, while binary controls compare horizons and recurrence behaviors.The right-side control results use ten seeds; the analytic logistic-decoder illustration is not sampled trajectory data.
  • Interpretation: The rare-switching family has q>0 but its quantitative margin is not bounded away from zero, so it cannot yield a nonvanishing asymptotic predictive lower bound.Predictive loss is evaluated after the logistic decoder under the HMM path law rather than in worst-case logit norm.

7 Binary Controls and Learned Transfer

The experiments compare learned and scalar recurrences against exact Bayes at theorem-matched and longer horizons. Saturating recurrences are more stable at long horizons, while learned transfer remains limited and architecture-dependent.

  • Learned transfer: Under distillation, every nonreference comparison arm meets the half-closure threshold at all six tested q values, but these experiments do not establish transfer to learned models.The closure definition calibrates the scalar S6 arm at 0 and analytic tanh at 1.
  • Long-horizon generalization: At long horizons, analytic tanh stays between 1.7 × 10^-4 and 6.4 × 10^-3 predictive KL, while scalar S6 degrades to 1.9×10^-1.The endpoint uses sequences of length ⌈8/q⌉, up to 2048 steps.
  • Learned transfer: At theorem-matched horizons, the end-to-end architecture-faithful arm reaches half closure at four of six tested q values.The passing values are q ∈ {2^-5, 2^-6, 2^-7, 2^-8}; the asymptotic theorem supplies no finite-q cutoff.
  • Binary controls: The binary controls show that saturating recurrences remain stable at long horizons, whereas affine and identity controls degrade; at q = 2^-8, affine loss reaches 5.53 nats.Hard clipping beats tanh at q = 2^-8, while tanh is better at q = 2^-3, so the theorem establishes one saturating instance rather than an optimum.

8 Limitations

The results are asymptotic and fixed in K, with assumptions on the Gaussian HMM and no universal compression criterion. The finite sweep and learned experiments illustrate the construction but do not establish broader transfer or converse guarantees.

  • Scope: The theorems fix finite K before q → 0+, allow constants to deteriorate with K, and prove no rate uniform in growing K.The moving witness gives neither a gap on every bounded input nor a lower bound on state dimension.
  • Scope: The construction compares two explicit maps and does not provide a universal criterion for pruning, quantization, distillation, low-rank adaptation, or other compression.The converse question—when an internal mismatch must incur nonvanishing task loss—remains open.

9 Conclusion

The paper proves that fixed-K Bayesian and radial updates can diverge in centered logits while decoded predictive KL vanishes, including in stationary expectation. Experiments illustrate these opposing trends without extending the theorem to growing K or learned architectures.

  • Conclusion: For fixed finite K, centered-logit separation can grow linearly while decoded categorical KL vanishes at the same witness.A stationary symmetric Gaussian HMM also yields convergence of expected predictive KL at the logarithmic horizon.
  • Conclusion: The sweep over K ∈ {2, 4, 8} illustrates increasing state separation alongside shrinking predictive KL, while binary controls compare short- and long-horizon recurrence behavior.These are controlled numerical illustrations rather than evidence for a uniform onset rate.

A Open Affine Control Lower Bound

The affine-control appendix analyzes why a nonsaturating recurrence can remain wrong after a switch for Θ(1/q) steps, while exact Bayes recovers sooner. The remaining long-horizon lower bound requires controlling a joint terminal event rather than summing worst-case errors step by step.

  • Mechanism: The affine recurrence is sluggish after a switch because its unclamped state remains near d/(2q) and needs Θ(1/q) steps to change sign.Exact Bayes recovers on a shorter confidence-building timescale.
  • Open step: The long-horizon separation remains unclaimed because the banded final row in the stepwise lower-bound table is still open.All other deterministic and probabilistic ingredients are proved.
  • Quantitative diagnostic: The affine crossing calculation yields y⋆ = 0.175866 and limiting run-age mass 1 − e^-y⋆ = 0.161270.The measured wrong-side fraction is approximately 0.18 to 0.20 across q = 2^-5 to 2^-12.
  • Proof obstacle: A stepwise union bound makes the switch-window probability vanish, whereas controlling the partial sum preserves a constant-probability window.The former also shrinks the guaranteed crossing horizon to Θ(1/(q log(1/q))).
  • Geometric setup: For K > 2, centered logits remove the additive-logit redundancy, and pairwise log odds provide coordinates on the meaningful K−1-dimensional belief space.The theorem fixes K while q → 0+, so the confidence scale diverges without increasing state dimension.
  • Geometric setup: Radial mixing preserves the direction of a centered state and changes only its radius, whereas exact probability mixing generally changes direction.Observation scores are added afterward and may rotate the trajectory.
  • Proof roadmap: Theorem 3 proves an explicit centered-logit gap of order L_K(q), while Theorem 4 proves vanishing expected categorical KL under the stationary path law.A common confidence cone and KL envelope connect internal divergence to low predictive cost.

D Proof Anatomy, Constants, and Quantifiers

For each fixed finite K, the proof separates a diverging map-level witness from stationary expected-KL recovery, using adaptive Gaussian control, path-mixture Bayes analysis, and confidence-cone curvature.

  • Quantifiers: For every fixed finite K ≥2, the theorem takes q →0+ only after fixing the means and variance, so it is not uniform in K or colliding means.The quantifier order permits finite unions over states and competitors but excludes a joint K,q limit.
  • Radial dynamics: The radial filter reduces each no-switch block to accumulated signal and Gaussian-noise coefficients, despite its adaptive recurrence.The state remains in the plane spanned by the signal vector and common noise direction; the analysis uses two scalar processes rather than reducing the model to two states.
  • Noise control: A pathwise Abel inequality and one Gaussian maximal event control all adaptive noise weights across the full horizon.The proof avoids separate timewise tail bounds and a horizon union bound.
  • Recovery cone: The tanh recurrence reaches signal scale O(L^2/3), which dominates the noise scale O(√L log L) after the transient.The exponent 2/3 balances the unit signal increment against tanh’s cubic correction.
  • Exact Bayes: Exact Bayes concentrates on the constant latent path because wrong constant paths lose exponentially and switched-path mass has expectation O(Hq), which vanishes.With H = Θ(L) and q = Θ(e^-L), the aggregate switched-path contribution disappears.
  • Expected KL: Softmax curvature suppresses the terminal disagreement once both filters share a confidence cone, while a deterministic 2L envelope controls the exceptional paths.The good-event KL bound decays polynomially times e^-c*L^2/3, and the complement probability makes its bounded contribution vanish.

E.8 Proof of diverging updates with vanishing predictive KL

The explicit witness makes centered update separation grow linearly in L while categorical KL vanishes, and the fixed-scale experiments illustrate these opposing slopes without serving as proof.

  • Explicit witness: The witness is centered and ignores common logit shifts, so the comparison concerns gauge-invariant pairwise output differences.The chosen input is zq = L(1/4, −1/4, 0, …, 0).
  • Explicit witness: DKL(softmax Φq(zq) ∥ softmax Rq(zq)) ≤ 4K(K −1)L^2e^-L/5 →0 at the same witness where the map gap diverges.Both maps place the same coordinate ahead by a margin proportional to L, allowing softmax curvature to suppress the categorical discrepancy.
  • Finite-scale illustration: Table 5 regresses centered exact-to-radial separation and log DKL(exact∥radial) on L for equally spaced Gaussian illustrations at K ∈ {2, 4, 8}.The table illustrates the theorem’s two directions at finite scale and is not used in the proof.
  • Learned models: The learned-model experiment does not establish transfer of the proved mechanism, although distilled comparison arms meet the stated half-closure threshold at all six switch probabilities.End-to-end selective SSM training meets that threshold at only four of six values.
  • Evaluation conventions: The scalar comparisons use paired observations and separately disclosed clipping conventions, so clipped and untruncated KL values are not interchangeable.The theorem and Table 7 use DKL(exact∥approximation), while the learned-model experiments use their own float32 clipping settings.

F.5 Errors relative to the latent state

The reported errors are relative to the latent state, not disagreement with the exact filter. Experiments find stable long-endpoint saturating controls, while nonsaturating controls degrade under the disclosed clipped metric.

  • Measurement distinction: Wtruth and Wjoint measure errors relative to the latent state, not disagreement with the exact filter.Both indicators are defined relative to the true latent state, so the experiment does not measure exact-filter disagreement.
  • Errors relative to the latent state: 0.1982±0.0064 to 0.1868±0.0036: Wtruth remains in this range across q = 2^-5, 2^-8, 2^-10, and 2^-12.The corresponding Wjoint values rise from 0.0794 ± 0.0027 to 0.1856 ± 0.0036, qualified by confidence.
  • Scope of the empirical conclusions: Long-endpoint saturating controls remain stable, whereas the two nonsaturating controls degrade under the disclosed clipped metric.These observations are deliberately limited and do not establish tanh as the unique stable recurrence.
Loading 2609.11132v1…