Source-linked AI summary

The Label Defines the Timescale: Trait-State Limits of Temporal-Aggregate Learning

Xizhe Zhang

arXiv:2608.01587v1stat.MLcs.AIcs.LG

TL;DR

The paper asks whether apparent benchmark ceilings reflect inadequate model capacity or short acquisition protocols that cannot capture long-horizon labels. It derives an exact protocol-conditioned Bayes-risk framework for Gaussian trait–state processes and label-dependent temporal aggregation. The results show that trait and state information scale differently, occupation labels require broader temporal calibration than means, and dispersed observations can outperform repeated same-time segments for state explainability.

  • Problem

    Long-horizon labels are often paired with snapshots or a few short windows, leaving unclear whether performance ceilings reflect model capacity or insufficient acquisition information.

  • Method

    The paper derives an exact protocol-conditioned Bayes-risk identity for temporal aggregates of a Gaussian process containing stable trait and correlated within-individual state.

  • Results

    Trait variation contributes O(1) label variance while state variation contributes O(T^-1); mean labels use ordinary correlation time, whereas occupation labels depend on higher-order correlation times.

  • Takeaways & Limitations

    The relevant temporal scale is defined by the label functional and latent dynamics, so temporally dispersed observations can provide state information that additional same-time segments cannot.

  • Takeaways & Limitations

    Trait-ceiling estimation can use repeated-measurement or test–retest data, but state-ceiling analysis requires dense short-lag temporal calibration or an externally justified kernel family.

Abstract

from arXiv · show

Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their apparent performance ceiling may therefore be an acquisition-protocol ceiling rather than a model-capacity ceiling. We study labels of the form $\Theta_{g,T}=T^{-1}\int_0^T g\{Z(t)\}\,\mathrm{d}t$ when the latent Gaussian process contains both a stable individual trait and a correlated within-individual state. An exact protocol-conditioned Bayes-risk identity provides a common tool. First, we decompose label variance into an $O(1)$ trait component and an $O(T^{-1})$ state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change. Second, we derive task-dependent effective temporal spans: mean labels depend on the ordinary correlation time, whereas occupation-time labels depend on an entire spectrum of higher-order correlation times. Third, state-driven occupation-label variance is maximal when the stable trait lies at the threshold; window efficiency decays much more slowly away from that boundary. Under an equal segment budget, exact risks and Monte Carlo experiments show that repeated segments at one time rapidly saturate, whereas temporally dispersed observations continue to increase state explainability. The trait ceiling uses quantities available from ordinary test-retest data; only the state ceiling requires short-lag temporal calibration. The results distinguish architectural limits from protocol limits and show that the label, rather than duration or segment count alone, defines the relevant timescale.

Introduction

The paper distinguishes model-capacity limits from acquisition-protocol limits when long-horizon labels are inferred from short temporal observations. It develops a trait–state framework showing that the label functional, latent dynamics, and temporal coverage jointly determine what information the protocol can contain.

  • Motivation: Long-horizon benchmarks can plateau because short-window inputs omit information required by the label, creating a sampling-protocol ceiling rather than an architectural ceiling.This distinction matters when labels summarize weeks, operating cycles, or seasons but inputs are snapshots or a few images.
  • Protocol dimensions: The protocol separates measurement precision, local temporal support, and temporal coverage: segments refine a fixed window, whereas window length and placement determine which state innovations are observed.These dimensions are not interchangeable; repeated segments can denoise a snapshot without revealing dynamics outside the observed support.
  • Label-dependent timescales: Occupation-time labels measure the fraction of a horizon above a threshold and depend on the complete thresholded correlation structure, unlike temporal means.Thus, an input’s statistical timescale is jointly defined by latent dynamics and the label functional.
  • Positioning: The paper’s broader contribution connects trait–state modeling, longitudinal sampling, measurement-error theory, and bag-level learning to a protocol ceiling that bounds learners using the same observations.The proposed ceiling is complementary to methods that attach labels to bags or study correlated repeated measurements.
  • Analytic framework: The framework uses correlated Gaussian state–trait processes, linear window observations, and an exact protocol-conditioned Bayes-risk identity to quantify prediction under a given acquisition design.The protocol-level explainability is distinct from the performance of any particular architecture and can be used to separate acquisition limits from model limits.

Protocol Ceilings and Learning Gaps

The paper derives an exact protocol-conditioned risk decomposition separating acquisition-limited information from architecture-dependent prediction error. It shows how stable traits preserve snapshot predictability even when finite-horizon state variation remains poorly observed.

  • The exact risk identity separates a fixed acquisition term from an architecture-dependent term, distinguishing protocol ceilings from model gaps.Only the second term can be reduced by architecture, optimization, or more training objects; the ceiling-utilization ratio is meaningful only under matching target, loss, and protocol assumptions.
  • High total R2 can coexist with weak within-person tracking because total explainability includes the stable trait channel.Conditioning on the trait isolates recoverable state variation and prevents cross-sectional prediction from being mistaken for successful dynamics modeling.
  • Temporal aggregation decomposes into a stable O(1) cross-sectional channel and a finite-horizon state channel.For a noisy snapshot, the stable channel yields a nonzero long-horizon prediction limit, whereas without a trait a fixed snapshot explains a vanishing fraction.
  • A conventional two-occasion test–retest design can identify the trait-channel ceiling, while state-channel quantities require short-lag temporal calibration.The trait ceiling uses cross-occasion covariance and observed variance; estimating the state kernel and higher-order integrals is needed only for state claims.
  • For mean labels, state variation depends on the ordinary correlation time, whereas occupation labels involve every higher-order correlation time.Matching the usual integral correlation time therefore does not guarantee matching temporal information for occupation-time targets.

Task-Dependent Effective Time

Effective temporal information depends on the label functional, not physical window duration alone. The analysis shows that repeated measurements improve local precision but temporally dispersed observations provide additional state support, with occupation labels especially sensitive near threshold boundaries.

  • The effective span of a window is determined by correlation-profile integrals and Hermite weights rather than physical duration alone.For a noise-free OU mean-label window, the effective span approaches 2τ for short windows and w for long windows.
  • Occupation labels use all higher-order correlation times and window profiles, so two kernels with equal ordinary correlation time can provide different temporal information.Mean labels retain only the first Hermite order, while occupation labels combine all nonzero weighted orders.
  • Repeated segments at one time exhaust local measurement noise but have a finite state-information limit; separated windows purchase additional state support.Under an equal raw-segment budget, same-time replication refines an already observed window, whereas separated windows add state observations.
  • State-driven occupation-label variance is maximized when the trait-conditioned threshold is at the boundary.Boundary-near individuals therefore have more absolute residual state error, even though window efficiency can decay much more slowly away from the boundary.

Implications for ML Benchmarks

The paper argues that benchmark progress should be evaluated against the information available under the acquisition protocol, while total cross-sectional prediction and state tracking should be reported separately. Acquisition design and model capacity are distinct experimental axes with asymmetric calibration requirements.

  • Benchmark scores should be compared with the protocol-specific information ceiling rather than the unattainable value R2 = 1.The recommended decomposition separates stable cross-sectional prediction from recoverable within-person change.
  • Increasing training-set size improves estimation of the optimal predictor but does not alter the protocol ceiling; changing the observation protocol does.A useful ablation holds the raw measurement budget fixed while reallocating it between same-time replication and temporal coverage.
  • The trait-channel ceiling can often be reported from ordinary test–retest data, whereas state claims require short-lag temporal calibration.A small densely sampled subset may be more informative for state calibration than another large cross-sectional sample using the same snapshot protocol.

Experiments and Benchmark Implications

The experiments show that acquisition protocol can impose benchmark ceilings: dispersed observations improve state explainability more than same-time replication, while stable traits preserve snapshot predictability over long horizons. The results also establish that temporal calibration requirements and effective spans depend on the label and latent correlation structure.

  • Equal segment budget: 0.808 versus 0.097 explainability at N = 64: dispersed occasions outperform same-time replication under an equal segment budget.The comparison uses an occupation label with α = 0; same-time replication uses D = 1, M = N, whereas dispersed coverage uses D = N, M = 1.
  • Trait-channel ceilings: Positive trait shares produce nonzero long-horizon snapshot ceilings, whereas explainability vanishes as the horizon grows when α = 0.For T/τ = 14, point-noise variance 0.2, and a zero-threshold occupation label, the ceilings are I = 0.119, 0.256, and 0.355 for α = 0, 0.20, and 0.35.
  • Task dependence: Figure 2 shows that equal ordinary correlation times do not guarantee equal occupation-label information, while efficiency decays more slowly away from the threshold.The OU and Matérn-3/2 kernels share τ1, but assign different relative value to the same window; magnitude is concentrated near a = 0, whereas efficiency declines more slowly.
  • Calibration requirements: Trait ceilings can be estimated from ordinary repeated-measurement data, but state-channel analysis requires short-lag temporal calibration or a justified parametric kernel family.Repeated measurements identify the trait parameters without resolving the short-lag kernel, whereas occupation labels depend on higher-order correlation times and the full window profile.
  • Benchmark implications: The experiments support separating acquisition limits from architecture limits when interpreting benchmark plateaus and cross-sectional accuracy.The study evaluates analytic predictions using exact posterior predictors, exact transitions, fine-grid labels, and repeated Monte Carlo experiments.

S1 Model, Notation, and Regularity Conditions

The model represents each object with an independent stable Gaussian trait and stationary Gaussian state process, then applies a square-integrable label functional under explicit summability and correlation assumptions. Linear Gaussian observations, including window averages, define the finite-dimensional acquisition protocol.

  • Latent process: Each latent process combines an independent standard-normal trait M with a zero-mean, unit-variance stationary Gaussian state process X.The state correlation is ρ(u), and the total-process correlation incorporates both trait and state dependence.
  • Label functional: Labels are formed by applying a square-integrable function g to the standardized latent process and aggregating over time.The framework uses Hermite expansions and Gaussian identities to analyze the resulting covariance structure.
  • Observation protocol: The observation protocol is a finite-dimensional linear Gaussian measurement independent of the latent process, with window averages as a special case.This protocol supplies the observations used for exact conditional prediction and Bayes-risk analysis.
  • Regularity conditions: The main asymptotic results assume integrability or summability conditions on the label covariance and its state-dependent differences.Stronger first-moment conditions provide O(T^-2) remainder control, while the effective-span theorem requires summability for each relevant Hermite order.
  • Correlation scope: The exact risk identity permits negative correlations, although the boundary theorem is stated for nonnegative correlations.The distinction concerns the assumptions used for the clean theorem statement, not the general applicability of the exact covariance-based risk identity.

S3 Posterior Predictor Used in Validation

Validation uses the exact conditional Gaussian posterior and protocol-risk identity to compute Bayes predictors and risks. Fixed-support refinement improves prediction monotonically, but additional segments cannot substitute for temporal support outside the observed window.

  • Posterior predictor: The posterior process remains Gaussian with conditional mean mπ(t) and variance 1 − qπ(t,t), providing the inputs to the exact risk calculation.The posterior mean is obtained from the observation–process covariance and the observation covariance matrix.
  • Exact risk: The Bayes predictor is the conditional expectation, and its exact risk is obtained from the protocol-conditioned posterior representation.The computational validation draws independent posterior-process replicas conditional on the observed data.
  • Fixed-support refinement: Refining observations on a fixed temporal support cannot worsen prediction and converges to the predictor based on the full support.The result follows from martingale convergence and the projection property of conditional expectation.
  • Protocol interpretation: Additional segments reduce measurement noise and can improve trait estimation, but they cannot reveal state innovations outside the observed temporal support.The limiting risk can vanish only in exceptional processes without relevant unobserved innovations.
  • Replication axes: Trait ceilings depend on separated occasions and within-occasion replication differently, and ordinary two-occasion test–retest data identify the trait parameters.Same-time replication removes measurement noise but leaves transient state variance; separated occasions additionally average the state component.

S6 Verification of Main Theorem 2: Effective Temporal Span

The effective temporal span for state prediction is derived by conditioning on the stable trait and analyzing window correlations with the occupation label. The resulting theorem allows the span to depend on the label-specific Hermite orders and the observation window.

  • Conditional state label: The occupation-label analysis conditions on the stable trait, converting the label into a trait-conditioned threshold problem for the unit state process.The observation is represented through a standardized noisy state average over a window centered within the long horizon.
  • Effective temporal span: Theorem S2 defines a task-dependent state-effective span under window-integrability and effective-span summability assumptions.The assumptions apply to every Hermite order carrying nonzero weight in the window-correlation expansion.

1. For the mean state label,

Mean-label state information is governed by the ordinary correlation time, while repeated measurements at one support eventually saturate and dispersed windows add temporal state support.

  • For mean labels, the long-horizon state-label variance is 2τ1/T + o(T^-1), and the corresponding effective-span expression follows from the explained state variance.
  • The mean-label effective span depends on the ordinary correlation time, with ℓ_mean(w) = 2τ + (2/3)w + O(w^2/τ) for short windows and ℓ_mean(w) = w + τ + O(τ^2/w) for long windows.
  • Repeated same-time measurements have a finite state-information limit as local measurement noise is exhausted, while separated windows contribute additional state information until sparse additivity saturates.
  • Increasing the number of segments within one window improves precision on fixed support, whereas increasing temporal dispersion can purchase additional state support.The distinction is between reducing measurement noise and expanding the observed temporal set.

S7 Verification of Main Theorem 3: Boundary Localization

For occupation labels, state-driven variance is largest when the trait-conditioned threshold is at the boundary, while effective-span dependence on threshold can be substantially flatter and is not generally monotone.

  • Occupation-label state variance is an even, nonincreasing function of the threshold magnitude and is maximized at the boundary a = 0.Strict decrease holds when the correlation function is positive on a set of nonzero measure.
  • The boundary result follows because the fixed-lag threshold-covariance integrand is even in a and strictly decreases with |a| before integration over lags.
  • No general monotonicity is claimed for the occupation effective span ℓ_a(w); numerical comparisons show it is substantially flatter than the boundary variance coefficient for the investigated OU and Matérn windows.

S9 Sparse Multi-Window Consequence

In the sparse multi-window regime, separated windows add state explainability under explicit spacing conditions, while non-sparse placement requires the full protocol calculation; state-channel conclusions also require richer temporal calibration than trait ceilings.

  • When windows are mutually separated and Dℓ_g(w)/T = o(1), their explained state variances add to first order.The result assumes cross-window correlations are o(1) in the asymptotic regime.
  • Non-sparse placement must use the full q_π expression and is closely related to excursion-set sequential-design problems.
  • The equal-budget experiment compares repeated same-time segments with evenly spaced occasions using exact explainability and Monte Carlo posterior estimates.The study uses N in {1, 2, 4, 8, 16, 32, 64}, with one midpoint occasion for replication and N evenly spaced occasions for coverage.
  • The trait ceiling requires α and measurement-noise variance available from ordinary test–retest data, whereas state-channel analysis requires dense short-lag calibration, longitudinal data, or a defensible parametric family.
  • A population average effective span is a ratio of expectations weighted toward individuals with larger state-driven occupation variance, so it should not be interpreted as the span of a typical object.
Loading 2608.01587v1…