Source-linked AI summary
Counterfactual Evaluation of Temporal Observation Protocols
Xizhe Zhang
TL;DR
Counterfactual protocol evaluation asks whether benchmark data collected under one observation protocol determine the predictive value of undeployed alternatives. This paper develops value-specific identification and calibration theory, showing that unlimited data under the realised protocol may still be insufficient, while targeted or dense additional observations support reliable comparison and design at data-resolved granularity.
Problem
The paper asks whether the joint measurement–target distribution from a realised protocol determines the predictive value of an undeployed observation protocol.
Method
The paper develops value-specific identification theory, covariance-based non-identification tests, finite-calibration error bounds, and cost-constrained target-aware observation design.
Results
Unlimited data under the realised protocol may not identify alternative value, while targeted measurements and dense calibration support reliable comparisons; broad temporal layouts are more reproducible than fine learned placements.
Takeaways & Limitations
Protocol optimisation should follow value identification, with calibration size and candidate granularity matched to the resolution at which protocol values can be compared.
Takeaways & Limitations
Beyond the Gaussian analysis, applying the framework requires characterising the latent laws compatible with realised data and shrinking the decision-relevant value range.
Abstract
from arXiv · showhide
We study counterfactual protocol evaluation: whether data collected under a realised observation protocol determine the predictive value of alternatives that were never deployed. Protocol value is the population $R^2$ of the Bayes-optimal predictor of a fixed trajectory-level target from the measurements an alternative would collect. We show that even infinite benchmark data need not determine this value: distinct latent covariance structures can induce the same benchmark measurement--target law while assigning different values to the same alternative. We develop a value-specific identification theory in which only latent ambiguity that changes the alternative's value matters. For linear targets, invisible covariance directions certify non-identification, while targeted measurements can restore identification without recovering the full latent covariance; an exact permutation construction extends the result to nonlinear aggregate targets. With finite dense calibration data, uniform error bounds control protocol-selection regret and distinguishable value gaps. Exact marginal gains then support cost-constrained, target-aware observation design. Simulations and retrospective analyses of Sleep-EDF and Long-Term AF show that broad temporal-layout differences can be more reliably distinguished than fine placements selected from finite data. Together, these results connect identification, calibration resolution and observation design for undeployed protocols.
1 Introduction
Counterfactual protocol evaluation asks whether data from a realised observation protocol determine the predictive value of undeployed alternatives, especially when targets span longer horizons than observed inputs. The paper develops value-specific identification, finite-calibration guarantees, and target-aware design principles for comparing and selecting such protocols.
- Problem: Counterfactual protocol evaluation asks what predictive performance an undeployed observation protocol could achieve when standard benchmarks hold the realised protocol fixed.The central issue is whether the realised measurement–target law determines the value of an alternative protocol.
- Motivation: Long-horizon targets may depend on temporal information omitted by the realised protocol, as in whole-night sleep-stage proportions and long-horizon atrial-fibrillation burden.Weak performance may therefore reflect limitations of the observation protocol rather than only the learner.
- Identification: An alternative protocol is identified precisely when its value is constant across all latent models inducing the same realised measurement–target law.This value-specific criterion allows protocol value to be identified even when the full latent covariance remains ambiguous.
- Calibration: Targeted augmentation can identify a specified alternative, whereas densely observed calibration trajectories identify the latent covariance needed to evaluate an entire candidate family.Finite calibration data then support uniform value-error bounds, selection-regret guarantees, and a calibration resolution below which value differences cannot be reliably distinguished.
- Design: Observation design should follow identification: collect information that removes ambiguity affecting candidate values, then match candidate-class granularity to finite calibration resolution.The information needed to evaluate protocols need not recover the full latent dependence structure.
2 Problem Formulation and Protocol Value
The section formalizes protocol value for predicting fixed trajectory-level targets from measurements collected by feasible observation protocols, including alternatives absent from benchmark data. It specifies temporal aggregate targets, a discrete Gaussian representation, and Bayes value as population R^2 explained by protocol measurements.
- 2 Problem Formulation and Protocol Value: An observation protocol is a finite collection of acquisition actions, while observation design selects a feasible protocol under a cost budget.The realised protocol generated the benchmark; an alternative is evaluated despite its complete measurement vector being unavailable.
- 2.1 Temporal Aggregate Targets and Observation Protocols: The target is a weighted temporal aggregate obtained by transforming the latent trajectory at each time and averaging over the observation horizon.Uniform weights produce ordinary time averages, while non-uniform weights emphasize selected periods.
- 2.1 Temporal Aggregate Targets and Observation Protocols: Mean and occupation targets represent average process level and weighted exceedance-time burden, respectively, with the target fixed when protocols are compared.Examples include average electricity demand, sleep-stage proportions, atrial-fibrillation burden, and symptomatic-state duration.
- 2.2 Discrete Gaussian Model: The discrete model represents latent states on a time grid with standardized marginal mean and variance, while K captures dependence between times.Quadrature weights encode the target, and protocol matrices with diagonal noise covariance encode observed point or window measurements.
- 2.3 Protocol Value under the Gaussian Model: Bayes protocol value is the population R^2 of the optimal predictor of a nonconstant scalar target based on protocol measurements and lies in [0, 1].The best-linear analogue uses second moments and is no larger than Bayes value, with equality when the conditional mean is affine.
- 2.3 Protocol Value under the Gaussian Model: Under the discrete Gaussian model, nonlinear target dependence enters protocol-value calculations through the covariance transform Cg applied to latent correlations.For the mean target, Cg(r) = r; occupation targets use a distinct transform based on the threshold function.
- 2.3 Protocol Value under the Gaussian Model: Gaussian conditioning identifies the trajectory component recoverable from protocol measurements, yielding total target variance, explained variance, and Bayes-optimal population R^2.The recovered covariance is QS(K), while protocol covariance includes measurement noise through ΣS(K) = ASKA⊤S + RS.
3 Value-Specific Identification
The realised benchmark law may leave latent covariance structures observationally indistinguishable while assigning different predictive values to an undeployed protocol. Identification is therefore value-specific: targeted additional measurements can identify a specified alternative without recovering the full covariance.
- Value-specific criterion: The alternative value is identified exactly if and only if it is constant across all covariance structures observationally equivalent under the realised benchmark.Additional benchmark sampling cannot distinguish members of the same equivalence class.
- Linear targets: For linear targets, covariance perturbations preserving the benchmark’s measurement covariance, measurement–target cross-covariance, and target variance are invisible.A value-changing invisible direction makes the alternative value locally non-identified, so no estimator based only on benchmark data can be consistent at both local alternatives.
- Stationary counterexample: For p ≤3 the stationary lag-correlation vector is identified, whereas for p ≥4 the stationary invisible-direction space has dimension p −3 ≥1.The four-point threshold applies specifically to the stationary, standardised, one-observation model with a uniform mean target.
- Stationary counterexample: For p = 4, observationally equivalent covariance matrices preserve benchmark moments to a maximum discrepancy of 10−16 but yield I(B; K+) = 0.6817 and I(B; K−) = 0.8274.The alternative observes (Z1, Z2).
- Nonlinear targets: Permutation-equivalent covariance structures extend non-identification to arbitrary nonconstant square-integrable aggregate targets whenever the permutation changes the alternative’s value.Without stationarity, three grid points suffice; in the stationary standardised one-point benchmark with a mean target, the sharp minimum is four.
- Targeted calibration: An augmented benchmark identifies I(B; K) exactly when row(B) ⊆ row(Ā), while full-column-rank calibration removes every invisible direction in the unrestricted model.Under covariance restrictions, fewer measurements may suffice: in the four-point stationary model, adding Z1 reduces the invisible space from dimension 1 to 0.
4 Finite Calibration and Protocol Comparison
Finite dense calibration data constrain covariance directions invisible under the realised protocol, and covariance-estimation error propagates to alternative protocol values through target-specific continuity and covariance conditioning. These bounds yield calibration rates, selection-regret guarantees, and distinguishable value gaps for protocol comparison.
- Calibration covariance: Dense calibration observations identify covariance through Cov(W^(i)) = K + R0, with independent units and separately known or estimated calibration-noise covariance.The finite-sample problem is how recovered-covariance error propagates to protocol comparisons.
- Target regularity: Target covariance transforms transmit covariance error linearly for smooth targets, while threshold targets have exponent β = 1/2 globally and β = 1 at interior models.At boundary correlations, the square-root envelope cannot be improved; correlations separated from ±1 recover linear behavior.
- Uniform stability: Uniform protocol-value stability requires target variance bounded below and controlled protocol conditioning, with constants depending on target and family bounds rather than the protocol itself.The resulting continuity exponent is β = 1 for smooth targets, β = 1/2 globally for threshold targets, and β = 1 for threshold targets at an interior model.
- Calibration rates: The global m^-1/4 rate is a worst-case guarantee over correlation matrices, whereas the ordinary root-m rate applies to every fixed threshold model satisfying the interior condition.The fixed-model result assumes positive-definite K, positive target variance, and the stated conditioning bound.
- Protocol comparison: Estimated or population protocol-value gaps larger than 2εm preserve their ordering, while smaller differences cannot be distinguished uniformly from calibration error.For nested protocol classes, the uncertainty bound is separated from approximation loss caused by restricting the search family.
5 Target-Aware Observation Design
Target-aware observation design selects feasible acquisition protocols by maximizing normalized Bayes value, using exact rank-one marginal gains for forward search and refinement. The objective is monotone when protocols are nested, but not generally submodular, so exhaustive search provides a benchmark for small catalogues.
- Design objective: Protocol selection maximizes Fg(S; K), equivalently normalized value Ig(S; K), because target variance is protocol-independent.Feasible protocols may be constrained by costs, catalogue restrictions, and nonsingularity of measurement covariance.
- Optimization: Exact rank-one marginal gains support forward selection with gain-per-cost ordering and one-swap refinement, while exhaustive search benchmarks optimization error for small catalogues.The rank-one update avoids re-forming the full observation-covariance inverse.
- Calibration: Calibration-based design replaces K with bK and maximizes Fg(S; bK), with Corollary 12 bounding the cost of that substitution.Candidate actions vary in acquisition times, window lengths, repetition counts, and noise levels.
- Optimization: Adding measurements cannot decrease Fg, but monotonicity does not imply diminishing returns because even the mean target yields non-submodular R^2 subset selection.Thus greedy search is not guaranteed to have the usual submodular diminishing-returns structure.
- Comparisons: Target-aware design is compared with target-free mutual information, integrated posterior variance, linear-target design, and noiseless kernel quadrature to isolate target and noise effects.These alternatives differ in whether they use the target transformation, actual weights, and measurement noise.
6 Empirical Evaluation
The empirical evaluation tests finite-calibration accuracy, target-aware observation design, and retrospective distinctions among temporal layouts and learned supports. It finds that near-optimal selection can be reliable before uniform value estimation is precise, target-aware search approaches exhaustive optima, and broad layouts are clearer than exact learned locations.
- Evaluation scope: The experiments evaluate family-uniform value error, selection regret, candidate-family resolution, target-aware design efficiency, and retrospective temporal-layout stability.Known data-generating laws support exact comparisons for calibration and design, while Sleep-EDF and Long-Term AF are studied retrospectively.
- Finite calibration: The uniform-error envelope approaches the fixed-model exponent −1/2 as calibration trajectories increase.Full-range log–log slopes are -0.4134 and -0.4142, while the three-largest-size slopes are -0.4618 and -0.4637.
- Finite calibration: Selection regret decreases faster than uniform error, with mean cellwise log–log slopes of -0.94 and -0.49, respectively.The median and maximum regret-to-envelope ratios are 0.024 and 0.37 across six covariance–target cells.
- Observation design: Target-aware one-swap refinement raises minimum heterogeneous-setting efficiency from 0.912 to 1.000 and reaches exhaustive optima in several settings.Remaining minimum efficiencies are 0.994 and 0.987, while evaluating only 0.09–0.23 of exhaustive sets.
- Retrospective evaluation: In retrospective data, dispersion generally outperforms contiguous observation for multi-window AF and several pooled Sleep comparisons, while learned Sleep-support differences include zero.At AF N = 4, contiguous and dispersed templates attain cross-fitted R2 values +0.696 and +0.971, respectively; target-aware versus kernel-quadrature differences have median -0.014 with range [−0.102, +0.078].
7 Related Work
Counterfactual protocol evaluation connects incomplete-record learning, model-based measurement design, acquisition-regime evaluation, and statistical identification. Its distinguishing question is whether an undeployed protocol’s predictive value is computable from the realised measurement–target law when alternative measurements may be unobserved.
- Incomplete temporal records: The paper differs from incomplete-record methods by asking whether an alternative protocol’s value is identified, rather than only estimating covariance functions or learning prediction rules from sparse observations.Functional data analysis uses sampling and smoothness assumptions to estimate mean and covariance functions, while irregular-time-series methods learn prediction or interpolation rules directly from incomplete records.
- Measurement and observation design: Classical, Bayesian, active-learning, sensor-selection, and kernel-quadrature designs optimise estimation, information, uncertainty, or integration criteria under a model, prior, fitted covariance, or kernel.Longitudinal and functional-data designs can use pilot data to estimate dependence and guide future observation times, including prediction of a scalar response.
- Acquisition-regime evaluation: Retrospective active-feature-acquisition and off-policy evaluation methods assess new policies from data generated by another process, typically requiring missing-data assumptions, weighting, or overlap between behaviour and target regimes.A deterministic realised protocol may assign zero acquisition probability to measurements demanded by an alternative, making reweighting insufficient for identifying their distribution or predictive value.
- Identification: The paper applies statistical identification and partial-identification principles to the predictive value of an undeployed protocol, using observational-equivalence geometry and non-identification certificates.The contribution concerns the value functional, not recovery of every latent feature of the data-generating process.
- Identification: Unlike Blackwell experiment comparisons and transportability or data-fusion methods, protocol value is a scalar criterion for one fixed target and loss within the same population under a regime change.The related comparison concerns predictive limits under a different feature-generating protocol, rather than ordering observation schemes across decision problems or transporting quantities across populations or sources.
- Contribution: The work extends a Gaussian identity for fixed-protocol evaluation to identification of undeployed-protocol value, measurement augmentation, finite-calibration comparison, and design.The fixed-protocol identity was derived by Zhang (2026).
8 Discussion
Discussion emphasizes that protocol value depends on information collected, so more data under a realized protocol cannot reveal undeployed alternatives when dependence remains unresolved. Calibration therefore determines comparison resolution and enables targeted or robust observation design across temporal and broader sensing settings.
- More training data or flexible predictors improve realized-protocol performance but cannot reveal undeployed measurement value when the observed law leaves decisive dependence unresolved.Targeted augmentation resolves ambiguity for a specified alternative, whereas dense calibration supports a broader family of protocols.
- Finite calibration limits comparison granularity, while richer candidate classes can reduce selection regret before all candidate values are accurately estimated.Useful resolution depends on value gaps among competitive protocols and the available information to distinguish them.
- Sleep results more reliably distinguish broad dispersed-versus-contiguous layouts than exact learned supports, while AF favors distributing matched budgets across the record over concentrating them in one block.Sleep support locations and held-out advantages vary across targets, source studies, and subsamples; the AF result applies to the multi-window budgets considered.
- The same calibration-before-optimization problem extends to dispersed ECG monitoring, multimodal panels, and new environmental sensor locations because existing measurements may not reveal relevant dependence.These settings are distinguished from ordinary feature selection by the need to evaluate measurements that were not collected.
- When point identification fails, compatible value ranges can still support partial comparisons and robust design using worst-case value or minimax regret.An alternative is evaluable when its value is constant across latent laws compatible with the realized data.
9 Conclusion
Counterfactual protocol value may remain unidentified even with unlimited data under the realised protocol, requiring new kinds of observation rather than more samples of the same kind. Finite calibration determines the scale at which candidate protocols can be compared reliably.
- 9 Conclusion: Unlimited data under the realised protocol may still fail to determine what could be achieved under another protocol.Resolving this failure requires new kinds of observation, not more samples of the same kind.
- 9 Conclusion: Finite calibration sets the scale at which candidate protocols can be compared reliably.
Appendix A. Proofs for Value-Specific Identification … B.2 Covariance Repair and Resolvent Bounds
The appendices prove the Gaussian protocol-value identity, construct covariance perturbations that preserve benchmark observations while changing alternative values, and establish when augmented measurements restore identification. They also derive regularity, covariance-repair, and resolvent bounds supporting finite-calibration analysis.
- A.1 Gaussian Protocol-Value Identity: Posterior-replica cross-covariances yield the Gaussian protocol-value identity, expressing explained variance as F_g(S; K) and minimum Bayes error as V_g(K) − F_g(S; K).The proof uses conditional independence, the Gaussian covariance identity, weighted summation, and total variance.
- A.2 Minimal Stationary Counterexample: The stationary counterexample has rank J = min{2, p − 1} and covariance ambiguity dimension dim ker J = max{0, p − 3}; at p = 4, the kernel is span{(1, −2, 1)}.Positive definiteness persists for sufficiently small perturbations, so the indistinguishable covariance alternatives remain admissible.
- A.2 Minimal Stationary Counterexample: The perturbation preserves benchmark covariances and target variance, while admissibility makes the alternative-value gap strict exactly when b ≠ 0.The cited proof identifies b = Cov(Z1, Θ) = Cov(Z2, Θ) as the condition governing strictness.
- A.3 General, Nonlinear and Augmented Protocols: For general and nonlinear protocols, invisible covariance directions preserve the benchmark law while a nonzero directional derivative changes value; augmented measurements identify the value when row(B) ⊆ row(Ā).The nonlinear example establishes a universal strict inequality for nonconstant g, while Proposition 8 gives the row-space condition for value identification.
- B.1 Regularity of the Target Covariance Transform: Hermite expansions and Gaussian Sobolev identities establish regularity of the target covariance transform, with the supremum attained at r = 1; threshold targets attain the limiting global Hölder behavior.For threshold functions, the sharp constant is attained at c = 0 and (r1, r2) = (−1, 1).
- B.2 Covariance Repair and Resolvent Bounds: Covariance diagonal rescaling preserves the raw estimation rate at an interior K when flooring is inactive and the perturbation is sufficiently small.Weyl’s inequality and operator-norm bounds control the repaired covariance relative to the original covariance.
- B.2 Covariance Repair and Resolvent Bounds: Uniform inverse and resolvent bounds hold when ae ≤ λ/2, with constants worsening only through latent covariance scale, measurement amplification a, and protocol conditioning λ.The bounds are uniform over the feasible family and apply to every non-empty S ∈ Π_B.
B.3 Value Error, Calibration Rates and Regret · B.4 Best-Linear Calibration
The paper derives uniform value-error rates and protocol-selection regret guarantees from calibration error, with root-m behavior under local threshold conditions. It also establishes root-m calibration for best-linear values under uniform conditioning assumptions.
- B.3 Value Error, Calibration Rates and Regret: The empty protocol has value zero under both covariance matrices, while non-empty protocols are handled through the target modulus and uniform value bounds.This separates the degenerate empty case from the non-empty protocol analysis.
- B.3 Value Error, Calibration Rates and Regret: The local threshold bound exploits identical unit diagonals, so only off-diagonal covariance arguments vary, and the ratio identity supplies an explicit uniform bound.The bound also uses 0 ≤ Fg(S; K) ≤ Vg(K).
- B.3 Value Error, Calibration Rates and Regret: For sufficiently small e, the denominator is at least v0/2 and the resulting bound has the stated dependencies through 2L(LβQ + 1)eβ/v0.The conditioning parameter κ remains bounded within a fixed neighborhood of K.
- B.3 Value Error, Calibration Rates and Regret: Value error is controlled uniformly over the finite candidate family when covariance and correlation neighborhoods remain suitably conditioned.The argument uses a common compact correlation subinterval and a local Lipschitz modulus for threshold targets.
- B.3 Value Error, Calibration Rates and Regret: The calibration rate yields root-m value error when β = 1, but only m−1/4 under the global threshold envelope β = 1/2.At a fixed threshold model satisfying (33), the local Lipschitz modulus restores the root-m rate.
- B.3 Value Error, Calibration Rates and Regret: Empirical protocol selection incurs at most 2εℓ regret within candidate class Π(ℓ), relative to that class’s optimum.Subtracting the classwise guarantee from the global optimum gives the corresponding comparison in (37).
- B.4 Best-Linear Calibration: Best-linear calibration achieves root-m value error uniformly when v, ΣS, and the relevant covariance mappings satisfy uniform nondegeneracy and boundedness conditions.The proposition assumes v ≥ v0 > 0, λmin(ΣS) ≥ λ > 0, bounded ∥cS∥2, and uniformly bounded protocol maps AS.
Appendix C. Design Derivations and Algorithms … E.2 Foldwise Selection and Held-Out Prediction
The appendices derive efficient rank-one updates and heuristic forward-plus-swap design algorithms, specify simulation and nonlinear-target evaluation settings, and document foldwise selection and held-out prediction for Sleep-EDF and AF. They also distinguish protocol-value estimands, selection criteria, and pooled out-of-sample performance measures.
- C.1 Rank-One Update: Rank-one covariance updates yield the stated Q and P recursions, with forward scoring and swap-sweep costs given explicitly in p, dB, |V|, and |S|.The forward pass costs O(dB|V|p^2), while trial swaps cost O{|S||V|(p^2|S| + |S|^3)}.
- C.2 Forward Selection and One-Swap Refinement: Forward selection chooses admissible best additions and then improving swaps, but stopping at zero marginal gain is heuristic because Fg need not be submodular.Exhaustive optima are reported for small families.
- C.3 Comparator Objectives: Comparator objectives separately vary target nonlinearity, measurement noise, and target-free latent coverage through mutual information, posterior variance, linear-target, and kernel-quadrature criteria.The linear-target criterion maximises ω⊤QS(K)ω; kernel quadrature uses the same linear criterion with RS = 0.
- D.1 Models, Targets and Protocol Families: Simulations use standardised Gaussian trajectories with OU, Matérn-3/2, mixture, and damped-periodic dependence, five temporal or aggregate targets, and enumerated target-specific protocol optima.Design comparisons use T = 10, p = 128, and α = 0 in the stated settings.
- D.2 Numerical Evaluation of Nonlinear Targets: Nonlinear target calculations use Plackett integrals for one-sided thresholds, covariance transforms for two-sided occupation, and truncated Hermite expansions with Gauss–Hermite quadrature for logistic transforms.These procedures provide the numerical evaluations used for the nonlinear targets.
- E.1 Annotation Mapping and Estimands: Sleep and AF protocols map records to common-grid measurements while defining targets over analysable intervals, with distinct weighting, exclusion, and template-averaging rules.Sleep uses midpoint epochs and AF averages bins; AF excludes unannotated prefixes and treats records as equally weighted analysis units.
- E.1 Annotation Mapping and Estimands: The analysis distinguishes population protocol value Ig(S), foldwise selection criterion Jτ|C(S), held-out predictive performance bR2cf(S), and AF-only full-sample bIL,p(S).These quantities serve different roles rather than representing interchangeable evaluation targets.
- E.2 Foldwise Selection and Held-Out Prediction: Sleep supports are selected within outer-training folds using weighted residual covariance, forward search, and at most 3 one-swap updates; pooled held-out predictions produce one out-of-sample R2 that can be negative.Training uses 79–81 independent subjects per fold, with median 80, and repeated nights do not increase that count.
E.3 Resampling and Sensitivity Analyses
Resampling quantifies conditional held-out variation and pipeline stability, while sensitivity analyses show that broad fixed-template ordering is stable even when fine support selection changes. An AF grid diagnostic finds limited numerical sensitivity across temporal templates.
- Resampling and pipeline stability: Conditional percentile ranges resample held-out pairs 2000 times by Sleep subject and 2000 times by AF record, while 1000 study-stratified 80% subsamples assess pipeline stability.Fixed predictors make the percentile ranges conditional on the realised predictors; each stability subsample reruns standardisation and moment estimation.
- Sensitivity analyses: Alternative weighting and covariance-floor choices preserve fixed-template ordering but can reduce support overlap to 0.14.Subject- versus recording-weighted moments preserve score signs in all 15 cells; floor multipliers {0.5, 1, 2} produce maximum outer-fold mean contrast range 0.0074 and overlap 0.33.
- Sensitivity analyses: SC-only and ST-only sensitivities rerun the full subject-grouped pipeline, while AF cohorts exclude records with AF burden near zero or one.The analyses report Sleep target contrasts at N = 16 and AF fixed-template results at N = 4 and N = 16.
- AF grid diagnostic: Across four AF temporal templates and grid resolutions p ∈{64, 128, 256}, the largest within-protocol range is 0.0073.The diagnostic measures discretisation sensitivity using centred contiguous and uniformly dispersed windows with fractional overlap against grid bins.