Source-linked AI summary
Revelation Control
Qinyou Wang
TL;DR
Learning systems can look equivalent under current summaries yet require different actions after future training. The paper develops Revelation Control to price decision-relevant revelation and distinguish information value from productive reuse, finding that deeper probes can reveal decision-changing distinctions while predictive variation that does not cross the action boundary has no value.
Problem
Current state summaries may merge learning states that look equivalent now but respond differently to future training and require different actions.
Method
The framework prices future-learning interventions by revelation depth, separating decision-relevant information from reusable progress and productive computation.
Results
Predictive variation that does not cross the declared decision boundary has zero binary decision value, while deeper observations can yield strict refinement value under local aliasing conditions.
Takeaways & Limitations
Decision-sufficient revelation depends on whether hidden distinctions change the action after future training, not merely on whether they improve prediction.
Takeaways & Limitations
The local separation is conditional on a declared frontier-ambiguity regime and does not assert that every empirical short-probe representation satisfies it.
Abstract
from arXiv · showhide
Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific.
1 Introduction · 2 Decision value and decision factorization · 3 Decision-sufficient revelation and revelation depth
The paper defines Revelation Control as choosing priced future-learning interventions that reveal only decision-relevant hidden state while separately accounting for reusable computation. It formalizes decision-sufficient revelation, dynamic decision value, and revelation depth, with explicit scope conditions on the resulting structural claims.
- 1 Introduction: Revelation Control chooses actions, future-learning interventions, revelation depth, and which tested future to promote, because current summaries can alias states that respond differently to later training.The framework separates information, execution technology, learnability, and cost, while distinguishing predictive information from information that changes the declared action.
- 1.1 Contributions: The paper contributes decision-sufficient revelation, productive probe-and-promote technology, dynamic continuation value, and a model-specific evaluation protocol.Its central distinction is between discard-and-restart probing and productive revelation that leaves reusable progress for execution.
- 1.2 Scope and organization: The claims are conditional rather than universal: finite-library results remain finite-library, structural refinement assumes a declared shallow ambiguity regime, and empirical conclusions are confined to the studied learning systems.The comparative claim uses one strengthened DynaMiCS-style short-probe frontier under a common evaluation setup.
- 2 Decision value and decision factorization: Decision factorization concerns only the action-relevant quotient: refinement has strict value exactly when no policy measurable under the coarse information is Bayes-optimal under the refined information.Predictive information can have zero decision value when it does not change the declared action, while a small quotient may suffice despite a high-dimensional latent state.
- 2.2 Dynamic decision-factorization obstruction: Dynamic insufficiency can persist despite zero current loss: with positive continuation weight, unresolved downstream distinctions make a coarse representation inadequate after the chosen action changes the training state.The decomposition isolates continuation obstruction from immediate aliasing and is not presented as a replacement for POMDP or dual-control theory.
- 3 Decision-sufficient revelation and revelation depth: Decision-sufficient revelation need not recover the full execution state; it must reveal the distinctions within current-information fibers that can change the terminal action.The companion work establishes predictive non-sufficiency [40], while this paper asks when those hidden distinctions become actionable decision information.
- 3.1 Exact binary aliasing identity: Binary refinement has positive value exactly when refined posterior states within a coarse-information fiber retain mass on opposite sides of the terminal decision boundary.Predictive variation that never crosses the declared boundary has zero value for the binary action menu.
- 3.2 Local geometry and revelation depth: Revelation depth is the first admissible probe depth sensitive to a decision-relevant direction, so a deeper probe can separate states that a short frontier observation locally aliases.The equal-budget separation is conditional on the stated smoothness, rank, depth, and persistence conditions; it does not assert that every empirical H4 representation satisfies them.
4 Productive revelation technology · C Decision Split
The paper distinguishes restart probing from productive revelation, showing that retaining a selected tested path converts equal compute into greater revelation depth while separating policy quality from productive reuse. Revelation matters only when it separates currently aliased states across the action boundary, not when it reconstructs irrelevant hidden state.
- 4 Productive revelation technology: A trial protocol of depth h produces an observation and leaves behind a state whose continuation consequences must be analyzed alongside the revealed information.This motivates treating revelation technology as both an information-acquisition process and a state-changing intervention.
- C Decision Split: Revelation has decision value only when a currently aliased distinction separates states requiring different terminal actions.Deeper future-learning responses need only separate the decision quotient; full hidden-state reconstruction is unnecessary.
- 4 Productive revelation technology: The framework explicitly distinguishes pure information refinement from technology expansion, preventing productive continuation gains from being misclassified as information value.Restart and promotion therefore represent different execution technologies rather than interchangeable implementations.
- C Decision Split: Productive revelation retains the selected tested path, so probing acquires information while partially executing the candidate action.This differs from restart probing, which discards the tested path and begins the selected action from the original anchor.
- C Decision Split: Promotion is credit for work that remains valid on the deployment path, not free compute; unsafe or unusable probes require restart instead.The productive-revelation identity therefore applies only when the tested path can be safely continued.
- 4.1 Exact equal-budget identity: Proposition 4.1 gives an exact equal-budget identity linking restart probing at depth c to productive probing at depth h when K_F(c) = K_AR(h).For two actions, the matching depths satisfy h = 2c.
- 4.1 Exact equal-budget identity: An H4 temporary probe and an H8 productive probe are compute-matched in the two-action H12 comparison, with both methods receiving 20 policy-visible updates.The equal-budget match is feasible within the declared horizon under the stated depth condition.
- 4.2 End-to-end value decomposition: The end-to-end value decomposition separates policy quality under common restart outcomes from the productive reuse value of the tested path.These components can be estimated on independent panels without mixing absolute outcomes from unrelated future banks.
5 Adaptive revelation control: the dynamic closure
This section closes revelation control dynamically by embedding deeper probing into conditional Bayes refinement and optimal stopping. It identifies when scalar summaries suffice for priced continuation decisions, when extra revealability coordinates are decision-nonredundant, and how model-specific proxies must be validated without target leakage.
- Dynamic closure: Dynamic revelation is conditional Bayes refinement: deeper probing has state-dependent continuation value, so static information gain becomes an optimal-stopping problem with priced sequential costs.The framework connects static refinement value to continuation decisions and finite-depth Bellman closure, rather than treating dynamics as a separate objective.
- Scalar-control factorization: A scalar shallow summary is decision-sufficient exactly when each scalar fiber lies entirely on the priced Stop or Continue side, with zero-gain ties value-neutral.Decision-nonredundancy requires cost-adjusted boundary crossing within a scalar fiber; conditional variation in continuation value alone is insufficient.
- Revealability geometry: Future revealability need not track current confidence: states with similar current margins can have different continuation values, and scalar-margin control is strictly suboptimal when conditional support crosses the priced boundary.Under the stated monotonicity conditions, positive conditional mass on both strict Stop and Continue sides establishes decision-nonredundancy.
- Model-specific instantiation: Legal model-specific revealability proxies must be fixed without target outcomes and evaluated for held-out control value beyond a scalar-margin comparator.Proxies may estimate revealability scale, conditional spread, or continuation gain, but noisy fitted scores do not directly identify latent Bayes quantities.
- Scope and limitations: Cross-model transport is structural rather than parametric: theory, cost accounting, continuation logic, and stopping decisions transfer, but proxies, coefficients, thresholds, and risk-to-utility bridges remain model-specific.Bounded decision-instability risk alone does not certify positive expected utility without an explicit severity or tail assumption.
6 From oracle revelation to learnable and certifiable control
Section 6 distinguishes oracle revelation value from finite-sample learnability, showing that richer telemetry can increase decision value while worsening estimation error. It also gives a cost-adjusted learned-frontier criterion and clarifies why bounded stop–flip risk alone cannot guarantee positive expected utility under unrestricted severity.
- Finite-sample revelation ledger: The finite-sample revelation ledger separates oracle information value, promotion value, learner approximation error, and acquisition cost in learned net value.This distinction makes clear that a refinement can improve oracle Bayes value while remaining a poor learned representation at finite sample size.
- Finite-sample revelation ledger: More telemetry can raise revelation value while increasing estimation complexity enough to worsen finite-sample estimation risk, motivating compact rather than maximal refinements.Development analyses with substantially broader telemetry exhibited this estimation obstruction.
- Minimal learnable refinement: A minimal learnable refinement is learner-, sample-size-, acquisition-technology-, and tolerance-relative, rather than necessarily the fewest features or the coarsest sufficient sigma-field.The construction minimizes a learner-relative complexity measure over refinements meeting the specified tolerance.
- Learned frontier victory: At equal resource cost, active learning is sufficient to win when Δrev + P_A > R_A − R_F, combining deeper revelation with productive promotion value.Equation (6.8) separates whether deeper future learning reveals decision-relevant structure from whether productive reuse improves end-to-end utility.
- Risk calibration and certification: Bounded Bernoulli calibration can control shallow/deep decision instability, but any nonzero stop–flip risk permits unboundedly negative worst-case expected net utility when severity is unrestricted.A strict expected-net-utility guarantee additionally requires control of the stopped-harm tail.
7 Related work and novelty boundary
The paper situates Revelation Control within established decision-theoretic, sensing, metareasoning, stopping, and probing traditions, while locating novelty in their coupling for hidden learning-state aliasing and productive future-training experiments. It claims a budget-indexed, finite-sample, utility-comparative framework that adaptively controls revelation depth without treating generic supporting tools as new.
- Terminology and foundations: The paper does not claim the phrase “value of revelation” or a new information-value theory, distinguishing its continuation value from established decision-analysis usage and Bayesian design traditions.It also distinguishes its controlled learning-state acquisition from mechanism-design revelation principles concerning strategic reporting and incentive compatibility.
- Partial observability and sensing: Revelation Control joins partial observability, dual control, and active sensing, but specifically studies future training as intervention on an internal learning execution state.Its target is controlled acquisition of decision-relevant state through system response, rather than a general replacement for those traditions.
- Local prospective probes: The framework asks whether local-probe information states factorize the priced decision, conceding DynaMiCS’s short prospective probes and local finite-difference summaries as prior art [14].Its distinct question concerns decision-relative sufficiency and non-factorization, not merely estimating local cross-domain slopes for data-mixture optimization.
- Finite-sample extraction: The paper treats representation complexity as part of the decision problem because richer telemetry can increase oracle information value while reducing learned value through estimation error.This minimal-learnable-refinement principle rejects the assumption that more legally available telemetry is automatically better.
- Novelty boundary: The strongest novelty claim couples hidden learning-state aliasing with productive future-training experiments, budget-indexed revelation depth, finite-sample extraction, and comparative utility.The paper explicitly disclaims inventing future probing, promotion, Bayes value, value of computation, selective early exit, finite-sample risk calibration, or dynamic control.
- Adaptive control: Its adaptive-control layer treats current decision margin and learning-state revealability as joint coordinates for purchasing deeper revelation, while characterizing when scalar control suffices.Generic metareasoning and risk-calibration tools remain prior art, including rational metareasoning, knowledge gradients [13], and finite-sample risk methods.
8 Transformer instantiation · 9 Fixed-depth revelation beyond current-information comparators · 10 Frontier comparison: DynaMiCS-style probing
Across independently instantiated Qwen2.5-7B and Mistral-7B-v0.3 systems, the framework preserves its decision geometry while allowing model-specific revelation proxies and heads. Qwen’s fixed-depth H8 policy beats every declared current-information comparator with positive simultaneous lower bounds, while the equal-compute frontier comparison tests productive reuse against strengthened DynaMiCS-style probing.
- 8 Transformer instantiation: The two model families share the structural experimental design but use independently prepared Mistral streams, histories, and development data, so transfer is structural rather than numerical.Both use LoRA adaptation and AdamW optimizer state, while the cross-model target does not assume numerical transfer of empirical proxies or coefficients.
- 8 Transformer instantiation: The common protocol probes candidate actions to H8, applies a development-fixed decision rule, and promotes the selected tested path to terminal evaluation at H12.The intervention matches the immediate adaptive field while changing hidden optimizer moments, enabling later common training to reveal differences absent from the immediate update.
- 8.1 Qwen2.5-7B prespecified active policy: Qwen’s active policy uses 24-dimensional raw H8 responses with Ridge(α= 10), a zero threshold, same-path promotion, and fallback to Keep for ties or nonfinite scores.The representation and rule were fixed before independent target panels; H8 was selected development-only, and broader telemetry increased estimation burden without reliable compensating gain.
- 8.2 Independent Mistral-7B-v0.3 instantiation: Mistral retains the H4/H8/H12 geometry, action menu, terminal utility, promotion technology, and 20-update accounting while fitting numerical heads and summaries on independent development data.Its instantiation explicitly does not require Qwen’s revealability proxy, coefficients, or stopping threshold to transfer numerically.
- 9 Fixed-depth revelation beyond current-information comparators: Every one of Qwen’s 22 current-information comparisons has a strictly positive point margin and simultaneous max-t lower bound in both independent panels.Figure 4 displays all comparisons without pooling the panels, and Table 2 defines the componentwise worst-case validation criterion.
- 9 Fixed-depth revelation beyond current-information comparators: Productive-path value is independently positive in both Qwen panels, but the claim is limited to the declared comparator library and same-path execution technology.The reported one-sided Student-t and bootstrap lower bounds are positive in both panels before panel-wise combination with the separate common-restart comparison.
- 10 Frontier comparison: DynaMiCS-style probing: The DynaMiCS-style comparator strengthens short-probe, local-slope, and restart information while adapting the source mixture-selection method to the present binary action gap.This avoids a deliberately weak mapping from slope information to action gap that could otherwise make a positive frontier result reflect head misspecification.
- 10.1 Equalized compute contract: Both methods receive exactly 20 policy-visible updates and identical terminal utility, while productive revelation reuses selected-probe updates and the comparator does not.Restore/save overhead is excluded from the primary comparator accounting, making the productive-revelation comparison conservative.
11 Structural decision refinement in the H4 frontier-ambiguity regime · 12 Productive revelation and equal-compute frontier performance
Across Qwen and Mistral, productive reuse of deeper probes—not restart-only selection—produces strict equal-compute gains, while independent-bank tests show that deeper revelation can refine shallow ambiguity into action-relevant distinctions. Mistral independently confirms positive deeper-revelation value, whereas stronger Qwen partition diagnostics are not treated as cross-model invariants.
- 11.1 Independent-bank refinement estimator: Within Qwen’s 226/432-anchor H4 ambiguity event, independent-bank refinement estimates agree near 1.93 × 10−4 and imply full-panel contribution 1.00952551446 × 10−4.The full-panel contribution has t lower bound 5.69892792476 × 10−5 and bootstrap lower bound 3.39715294723 × 10−5.
- 11.1 Independent-bank refinement estimator: The refinement estimator uses one bank to select the best coarse constant action and an independent bank to evaluate fixed H8 decisions, with symmetric reversal and bootstrap reselection.This prevents terminal outcomes from being reused for both choice and evaluation.
- 11.2 Opposite-action sign replication: The Qwen H8 partition separates states favoring opposite actions, with sample Bayes gap 1.63115121441 × 10−4 and bootstrap lower bound 6.04584733171 × 10−5.An H8 representation also improves cross-bank terminal-gap mean squared error by 1.24677 × 10−7, but the learned-policy value interval crosses zero.
- 11.3 Independent Mistral aliasing and deeper-revelation value: On Mistral’s independent 480-anchor panel, the same shallow H4 ambiguity regime contains stable subsets requiring opposite terminal actions across both terminal banks.Bank A yields 216/480 ambiguous anchors, including 88 stable-Xi and 58 stable-Keep anchors; Bank B yields 210/480, including 83 and 58.
- 11.3 Independent Mistral aliasing and deeper-revelation value: Mistral’s fixed H8 action has strictly positive terminal decision value relative to H4, with Student-t lower bound 1.56391515827 × 10−4 and bootstrap lower bound 1.58226436880 × 10−4.H4 and H8 actions differ on 100/480 exact anchors in at least one bank, reproducing positive deeper-revelation value on nontrivial mass.
- 12.1 Qwen equal-compute closure: Qwen’s two productive-path panels yield end = 8.95147503879 × 10−5 and end = 1.47365323110 × 10−4, both with positive 95% lower bounds.The productive-path values are 8.9913 × 10−5 and 1.4776 × 10−4, respectively, with positive one-sided t and bootstrap lower bounds.
- 12.2 Independent Mistral equal-compute replication: The Mistral equal-compute end-to-end advantage has one-sided Student-t lower bound 2.00162462592 × 10−4 and bootstrap lower bound 2.04870462705 × 10−4.Same-path reuse contributes with one-sided Student-t and bootstrap lower bounds 1.96017743591 × 10−4 and 2.01665987607 × 10−4.
- 12.2 Independent Mistral equal-compute replication: Across both Transformer families, productive reuse of deeper tested computation creates strict equal-budget end-to-end advantage, while restart-only selector differences are not the main gain.This pattern is consistent with the productive-revelation mechanism formalized in Equation (4.6).
13 Adaptive Revelation Control across model families · 14 Cross-model structural validation · 15 Potential application domains
Across Qwen2.5-7B and Mistral-7B-v0.3, Revelation Control transfers structurally: deeper revelation and productive reuse improve compute-accounted utility, while shallow decision sufficiency is model-specific. The framework may extend beyond learning systems when interventions are priced, responses refine consequential decisions, and terminal utility and safety costs are declared.
- 13.1 Qwen2.5-7B: decision-nonredundant revealability and held-out safe compute: Qwen’s two-coordinate controller improves paired cost-aware value over the scalar |𝑓H4| rule by 3.2514×10−7, with positive one-sided Student-t and bootstrap lower bounds and no harmful held-out stops.The two-coordinate result supports decision-nonredundant shallow revealability, while its strict net-value lower bounds remain sensitive to finite-sample tails.
- 13.1 Qwen2.5-7B: decision-nonredundant revealability and held-out safe compute: Qwen’s H8-over-H4 terminal decision value is 1.229 × 10−4, with 95% Student-t and bootstrap lower bounds of 8.687 × 10−5 and 8.769 × 10−5.This independent 480-history full-grid study fixed the continuation architecture, H4/H8 policies, update price, and compute ledger prospectively.
- 13.1 Qwen2.5-7B: decision-nonredundant revealability and held-out safe compute: Qwen’s calibrated rule stops on 4/480 histories, saves 0.1667% of visible updates, keeps the 95% joint-risk bound below 2%, and has positive net value 2.890 × 10−7.The bound is 0.6222%, while Student-t and bootstrap lower bounds are 5.161 × 10−8 and 7.225 × 10−8.
- 13.2 Mistral-7B-v0.3: scalar continuation control at the tested resolution: Mistral’s scalar utility-aware controller retains positive familywise-adjusted lower bounds on an independent 480-anchor target panel, supporting scalar decision sufficiency at the tested resolution.The controller was fit only on a separate 336-anchor development panel, and the raw Qwen-style R4 coordinate did not improve adaptive value.
- 13.3 Risk–severity separation in Mistral: Mistral’s risk-calibrated rule illustrates that low action-instability does not ensure positive utility: four stopped action changes coexist with a largest continuation value of 9.811 × 10−3 and negative net value.Rare high-severity misses can outweigh many cheap correct stops, separating risk control from utility control.
- 14 Cross-model structural validation: Across both families, future learning supplies decision-relevant information, deeper paths carry productive value, and Revelation Control improves compute-accounted utility, but fitted controllers do not transfer unchanged.The transferable object is the relation among legal shallow information, continuation value, productive cost, and terminal decision utility.
- 15 Potential application domains: Potential applications include staged learning updates, adaptive numerical computation, robotics, process diagnostics, scientific experiments, operations pilots, and high-stakes human decisions, but these remain prospective rather than established deployment claims.Applicability requires a consequential terminal action, hidden state, priced intervention, observable response, and defined utility and safety costs.
- 15.3 A domain-level applicability test: For a new domain, declare legal information, terminal actions, utility, admissible probes, and costs, then use development-only responses to estimate continuation value and test whether a scalar summary is decision-sufficient.The central transfer test is structural alignment with Revelation Control, not superficial resemblance to model fine-tuning.
16 Limitations · 17 Conclusion · A Conservative replicate-denoised Bayes-regret bridge
The paper concludes that controlled future learning can reveal decision-relevant hidden distinctions and reuse the resulting work for higher utility, while empirical claims remain bounded by system, comparator, resource, and tail-certification limitations. An optional replicate-denoised bridge provides conservative Bayes-regret certification under explicit assumptions, but clearing it is sufficient rather than necessary.
- 16 Limitations: Productive revelation is appropriate only when tested work remains legally and scientifically reusable, and the main frontier does not charge restore/save overhead or other deployment resources.If probes corrupt deployment, create irreversible risk, change the target distribution, or add state-I/O or safety costs, restart may be appropriate and the budget identity must be repriced.
- 16 Limitations: The location–scale theorem is conditional: a second shallow coordinate is required only when a positive-mass scalar fiber crosses the priced Stop/Continue boundary.One-step decision sufficiency also does not imply multi-stage Markov closure; using (m_h, σ_h) as a complete dynamic state requires an additional closure condition.
- 16 Limitations: Positive held-out utility or bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity; population certification requires a predeclared severity, moment, or integrable-tail class.The limitation is identified as an identification boundary rather than merely a power limitation.
- 16 Limitations: The conclusions do not establish universal method optimality, dominance over every measurable current-information policy, or transfer beyond the evaluated learning-system configurations.The empirical systems share LoRA, AdamW-style optimization, a binary intervention menu, and H4/H8/H12 geometry; policies were fixed before target evaluation, and the comparator was one strengthened DynaMiCS-style short-probe/restart class.
- 17 Conclusion: The framework identifies a decision quotient as the relevant state representation: a coarse observation suffices when the refined optimal action factorizes through it, while productive revelation separates information gain from reused computation.The fork–probe–promote design treats future learning as an experiment on hidden learner state and embeds static Bayes refinement into continuation value.
- 17 Conclusion: Across Qwen2.5-7B and Mistral-7B-v0.3, controlled future learning shows action value, productive reuse yields strict equal-compute advantages, and adaptive revelation supports state-dependent control.Qwen also shows a decision-nonredundant shallow revealability regime, while Mistral supports scalar continuation sufficiency only within the tested architecture family and resolution.
- A Conservative replicate-denoised Bayes-regret bridge: The optional replicate-denoised bridge uses independently generated future outcomes to separate persistent conditional signal from one-bank noise and conservatively upper-bound coarse-policy Bayes regret under square-integrability, shared conditional means, and c_X≥0 assumptions.The bridge also applies conditionally on any positive-probability event measurable in the coarse information.
- A Conservative replicate-denoised Bayes-regret bridge: The bridge is sufficient rather than necessary: failing to clear it does not imply that active value is below the coarse-information Bayes value.Persistent hidden heterogeneity contributes a nonnegative floor even when the score equals the coarse-information conditional mean.
B Adaptive-depth certification details · C Proofs
The appendix proves that productive-prefix reuse preserves decisions when deeper and shallow policies agree, while showing that stop–flip risk alone cannot certify utility without controlling the severity of harmful flips. It therefore requires prospective certification to combine bounded-risk calibration with a predeclared severity, moment, or tail condition.
- B Adaptive-depth certification details: Risk–severity factorization shows productive-prefix reuse makes no decision difference when shallow and deep policies agree, with |Y| = |D|F otherwise.This isolates decision changes to stopped flips and ties their magnitude to terminal decision severity.
- B Adaptive-depth certification details: For any nonzero stop–flip risk and unrestricted stopped-flip severity, the same risk rates can yield arbitrarily poor expected net utility, so risk alone has no finite robust lower bound.The impossibility holds for every fixed stop probability and stop–flip probability with finite positive-harm expectation.
- B Adaptive-depth certification details: The impossibility concerns deriving a nontrivial lower bound on unbounded mean severity from event probability alone, aligning with classical nonparametric impossibility results for unrestricted mean inference.It does not imply that bounded-risk calibration itself is impossible.
- B Adaptive-depth certification details: Bernoulli-risk procedures such as fixed-sequence calibration, Learn-then-Test, and Conformal Risk Control can prospectively control how often early stopping changes the deeper action, but not how severe those changes are.The procedures require a fixed risk model, threshold family, calibration dataset, and risk target.
- B Adaptive-depth certification details: A probability budget is not a utility budget: lower flip frequency can still produce negative net value when one harmful flip has sufficiently large severity.Table 6’s fixed-data nested evaluation contrasts a positive-net-value direction with a lower-risk direction whose single harmful flip makes realized net value negative.
- B Adaptive-depth certification details: A valid expected-utility certificate must combine bounded stop–flip-risk calibration with control of the integrated stopped-harm tail or a declared severity, moment, or envelope surrogate.The resulting finite-sample lower bounds are assumption-matched to the declared severity class or an independently valid upper bound on H+.
C.1 Proof of Proposition 2.2 … C.8 Proof of the replicate-denoised bridge
The appendix proves the paper’s main decision, cost, value, refinement, geometric, and denoising results by reducing each claim to conditional Bayes identities, budget accounting, sign changes, or projection inequalities. Together, these arguments establish the stated equivalence conditions, resource comparisons, lower bounds, and policy-improvement guarantees.
- C.1 Proof of Proposition 2.2: The proof of Proposition 2.2 uses conditional Bayes actions and nonnegative regret to characterize when an H-measurable action is also G-optimal.A measurable tie rule ensures a G-Bayes action exists; nonnegative integrands must vanish almost surely.
- C.2 Proof of Proposition 2.3: Proposition 2.3 follows because the nonnegative maximum term vanishes exactly when coarse loss is zero and every competing term is nonpositive.The proof uses that coarse-optimal actions have zero decision gap and nonnegative weight.
- C.3 Proof of Proposition 4.1: Proposition 4.1 equates temporary probing and probe-and-promote costs through (m−1)h = mc.Temporary probing costs mc + H, whereas probe-and-promote costs H + (m−1)h.
- C.4 Proof of Proposition 6.1: Proposition 6.1 begins from the best value in the learner class and defines the corresponding benchmark quantity W_s,n = V★.The supplied proof passage identifies the learner-class optimum as the starting point.
- C.5 Proof of Corollary 6.2: Corollary 6.2 derives its value comparison by substituting the policy decompositions and then subtracting priced resource costs.The identities express values through the coarse or refined Bayes value, regret, and productive term before forming net value.
- C.6 Proof of Theorem 3.2: Theorem 3.2 relates refined and coarse Bayes increments by conditioning on H, then integrates a δ-scaled lower bound over the event A.The proof establishes Equation (3.8), bounds min{a_H,b_H} by δ min(α,β) on A, and obtains Equation (3.10).
- C.7 Proof of Proposition 3.3: Proposition 3.3 constructs a curve within a summary level set whose terminal gap changes sign in opposite directions, while another output changes to first order.The level-set tangent space is ker DΦ_c(x_0), and the condition DY_h(x_0)v ≠ 0 supplies the first-order change.
- C.8 Proof of the replicate-denoised bridge: The replicate-denoised bridge bounds binary-policy regret using disagreement, |m_H| ≤ |m_H−f|, and Cauchy–Schwarz, yielding conditional policy improvement.Adding and subtracting V(π_f) proves the policy-improvement inequality, and conditioning on A ∈ H gives its conditional form.
C.9 Proof of Proposition 5.1 … C.15 Proof of Corollary 5.7
The proofs establish the dynamic recursion, continuation boundary, and exact scalar-sufficiency criterion for revelation decisions. They also characterize when refinement strictly improves value and when scalar summaries or threshold rules fail.
- C.9 Proof of Proposition 5.1: Bayes value under each information state gives the final equality in Proposition 5.1.
- C.10 Proof of Proposition 5.4: At the deepest admissible depth, value equals commitment value; earlier decisions maximize immediate commitment against paid refinement and conditional continuation value.Backward induction yields Equation (5.14).
- C.11 Proof of Proposition 5.2: Continuation is optimal exactly when conditional expected gain from refinement exceeds its cost.
- C.12 Proof of Theorem 5.3: Scalar sufficiency holds exactly when no positive-mass scalar fiber contains both strictly positive and strictly negative net continuation gains.Zero-gain ties do not affect the equality condition.
- C.13 Proof of Theorem 5.5: Refinement yields a strict value improvement whenever it retains positive tail mass beyond the relevant threshold.
- C.14 Proof of Corollary 5.6: Scalar sufficiency for one-step continuation does not by itself imply a threshold representation, which additionally requires monotonicity of scalar value relative to cost.
- C.14 Proof of Corollary 5.6: A nondegenerate strictly monotone continuation transform cannot be represented by any scalar function of the coarse margin summary.
- C.15 Proof of Corollary 5.7: Under continuity and strict monotonicity, scalar sufficiency corresponds to separating gains around a cost threshold; mixed positive and negative gains on margin fibers make the scalar representation strictly suboptimal.
C.16 Proof of Proposition 5.8 … D Notation and decision objects
The appendix proves the stated propositions by linking terminal gains to stopped-flip events and derives impossibility and finite-sample certification results. It also defines the decision objects and cost-adjusted continuation quantities used throughout the framework.
- C.16 Proof of Proposition 5.8: The proof of Proposition 5.8 uses Equation (5.16), m_h′−m_h=σ_hZ, with σ_h measurable under F_h and Z independent of F_h.This establishes the decomposition through conditional measurability and independence.
- C.17 Proof of Proposition B.1: The proof of Proposition B.1 shows that the two policies have identical terminal outcomes when F=0, while |Y|=|D|F when their action indicators differ.Thus terminal gains are confined to the event where the stopped-flip indicator is active.
- C.18 Proof of Proposition B.2: Proposition B.2 follows because S_Y=0 outside E={S=1,F=1}, with the r=0 case reducing to ΔV(S)=cρ.The value contribution therefore occurs only where both stopping and flipping events coincide.
- C.19 Proof of Theorem B.3: Theorem B.3 uses a layer-cake identity and Equation (B.3) to establish its bound, then constructs a law showing no finite lower bound follows from (c,ρ,r) alone.The construction fixes 0<r≤ρ≤1 while allowing stopped-flip severity M to grow arbitrarily, despite finite E[Y+].
- C.19 Proof of Theorem B.3: Because M is arbitrary while (ρ,r) remain fixed, bounded stop-flip risk alone cannot certify positive expected utility under unrestricted stopped-flip severities.This is the theorem’s impossibility limitation, not a claim that every distribution has negative net value.
- C.20 Proof of Corollary B.4: Corollary B.4 applies Hölder’s inequality and tail envelopes to derive finite-sample lower certificates using simultaneous confidence bounds for ρ, r, or H+.The proof first assumes E[|D|^q]≤M_q for q>1 before substituting the confidence bounds.
- D Notation and decision objects: The notation defines D as the binary terminal gap and Δrev as pure information-refinement value, while G_h→h′ and Γ_h→h′ denote deeper policy gain and its cost-adjusted continuation gain.The cost-adjusted quantity is Γ_h→h′=G_h→h′−c_h→h′.
- D Notation and decision objects: The decision objects also define σ_c(m) as the critical revealability changing the priced Stop/Continue decision and h as a target-independent, model-specific empirical proxy.The proxy is fixed without using target outcomes.
E Statistical estimands and simultaneous inference · F Experimental protocol details
The appendix defines exact-anchor inference, family-specific multiplicity control, and independent-panel evaluation protocols. It reports positive familywise-adjusted evidence for the scalar Mistral continuation rule while limiting claims to the declared sampling designs and estimands.
- E Statistical estimands and simultaneous inference: Inference uses the exact anchor as the resampling unit, averaging conditionally independent future-bank contributions within each anchor before estimating population means.This avoids treating bank-level realizations, readouts, or policy actions as independent inferential units.
- E Statistical estimands and simultaneous inference: Bootstrap uncertainty resamples exact anchors with replacement using 50,000 draws and targets the declared anchor-sampling design rather than universal guarantees across models, tasks, or deployments.The lower bound is the empirical α quantile of bootstrap means.
- E Statistical estimands and simultaneous inference: The Qwen finite-library conjunction requires positive margins and both one-sided t and exact-anchor bootstrap lower bounds for all 22 comparator components, with simultaneous 95% max-t bounds.The two 336-anchor panels are independent replications and are not pooled.
- E Statistical estimands and simultaneous inference: Multiplicity is controlled within each prespecified family, while the manuscript does not claim one paper-wide familywise-error guarantee across heterogeneous estimands.The 22-component current-information library and the two-architecture Mistral adaptive family are treated separately.
- E Statistical estimands and simultaneous inference: 2.37141363069 × 10^-6 and 5.41900958243 × 10^-6 are the scalar Mistral rule’s ordinary one-sided 95% t and bootstrap lower bounds, with Bonferroni-adjusted familywise bounds remaining positive.The scalar and two-coordinate architectures form a prespecified two-element family, with parameters fit only on the independent 336-anchor development panel.
- E Statistical estimands and simultaneous inference: Stop–flip risk uses bankwise one-sided Bernoulli–KL bounds for the probability of changing the deeper action, but these bounds do not control the utility severity of such changes.This limitation follows from Theorem B.3 and prevents bounded risk alone from certifying utility.
- F Experimental protocol details: The Qwen comparative panel has 432 disjoint exact anchors, while Mistral uses a separate 336-anchor development panel and an independent 480-anchor target panel with two future banks and a fixed 20-update compute contract.Mistral target evaluation fixes the H4 ambiguity threshold at 0.00236920914414, estimated from development data, and uses model revision c03fc1dabc3d31b96271626f15a76a6779fb4037.