Source-linked AI summary

Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings

Jiangang Han

arXiv:2607.13918v1math.STcs.AIcs.LG

TL;DR

Partially correlated verifier cascades lack a tight reliability theory despite shared blind spots in practical harnesses. This note models correlation with latent false-accept rates and shows that cascade evidence is concave, reliability can plateau or worsen, and decorrelation is the practical lever.

  • Problem

    Conditional independence is a limited assumption for verifier cascades because generators and verifiers can share blind spots.

  • Method

    The paper uses a de Finetti latent-variable model, moment identities, and beta-binomial or nonparametric maximum-likelihood estimation to characterize and measure cascade reliability.

  • Results

    Correlated cascades have concave log-odds, while upper-tail exponents determine whether gates eventually help, plateau, or harm, with crossover k† in the harmful regime.

  • Takeaways & Limitations

    When verifier correlation is high, adding gates buys little; changing model family, modality, or evidence source is the practical lever for reliability.

  • Takeaways & Limitations

    The validation is synthetic, and the scalar exchangeable latent is a mean-field idealization that omits heterogeneous verifier families and real-pair measurement.

Abstract

from arXiv · show

Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if $k$ verifier calls all accept it. Under conditionally independent gates, the recent Odds Law (arXiv:2606.15712) shows that posterior log-odds grow linearly in $k$, so failure decays exponentially, and states that "a tight theory of partially correlated verifier cascades remains open." This note gives a minimal such theory. Modeling the per-instance false-accept rate on the generator's own errors as a latent variable $α\sim G$ (de Finetti), the exact cascade posterior is $\ell_k = \ell_0 - \ln m_k$, with $m_k$ the $k$-th moment of $G$. Then: (i) $\ell_k$ is concave in $k$ for every non-degenerate $G$ -- the Odds Law is its tangent at the first gate and an upper bound; (ii) for Beta$(a,b)$ latents, failure decays polynomially, $1-r_k \asymp k^{-b}$, with correlation parameter $ρ_v = 1/(a+b+1)$; (iii) a blind-spot atom of mass $1-π$ at $α=1$ caps the evidence extractable from any number of gates at $-\ln(1-π)$ nats, so reliability saturates below 1; (iv) letting the true-accept rate also vary ($β\sim H$) yields a trichotomy -- gates eventually always help, plateau, or actively harm -- decided by the upper-tail exponents of $G$ and $H$, with closed-form crossover $k^\dagger$. The mechanism is survivorship: errors surviving gates are the high-$α$ ones. The theory is measurable: $R$ repeated verdicts per instance identify the first $R$ moments of $G$, so two verdicts identify $ρ_v$; beta-binomial likelihood and NPMLE recover the reliability curve and the ill-posed ceiling. In synthetic tests, independence-based extrapolation underestimates failure by 20x at $k=5$ and ~3000x at $k=10$; the correlated fit at $R=8$ tracks held-out depths. The practical lever is decorrelation -- changing model family, modality, or evidence source -- not adding gates.

1 Introduction

This note develops a minimal theory of correlated verifier cascades, showing that survivorship makes evidence gains concave, reliability polynomial or capped, and gate effects dependent on both false- and true-accept tails. It also provides moment-based identification and argues that decorrelation, rather than additional gates, is the practical lever.

  • Exact cascade posterior and concavification: Latent false-accept rates α ∼ G yield the exact posterior ℓ_k = ℓ_0 − ln m_k, with concave depth-scaling because surviving errors are enriched for high α.The Odds Law is the degenerate case, matches the true curve only at the first gate, and upper-bounds it thereafter.
  • Polynomial reliability and a one-parameter family: For G = Beta(a, b), failure decays polynomially as 1 − r_k ∼ κ k^−b, with within-instance verdict correlation ρ_v = 1/(a + b + 1).Correlation therefore replaces the independent cascade’s exponential convergence with polynomial convergence.
  • Blind-spot ceiling: A blind-spot atom of mass 1 − π at α = 1 caps cascade evidence at −ln(1 − π) nats and keeps limiting reliability below 1.The limiting reliability is r_∞ = p_0/(p_0 + (1 − p_0)(1 − π)).
  • Two-sided trichotomy with closed-form crossover: When β also varies, gates eventually help, plateau, or harm according to whether G’s upper-tail exponent exceeds, equals, or falls below H’s.In the harmful regime, reliability peaks at the closed-form crossover k† and then decays to zero even when mean Λ̄ > 1.
  • Identification and inversion of ρ_v: Repeated verdict logs identify G’s first R moments, so R = 2 identifies ρ_v; beta-binomial likelihood and nonparametric MLE recover reliability curves and latent atoms.The ceiling remains intrinsically ill-posed because it depends on an upper-tail functional.
  • Practical implication: Once ρ_v is large, adding gates buys almost nothing; decorrelating verifier and generator through different models, modalities, or external checks raises b, shrinks blind spots, and lifts the ceiling.The note’s practical correction is that correlation taxes extrapolation rather than the first gate.

2 Setup

The setup models all-accept verification cascades through survivor reliability: returned answers are those accepted by every gate. It treats false-accept behavior as an instance-level latent propensity, capturing shared blind spots across exchangeable verdicts.

  • Gates and survivor reliability: All-accept gating returns a candidate only when all k verification gates accept, with survivor reliability measuring the correctness of returned answers.In generate–verify–retry, rejected correct answers reduce throughput rather than precision; Section 3 assumes β = 1 before removing that idealization in Section 4.
  • Latent false-accept rate and the generator–verifier bridge: The false-accept rate α is modeled as an instance-level latent variable distributed across the generator’s erroneous outputs.This makes G’s upper tail the population of generator errors that the verifier systematically fails to catch.
  • Latent false-accept rate and the generator–verifier bridge: G’s upper tail near α = 1 represents generator–verifier blind-spot alignment, which can persist even when verification performs well on random errors.The supplied passage states that this tail thickens as generators strengthen because their errors become more systematic.
  • de Finetti idealization: Conditioned on α, gate verdicts are i.i.d. Bernoulli(α) for errors and Bernoulli(β) for correct answers under the de Finetti idealization.The assumption is exact for exchangeable verdicts and serves as a mean-field model of shared gate structure.
  • de Finetti idealization: The same latent model covers repeated stochastic samples from one verifier and verifier families that share blind spots.In the first reading α is a verifier’s per-instance acceptance propensity; in the second it is a family-level propensity.

3 One-sided theory: concavity, polynomial decay, ceiling

Correlated verifier cascades have an exact posterior governed by the latent false-accept moments, making log-odds concave rather than linear. Upper-tail heterogeneity causes polynomial reliability gains, while blind-spot mass imposes a finite evidence ceiling below perfect reliability.

  • Exact cascade posterior: The exact cascade posterior is ℓ_k = ℓ_0 − ln m_k, where m_k = E[α^k], and reliability r_k is nondecreasing in k.Conditionally surviving errors contribute m_k to the all-accept probability, so Bayes’ rule gives the posterior and monotonicity.
  • Concavity: For non-degenerate G, ℓ_k is strictly concave: per-gate evidence decreases, and the first-gate Odds Law line upper-bounds later log-odds.Equality with the tangent occurs only at k ∈ {0, 1}; the independent case G = δ_ᾱ is the linear degenerate case.
  • Polynomial decay: Regularly varying upper tails degrade independent exponential failure decay to polynomial k^-b; for Beta(a,b), the exponent is b and ρ_v = 1/(a+b+1).The surviving-error population is tilted toward high-α instances, so late conjunctive gates face increasingly difficult survivors.
  • Blind-spot ceiling: A blind-spot atom of mass 1 − π caps total cascade evidence at −ln(1 − π) nats and leaves limiting reliability r_∞ < 1.The ceiling depends on blind-spot mass alone, not on the number of gates or the independent-case evidence factor.
  • Empirical implication: At k = 5, independence underestimates failure by 20×; at k = 10, it underestimates failure by 3000× in the moderate correlated regime tested.The comparison uses p_0 = 0.5, ᾱ = 0.3, and ρ_v = 0.3, with curves agreeing at k = 1 by construction.

4 Two-sided theory: when gates help, plateau, or harm

When both false accepts and false rejects vary across instances, cascade behavior is governed by the upper-tail exponents of their latent rates: gates eventually help, plateau, or harm. In the harmful regime, reliability can peak at a finite depth before declining toward zero, even when the first gate helps.

  • Trichotomy: Under regularly varying upper tails, exactly one regime obtains: gates eventually help, plateau below 1, or eventually harm.The trichotomy compares exponents bα and bβ for latent false-accept and true-accept rates.
  • Trichotomy: The tail comparison decides the eventual direction: bα > bβ yields help, bα = bβ yields a plateau, and bα < bβ yields harm.Errors with near-universal verifier acceptance and correct answers with near-universal acceptance compete through survivorship.
  • Correlation changes the criterion: A mean likelihood ratio ¯Λ > 1 guarantees only that the first gate helps; correlated cascades can still have rk →0 when tail exponents favor harm.The eventual direction is logically independent of ¯Λ, and Table 2 exhibits ¯Λ = 2.2 with rk →0.
  • Crossover: In the harmful regime with δ1 > 0, reliability is unimodal and peaks at the discrete optimum kopt = ⌈k†⌉.The crossover k† is driven by selection effects alone, even with zero gate cost.
  • Numerical example: 78.7% is the peak reliability at k = 5 in the numerical example, declining to 62% by k = 20 and heading to zero.Here ¯Λ = 2.2 and δ1 = 0.80 > 0, yet bα = 1.63 < bβ = 4 gives k† ≈4.3; independence extrapolation instead reports five nines and rising.
  • Blind-spot ceiling: With blind-spot atoms on both sides, the limiting reliability equals the ratio of surviving blind-spot masses: r∞= p0(1−πβ) / [p0(1−πβ)+(1−p0)(1−πα)].The two-sided ceiling generalizes the one-sided blind-spot limit.

5 Measuring ρv: an inversion protocol

The section presents an inversion protocol that estimates verifier-cascade correlation from repeated verdicts on the generator’s own errors, while exposing the resolution limits of blind-spot ceiling estimates. Beta-binomial and nonparametric fits support curve extrapolation and guide when to pursue decorrelation instead of additional gates.

  • Inversion protocol: Repeated verdicts on the generator’s own erroneous outputs identify the first R moments of G, with R = 2 sufficient for consistent estimation of ρv.The protocol samples each erroneous instance R times and uses binomial deconvolution or low-order functionals.
  • Inversion protocol: Beta-binomial maximum likelihood predicts the full reliability curve and ceiling, while nonparametric MLE can reveal α = 1 atoms and multimodality without assuming Beta.These are higher-resolution estimators than direct low-order moment recovery.
  • Ill-posedness of the ceiling: R verdicts cannot distinguish α = 1 from α = 1 − ϵ below boundary resolution ϵ ∼1/R, creating an R-dependent identifiability floor for ceiling estimates.The section notes that ceilings 0.91 and 1.00 can produce nearly identical small-R data, so regularization is required.
  • Falsification loop: Out-of-sample extrapolation compares low-order fits of G against held-out gate depths, separating exponential predictions with ρv = 0 from polynomial predictions with ρv > 0.The falsification loop tests the entire predicted reliability curve rather than only fitted depths.
  • Decorrelation and exchangeability: Decorrelation thins G’s effective upper tail, raises exponent b, shrinks blind-spot mass 1 −π, and lifts the ceiling −ln(1 −π).Changing model family, modality, or evidence source makes gates heterogeneous, but the scalar model interprets this as moving G in the reliability-improving direction.

6 Synthetic-recovery validation

Synthetic experiments recover verifier-correlation parameters and reliability curves from repeated verdicts, while showing that blind-spot ceilings require substantially more data to identify. Correlated extrapolation matches held-out truth, unlike independence-based extrapolation.

  • Experimental validation: Across N = 4000 fixed-seed instances, synthetic recovery tests evaluate known verifier-correlation ground truth before applying the pipeline to real logs.The experiments use reproducible code and span ρv values from 0.05 to 0.50 with R = 10.
  • Full-spectrum recovery: The moment estimator recovers ρv as 0.050/0.291/0.502 for true values 0.05/0.30/0.50, with M2 agreeing.This demonstrates full-spectrum recovery under the stated synthetic setting.
  • Two verdicts suffice: For true ρv = 0.30, R = 2 estimates ρv at 0.274, while R ≥3 is essentially exact.The result operationalizes the claim that two verdicts suffice to estimate verifier correlation.
  • Ceiling ill-posedness: At R = 5, worlds with 10% atom mass at α = 1.00 versus α = 0.97 have nearly indistinguishable histograms; at R = 50, NPMLE assigns masses 0.100 versus 0.000.Their true ceilings are r∞= 0.91 versus 1.00, illustrating that ceiling identification is more data-intensive than estimating ρv.
  • Falsification and real-log validation: At k = 5, correlated extrapolation predicts 0.954 versus truth 0.953, while independence extrapolation from the same first-gate data predicts 0.998.The beta-binomial fit uses accept counts through R = 8 and estimates ρv at 0.30; the next step is applying the estimators to real logs.

7 Related work

This work extends reliability-algebra and correlated-voting research by replacing conditional independence with correlated checking, formalizing blind-spot ceilings, and providing measurable finite-k reliability results. It distinguishes this mechanism from correlated generators, misspecified judges, and prior polynomial-scaling findings.

  • Reliability algebra for harnesses: The paper modifies conditional independence in the Odds Law/Maestro Order framework and addresses its stated open problem on partially correlated verification cascades.Its corollary recovers prior independent-gate results in the appropriate limit.
  • Correlated generators and ceilings: Repeated same-instance verdicts identify verifier correlation, unlike cross-model pairwise error statistics, and expose the ill-posed tail quantity governing the reliability ceiling.This observable distinction separates Proposition 5.1 from the non-identifiability result for model-pool selection.
  • Self-verification blind spots: The contribution turns self-verification blind spots into explicit finite-k posteriors, concavity, polynomial decay, ceiling structure, closed forms, and a measurement protocol.Prior empirical and information-theoretic work documents missed self-errors or shared latent causes but lacks this measurement protocol.
  • Verifier/judge scaling and the exponential–polynomial dichotomy: The correlation-driven finite-optimum phenomenon is distinct from misspecified reward models because verifier correlation ρ_v is directly measurable in data.This makes the two mechanisms empirically separable despite similar exponential–polynomial behavior.
  • Verifier/judge scaling and the exponential–polynomial dichotomy: Prior test-time reliability studies likewise report exponential, polynomial, or crossover behavior depending on the underlying scaling or generative mechanism.Examples include knockout tournaments, best-of-k with misspecified rewards, and attack-success curves.

8 Limitations and outlook

The theory is deliberately minimal: it uses a scalar exchangeable latent and all-accept semantics, while several extensions and real-world validations remain open.

  • Model limitations: The scalar exchangeable latent is a mean-field idealization; heterogeneous verifier families require a vector latent.Gate-specific systematic differences are not represented by the current model.
  • Distributional assumptions: Beta is a convenience for closed forms; the asymptotics require only regularly-varying tails, and M3 removes the parametric assumption in estimation.This broadens the asymptotic and estimation frameworks beyond the Beta specification.
  • Validation outlook: Validation currently demonstrates synthetic recovery, leaving measurement on real generator–verifier pairs as the quantitative bridge.The supplied passage identifies real-pair measurement as an outstanding validation step.

A Deferred proofs · A.1 Theorem 3.5 (polynomial decay)

The deferred proof shows that the verifier-survival moment is dominated near s = O(1/k), yielding polynomial decay. For G = Beta(a, b), this gives mk ∼ Γ(a+b)/Γ(a) k^-b and 1 − rk ∼ (1−p0)/p0 · mk.

  • A.1 Theorem 3.5 (polynomial decay): The moment integral is dominated by s = O(1/k).This identifies the boundary layer governing the large-k asymptotics.
  • A.1 Theorem 3.5 (polynomial decay): Near that scale, g(1 − s) ∼ c s^(b−1).The density’s behavior near α = 1 determines the decay exponent.
  • A.1 Theorem 3.5 (polynomial decay): The factor (1 − s)^k satisfies (1 − s)^k ∼ e^−ks on the dominant scale.This approximation converts the moment asymptotics into a Gamma integral.
  • A.1 Theorem 3.5 (polynomial decay): Watson’s lemma yields mk ∼ c Γ(b) k^−b.The resulting power law follows from integrating e^−ks s^(b−1).
  • A.1 Theorem 3.5 (polynomial decay): For G = Beta(a, b), mk = B(a+k,b) B(a,b) = Γ(a+b) Γ(a) k^−b by Stirling.The beta moment has the same polynomial exponent b.

A.2 Corollary 3.8 (cost-optimal gates) · A.3 Theorem 4.1 (trichotomy) · A.4 Proposition 4.3 (crossover)

The merged results characterize cost-optimal gate depth under correlated versus independent errors, and establish a tail-exponent trichotomy with at most one reliability crossover. For Beta latents, the crossover has a closed-form linear equation and can produce a unimodal reliability curve.

  • A.2 Corollary 3.8 (cost-optimal gates): Correlated-gate marginal benefit decays as dk ≈ Uκb k^−(b+1), yielding k∗ = (Uκb/c)^(1/(b+1)) when equated to cost c.This is the cost-optimal depth under the correlated asymptotic marginal-gain scaling.
  • A.2 Corollary 3.8 (cost-optimal gates): Under independence, failure behaves as 1−rk ≈ (1−p0)ᾱ^k, so the optimal depth scales as Θ(ln(U/c) / ln(1/ᾱ)).The independent-gate calculation comes from equating U times the derivative of the exponentially decaying failure curve to c.
  • A.3 Theorem 4.1 (trichotomy): The trichotomy is determined by the sign of the coefficient multiplying ln k in the asymptotic log-odds expression, namely bα−bβ.The supplied theorem passage identifies the three regimes through this sign.
  • A.4 Proposition 4.3 (crossover): For Beta latents, the xj-tilted distribution of Beta(a,b) is Beta(a+j,b), with tilted mean (a+j)/(a+b+j).These tilted distributions provide the quantities used to compare gate behavior across depths.
  • A.4 Proposition 4.3 (crossover): The crossover depth k† is obtained by equating the two Beta tilted means for the α and β latents.The equality is written as (aβ+j)/(aβ+bβ+j) = (aα+j)/(aα+bα+j).
  • A.4 Proposition 4.3 (crossover): Cross-multiplication cancels quadratic terms and gives the linear equation j(bα−bβ) = aαbβ−aβbα, implying at most one sign change.The linearity of the equation is the reason the crossover is unique when it exists.
  • A.4 Proposition 4.3 (crossover): When bα < bβ and δ1 > 0, increments are positive before k† and negative afterward, making ℓk and rk unimodal with integer maximizer ⌈k†⌉.The result specifies the direction of the crossover and the maximizing integer depth.
Loading 2607.13918v1…