Source-linked AI summary
A Closed-Form Estimator and Diagnostic Battery for Anchor-Judge Error Correlation, Under a Single-Common-Factor Model
Veerendra Kumar Sunkavalli
TL;DR
Existing anchored LLM-judge decompositions assume external anchors are uncontaminated, leaving anchor–judge error correlation unmeasured. This paper derives a closed-form estimator under a single-common-factor model and surrounds it with diagnostics, while identifying ordinal-score limits and reporting that no tested real panel has yet passed the model pre-test.
Problem
Existing anchored decompositions assume anchor error is uncorrelated with judges’ shared error, so contamination can silently invalidate the decomposition.
Method
The paper derives a closed-form estimator under a single-common-factor model and gates it with diagnostics, bootstrap confidence intervals, weak-identification screening, and ordinal-case extensions.
Results
Under the model, ≥2 judges and ≥2 anchors identify the target variances and contamination correlations, while asymmetric violations bias ρk and uniform second factors leave it unbiased.
Takeaways & Limitations
The estimator is a transparent tool for qualifying judge/anchor systems, with exact boundary diagnosis and diagnostics for violations that matter.
Takeaways & Limitations
No tested real panel has passed the model pre-test, so the estimator is validated in simulation and semi-synthetic oracle-calibrated stress tests rather than on a qualifying real panel.
Abstract
from arXiv · showhide
When an external reference set (an anchor) is used to decompose an LLM-judge panel's error into a quality signal and a shared common-mode error, standard practice assumes the anchor is uncontaminated: its error uncorrelated with the judges' shared error. We study when that assumption can be dropped and replaced by an estimate. Under a single-common-factor model, >=2 judges and >=2 anchors point-identify the quality variance, the common-mode variance, and each anchor's contamination correlation rho_k in closed form, with an exact per-anchor-pair failure boundary; a designated clean-anchor estimator, by contrast, reports a contaminated companion anchor as fully clean once its trusted anchor is itself contaminated. Because the single-common-factor assumption is itself untestable, the estimator ships gated behind a calibrated diagnostic battery (judge-covariance dispersion; over-identification; a family-block test from judge metadata, with a family-blocked estimator that removes family-level shared-residual bias exactly), bootstrap confidence intervals with measured coverage, and a weak-identification screen. A proposition maps which violations bias rho_k, in which direction, and which evade detection. For ordinal scores we show an identification hierarchy: with all variables ordinal, rho_k is not identified at any number of anchors; with ordinal judges and >=3 continuous anchors it is, and we give an estimator for that case. On real data the validation is asymmetric, and we say so plainly: the diagnostics are validated in the rejecting direction (both real panels we test are correctly rejected by the model-adequacy pre-test), while the estimator is validated in simulation and stress-tested semi-synthetically under oracle calibration; no real panel has yet passed the pre-test, and the pre-test exists precisely to say so. All results replay offline from shipped, checksummed artifacts.
1 Introduction
The paper replaces the clean-anchor assumption with a closed-form contamination estimator under a single-common-factor model, while gating use behind diagnostics because that model is untestable. It characterizes exact failure conditions, relevant violations, ordinal identification limits, and the method’s current validation boundary.
- Estimator and identification: Under the single-common-factor model, ≥2 judges and ≥2 anchors point-identify quality variance, common-mode variance, and each anchor’s contamination correlation in closed form.The estimator requires neither a clean anchor nor an anchor-difference assumption.
- Estimator and identification: A measure-zero condition βk = σ2_c is the exact non-identification boundary, with a practically relevant weak-identification neighbourhood.The method includes a studentized-denominator screen that flags proximity to the boundary from the data alone.
- Violations and diagnostics: Only asymmetric second-factor loadings bias ρk; uniformly loaded factors leave it unbiased by being absorbed into the quality variance.The uniform case is observationally equivalent to inflating the quality variance and is not detected by second-moment diagnostics.
- Violations and diagnostics: A calibrated three-test battery targets judge-loading dispersion, anchor over-identification, and shared family-level judge residuals, with a family-blocked estimator removing the latter bias exactly across ≥2 families.Null-calibrated thresholds use a 5% false-positive rate under the correctly specified model; power increases with signal strength and sample size, but mild violations can remain invisible.
- Ordinal scores: With all variables ordinal, ρk is not identified at any number of anchors; with ordinal judges and ≥3 continuous anchors, it is identified and estimable.This establishes an identification hierarchy rather than a universal ordinal-score solution.
- Positioning and scope: The closed form is statistically indistinguishable from full-information ML on identical data, but its contribution is transparency and exact boundary diagnosis rather than efficiency.The paper also reports finite-sample behaviour, weak identification near the boundary, and upward bias of the mixed-scale estimator near ρ = 0 at small N.
2 Related work and positioning
The paper situates its model within established MTMM and reference-method work, distinguishing its contribution as a judge/anchor-specific closed form, failure characterization, violation analysis, and diagnostic battery. It also contrasts with LLM aggregation and gold-standard-free error-rate literature that does not parameterize anchor–judge error correlation.
- MTMM positioning: The underlying identifiability is established in MTMM common-method-factor models with correlated errors; the paper adds a closed form specialized to the unit-loading judge parameterization.It also adds the exact non-identification boundary, a proposition on biasing violations, and diagnostics for the judge/anchor setting.
- Adjacent literature: Prior LLM-judge aggregation recovers quality without ground truth but treats human anchors as trusted labels and has no parameter for anchor–judge error correlation.The paper cites overlapping omitted-confounder analysis rather than claiming that result as novel.
3 Model and identification
Under a single-common-factor model, a judge panel with at least two judges and two anchors identifies quality variance, common-mode variance, and anchor contamination without assuming a clean anchor. The estimator has an exact failure boundary, while shared-residual violations can bias contamination estimates and require diagnostics or family blocking.
- Model: The model uses latent quality and common-mode factors, with judges sharing unit loadings and anchors carrying anchor-specific common-mode loadings.Judge residuals are assumed mutually uncorrelated; this is the empirically fragile part of the model.
- Diagnostics and safeguards: Shared judge residuals can bias contamination estimates, and the closed form is therefore gated by adequacy tests, weak-identification screening, and family blocking where applicable.The family-blocked estimator removes family-level shared-residual bias exactly when the panel spans at least two families; out-of-range estimates are reported rather than clipped.
- Identification: At least two judges and two anchors point-identify the model parameters and each anchor's contamination correlation when the identifying denominator is nonzero.No clean-anchor or anchor-difference assumption is required; fewer than two anchors leaves the parameters unidentified.
- Identification: Identification fails exactly when an anchor's contamination covariance equals the common-mode variance, with a positive-measure neighbourhood of weak identification.Near the boundary, the estimator's variance grows even when point estimates remain approximately unbiased.
- Violations: A uniformly loaded second factor is observationally absorbed into quality variance and leaves contamination correlations unbiased, whereas asymmetric loadings bias them.The quality variance itself remains biased under the uniformly loaded case, affecting signal-to-noise assessments.
4 Simulation study
Simulations show that the closed form recovers contamination correlations accurately under the stated model, while finite samples increase variance and near-boundary settings create weak identification. The simulation study also characterizes the estimator's operating range and diagnostic context.
- Recovery: Simulation recovery was within 0.02 for four contaminated-anchor configurations at N=40,000, including statistically identical anchors.The results verify moment inversion under simulation, not model validity for real judges.
- Finite-sample behavior: Recovery of ρ2=0.7 improved from 0.718 ± 0.067 at N=100 to 0.701 ± 0.007 at N=40,000.No out-of-range estimates arose in this finite-sample experiment.
- Weak identification: At ρ1=0.99 near the identification boundary, the standard deviation grew fifteen-fold despite approximately unbiased point estimates through ρ1=0.98.The weak-identification screen is designed to flag this regime.
5 Misspecification diagnostic battery
The diagnostic battery targets asymmetric judge or anchor loadings and family-level shared residuals, while explicitly flagging harmful panel-wide residuals that remain undetectable. Simulation shows calibrated tests and a family-blocked estimator can preserve or restore accurate contamination estimates under supported conditions.
- Which violations matter: Asymmetric anchor loadings bias ˆρ2, but the direction and magnitude are non-monotone rather than determined by asymmetry alone.With g = 0.6, bias is 0.001 at h = g; estimates rise 0.761 →0.780, dip to 0.769, then fall to 0.565 against true 0.7.
- Diagnostic tests: Test A measures dispersion in off-diagonal judge covariances and requires at least three judges because two provide only one off-diagonal.Non-uniform judge loadings violate the equal-covariance implication of the single-common-factor model.
- Diagnostic tests: Test B measures disagreement among anchor-pair estimates of σ2_t and detects non-uniform anchor loadings where Test A is blind.It requires at least three anchors for over-identification.
- Operating characteristics: Null thresholds deliver a 5% false-positive rate by construction, but mild violations can remain invisible and heavy tails inflate false positives.At spread 0.2, Test A power is 0.04 at N=500 and 0.11 at N=2000; heavy tails raise it to 0.083.
- Blind spots: A shared judge residual biases ˆρk toward a false clean estimate in the typical regime, while panel-wide residuals remain inseparable from common mode and make the estimator unusable.Single-family panels should treat small ˆρk as unverified, regardless of diagnostic results.
- Family-blocked estimator: The family-blocked estimator uses cross-family judge covariances to remove family-level residual bias exactly, while within-minus-cross differences estimate family-residual variance.The construction needs at least two families, with one family containing at least two judges; singleton residuals enter only through cross terms.
- Simulation: In six judges across three families, the naive estimate shifts 0.697 →0.650 as residual s.d. reaches 0.7, while the blocked estimate stays at 0.696–0.697.Test C flags 4/100 null datasets and 100/100 datasets at every tested strength, recovering σ2_f to the second decimal.
- Score types: Under heteroscedastic noise, recovery remains essentially unaffected at ˆρ2 = 0.703 ± 0.008, whereas ordinal scores change identification rather than merely bias.The ordinal identification hierarchy is treated separately: all-ordinal variables do not identify ρk.
6 Inference, and a comparison with maximum likelihood
Bootstrap intervals achieve measured nominal coverage, and a studentized denominator screen identifies boundary weakness. The closed form matches full-information ML statistically while adding transparent boundary diagnostics, but robustness depends on the model and panel configuration.
- Confidence intervals: 95% bootstrap intervals for ρ2 achieve 0.953 coverage at N=2000 and N=500, with mean widths 0.118 and 0.244 respectively.Near the boundary, coverage remains conservative, reaching 0.992 at ρ1=0.97, although only 118 of 150 replicates return an estimate.
- Weak identification: The weak-identification screen flags boundary proximity when T < 4, firing on 46% of datasets at ρ1=0.90 and 100% at ρ1=0.97.At mid parameters it fires on 0% of datasets for both N=500 and N=2000; flagged estimates should be reported as intervals only.
- Comparison with maximum likelihood: Closed-form and full-information ML estimates are statistically indistinguishable, including over-identified and shared-residual misspecified settings.At mid parameters both give ˆρ2 = 0.704 ± 0.016; with shared residual strength 0.5, estimates are 0.585 versus 0.597 against true 0.7.
- Comparison with maximum likelihood: The closed form is faster but its main advantage is transparency: the exact failure boundary and weak-identification screen follow directly from the formula.The paper reports roughly 0.2ms versus 4ms per fit, while emphasizing that both methods are computationally trivial at this scale.
- Non-Gaussian robustness: Under skewed and heavy-tailed latents, CI coverage is 0.975 and 0.908, while Gaussian-calibrated tests show mild false-positive inflation.The paper recommends recalibrating null distributions on matched marginals for real panels.
- Practical operating point: At N ≈500, coverage is 0.953 with mean width 0.244, while Test A and Test B power depends strongly on loading spread.Power is 0.23/0.91 for Test A and 0.50/0.99 for Test B at loading spreads 0.4/0.6.
- Configuration guide: Two judges and two anchors permit estimation with only the weak-identification screen; larger panels enable Test A, Test C, or the full battery depending on the misspecification.Panel-wide residuals and uniform factors remain outside what the full battery can diagnose, although the uniform factor is harmless for ρk.
7 Ordinal scores: an identification hierarchy
Ordinal observation changes identification: contamination is not identified when all variables are ordinal, but becomes locally identified with ordinal judges and at least three continuous anchors. The paper supplies an estimator for that mixed-scale case and tests it on semi-synthetic real-rater and real-LLM settings.
- Identification hierarchy: With all variables ordinal, ρk is not identified at any number of anchors.The paper attributes this to latent scale information being absorbed by ordinal thresholds.
- Identification hierarchy: With ordinal judges and at least three continuous anchors, ρ is locally identified with one over-identifying restriction.The estimator combines polychoric judge correlations, polyserial judge–anchor correlations, and the raw anchor covariance block.
- Estimator performance: Treating ordinal judge codes as continuous was erratic across discretizations, while the mixed-scale estimator was essentially unchanged across 3-, 5-, and 7-level discretization.Ordinal judges also incurred roughly an order of magnitude more variance than continuous judges at the same sample size.
- Validation: Under real rater textures, mean recovery was monotone for injected ρ=(0.0, 0.4, 0.7), but strict ordering held in only 8/24 runs at N=431 and 14/24 at N=2000.High contamination was recovered well, while near-zero contamination was less precise.
- Validation: The real HANNA panel triggered Test A at 0.589 versus a 0.247 null threshold, so the diagnostic pre-test rejected the panel before estimator use.The real-LLM experiment was likewise a robustness observation under oracle calibration rather than clean-model validation.
8 What the study does and does not establish
The study establishes estimator correctness, identification boundaries, violation behavior, and diagnostic calibration, but does not establish that real judge–anchor systems satisfy the single-common-factor model. Real-panel evidence currently supports rejection rather than estimator qualification.
- Established: The paper establishes that its closed form inverts the model’s moments and characterizes the exact boundary, weak-identification neighborhood, and violation biases.It also establishes that identification is supplied by the judge panel and documents the battery’s calibrated detection profile.
- Not established: The paper does not establish that real LLM-judge and human-anchor errors follow the single-common-factor assumption.Passing the battery is necessary but not sufficient.
- Not established: A shared bias with the exact observationally equivalent shape would be miscredited, while the battery detects only its asymmetric departures.This limits what the diagnostics can rule out.
9 Conclusion
The paper presents closed-form anchor-contamination estimation under a single-common-factor model, together with its failure boundary, violation analysis, and diagnostic battery. Whether the assumption holds for any given judge–anchor system remains an empirical question.
- Conclusion: Under a single-common-factor model, anchor contamination can be estimated in closed form from a judge panel and two anchors.The paper also provides an exact identifiability boundary, a violation-bias proposition, and diagnostics for relevant departures.
- Conclusion: The contribution does not resolve whether the model assumption holds for a given judge–anchor system.It instead provides tools with which practitioners can probe that assumption.
Reproducibility and data statement
All results replay offline through deterministic code, a pinned environment, and two frozen, checksummed inputs.
- Reproducibility: One command regenerates every number within numeric tolerance from a pinned environment.The replay uses deterministic, seeded code and two frozen inputs.
- Data statement: The frozen inputs are the public HANNA benchmark and a SHA-256-checksummed six-judge LLM verdict cache.The cache contains 6,000 calls and 5,996 parsed verdicts and is shipped with the artifact.
A Proofs
The section derives a closed-form relation by expanding observable moments and notes an analytically verified reparameterization under uniform second factors that leaves ρ_k unchanged.
- Theorem 1: Expanding the moment equation cancels the x^2 terms and yields a linear equation in x.The passage gives the resulting coefficient relation but is truncated before the full structural expression is shown.
- Theorem 1: The derivation substitutes structural moments into the denominator of the resulting expression.
- Proposition 1: The uniform-second-factor case is analyzed through transformed observable moments for judges and anchors.The displayed moment formulas are truncated in the supplied passage.
- Proposition 1: Solving for an equivalent single-common-factor parameterization and setting g = h preserves ρ_k.The result is marked analytic and SymPy-verified.