Source-linked AI summary
Aggregate Disambiguation Systems
José María Lago, Albert Castellana, Edgars Nemše
TL;DR
Natural-language evaluators can disagree on identical declared information, so the paper asks when finite panels reproduce a declared evaluator population rather than semantic truth. It formalizes regime-specific error and estimates reproducible-resolution mass with a nested finite-sample certificate that separates evaluator and generator sampling. The certificate allows shared-row dependence, while experiments verify validity and expose limited power; the pilot remains constrained by its small, manually stratified dataset and lack of independent semantic adjudication.
Problem
Natural-language tasks can produce different verdicts from protocol-following evaluators, creating a need to assess finite-panel reproduction relative to a declared evaluator reference rather than semantic truth.
Method
The paper separates finite-census, probabilistic-population, and growing-census regimes, then combines exact evaluator intervals with one-sided generator inversion to lower-bound reproducible-resolution mass under shared-row dependence.
Results
The experiments report no violations in 8,000 simulations, while the baseline design is valid but usually underpowered to certify true coverage 0.70 against target 0.60; the LLM pilot found high coordination on both clear and underdetermined constructed cases.
Takeaways & Limitations
Successful certification establishes reproducibility rather than truth, so deployment should report panel-reproduction guarantees alongside independent semantic metrics when adjudicated truth is available.
Takeaways & Limitations
The method assumes a predeclared finite grid, i.i.d. evaluator rows, and correct binomial marginals, while the pilot has 50 manually stratified units and no independent semantic adjudication.
Abstract
from arXiv · showhide
Natural-language tasks can elicit different verdicts from protocol-following evaluators that receive the same declared information. We study aggregate disambiguation systems (ADSs). Given a task and a candidate solution, each evaluator casts a binary vote on whether the solution should be accepted, and the system aggregates the votes of a finite panel. The target is protocol reproducibility relative to an explicitly declared evaluator reference, not semantic truth. We separate fixed finite censuses, probabilistic evaluator populations, and growing-census limits, since their endpoint laws and guarantees are not interchangeable. In the population setting, we use finite samples to estimate how often a finite panel reaches the same decision as the declared evaluator population. We provide a lower confidence bound on the fraction of candidate solutions for which the disagreement probability is at most a chosen tolerance. The calculation accounts separately for sampling candidate solutions and sampling evaluators. The construction permits arbitrary dependence among columns induced by shared evaluator rows and uses exact binomial intervals at the evaluator layer and an exact one-sided binomial inversion at the generator layer. Simulations check the implementation against known population coverages and expose power limitations.
1 Introduction
The paper defines aggregate disambiguation systems around reproducibility relative to a declared evaluator reference, rather than semantic truth. It develops regime-specific panel guarantees, a finite-sample certificate, and sensitivity analyses that preserve distinctions among evaluator populations and sampling laws.
- Motivation: Natural-language tasks can yield different evaluator verdicts despite identical declared information, making finite-panel reproduction of a declared population decision the operational target.
- Definition: An aggregate disambiguation system generates a candidate, evaluates it across episodes, and aggregates binary verdicts while keeping reproducibility separate from semantic correctness.
- Inference regimes: Finite frozen censuses, growing catalogues, and probabilistic evaluator populations induce different sampling laws and therefore cannot share panel guarantees without qualification.
- Contributions: The paper studies exact finite-panel error, certification of reproducible-resolution mass from finite generator–evaluator experiments, and adversarial sensitivity indexed by panel size and task clarity.
- Contributions: The certificate combines exact finite-census or superpopulation errors with generator-mass inference, allowing shared-row dependence through familywise or mass-controlled guarantees.
- Evaluation: Synthetic experiments validate implementation and examine power, while the real-data pilot is explicitly limited and is not treated as a separate methodological contribution.
2 Methodology
The methodology defines panel decisions under three reference regimes and estimates workload-level reproducibility from independent generator and evaluator samples. Its certificate combines evaluator uncertainty with generator sampling while allowing dependence across resolution columns and distinguishing validity from power.
- System construction: The system computes one global binary vote per evaluator before aggregating evaluator votes; aggregating component majorities first generally defines a different mechanism.
- Population reference: The declared evaluator law determines the population decision and its operational clarity margin, which measures distance from the decision boundary rather than semantic truth.
- Three inference regimes: The three regimes use hypergeometric tails for finite frozen censuses, binomial tails for fresh population panels, and approximation envelopes for transferring growing censuses to ideal populations.
- Workload-level error: Population coverage is the generator-law fraction of resolutions whose fresh-panel disagreement probability is at most δ; resolvability requires coverage at least 1 − β.
- Sampling design: Independent generator configurations and evaluator rows form a binary matrix whose columns may be arbitrarily dependent through shared row-level randomness, while rows remain i.i.d. with declared binomial marginals.
- Certificate: The mass-controlled certificate uses exact evaluator-layer binomial intervals and exact generator-layer inversion, selecting panel sizes from a predeclared grid with confidence at least 1 − η_E − η_G.
- Alternatives and adversaries: Hoeffding provides a valid but generally more conservative alternative, and adversarial analyses distinguish fixed-share from targeted contamination through exact binomial tails.
3 Synthetic Results
Synthetic experiments validate the finite-sample certificate and map how panel size, clarity, dependence, and adversarial participation affect reproducibility. Exact finite-panel calculations expose finite-population effects, while phase diagrams show that adversarial tolerance is conditional rather than universal.
- Exact finite-panel behavior: Exact hypergeometric disagreement for a frozen census falls quickly when census clarity is far from the threshold and slowly when it is close.The census contains 1,500 votes under an inclusive majority rule; binomial approximations are accurate only for small sampling fractions and are not exact at the full census.
- Exact finite-panel behavior: Panel size alone is not a security or quality parameter: required K scales inversely with the square of clarity away from the threshold.At the superpopulation boundary decisions do not stabilize as K→∞, whereas finite-census sampling becomes exact at K = M.
- Finite-sample validity and power: No violations occurred in 8,000 experiments for either exact outer binomial inversion or the Hoeffding benchmark, confirming implementation validity but not proving the certificate.Experiments used 2,000 repetitions per configuration and included shared-row dependence that preserves exact binomial marginals while inducing strong column dependence.
- Finite-sample validity and power: A baseline design usually cannot certify true coverage of 0.70 against a target of 0.60, whereas the powered design usually can.This comparison separates statistical validity from practical power.
- Panel size, clarity, and adversarial participation: Adversarial tolerance varies with task clarity, panel size, attack model, and error target rather than forming a single safe percentage.The fixed-share and targeted-contamination models produce different sensitivity surfaces, and the nominal benchmark is not a deployed-network security guarantee.
4 Applied Case: A Real-Data LLM Pilot
The LLM pilot compares fixed-catalogue, frozen-census, and hypothetical outer-sampling interpretations of panel reproducibility, showing that their guarantees differ. The manually stratified catalogue supports diagnostic and conditional claims, not workload-population inference.
- Pilot design: The pilot used 50 manually stratified frozen workload units, forming a complete seed catalogue rather than an i.i.d. workload sample.The units were labelled by construction stratum, not independently adjudicated for semantic truth.
- Fixed catalogue: At K=47, evaluator intervals certified 45 of 50 fixed-catalogue units, giving lower coverage 0.900 under uniform weights and 0.8775 under declared weights.The stated confidence was 0.975 under the declared i.i.d. evaluator-row model.
- Frozen census: Under the frozen evaluator census, 47 of 50 units had stable panel size K≤7 at δ=0.01, while mkt-03, ins-03, and air-03 required 26, 34, and 39.Without-replacement panels necessarily satisfy K≤40.
- Reference-object contrast: The same matrix yielded 47 of 50 frozen-census stability results at K=7 but zero evaluator-superpopulation interval certificates, because the reference objects differ.This contrast motivates keeping finite-census and superpopulation regimes separate.
- Outer-sampling diagnostic: Under a hypothetical i.i.d. outer design, exact binomial inversion first crossed the 0.60 pilot target at K=23, whereas the retained Hoeffding audit reference crossed at K=47.At K=23, 44 units passed and the exact lower bound was 0.6372; these values are not workload-population estimates for the actual stratified catalogue.
- Contamination sensitivity: At K=7, non-adaptive frozen-census sensitivities fell from 47 qualifying units with no flips to 46 after two and 44 after four worst-direction flips per column.These are conditional sensitivities, not guarantees against adaptive corruption or a deployed network.
- Interpretation: Nine of 13 units constructed as underdetermined passed the K=47 pointwise rule, but the pilot cannot establish semantic correctness without independent adjudication.Construction agreement and reproducibility relative to a declared population are distinct from externally validated accuracy.
5 Discussion
The discussion frames ADS certification as a reproducibility guarantee whose meaning depends on the evaluator and workload reference objects. It also emphasizes power, dependence, adversarial-model, and deployment limitations.
- Interpretation: A frozen census supports exact design-based claims, a probabilistic evaluator population supports fresh-panel claims, and a growing catalogue needs justified approximation.The certificate combines evaluator uncertainty with generator sampling without pooling resolution-specific margins.
- Dependence: Shared model families can produce strong cross-task dependence and common bias simultaneously, so evaluator-row dependence must remain part of the analysis.The paper allows shared-row dependence in its certification construction.
- Interpretation: Failure to certify may reflect genuine ambiguity or inadequate experimental power, while successful certification establishes reproducibility rather than truth.The baseline Monte Carlo design is valid but underpowered.
- Adversarial tolerance: Adversarial tolerance depends on task clarity, panel size, sampling mechanism, and attacker model rather than forming a universal percentage.Sensitivity curves require a pinned deployment population and inclusion law to become finite-sample guarantees.
- Limitations: The pilot is limited to 50 manually stratified units, the N=1 identity map, no independent semantic adjudication, and assumptions including a predeclared finite grid and i.i.d. rows.The discussion identifies clustered models, time-uniform confidence sequences, adaptive adversaries, and adjudicated workloads as extensions.
- Scope: No production deployment outcomes or user data are analyzed, and the statistical results do not depend on the application implementation used for the sensitivity grid.This bounds the applied evidence to the diagnostic pilot.
Appendices: Proofs and Technical Results
The appendices formalize ADS objects, evaluator laws, global aggregation, finite-census sampling, and dependence assumptions. They distinguish operational reproducibility from semantic truth and derive measurable, exact probability models.
- Certification objects: The technical framework defines empirical resolvability coverage, evaluator interval certificates, generator-level lower bounds, and contamination profiles over a predeclared panel-size grid.These quantities separate evaluator- and generator-level calibration budgets and distinguish familywise from mass-controlled constructions.
- Formal ADS model: An ADS freezes a generated resolution, obtains evaluator component verdicts, applies a deterministic global map, and aggregates the resulting global votes.The global map is applied after each evaluator’s complete vector is formed, not to component-wise population majorities.
- Evaluator episodes: Evaluator episodes combine a declared configuration space with private randomness, while prompt interpretation and affirmative-response parsing define binary component verdicts.Malformed and non-affirmative responses count as zero, and private evaluator behavior is included in the measurable interpreter.
- Dependence: The model permits dependence among one evaluator’s component coordinates, but independent episodes produce i.i.d. binary global votes under the declared product construction.A shared external random state would require conditioning or an added common coordinate and could induce dependence.
- Population decision: Population decisions use an inclusive threshold and define clarity as distance from that threshold; at zero clarity, finite panel decisions need not stabilize on one side.The tie policy determines the target at the boundary.
- Finite census: Finite-census analysis conditions on frozen evaluator verdicts, leaving only the without-replacement sampling design as random.This regime is constructed directly rather than importing independence from the superpopulation model.
C.2 Exact Panel Law
The exact panel law conditions on a realized finite evaluator census and computes panel–census disagreement through hypergeometric sampling. This avoids approximations while making the finite-census question distinct from binomial superpopulation inference.
- Exact panel law: Conditional on a frozen census with C_M positive verdicts among M evaluators, a K-member panel’s positive count follows the exact hypergeometric law.The law counts subsets containing h positives and K−h negatives.
- Exact error: The resulting panel–census disagreement probability is computed exactly from the hypergeometric distribution, with infeasible terms set to zero.Once C_M is known, no normal approximation or Monte Carlo simulation is required for uniform finite-panel certification.
- Finite-population correction: The finite-census panel mean is the hypergeometric positive-count proportion, whose variability includes a finite-population correction relative to independent sampling.The correction depends on M and K, including the factor involving (M−K)/(M−1).
- Numerical implication: For the illustrated setting, exact disagreement probabilities are approximately 0.997, 0.984, and 0.966 at K=10, 50, and 100, respectively, and zero at the full census.A binomial calculation may be numerically close for a small sampling fraction but answers a different conditional question.
C.3 Finite-Sample Certificates
Finite-census certificates quantify exact panel disagreement relative to a declared frozen evaluator census, while growing-census guarantees require explicit convergence and rate assumptions.
- Finite-census certificates: Exact hypergeometric tails quantify finite-panel disagreement relative to a frozen evaluator census.Serfling bounds provide a distribution-free alternative, but exact hypergeometric inversion is preferred when census counts are available.
- Finite-census certificates: Minimum certified panel size uses the exact error sequence and need not decrease at every consecutive panel size.Parity effects from integer quotas and tie conventions can cause nonmonotonicity; nested-panel deployment requires the stable-size definition.
- Scope: A frozen census certifies agreement with that specified census, not representativeness of a larger population or reproducibility under changed conditions.Those claims require separate population or campaign-level assumptions.
- Growing designs: Growing finite designs require a declared evaluator law and integral convergence, because weak convergence alone may fail when the binary vote map is discontinuous.A sufficient route is weak convergence plus almost-everywhere continuity of the vote map.
- Growing designs: Deterministic census consistency holds eventually when the limiting acceptance margin is positive, but convergence alone supplies no numerical census size.A finite-to-limit certificate needs a justified rate or error envelope.
D.2 Probabilistic Superpopulations
Probabilistic superpopulations define panel behavior through i.i.d. evaluator sampling and distinguish unconditional population laws from conditional realized-census laws.
- Sampling laws: The unconditional subsample law is binomial, whereas conditioning on realized census verdicts yields the hypergeometric law.The finite-population correction is therefore conditional on a realized census, while the binomial law describes a randomly generated census.
- Finite-to-limit distinction: A finite census used as a proxy for a limiting population contains separate panel-versus-census and census-versus-limit discrepancies.The second discrepancy requires a deterministic approximation rate or a probabilistic census-generation model.
- Resolution-level analysis: Generator and evaluator populations are distinct, and each concrete resolution requires its own acceptance mean and panel error before aggregation.Pooling all matrix entries can hide resolution-level clarity and produce misleading decisions.
- Population model: Under a declared task-specific evaluator law, a fresh panel consists of i.i.d. evaluator episodes and has an exact binomial disagreement tail.This is not a CLT approximation and can also represent the fixed-panel limit of admissible finite censuses.
- Population model: Population resolvability requires that most generator-produced resolutions have panel disagreement at most δ relative to the limiting population decision.The generator mass of exceptions is bounded by β, and resolutions are evaluated separately rather than pooled.
E.4 An End-to-End Finite-Sample Certificate
The end-to-end certificate combines evaluator-mean intervals with generator-level binomial inversion to lower-bound population resolvability from a finite matrix.
- Certificate construction: A simultaneous certificate accounts for uncertainty in both evaluator rows and generator columns without a CLT.It can reuse evaluator rows across columns and does not require cross-column independence.
- Certificate construction: Exact evaluator confidence intervals certify pointwise resolutions, after which exact one-sided binomial inversion lower-bounds generator-level coverage.The evaluator and generator failure budgets are ηE and ηG, yielding confidence at least 1 − ηE − ηG.
- Interpretation: The certificate separates δ for fresh-panel disagreement, β for uncertified generator mass, and ηE and ηG for calibration confidence.The latter two are not runtime panel-error probabilities.
- Mass-controlled alternative: The mass-controlled construction replaces the familywise requirement that no sampled interval fails with a bound on the generator mass of interval failures.Its evaluator interval level is independent of the number of generator columns, while ξE contributes to the unresolved-mass budget.
- Power and design: For |K| = 8 and ηG = 0.025, sufficient generator counts are A ≥ 513 for (β, ξE) = (0.10, 0.025) and A ≥ 1803 for (0.05, 0.01).These thresholds prevent vacuous bounds when every sampled column certifies; near-boundary resolutions still reduce power.
E.5 The Finite-to-Ideal Bridge
The finite-to-ideal bridge shows when finite evaluator censuses approximate limiting population decisions and panel errors, while emphasizing that rates and population design choices remain essential.
- Consistency: For fixed K and positive limiting margin, admissible census sequences eventually produce the same decision as the limiting evaluator population.The result is qualitative unless a convergence rate is supplied.
- Quantitative bridge: Under the stated variance condition and M = 1500, hypergeometric-to-binomial envelopes are approximately 0.0060, 0.0327, and 0.0660 for K = 10, 50, and 100.These are worst-case distributional envelopes, not the exact tail errors.
- Coverage: Finite experimental coverage converges for finitely many resolutions when limiting errors avoid the certification boundary.For i.i.d. generator samples, the empirical coverage is an average of binary variables and receives an additional generator-sampling bound.
- Illustration: Synthetic Figure 6 computes each resolution’s census acceptance rate and exact minimum panel size before obtaining family-level coverage.Its values illustrate the protocol and are not empirical claims about a deployed LLM population.
- Population robustness: Plug-in reweighting of one evaluator matrix is descriptive sensitivity analysis, not a certificate for an alternative population.Valid alternative-population inference requires appropriate sampling or weighting design.
F.1 Census-Level Robustness
A corruption budget of b entries can reduce census acceptance clarity by at most b/M. Any strictly smaller reduction cannot cross the threshold, while equality depends on the inclusive decision rule.
- F.1 Census-Level Robustness: b/M bounds the reduction in effective census clarity under b-coordinate corruption.Changing at most b binary entries changes the acceptance proportion by at most b/M.
- F.1 Census-Level Robustness: A perturbation strictly smaller than the honest clarity preserves the census decision.The perturbed mean remains on the same side of the threshold.
- F.1 Census-Level Robustness: At equality, an inclusive threshold may preserve acceptance but can flip a rejecting decision.Equality can reach the boundary from the rejecting side, whereas acceptance may remain preserved.
F.2 Exact Worst-Case Panel Error
The paper derives exact worst-case panel disagreement under non-adaptive population corruption and distinguishes it from post-sampling vote changes. Robustness holds below the clarity margin, while contamination at or above that margin admits no population-independent guarantee.
- F.2 Exact Worst-Case Panel Error: The exact worst case is obtained by changing as many positive census entries as the corruption budget permits.For an accepting census, the adversary minimizes positives; the rejecting case is symmetric.
- F.2 Exact Worst-Case Panel Error: The robust finite-panel certificate remains an exact hypergeometric tail based on census positives and the integer corruption budget.It requires no evaluator-accuracy notion relative to the population decision being estimated.
- F.2 Exact Worst-Case Panel Error: If b/M is below census clarity, every admissible corrupted census remains on the honest side of the threshold.The same margin argument yields concentration around the honest decision.
- F.2 Exact Worst-Case Panel Error: For convergent honest censuses, contamination below limiting clarity preserves the honest limiting decision, while contamination appears as deterministic drift bounded by its share.This asymptotic statement assumes the honest clarity is positive and corrupted proportions converge.
- F.2 Exact Worst-Case Panel Error: At contamination share at least equal to clarity, no population-independent guarantee is possible because corruption can reach or cross the threshold.This obstruction is deterministic and precedes any central-limit approximation.
- F.2 Exact Worst-Case Panel Error: Post-sampling corruption of panel votes produces shifted exact tails and is a different threat model from non-adaptive corruption of the eligible population.Deployments must specify which corruption mechanism is being analyzed.
F.5 From a Corruption Fraction to a Sensitivity Profile
The sensitivity profile distinguishes random network ownership from targeted contamination and reports workload-level attack-risk guarantees. Larger panels reduce sampling variation but can amplify a contaminated population decision, while stabilization and pilot results remain scope-limited.
- F.5 From a Corruption Fraction to a Sensitivity Profile: Committee-membership capture curves are independent of task clarity, whereas outcome-manipulation frontiers depend on clarity and the attack model.A task–panel pair above a curve meets the declared pointwise attack-risk target under the benchmark.
- F.5 From a Corruption Fraction to a Sensitivity Profile: For fixed network share, the population frontier depends on honest-subpopulation clarity through the mixture model.The network-ownership interpretation requires the stated mixture assumption.
- F.5 From a Corruption Fraction to a Sensitivity Profile: Larger panels suppress random committee variation but increasingly concentrate on the adversarially shifted population decision.Below the boundary, larger panels reduce sampling error; above it, they amplify the remaining population side.
- F.5 From a Corruption Fraction to a Sensitivity Profile: Targeted contamination is at least as damaging as fixed network share for the same numerical contamination and clarity.The targeted model lets the attacker select entries supporting the honest decision rather than merely control random identities.
- F.5 From a Corruption Fraction to a Sensitivity Profile: The certificate lower-bounds the fraction of workload units whose pointwise attack probability is at most a chosen tolerance.It combines nested evaluator evidence with an outer finite-sample bound over workload units.
- F.5 From a Corruption Fraction to a Sensitivity Profile: The finite-census, superpopulation, and growing-census analyses retain distinct laws and require explicit convergence or sampling assumptions before transfer.Observed stabilization is evidence for a limit, not proof of a universal population.
- F.5 From a Corruption Fraction to a Sensitivity Profile: At the superpopulation boundary, finite i.i.d. decisions need not stabilize, whereas positive clarity yields eventual stabilization.The almost-sure stabilization result excludes equality with the threshold.
- F.5 From a Corruption Fraction to a Sensitivity Profile: The pilot had saturated margins and was not a high-power test of intermediate clarity, multicomponent mappings, or hostile presentation.Its design also used a stratified seed catalogue rather than an identified i.i.d. workload law.