Source-linked AI summary

Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements

Naoki Egami, Sooahn Shin

arXiv:2608.18294v1stat.MEcs.AIcs.CLcs.LGstat.ML

TL;DR

Ignoring errors in AI-generated measurements can bias downstream inference and invalidate confidence intervals. DMM combines multiple imperfect measurements under conditional independence assumptions, and simulations and empirical validation show small bias, nominal coverage, and estimates near an oracle benchmark without expert labels.

  • Problem

    Ignoring non-random, heterogeneous errors in AI annotations can bias downstream inference and invalidate confidence intervals even when measurement accuracy is high.

  • Method

    DMM combines three or more imperfect measurements and conditions on observed features to support valid downstream inference without gold-standard labels.

  • Results

    DMM achieves small bias and nominal coverage in simulations and produces estimates and confidence intervals close to an oracle benchmark in empirical validation.

  • Takeaways & Limitations

    Multiple imperfect AI measurements can support valid downstream inference without expert labels when the conditional independence framework is appropriate.

  • Takeaways & Limitations

    More proxies do not mechanically improve the particular DMM estimator’s efficiency, although alternative weighting can address this issue.

Abstract

from arXiv · show

An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. Although AI-measured variables are often analyzed as if observed without error, ignoring prediction errors in automated measurement leads to substantial bias and invalid confidence intervals in downstream analyses, even if AI measurement accuracy is high, e.g., above 90%. Existing solutions, such as design-based supervised learning and prediction-powered inference, combine error-prone AI-based measurements with gold-standard labels, which may be costly and difficult to obtain in some application areas. In this paper, we propose debiased inference with multiple imperfect measurements (DMM), a framework that combines multiple error-prone AI measurements to enable valid downstream inference without gold-standard labels. Building on the established results on CP decomposition, DMM assumes that these measurements are independent conditional on the latent true label and observed unit-level features, such as text features represented by embeddings. This framework allows for unknown misclassification rates to vary across annotation methods (e.g., large language models) and across units of annotation (e.g., texts). Under this assumption, we use semiparametric inference theory to prove that the DMM estimator is consistent and asymptotically normal, enabling valid inference for a wide range of downstream statistical analyses common in the social sciences. Our simulation results show that DMM yields valid inference and that adding accurate, though imperfect, measurements can improve efficiency. Focusing on common applications of large language model annotations, we also develop diagnostics to assess the conditional independence assumption.

1 Introduction

The paper introduces DMM, which combines multiple imperfect AI measurements to enable debiased downstream inference without gold-standard labels. It identifies latent labels under conditional independence given observed features, extends applicability across downstream analyses, and develops diagnostics for LLM annotation settings.

  • Ignoring heterogeneous AI measurement errors can produce substantial downstream bias and invalid confidence intervals, even when measurement accuracy exceeds 90%.
  • DMM combines three or more imperfect measurements to identify latent labels and conduct valid downstream inference without assuming access to gold-standard labels.Identification builds on nonparametric latent-variable models and CP decomposition.
  • Conditional independence is imposed given the latent true label and observed unit-level features, allowing measurement errors to share sources captured by those features.The framework makes this identifying assumption explicit and discusses strategies for assessing its plausibility.
  • The DMM estimator uses semiparametric inference theory because estimating a latent-variable model and directly including learned latent variables downstream is insufficient for valid inference.
  • DMM covers linear and logistic regression, most maximum likelihood estimators, categorical error-prone dependent or independent variables, and more than three proxies.
  • For LLM annotations, the paper proposes choosing conditioning variables, proxy model families, and prompts to make conditional independence more plausible and evaluates its observational implications.

2 The Problem Setting and Existing Approaches

Measurement error in machine or human annotations can bias downstream inference even when annotation accuracy exceeds 90%. Existing approaches either aggregate imperfect labels without eliminating nonclassical error or use gold-standard labels to recover valid inference, often sacrificing efficiency.

  • Problem setting: Annotation errors can bias downstream inference even when accuracy exceeds 90%.The section first focuses on binary error-prone independent variables and later generalizes to categorical variables used as independent or dependent variables.
  • Problem setting: Multiple proxies may come from human annotators, machine-learning models, or LLMs, with accuracies and errors varying across methods and units.The framework allows systematic or heterogeneous errors, and multiple LLM-generated proxies are often readily available.
  • Proxy-label approaches: Aggregating imperfect labels, such as by majority vote, does not justify treating the resulting proxy as error-free downstream.Nonclassical classification errors can remain correlated with observed and unobserved variables relevant to downstream regression, even when aggregate classification accuracy is high.
  • Gold-standard approaches: Gold-standard-only estimation is consistent and asymptotically normal when the sampling probability is known, but it discards imperfect-label information and can be inefficient.GSO runs the downstream regression only on units with gold-standard labels.
  • Bias-correction approaches: Design-based supervised learning combines large-scale predictions with sampled gold-standard labels to bias-correct downstream moments and remains valid under arbitrary misspecification of predictors and LLM annotations.DSL uses cross-fitting and a doubly robust procedure, assumes a known sampling design, and becomes more accurate when predictions improve.

3 Debiased Inference with Multiple Imperfect Measurements

DMM uses multiple imperfect measurements and conditional-independence-based latent-variable identification to recover unbiased downstream moments without gold-standard labels. It delivers consistent, asymptotically normal inference, while additional proxies improve precision only under suitable quality conditions.

  • Framework: DMM recovers unbiased downstream moments from the joint distribution of multiple imperfect labels without requiring gold-standard labels.The approach relies on latent-variable identification under conditional independence.
  • Assumptions: The framework permits proxy-label misclassification rates to vary across annotators and units, rather than requiring perfect or identically distributed measurements.An anchor proxy can be oriented by requiring its true-positive rate to exceed its false-positive rate.
  • Assumptions: Conditional independence may accommodate shared error sources when input-, annotation-, and downstream-level information is included among conditioning variables.These variables can capture heterogeneous accuracy and dependence without necessarily entering the downstream regression.
  • Inference: DMM estimation is consistent and asymptotically normal under the identification and regularity conditions.Consistency requires nuisance-function convergence, with multiple robustness requiring consistent estimation for only J−1 proxies.
  • Efficiency: Adding proxies improves precision when the new proxy’s conditional bridge-function variance is sufficiently small, but more proxies do not mechanically improve this estimator.Under the stated result, the relevant remainder decreases with J at rate J−2.

4 Extensions

Section 4 extends DMM from binary independent variables to binary dependent variables and multicategory outcomes. The extensions preserve identification, unbiased bridge-based moments, and large-sample inference under conditional independence and stronger rank or anchoring conditions.

  • Binary dependent variables: DMM extends to binary dependent variables by conditioning independently measured outcome labels on the latent outcome and measurement covariates.Identification uses analogs of the overlap, relevance, and anchoring conditions, with outcome-specific proxy bridges yielding an identified downstream parameter.
  • Binary dependent variables: The binary-outcome estimator is constructed by fitting conditional latent-class models within training folds and applying the anchoring rule to out-of-fold bridges.The resulting estimator inherits the large-sample theory and covariance estimation strategy of the binary independent-variable case.
  • Multicategory outcomes: For multicategory latent outcomes, class-specific bridges combine tensor-decomposition identification with left inverses of the joint measurement matrix.Under conditional independence and corresponding rank and anchoring conditions, the observed-data moment is unbiased for the oracle full-data moment.
  • Multicategory outcomes: At least three conditionally independent measurements retain local robustness, and composite measurements can supply enough categories to distinguish all latent classes.The robust bridge recovers latent class indicators in conditional expectation and preserves first-order cancellation of first-stage estimation error.
  • Multicategory outcomes: Multicategory identification requires richer proxy supports, sufficiently independent measurement-matrix columns, and anchoring conditions matching all latent classes.Three full-column-rank measurement matrices provide a sufficient condition because Kruskal’s requirement reduces to 3K∗≥2K∗+ 2 for K∗≥2.

5 Designing and Assessing Conditional Independence

The paper recommends designing multiple measurements and conditioning sets to make conditional independence plausible, then assessing the assumption with observable diagnostics. When the assumption remains doubtful, a small gold-standard sample enables direct testing and bias correction while DMM can improve precision.

  • Designing conditional independence: Researchers should condition on rich annotation- and unit-level information, including task difficulty or disagreement among auxiliary proxies, to make conditional independence more plausible.The conditioning set may include input or annotation information D_i, downstream variables W_i, and disagreement scores from additional proxies excluded from the main analysis.
  • Designing conditional independence: Measurements should use substantively different error mechanisms, such as different model families or prompts, rather than being selected solely for predictive accuracy.Randomly sampling prompts from an admissible pool can reduce shared sources of correlated error when one validated prompt cannot be fixed.
  • Assessing conditional independence: Observable diagnostics compare fitted product-mixture distributions with empirical proxy-label distributions, using bootstrap deviance, conditional moments, residual interactions, or held-out predictive checks.For continuous or high-dimensional conditioning variables, cross-fitted residual interaction moments and held-out checks are more practical than cell-by-cell tests.
  • Assessing conditional independence: Subset-specific DMM estimates should agree up to sampling error, but failure to reject their equality is only a specification diagnostic and does not establish conditional independence.Joint bootstrap or multiplier-bootstrap tests can assess heterogeneity across estimates based on informative proxy subsets.
  • Gold-standard validation: A small gold-standard sample permits direct testing and bias correction without assuming conditional independence, although limited power means nonrejection cannot establish the assumption.DMM-assisted DSL uses DMM to impute the full-data moment for all observations, while validation observations estimate and correct remaining imputation error; the validation design provides validity and DMM can improve precision.

6 Simulation and Empirical Validation

Simulation and real-world validation show that DMM corrects substantial measurement-error bias, restores confidence-interval coverage, and can improve precision as informative proxy labels are added. In an empirical application without gold-standard labels, DMM achieves coverage closer to the expert-label benchmark than naive estimators.

  • 6.1 Simulation: With the top three proxies, naive majority voting has bias 0.186 and coverage 0.018, whereas DMM has bias 0.002 and coverage 0.948.DMM’s RMSE is 0.063, compared with 0.191 for majority voting and 0.045 for the oracle benchmark.
  • 6.1 Simulation: Across evaluated proxy sets, DMM’s bias remains 0.002–0.003 and coverage 0.946–0.962, while majority-vote bias is 0.085–0.226 and coverage is 0–0.538.With 15 labels, DMM has RMSE 0.047 versus 0.096 for majority voting.
  • 6.1 Simulation: DMM’s RMSE falls from 0.063 with three labels to 0.046 with 13 labels, approaching the oracle RMSE of 0.045 before gains flatten.The intermediate RMSE values are 0.053, 0.048, and 0.047 with the top 5, 7, and 10 labels, respectively.
  • 6.2 Empirical Validation: In the Pan and Chen application, the expert-label benchmark is −1.039; DMM estimates −0.844 without gold-standard labels, and its confidence interval contains the benchmark.DSL uses 500 gold-standard labels and estimates −1.279, while naive estimates range from −1.869 to −0.388.
  • 6.2 Empirical Validation: Empirical coverage is 0.862 for DMM and 0.981 for DSL, compared with 0.154–0.572 for the three naive estimators.Coverage is computed as the proportion of 95% confidence intervals containing the benchmark over 500 resamples.

7 Discussion

DMM enables valid downstream inference without gold-standard labels by combining multiple imperfect measurements under conditional independence. Its precision benefits from additional informative measurements, but the framework depends on an assumption that may be violated and should be complemented by validation-based approaches when available.

  • Core contribution: DMM identifies the conditional classification-rate model needed for valid downstream inference without gold-standard labels, assuming measurements are conditionally independent given the latent variable and observed features.Individual measurements may still have systematic errors.
  • Core contribution: The framework supports downstream analyses where the latent variable is either an independent or dependent variable by recovering oracle full-data moments through a robust bridge.The bridge’s conditional expectation recovers the latent indicator.
  • Efficiency and design: DMM accommodates measurement collections ranging from three imperfect measurements to much larger sets, with informative additions reducing measurement-induced asymptotic variance.Additional measurements can bring DMM closer in efficiency to the oracle estimator.
  • Validation: Simulations show small bias and nominal coverage across label counts, while naive estimators remain biased and undercover even when individual proxies have high F1 scores.The simulations were calibrated to data from Fowler et al. (2021).
  • Limitations and complements: DMM’s main limitation is conditional independence, which high classification accuracy does not guarantee; limited validation samples can assess the assumption and support combining DMM with DSL.Violations may arise when annotations interact with omitted input features.
  • Limitations and complements: DMM and validation-based methods are complementary, with DMM suited to multiple reasonably accurate measurements having non-random, nonclassical errors.Even small uncorrected measurement errors can substantially bias downstream inference.

A Proofs

The proof partitions observations into fixed-proportion folds and defines the DMM score using nuisance-adjusted interpolation between latent-label-specific scores. Fold-specific nuisance estimates are fitted on complementary training observations.

  • Cross-fitting: The proof partitions {1, . . . , n} into K folds with fixed fold proportions and uses complementary indices I−k for training.For fold k, Ik contains nk observations and I−k denotes the corresponding training indices.
  • Score construction: The DMM score ψDMM combines ψ0(Y, W; β) and ψ1(Y, W; β) using weights 1 − HR and HR, respectively.The weights depend on observed features and measurements through HR(e X, eD; η).
  • Nuisance functions: The weighting function is defined as HR = 3H2 − 2H3, while ψa(Y, W; β) evaluates the downstream score at the latent label X∗ = a.The construction uses η∗ for the population weighting function and fold-specific fitted nuisance objects based on bη−k.

A.1 Proof of Proposition 3.1: Bounds for the Robust Bridge

Proposition 3.1 establishes robust bridge error identities whose terms involve errors from at least two distinct proxy labels, eliminating linear dependence on any single first-stage error. Consequently, the bridge remains robust when nuisance functions for at most one proxy are misspecified.

  • Error decomposition: The conditional bridge error contains no term linear in a single first-stage error.The proof derives this from cancellation among averaged linear terms in the robust bridge construction.
  • Error decomposition: Every error term in the displayed identities is a product involving errors from at least two distinct labels.This structure follows from conditional independence of the fitted single-proxy bridges and cancellation of the averaged linear terms.
  • Norm bounds: Uniformly bounded-away-from-zero class contrasts ensure single-bridge errors are uniformly bounded, allowing cubic terms to be controlled by pair products.With fixed J, summing the finite collection of terms establishes the corresponding norm bound.
  • Robustness: If nuisance functions are correctly specified for all but one proxy, all pair and triple error products vanish, yielding multi-proxy robustness.The proposition’s proof applies the lemma at the true nuisance functions, where the single-bridge errors equal zero.

A.2 Proof of Theorem 3.2: Second-Order Remainder for the DMM Moment Function and Large-Sample Theory

The proof establishes that the DMM moment-function remainder is second order under the theorem’s regularity conditions, yielding consistency, robustness, and asymptotic normality of the cross-fitted estimator.

  • Second-order remainder: Under Lemma A.1 and envelope conditions, Lemma A.2 bounds the conditional population remainder uniformly over folds and β in B0.The same rate applies to the aggregate remainder because the number of folds is fixed and fold weights sum to one.
  • Consistency: The cross-fitted sample moment converges uniformly to its population counterpart because empirical, oracle, and remainder terms are each uniformly op(1).The argument combines conditional concentration, a uniform law of large numbers, and Lemma A.2 with δn = op(1).
  • Consistency and robustness: The DMM estimator is consistent, bβDMM p→β∗, by well-separated-root Z-estimation.The proof also shows robustness when nuisance functions for all but one proxy are consistently estimated, because each bridge-error product contains an op(1) factor.
  • Asymptotic Normality: At β∗, nuisance-estimation contributions are op(n−1/2), allowing the mean-value expansion and central limit theorem to establish asymptotic normality.The limiting Jacobian is −A0, with A0 nonsingular, and Slutsky’s theorem yields the stated asymptotic distribution.
  • Asymptotic Normality: Consistency and foldwise laws of large numbers establish convergence of the estimated variance components to the population sandwich form.The cross-fitted moment function converges in L2(P) to ψDMM(O; β∗, η∗).

B Details for the Efficiency Discussion

The efficiency discussion characterizes when adding an imperfect proxy improves precision. Under conditional independence, the threshold depends on the new proxy’s conditional bridge variance relative to existing proxies, with especially clear gains for equally accurate proxies.

  • Efficiency gains: Adding a proxy improves conditional bridge precision when its conditional bridge variance is below the existing threshold; if all current variances are zero, both variances remain zero.The threshold is stated through the variance comparison, while the zero-denominator case yields no change in either conditional bridge variance.
  • Efficiency gains: For every coefficient contrast, satisfying the threshold for both latent classes is sufficient for weak improvement, while a weighted-average condition is sufficient and necessary for a fixed contrast.The fixed-contrast criterion is exact, whereas the across-contrast criterion is sufficient.
  • Special case: An additional proxy with the same conditional bridge variance as existing proxies always satisfies the improvement condition.This special case assumes existing proxies share a common positive conditional bridge variance and the added proxy has matching variance.
  • Special case: With common variance κa(d) > 0, conditional bridge variance decreases in J at order J^-2, and the threshold converges to 2κa(d), allowing somewhat noisier proxies to improve precision.The improvement comes from additional averaging across proxies.
  • Limitations: These efficiency comparisons require conditional independence and fixed J; dependence adds covariance terms, while J = Jn →∞ requires control of nuisance dimension and minimum class contrast.The stated threshold therefore does not generally apply when proxy conditional independence fails.

C Additional Details on the Experiments · C.1 Scope and Relationship to the Original Applications

The appendix adds details on two numerical illustrations of DMM, clarifying their relationship to the original applications and explaining the shared experimental setup. The analyses simplify the original substantive studies, while expert labels serve benchmarking and calibration roles rather than entering feasible DMM estimation.

  • C Additional Details on the Experiments: The appendix supplements the two Section 6 studies with application context, the logistic latent-class model, the EM algorithm, and application-specific details.An “expert label” denotes the benchmark annotation throughout.
  • C.1 Scope and Relationship to the Original Applications: The two analyses are numerical illustrations of DMM rather than exact replications of the original substantive studies.This scope distinction frames how results from the applications should be interpreted.
  • C.1 Scope and Relationship to the Original Applications: The first application replaces the original candidate-adjusted analysis of political-advertising tone with an ad-level logistic regression for a binary Promote indicator.The original studies compared Facebook and television using an expert-coded corpus, whereas this analysis uses a transparent simplified specification.
  • C.1 Scope and Relationship to the Original Applications: The second application focuses on prefecture-level wrongdoing as a single main effect rather than modeling both prefecture- and county-level wrongdoing.This avoids introducing a second latent variable and the county-wrongdoing interaction used in the original analysis.
  • C.1 Scope and Relationship to the Original Applications: Expert labels define the full-sample benchmark in both applications, calibrate the synthetic Fowler data-generating process, and rank proxy labels by F1 score.These roles support benchmarking, calibration, and construction of interpretable nested proxy sets.
  • C.1 Scope and Relationship to the Original Applications: Expert labels do not enter the feasible DMM moment equations within a Monte Carlo replication or the DMM fit.Thus, their retrospective F1-based ranking is separate from feasible DMM estimation.

C.2 Model Specification … C.3 Monte Carlo Simulations

The applications use a shared low-dimensional logistic latent-class model estimated by EM, then construct DMM estimates from fitted nuisance components. Monte Carlo simulations calibrate this framework to Facebook-ad annotation data under exact conditional independence and compare feasible estimators with an oracle.

  • C.2 Model Specification: Both applications use a low-dimensional logistic latent-class model for latent labels, proxy labels, and nuisance functions.The nuisance-model design vector includes an intercept and observed covariates.
  • C.2 Model Specification: Conditional independence yields the observed-data likelihood while excluding the downstream dependent variable from nuisance estimation.The implementation imposes conditional independence among proxy measurements given the latent label and covariates, and between proxy measurements and the downstream outcome given the latent label and covariates.
  • C.2.1 Expectation–Maximization Algorithm: The EM algorithm updates latent responsibilities and separates the M-step into logistic regressions for latent prevalence and proxy response models.Proxy-specific regressions use class-specific effective weights.
  • C.2.1 Expectation–Maximization Algorithm: Six candidate starts are evaluated, and the solution with the largest final observed-data log-likelihood is retained before orienting latent classes by proxy contrast.The starts include the row-wise proxy mean, its complement, single-proxy starts, and additional Uniform(0.2, 0.8) draws when needed.
  • C.2.2 From the EM Nuisance Fit to DMM: DMM converts fitted proxy response probabilities into stabilized bridges and solves a generated-outcome quasi-score with sandwich standard errors.The bridge denominator uses a floor of ϵ = 0.05, described as a finite-sample stabilization device rather than part of population identification.
  • C.3 Monte Carlo Simulations: The Monte Carlo design holds 12,973 observed covariate rows fixed and generates conditionally independent proxy labels from calibrated logistic models.The simulated population satisfies conditional independence exactly, and the EM nuisance family is correctly specified; each proxy-count cell uses 500 replications.
  • C.3 Monte Carlo Simulations: Simulations compare naive majority vote, DMM, and an oracle across nested proxy sets ranked by F1 score, while additional outcomes include Contrast and Attack.Promote accounts for 75.6% of advertisements, versus 17.2% for Contrast and 7.1% for Attack; their target Facebook coefficients are −1.029 and −1.334, respectively.

C.4 Empirical Validation

The empirical validation uses 1,412 Chinese citizen complaints to compare DMM with naive, expert-label, and DSL estimators in a downstream model of complaint escalation. It also evaluates proxy-label quality and fixed-target bootstrap coverage against the expert-label coefficient.

  • Data and downstream model: The analysis studies 1,412 Chinese citizen complaints, with expert-coded wrongdoing prevalence 0.055 and mean escalation outcome 0.418.The latent independent variable indicates prefecture-level wrongdoing, while the outcome records whether complaints were sent to higher-level authorities.
  • Data and downstream model: The downstream model estimates τ using controls for complaint characteristics, with the full-sample expert-label benchmark equal to bτexpert = −1.0388.Controls include connect2b, prevalence, region, groupIssue, realWorldCollectiveAction, petitioning, sentiment_indico, and personal_experience.
  • Proxy labels: Three proxy labels—GPT-4.1 5-shot, GPT-4 5-shot, and Llama-4 0-shot—are evaluated against expert labels using descriptive quality metrics.Table C.3 reports class-weighted F1, positive-class F1, and sensitivity for the rare wrongdoing class.
  • DMM and benchmark estimators: DMM combines the three LLM labels through EM-estimated nuisance components and a robust bridge, while naive estimators use one label and the expert benchmark uses X∗i.DSL instead uses a simple random sample of 500 expert labels, about 35.4% of the sample, for supervised prediction and design-based correction.
  • Fixed-target bootstrap diagnostic: Empirical coverage is assessed by nonparametric bootstrap intervals tested against the fixed full-sample expert-label coefficient.Each draw resamples 1,412 complaints with replacement, refits the estimator, and records its interval.
Loading 2608.18294v1…