Source-linked AI summary
The Measurement Revolution? Credible Measurement and Inference in the Age of AI
Melissa Dell, Ashesh Rambachan
TL;DR
AI makes measurement from unstructured data scalable, but shifts the central challenge toward choosing among many plausible measures that may yield different conclusions. This review organizes AI's role across discovery, construct definition, and observation, and argues that credible inference depends on explicit validation criteria. Random validation samples can support valid inference even with arbitrarily biased predictions; without them, credibility depends on assumptions about model behavior.
Problem
AI creates many plausible measurement choices, raising credibility, selection, interpretation, and reproducibility challenges for empirical analysis.
Method
The review analyzes discovery, construct definition, and observation, then develops validation-based approaches for AI-generated variables with and without random validation samples.
Results
Credible inference can use random validation samples to learn and correct prediction errors without assuming AI predictions are unbiased.
Takeaways & Limitations
Researchers should specify verifiable measurement aims, evaluate candidate measures against explicit criteria, and design measurement to support valid, interpretable inference.
Takeaways & Limitations
When random validation is impossible and model-behavior conditions for identification fail, remaining validation may be only suggestive and requires awareness of potential pitfalls.
Abstract
from arXiv · showhide
Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as text and images, into structured variables at low cost, making previously prohibitive measurement feasible at scale. This shifts the bottleneck from finding any scalable measure of a phenomenon to choosing among many plausible ones, which may support different empirical conclusions. This review provides guidance for navigating that shift. We describe three stages at which AI enters the measurement pipeline---discovery, construct definition, and observation---and what each demands of researchers. We argue that credible inference with AI-generated variables requires appropriately designed validation: anchoring measurement to explicit criteria, rather than informal claims that a proxy is reasonable. We then examine how validation samples support valid inference even when AI predictions are arbitrarily biased, and what can be done when a random validation sample is unavailable.
1 Introduction
AI makes measurement from unstructured data cheap and scalable, shifting the bottleneck toward choosing among many plausible measures. The review argues that this abundance creates credibility and reproducibility challenges requiring explicit validation and disciplined inference.
- AI models convert unstructured data into low-dimensional structured variables at low marginal cost, extending measurement to previously difficult domains.Examples include LLM coding of text and computer-vision measures from satellite imagery.
- AI creates many implementation choices, including rubrics, models, prompts, training-data distributions, tuning strategies, and parameters.
- The measurement bottleneck is shifting from finding any scalable measure to choosing among many plausible measures that may support different empirical conclusions.
- Black-box model opacity makes it difficult to interpret which features are preserved in AI-generated measurements, especially when measurement choices affect empirical conclusions.
- Credible inference requires validating AI-generated variables against explicit, observable criteria rather than relying on informal claims that a proxy is reasonable.
- The review focuses on observation, examining how validation samples support valid inference under arbitrarily biased predictions and how assumptions matter without random validation samples.
2 The Measurement Pipeline: Discovery, Definition,
The measurement pipeline has three stages—discovery, construct definition, and observation—and AI can transform each by finding patterns, generating operationalizations, and applying definitions at scale. Each stage has distinct uncertainty and validation requirements, while theory and held-out evidence discipline the resulting choices.
- The Measurement Pipeline: Measurement proceeds through discovery, construct definition, and observation, each involving different uncertainty and requiring a different form of validation.Discovery concerns meaningful structure, construct definition concerns capturing the intended concept, and observation concerns accurate implementation at scale.
- The Measurement Pipeline: AI can surface previously unspecified patterns, generate and compare alternative operationalizations, and apply chosen definitions at scale.
- Discovery: Discovery procedures generate candidate constructs from unstructured data, but their quality must be evaluated on held-out data untouched by the generating procedure.
- Construct Definition: Construct validity for latent concepts is assessed through patterns of theoretically motivated relationships, including convergent, discriminant, and cross-setting evidence.
- Construct Definition: AI lowers the cost of comparing operationalizations, but theory and substantive expertise remain necessary to determine which construct answers the research question.
- Construct Definition: When credible external criteria are unavailable, discovery may be preferable because interpretable patterns can inform subsequent theory and construct definition.
3 From AI Outputs to Credible Measurements
Credible AI measurement requires explicit criteria that guide validation and comparison among candidate measures. The review shows that random validation samples can support correction without assuming unbiased predictions, while identification becomes assumption-dependent when such samples are unavailable.
- Validation as an Anchor: Validation specifies observable criteria, compares candidate measures against them, and adjusts target estimates for systematic measurement error.Criteria can be grounded in theory and domain expertise and used to trace disagreements to specific observations.
- Validation as an Anchor: Explicit criteria make black-box predictions more interpretable and discipline comparisons among the many measures AI can cheaply produce.
- Validation as an Anchor: Residual disagreement and ambiguous edge cases identify where construct definitions, procedures, or validation criteria require refinement rather than eliminating validation's value.
- Measurement Choices: Model and prompt choices can change estimates in magnitude, significance, and even sign, with prompt paraphrases potentially making virtually any hypothesis appear statistically significant.
- Routes to Credible Measurement: Credible routes therefore either validate model outputs as a black box or open the black box enough to justify assumptions about model behavior.
- Routes to Credible Measurement: A random validation sample reveals model errors relative to the intended construct, allowing downstream estimates to be corrected without assuming the AI predictions are unbiased.
- Routes to Credible Measurement: When random validation is impossible and behavioral identification conditions fail, researchers may pursue suggestive validation but must remain aware of the resulting pitfalls.
A Missing-
The review frames AI-generated measurement as inference with missing structured data: validation samples provide high-quality labels for debiasing scalable imputations. Its framework supports valid inference under arbitrary prediction bias, while prediction accuracy affects precision and practical power.
- Inference with AI Predictions as Inference with Missing Structured Data: Validation data supply high-quality labels from a well-specified construct definition, obtained through non-scalable processes such as expert annotation or scientific instruments.The validation sample is the basis for estimating the difference between imputed data and carefully measured labels.
- The Formal Framework: MAR-S requires annotated and unannotated observations to be comparable in ground-truth values after adjusting for observables, with annotation status well defined and consistently applied.The framework also uses an annotation score π(x), assumed in the baseline to be known and bounded away from zero, while that assumption can be relaxed.
- Inference with AI Predictions as Inference with Missing Structured Data: AI measurement can be treated as missing structured data, with neural networks imputing low-dimensional variables that are unavailable across the full unstructured dataset.Structured variables are directly usable in estimating equations, whereas unstructured data generally are not; imputation enables use of datasets much larger than validation samples.
- The Formal Framework: AI predictions can be arbitrarily inaccurate without compromising validity when estimates are properly debiased, but greater prediction accuracy reduces variance and improves precision.The imputation estimator’s mean squared error consistency is needed for efficiency, not unbiased inference.
- Aggregating AI Predictions: Debiasing frameworks require validation data for the imputed variables entering the estimating equation, but applications often observe ground truth only for granular texts or images.Carlson and Dell address this mismatch when the estimating variable is an aggregate or nonlinear functional of AI predictions.
- Annotation in Practice: More predictive imputations increase effective sample size, potentially reaching the full dataset size, whereas noisier predictions contribute less beyond the validation sample.The relevant asymptotic correlation is unknown, so effective sample size cannot be calculated directly without assumptions or scenario analysis; reasonably accurate binary classifiers may need only a couple hundred high-quality annotations.
5 Observation without Random Validation Samples
When random validation samples are unavailable, credible inference requires explicit assumptions about how AI-generated measurements behave, either across repeated protocols or across samples. These assumptions can support identification, but they must be defended and empirically assessed because model errors and relationships may change across settings.
- Multiple Measurements: With at least three conditionally independent, informative measurements of a binary construct, the joint distribution identifies the latent construct distribution and each protocol’s error rates without validation data.Conditional independence means different protocols err independently on the same input; informativeness requires each protocol’s true positive rate to exceed its false positive rate.
- Multiple Measurements: Frontier models often make correlated errors, so conditional independence must be evaluated against benchmark or domain evidence rather than assumed.The review notes that models can err on the same items and that economic measurement errors may be strongly correlated.
- Data Combination: When validation labels come from an auxiliary sample, inference requires either outcome stability or measurement stability across the auxiliary and analysis samples.Outcome stability transfers the distribution of truth given predictions; measurement stability transfers the distribution of predictions given truth, including model error rates.
- Data Combination: Under measurement stability, an informative model’s auxiliary-sample error rates can correct the analysis sample’s prediction rate to identify construct prevalence.The two stability assumptions use the same data but imply different formulas and can produce different answers.
- Data Combination: Existing predictive-performance benchmarks do not establish outcome or measurement stability, motivating benchmarks that directly assess whether these conditional distributions transfer across settings.Relevant shifts include regions, periods, sensors, and model versions.
- Implications: Without random validation, researchers must state assumptions explicitly, defend and evaluate them where possible, and report how conclusions change as assumptions are relaxed.This shifts the discipline of measurement from validation-sample design to the defense of identification assumptions.
6 Conclusion
AI expands what economics can measure by converting previously difficult-to-structure data into variables at scale. The article focuses on observation and argues that researcher-controlled validation can support interpretable inference even with black-box or biased AI predictions.
- AI transforms unstructured data into structured variables, allowing measurement processes once limited to small samples to scale across entire corpora.This widens the range of economic phenomena that can be measured systematically.
- The article focuses primarily on observation, where AI use is most widespread and the statistical challenges are most developed.
- A random validation sample can compare defensible measurements with AI-generated measures to correct downstream estimates.This logic parallels statistical agencies’ use of expensive, high-quality validation samples to correct cheaper census measurements.
- Validation-based inference need not model the neural network, assume unbiased predictions, or require accurate predictions.Prediction errors are learned from a validation design controlled by the researcher, supporting interpretation and defense of estimates.
- When random ground truth cannot be collected, stronger assumptions become unavoidable, and connecting frontier-model evaluations to credible-research assumptions remains future work.
- The article’s checklist identifies the cases, assumptions, requirements, and measurement details that researchers should report for transparent use of AI-generated variables.It covers construct definition, validation-label generation, and the model and prompt used.
A A Checklist for Observation with AI-Generated Vari-
The checklist helps authors report the measurement details readers need to assess AI-generated variables and recommends placing those details in supplemental or article materials.
- The checklist is intended to support transparent reporting of the details needed to assess AI-generated measures.
- Authors should answer the checklist questions explicitly in supplemental materials or document the details elsewhere in the article or replication materials.They should state “unknown” when a detail is unavailable.
Provenance and Reproducibility
The provenance and reproducibility checklist distinguishes constructed AI variables from externally obtained ones and asks researchers to preserve the materials needed to reproduce measurements.
- Researchers should state whether they constructed the AI-generated variables or obtained them from an external source, identifying the source when applicable.
- For constructed variables, researchers should release code, prompts, model versions, and permissible data no later than publication.
- Replication files should include model outputs so measurements can be reproduced without re-querying the model.
Construct Definition
The construct-definition checklist asks researchers to specify what AI measures and document the evidence used to validate that construct.
- Researchers should provide the definition of each construct measured with AI.
- Researchers should report whether construct validation was performed and what evidence it used.
- Validation evidence may assess whether the construct correlates with relevant variables or remains distinct from them.
Validation Regime
Researchers should identify which validation regime applies and document how validation data were obtained, sampled, and related to the sample of interest. When validation is unavailable or comes from another sample, they should explain feasibility constraints and justify the assumptions supporting measurement validity.
- The checklist distinguishes three regimes: no validation data, representative validation data, and validation data from a different sample.
- Without a validation sample, researchers should explain why obtaining one was infeasible and identify other evidence supporting the AI-generated measure.
- For representative validation, researchers should report how labels were sampled, whether annotation probabilities were known, and whether all observations had positive annotation probability.
- When validation labels come from another sample, researchers should describe sample differences and state the stability assumption linking the samples.
Validation Dataset Description (if applicable)
A validation dataset description should make label production and quality assessable. It should cover sample size and timing, labeling technology and procedures, consistency, and agreement among annotators or instruments.
- Researchers should report the validation sample size and when its labels were collected.
- They should describe whether labels came from human annotators or instruments, including relevant training, instructions, or technical specifications.
- The checklist asks whether a consistent rubric was used to produce validation labels.
- Researchers should report inter-annotator or inter-instrument agreement, the number of observations examined, and how disagreements were resolved.
Construction of the AI Predictions
Researchers should document how AI predictions were generated, including model and prompt choices, training data, generation settings, coverage, and transformations. They should also align any correction for measurement error with the estimand, the variable’s analytical role, and the validation regime.
- Model documentation should identify model versions, whether models were off-the-shelf or fine-tuned, citations, access dates, and hosted-model version strings.
- Researchers should report how models and prompts were selected, which alternatives were considered, and whether estimates were sensitive to those choices.
- For language-model prompts, researchers should provide the full prompt and generation settings, including temperature, top-p, maximum tokens, random seed, and samples per observation.
- Training-based predictions require documentation of training-data creation, training-sample size, and relevant training details, alongside the number of observations with AI-generated measurements.
- Researchers should report any aggregation or transformation of AI measurements and its level relative to the unit of analysis.
- Measurement-error correction depends on the estimand, the variable’s role as outcome, regressor, instrument, or otherwise, and the applicable validation regime.