Source-linked AI summary

Identification of and correction for publication bias

Isaiah Andrews, Maximilian Kasy

arXiv:1711.10527v1econ.EM

TL;DR

Selective publication can bias estimates and distort inference because study results are not equally likely to appear. The paper identifies conditional publication probabilities using systematic replications and meta-studies, then develops bias-corrected inference and applies it across several literatures. Applications find especially strong selectivity for significant replication results, while other literatures show different patterns and varying precision.

  • Problem

    Selective publication makes published estimates and inference potentially biased, but the conditional probability of publication as a function of study results is not directly known.

  • Method

    The paper identifies conditional publication probabilities from systematic replications and meta-studies, then constructs median-unbiased estimators and confidence sets for known selectivity.

  • Results

    Results significant at the 5% level were over 30 times more likely to be published than insignificant results in experimental economics and psychology replications, with similar conclusions from meta-studies.

  • Takeaways & Limitations

    The methods let readers draw valid inference from published research while taking the publication rule as given.

  • Takeaways & Limitations

    The framework assumes publication decisions are independent of true effects conditional on reported results; if this fails, the main selection interpretation is constrained.

Abstract

from arXiv · show

Some empirical results are more likely to be published than others. Such selective publication leads to biased estimates and distorted inference. This paper proposes two approaches for identifying the conditional probability of publication as a function of a study's results, the first based on systematic replication studies and the second based on meta-studies. For known conditional publication probabilities, we propose median-unbiased estimators and associated confidence sets that correct for selective publication. We apply our methods to recent large-scale replication studies in experimental economics and psychology, and to meta-studies of the effects of minimum wages and de-worming programs.

1 Introduction

Publication bias can make published estimates and confidence sets severely misleading because significant, confirmatory, or surprising findings are more likely to be published. The paper identifies publication selectivity, develops corrections for inference, and applies them across four empirical literatures.

  • Motivation: Publication selectivity can severely bias published estimates and confidence sets when editors, referees, or researchers favor particular findings.Favored findings may be statistically significant, confirm prior beliefs, or be surprising.
  • Contributions: The paper provides nonparametric identification results for conditional publication probabilities as functions of empirical study results.These results are presented as the paper’s first contribution of this kind.
  • Identification: Systematic replications identify publication probabilities from asymmetries in the joint distribution of initial and replication estimates, assuming selection depends only on the initial estimate.Absent selectivity, the joint distribution is symmetric.
  • Identification: Meta-studies identify publication probabilities from deviations between the predicted distributions of estimates with different variances, under a standard independence assumption.Absent selectivity, high-variance estimates should be noisier versions of low-variance estimates.
  • Identification: Publication probabilities are identified only up to scale, but this does not affect publication bias or size distortions.Multiplying all publication probabilities by a constant leaves the published-estimate distribution unchanged.
  • Inference: The paper proposes median-unbiased estimators and valid confidence sets when selectivity is known, plus Bonferroni corrections when the selection model is estimated.The methods support inference on individual-study parameters rather than only average effects across a literature.
  • Applications: Results significant at the 5% level were estimated to be over 30 times more likely to be published than insignificant results in experimental economics and psychology replications.The meta-study approach using only originally published results produced similar conclusions.
  • Applications: Negative significant minimum-wage effects were about 3 times more likely to be published than insignificant results, while deworming results appeared more likely included when nonsignificant.The deworming estimate has large standard errors, and no selectivity cannot be rejected.

2 Setting

The paper models publication as selective observation of latent study results, derives the resulting truncated likelihood, and uses replication or meta-study information to identify selectivity. Its illustrative example shows how significance-based selection distorts conventional estimates and confidence-interval coverage, while alternative observability and manipulation settings remain within the framework.

  • Data-generating process: Latent studies have true effects drawn from a population distribution, and their observed estimates are generated from a known conditional distribution.Different latent studies may estimate different true parameters; a common parameter is a degenerate special case.
  • Data-generating process: Studies are published with probability p(X∗), and researchers observe only the estimates from published studies.Publication decisions combine researcher and journal decisions, which the model does not disentangle.
  • Likelihood: The truncated likelihood reweights the distribution of published results by the publication probability.This likelihood is central for both identifying selectivity and conducting inference on true effects.
  • Extensions: Study-level covariates can condition the publication model when publication decisions and effect distributions differ across observable characteristics.Applications condition, for example, on publication journal and year of initial circulation.
  • Illustrative example: Under significance-based selection, published estimates tend to overestimate treatment-effect magnitude, while conventional confidence intervals under-cover for small true effects and over-cover for somewhat larger effects.Figure 1 plots median bias and true coverage under the specified publication rule.
  • Alternative observability: With censoring or observed unpublished working-paper results, the truncated likelihood remains a valid limited-information likelihood for identification and inference.Additional information may provide further insight without invalidating the results based on the truncated likelihood.
  • Manipulation of results: The framework assumes the conditional distribution of results given true effects is known, although many forms of p-hacking can be represented through publication selection.More general selection can depend on both the empirical result and the true effect.

3 Identifying selection

The paper identifies conditional publication probabilities up to scale using asymmetries in systematic replication data or deviations from variance-based predictions in meta-studies. Extensions allow selection to depend on true effects and additional variables, while retaining identification of key objects under stated assumptions.

  • Systematic replication studies: Systematic replications identify publication probabilities from asymmetries in the joint distribution of original and replication estimates, assuming selection acts only on the original estimate.Latent original and replication estimates are symmetric absent selectivity; observed deviations from symmetry reveal the selection function up to scale.
  • Meta-studies: Meta-studies identify publication probabilities from deviations between the distributions of estimates with different standard deviations, assuming standard deviation is independent of true effects.Without selectivity, higher-variance estimates are noisier versions of lower-variance estimates; departures from this prediction identify selection.
  • Additional identified objects: Under normality, replication data can also identify the latent distribution of true effects, while symmetric meta-study models identify the distribution of absolute true effects.These results support recovering population-level objects alongside the selection function in the corresponding settings.
  • Generalized selection: Allowing publication to depend on true effects or other manuscript findings changes the selection model, but the generalized setup can still identify fX|Θ for bias-corrected inference.Additional variables may include alternative specifications and unreported robustness checks.
  • Relation to existing approaches: Meta-regression coefficients provide tests for no publication bias through β0 = 0 and γ1 = 0, although some forms of selectivity have no power against these tests.Absent publication bias, β1 and γ0 recover the average true effect in the latent-study population.

4 Corrected inference

The paper corrects inference after selective publication by inverting the publication-adjusted distribution of reported estimates. It develops median-unbiased estimators and confidence sets, including procedures that account for estimated selection probabilities.

  • Correction principle: Selective publication reweights the distribution of reported estimates, so valid inference requires correcting the published-result distribution.The correction is defined using the cdf and density of published results conditional on the true effect.
  • Median-unbiased estimation: When the published-result cdf is strictly decreasing in the parameter, inverting it yields a quantile-unbiased estimator.Under continuity and tail conditions, the estimator exists uniquely and is continuous and strictly increasing in the observed result.
  • Inference guarantees: The resulting median-unbiased estimator and confidence set fully correct bias and coverage distortions induced by selective publication.The frequentist construction targets inference for a study-specific scalar parameter rather than only an average effect across a literature.
  • Illustrative example: In the illustrative normal model, the corrected interval has exact coverage, is narrower for small X, wider for moderate X, and essentially matches the usual interval for X ≥5.The corrected estimator lies below the usual estimator for small positive X, with the difference eventually decreasing.
  • Interpretation: The corrections are intended to support valid inference under a given publication rule, not to justify changing publication standards.Different critical values would require corresponding adjustments to the correction.
  • Estimated selection probabilities: With an estimated selection model, Bonferroni-corrected confidence intervals account for estimation error and achieve coverage of at least 1 − α in large samples when the model is correctly specified.Median-unbiased estimation is more difficult when publication probabilities are estimated with error, whereas valid confidence sets remain straightforward.

5 Applications

The applications find substantial publication selectivity in experimental economics and psychology, while evidence is weaker and more mixed for minimum-wage and de-worming literatures. Bias corrections often reduce estimated effects and markedly change statistical significance.

  • 5.1 Economics laboratory experiments: Adjusting Camerer et al. estimates for selection increased adjusted confidence sets containing zero from 2 to 12 of 18 studies.Adjusted estimates tracked replication estimates fairly well but were smaller than original estimates in many cases.
  • 5.2 Psychology laboratory experiments: Psychology data show publication-probability jumps near the 5% and possibly 10% significance thresholds, confirmed by both replication and meta-study estimates.Results significant at 5% were over 100 times more likely to be published than results insignificant at 10%, and nearly five times more likely than results significant at 10% but not 5%.
  • 5.2 Psychology laboratory experiments: Psychology corrections reduced the number of original 95% confidence sets excluding zero from 62 of 73 to 28 adjusted sets.Adjusted estimates tracked replication estimates for small original z-statistics, with a worse fit for larger original z-statistics.
  • 5.3 Effect of minimum wage on employment: Minimum-wage estimates suggest selection favoring negative employment effects, but initial meta-regressions cannot reject no selection and the estimated sign dependence is noisy.Insignificant results were about 30% as likely to be published as significant results finding a negative employment effect; the selection-adjusted mean effect was about half the naive estimate.
  • 5.4 Effect of de-worming programs: The de-worming application suggests insignificant results may be more likely to appear than significant ones, but large standard errors and specification sensitivity make the finding difficult to interpret strongly.With only 22 estimates, the apparent density jump at zero should not be interpreted too strongly, and regression checks do not reject no selection.

6 Conclusion

The conclusion presents identification, bias-correction, and empirical-application contributions, and recommends assessing selectivity when synthesizing or reading empirical evidence. It also cautions that publication rules may optimally favor surprising findings and that corrected critical values should not become publication standards.

  • The paper contributes nonparametric identification of conditional publication probabilities, bias-corrected estimators and confidence sets, and applications documenting varied selectivity across literatures.
  • Researchers conducting meta-analyses may assess selectivity and apply corrections to individual estimates, tests, and confidence sets.The authors provide code implementing the proposed methods for a flexible family of selection models.
  • Readers should adjust published magnitudes downward for intermediate z-statistics, especially around 2, when publication rises at the 5% threshold but otherwise depends little on findings.The paper states that estimates close to zero or far above the threshold may not require the same adjustment.
  • Corrected critical values should not be adopted as publication standards because doing so would create an arms race of selectivity and invalidate the stricter thresholds.
  • Optimal publication need not be non-selective: when publication informs policy under limited attention, welfare may favor surprising findings that update prior beliefs most.

Supplement to the paper

The supplement supplies proofs and supplementary results for the paper, including analyses of meta-regression behavior, empirical likelihoods, expanded selection models, and additional technical material.

  • The supplement contains proofs for the main-text results and supplementary analyses of meta-regression coefficients and empirical-application likelihoods.
  • It also develops a model allowing selection on both the normalized estimate and a latent variable, alongside further supplementary results.

A Proofs

The proofs establish nonparametric identification of publication probabilities up to scale under replication and meta-study designs, then derive identification of latent-effect distributions and model parameters under additional assumptions.

  • Replication identification: Publication probabilities are identified up to scale from asymmetries in the joint distribution of original and replication estimates.The replication argument uses symmetry that would hold absent selective publication.
  • Latent distributions: Deconvolution recovers the latent-effect distribution conditional on observed estimates, after which Bayes’ rule identifies the publication function.The proof treats deconvolution as the step connecting observable densities to latent effects.
  • Parameter identification: The differential-equation proof is overidentified, yielding infinitely many restrictions on finitely many free parameters.Higher-order derivatives or evaluations at different z values provide additional restrictions.

B Interpretation of meta-regression coefficients

Meta-regression coefficients are generally not selection-corrected effect estimates because selective publication makes the conditional mean nonlinear in precision. The appendix also describes likelihood-based specifications and normalization choices for estimating publication rules.

  • Interpretation: Meta-regression slopes cannot generally be interpreted as selection-corrected estimates of average effects when selectivity is present.The conditional expectation of the statistic given precision is generally nonlinear.
  • Interpretation: Under strict selection on significant positive results, the conditional mean combines the scaled effect with an inverse Mills ratio term.Both the best linear predictor’s slope and intercept depend on the distribution of standard errors.
  • Estimation: The estimation framework uses step-function publication rules with a normalization that interprets coefficients relative to a reference category.Sign-normalized applications impose symmetry and estimate using absolute or sign-adjusted statistics.
  • Specification checks: Replication data can identify models in which publication probabilities depend on both observed results and latent effects, enabling specification checks beyond the baseline model.A parametric example integrates over an unobserved independent estimate to obtain publication probabilities of the form p(Z∗, Θ∗).
  • Application: In the economics application, journal-specific estimates are noisy and do not reject a common publication rule across journals.The evidence includes one significant and one insignificant QJE result versus fifteen significant and one insignificant AER result.

F.2 Additional results for psychology laboratory experiments

Additional psychology analyses examine restricted samples, approved replication protocols, and journal-specific rules. The restricted-degree-of-freedom results remain broadly similar, while journal differences are not statistically significant.

  • Denominator degrees of freedom: 52 observations with denominator degrees of freedom at least 30 produce results broadly similar to those from the full psychology replication data.The restriction addresses concerns that small-degree-of-freedom t distributions have heavier tails than normal distributions.
  • Sample construction: Screening on replication degrees of freedom could introduce additional selection on original-study results because replication sample sizes depend on initial results.The analysis therefore screens only on denominator degrees of freedom in the original study.
  • Publication rule varies by journal: Journal-specific publication probabilities show roughly the same pattern as the baseline estimates, but none of the journal differences is significant.The joint-test p-values are .78 for replication estimates and .84 for meta-study estimates.
  • Publication rule varies by journal: The journal-specific model does not reject the null that all journals use the same publication rule.The model allows publication probabilities to vary across Psychological Science, JPSP, and JLMC.

F.3 Additional results for minimum wage meta-study

Additional minimum-wage and deworming analyses test sample, time-trend, specification, and distributional robustness. Minimum-wage results are stable across published-only samples and over time, whereas deworming estimates are sensitive to asymmetric specifications and sparse negative evidence.

  • Published studies: 705 estimates from 31 published studies yield estimates broadly similar to those from the full minimum-wage sample.The published-only analysis clusters standard errors at the study level.
  • Time trends: A joint test with p-value 0.7 is consistent with publication rules being constant over time.The time-trend specification measures years relative to 2013 and uses a logistic function to keep probabilities between zero and one.
  • Deworming specification: A flexible deworming specification suggests strong selectivity against negative estimates, particularly negative and significant estimates.Only one negative and statistically significant estimate appears in the sample, making conventional large-sample approximations highly suspect.
  • Deworming specification: A restricted asymmetric deworming specification estimates positive effects as ten times more likely to be published than negative effects.The specification choice was driven by the preceding results, so it constitutes specification search and raises concerns about conventional asymptotic approximations.
  • Distributional robustness: Moment-based estimators requiring a functional form only for publication probabilities generally produce less precise conclusions but preserve the main findings.These estimators leave the distribution of true effects fully nonparametric.

G.2.1 Economics laboratory experiments

The appendix applies the paper’s selection framework to laboratory experiments in economics and psychology, minimum-wage studies, and deworming research. Identification is informative in some settings but weak in others, while selective publication distorts conventional inference.

  • Economics laboratory experiments: Replication-based estimates for economics and psychology model publication selection using Camerer et al. (2016) and Open Science Collaboration (2015) data.The appendix also compares replication-based moments with meta-study moments based on initially published estimates.
  • Psychology laboratory experiments: Identification of βp2 is weak in the psychology application, although replication-based confidence sets restrict βp,1 to small values.Meta-study confidence sets allow a wide range of values for either parameter.
  • Minimum wages: Meta-study estimates for minimum-wage studies favor publication of significant negative employment effects over insignificant results.Estimated selection favoring positive significant effects is noisy and remains consistent with selection on statistical significance alone.
  • Deworming: The deworming application produces a point estimate that differs from the baseline specification, but its robust confidence set is unbounded above.This indicates weak identification and makes the point estimate likely unreliable.
  • Multivariate inference: Selective publication creates substantial bias and coverage distortions for conventional estimators in the multivariate difference-in-differences example.The corrected estimator and test depend on both the difference-in-differences statistic G and the pretest statistic L.

I.3 Optimal quantile-unbiased estimates

The paper extends median-unbiased inference to a multivariate setting with nuisance parameters. The resulting estimator is quantile-unbiased and optimally concentrated, while corrected estimates and tests can depend on auxiliary statistics used in the selection process.

  • Multivariate extension: The multivariate extension derives optimal quantile-unbiased estimators and equal-tailed confidence sets using the conditional distribution of G given W.The construction treats the remaining elements of Θ as nuisance parameters.
  • Optimal quantile-unbiased estimates: The proposed estimator chooses γ so the observed G lies at the α quantile of its conditional distribution given W.This construction delivers a quantile-unbiased estimator under the stated regularity conditions.
  • Optimal quantile-unbiased estimates: The estimator satisfies Pr{ˆγα(X) ≤ γ | Θ=(γ,ω)} = α for all γ and ω.This quantile-unbiasedness property supports equal-tailed confidence intervals.
  • Optimal quantile-unbiased estimates: The estimator is uniformly most concentrated among level-α quantile-unbiased estimators for loss functions minimized at the true parameter.The theorem compares expected loss against any other estimator in that class.
  • Difference in differences example: In the difference-in-differences example, the corrected estimator depends on both G and L rather than only the target parameter γ.The corresponding 5% test’s rejection region also depends on both statistics.

J Bayesian inference

The appendix shows that selective publication affects Bayesian inference differently under two classes of priors. With unrelated-parameters priors, selection does not alter the posterior; with common-parameters priors, the posterior uses the truncated likelihood, and publication may be justified under a policy objective.

  • Two classes of priors: The unrelated-parameters prior treats each latent study as targeting a different parameter, whereas the common-parameters prior assumes all studies target the same parameter.These prior classes determine how information and selection enter posterior inference.
  • Bayesian posterior: Under unrelated-parameters priors, the posterior after observing a published estimate is the same as without selection.The publication rule has no effect on the posterior in this case.
  • Bayesian posterior: Under common-parameters priors, the posterior updates the marginal prior using the truncated likelihood implied by selective publication.Selection changes the distribution of true effects among published studies under this prior.
  • Optimal publication: When publication decisions are inputs to policy decisions and publication capacity is constrained, selective publication may be justified by the journal’s objective.The appendix develops this argument in a stylized poverty-reduction setting.
  • Optimal publication: The optimal publication rule can select positive results significant relative to a critical value xc, while still generating publication bias requiring corrected inference.The critical value may vary across studies with treatment costs, population sizes, and outcome variances.
Loading 1711.10527v1…