Source-linked AI summary

Measuring reproducibility of high-throughput experiments

Qunhua Li, James B. Brown, Haiyan Huang, Peter J. Bickel

arXiv:1110.4705v1stat.AP

TL;DR

High-throughput experiments require reproducibility measures because individual assay outputs can be substantially variable. The paper introduces a rank-based correspondence curve and copula mixture model, derives the irreproducible discovery rate, and demonstrates the approach through simulations and ChIP-seq data.

  • Problem

    High-throughput assay outputs can be substantially variable, creating a need for objective metrics that assess reproducibility and support reliable discoveries.

  • Method

    The paper models rank consistency across replicate scores with a copula mixture, using a correspondence curve and IDR-based selection criterion without assuming score scales or marginal distributions.

  • Results

    Simulated and real data illustrate that the method can assess reproducibility independently of prespecified thresholds, identify biologically relevant thresholds, improve signal identification accuracy, and identify suboptimal results.

  • Takeaways & Limitations

    The reproducibility criterion can select signals from continuous rankings, complement individual-replicate thresholds, and support comparable standards across data sources and platforms.

  • Takeaways & Limitations

    Systematic bias that creates false association can make the procedure underestimate the empirical false discovery rate, so replicate independence requires care.

Abstract

from arXiv · show

Reproducibility is essential to reliable scientific discovery in high-throughput experiments. In this work we propose a unified approach to measure the reproducibility of findings identified from replicate experiments and identify putative discoveries using reproducibility. Unlike the usual scalar measures of reproducibility, our approach creates a curve, which quantitatively assesses when the findings are no longer consistent across replicates. Our curve is fitted by a copula mixture model, from which we derive a quantitative reproducibility score, which we call the "irreproducible discovery rate" (IDR) analogous to the FDR. This score can be computed at each set of paired replicate ranks and permits the principled setting of thresholds both for assessing reproducibility and combining replicates. Since our approach permits an arbitrary scale for each replicate, it provides useful descriptive measures in a wide variety of situations to be explored. We study the performance of the algorithm using simulations and give a heuristic analysis of its theoretical properties. We demonstrate the effectiveness of our method in a ChIP-seq experiment.

1. Introduction.

The paper addresses reproducibility concerns in high-throughput experiments by replacing threshold-dependent assessment with rank-based consistency analysis and a copula-mixture reproducibility criterion. It applies the approach to simulations and ChIP-seq data to assess and select reliable signals.

  • High-throughput assays can produce substantially variable outputs, making reproducibility metrics important for reliable discoveries and monitoring data-generating procedures.
  • Spearman rank correlation commonly assesses agreement among signals passing prespecified significance thresholds, but the paper takes an alternative threshold-free approach.
  • The method visualizes where signal ranks lose consistency across replicates as significance decreases, then models reproducible and irreproducible signal groups with a copula mixture.
  • The irreproducible discovery rate (IDR) estimates average irreproducibility among selected signals and supports ranking and selection procedures analogous to multiple-testing criteria.
  • Because it makes no parametric assumptions about marginal score distributions, the approach applies to probabilistic and heuristic scores and can set thresholds when scoring thresholds are difficult to determine.
  • The paper evaluates the method through simulations and a ChIP-seq application comparing reproducibility and signal reliability across algorithms.

2. Statistical methods.

The statistical framework treats replicate scores for putative signals as jointly informative while allowing unknown, replicate-specific score scales. It uses rank consistency to distinguish more reproducible signals from spurious ones and derive a reproducibility-based selection criterion.

  • The data contain many putative signals measured on few replicates, each assigned a distinct score reflecting evidence strength through heuristic or probabilistic methods.
  • Score distributions and scales may vary across replicates, while the scores are assumed to preserve the relative ranking of signals.
  • For two replicates, the method jointly considers each signal’s score vector to measure reproducibility and select reliable signals; extensions to more replicates are possible.
  • Genuine signals are expected to have higher and more consistent ranks across replicates, whereas consistency may degrade toward the noise level.
  • The framework first visualizes changing rank association without model assumptions, then uses a model-based approach to quantify association heterogeneity and guide threshold selection.

2.1. Displaying the change of association.

The correspondence curve displays how agreement between replicate rankings changes as progressively less significant signals are included. Its shape can reveal transitions from strongly associated signals to weakly associated or noisy signals.

  • The correspondence curve is a rank-based graphical tool designed to display changing association and localize where consistency breaks down.
  • Ψn(t,v) measures the proportion of signals ranked in the upper t% of one replicate and upper v% of the other; the analysis focuses on the symmetric case t = v.
  • The curves depend only on replicate ranks and remain invariant under location and scale transformations, making them scale free.
  • Perfect rank correlation produces a diagonal Ψn curve with slope 1, whereas independence produces a parabolic Ψn(t) = t^2 curve.
  • When the top t0n signals are perfectly correlated and the remainder independent, the curve changes form at t0, marking the transition between reproducible and weakly associated groups.
  • In an idealized example, the curve transition occurs at 50%, matching the boundary between perfectly corresponding top-ranked signals and independent lower-ranked observations.
  • The derivative curve can make transitions more visible because it separates zero-slope and positive-slope regions, helping localize less-sharp breakdowns.
  • A real salmon dataset showed a transition near t = 0.5, suggesting roughly equal strongly and weakly associated groups and agreeing with earlier speculation.

2.2. Inferring the reproducibility of signals.

The paper models replicate scores with a semiparametric copula mixture to quantify dependence and infer which signals are reproducible. The model accommodates unknown, varying marginal score scales and assigns signals reproducibility-related classification probabilities.

  • Assumptions and estimation: The model assumes independent and identically distributed bivariate observations, although genome-wide profiling signals are often dependent.The authors retain the method’s descriptive and graphical value because their focus is on first-order effects.
  • Copula mixture model: The model treats signals as genuine or spurious groups that differ in reproducibility, significance, and dependence between replicate scores.Genuine signals are modeled as the more reproducible group, while spurious signals form the less reproducible group.
  • Copula mixture model: The copula mixture model separates dependence between replicates from their unknown marginal distributions.This semiparametric construction is useful when score distributions and scales vary across replicates or datasets.
  • Inference: For two replicates, each signal is represented by a bivariate score vector, and the method estimates posterior probabilities of membership in the irreproducible group.The posterior probabilities are obtained by estimating model parameters and substituting them into the classification formulas.
  • Assumptions and estimation: The estimation algorithm is practically convergent, but its theoretical convergence has not yet been established.The authors report a heuristic asymptotic-efficiency argument for a limit point near the true parameter and identify a convergent modification under investigation.

2.3. Irreproducible identification rate.

The paper defines reproducibility measures by analogy with false-discovery quantities, classifying replicate observations through a mixture model and ranking them by estimated irreproducibility. These measures support controlled selection and can also combine replicate scores when those scores are p-values.

  • Mixture-based classification: The model’s two mixture classes represent irreproducible and reproducible measurements rather than null and nonnull hypotheses.This distinction preserves the analogy to multiple testing while assigning the classes a different scientific interpretation.
  • Local IDR: The local irreproducible discovery rate is the posterior probability that a signal is not reproducible across a pair of replicates.It is estimated from the copula mixture model.
  • Global IDR: The irreproducible discovery rate, or IDR, summarizes the expected proportion of irreproducible signals among selected observations.Signals are selected using local IDR values and a desired control level, analogously to an adaptive multiple-testing procedure.
  • Selection procedure: The procedure re-ranks identifications by the likelihood ratio of the joint replicate distribution rather than by either replicate’s original significance scores.The resulting ranking can differ from the rankings on the individual replicates.
  • Score flexibility: The method accepts arbitrary continuous scores instead of requiring p-values, while p-value inputs allow it to function as a p-value combination method.The paper compares this use with two commonly used p-value combinations in simulations.

3. Simulation studies.

Simulation studies examine correspondence curves, model estimation, calibration, and discriminative power under varied signal and noise configurations. The method is reasonably calibrated in S1–S3 and identifies more true signals than the compared methods at a given number of false calls.

  • 3.1. Illustration of correspondence curves.: Homogeneous association lacks the characteristic curve transition, whereas mixtures with independent lower-ranked and positively correlated higher-ranked components produce it.The transition is proposed as an indicator of changing association across ranked signals.
  • 3.2. Simulation studies.: The simulations assess classification accuracy, replicate-information benefits, robustness to model-assumption violations, and comparisons with existing p-value combination methods.Four simulations use parameters estimated from a ChIP-seq data set, with 10,000 pairs per data set in the simulation studies.
  • 3.2.1. Simulation setup.: Signals are generated from a normal mixture model, transformed to a Student’s t distribution, and converted into p-values used as significance scores.The resulting p-values are intentionally not calibrated but reflect relative evidence strength.
  • 3.2.1. Simulation setup.: In S1–S3, estimated parameters are close to their true values, except that σ1 is underestimated when the true-signal proportion is π1 = 0.05.The small-signal-proportion case is described as difficult to distinguish from a single component.
  • 3.2.2. Comparison of discriminative power.: The method is reasonably calibrated in S1, S2, and S3, while reproducible noise in S4 causes slight overestimation of real-signal proportion and correlation and underestimation of empirical FDR.Original scores and other combination methods are overly conservative in estimated FDR across simulations.
  • 3.2.2. Comparison of discriminative power.: At a given number of false calls, the method consistently identifies more true signals than the original significance score, Fisher’s method, and Stouffer’s method.This advantage persists when reproducible artifacts are present or genuine signals are rare.

4. Applications on real data.

The real-data application compares nine ChIP-seq peak callers using correspondence profiles, copula-mixture inference, IDR-based selection, and motif overlap. Peakseq, MACS, and SPP are consistently identified as the most reproducible, and their selected peaks show the highest motif-occurrence rates.

  • 4. Applications on real data.: The ENCODE application compares reproducibility across multiple peak-calling algorithms, supports uniform selection, and helps identify poor-quality results.The data come from two biological replicates of a CTCF ChIP-seq experiment.
  • 4.1. Comparing the reproducibility of multiple peak callers for ChIP-seq experiments.: Peak widths are normalized to 40 bp centered at reported summits because wider peaks are more likely to overlap true binding sites by chance.SPP and SISSRS originally produce fixed widths of 100 bp and 40 bp, respectively; other callers have median widths of 130–760 bp.
  • 4.2.1. Correspondence profiles.: Five callers show the characteristic transition from strong association to near independence, with Peakseq, MACS, and SPP showing the highest reproducibility by transition position.Erange, Cisgenome, Quest, and SISSRS show less clear transitions and substantially fewer reproducible peaks.
  • 4.2.2. Inference from the copula mixture model.: The one-component model fits SISSRS, Quest, and Cisgenome better, whereas the two-component model fits the other peak callers, consistent with the correspondence profiles.Model selection uses a likelihood-ratio test with parametric-bootstrap p-values.
  • 4.2.2. Inference from the copula mixture model.: Peakseq, MACS, and SPP have about 3% irreproducible peaks among the top 25,000, while most other callers reach higher IDR before their top 10,000 peaks.The ordering by peaks identified before reaching 5% IDR is Peakseq, MACS, SPP, Fseq, then the others.
  • 4.2.3. Evaluating the biological relevance of the reproducibility assessment.: For two-component callers, motif occurrences plateau near the 5% IDR selection point, whereas one-component callers continue increasing at default thresholds.The latter pattern indicates that default thresholds for those callers are likely overly stringent for this data set.
  • 4.2.3. Evaluating the biological relevance of the reproducibility assessment.: SPP, Peakseq, and MACS have both the highest reproducibility and the highest motif-occurrence rates among the compared algorithms.The agreement illustrates the potential of the method as a quality measure.

5. Discussion.

The method measures reproducibility and sets selection thresholds without relying on prespecified or platform-dependent score thresholds. Its interpretation requires attention to replicate independence and to the data-specific behavior of analytical methods.

  • Simulated and real data illustrate reproducibility assessment independent of prespecified thresholds, biologically relevant threshold selection, improved signal identification, and detection of suboptimal results.
  • The method applies to continuous ranking systems regardless of score scale and supports principled selection for heuristic scores.
  • Consistency between replicates can provide uniform selection standards across multiple data sources and facilitate inter-platform consistency assessment.
  • Reproducibility is necessary but not sufficient for accuracy: systematic bias in replicates can make derived thresholds underestimate the empirical false discovery rate.
  • Reproducibility assessments of analytical outputs reflect both the method and samples, so the example is specific to the studied data set rather than a general conclusion.
  • An R package and the estimation algorithm are provided as supplemental implementation resources.

SUPPLEMENTARY MATERIAL

The supplementary materials provide algorithmic, theoretical, and correspondence-curve details supporting the copula mixture model and its estimation.

  • The supplement describes parameter estimation, asymptotic estimator efficiency, correspondence-curve properties, and an extension of the model.
Loading 1110.4705v1…