Source-linked AI summary

Data Exploration, Quality Control and Testing in Single-Cell qPCR-Based Gene Expression Experiments

Andrew McDavid, Greg Finak, Pratip K. Chattopadyay, Maria Dominguez, Laurie Lamoreaux, Steven S. Ma, Mario Roederer, Raphael Gottardo

arXiv:1210.1226v1stat.APq-bio.QM

TL;DR

Single-cell qPCR data present analytical challenges because biological and technical variability coexist and expression can be dichotomously absent. The paper develops quality-control procedures and a discrete/continuous statistical model with a combined likelihood-ratio test. The framework addresses both expression frequency and conditional expression while filtering technical artifacts.

  • Problem

    Single-cell qPCR requires specialized analysis because biological and technical variability can be comparable and measurements may be dichotomously absent.

  • Method

    The paper combines outlier filtering with a mixture model representing expression as a zero point mass plus a log-normal distribution, and derives a joint likelihood-ratio test.

  • Results

    The framework’s filtering removes technical artifacts, while its combined test incorporates changes in both expression frequency and conditional mean expression.

  • Takeaways & Limitations

    Both the discrete zero-inflated component and the continuous expression component are meaningful for quality control and differential-expression detection.

  • Takeaways & Limitations

    Filtering defaults are conservative and may need tuning to balance eliminating technical error against missing biological heterogeneity.

Abstract

from arXiv · show

Cell populations are never truly homogeneous; individual cells exist in biochemical states that define functional differences between them. New technology based on microfluidic arrays combined with multiplexed quantitative polymerase chain reactions (qPCR) now enables high-throughput single-cell gene expression measurement, allowing assessment of cellular heterogeneity. However very little analytic tools have been developed specifically for the statistical and analytical challenges of single-cell qPCR data. We present a statistical framework for the exploration, quality control, and analysis of single-cell gene expression data from microfluidic arrays. We assess accuracy and within-sample heterogeneity of single-cell expression and develop quality control criteria to filter unreliable cell measurements. We propose a statistical model accounting for the fact that genes at the single-cell level can be on (and for which a continuous expression measure is recorded) or dichotomously off (and the recorded expression is zero). Based on this model, we derive a combined likelihood-ratio test for differential expression that incorporates both the discrete and continuous components. Using an experiment that examines treatment-specific changes in expression, we show that this combined test is more powerful than either the continuous or dichotomous component in isolation, or a t-test on the zero-inflated data. While developed for measurements from a specific platform (Fluidigm), these tools are generalizable to other multi-parametric measures over large numbers of events.

1 Introduction

Single-cell qPCR reveals heterogeneity that population-level assays obscure, but its biological and technical variability and zero-inflated measurements require specialized analysis. The paper introduces quality-control filtering and a discrete/continuous model for differential-expression testing.

  • Single-cell measurements can expose heterogeneous gene expression within populations that appear homogeneous under established surface-marker sorting.
  • Biological and technical variability can be similarly large in single-cell data, complicating their separation.
  • Zero-valued measurements may reflect biologically absent expression, so single-cell data are not always continuous.
  • Traditional normalization schemes are not directly applicable because the individual cell is the normalization unit and expression is dichotomous.
  • The framework filters outlying single-cell measurements and uses concordance with 100-cell measurements to assess normalization and technical artifacts.
  • A mixture of a zero point mass and log-normal distribution supports a likelihood-ratio test combining changes in mean expression and the percentage of expressed cells.

2 Methods

The paper develops a Fluidigm single-cell qPCR framework for notation, expression modeling, quality control, filtering, and differential-expression testing. Its model and likelihood-ratio test jointly represent whether genes are expressed and their positive expression levels.

  • Data sets and notations: The Fluidigm assay measures 96 genes across 96 cells after multiplexed pre-amplification and gene-specific quantification.The study also compares single-cell measurements with 100-cell aggregates.
  • Data sets and notations: Undetected genes are treated as unexpressed, with expression threshold set to −∞ and mRNA abundance set to zero.This interpretation is supported by detected genes typically having cycle thresholds well below the maximum and by improved concordance with 100-cell measurements.
  • A model for single cell expression: Single-cell expression is modeled with a discrete on/off component and a log-normal distribution for positive expression values.The model uses the expression frequency πj and conditional mean and variance parameters for expressed cells.
  • Testing for ET differences between experimental groups: Differential expression between two biological units is tested by comparing both changes in expression frequency π and conditional mean expression µ.The likelihood-ratio test uses null and alternative maximum-likelihood estimates for the model parameters.
  • Testing for ET differences between experimental groups: The combined likelihood-ratio statistic decomposes into Bernoulli and log-normal components, combining discrete and continuous information.The authors report that this combined test is more powerful than either component alone.

3 Results

The results validate treating undetected single-cell measurements as zeros, show that filtering improves concordance with hundred-cell measurements, and demonstrate advantages of the combined likelihood-ratio test for differential expression.

  • Concordance and filtering: Positive expression values agree with a log-normal model, while undetected genes are consistent with null or negligible RNA abundance.The positive expression distributions match the postulated normal distribution on the log scale.
  • Concordance and filtering: Treating undetected wells as zeros produces good concordance with hundred-cell measurements, unlike treating them as missing values.The comparison uses an in-silico average of 100 single-cell measurements against an in-vitro hundred-cell aggregate.
  • Concordance and filtering: Filtering parameters tz = 9 and tζ = 9 achieve the best WSS reduction across three data sets and move average per-unit estimates toward the diagonal.The WSS is computed on the log2(y + 1) scale to reduce the effect of extreme outliers.
  • Normalization and housekeeping genes: Housekeeping genes provide little normalization utility: GAPDH and POLR2A have R2 = .027 in data set A, and normalization could mask cellular artifacts.After filtering, their regression has R2 = .017, while their apparent correlation is concentrated in cells suspected of technical error.
  • Differential expression testing: At an FDR of 1%, the combined test detects more than 20 additional gene×unit changes across four stimulations than the comparison tests.Across a wide range of FDR values, the combined likelihood test produces the greatest number of discoveries compared with Bernoulli-only, normal-theory-only, and t-tests.
  • Differential expression testing: The combined likelihood-ratio test remains robust across gene-expression frequencies by combining evidence from Bernoulli and normal components.Figure 4 plots −log10 p-values against expression frequency πj for the Bernoulli, normal, and combined tests.
  • Differential expression testing: At an FDR of 1%, 65 genes are detected in at least one individual, with clustered up-regulated and down-regulated modules and substantial response variability.The heatmap uses signed log10 p-values, clustering selected genes by rows and individuals by stimulation-ordered columns.

4 Conclusion

The paper presents a framework for exploring, quality-controlling, and testing single-cell gene-expression data, combining discrete and continuous evidence in a likelihood-ratio test. It positions these tools as a coherent framework for researchers using this nascent technology.

  • The framework combines the discrete zero-inflated and continuous portions of single-cell expression data in a likelihood-ratio test.Both components can provide meaningful evidence for detecting outliers and biologically interesting expression changes.
  • The conclusion emphasizes a framework for data exploration, quality control, and differential-expression testing in single-cell assays.
  • The selected-gene heatmap organizes genes by rows and individuals by columns, with column colors indicating antigen stimulation.
  • The proposed default outlier-filtering parameters are conservative and may require tuning to balance technical-error removal against missing biological heterogeneity.
  • The likelihood-ratio test may not apply when traits are not blocked within individuals, although mixed-effects extensions could address between-individual variability.
  • The approach is intended to support researchers using single-cell assays as the technology becomes more routine.The paper frames effective statistical methods as increasingly important for this developing technology.
Loading 1210.1226v1…