Source-linked AI summary
Data Exploration, Quality Control and Testing in Single-Cell qPCR-Based Gene Expression Experiments
Andrew McDavid, Greg Finak, Pratip K. Chattopadyay, Maria Dominguez, Laurie Lamoreaux, Steven S. Ma, Mario Roederer, Raphael Gottardo
TL;DR
Single-cell qPCR data present analytical challenges because biological and technical variability coexist and expression can be dichotomously absent. The paper develops quality-control procedures and a discrete/continuous statistical model with a combined likelihood-ratio test. The framework addresses both expression frequency and conditional expression while filtering technical artifacts.
Problem
Single-cell qPCR requires specialized analysis because biological and technical variability can be comparable and measurements may be dichotomously absent.
Method
The paper combines outlier filtering with a mixture model representing expression as a zero point mass plus a log-normal distribution, and derives a joint likelihood-ratio test.
Results
The framework’s filtering removes technical artifacts, while its combined test incorporates changes in both expression frequency and conditional mean expression.
Takeaways & Limitations
Both the discrete zero-inflated component and the continuous expression component are meaningful for quality control and differential-expression detection.
Takeaways & Limitations
Filtering defaults are conservative and may need tuning to balance eliminating technical error against missing biological heterogeneity.
Abstract
from arXiv · showhide
Cell populations are never truly homogeneous; individual cells exist in biochemical states that define functional differences between them. New technology based on microfluidic arrays combined with multiplexed quantitative polymerase chain reactions (qPCR) now enables high-throughput single-cell gene expression measurement, allowing assessment of cellular heterogeneity. However very little analytic tools have been developed specifically for the statistical and analytical challenges of single-cell qPCR data. We present a statistical framework for the exploration, quality control, and analysis of single-cell gene expression data from microfluidic arrays. We assess accuracy and within-sample heterogeneity of single-cell expression and develop quality control criteria to filter unreliable cell measurements. We propose a statistical model accounting for the fact that genes at the single-cell level can be on (and for which a continuous expression measure is recorded) or dichotomously off (and the recorded expression is zero). Based on this model, we derive a combined likelihood-ratio test for differential expression that incorporates both the discrete and continuous components. Using an experiment that examines treatment-specific changes in expression, we show that this combined test is more powerful than either the continuous or dichotomous component in isolation, or a t-test on the zero-inflated data. While developed for measurements from a specific platform (Fluidigm), these tools are generalizable to other multi-parametric measures over large numbers of events.
1 Introduction
Single-cell qPCR reveals heterogeneity that population-level assays obscure, but its biological and technical variability and zero-inflated measurements require specialized analysis. The paper introduces quality-control filtering and a discrete/continuous model for differential-expression testing.
- Single-cell measurements can expose heterogeneous gene expression within populations that appear homogeneous under established surface-marker sorting.
- Biological and technical variability can be similarly large in single-cell data, complicating their separation.
- Zero-valued measurements may reflect biologically absent expression, so single-cell data are not always continuous.
- Traditional normalization schemes are not directly applicable because the individual cell is the normalization unit and expression is dichotomous.
- The framework filters outlying single-cell measurements and uses concordance with 100-cell measurements to assess normalization and technical artifacts.
- A mixture of a zero point mass and log-normal distribution supports a likelihood-ratio test combining changes in mean expression and the percentage of expressed cells.
2 Methods
The paper develops a Fluidigm single-cell qPCR framework for notation, expression modeling, quality control, filtering, and differential-expression testing. Its model and likelihood-ratio test jointly represent whether genes are expressed and their positive expression levels.
- Data sets and notations: The Fluidigm assay measures 96 genes across 96 cells after multiplexed pre-amplification and gene-specific quantification.The study also compares single-cell measurements with 100-cell aggregates.
- Data sets and notations: Undetected genes are treated as unexpressed, with expression threshold set to −∞ and mRNA abundance set to zero.This interpretation is supported by detected genes typically having cycle thresholds well below the maximum and by improved concordance with 100-cell measurements.
- A model for single cell expression: Single-cell expression is modeled with a discrete on/off component and a log-normal distribution for positive expression values.The model uses the expression frequency πj and conditional mean and variance parameters for expressed cells.
- Testing for ET differences between experimental groups: Differential expression between two biological units is tested by comparing both changes in expression frequency π and conditional mean expression µ.The likelihood-ratio test uses null and alternative maximum-likelihood estimates for the model parameters.
- Testing for ET differences between experimental groups: The combined likelihood-ratio statistic decomposes into Bernoulli and log-normal components, combining discrete and continuous information.The authors report that this combined test is more powerful than either component alone.
3 Results
The results validate treating undetected single-cell measurements as zeros, show that filtering improves concordance with hundred-cell measurements, and demonstrate advantages of the combined likelihood-ratio test for differential expression.
- Concordance and filtering: Positive expression values agree with a log-normal model, while undetected genes are consistent with null or negligible RNA abundance.The positive expression distributions match the postulated normal distribution on the log scale.
- Concordance and filtering: Treating undetected wells as zeros produces good concordance with hundred-cell measurements, unlike treating them as missing values.The comparison uses an in-silico average of 100 single-cell measurements against an in-vitro hundred-cell aggregate.
- Concordance and filtering: Filtering parameters tz = 9 and tζ = 9 achieve the best WSS reduction across three data sets and move average per-unit estimates toward the diagonal.The WSS is computed on the log2(y + 1) scale to reduce the effect of extreme outliers.
- Normalization and housekeeping genes: Housekeeping genes provide little normalization utility: GAPDH and POLR2A have R2 = .027 in data set A, and normalization could mask cellular artifacts.After filtering, their regression has R2 = .017, while their apparent correlation is concentrated in cells suspected of technical error.
- Differential expression testing: At an FDR of 1%, the combined test detects more than 20 additional gene×unit changes across four stimulations than the comparison tests.Across a wide range of FDR values, the combined likelihood test produces the greatest number of discoveries compared with Bernoulli-only, normal-theory-only, and t-tests.
- Differential expression testing: The combined likelihood-ratio test remains robust across gene-expression frequencies by combining evidence from Bernoulli and normal components.Figure 4 plots −log10 p-values against expression frequency πj for the Bernoulli, normal, and combined tests.
- Differential expression testing: At an FDR of 1%, 65 genes are detected in at least one individual, with clustered up-regulated and down-regulated modules and substantial response variability.The heatmap uses signed log10 p-values, clustering selected genes by rows and individuals by stimulation-ordered columns.
4 Conclusion
The paper presents a framework for exploring, quality-controlling, and testing single-cell gene-expression data, combining discrete and continuous evidence in a likelihood-ratio test. It positions these tools as a coherent framework for researchers using this nascent technology.
- The framework combines the discrete zero-inflated and continuous portions of single-cell expression data in a likelihood-ratio test.Both components can provide meaningful evidence for detecting outliers and biologically interesting expression changes.
- The conclusion emphasizes a framework for data exploration, quality control, and differential-expression testing in single-cell assays.
- The selected-gene heatmap organizes genes by rows and individuals by columns, with column colors indicating antigen stimulation.
- The proposed default outlier-filtering parameters are conservative and may require tuning to balance technical-error removal against missing biological heterogeneity.
- The likelihood-ratio test may not apply when traits are not blocked within individuals, although mixed-effects extensions could address between-individual variability.
- The approach is intended to support researchers using single-cell assays as the technology becomes more routine.The paper frames effective statistical methods as increasingly important for this developing technology.