Source-linked AI summary

Initial recommendations for performing, benchmarking, and reporting single-cell proteomics experiments

Laurent Gatto, Ruedi Aebersold, Juergen Cox, Vadim Demichev, Jason Derks, Edward Emmott, Alexander M. Franks, Alexander R. Ivanov, Ryan T. Kelly, Luke Khoury, Andrew Leduc, Michael J. MacCoss, Peter Nemes, David H. Perlman, Aleksandra A. Petelski, Christopher M. Rose, Erwin M. Schoof, Jennifer Van Eyk, Christophe Vanderaa, John R. Yates, Nikolai Slavov

arXiv:2207.10815v2q-bio.OT

TL;DR

Single-cell MS proteomics can quantify many proteins across individual cells, but artifacts and analytical choices threaten accuracy and reproducibility. The paper proposes experimental, computational, quality-control, and reporting recommendations, concluding that community adoption can establish stronger foundations for the field and improve data reuse and replication.

  • Problem

    Single-cell MS proteomics is vulnerable to experimental and computational artifacts that can compromise quantitative accuracy, reproducibility, and interpretation.

  • Method

    The paper proposes guidelines covering experimental design, sample preparation, data evaluation, uncertainty assessment, reproducibility, and reporting of single-cell proteomics workflows.

  • Results

    The proposed reporting standards and evaluation practices are intended to support replication, data reuse, and more rigorous adoption of single-cell MS proteomics.

  • Takeaways & Limitations

    Community adoption of these guidelines can establish stronger foundations for single-cell MS proteomics and promote dissemination, improvement, and adoption of its technologies and analyses.

  • Takeaways & Limitations

    Reproducibility is bounded by cellular uniqueness and batch effects, so reports must specify what is reproduced and how batch effects influence it.

Abstract

from arXiv · show

Analyzing proteins from single cells by tandem mass spectrometry (MS) has become technically feasible. While such analysis has the potential to accurately quantify thousands of proteins across thousands of single cells, the accuracy and reproducibility of the results may be undermined by numerous factors affecting experimental design, sample preparation, data acquisition, and data analysis. Broadly accepted community guidelines and standardized metrics will enhance rigor, data quality, and alignment between laboratories. Here we propose best practices, quality controls, and data reporting recommendations to assist in the broad adoption of reliable quantitative workflows for single-cell proteomics.

Experimental design

Reliable single-cell proteomics requires minimizing variation and contamination from cell isolation through MS analysis, while using controls and randomized designs to assess accuracy, precision, and batch effects.

  • Experimental design: Randomize treatments, biological and technical replicates, reagent batches, and analytical batches to distinguish biological variation from technical artifacts.Staggering treatments, processing, and analyses also helps account for sources of variation during interpretation.
  • Sample preparation: Preserve cellular state during isolation by using gentle procedures, rapid low-temperature processing, and documented isolation parameters.FACS can be effective but may be less robust without high-performing sorters and expert operators.
  • Sample preparation: Minimize handling, contaminants, and sample losses because single cells contain only picogram-level protein amounts.Small preparation volumes of 2-20 nl reduce reagent amounts per cell and can proportionally reduce contaminant ions.
  • Data acquisition: Use low-flow nanoLC and maximize ionization efficiency and ion transmission to improve delivery of sample ions to the mass analyzer.Lower flow rates create smaller droplets that desolvate more readily, increasing ionization efficiency.
  • Controls: Assess precision with repeat measurements, accuracy with proteins of known abundance, and matching-between-runs errors with mixed- or single-species controls.Empty samples can underestimate false discoveries from propagated sequence identifications, so they are insufficient as the only MBR control.

Data evaluation and interpretation

Single-cell proteomics interpretation must separate accuracy, precision, biological variation, and technical effects while explicitly accounting for missingness, normalization, dimensionality reduction, and analytical choices.

  • Reproducibility: Reproducibility should specify whether results are repeated, reproduced, or replicated and document software, databases, parameters, metadata, and processing workflows.Containerized workflows and version-controlled protocols can facilitate replication and standardize analyses.
  • Quantitative evaluation: Quantitative accuracy and precision are distinct: clustering may prioritize precision, whereas protein covariation and biophysical modeling depend more on accuracy.Reproducibility alone cannot establish accuracy because systematic errors can produce highly repeatable but erroneous measurements.
  • Data evaluation and interpretation: Absolute-intensity correlations can mislead relative-quantification benchmarks, while cellular uniqueness and batch effects limit reproducibility claims to specified trends or clusters.Subjective thresholds and analytical choices should be treated as uncertainty sources, with sensitivity of conclusions reported when possible.
  • Data evaluation and interpretation: Normalize for cell size and variable protein detection, using common protein subsets or references when the quantified-protein subset differs across cells.Cell size is a major confounder of protein-intensity differences, and total protein content or forward scatter can be included as a downstream covariate.
  • Data evaluation and interpretation: Missing values are common near the LC-MS detection limit, so their uncertainty should be propagated and imputation should reflect whether missingness is random.Observed-data analyses assume observed values represent missing values, whereas imputation methods should account for the missing-data mechanism.
  • Data evaluation and interpretation: Low-dimensional projections are useful for visualization but inadequate for benchmarking quantification because they can omit variance and distort distances through method-specific tuning.PCA generally preserves global distances better than tSNE or UMAP, but any clustering based on a manifold should be checked against higher-dimensional data and technical covariates.

Reporting standards

The recommendations define reporting standards that make single-cell proteomics data reproducible, reusable, and interpretable by documenting experimental design, metadata, acquisition, processing, and file organization. They emphasize that data sharing alone is insufficient without detailed metadata and reproducible analysis records.

  • Experimental design and metadata: Replication requires detailed metadata describing each analyzed cell, including biological and technical descriptors, treatment groups, sampling methods, and analysis batches.The experimental design should be represented as a table with one row per single cell and one column per descriptor.
  • Data sharing: Raw and processed data should be shared in open formats through dedicated repositories, with proprietary binary files converted when possible.Recommended formats include mzML for raw data, mzIdentML for search results, and mzTab or text-based spreadsheets for quantitative data.
  • Scale-dependent reporting: Large single-cell studies heighten the need to report preparation and acquisition batches, cell phenotypes, and intermediate datasets such as search results and peptide-by-cell matrices.These records reduce the computational burden of reproducing individual analysis steps.
  • Analysis reproducibility: Annotated scripts, notebooks, commands, and parameters must accompany shared data because processing choices can profoundly influence final biological interpretations.Standardized software and audit logs can help compare workflows and reproduce individual processing steps.
  • File organization: Consistent folder structures, README files, and machine- and human-readable filenames make multiple raw, identification, and quantitation file sets easier to interpret and reuse.Each folder should describe the research question or file set it contains, while naming conventions and abbreviations should be documented.
  • Implementation: Reporting recommendations should be implemented throughout study design, acquisition, processing, and interpretation so they become routine practices rather than an added burden.The goal of reporting is to enable other researchers to repeat, reproduce, assess, and build upon published data and interpretations.

Conclusions and perspectives

The authors call for community and infrastructure support to establish rigorous foundations for single-cell MS proteomics and improve reuse of its data and results. They present the guidelines as an evolving framework that should adapt to technological advances, larger biological questions, and integration with other modalities.

  • Community adoption: Adoption of the guidelines by researchers, journals, and data archives is presented as essential for scientific rigor in single-cell MS proteomics.The standards are intended to facilitate replication and support broader dissemination and adoption of single-cell technologies and analyses.
  • Future development: The guidelines are expected to evolve as single-cell technologies advance, study scales and biological questions grow, and additional data modalities become integrated.The authors specifically anticipate connections with transcriptomics, spatial transcriptomics, imaging, electrophysiology, prioritized MS, and PTM- or proteoform-level methods.
Loading 2207.10815v2…