Source-linked AI summary

Interpreting 16S metagenomic data without clustering to achieve sub-OTU resolution

Mikhail Tikhonov, Robert W. Leach, Ned S. Wingreen

arXiv:1312.0570v2q-bio.QMq-bio.GN

TL;DR

Clustering 16S sequences into OTUs can obscure ecologically distinct subpopulations, motivating finer-grained analysis of community diversity and dynamics. The paper uses cross-sample correlation analysis of denoised, unclustered 16S data to identify sub-OTU structure, resolving up to 20 subpopulations within standard 97% OTUs and finding strong ecological correspondence for exact shared tags.

  • Problem

    Clustering similar 16S sequences into OTUs can obscure questions about community diversity, stability, resilience, and assembly.

  • Method

    The approach analyzes cross-sample abundance correlations after denoising, while retaining distinct true sequences instead of merging them into OTUs.

  • Results

    Up to 20 distinct subpopulations were identified within standard 97% similarity OTUs, while exact shared 16S tags showed dynamic similarity across cohabiting hosts and a single-nucleotide mismatch degraded it.

  • Takeaways & Limitations

    Sub-OTU structure can provide ecological insight, including for similar habitats, repeated sampling, environmental cross-sectional data, and datasets generated with 454 sequencing.

  • Takeaways & Limitations

    The approach is intended for subtle differences between broadly similar communities and requires error structure to be consistent across samples.

Abstract

from arXiv · show

The standard approach to analyzing 16S tag sequence data, which relies on clustering reads by sequence similarity into Operational Taxonomic Units (OTUs), underexploits the accuracy of modern sequencing technology. We present a clustering-free approach to multi-sample Illumina datasets that can identify independent bacterial subpopulations regardless of the similarity of their 16S tag sequences. Using published data from a longitudinal time-series study of human tongue microbiota, we are able to resolve within standard 97% similarity OTUs up to 20 distinct subpopulations, all ecologically distinct but with 16S tags differing by as little as 1 nucleotide (99.2% similarity). A comparative analysis of oral communities of two cohabiting individuals reveals that most such subpopulations are shared between the two communities at 100% sequence identity, and that dynamical similarity between subpopulations in one host is strongly predictive of dynamical similarity between the same subpopulations in the other host. Our method can also be applied to samples collected in cross-sectional studies and can be used with the 454 sequencing platform. We discuss how the sub-OTU resolution of our approach can provide new insight into factors shaping community assembly.

I Introduction

16S studies commonly cluster reads into OTUs, but sequence similarity is an unreliable proxy for ecological similarity and clustering depends on computational choices. The paper introduces clustering-free, cross-sample analysis to resolve ecologically distinct subpopulations within OTUs.

  • OTU clustering is the de facto standard for analyzing 16S tag-sequencing data.
  • OTU membership is operational rather than biologically motivated, and 16S similarity may not predict ecological similarity.
  • Denoising improves resolution but cannot fully handle errors outside its model, especially when identifying low-abundance species.
  • The proposed method combines error-model denoising with cross-sample comparisons without clustering similar sequences together.

II Materials and Methods

The study analyzes longitudinal Illumina 16S data from human body sites, focusing on tongue samples, and uses cross-sample abundance patterns to distinguish true closely related sequences from noise. This strategy supports sub-OTU resolution while limiting interpretation to sequences with sufficient abundance.

  • The dataset contains longitudinal V4 16S Illumina samples from four body sites in one male and one female individual.
  • The analysis focuses primarily on tongue samples because they most closely probe dynamics in a defined body location.
  • Denoising alone cannot reliably resolve close sequences because unmodeled errors may be labeled as true sequences.
  • Cross-sample comparisons use abundance dynamics to identify closely related sequences that behave as independent community members.
  • The workflow applies cross-sample correlation analysis after individual-sample denoising to achieve sub-OTU resolution.
  • The simplified denoiser runs two orders of magnitude faster than DADA while matching its performance for moderate-abundance sequences in mock-community data.

C Cluster-free filtering – the denoiser

The denoiser estimates substitution-error structure from abundant sequences, conservatively retains sequences repeatedly classified as real across samples, and avoids remapping noisy reads to preserve relative abundances.

  • Illumina substitution errors generate many unique sequences and have an error structure that can be estimated directly from data.
  • Candidate sequences that survive chimera filtering are labeled real under conservative criteria designed to avoid false positives.
  • Sequences marked real in at least 2 of 507 samples are retained, adapting the denoiser to multi-sample analysis.
  • The method does not remap noisy reads to probable sources, prioritizing speed and robustness of relative-abundance measurements.
  • Because non-identical reads are never clustered together, the workflow provides single-nucleotide resolution.

III Results

The analysis resolves distinct bacterial subpopulations by combining sequence information with cross-sample abundance dynamics, showing that sequence similarity alone does not determine ecological similarity. It also distinguishes related sequences arising from different cells from genomic variants within a single bacterium.

  • III Results: 307 filtered sequences included 184 sequence pairs with 99.2% similarity that were resolved as independently present in the community.The smallest resolvable unit is a subpopulation sharing an identical sequenced 16S fragment.
  • A Sequence similarity need not imply ecological similarity, and vice versa: 1 nucleotide differences can separate ecologically distinct subpopulations, while 81% sequence identity can accompany nearly identical abundance dynamics.The contrasting sequence pairs demonstrate that sequence similarity need not imply ecological similarity, and vice versa.
  • A Sequence similarity need not imply ecological similarity, and vice versa: Dynamical similarity is measured as abundance-trace Pearson correlation normalized by the maximum expected correlation cmax, accounting for Poisson noise in low-abundance sequences.Sequence distance is measured by Hamming distance after pairwise alignment.
  • B Cluster-free filtering can resolve distinct subpopulations with high dynamical similarity.: Autocorrelation analysis supports the use of cross-sample dynamics because individual subpopulations fluctuate slowly enough for correlations between consecutive samples to be detected.The analysis treats samples as approximately equally spaced, with a mean separation of 1.1 days.
  • B Cluster-free filtering can resolve distinct subpopulations with high dynamical similarity.: Persistence of difference PD distinguishes same-cell variants from distinct subpopulations by measuring the 1-day autocorrelation of normalized abundance-ratio differences.Same-cell variants should have weak or negligible PD, whereas distinct subpopulations can retain persistent relative differences.
  • B Cluster-free filtering can resolve distinct subpopulations with high dynamical similarity.: High-dynamical-similarity pairs show the predicted bimodal PD distribution, identifying a subset as likely genomic 16S variants within single bacteria.Dynamically dissimilar pairs are unimodal and consistent with the null model; low-PD pairs occupy the bottom-right region.

C Clustering reads into OTUs vastly underestimates ecological richness

The analysis shows that similarity-based OTUs can combine subpopulations with distinct abundance dynamics, while exact 16S tag identity is more informative across cohabiting hosts than 99.2% near-identity.

  • OTU dynamical diversity: OTUs combine sequences from subpopulations with high dynamical diversity, despite their traces remaining representative of the most abundant member.The weighted score Qw is relatively high, whereas the unweighted score Qu is dramatically lower.
  • OTU dynamical diversity: High-abundance members were analyzed because abundance time-trace correlations become noisy for low-abundance sequences.The Fig. 4 analysis considered only sequences among the top 200 by overall abundance.
  • OTU dynamical diversity: The method has finite resolution because analyzed sequences are only 130 nt long, so distinct 16S genes can remain unresolved.This limitation artificially inflates OTU quality scores near 100% similarity, implying that true quality scores are likely lower.
  • Exact identity across hosts: 75% of top-N oral 16S sequences were shared at 100% identity between two cohabiting individuals.The high proportion of perfect matches was interpreted as evidence for more recent common ancestry than close, non-identical sequences within either community.
  • Exact identity across hosts: Dynamical similarity was correlated across hosts for 73 sequences shared within the top 100, but pairing sequences with one nucleotide mismatch substantially degraded that correlation.The inexact pairing represented 99.2% sequence identity.
  • Exact identity across hosts: Exact 100% identity therefore had qualitatively different implications from 99.2% near-identity for predicting subpopulation dynamics.The comparison used independently measured pairwise dynamical similarities in the two individuals.

IV Discussion

The discussion distinguishes sub-OTU resolution from OTU coarse-graining and argues that multi-sample abundance dynamics can reveal ecologically meaningful sequence-level structure. The approach is intended for moderate-to-high abundance members of broadly similar communities, not as a replacement for diversity-oriented OTU analyses.

  • Up to 20 distinct subpopulations were identified within standard 97% similarity OTUs using cross-sample correlation analysis of denoised 16S data.
  • Exact 16S-tag identity between cohabiting individuals was associated with dynamically similar subpopulations, whereas a single nucleotide mismatch degraded that similarity.
  • The approach deliberately avoids merging similar true sequences, distinguishing denoising erroneous reads from application-dependent OTU coarse-graining.
  • The method is intended for broadly similar communities and can inform studies of strain exchange, invasion or extinction dynamics, and ecological divergence relative to sequence divergence.
  • Multi-sample data make abundance ratios more informative than absolute abundances and give sub-OTU resolution its greatest power in time-series contexts.
  • Because it discards low-abundance sequences, the approach is unsuitable for population-level alpha- or beta-diversity and does not replace OTU clustering.

A Supplementary methods. Cluster-free filtering: details and applications

The method begins from the observation that stable, low-noise communities should repeatedly produce dominant sequences rather than diffuse clouds of similar reads. Tongue microbiome data support this premise through consistent dominance of strongly distinct sequences across samples.

  • Low sequencing noise predicts that community members are predominantly represented by the same sequences, surrounded by low-abundance error variants.
  • The five most abundant tongue-community sequences were strongly distinct and remained consistently among the top 10 ranks across samples.
  • The observed rank stability supports treating dominant sequences as biological representatives rather than diffuse sequencing-error clouds.

2 Estimating rates of one-nucleotide substitutions

Substitution error rates are inferred directly from reproducible error clouds around abundant sequences, then used to estimate each sequence’s expected error-derived abundance. This provides a data-specific basis for denoising and retaining genuine close variants.

  • 390 first neighbors exist for each 130-nt sequence, enabling error-rate estimation from observed Hamming-distance-one neighbors around abundant sequences.
  • Substitution type explains most neighbor-abundance variation, while position along these 130-nt sequences contributes little.
  • Different substitution rates can differ by up to 50-fold, yet estimates from the top 50 sequences are highly reproducible.
  • Directly estimating rates avoids a conservative global error bound that could overestimate some error types by up to 50-fold and reduce resolution of close sequences.
  • The estimated average total error rate is 1.0 10^-3 per nucleotide, corresponding under independent-error assumptions to an 88% probability of recording a 130-nt sequence without errors.
  • The denoiser orders sequences by abundance and calculates each sequence’s null-model abundance from spillover probabilities contributed by more abundant first neighbors.

4 Other error types, including chimeras and PCR indels

The pipeline models substitution errors but treats chimeras and indels separately because they are not captured by that model. It provides optional or platform-specific filtering and a pooled multi-sample workflow for scalable processing.

  • Substitution errors are modeled quantitatively, but retained sequences may also include chimeras, PCR indels, and context-dependent substitutions outside the model.
  • PCR indel errors are context-specific and lack a quantitative model; conservative filtering cannot resolve true biological sequences differing only by indels.
  • Indel filtering is optional for Illumina data but necessary for 454 data, where homopolymer indel errors are frequent.
  • Chimeras are filtered with UCHIME de novo, while the software package implements the denoising workflow as open-source Perl scripts.
  • For large multi-sample datasets, samples are pooled, error rates are estimated from the pooled library, and the neighbor structure is constructed once.

6 Mock community validation and comparison with DADA

Mock-community tests show that cluster-free filtering recovers reference sequences comparably to DADA while retaining some sequences DADA rejects. The approach is designed for efficient cross-sample analysis, especially of moderate-to-high-abundance sequences.

  • The reference comparisons were complicated by errors and missing matches in the Sanger clone reference sets.
  • 23 reference sequences were correctly identified by both DADA and cluster-free filtering in the Divergent dataset.
  • 1 of 90 Artificial-set reference sequences was missed by DADA but correctly identified by cluster-free filtering.
  • Cluster-free filtering retained three additional detections just above its 10-count threshold, including a sequence DADA classified as a possible error.
  • The method is optimized for large datasets and cross-sample comparisons of individually denoised samples, focusing on moderate-to-high-abundance sequences.

8 Example of other applications: environmental cross-sectional 454 data

The method extends beyond longitudinal Illumina data to cross-sectional environmental samples and 454 sequencing, resolving ecologically distinct sequences that differ by one nucleotide. Its resolution depends on sample number and data quality, with several scope and interpretation caveats.

  • Cross-sample comparisons can be applied to cross-sectional or location-series datasets when samples are collected and processed similarly enough to share an error structure.
  • 454-data error rates depended on quality filtering and experimental protocol, supporting direct dataset-specific error estimation rather than separate calibration.
  • Single-nucleotide sequence differences in Lake Mystic samples corresponded to ecologically significant distinctions across depth profiles.
  • More samples permit finer resolution: two samples resolved one pair, whereas subtler depth-trace differences required more samples.
  • Autocorrelation-based analyses are sensitive to sample number, and applying persistence of difference to spatial heterogeneity requires further investigation.
  • Reported OTU quality scores likely underestimate true diversity because analyses emphasized abundant members but could overcount paralogs or unmodeled errors.

1 Cross-individual analysis of fecal samples

Fecal samples from two individuals show that nearly identical 16S sequences can have distinct ecological dynamics, while shared sequences exhibit correlated dynamics across hosts. The findings support recurring exchange of community members but leave some comparisons statistically limited.

  • Single-nucleotide differences separated dominant sequences in Male 3 and Female 4, with virtually no cross-contamination observed.
  • Both dominant sequences mapped to Bacteroides sp. in GreenGenes despite their distinct host distributions.
  • The two gut communities shared many sequences at 100% identity, supporting non-negligible exchange of community members.
  • Dynamical similarity of shared sequences, measured independently in the two individuals, was clearly correlated.
  • Fewer shared fecal sequences than tongue sequences left the statistics insufficient for comparing the two body sites.

2 Cross-individual analysis at 97% OTU level

At the conventional 97% OTU level, the cross-individual analysis examined dynamical similarity among 78 common OTUs within the top 100. These OTUs were assigned using closed-reference QIIME matching to GreenGenes.

  • 78 common 97% OTUs within the top 100 were compared for dynamical similarity measured independently in the two individuals.
  • The analysis used a scatter plot of dynamical similarity between pairs of common OTUs across the two individuals.
  • The OTUs were constructed by closed-reference OTU picking in QIIME against GreenGenes at 97% sequence similarity.
Loading 1312.0570v2…