Source-linked AI summary
A Characterization of the DNA Data Storage Channel
Reinhard Heckel, Gediminas Mikutis, Robert N. Grass
TL;DR
DNA storage offers longevity and high density, but its short, unordered molecules and imperfect synthesis, handling, storage, and sequencing create losses and within-molecule errors. The paper characterizes this channel using datasets from the authors and two other groups, estimating sampling and error distributions across experimental conditions. It finds that within-molecule errors are mainly associated with synthesis and sequencing, while handling and storage cause significant sequence loss.
Problem
DNA-storage design lacks a quantitative and qualitative understanding of molecule sampling, sequence loss, and within-molecule errors across synthesis, storage, handling, and sequencing.
Method
The paper estimates channel statistics from the authors’ experiments and datasets from Goldman’s and Erlich’s groups, using aligned reads to measure edit-based substitution, insertion, and deletion probabilities.
Results
Within-molecule errors are mainly due to synthesis and sequencing, while PCR bias and storage cause significant changes in read distributions and loss of sequences.
Takeaways & Limitations
DNA-storage error-correcting designs should account for unavoidable sequence loss and for storage time, physical redundancy, and anticipated DNA interactions such as PCR steps.
Takeaways & Limitations
The analysis includes a filtering artifact in Erlich’s data and an expectation about PCR saturation in the Erlich and Zielinski experiment.
Abstract
from arXiv · showhide
Owing to its longevity and enormous information density, DNA, the molecule encoding biological information, has emerged as a promising archival storage medium. However, due to technological constraints, data can only be written onto many short DNA molecules that are stored in an unordered way, and can only be read by sampling from this DNA pool. Moreover, imperfections in writing (synthesis), reading (sequencing), storage, and handling of the DNA, in particular amplification via PCR, lead to a loss of DNA molecules and induce errors within the molecules. In order to design DNA storage systems, a qualitative and quantitative understanding of the errors and the loss of molecules is crucial. In this paper, we characterize those error probabilities by analyzing data from our own experiments as well as from experiments of two different groups. We find that errors within molecules are mainly due to synthesis and sequencing, while imperfections in handling and storage lead to a significant loss of sequences. The aim of our study is to help guide the design of future DNA data storage systems by providing a quantitative and qualitative understanding of the DNA data storage channel.
1 Introduction
DNA is promising for archival storage because of its longevity and density, but practical systems rely on many short, unordered molecules that are sampled and affected by synthesis, storage, handling, and sequencing errors. This paper models and quantifies these channel effects to inform storage-system and encoder/decoder design.
- Motivation: DNA has attracted archival-storage interest because conventional media guarantee lifetimes of only a few years, whereas DNA offers longevity and enormous information density.The paper cites CERN’s more than 100 petabytes of archived physical data as an example of growing archival demand.
- Technological constraints: Current systems use short DNA molecules stored in an unordered pool, preventing spatial random access and requiring information retrieval through sampling and sequencing.Synthesis is difficult for strands significantly longer than one-two hundred nucleotides, and sequencing is preceded by PCR amplification.
- Error and loss sources: Synthesis may fail or produce uneven copy numbers, storage can cause molecular loss, and sampling reveals only a fraction of the molecules in the pool.The observed fraction depends on the pool’s molecule distribution and the number of reads.
- Channel model: The channel abstraction takes a multiset of M length-L molecules, samples N times independently, and disturbs sampled molecules with insertions, deletions, and substitutions.The sampling and error distributions determine the channel and how much data can be stored.
- Design objective: Understanding channel parameters such as storage time and physical density is important for maximizing bits per total nucleotide while guaranteeing successful decoding.The relevant storage cost includes substantially more nucleotides than the ML nucleotides in the input multiset.
- Paper aim: The paper quantifies sampling and within-molecule error distributions across experimental setups and assigns errors, when possible, to reading or writing processes.The analysis combines DNA-storage literature data with the authors’ own experiments because fully controlled experiments are costly.
2 Error sources
DNA-storage errors arise from synthesis, PCR and other handling, storage decay, and sequencing. These processes can either remove molecules from observation or introduce substitutions, insertions, and deletions within molecules.
- Synthesis: Chemical synthesis can produce uneven copy numbers across sequences because synthesis yield depends on chip location, sequence properties, and surface imperfections.Current chip-based methods generate sets containing millions of copies rather than individual DNA strands.
- PCR: PCR is generally high fidelity but preferentially amplifies some sequences, distorting their copy-number distribution and thereby changing the observed read distribution.Each cycle multiplies molecules by a sequence-dependent factor typically slightly below two.
- Storage and handling: Storage and processing cause hydrolytic damage, including strand-breaking depurination and cytosine deamination that forms uracil.Broken strands lack required primers and are diluted out during amplification, while deamination can produce substitutions depending on the PCR enzyme.
- Summary: Storage can cause significant loss of whole molecules as well as substitution errors within DNA molecules.The paper summarizes these as the two principal consequences of storage-related damage.
- Sequencing: Illumina sequencing errors are strand-specific, increase toward read ends, and are dominated by substitutions, while insertions and deletions are much less likely.The cited order for insertions and deletions is 10^-6, although the exact number depends on the dataset.
- Deamination consequences: Proof-reading PCR enzymes can prevent amplification of strands containing uracil, whereas non-proof-reading enzymes can convert U-G damage into incorrect T-A basepairs and C2T errors.The figure also associates the complementary change with G2A errors.
- Sequence-dependent effects: High GC content and long homopolymer stretches increase substitution and deletion error rates, while high-GC molecules also exhibit high dropout and PCR-error rates.Substitution and deletion rates increase significantly for homopolymers longer than six.
3 Error statistics
The paper estimates molecule-level error probabilities and analyzes how observed errors can be allocated among writing, storage, and reading processes across DNA-storage experiments.
- Analysis scope: Error statistics are estimated for substitutions, insertions, deletions, and the output distribution of molecules across datasets from three research groups.The analysis compares errors at the molecule level with the distribution of molecules observed at the channel output.
3.1 Data sets
The study compares DNA-storage datasets that vary in synthesis scale, molecule length, synthesis method, storage treatment, physical redundancy, PCR cycles, and read coverage. These differences provide the basis for analyzing how experimental conditions affect channel behavior.
- Common setup: Across all datasets, the input consists of synthesized molecules and the output consists of forward and backward Illumina reads after storage and processing.The datasets therefore expose both channel inputs and sequencing-derived outputs for comparison.
- Dataset parameters: The datasets differ in synthesized-strand count, target sequence length, synthesis method, expected thermal decay, physical redundancy, PCR cycles, and read coverage.Physical redundancy is the expected number of copies of each strand in the DNA pool, while read coverage is the number of reads per synthesized strand.
- Goldman et al.: Goldman’s dataset contains 153,335 length-117 molecules without homopolymers and uses Agilent OLS Sureprint synthesis with Illumina HiSeq 2000 sequencing.Forward and backward reads have length exactly 104.
- Erlich and Zielinski: Erlich and Zielinski’s dataset contains 72,000 length-152 molecules with no homopolymers longer than two and includes tenfold physical-redundancy dilutions.The dilution series starts from approximately 10^7 copies.
- Grass et al.: Grass et al.’s high-redundancy datasets contain 4,991 length-117 molecules without homopolymers longer than three, measured before and after thermal treatment corresponding to four half-lives.The DNA was synthesized with Customarray and sequenced using Illumina Miseq 2x150bp Truseq.
- Low-redundancy datasets: A corresponding low-redundancy dataset uses 4,991 length-117 molecules and compares untreated DNA with DNA exposed to thermal treatment corresponding to four half-lives.These datasets are denoted Low PR and Low PR 4t1/2.
3.2 Errors within molecules
This section estimates substitution, insertion, and deletion probabilities within individual DNA molecules across multiple datasets, separating synthesis, storage, and sequencing contributions. It finds that synthesis likely accounts for most indels, while substitutions dominate sequencing-related errors and storage decay produces characteristic substitution patterns.
- Overall error probabilities: The analysis estimates substitution, insertion, and deletion probabilities by aligning each read to the closest original molecule using Levenshtein distance.Probabilities are averaged as the number of each error type per nucleotide required for alignment.
- Overall error probabilities: Sequencing error is estimated independently from two-sided reads, whereas overall error combines synthesis, storage, and sequencing errors.This separation is possible because two-sided reads provide an independent estimate of sequencing error.
- Errors by source: Most deletions and insertions in the overall reads are likely due to synthesis, while substitution errors dominate among reads with the correct target length.The datasets examined include Goldman’s and Erlich’s data, High PR, and heat-treated High PR 4t1/2 DNA.
- Substitution patterns: Conditional substitution probabilities show that T-to-C and A-to-G mistakes are much more likely than other substitution errors.The paper links these transitions to cytosine deamination and uracil misinterpretation during PCR with non-proofreading polymerases.
- Substitution patterns: Storage decay increases C-to-T and G-to-A errors because cytosine deamination can occur during DNA processing, PCR, handling, and storage.Heat-treated DNA is included to examine substitution errors introduced by decay.
3.3 Distribution of molecules
The observed molecule distribution reflects physical redundancy, PCR amplification bias, storage-related decay, and sequencing, shaping how many synthesized sequences are recovered. Low redundancy and PCR bias produce uneven read counts and sequence loss, while storage further increases dropout and substitution errors.
- Distribution of molecules: The output distribution depends on molecule proportions in the pool, sequencing, and read count, while observed-sequence coverage is central to coding design.The fraction of synthesized molecules observed at least once affects how many bits can be stored on a given number of DNA sequences.
- PCR bias: Erlich’s and Goldman’s datasets are approximately negative-binomial distributed, whereas the original DNA and dry control have long tails and peaks at zero.Reading coverage is approximately the same across datasets; differences are associated with PCR cycles, physical redundancy, and dilution.
- Distribution of molecules: 56.6%, 57.7%, 56.7%, 83%, and 78.6% of molecules have the correct target length across the five datasets.Previous works recovered information only from molecules with the correct length.
- PCR bias: PCR amplification bias can strongly reshape molecule proportions because small efficiency differences compound over many cycles.With efficiencies of 80% and 90%, 60 cycles change equal proportions to (1.8/1.9)^60 = 0.039.
- Low physical redundancy and loss of sequences: At physical redundancy 1, uniform random subsampling is expected to lose about 36% of distinct sequences, producing many molecules with few or no reads.The dilution experiment shows increasingly long-tailed read distributions as physical density decreases.
- Errors due to storage: Storage increases sequence dropout and substantially raises substitution errors, while having only a marginal effect on insertion and deletion probabilities.In High PR, unseen molecules rise from 1.6% to 8% after four half-lives; in Low PR, they rise from 34% to 53%.
4 Conclusion
The paper characterizes DNA storage errors and shows that sequence loss, especially from PCR bias and storage, is central to system design. It recommends coding strategies that account for lost sequences while exploiting the generally low within-sequence error rate.
- Errors within sequences are mainly caused by synthesis and sequencing, with storage contributing to a smaller extent.Synthesis mainly introduces deletions and some insertions; sequencing adds substitutions when next-generation sequencing is used.
- PCR bias significantly changes read distributions, and many PCR cycles increase the number of sequences that are never read.Large numbers of cycles produce a very long-tailed distribution.
- Storage causes significant sequence loss and increases substitution error probabilities through cytosine deamination.
- Low physical redundancy increases information density but can cause sequence loss, altered read distributions, and greater exposure to PCR bias.These effects should inform error-correcting scheme parameters and depend on anticipated interactions with the DNA.
- An outer code is crucial for correcting lost sequences, while inner codes, read combining, and outer-code correction can address relatively infrequent within-sequence errors.The number of erasures the outer code corrects should reflect expected losses from storage, redundancy, and DNA interactions.
A Detailed description of Low PR and Low PR 4t1/2
The experiment synthesized 4,991 DNA molecules, amplified them for 30 PCR cycles, and exposed high-redundancy samples to accelerated decay conditions.
- 4,991 DNA molecules of length 117 were synthesized without homopolymers longer than three.
- The initial library was amplified for 30 PCR cycles before preparing buffered samples for storage.Samples contained 70 ng/mL DNA in Tris-EDTA buffer at pH 8.0.
- High-redundancy samples were exposed to 65 C for 20 days to accelerate decay to 4t1/2.