Source-linked AI summary

Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species

Keith R. Bradnam, Joseph N. Fass, Anton Alexandrov, Paul Baranay, Michael Bechner, İnanç Birol, Sébastien Boisvert, Jarrod A. Chapman, Guillaume Chapuis, Rayan Chikhi, Hamidreza Chitsaz, Wen-Chi Chou, Jacques Corbeil, Cristian Del Fabbro, T. Roderick Docking, Richard Durbin, Dent Earl, Scott Emrich, Pavel Fedotov, Nuno A. Fonseca, Ganeshkumar Ganapathy, Richard A. Gibbs, Sante Gnerre, Élénie Godzaridis, Steve Goldstein, Matthias Haimel, Giles Hall, David Haussler, Joseph B. Hiatt, Isaac Y. Ho, Jason Howard, Martin Hunt, Shaun D. Jackman, David B Jaffe, Erich Jarvis, Huaiyang Jiang, Sergey Kazakov, Paul J. Kersey, Jacob O. Kitzman, James R. Knight, Sergey Koren, Tak-Wah Lam, Dominique Lavenier, François Laviolette, Yingrui Li, Zhenyu Li, Binghang Liu, Yue Liu, Ruibang Luo, Iain MacCallum, Matthew D MacManes, Nicolas Maillet, Sergey Melnikov, Bruno Miguel Vieira, Delphine Naquin, Zemin Ning, Thomas D. Otto, Benedict Paten, Octávio S. Paulo, Adam M. Phillippy, Francisco Pina-Martins, Michael Place, Dariusz Przybylski, Xiang Qin, Carson Qu, Filipe J Ribeiro, Stephen Richards, Daniel S. Rokhsar, J. Graham Ruby, Simone Scalabrin, Michael C. Schatz, David C. Schwartz, Alexey Sergushichev, Ted Sharpe, Timothy I. Shaw, Jay Shendure, Yujian Shi, Jared T. Simpson, Henry Song, Fedor Tsarev, Francesco Vezzi, Riccardo Vicedomini, Jun Wang, Kim C. Worley, Shuangye Yin, Siu-Ming Yiu, Jianying Yuan, Guojie Zhang, Hao Zhang, Shiguo Zhou, Ian F. Korf

arXiv:1301.5406v3q-bio.GN

TL;DR

High-quality vertebrate genome assembly remains difficult, and assemblies vary substantially in their strengths across metrics and species. Assemblathon 2 evaluated assemblies using complementary sequence-based and statistical measures, finding both strong performers and important limitations in interpreting assembly quality.

  • Problem

    De novo assembly of large vertebrate genomes remains challenging, with limited transcriptome resources available for directly assessing gene completeness.

  • Method

    The study evaluated assemblies using core-gene detection, cumulative length plots, Fosmid tag-pair accuracy, and combined metric rankings.

  • Results

    Assembly performance was heterogeneous: SGA ranked first across the snake metrics, while bird rankings varied depending on the metric set and scoring choices.

  • Takeaways & Limitations

    Assemblies require careful inspection because performance can differ across metrics, and complementary resources can help validate assembly quality.

  • Takeaways & Limitations

    The study did not assess whether assemblers resolved heterozygous regions into separate haplotypes, limiting interpretation of unexpectedly large assemblies.

Abstract

from arXiv · show

Background - The process of generating raw genome sequence data continues to become cheaper, faster, and more accurate. However, assembly of such data into high-quality, finished genome sequences remains challenging. Many genome assembly tools are available, but they differ greatly in terms of their performance (speed, scalability, hardware requirements, acceptance of newer read technologies) and in their final output (composition of assembled sequence). More importantly, it remains largely unclear how to best assess the quality of assembled genome sequences. The Assemblathon competitions are intended to assess current state-of-the-art methods in genome assembly. Results - In Assemblathon 2, we provided a variety of sequence data to be assembled for three vertebrate species (a bird, a fish, and snake). This resulted in a total of 43 submitted assemblies from 21 participating teams. We evaluated these assemblies using a combination of optical map data, Fosmid sequences, and several statistical methods. From over 100 different metrics, we chose ten key measures by which to assess the overall quality of the assemblies. Conclusions - Many current genome assemblers produced useful assemblies, containing a significant representation of their genes, regulatory sequences, and overall genome structure. However, the high degree of variability between the entries suggests that there is still much room for improvement in the field of genome assembly and that approaches which work well in assembling the genome of one species may not necessarily work well for another.

Affiliations

The paper’s listed affiliations span genome centers, universities, and research institutes in the United States, Canada, the United Kingdom, and Russia.

  • The affiliations include Genome Center, UC Davis, and computational biology, bioinformatics, and computer science departments.
  • Cold Spring Harbor Laboratory, the University of California, Berkeley, and Wayne State University are also represented.
  • Additional affiliations include Baylor College of Medicine, the British Columbia Cancer Agency, the Wellcome Trust Sanger Institute, and institutions in Portugal and Russia.

Keywords

The listed keywords identify genome assembly, assembly statistics, scaffolding, assessment, heterozygosity, and COMPASS.

  • Genome assembly and scaffolds identify the paper’s central assembly objects.
  • Assessment, heterozygosity, and COMPASS identify evaluation themes and an analysis framework.

Background

Advances in sequencing have expanded de novo genome assembly, but evaluating assembly quality remains difficult because assemblies can vary in completeness and structure. Assemblathon 2 addresses this gap using real multi-technology data and validation resources across three vertebrate species.

  • Sequencing technologies have become faster, easier, and more accurate, with read lengths increasing substantially on Illumina platforms.
  • De novo assemblers now target large vertebrate genomes, but assemblies can be shorter than references and depleted in segmental duplications and larger repeats.
  • The field has developed benchmarking efforts including dnGASP, GAGE, and Assemblathon using standardized datasets.
  • Assemblathon 2 uses real sequencing reads from multiple NGS technologies for a budgerigar, Lake Malawi cichlid, and boa constrictor.
  • Assemblies are assessed with multiple metrics and experimental datasets, including Fosmid sequences and optical maps, because correct reference sequences are unavailable.
  • Many assemblers perform well on individual metrics but few perform consistently across metrics, and performance can differ between species.

Data Description

Assemblathon 2 collected assemblies generated by participating teams from varied sequencing data, software, and computing environments. The submitted assemblies and associated sequence data, analysis code, and results were made available through public repositories and files.

  • Teams had four months to assemble three vertebrate genomes from varied NGS data and could submit competitive and evaluation assemblies.
  • Assemblies used diverse software with substantially varying hardware and time requirements.
  • Assemblies smaller than 25% of the expected species genome size were excluded from detailed analysis.
  • Submitted assemblies and input reads were deposited through the Assemblathon website, GigaDB, and sequence read archives.
  • Analysis scripts, assembly statistics, and results were provided through a GitHub repository, spreadsheet, and CSV file.

Additional Files

The supplementary files provide sequencing-data descriptions, assembly instructions, comprehensive metric results, validation alignments, validated Fosmid regions, and accession information.

  • The supplementary data description details the Illumina, Roche 454, and Pacific Biosciences sequencing data made available to participating teams.
  • Assembly instructions document how participating teams could use software to recreate their assemblies.
  • The master spreadsheet contains 102 metrics for every assembly, plus z-scores and average rankings for ten key metrics.
  • A CSV file provides the master spreadsheet’s results in a format more suitable for computer-script parsing.
  • Bird and snake Fosmid validation files include BLAST alignments, read coverage, repeats, and assembly tracks, while separate files provide validated regions and 100 nt tags.
  • Accession spreadsheets identify Project, Study, Sample, Experiment, and Run records for bird, fish, and snake input reads.

Analyses

The analyses combine assembly-size and contiguity statistics with gene-content, reference-alignment, Fosmid, optical-map, and error-based assessments. Results show that metrics can disagree substantially, so assembly suitability depends on the intended use and species.

  • Statistical description of assemblies: NG50 replaces total assembly size with estimated genome size, whereas N50 uses the assembly’s total sequence length when identifying the 50% cumulative-length threshold.
  • Statistical description of assemblies: Scaffold NG50 and contig NG50 correlated modestly in bird and snake (r = 0.50 and 0.55) but more strongly in fish (r = 0.78, P < 0.01).
  • Statistical description of assemblies: 246% of the estimated genome size was contained in the IOBUGA fish evaluation assembly, illustrating that unusually large assemblies may reflect assembly errors or successfully included sequence.
  • Statistical description of assemblies: 99.2% of the estimated genome size in the Ray bird assembly occurred in scaffolds at least 25 Kbp long despite its third-lowest NG50 scaffold ranking.
  • Presence of core genes: Nearly all of the 458 CEGs appeared in at least one assembly, while per-assembly fractions generally varied between 85–95%.
  • Presence of core genes: The secondary 248-gene CEG subset revealed additional core genes too fragmented for detection in the original analysis, although both CEG sets produced highly correlated results.
  • COMPASS analysis of VFRs: COMPASS evaluates reference alignments using coverage, validity, multiplicity, and parsimony to account for aligned sequence, duplication, and repeat expansion or collapse.
  • COMPASS analysis of VFRs: The Ray snake assembly ranked first overall in COMPASS and first on every individual measure except multiplicity, where it still performed better than average.

Discussion

Assemblathon 2 found substantial variation in assembly quality across species, assemblers, metrics, and data combinations. No single assembly strategy performed consistently across genomes or evaluation criteria.

  • Interspecific variation: Bird assemblies generally had longer contigs and scaffolds, with more assemblies reaching at least the estimated 1.2 Gbp genome size.Bird assemblies also performed better than fish and snake in optical-map assessments, although coverage differences may contribute.
  • Interspecific variation: 383, 418, and 424 core eukaryotic genes were detected on average in competitive bird, fish, and snake assemblies, respectively.REAPR estimated that 68% of fish-assembly bases were error free, compared with 73% for bird and 75% for snake.
  • Interspecific variation: Genome size alone did not explain quality differences: snake had the largest estimated genome at 1.6 Gbp, yet fish assemblies generally scored lowest.Differences in heterozygosity, repeat content, and sequencing data may account for interspecific variation.
  • Species-specific rankings: The BCM-HGSC competitive assembly ranked highest for bird by summed z-score, while BCM-HGSC and Allpaths exchanged first and second place under average-rank scoring.The ranking depended partly on included metrics, because excluding CEGMA, scaffold NG50, or contig NG50 would change the competitive winner.
  • Species-specific rankings: SGA produced the clearest snake result, ranking first overall and scoring highly in eight of ten key metrics, while remaining first when any one metric was removed.However, SGA ranked first in only one individual metric and seventh in gene-sized scaffolds, illustrating the difficulty of aggregating metrics.
  • Data combinations and consistency: Assemblers rarely performed consistently across metrics or species, and combining sequencing technologies could improve some measures while lowering overall rank or increasing cost.PacBio gap filling lengthened BCM-HGSC contigs but lowered its overall rank; adding Roche 454 improved SOAPdenovo mainly through coverage and validity.

Methods

Assemblathon 2 evaluated submitted assemblies with statistical descriptions and independent sequence-based reference analyses, including CEGMA, Fosmid sequences, and COMPASS alignments. The evaluation also recognized that a 98% identity cutoff and treatment of ambiguous bases might not suit every case.

  • Assemblies were submitted as FASTA scaffold sequences, with runs of at least 25 Ns marking contig boundaries.
  • Basic scaffold and contig statistics were generated by splitting scaffolds at runs of 25 or more Ns.
  • CEGMA assessed each assembly’s gene complement, using its vertebrate-intron option to detect longer introns.
  • Fosmid paired-end libraries with approximately 35 Kbp inserts provided independent reference data for bird and snake assemblies.
  • COMPASS aligned scaffolds to a reference and calculated coverage, validity, multiplicity, and parsimony metrics from the parsed alignments.
  • The COMPASS analysis used a minimum 98% alignment identity and treated N characters as ambiguous, options that may not be appropriate for all cases.

VFR distance analysis

The analysis combined validated-region, optical-map, and read-mapping assessments to characterize assembly accuracy and structure. REAPR-derived statistics were normalized within species and combined into a summary score.

  • Validated Fosmid-region analysis searched successive 1,000 nt regions using paired 100 nt tags, retaining matches with at least 95 nt aligned per tag.
  • Optical-map validation aligned scaffolds of at least 300 Kbp with at least 9 restriction sites to map supercontigs across three alignment levels.
  • Optical-map level 1 represented strict global concordance, level 2 allowed gaps or minor differences, and level 3 captured local alignments after global failure.
  • REAPR used independently mapped reads and proper-pair, insert-size, alignment-score, and mapping-quality filters to generate perfect and uniquely mapping coverage.
  • The final REAPR score combined normalized error-free bases, N50, and broken N50 as Number of error free bases * (broken N50)^2 / (original N50).

Dataset 1

The first supporting dataset contains Assemblathon 2 genome assemblies and associated CEGMA gene-prediction outputs for three vertebrate species. Assemblies were distributed as scaffolded or contig FASTA files, with CEGMA results generated using a vertebrate-intron setting.

  • Assemblathon 2 submitted 43 assemblies for three species: 15 bird, 16 fish, and 12 snake.
  • The assemblies were assessed with statistical approaches and experimental data from Fosmid sequences and optical maps.
  • The dataset covers budgerigar, Lake Malawi cichlid, and boa constrictor assemblies, totaling 27 GB in 86 compressed FASTA files.
  • Assemblies were scaffolded contigs in FASTA format, with runs of at least 25 Ns marking gaps; unscaffolded submissions had identical contig and scaffold files.
  • CEGMA gene-content assessment searched for nearly full-length genes matching HMMs built from 458 highly conserved genes.
  • CEGMA was run with the --vrt option to allow vertebrate-sized introns, producing DNA, protein, GFF, identifier, and completeness-report outputs.

Dataset 3

The third supporting dataset contains assembled Fosmid sequences used to assess bird and snake genome assemblies. It includes validated Fosmid regions and their associated sequencing-based construction data.

  • Bird and snake assemblies were assessed using high-confidence assembled Fosmid regions supplied with the Assemblathon 2 manuscript.
  • The dataset contains 47 complete assembled Fosmid sequences for bird and 29 for snake.
  • The Fosmid data are provided as FASTA DNA-sequence files for budgerigar and boa constrictor.
  • Pooled Fosmid clone Illumina paired-end libraries used approximately 35 Kbp inserts and ten pools with varying pooling levels.

Competing Interests

The study involved multiple teams contributing genome assemblies, evaluation analyses, sequence data, and manuscript support.

  • Participating teams generated genome assemblies from the supplied sequencing data using various software packages and optional preprocessing steps.
  • Contributors generated optical maps and analyzed assemblies against them for each species.
  • Other contributors provided bird, bird-and-snake Fosmid, and manuscript-review support.
  • Contributors assembled Fosmid sequences, performed COMPASS and REAPR analyses, and calculated comparative assembly metrics.

Figure Legends

The figures visualize assembly lengths, gene content, Fosmid and optical-map validation, REAPR scores, and rankings across key metrics.

  • Figures 1–3 show NG graphs summarizing scaffold lengths for bird, fish, and snake assemblies.
  • Figures 4–6 depict gene-sized scaffold representation, core eukaryotic gene presence, and annotated Fosmid sequences.
  • Figures 7–9 define and report COMPASS Coverage, Validity, Multiplicity, and Parsimony metrics for bird and snake assemblies.
  • Figures 10–12 use cumulative scaffold or alignment lengths and Validated Fosmid Regions to assess assembly structure and short-range accuracy.
  • Figures 13–15 present optical-map alignment results, while Figure 16 summarizes REAPR scores and notes unavailable analyses for some assemblies.

Tables

The tables document participating teams, sequencing inputs, assembly software, and identifiers used throughout the evaluation.

  • Table 1 lists participating teams, assembly identifiers, affiliations, and software information.
  • Assembly entries include tools such as ABySS, SOAPdenovo, SGA, Ray, Symbiose, Meraculous, Phusion, and PRICE.
  • Some entries combined multiple tools, including Monument, SSPACE, SuperScaffolder, and GapCloser.
  • Teams could submit competitive and evaluation assemblies, while some assemblies used distinct or mislabeled library combinations.
  • Table 2 summarizes the sequencing data provided to participants, with additional sequence descriptions available in supplementary methods.
Loading 1301.5406v3…