Source-linked AI summary

BUSCO update: novel and streamlined workflows along with broader and deeper phylogenetic coverage for scoring of eukaryotic, prokaryotic, and viral genomes

Mosè Manni, Matthew R Berkeley, Mathieu Seppey, Felipe A Simao, Evgeny M Zdobnov

arXiv:2106.11799v1q-bio.GNq-bio.PE

TL;DR

Assessing genomic-data quality is challenging as sequencing data accumulate and genome origins may be unknown. BUSCO v5 updates workflows and datasets, improving runtimes and supporting broader quality assessment across genomic data.

  • Problem

    Growing genomic datasets and unknown metagenome-assembled genome origins motivate improved quality-assessment methods.

  • Method

    BUSCO v5 upgrades workflows and underlying datasets, including dataset details on species, orthologous groups, and genes.

  • Results

    BUSCO v5 provides improved runtimes and user experience, with a representative BUSCO score of C:99.1%.

  • Takeaways & Limitations

    The updated BUSCO workflows support quality assessment across genomic data with more efficient analyses and renewed datasets.

  • Takeaways & Limitations

    High completeness scores can make the auto-lineage workflow less useful for spotting contaminants.

Abstract

from arXiv · show

Methods for evaluating the quality of genomic and metagenomic data are essential to aid genome assembly and to correctly interpret the results of subsequent analyses. BUSCO estimates the completeness and redundancy of processed genomic data based on universal single-copy orthologs. Here we present new functionalities and major improvements of the BUSCO software, as well as the renewal and expansion of the underlying datasets in sync with the OrthoDB v10 release. Among the major novelties, BUSCO now enables phylogenetic placement of the input sequence to automatically select the most appropriate dataset for the assessment, allowing the analysis of metagenome-assembled genomes of unknown origin. A newly-introduced genome workflow increases the efficiency and runtimes especially on large eukaryotic genomes. BUSCO is the only tool capable of assessing both eukaryotic and prokaryotic species, and can be applied to various data types, from genome assemblies and metagenomic bins, to transcriptomes and gene sets.

Introduction

BUSCO v5 expands genomic-data quality assessment through renewed datasets, broader phylogenetic coverage, and revised workflows for eukaryotic, prokaryotic, viral, and unknown-origin sequences. These changes improve assessment scope, efficiency, and applicability to heterogeneous datasets.

  • Datasets and phylogenetic coverage: New prokaryotic, eukaryotic, and viral functionalities broaden BUSCO v5 assessments across the three domains of life.BUSCO v5 includes 27 viral datasets and adds 83 bacterial and archaeal datasets relative to v3.
  • Datasets and phylogenetic coverage: BUSCO v5 renews and expands its underlying datasets with OrthoDB v10, covering more lineages and substantially increasing marker-gene representation.The datasets include more than threefold more sets than odb9 and a fivefold increase in BUSCO marker genes, derived from more than twice as many species.
  • Automated lineage selection: The --auto-lineage function automatically selects an appropriate BUSCO dataset, enabling assessment of sequences with unknown taxonomic origin and metagenomic data.BUSCO can also detect subsets of viruses belonging to clades represented by the newly introduced viral datasets.
  • Scope and applicability: BUSCO v5 supports comprehensive analyses of large heterogeneous datasets, including eukaryotic genomes and microbial-eukaryote, prokaryote, and viral MAGs.It is described as the only available tool assessing genomic data from the three domains in a single analysis.

Materials and Methods

The study used BUSCO datasets documenting the species, orthologous groups, and genes used to construct each set, with analyzed assemblies, gene sets, and main results listed in supplementary tables. Further analytical details were provided in supplementary notes, and plots were generated in R with ggplot2.

  • Datasets: BUSCO datasets are available online and document the species, orthologous groups, and genes used to construct each set.The datasets are hosted at https://busco-data.ezlab.org/v5/data/lineages/.
  • Study materials: Versions and accessions of all genome assemblies and gene sets, together with the analyzed BUSCO main results, are listed in supplementary tables.
  • Analysis and visualization: Further analytical details are described in supplementary notes, while analysis plots were made using ggplot2 in R.

Data Availability

BUSCO is freely available under the MIT Licence, with source code, datasets, mappings, analyzed-accession information, and benchmark data distributed through public repositories.

  • BUSCO is freely distributed under the MIT Licence, with source code available through GitLab, Bioconda, and a Docker container.
  • BUSCO datasets are available online with details on their species, orthologous groups, genes, and mappings to OrthoDB information.
  • Protein IDs used to build the datasets, analyzed assemblies and gene sets, and benchmark data are provided through the BUSCO data site, supplementary tables, NCBI, and Zenodo.

Tables

The tables document the expanded BUSCO odb10 datasets and supplementary evaluations across taxonomic groups, workflows, runtimes, and cross-domain matches. They cover comparisons with earlier BUSCO versions and CheckM, including bacterial, archaeal, eukaryotic, fungal, protist, arthropod, and viral data.

  • Datasets: odb10 greatly expands the number of BUSCO datasets relative to odb9, with supplementary tables reporting marker and species counts for BUSCO versions 4 and 5.Equivalent odb9 datasets are also provided for comparison.
  • Version comparisons: Supplementary scores compare BUSCO v3 and v5 on bacterial, fungal, and metazoan gene sets using corresponding odb9 and odb10 datasets.A separate table summarizes the main differences among BUSCO v3, v4, and v5.
  • Workflow evaluations: Supplementary evaluations report BUSCO_MetaEuk and BUSCO_Augustus scores on fungi, protists, and arthropod genomes, including scores and runtimes across MetaEuk sensitivity settings.Viral genomes and RefSeq gene sets are also assessed where a BUSCO viral dataset is available.
  • Cross-domain matches: Additional supplementary tables examine cross-domain BUSCO matches and marker frequencies across bacterial genomes, fungal genomes, and invertebrate gene sets from RefSeq.The datasets include bacteria_odb10, archaea_odb10, and eukaryota_odb10, with frequency analyses covering 2’779 bacterial genomes, 370 fungal genomes, and 235 invertebrate gene sets.
Loading 2106.11799v1…