Source-linked AI summary

Variant interpretation using population databases: lessons from gnomAD

Sanna Gudmundsson, Moriel Singer-Berk, Nicholas A. Watts, William Phu, Julia K. Goodrich, Matthew Solomonson, Genome Aggregation Database Consortium, Heidi L. Rehm, Daniel G. MacArthur, Anne ODonnell-Luria

arXiv:2107.11458v3q-bio.GN

TL;DR

Distinguishing disease-causing variants from benign human variation requires large, diverse population databases. This review guides use of the gnomAD browser and its features for interpretation, highlighting that the database contains only a subset of possible synonymous and nonsense variants.

  • Problem

    Variant and gene interpretation requires larger, more diverse reference population databases to distinguish Mendelian-disorder causes from benign variation.

  • Method

    This review guides use of the gnomAD browser and features including allele frequency, expression levels, constraint scores, and variant co-occurrence.

  • Results

    gnomAD currently contains only a subset of possible synonymous and nonsense variants.

  • Takeaways & Limitations

    The gnomAD browser features support interpretation of candidate variants and novel genes in rare disease.

  • Takeaways & Limitations

    The review notes caveats in using constraint scores for variant interpretation.

Abstract

from arXiv · show

Reference population databases are an essential tool in variant and gene interpretation. Their use guides the identification of pathogenic variants amidst the sea of benign variation present in every human genome, and supports the discovery of new disease-gene relationships. The Genome Aggregation Database (gnomAD) is currently the largest and most widely used publicly available collection of population variation from harmonized sequencing data. The data is available through the online gnomAD browser (https://gnomad.broadinstitute.org/) that enables rapid and intuitive variant analysis. This review provides guidance on the content of the gnomAD browser, and its usage for variant and gene interpretation. We introduce key features including allele frequency, per-base expression levels, constraint scores, and variant co-occurrence, alongside guidance on how to use these in analysis, with a focus on the interpretation of candidate variants and novel genes in rare disease.

GRANT NUMBERS … 2. DATA COMPOSITION

The review presents gnomAD as a large, widely used reference population database for distinguishing potentially pathogenic variation and interpreting rare-disease variants and genes. It describes gnomAD’s composition, version-specific coverage, quality controls, representation limits, and available variant classes.

  • GRANT NUMBERS: The work was supported by National Institutes of Health awards UM1HG008900, U01HG011755, U24HG011450, and U54DK105566.The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
  • KEYWORDS: The paper focuses on reference populations, gnomAD, variant interpretation, and allele frequency.
  • 1. INTRODUCTION: Population frequency data help distinguish rare variants more likely to cause Mendelian disorders from common, largely benign variation.gnomAD contains data from more than 195,000 individuals and is the most widely accessed reference population dataset, with over 150,000 weekly page views.
  • 2.2 gnomAD version 2 and version 3: gnomAD v2 includes 125,748 exomes and 15,708 genomes aligned to GRCh37 and is preferable for coding-variant interpretation, whereas v3 provides 76,156 GRCh38 genomes for noncoding or poorly captured regions.The versions substantially overlap: 81% of v2 genomes are also in v3, and 6% of v2 exomes are represented as v3 genomes.
  • 2.3 Structural and mitochondrial variants: The browser provides allele frequencies for structural and mitochondrial variants, including ~445,000 structural variants from 10,738 genomes and mitochondrial data from 56,434 v3 genomes.Mitochondrial annotations include homoplasmic and heteroplasmic calls plus population- and haplogroup-specific allele frequencies.

3. NAVIGATING THE gnomAD BROWSER

The gnomAD browser gene page integrates coverage, transcript and tissue-expression profiles, constraint metrics, per-base expression, ClinVar distributions, and variant tracks to support gene and variant interpretation. These views reveal transcript representation, regional variation patterns, and potential limitations affecting interpretation.

  • Gene-page overview: Gene pages combine exome and genome coverage, union all exons across transcripts, transcript models, and GTEx tissue-expression profiles for gene-level assessment.Mean gene coverage reveals differences between exome and genome sequencing, while transcript views identify canonical models and expression patterns.
  • Constraint metrics: Constraint metrics compare observed and expected very rare SNVs, corrected for sequence context and coverage, across synonymous, missense, and predicted loss-of-function classes.Positive Z-scores indicate fewer variants than expected and increased constraint, whereas negative Z-scores indicate more variants than expected.
  • Constraint metrics: Mapping challenges from paralogs, pseudogenes, or segmental duplications can inflate observed synonymous variation, making constraint interpretation unreliable when synonymous Z-scores deviate substantially from zero.The paper advises caution with all constraint data for genes showing this pattern.
  • Per-base expression: The pext score provides normalized expression for each gene position, visualizing aggregate exon expression or expression across 38 tissues.In NSD1, pLoF variants occur in low-pext regions, while pathogenic ClinVar variants are largely absent there; pext may not reflect developmental expression.
  • ClinVar and variant tracks: ClinVar tracks display coding variants with pathogenicity-based colors, include variants regardless of gnomAD presence, and support filtering to variants present in gnomAD.Expanded tracks can reveal missense hotspots and pathogenic variant enrichment, including NSD1 pLoF variants and missense variants in SET and PWWP domains.

4. VARIANT INTERPRETATION USING gnomAD

gnomAD supports rare-disease variant and gene interpretation through population-frequency evidence, constraint metrics, and variant co-occurrence. Its evidence must be interpreted cautiously because rarity is not sufficient for pathogenicity and constraint varies by disease mechanism and phenotype timing.

  • Variant classification: ACMG/AMP standards incorporate reference databases such as gnomAD into variant classification across benign, likely benign, uncertain significance, likely pathogenic, and pathogenic categories.The standards aggregate population, computational, functional, and segregation evidence.
  • Population-frequency evidence: Rarity is essential for Mendelian analysis but is not sufficient for pathogenicity, because most rare variants are not pathogenic.Rarity or absence from population databases has been recommended as supporting rather than moderate evidence, since unrepresented variation is often benign and databases are unsaturated.
  • Population-frequency evidence: The popmax allele frequency, defined as the maximum frequency across continental populations, is recommended for filtering common variants, with caution for population-enriched conditions.A variant common in one population can generally be considered benign across populations, but this assumption may fail when a condition is much more common in that population.
  • Variant co-occurrence: Haplotype prediction from variant co-occurrence helps determine whether variants share a haplotype and whether combinations are likely deleterious for expected phenotypes.Variants predicted to be on different haplotypes are unlikely to form a deleterious combination for phenotypes not expected in gnomAD.

5. LIMITATIONS OF REFERENCE POPULATION DATABASES

gnomAD is a useful public sequence resource, but caveats limit inferences about variant pathogenicity. Its incomplete variation, missing individual-level phenotypes, population imbalance, and sequencing or annotation artifacts require caution in interpretation.

  • Incomplete disease exclusion: Individuals with Mendelian disease may be included in gnomAD, so variants observed in one or a few individuals should not be automatically excluded as disease candidates.The database remains far from representing all possible variation, and absence from gnomAD is consistent with—but insufficient for—disease involvement.
  • Phenotype limitations: Aggregate gnomAD data lack individual-level phenotypes, and known Mendelian phenotypes have been removed, making it unsuitable for patient-phenotype matching.Biobanks and resources such as Geno2MP and VariantMatcher provide phenotype-linked rare-disease data and researcher contact options.
  • Population representation: European participants are over-represented, while African, Middle Eastern, and Oceanian populations are poorly represented, leaving patients from these communities with more rare variants of uncertain significance.Improving population diversity is a high priority, including reprocessing available datasets for gnomAD inclusion.
  • Data quality: Despite extensive quality control, gnomAD contains sequencing and annotation artifacts, requiring approaches to evaluate variant quality.These limitations apply to gnomAD as to any large genomics resource.

6. RESOURCES

The gnomAD browser provides ongoing feature updates, interactive metric views, and user support resources. Additional learning materials include open-access primary publications and video tutorials on browser use and population genetics.

  • Additional gnomAD browser features are announced through the News section, Changelog, and @gnomAD_project on Twitter.
  • Users can consult the browser’s “?” buttons and Help page for additional information or contact the team directly.
  • Interactive browser views expose gnomAD metrics including allele frequency, constraint scores, and pext.
  • Open-access primary gnomAD publications are listed on the browser’s Publications page for deeper understanding of the dataset.
  • Video tutorials cover gnomAD browser use and medical and population genetics through workshop and Broad Institute resources.

7. CONCLUDING REMARKS

gnomAD’s scale enables highly accurate estimates of rare allele frequencies and broad analyses of variation, supporting rare-disease variant interpretation, gene discovery, and diverse scientific applications. Future releases will expand variant-class coverage and population representation, improving analytical power and reducing uncertainty for under-represented ancestries.

  • Current contributions: gnomAD’s size enables accurate allele-frequency estimates for incredibly rare variation and analysis of variation across genes and genomic regions.These capabilities have proven invaluable for interpreting variants in patients with rare genetic disorders.
  • Current contributions: The resource supports scientific applications including cross-species mutational-intolerance comparisons, selection-coefficient estimation, disease-gene assessment, and dominant-variant penetrance determination.
  • Current contributions: gnomAD has aided discovery of genes associated with diseases including neurodevelopmental and congenital heart disorders.
  • Future directions: Increasing representation of diverse ancestries is needed to improve applicability to under-represented populations and decrease variants of uncertain significance in affected patients.This requires inclusion in global genomics projects and data sharing that balances accessibility with community and individual preferences, especially for Indigenous peoples.

DATA AVAILABILITY STATEMENT

gnomAD data are displayed through the gnomAD browser and available for download through multiple public data platforms. The dataset incorporates data from TOPMed or dbGaP projects, including TCGA, GTEx, and ADSP.

  • The gnomAD data are displayed on the gnomAD browser at https://gnomad.broadinstitute.org/.
  • gnomAD data are available for download through Google Cloud Public Datasets, the Registry of Open Data on AWS, Azure Open Datasets, and the UCSC genome browser.
  • The gnomAD data are partly based on TOPMed or dbGaP data from TCGA, GTEx, and ADSP.The cited projects are managed by NCI and NHGRI, the NIH Common Fund and NHGRI, and NIA and NHGRI, respectively.

CONFLICT OF INTEREST · WEB RESOURCES

The section reports authors’ conflicts of interest and lists web resources, supplementary materials, authorship information, and institutional affiliations associated with the gnomAD review.

  • CONFLICT OF INTEREST: DGM is a founder with equity of Goldfinch Bio and serves as a paid advisor to GSK.
  • CONFLICT OF INTEREST: DGM’s paid-advisor roles include Variant Bio, Insitro, and Foresite Labs.
  • WEB RESOURCES: UCSC genome browser gnomAD tracks are provided for GRCh37 and GRCh38.
  • WEB RESOURCES: The web resources include gnomAD links in the manuscript, an NIH Statement on Sharing Research Data, and a gnomAD educational video.
  • WEB RESOURCES: Supplementary data accompany the paper, titled “Variant interpretation using population databases: lessons from gnomAD.”
  • WEB RESOURCES: The listed authors include Sanna Gudmundsson, Moriel Singer-Berk, Nicholas A. Watts, William Phu, Julia K. Goodrich, Matthew Solomonson, and the Genome Aggregation Database Consortium.
  • WEB RESOURCES: Author affiliations span the Broad Institute, Boston Children’s Hospital, Massachusetts General Hospital, Garvan Institute of Medical Research, UNSW Sydney, and Murdoch Children’s Research Institute.

CONTENT

The paper’s supplementary content includes methods, figures and tables, and references, followed by consortium authorship and funding statements.

  • Supplementary material: Supplementary methods are provided on page 36.
  • Supplementary material: Supplementary figures and tables are provided on page 37.
  • Supplementary material: Supplementary references are provided on page 48.
  • Consortium information: Genome aggregation database consortium authors are listed on page 49.
  • Consortium information: Genome aggregation database consortium funding statements are provided on page 62.

SUPPLEMENTARY METHODS

The analysis randomly sampled 100 individuals from five continental populations in gnomAD v2.1.1 exomes and filtered variants using rarity, quality, and problematic-region criteria. Variants were classified by their most severe canonical-transcript VEP consequence into pLoF, missense/inframe indel, or synonymous categories.

  • Sampling and variant filtering: 100 individuals were randomly selected from each of five populations in the gnomAD v2.1.1 exome dataset.The populations were African/African American, Latino/Admixed American, East Asian, Non-Finnish European, and South Asian.
  • Sampling and variant filtering: Variants were restricted to very rare variants with popmax allele frequency < 0.1% across the entire gnomAD dataset.The dataset included v2 exomes, v2 genomes, and v3 genomes; popmax is the maximum frequency across continental populations, excluding Ashkenazi Jewish, Finnish, and remaining samples.
  • Sampling and variant filtering: Variants failing gnomAD quality control or located in problematic regions were excluded.Problematic regions included low complexity, decoy, and segmental duplication regions.
  • Variant annotation: VEP version 85 assigned variants to pLoF, missense/inframe indel, or synonymous categories using the most severe canonical-transcript consequence.pLoF variants required high-confidence LOFTEE support and included splice-acceptor, splice-donor, stop-gained, and frameshift consequences.
  • Implementation: The analysis used Hail for processing and exported variants to TSV for visualization in R.The code selected individuals and filtered variants.

SUPPLEMENTARY FIGURES AND TABLES

The supplementary materials show that larger, more diverse gnomAD datasets reduce apparently unique coding variation and provide examples of variant co-occurrence analysis. They also illustrate transcript expression, regional constraint, loss-of-function curation, complex variants, coverage limitations, and allele-balance warnings.

  • Supplementary tables: Mean unique coding variants decreased from 36 (± 16) to 27 (± 13) across populations when comparing v2 exomes with the entire gnomAD dataset.The difference was significant (p < 2.2e-16), demonstrating the importance of increased sample size and more diverse population representation.
  • Gene and variant interpretation: Supplementary examples demonstrate transcript tissue-expression sorting, regional missense constraint, and manual or LOFTEE-based loss-of-function curation.The NSD1 examples identify the highest-expressed transcript in cerebellar hemisphere tissue and show pathogenic-variant clustering in missense-constrained regions.
  • Variant co-occurrence: Variant co-occurrence analysis predicted two LAMA1 variants to reside on the same haplotype in informative populations.The variants occurred in isolation in 68 and 170 individuals and co-occurred in 304 individuals; East Asian and Finnish populations were uninformative.
  • Complex variants: Complex-variant examples show that separately predicted frameshift or stop-gained effects can combine into in-frame or alternative consequences.A GAA frame-restoring indel forms a 9 base-pair (3 amino acid) in-frame deletion, while an MBD5 stop-gained and missense combination produces a missense variant.
  • Quality caveats: Coverage and allele-balance warnings identify situations in which low coverage or contamination can distort variant calls and allele frequencies.Examples include coverage in fewer than 50% of individuals and inflated alternate reads causing homozygous variants to be called heterozygotes.

SUPPLEMENTARY REFERENCES

The supplementary references identify the Genome Aggregation Database Consortium authors and provide their institutional affiliations.

  • Supplementary references: The paper credits the Genome Aggregation Database Consortium authors.The supplementary section begins with the consortium heading and an extensive author list.
  • Supplementary references: Contributors represent institutions including the Broad Institute of MIT and Harvard, Massachusetts General Hospital, Harvard Medical School, and Columbia University Medical Center.Affiliations span medical, academic, and research organizations in the United States and internationally.

STATEMENTS

The statements acknowledge funding from cardiovascular, diabetes, neuroscience, complex-disease genetics, and international research organizations. They also document substantial support for TOPMed, MESA, and related genomic and clinical research activities.

  • Funding acknowledgments: Authors report support from the British Heart Foundation, Academy of Finland, National Institute of Diabetes and Digestive and Kidney Diseases, and Medical Research Council UK.Named awards include British Heart Foundation awards CS/14/2/30841 and RG/18/10/33842, Academy of Finland grants 338182, 312073, and 336823, and Medical Research Council UK Centre Grant No. MR/L010305/1.
  • Funding acknowledgments: Additional support came from the Academy of Finland Center of Excellence for Complex Disease Genetics and the National Institute of Diabetes and Digestive and Kidney Diseases.The acknowledgments also name Sigrid Jusélius Foundation support and grants 312074 and 336824 for complex disease genetics.
Loading 2107.11458v3…