Source-linked AI summary
Fast and accurate imputation of summary statistics enhances evidence of functional enrichment
Bogdan Pasaniuc, Noah Zaitlen, Huwenbo Shi, Gaurav Bhatia, Alexander Gusev, Joseph Pickrell, Joel Hirschhorn, David P Strachan, Nick Patterson, Alkes L. Price
TL;DR
Existing imputation methods often require individual-level genotypes, whereas summary association statistics are increasingly available. The paper develops Gaussian imputation from summary statistics, recovering much of the effective sample size of HMM-based imputation and increasing functional enrichment evidence.
Problem
Existing HMM-based imputation requires individual-level genotypes, while summary association statistics are becoming widely available and individual-level data may be inaccessible.
Method
The paper uses a conditional multivariate Gaussian approach to impute association statistics at untyped variants from typed-variant summary statistics.
Results
The method recovers much of HMM-based imputation’s effective sample size and increases genic-versus-non-genic enrichment magnitude and statistical evidence for four lipid traits.
Takeaways & Limitations
Imputation of summary statistics is presented as a tool for future functional enrichment analyses.
Takeaways & Limitations
The approaches have limitations when summary LD information is involved, as discussed by the authors.
Abstract
from arXiv · showhide
Imputation using external reference panels is a widely used approach for increasing power in GWAS and meta-analysis. Existing HMM-based imputation approaches require individual-level genotypes. Here, we develop a new method for Gaussian imputation from summary association statistics, a type of data that is becoming widely available. In simulations using 1000 Genomes (1000G) data, this method recovers 84% (54%) of the effective sample size for common (>5%) and low-frequency (1-5%) variants (increasing to 87% (60%) when summary LD information is available from target samples) versus 89% (67%) for HMM-based imputation, which cannot be applied to summary statistics. Our approach accounts for the limited sample size of the reference panel, a crucial step to eliminate false-positive associations, and is computationally very fast. As an empirical demonstration, we apply our method to 7 case-control phenotypes from the WTCCC data and a study of height in the British 1958 birth cohort (1958BC). Gaussian imputation from summary statistics recovers 95% (105%) of the effective sample size (as quantified by the ratio of $χ^2$ association statistics) compared to HMM-based imputation from individual-level genotypes at the 227 (176) published SNPs in the WTCCC (1958BC height) data. In addition, for publicly available summary statistics from large meta-analyses of 4 lipid traits, we publicly release imputed summary statistics at 1000G SNPs, which could not have been obtained using previously published methods, and demonstrate their accuracy by masking subsets of the data. We show that 1000G imputation using our approach increases the magnitude and statistical evidence of enrichment at genic vs. non-genic loci for these traits, as compared to an analysis without 1000G imputation. Thus, imputation of summary statistics will be a valuable tool in future functional enrichment analyses.
2. Bioinformatics Interdepartmental Program, University of California Los Angeles
The listed affiliation is the Department of Medicine, Lung Biology Center, at the University of California San Francisco.
- The affiliation is the Department of Medicine, Lung Biology Center, University of California San Francisco.
- The institution identified is the University of California San Francisco.
- The department identified is the Department of Medicine.
5. Departments of Epidemiology and Biostatistics, Harvard School of Public Health,
The listed affiliations include the Broad Institute of Harvard and MIT, Harvard Medical School, and St George’s University of London.
- One listed affiliation is the Broad Institute of Harvard and MIT in Cambridge, MA, USA.
- Another listed affiliation is Harvard Medical School, MA, USA.
- A third listed affiliation is the Division of Population Health Sciences and Education at St George’s, University of London, UK.
Author Summary
The paper develops summary-statistics imputation to extend association analyses beyond typed variants without individual-level data. Simulations and enrichment analyses indicate that it captures much of the signal of standard imputation and strengthens functional enrichment evidence.
- Summary-statistics imputation addresses the need for methods that work without individual-level genotype data.
- The approach captures the same signal as standard individual imputation for common and, to a lesser extent, low-frequency variants.
- The method supports enrichment analyses by enhancing signal for functionally relevant variant classes compared with non-imputed analyses.
Introduction
Genotype imputation expands association testing, but established HMM methods require individual-level data. This work introduces Gaussian imputation from typed-SNP summary statistics, evaluates its accuracy, and applies it to association and enrichment analyses.
- Motivation: Genotype imputation leverages public haplotypic diversity to increase the number of variants tested in genome-wide studies.
- Motivation: Imputation also supports meta-analysis by harmonizing variants across studies using different genotyping platforms.
- Problem: Existing high-accuracy HMM approaches require individual-level genotype data, which privacy and logistical constraints can restrict.
- Approach: The proposed method tests association at untyped SNPs using summary association statistics from typed variants.
- Approach: The method approximates locus-level association-statistic distributions with a multivariate Gaussian informed by linkage disequilibrium.
- Simulation results: 84% (54%) of effective sample size was recovered for common (>5%) and low-frequency (1-5%) variants versus 89% (67%) for HMM imputation.
- Simulation results: 87%(60%) of effective sample size was recovered when summary pairwise LD information from GWAS samples was available, without increased false-positive associations.
- Applications: Masked lipid-trait summary statistics correlated 0.98 (0.95) for common (low-frequency) variants, while imputation increased genic-versus-non-genic enrichment evidence.
Methods
The method models GWAS z-scores with a multivariate Gaussian distribution and imputes unobserved SNP statistics conditionally on typed SNPs. It uses reference-panel LD, windowing, noise adjustment, and variance normalization to produce calibrated and computationally efficient summary-statistic imputation.
- Imputation estimates the posterior mean of unobserved SNP z-scores conditional on observed typed z-scores and reference-panel covariance.
- Finite reference-panel size introduces statistical noise, so the method uses windowing and adds λI to the estimated LD matrix.Windows model distant SNPs as uncorrelated, while λI stabilizes matrix inversion and controls noise at proximal SNPs.
- The imputed z-scores are linear combinations of typed z-scores with weights precomputed from the reference panel.
- When target-sample summary LD is available, ImpG-SummaryLD directly estimates imputed-statistic variance without reference-panel noise adjustment.This produces calibrated null statistics and increased power relative to ImpG-Summary.
- GWAS z-scores are approximated by a multivariate Gaussian distribution whose covariance reflects LD between SNPs.
- Conditional variance yields r2pred, an imputation-accuracy measure, and imputed z-scores are normalized to have theoretical null variance 1.The variance is estimated from predictor weights and LD among typed SNPs.
1958 Birth Cohort data
The British 1958 birth cohort provided nationally representative DNA and phenotype data for empirical imputation analyses. The study combined overlapping SNPs across genotyping panels and evaluated imputation against masked association statistics.
- The British 1958 birth cohort followed people born in England, Scotland, and Wales during one week in 1958.
- At age 44–45 years, participants underwent biomedical examination and blood sampling, establishing a nationally representative DNA collection.
- Non-overlapping DNA subsets were genotyped by WTCCC, T1DGC, and GABRIEL using several array platforms.
- A combined dataset retained SNPs common across the three panels, with SNPs failing frequency, call-rate, HWE, or cross-dataset checks excluded.
- Public lipid summary data comprised roughly 2.7 million SNP statistics from approximately 100,000 samples per phenotype, reduced to about 2.0 million SNPs after filtering.
Results
Across simulations and empirical datasets, summary-statistic Gaussian imputation recovered substantial association power, maintained calibration when reference-panel noise was adjusted, and ran far faster than HMM-based approaches. Applying it to lipid meta-analysis data increased functional-enrichment evidence at genic relative to intergenic loci.
- λGC = 0.94 was observed for ImpG-Summary, with no increase in false-positive rate at the distribution tail.Adjustment shrank predictor weights slightly but was necessary to avoid false positives.
- 84% (54%) of effective sample size was recovered for common (>5%) and low-frequency (1–5%) variants by ImpG-Summary, versus 89% (67%) for Beagle.ImpG-SummaryLD increased recovery to 87% (60%) when typed-variant LD from the GWAS was available.
- ImpG-Summary took less than one CPU day for a 10,000-sample GWAS, whereas IMPUTE2 with pre-phasing took more than 200 CPU days.The runtime difference increases for larger studies.
- When individual-level genotypes are available, HMM-based imputation remains the approach of choice, while ImpG-SummaryLD is recommended for rapid prioritization.
- Association statistics from summary-statistic imputation correlated at r=0.98 (0.95) with HMM-based results at common (low-frequency) variants.Array-based masking and re-imputation yielded r=0.97 (0.91) at common (low-frequency) variants.
- 1000G imputation increased statistical evidence of enrichment at genic versus intergenic SNPs, with larger variance increases for genic classes.
- Across all four lipid phenotypes, median KS statistics were more significant after 1000G imputation than in the original data.For HDL, the values were 4.75E-08 versus 7.63E-05.
Discussion
The authors introduce summary-statistics imputation methods that avoid individual-level genotypes, address false positives from limited reference panels, and support functional-enrichment analyses. Across simulations and empirical applications, the approach is nearly as powerful as individual-level HMM imputation and strengthens enrichment signals.
- Simulation and empirical performance: The approach is almost as powerful as individual-level genotype imputation across simulations and real data involving discrete and continuous phenotypes.The evaluations assess both power and false-positive associations.
- Method and scope: ImpG-Summary and ImpG-SummaryLD impute association statistics from typed-variant summary statistics, using reference haplotype panels such as 1000 Genomes.ImpG-SummaryLD additionally uses summary LD statistics, which are not currently widely shared.
- Functional enrichment: For four lipid traits, imputed 1000 Genomes summary statistics show larger and more statistically significant enrichment in genic versus non-genic regions than the original data.The authors publicly released imputed association statistics at 1000 Genomes variants and conclude that summary-statistics imputation can increase power in functional-enrichment studies.
- False-positive control: A ridge-like adjustment removes false-positive associations caused by using unadjusted LD estimates and accounts for the limited size of the reference panel.The authors expect larger reference panels to require a smaller adjustment factor and improve accuracy.
- Limitations and future work: Accuracy is higher for common than low-frequency variants and may be lower for very rare variants; appropriate quality control is needed to avoid false positives.The authors also identify extension to low-coverage sequencing as future work.
Web Resources
The paper provides software, reference-panel weights, and imputed summary association statistics for four lipid traits.
- The released resources include ImpG-Summary and ImpG-SummaryLD software.
- Precomputed weights from the 1000 Genomes reference panel are provided.
- Imputed summary association statistics at 1000G SNPs are available for four lipid traits.
Figure Titles and Legends
The figures compare HMM-imputed and ImpG-Summary association statistics across disease, height, lipid, and functional-class analyses.
- Figure 1 compares HMM-imputed and ImpG-Summary z-scores for BD in WTCCC and height in the 1958 Birth Cohort.
- Figure 2 compares HMM-imputed and ImpG-Summary z-scores at known NHGRI GWAS Catalog SNPs for WTCCC and height.
- Figure 3 compares HMM-imputed and ImpG-Summary z-scores for triglycerides using masked z-scores or variants present on the Illumina 610 array.The masked analysis imputes 10% from the remaining 90%; array-based imputation starts from all Illumina 610 variants.
- Figure 4 bins average variance per SNP by functional class for four blood phenotypes, comparing original data with 1000 Genomes-imputed statistics.The bottom panel shows absolute differences between original and imputed association statistics.
Tables
The tables define metrics and compare imputation performance across simulated effective sample size, reference-panel settings, runtime, and association statistics.
- Relative effective sample size is the ratio of average χ2 statistics at imputed versus typed SNPs, with R>1 indicating non-null effect sizes.
- Reference-panel comparisons evaluate imputation in random 1000 Genomes European subsamples versus Great Britain haplotypes at odds ratio 1.5.
- Runtime estimates report CPU days required for whole-genome imputation across 11.6 million European-polymorphic SNPs for varying numbers of individuals.
- Association-performance comparisons use average χ2 statistics across known NHGRI GWAS Catalog SNPs for eight phenotypes.The WTCCC average excludes the HLA region and covers 216 SNPs.