Source-linked AI summary
Accurate Genomic Prediction Of Human Height
Louis Lello, Steven G. Avery, Laurent Tellier, Ana Vazquez, Gustavo de los Campos, Stephen D. H. Hsu
TL;DR
The paper addresses the gap between SNP heritability and variance captured by existing genomic predictors for complex quantitative traits. It uses LASSO-based whole-genome regression on large UK Biobank data, with validation in held-out and external samples. The predictors capture substantial variance for height and heel bone density, while educational-attainment prediction remains limited and height performance declines under cross-dataset validation.
Problem
Common-SNP associations previously explained only a small fraction of complex-trait heritability relative to variance suggested by SNP heritability estimates.
Method
The authors use L1-penalized regression to construct genomic predictors from UK Biobank data and assess them in held-out and external validation samples.
Results
The approach captures much of the expected common-SNP heritability, with prediction correlations approaching 0.7 for height and about 0.45 for heel bone mineral density, while educational attainment reaches ∼0.3.
Takeaways & Limitations
Accurate genomic prediction can improve phenotype prediction and support inference about genetic architecture, with potential applications to identifying individuals at elevated genetic risk.
Takeaways & Limitations
GCTA estimates of SNP heritability may be biased, so they should be treated only as a rough guide for evaluating predictor performance.
Abstract
from arXiv · showhide
We construct genomic predictors for heritable and extremely complex human quantitative traits (height, heel bone density, and educational attainment) using modern methods in high dimensional statistics (i.e., machine learning). Replication tests show that these predictors capture, respectively, $\sim$40, 20, and 9 percent of total variance for the three traits. For example, predicted heights correlate $\sim$0.65 with actual height; actual heights of most individuals in validation samples are within a few cm of the prediction. The variance captured for height is comparable to the estimated SNP heritability from GCTA (GREML) analysis, and seems to be close to its asymptotic value (i.e., as sample size goes to infinity), suggesting that we have captured most of the heritability for the SNPs used. Thus, our results resolve the common SNP portion of the "missing heritability" problem -- i.e., the gap between prediction R-squared and SNP heritability. The $\sim$20k activated SNPs in our height predictor reveal the genetic architecture of human height, at least for common SNPs. Our primary dataset is the UK Biobank cohort, comprised of almost 500k individual genotypes with multiple phenotypes. We also use other datasets and SNPs found in earlier GWAS for out-of-sample validation of our results.
1 Introduction
The paper addresses missing heritability by moving from genome-wide association discovery to accurate whole-genome prediction. Using UK Biobank data and L1-penalized regression, it argues that many relevant common-SNP effects can be captured jointly.
- Common SNPs explain substantial heritability for height, heel bone density, and educational attainment, but identified SNPs previously captured only a small fraction of that heritability.
- The authors attribute missing heritability primarily to insufficient statistical power for detecting variants with small effects, low minor-allele frequency, or both.
- Whole-genome prediction seeks the most accurate phenotype predictor and can tolerate a small fraction of false-positive SNPs, unlike GWAS’s focus on reliable individual associations.
- L1-penalized regression constructs predictors by globally optimizing SNP effects in a high-dimensional space, particularly when effects are sparse or approximately sparse.
- The activated SNPs can improve prediction accuracy while helping characterize overall genetic architecture, even when individual variants do not reach genome-wide significance.
2 Data and Methods
The study builds linear genomic predictors from nearly 500k UK Biobank genotypes and phenotypes using LASSO after screening candidate SNPs. The resulting coefficient vector defines a predictive model intended to capture heritable genetic variance.
- Nearly 500k UK Biobank genotypes and associated phenotypes comprise the main dataset.
- LASSO estimates a vector of linear SNP effects by minimizing an objective with an L1 penalty, using standardized phenotypes and genotype values.
- The resulting effects vector defines a linear predictive model that captures a large portion of heritable genetic variance.
- A preliminary single-marker regression screen reduces 645,589 quality-controlled SNPs to the top 50k or 100k candidates by statistical significance.
3 Results
LASSO genomic predictors achieved substantial validation accuracy for height and heel bone density, with height performance approaching an asymptotic limit. The results also show broadly distributed activated SNPs, out-of-sample validity, and uncertainty about GCTA heritability estimates.
- Predictive accuracy: A correlation somewhat below 0.7, corresponding to roughly 50 percent of total variance, was approached for height as training size increased.The correlation was measured in 5,000 validation individuals not used for training optimization.
- Predictive accuracy: Most individuals’ actual heights were within about 3 cm of their predicted values in a 2,000-person held-out sample.
- Predictive accuracy: Educational attainment reached a maximum correlation of ∼0.3 with about 10k activated SNPs, without evidence of approaching a limiting value.The authors state that significantly more or higher-quality data may be required to capture most of its SNP heritability.
- Genetic architecture: Roughly 20k SNPs were activated in the optimal predictors for height and bone density, while expanding candidates from p = 50k to p = 100k increased maximum correlation somewhat.The number of activated SNPs did not change significantly after expanding the candidate set.
- Validation: Out-of-sample testing on ARIC reduced height correlation to ∼0.54, while most ARIC individuals remained within 4 cm or less of predicted height.The authors distinguish this from a ∼5% reduction caused by SNP restrictions and imputation, leaving an additional ∼7% decrease attributed to out-of-sample effects.
- Genetic architecture: Height-model activated SNPs were roughly uniformly distributed across the genome, including many regions not previously identified by earlier GWAS.The specific height predictor based on 50k candidate SNPs achieved a correlation of ∼0.61.
4 Discussion
The discussion contrasts association discovery with optimal phenotype prediction and argues that common-SNP heritability can be captured for complex traits. It also estimates data requirements for extending prediction to highly heritable disease risk.
- 4 Discussion: Much of the expected heritability from common SNPs can be captured even for complex traits affected by thousands of variants.The paper contrasts this prediction goal with earlier emphasis on marker–phenotype association discovery.
- 4 Discussion: The authors suggest that similar predictors may be obtainable for other quantitative traits and highly heritable diseases given enough data and high-quality phenotypes.Examples include cognitive ability and specific disease risks such as Alzheimer’s, Type I Diabetes, obesity, ovarian cancer, and schizophrenia.
- 4 Discussion: About 100,000 cases plus a similar number of controls might enable good prediction of highly heritable disease risk.This estimate applies even when disease risk depends on a thousand or more genetic variants.
- 4 Discussion: For quantitative traits with h2 ∼0.5, simulations predict a LASSO phase transition at n ∼30s.Here, n is the sample size and s is the number of variants with non-zero effects.
A.1 UKBB Dataset QC
The UK Biobank dataset began with nearly half a million genotyped individuals and was progressively quality-controlled before imputation-based analyses and validation.
- A.1 UKBB Dataset QC: 488,377 individuals were initially genotyped at 805,426 SNPs on two Affymetrix platforms.The release combined approximately 50,000 UK BiLEVE Axiom samples with the remainder on the UK Biobank Axiom array.
- A.1 UKBB Dataset QC: 92,060,613 SNPs and 487,411 individuals remained after imputation and initial quality control.Further filtering excluded SNPs and samples with missing call rates exceeding 3% and SNPs with MAF below 0.1%.
- A.1 UKBB Dataset QC: 632,155 SNPs remained after restricting the imputed UK Biobank data to variants shared with the out-of-sample validation data.This filtering supported height validation outside the primary UK Biobank analysis.
A.2 Confounding variables: age, sex and family structure
Phenotypes were adjusted for age, sex, and year of birth, while family structure and cryptic relatedness were evaluated as potential confounders in UK Biobank analyses.
- A.2 Confounding variables: age, sex and family structure: Phenotypes for self-identified Caucasians were sex-specific z-scored and adjusted for year of birth using univariate linear regression.The adjusted phenotype was defined as the residual from the regression on year of birth.
- A.2 Confounding variables: age, sex and family structure: 107,163 familial relationships at the level of third cousins or higher were identified among UK Biobank participants.Filtering these related individuals would substantially reduce the data available for model selection.
- A.2 Confounding variables: age, sex and family structure: LASSO training was performed with and without filtering related individuals to investigate the relevance of family structure and cryptic relatedness.
A.3 L1-penalized regression
The method models phenotype from genome-wide SNP data with a linear, sparse-effects regression and selects a regularized predictor through cross-validation. Its performance depends on sparse effects, sufficient heritability, and enough observations relative to problem complexity.
- A.3 L1-penalized regression: The genotype data are represented as an n × p design matrix, with each SNP coded as 0, 1, or 2 copies of the most frequent minor allele.Missing genotype values are mean-imputed.
- A.3 L1-penalized regression: The model assumes a linear relationship between phenotype and SNP data, with normally distributed errors incorporating environmental and nonlinear genetic effects.The error variance is unknown, and gene–gene and gene–environment nonlinear effects are included in the residual component.
- A.3 L1-penalized regression: LASSO estimates the vector of linear SNP effects by minimizing a standardized regression objective with an L1 penalty controlled by λ.The L1 norm is the sum of the absolute coefficient values.
- A.3 L1-penalized regression: The L1 penalty favors sparse solutions by shrinking nonzero coefficients toward 0, reducing variance and improving expected fit for small samples.This is motivated by the expectation that most SNPs have no effect on a given phenotype.
- A.3 L1-penalized regression: LASSO can recover accurate effects even when n ≪ p if the effects vector is sparse and trait heritability is sufficiently high.Equivalently, the noise variance must be bounded.
- A.3 L1-penalized regression: The cross-validation procedure ranks GWAS SNPs, applies LASSO across penalty values, selects λ by test-set correlation, and evaluates the optimized predictor on validation sets.The workflow splits data into training, test, and validation sets before final evaluation.
A.4 Coordinate Descent
The paper solves the LASSO objective by cycling through coefficients, updating one coordinate at a time while holding the others fixed. Sign checks determine whether the coordinate update is valid or should be set to zero.
- Coordinate descent cycles sequentially through all p coefficients, minimizing the objective with respect to each β_j while holding the others fixed.
- A sign flip during the update indicates a spurious solution, so the optimal coefficient is placed at the boundary.
- The paper presents this procedure as Algorithm 1, the basic coordinate descent algorithm for LASSO.
A.5 Out-of-sample Validation
The authors construct LASSO genomic predictors in UK Biobank and validate them both within UK Biobank and out of sample in ARIC. For height, the predictor achieved correlations of approximately 0.54 in ARIC, with most predictions within 4 cm of actual height.
- Validation design: LASSO predictors were trained on UK Biobank data and evaluated using withheld UK Biobank individuals and an independent ARIC dataset.The independent dataset was used to test validity and check against overtraining.
- Height validation results: 0.61 was the within-UK Biobank correlation using the un-imputed dataset, falling to 0.58 with SNPs shared with ARIC and reaching 0.54 in ARIC participants.Most individuals in the ARIC validation set had actual heights within 4 cm or less of their predictions.
- Validation design: The ARIC validation set contained 9,618 filtered individuals and 632,155 SNPs shared with UK Biobank imputed data.The original ARIC dataset contained 12,772 individuals and 841,820 SNPs before filtering.
- Genetic architecture: The height predictor built from 50,000 candidate SNPs achieved an actual–predicted height correlation of approximately 0.61, while activated SNPs were broadly distributed across the genome.Activated SNP regions significantly overlapped regions near previously known SNPs, despite the genome-wide distribution.
- Height validation results: Figure 6 compares actual and predicted height in 2,000 ARIC individuals, with error bars showing ±1 SD from a larger validation set.The plotted sample contained roughly equal numbers of males and females without age or gender corrections.
- Genetic architecture: Activated height SNPs showed no statistically significant deviation from random effect signs: minor alleles were equally likely to increase or decrease height.The predictor’s activated SNPs were compared with known GWAS hits using variance explained and proxy matching.