Source-linked AI summary

Testing for Associations between Loci and Environmental Gradients Using Latent Factor Mixed Models

Eric Frichot, Sean Schoville, Guillaume Bouchard, Olivier François

arXiv:1205.3347v3q-bio.PEstat.CO

TL;DR

Detecting weak-effect loci involved in local adaptation requires associating genetic polymorphisms with environmental variables while accounting for population structure. The paper develops LFMM algorithms using latent factors for this correction, finds reduced false-positive associations under isolation by distance, and applies the models to plant and human data.

  • Problem

    Classical locus–environment association tests can produce high false-positive rates without correction for population structure, and neutral loci are difficult to identify a priori when selection affects many loci.

  • Method

    LFMMs model environmental variables as fixed effects while using latent factors to estimate background population structure during genome-wide association testing.

  • Results

    LFMMs reduced false-positive associations in simulations of isolation-by-distance patterns and identified 2,624 SNPs, or 0.4% of polymorphisms, associated with temperature gradients in HGDP data.

  • Takeaways & Limitations

    LFMM provides a computationally efficient genome-scanning approach for detecting local-adaptation signals while accounting for population history and isolation by distance.

  • Takeaways & Limitations

    The approach is difficult to apply when selection acts on phenotypes at many loci because selectively neutral SNPs are difficult to choose a priori for comparison methods.

Abstract

from arXiv · show

Adaptation to local environments often occurs through natural selection acting on a large number of loci, each having a weak phenotypic effect. One way to detect these loci is to identify genetic polymorphisms that exhibit high correlation with environmental variables used as proxies for ecological pressures. Here, we propose new algorithms based on population genetics, ecological modeling, and statistical learning techniques to screen genomes for signatures of local adaptation. Implemented in the computer program "latent factor mixed model" (LFMM), these algorithms employ an approach in which population structure is introduced using unobserved variables. These fast and computationally efficient algorithms detect correlations between environmental and genetic variation while simultaneously inferring background levels of population structure. Comparing these new algorithms with related methods provides evidence that LFMM can efficiently estimate random effects due to population history and isolation-by-distance patterns when computing gene-environment correlations, and decrease the number of false-positive associations in genome scans. We then apply these models to plant and human genetic data, identifying several genes with functions related to development that exhibit strong correlations with climatic gradients.

1 Introduction

The paper introduces LFMMs to detect loci correlated with environmental variables while accounting for hidden population structure. This addresses false positives and circularity concerns in genome scans for local adaptation.

  • Motivation: Environmental correlations can reveal loci involved in local adaptation when beneficial alleles have weak phenotypic effects.The approach uses environmental variables as proxies for ecological pressures.
  • Motivation: Geographic structure, gene flow, drift, and demographic history can confound locus–environment associations and produce false positives.Classical regression tests are especially prone to high false-positive rates without population-structure or isolation-by-distance corrections.
  • Motivation: Empirical covariance-based corrections can be circular because they require identifying neutral loci before testing for adaptation.This prerequisite reduces the power to reject neutrality when selection affects many loci.
  • Contribution: LFMMs model environmental effects while estimating hidden factors that represent residual population structure.The models extend probabilistic principal component analysis and statistical learning methods.
  • Contribution: The algorithms control for population history and spatial autocorrelation when estimating gene–environment associations and are applied to plant and human data.The implementation is designed for fast analysis of very large polymorphism datasets.

2 Method

LFMM represents allele-frequency data with environmental fixed effects and latent factors for background structure. Low-rank factorization and computational algorithms enable scalable estimation of associations across many loci.

  • Model: The model treats allele frequencies as responses in a regression mixed model with locus effects, environmental coefficients, latent factors, and residual errors.The latent factors capture variation not explained by measured environmental variables.
  • Model: LFMM introduces environmental variables as fixed effects and population structure as latent factors.The model is named the Latent Factor Mixed Model.
  • Factorization: Matrix factorization decomposes the centered genetic data into individual scores and locus loadings, with K chosen below the number of loci or individuals.For K below 50, the approach is essentially a low-rank approximation of covariance structure.
  • Computation: The Gibbs sampler estimates factor scores, loadings, environmental effects, and biases using products of low-dimensional matrices.Its computational design targets datasets of roughly n ≈ 1,000 individuals and L ≈ 500,000 SNPs.
  • Comparison: LFMM differs theoretically from Bayenv because it infers the factor matrix with a fully Bayesian algorithm rather than an empirical Bayes algorithm.Both approaches model background covariance, but their factor-inference procedures differ.

3 Simulation Study

Simulations evaluated LFMM calibration, estimation error, and outlier detection against regression, principal-components, and Bayenv approaches under varying population structure. LFMM produced well-calibrated tests, smaller estimation errors, and stronger detection performance across the tested settings.

  • Simulation design: The simulations tested LFMM calibration and performance against standard regression, PC regression, and other mixed models under varying latent-factor structure.The experiments examined null P-value distributions, environmental associations, estimation errors, and detection of selected loci.
  • Null calibration: For K = 5 and K = 20, LFMM tests produced small numbers of false positives, with slightly conservative behavior at K = 20.For values of K lower than 5, the ECDF was close to uniform; at K = 20, tests were slightly conservative.
  • Null calibration: Under neutral variation, LM and GLM P-values departed strongly from uniformity and produced many false positives, whereas LFMM and PCRM were much closer to uniform.With K = 7 factors or principal components, LFMM and PCRM showed improved null calibration; Tracy-Widom-based K selection was slightly conservative in these simulations.

4 Data Analysis

LFMM was applied to loblolly pine and human genomic data using climatic variables summarized from environmental measurements. The analyses identified SNPs and genes correlated with climatic gradients, including genes involved in stress responses, disease or traits, and development.

  • Loblolly pine: In loblolly pine, 113, 30, and an unspecified lower count of SNPs exceeded |z|-scores of 3, 4, and 5, respectively.The analysis treated |z| > 4 as significant and found 17 shared loci among the 50 highest-scoring loci compared with Bayenv.
  • Loblolly pine: Seven of 10 SNPs with Bayes factors greater than 10^3 were confirmed by LFMM, including the two highest-Bayes-factor SNPs for the first and second environmental variables.LFMM also identified additional associations, including proteins involved in photosynthesis, oxidative and salt stress, and temperature-stress responses.
  • Human data: Among human SNPs with |z|-scores greater than 5, 28 were GWAS-SNPs with known disease or trait associations, including variants near OCA2 and SLC45A2.Reported associations included celiac disease, height, and vitamin D synthesis or activation.
  • Human data: Among 65 human SNPs with |z|-scores greater than 7, 31 were located within genes, including several associated with multicellular organismal development.Examples included EPHB4, NRG1, RBM19, EYA2, and POLA1, with reported |z|-scores from 7.04 to 8.90.

5 Discussion

The discussion presents LFMMs as matrix-factorization gene–environment tests that jointly model environmental effects and latent population structure, reducing false positives under isolation-by-distance. It also highlights applications to plant and human data, while noting sensitivity to latent-factor choice and assumptions about population structure.

  • Simulation results show that LFMM tests reduce false-positive associations in the presence of isolation-by-distance patterns.Simple linear or logistic regression can be misleading under isolation by distance.
  • LFMMs estimate latent factors and regression coefficients simultaneously, providing a unified framework for environmental and demographic effects.
  • Bayenv requires a priori separation of neutral and adaptive SNPs, which is difficult when selection affects many loci and may cause associations to be overlooked.The authors report that selecting neutral SNPs was extremely difficult for the Loblolly pine data.
  • The low-rank covariance approximation makes LFMM computationally faster than Bayenv for large analyses.LFMMs model correlations between environmental predictors and allele frequencies while hidden factors explain residual genetic variation.
  • Choosing too many latent factors makes tests highly conservative and reduces power to reject neutrality.The authors select K using Tracy–Widom theory rather than computationally intensive cross-validation, obtaining slightly conservative tests.
  • Applications to Pinus taeda and human HGDP data identified climate-linked variants, including genes related to development, pigmentation, stress responses, and other biological functions.In the HGDP data, 0.4% of polymorphisms, or 2,624 SNPs, were significantly associated with temperature gradients at |z| > 5.

Supplementary Text 2

The supplementary figures compare empirical cumulative distributions for LFMM and PC regression tests across different numbers of latent factors.

  • LFMM test distributions are evaluated under generative-model simulations with K = 1, 3, 5, 10, and 20 latent factors.
  • LFMM test distributions are also examined for K = 1, 3, 5, 7, 10, and 20 latent factors.
  • PC regression test distributions are examined for K = 1, 3, 5, 7, 10, and 20 latent factors.

Tables

The tables report simulation error and false-positive or false-negative rates, along with SNP annotations and climatic variables for loblolly pine and human analyses.

  • Table 1 reports mean squared errors for estimates of environmental effects.
  • Table 2 reports false-negative and false-positive percentages for linear, PCR, and LFM models across type I error levels.
  • Tables 3 and S1 annotate loblolly-pine SNPs with absolute z-scores greater than 4.
  • Tables 4 and S3 list HGDP SNPs with high z-scores associated with phenotypic traits or genes with molecular and biological functions.
  • Table S2 lists the climatic variables used in the HGDP analysis.
  • Table S3 (bis) is included as an additional human-data table.

Gibbs Sampling algorithm for the LFMM

The LFMM Gibbs-sampling algorithm specifies hierarchical prior and conditional distributions, then iteratively samples parameters and averages post-burn-in draws.

  • Model notation: The model defines D as the number of environmental variables and Ii,ℓ as an indicator for whether data are observed.
  • Prior distributions: The LFMM priors include normal distributions for genetic effects, latent factors, locus factors, and intercepts, with variance parameters updated during sampling.
  • Conditional distributions: The hierarchical model is described through conditional distributions for its parameters.
  • Main algorithm: After burn-in, the algorithm computes posterior means for latent factors and environmental-effect statistics from iterations through nIter.
  • Main algorithm: Each iteration updates the residual variance and samples environmental effects, individual latent factors, and locus-specific factors.
Loading 1205.3347v3…