Source-linked AI summary
Using haplotype differentiation among hierarchically structured populations for the detection of selection signatures
Marìa Inès Fariello, Simon Boitard, Hugo Naya, Magali SanCristobal, Bertrand Servin
TL;DR
Selection scans often rely on allele-frequency differentiation and may not capture haplotype structure or hierarchical relationships among populations. The paper introduces hapFLK, which combines both features, and shows improved detection across selection scenarios, including soft and incomplete sweeps, with applications to six sheep breeds. Its scope includes assumptions about population structure and simulation settings that motivate further modeling extensions.
Problem
Selection scans need to detect differentiated loci while accounting for haplotype information and hierarchical structure among sampled populations.
Method
hapFLK extends FLK by using local haplotype clusters from a multipoint linkage-disequilibrium model to measure differentiation across hierarchically structured populations.
Results
hapFLK provides greater detection power than single-SNP FST and XP-EHH in dense-genotyping simulations, including hard and soft sweeps, and detects a wide range of selection events.
Takeaways & Limitations
The combined use of haplotype information and hierarchical population structure supports detection of soft, incomplete, multi-population, and haplotype-targeted selection.
Takeaways & Limitations
The simulations used one selected site, a setting favorable to single-SNP approaches, and the population-frequency model could be extended with more relaxed hypotheses.
Abstract
from arXiv · showhide
The detection of molecular signatures of selection is one of the major concerns of modern population genetics. A widely used strategy in this context is to compare samples from several populations, and to look for genomic regions with outstanding genetic differentiation between these populations. Genetic differentiation is generally based on allele frequency differences between populations, which are measured by Fst or related statistics. Here we introduce a new statistic, denoted hapFLK, which focuses instead on the differences of haplotype frequencies between populations. In contrast to most existing statistics, hapFLK accounts for the hierarchical structure of the sampled populations. Using computer simulations, we show that each of these two features - the use of haplotype information and of the hierarchical structure of populations - significantly improves the detection power of selected loci, and that combining them in the hapFLK statistic provides even greater power. We also show that hapFLK is robust with respect to bottlenecks and migration and improves over existing approaches in many situations. Finally, we apply hapFLK to a set of six sheep breeds from Northern Europe, and identify seven regions under selection, which include already reported regions but also several new ones. We propose a method to help identifying the population(s) under selection in a detected region, which reveals that in many of these regions selection most likely occurred in more than one population. Furthermore, several of the detected regions correspond to incomplete sweeps, where the favourable haplotype is only at intermediate frequency in the population(s) under selection.
INTRODUCTION
Detecting selection signatures helps explain population divergence and identify genes and traits shaped by adaptation or artificial selection. Existing approaches use allele frequencies, haplotypes, or differentiation among populations, but the paper develops hapFLK to combine haplotype information with hierarchical population structure.
- Motivation: Selection-signature detection informs studies of population divergence, disease resistance, climate and altitude adaptation, and livestock traits.These applications span biomedical and agricultural population genetics.
- Existing approaches: Selection-detection methods target derived-allele frequency, haplotype structure, or genetic differentiation, with efficiency varying across selection time scales.Haplotype-based methods include EHH and related statistics, while differentiation-based methods compare populations.
- Existing approaches: FST scans can be biased and produce false positives when populations differ in effective size or share hierarchical ancestry.Raw FST implicitly assumes equal effective sizes and independent derivation from one ancestral population.
- Research gap: Dense genotyping makes correlations among adjacent markers relevant, motivating methods that combine multiple populations with haplotype information.The paper notes that existing combined strategies did not account for hierarchical population structure.
- Contribution: hapFLK extends FLK by clustering chromosomes into local haplotypes and measuring differentiation from their estimated frequencies across populations.The method uses a multipoint linkage disequilibrium model to form local haplotype clusters.
- Study design: The study evaluates haplotype-level differentiation through simulations and applies the approach to six sheep breeds, including a strategy for identifying populations under selection.The analyses compare hapFLK with FST, FLK, and XP-EHH and seek selection signals in dense genotyping data.
METHODS
The methods extend differentiation tests from SNPs to haplotype clusters while accounting for hierarchical population structure, and evaluate performance through simulations and empirical calibration.
- Population-structure correction: FLK models population differentiation through a kinship matrix whose entries represent drift accumulated along the population tree.The matrix is estimated from genome-wide Reynolds’ distances and a Neighbor-Joining tree when demographic parameters are unknown.
- Multiallelic markers: The multiallelic FLK statistic treats haplotypes or other multiallelic markers as allele-frequency vectors across populations.Ancestral allele frequencies and cross-allele covariance terms are incorporated, with the covariance matrix inverted using a Moore-Penrose generalized inverse.
- Haplotype test: hapFLK obtains local haplotype frequencies by clustering similar haplotypes with a multipoint linkage-disequilibrium model.Posterior cluster probabilities are averaged across individuals within each population, allowing clusters to serve as alleles while retaining correlations among nearby markers.
- Haplotype test: The haplotype statistic modifies the multiallelic FLK covariance calculation by replacing B0 with the identity matrix because small ancestral frequencies caused numerical instability.Simulations found this version more powerful than the formulation including B0.
- Simulation evaluation: The study compares hapFLK with FST, FLK, hapFST, and XP-EHH using simulations matched to dense Sheep HapMap genotyping data.Scenarios vary population number, hierarchical structure, selection, bottlenecks, and migration; real-data p-values use an empirical approach because neutral simulation calibration was considered unsuitable under ascertainment bias.
- Simulation evaluation: Simulation power was defined as the proportion of selected replicates whose regional maximum statistic exceeded the corresponding null-distribution quantile.The maximum statistic was recorded across the 5Mb region for each replicate.
RESULTS
Simulations show that haplotype information and accounting for hierarchical population structure each improve selection detection, with hapFLK combining both advantages. The method remains robust to bottlenecks and migration and identifies interpretable selection signals in Northern European sheep.
- Haplotype information: hapFLK provides more detection power than single-SNP FST for hard and soft sweeps in dense-genotyping simulations.XP-EHH is more powerful than FST but less powerful than hapFLK for hard sweeps, while its power decreases more strongly for soft sweeps.
- Haplotype information: At an initial selected-allele frequency of 20% and final frequency of 90%, hapFLK detection power exceeds 75% at a 1% type I error rate.Power remains around 60% for low-frequency mutations reaching only 50–60% frequency, despite the resulting incomplete sweep.
- Hierarchical structure: In four hierarchically structured populations, combining haplotype information and population structure in hapFLK provides greater power than using either feature alone.The haplotype-based and structure-aware gains are each similar in size, while classical FST is least powerful.
- Hierarchical structure: Testing all population pairs is always less powerful than applying hapFLK jointly to the four populations.This supports analyzing jointly sampled populations rather than relying only on pairwise tests.
- Robustness: Demographic deviations from pure drift have little effect on the hapFLK distribution after conditioning on the estimated covariance matrix.A severe bottleneck increases an estimated branch length by 10%, while one-way and two-way migration reduce Reynolds distance by 5% and 10%, respectively.
- Sheep application: In six Northern European sheep breeds, hapFLK identifies seven genome-wide significant regions, whereas FLK provides little evidence for sweeps.Local population trees help identify the populations under selection and reveal both shared selection and signals without hard sweeps.
DISCUSSION
The discussion emphasizes hapFLK’s combined use of haplotype information and hierarchical population structure, while examining its applicability, empirical findings, and limitations. Results include stronger detection of incomplete sweeps and selection signatures in sheep, but interpretation remains sensitive to model assumptions and SNP ascertainment bias.
- Haplotype versus single marker differentiation tests: Haplotype-based tests improve detection power for dense genotyping data when the selected site itself is generally unobserved.The simulations did not show the same improvement for sequencing data, although multi-locus selection could favor haplotype-based tests.
- Different strategies for the inclusion of haplotype information in differentiation: hapFLK estimates local haplotype clusters from genotype data and treats those clusters as alleles in haplotype-based differentiation tests.The implementation uses fastPHASE, while other clustering models such as Beagle could also be used.
- Multiple-population analysis: Joint analysis of multiple populations avoids the multiple-testing penalty of pairwise comparisons and can improve ancestral-frequency estimation.A meta-population approach can reveal selection shared by closely or distantly related populations, although identifying the selected populations becomes more difficult.
- Robustness and p-values: The neutral distribution of hapFLK can be affected by both demography and SNP ascertainment bias, whose effects are more complex for haplotype tests.The authors estimate the null distribution empirically with an outlier-robust estimator, but state that validity must be checked for each dataset.
- Empirical sheep results: In the sheep data, hapFLK produced p-values down to 10^-13, compared with approximately 10^-4 for FLK on the same dataset.The authors caution that distributional choices may artificially inflate the difference, but argue that the data support those choices and simulations show FLK can reach at least 10^-11.
- Soft or incomplete sweeps: hapFLK is reported to retain reasonable power for favorable alleles initially reaching frequencies up to 30% and for incomplete sweeps at intermediate frequency.Several selection signatures in the Sheep HapMap data involved intermediate-frequency selected haplotypes.
- Interpretation of sheep signals: The discussion notes that few hard sweeps were detected in sheep, potentially because artificial selection on quantitative traits is generally polygenic.Other proposed explanations include short divergence times and changes in selection intensity or direction over time.
- General limitations: Selection in functional and non-functional regions may involve more complex constraints than those represented in typical simulations.The authors specifically mention purifying, background, polygenic, and balancing selection as factors requiring further study.