Source-linked AI summary
Bayesian Test for Colocalisation Between Pairs of Genetic Association Studies Using Summary Statistics
Claudia Giambartolomei, Damjan Vukcevic, Eric E. Schadt, Lude Franke, Aroon D. Hingorani, Chris Wallace, Vincent Plagnol
TL;DR
Interpreting the molecular basis of genetic associations requires testing whether signals at the same locus share a causal variant. This paper develops a Bayesian colocalisation test using summary statistics and finds six previously reported lipid–eQTL pairs unsupported by colocalisation.
Problem
Although genetic associations are replicated, identifying their underlying molecular mechanisms and determining whether signals share a causal variant remains challenging.
Method
The paper develops a Bayesian colocalisation test that uses single-SNP summary statistics and analytically computed posterior probabilities to assess shared causal variants.
Results
Six previously reported lipid–eQTL pairs were identified as potential false positives because the analysis did not support colocalisation.
Takeaways & Limitations
The framework supports systematic comparison of association studies to identify candidate shared mechanisms across diseases, biomarkers, and eQTLs.
Takeaways & Limitations
A low posterior probability for colocalisation may reflect limited power rather than evidence against colocalisation, and a high posterior probability measures correlation rather than causality.
Abstract
from arXiv · showhide
Genetic association studies, in particular the genome-wide association study design, have provided a wealth of novel insights into the aetiology of a wide range of human diseases and traits. The next challenge consists of understanding the molecular basis of these associations. The integration of multiple association datasets, including gene expression datasets, can contribute to this goal. We have developed a novel statistical methodology to assess whether two association signals are consistent with a shared causal variant. An application is the integration of disease scans with expression quantitative trait locus (eQTL) studies, but any pair of GWAS datasets can be integrated in this framework. We demonstrate the value of the approach by reanalysing a gene expression dataset in 966 liver samples with a published meta-analysis of lipid traits including >100, 000 individuals of European ancestry. Combining all lipid biomarkers, our reanalysis supported 29 out of 38 reported colocalisation results with eQTLs and identified 14 new colocalisation results, highlighting the value of a formal statistical test. In two cases of reported eQTL-lipid pairs (IFT172, TBKBP1) for which our analysis suggests that the eQTL pattern is not consistent with the lipid association, we identify alternative colocalisation results with GCKR and KPNB1, indicating that these genes are more likely to be causal in these genomic intervals. A key feature of the method is the ability to derive the output statistics from single SNP summary statistics, hence making it possible to perform systematic meta-analysis type comparisons across multiple GWAS datasets (http://coloc.cs.ucl.ac.uk/coloc/). Our methodology provides information about candidate causal genes in associated intervals and has direct implications for the understanding of complex diseases and the design of drugs to target disease pathways.
1 UCL Genetics Institute, University College London (UCL), Darwin Building, Gower Street, London WC1E 6BT, UK · Introduction
GWAS and eQTL studies have generated extensive association data, motivating formal assessment of whether independent signals at one locus share a causal variant. The paper develops a Bayesian colocalisation test that models local linkage disequilibrium and uses summary statistics for rapid genome-wide comparisons.
- Introduction: Hundreds of genomic loci affecting complex diseases and intermediate phenotypes have been identified and replicated through GWAS, while gene expression is studied as an eQTL outcome trait.eQTL analyses use expression measurements from microarray or RNA sequencing studies.
- Introduction: Colocalisation asks whether two independent association signals at the same locus are consistent with a shared causal variant.A positive result increases the likelihood that both traits share a causal mechanism and can link disease associations to causal genes and tissues.
- Introduction: Visual overlap of association signals is insufficient because abundant eQTLs across the genome and tissues make accidental overlaps likely.This creates a need for statistical methods that distinguish genuine shared signals from coincidental overlap.
- Introduction: Existing approaches can account for local LD or match expression and GWAS patterns, but they do not directly provide the formal colocalisation test sought here.Nica et al. rank SNPs using residual association conditional on the most associated SNP, whereas Sherlock addresses a different hypothesis and can use p-values.
- Introduction: The authors derive a novel Bayesian statistical test for colocalisation between pairs of traits, focusing on one genomic region at a time and interpreting its LD pattern.The method is designed to address shortcomings of existing tools.
- Introduction: The approach requires single-SNP p-values and MAFs, or estimated allelic effects and standard errors, and uses closed-form analytical results for rapid genome-wide comparisons.Its model is closely related to a method for discovering eQTLs across multiple tissues.
Results
The Bayesian method assigns posterior support across five causal configurations, distinguishing shared from independent variants and enabling rapid summary-statistic analysis. Simulations and lipid–eQTL reanalysis demonstrated strong performance, identified new colocalisations, and highlighted alternative candidate genes and conditional signals.
- Bayesian colocalisation framework: PP4 > 80% supports a single shared variant, whereas PP3 > 90% supports two independent causal SNPs affecting the traits.The method integrates likelihoods and priors across configurations to produce posterior probabilities PP0–PP4.
- Computational performance: Q^d computation using approximate Bayes factors requires no iterative scheme or specialised infrastructure, enabling rapid analysis of genomic regions.Here, Q is the number of variants and d the number of distinct associations, typically d = 2 for two traits.
- Simulation comparison: PP4 or PP3 > 0.9 correctly identified simulated shared or distinct causal-variant hypotheses, with PP3 > 0.9 slightly outperforming proportional testing at p-value < 0.05.The comparison used a biomarker dataset with 10,000 samples.
- Lipid–eQTL reanalysis: Three reported genes failed colocalisation, while nearby SORT1, GCKR, and KPNB1 showed high colocalisation probability and became more likely regional candidates.Conditional analysis also found evidence for expression colocalisation at two loci involving secondary signals.
- Lipid–eQTL reanalysis: 14 colocalised loci involving 15 genes were newly identified with PP4 > 75% beyond previously reported lipid–eQTL results.Eleven of the 15 genes were described as strong or previously suggested candidates for lipid metabolism.
Discussion
The discussion highlights the method’s efficient summary-statistic framework, while emphasizing data-quality requirements, modeling assumptions, cautious posterior interpretation, and limits imposed by biological context. It also describes extensions to complex loci and broader applications across association studies.
- Method: The Bayesian procedure tests whether two association signals plausibly share a causal variant, using analytical forms and single-variant p-values without identifying the causal variant itself.This distinguishes colocalisation testing from fine-mapping, which seeks the most likely causal variant.
- Data quality: Accurate colocalisation requires high-density genotyping or accurate imputation, because imputation uncertainty is not captured by variance estimates based only on allele frequency and sample size.For binary outcomes, the case-control ratio also enters the variance estimate.
- Assumptions: The framework currently assigns equal variant priors and uses a 10−4 prior probability for both cis-eQTLs and lipid associations.Location-specific priors could incorporate gene distance or functional elements, while trans-eQTLs may require a more stringent threshold.
- Interpretation: Posterior probabilities require caution: low PP4 can reflect limited power when PP3 is also low, whereas high PP4 indicates correlation rather than causality.High PP0, PP1, and/or PP2 can evidence limited power.
- Multiple associations: Conditional p-values can address multiple independent associations at one locus, but top-SNP selection may introduce bias, although this bias is small with large samples and/or strong effects.For difficult loci with multiple associations for both traits and available genotype data, estimating Bayes factors may be more appropriate.
- Biological scope: Colocalisation depends on expression biology and tissue choice: GWAS signals are explained by eQTLs when causal variants alter mRNA amount, and results vary with the expression dataset.The method is intended for systematic use across multiple GWAS and eQTL studies, cell types, traits, and association-study pairs to clarify GWAS mechanisms and disease aetiology.
Materials and Methods
The analysis integrated gene expression and genotype data from 966 human liver samples with eQTL and biomarker association statistics. Colocalisation was evaluated using Bayesian hypothesis posterior probabilities derived from configuration likelihoods, Bayes factors, and SNP-level priors.
- Data and samples: 966 human liver samples from unrelated European-American subjects were genotyped with Illumina 650Y BeadChip arrays, while 39,000 expression probes were profiled on Agilent arrays.Samples came from two non-overlapping studies and were collected post-mortem or during surgical resection.
- Genotype processing: Quality-control filters removed individuals and SNPs with more than 10% missing genotypes, SNPs with HWE p-values below 0.001, and post-imputation monomorphic SNPs.Imputation was performed on separately phased genomic chunks that were subsequently re-assembled.
- Association analyses: eQTL statistics were estimated by linear trend regression for variants within 200 kilobases of each probe after filtering rare, monomorphic, multi-allelic, and poorly imputed variants.Variants with MAF < 0.001 or imputation Rsq < 0.3 were excluded, along with variants not sufficiently well imputed between both datasets.
- Bayesian colocalisation model: The Bayesian model grouped SNP-association configurations into five sets corresponding to hypotheses H0 through H4 and computed posterior probabilities from configuration likelihoods weighted by prior probabilities.The posterior probability for each hypothesis was obtained by comparing its likelihood with the summed likelihoods across the five hypotheses.
- Bayes factors and priors: Approximate Bayes Factors measured each SNP’s relative support for association with either trait versus no association, using p-values or regression estimates and variances.SNP-level priors were assigned as 1 × 10^-4 for p1 and p2 and 1 × 10^-6 for p12.
Figures
The figures illustrate the hypotheses, visual interpretation, and simulation behavior underlying Bayesian colocalisation analysis, alongside an LDL–SYPL2 locus example with conditional eQTL results.
- Colocalisation results: Figure 2 contrasts negative FRK–LDL colocalisation (PP3 > 90%) with positive SDC1–total cholesterol colocalisation (PP4 > 80%).The plots show -log10(p) association values for biomarker and expression datasets across a 1Mb locus range.
- Simulation analyses: Figure 3 summarizes PP4 distributions for simulations with a shared causal variant between a 966-sample eQTL study and a biomarker study.The biomarker’s variance explained is colour coded, sample size is on the x-axis, and the y-axis reports median, 10% and 90% PP4 quantiles.
- Simulation analyses: Figure 4 evaluates PP4 across 1,000 simulations while assessing the impact of restricting variants from the MetaboChip Illumina array to the Illumina 660W array.Both studies have equal variance explained, with an eQTL sample size of 966 and biomarker sample size of 4,000.
- Simulation analyses: Figure 5 compares proportional and Bayesian colocalisation across simulated scenarios with varying causal-variant configurations and typed or imputed variants.Full, top-shaded, and bottom-shaded circles denote variants affecting both traits, the eQTL trait only, or the biomarker trait only, respectively.
- Locus example: Figure 6 displays LDL and SYPL2 eQTL association signals, including SYPL2 expression conditional on the top eQTL SNP rs2359653.The LDL association uses a published meta-analysis of > 100,000 individuals, while the eQTL data comprise 966 liver samples.
Tables
The tables identify previously reported liver-eQTL colocalisations not supported by the method and novel loci supported by the analysis. They define these categories using posterior-probability thresholds and a strict prior for shared associations.
- Previously reported but unsupported loci: Table 1 lists loci previously reported to colocalise with liver eQTLs but not supported by the analysis.These associations had PP3, the posterior probability for distinct signals, greater than 75%.
- Previously reported but unsupported loci: The analysis reports other genes in the same regions when they colocalise using the method.Secondary signals are included only when a secondary eQTL has p-value greater than 10−4, with tests computed after conditioning expression data on the listed SNP.
- Novel colocalising loci: Table 2 identifies novel loci not previously reported to colocalise with liver eQTLs but supported by the analysis.The listed signals had PP4 greater than 75% for colocalisation between liver eQTLs and the Teslovich meta-analysis of LDL, HDL, TG, and TC, using p12 = 10−6.
- Novel colocalising loci: For 11 genes with strong candidate status for lipid metabolism, Table 2 lists a key reference describing their function.These genes were among the signals not reported by Teslovich et al.
1 UCL Genetics Institute, University College London (UCL), Darwin Building, Gower Street, London WC1E 6BT, UK
The paper is affiliated with the UCL Genetics Institute at University College London, and the cited version is dated 4 November 2013.
- Affiliation and version: The cited manuscript version is arXiv:1305.4022v3, dated 4 Nov 2013.The affiliation is UCL Genetics Institute, University College London, Darwin Building, Gower Street, London WC1E 6BT, UK.
Supplementary material … Relationship between model parameters and LD using imputed genotypes
The supplementary derivations define the coloc prior model and ABF calculations from summary statistics, specify prior and simulation settings, and relate causal-marker effects to LD with observed or imputed genotypes. Under imputation, posterior-mean genotype substitution and stated conditional-linearity assumptions yield a formula analogous to the observed-genotype result.
- Supplementary material: The model partitions each SNP into mutually exclusive probabilities p0, p1, p2, and p12 for neither, trait-specific, or shared association.When priors are constant across SNPs, configurations within each hypothesis set share the same prior probability, enabling the simplification.
- Supplementary material: Conditioning on at most one association per trait requires normalized configuration probabilities, but the normalizing constant cancels in prior-odds ratios.The derivation uses p0 ≈1 when dividing configuration probabilities by P(S0).
- ABF 2: ABFs combine summary-statistic evidence under β = 0 with a normally distributed alternative β ~ N(0, W), using asymptotic regression estimates.The derivation assumes unrelated individuals and independent effect sizes for the two traits, with the intercept removed as a nuisance parameter.
- ABF 2: The variance V is approximated from allele frequency, sample size, and case-control ratio, under Hardy–Weinberg equilibrium and normalized continuous-trait variance.Case-control outcomes are modeled as binomial, whereas continuous outcomes are normalized so Var(Y) = 1.
- Choice of priors for the probability that each variant affects the traits: The conditional prior p12/(p12 + p1) gives the probability of trait-2 association among trait-1-associated SNPs; with p12 = 1 × 10−6 and p1 = p2 = 1 × 10−4, this is 1 in 100.The corresponding null prior is p0 = 1 −(p1 + p2 + p12) ≈0.9998 ≈1.
- Choice of priors for the standard deviation W of the effect size parameter β: The effect-size prior uses standard deviation 0.15 for continuous traits and W = 0.2 for case-control log-odds ratios, corresponding to an odds-ratio interquartile range of 0.74–1.35.For a continuous trait, the 0.15 prior corresponds to approximately 0.01 variance explained at MAF 30%.
- Simulation procedure: Simulations select a causal SNP with Rsq > 0.8, generate phenotypes from its effect and unit-variance noise, and rescale estimation error to emulate larger samples and derive new p-values.The simulations assess sample-size requirements for colocalisation and consequences of limited variant density.
Simulations to find sample size required for colocalisation analysis · Simulations to find consequence of using limited variant density
Simulations assessed how biomarker sample size and causal-variant variance explained affect PP4, using the original eQTL sample size and expression variance explained. Additional simulations evaluated reduced variant density and omission of the causal SNP by comparing dense imputed data with Illumina 660K variants.
- Simulations to find sample size required for colocalisation analysis: The simulations compared PP4 distributions across randomly sampled regions under varying biomarker sample sizes and causal-variant variance explained.The eQTL dataset was fixed at n=966, with expression variance explained set to 10%.
- Simulations to find sample size required for colocalisation analysis: The eQTL simulation design retained the original sample size of n=966.Expression variance explained was fixed at 10% while biomarker-study parameters varied.
- Simulations to find sample size required for colocalisation analysis: The biomarker simulation varied sample size to examine its effect on PP4 across randomly sampled regions.The procedure also varied the proportion of biomarker variance explained by the causal variant.
- Simulations to find consequence of using limited variant density: PP4 was compared before and after filtering each region to SNPs present on the Illumina 660K Chip.This comparison used the same randomly sampled regions and the simulation procedure described above.
- Simulations to find consequence of using limited variant density: The density simulation contrasted > 990genomes imputed data with the sparser Illumina dataset.The imputed dataset was expected to contain substantially denser variant coverage than the Illumina dataset.
- Simulations to find consequence of using limited variant density: A separate simulation excluded the causal SNP only from the Illumina dataset to model its absence from the available data.This procedure evaluated the consequence of missing the causal variant under limited variant density.
- Simulations to find consequence of using limited variant density: All analyses for the limited-density simulations were conducted in R [11].The causal-SNP omission scenario followed the same simulation procedure while excluding that SNP only from the Illumina dataset.
Simulations for comparison with existing colocalisation tests
Simulations compared proportional and Bayesian colocalisation approaches using sampled haplotypes from 49 type 1 diabetes susceptibility regions outside the MHC. The regions reflected typical GWAS-hit variation in size and genomic topography, while excluding the MHC because of its complex genetic architecture.
- Simulation design: Simulations sampled, with replacement, phased haplotypes of SNPs with minor allele frequency ≥5% from >990 Genomes Project data.The haplotypes came from data spanning genomic regions outside the MHC.
- Simulation design: 49 type 1 diabetes susceptibility regions outside the MHC were included, representing a range of region sizes and genomic topography typical of GWAS hits.The regions were identified as T1D susceptibility loci and summarized in T1DBase.
- Simulation design: The MHC was excluded because its high variation, strong LD, and multiple-locus genetic influence on autoimmune disease risk require individual treatment in GWAS.For each trait, one or two causal variants were selected at random.
Supplementary Figures and Tables
Supplementary figures assess colocalisation performance under simulated missing or imputed variants and examine the relationship between PP4 and proportional-testing p-values. Additional regional plots and tables document reported, unsupported, and main-text colocalisation loci.
- Simulation analyses: Simulation analyses compare imputed and non-imputed data for shared causal variants, including scenarios where the causal SNP is absent from one dataset.The simulations use an eQTL study with 966 samples and a biomarker study across different sample sizes, with causal-variant variance explained colour coded.
- Proportional testing: Supplementary analyses relate PP4 to posterior predictive p-values from proportional testing using Bayesian model averaging across possible two-SNP models.A second analysis restricts SNPs to those on the Illumina HumanOmniExpress array and specifies eQTL and biomarker sample sizes and variance explained.
- Regional analyses: Regional Manhattan plots show loci reported as having probable shared variants but not supported by the method at PP4 < 75%.The plots span approximately 400 kilobases around each specified expression probe and compare published lipid-association and liver-eQTL signals.
- Colocalisation tables: Supplementary tables report colocalisation results for liver eQTLs with LDL, HDL, TG, and TC, including posterior probabilities for different and common signals.Table S1 covers reported loci correlating with liver expression and one of the four lipid traits, with multiple probes represented for some loci.
- Regional analyses: Additional regional Manhattan plots correspond to loci listed in Table 2 of the main text, with genomic ranges extended beyond approximately 400 kilobases when useful for visualisation.The supplementary figures also label regional examples involving genes and lipid traits such as SDC1/TC, TGOLN2/HDL, and INHBB/LDL.
Overview of gene function of new colocalisation results associated with blood lipid levels and liver expression
The new colocalisation results implicate genes with established roles in lipid transport, metabolism, homeostasis, and related liver or plasma processes. Other candidates connect to cardiovascular risk, innate immunity, oxygen adaptation, or drug-responsive lipid regulation.
- Lipid transport and metabolism: SDC1 mediates hepatic clearance of triglyceride-rich lipoproteins, while TGOLN2 regulates cholesterol transport to the trans-Golgi network and plasma membrane caveolae.These functions directly connect both genes to lipoprotein and cholesterol handling.
- Other biological links: INHBB has been associated with elevated lipid levels and cardiovascular disease risk traits, while HP and HPR have prior lipid associations and plasma or innate-immune functions.HP transports free hemoglobin to the liver, whereas HPR associates with apolipoprotein-L-I-containing HDL particles and innate immunity.
- Lipid homeostasis and cellular regulation: UBXN2B participates in endoplasmic reticulum-associated degradation, crucial for lipid-droplet maintenance, while CYP26A1 regulates retinoic acid and lipid homeostasis.CYP26A1 also belongs to the cytochrome P450 superfamily involved in lipid homeostasis and drug metabolism.
- Lipid transport and metabolism: VLDLR has important roles in VLDL-triglyceride metabolism and correlates with glucose and triglyceride plasma levels; VIM contributes to metabolism of lipoprotein-derived cholesterol.Both genes have previously reported links to circulating lipid-related processes.
- Other biological links: OGFOD1 supports cellular adaptation to oxygen changes and has been reported to function in ischaemic signalling.Its reported function provides a cellular-stress connection distinct from the candidates’ lipid-specific roles.
- Lipid homeostasis and cellular regulation: PPARA encodes a major hepatic lipid-metabolism regulator and the cellular receptor for fibrates, which lower serum triglycerides and raise serum HDL-cholesterol levels.This links the candidate to both lipid regulation and an anti-dyslipidemia drug mechanism.