Source-linked AI summary
Integrating diverse datasets improves developmental enhancer prediction
Genevieve D. Erwin, Rebecca M. Truty, Dennis Kostka, Katherine S. Pollard, John A. Capra
TL;DR
Identifying developmental enhancers is difficult because no single data type captures all active enhancers across cellular contexts. EnhancerFinder integrates diverse genomic evidence and improves validated enhancer prediction while producing tissue-specific genome-wide annotations.
Problem
Comprehensive developmental enhancer prediction remains difficult because no single data type captures all active enhancers or transfers predictably across cellular contexts.
Method
EnhancerFinder uses two-step multiple-kernel learning to integrate sequence, evolutionary, and functional genomics features, then classify enhancer tissue specificity.
Results
Integrating diverse data from many cellular contexts improved validated enhancer prediction and produced more than 80,000 tissue-specific developmental enhancer predictions genome-wide.
Takeaways & Limitations
The predictions annotate human non-coding DNA and are enriched near relevant genes and genome-wide association study hits, with a validated cranial nerve enhancer at ZEB2.
Takeaways & Limitations
The VISTA training enhancers are experimentally validated in vivo but represent a narrow definition and are not an exhaustive catalog of enhancers.
Abstract
from arXiv · showhide
Gene-regulatory enhancers have been identified by many lines of evidence, including evolutionary conservation, regulatory protein binding, chromatin modifications, and DNA sequence motifs. To integrate these different approaches, we developed EnhancerFinder, a novel method for predicting developmental enhancers and their tissue specificity. EnhancerFinder uses a two-step multiple-kernel learning approach to integrate DNA sequence motifs, evolutionary patterns, and thousands of diverse functional genomics datasets from a variety of cell types and developmental stages. We trained EnhancerFinder on hundreds of experimentally verified human developmental enhancers from the VISTA Enhancer Browser, in contrast to histone mark or sequence-based enhancer definitions commonly used. We comprehensively evaluated EnhancerFinder, and found that our integrative approach improves enhancer prediction accuracy over previous approaches that consider a single type of data. Our evaluation highlights the importance of considering information from many tissues when predicting specific types of enhancers. We find that VISTA enhancers active in embryonic heart are easier to predict than enhancers active in several other tissues due to their uniquely high GC content. We applied EnhancerFinder to the entire human genome and predicted 84,301 developmental enhancers and their tissue specificity. These predictions provide specific functional annotations for large amounts of human non-coding DNA, and are significantly enriched near genes with annotated roles in their predicted tissues and hits from genome-wide association studies. We demonstrate the utility of our enhancer predictions by identifying and validating a novel cranial nerve enhancer in the ZEB2 locus. Our genome-wide developmental enhancer predictions will be freely available as a UCSC Genome Browser track.
AUTHOR SUMMARY
Enhancers are short regulatory DNA elements that control when and where genes are activated during development. Mapping their genomic locations can illuminate developmental processes, genetic diseases, and responses to drugs.
- Motivation: Enhancers are short DNA regulatory elements that switch genes on or off at specific developmental times in specific cells or tissues.They help regulate when and where genes are turned on during development.
- Motivation: Locating enhancers in the human genome can provide insight into development, genetic diseases, and responses to drugs.This is especially important because much human non-coding DNA has unknown function.
INTRODUCTION
EnhancerFinder integrates sequence, evolutionary, and functional-genomics data from diverse cellular contexts in a two-step machine-learning framework to predict developmental enhancers and tissue specificity. The authors report improved validated-enhancer prediction and more than 80,000 genome-wide developmental-enhancer predictions with tissue-specific annotations.
- Motivation: Comprehensive enhancer prediction is difficult because no single data type identifies all context-specific enhancers and functional-genomics data span many tissues and cellular contexts.Single marks or motifs provide incomplete representations of biologically active enhancers.
- EnhancerFinder: EnhancerFinder uses two classifiers: one separates VISTA developmental enhancers from genomic background, and the other distinguishes enhancers by tissue specificity.The first classifier identifies known human developmental enhancers, while the second classifies their tissue specificity.
- EnhancerFinder: EnhancerFinder integrates diverse data from many cellular contexts and significantly improves prediction of validated enhancers over previous approaches.The method combines evolutionary conservation, sequence motifs, and functional-genomics data within a two-step multiple-kernel-learning framework.
- Genome-wide application: The method predicts more than 80,000 developmental enhancers genome-wide, including tissue-specific predictions for brain, limb, and heart.These predictions provide tissue-specific annotations across the human genome.
RESULTS
EnhancerFinder integrates functional genomics, evolutionary conservation, and DNA sequence motifs to predict developmental enhancers accurately. Its broad functional-genomics data integration outperformed narrower embryonic or relevant-data subsets, showing that diverse contexts improve prediction.
- EnhancerFinder integrates diverse data types to accurately identify developmental enhancers: AUC = 0.96 when EnhancerFinder combined functional genomics, evolutionary conservation, and DNA sequence motifs to distinguish enhancers from genomic background.The integrated classifier used three distinct data types.
- EnhancerFinder performs significantly better than previous enhancer prediction approaches: EnhancerFinder performed significantly better than previous enhancer-prediction approaches, with complementary component classifiers sharing only roughly two-thirds of their predictions.The comparison included classifiers based on all functional genomics data, DNA motifs, and evolutionary conservation, as well as prior approaches.
- Integrating diverse functional genomics data improves enhancer prediction: All Functional Genomics achieved AUC=0.89, slightly exceeding Relevant Functional Genomics at AUC=0.87 and outperforming Embryonic Functional Genomics at AUC=0.83.The comparison used classifiers trained on different subsets of functional-genomics features.
- Integrating diverse functional genomics data improves enhancer prediction: Adding functional-genomics datasets from less obviously relevant contexts improved prediction beyond datasets selected specifically for embryonic or relevant tissues.At low false-positive rates, the differences in power were modest.
Histone marks and p300 provide complementary information about enhancer activity
Individual functional-genomics classifiers performed better than random, while EnhancerFinder assigned biologically meaningful feature weights and generated tissue-specific and genome-wide enhancer predictions with quantified performance and limitations.
- Functional genomics classifiers: All three single-dataset SVM classifiers—H3K27ac, H3K4me1, and p300—performed better than random when distinguishing VISTA positives from genomic background.The classifiers used ENCODE data spanning multiple cell types and tissues.
- Functional genomics classifiers: EnhancerFinder generally assigned positive weights to features associated with enhancer activity in relevant cellular contexts and negative weights to features associated with other element types.These learned weights reflect each feature’s contribution to enhancer predictions.
- Tissue-specific enhancer prediction: Heart tissue-specific predictions performed best (AUC=0.85), followed by limb (0.74), forebrain (0.72), midbrain (0.72), hindbrain (0.69), and neural tube (0.62).All tested tissue classifiers performed better than random, but performance varied substantially between tissues.
- Genome-wide predictions: At a 5% false-positive-rate threshold, the predictions corresponded to a 65% true-positive rate, with estimated false-discovery rates ranging from 9% to 47% depending on enhancer prevalence.The 9% estimate assumes that 50% of genomic windows harbor developmental enhancers, whereas the 47% estimate assumes 10%.
- Genome-wide predictions: Genome-wide analysis used a smaller functional-genomics dataset and excluded evolutionary conservation because 706 of 711 VISTA enhancers overlapped conserved elements.These choices reduced computational time and addressed strong conservation bias in the training positives.
- Genome-wide predictions: Step 2 predicted 7,400 limb enhancers, 19,051 heart enhancers, and 11,693 brain enhancers at tissue-specific 5% false-positive-rate thresholds.Predictions were made independently for each tissue, so regions could be assigned to multiple tissues.
Predicted enhancers are associated with relevant functional genomic regions
Genome-wide EnhancerFinder predictions are enriched near genes with relevant tissue expression and biological functions. A ZEB2 case study further identified a dense enhancer landscape and experimentally validated a predicted cranial nerve enhancer.
- Predicted brain and heart enhancers are near genes enriched for expression in the corresponding tissues, with nearby genes also enriched for relevant GO biological processes.These analyses used independent indicators of enhancer function.
- Case Study: EnhancerFinder predictions highlight a novel enhancer near ZEB2: 54 predicted enhancers have ZEB2 as their nearest TSS, placing ZEB2 in the top 0.2% of genes by adjacent enhancer predictions.Known VISTA enhancers overlap EnhancerFinder predictions near ZEB2.
- Case Study: EnhancerFinder predictions highlight a novel enhancer near ZEB2: All seven embryos with staining showed cranial nerve expression from the tested enhancer region, regardless of whether the construct contained the human or chimpanzee sequence.The enhancer was selected for its high EnhancerFinder score and overlap with a human accelerated region.
DISCUSSION
EnhancerFinder integrates evolutionary, sequence, and functional-genomics data to predict developmentally active enhancers, while its analyses reveal why diverse contexts help and why tissue-specific performance varies. The framework supports genome annotation and future integration of additional genomic data, but remains constrained by training and data-availability biases.
- Contribution: EnhancerFinder flexibly integrates evolutionary, DNA-sequence, and functional-genomics data rather than relying on one or a few data types.The framework was developed for predicting regulatory enhancers from diverse data sources.
- Limitations: The biologically active in vivo enhancer definition is preferable for exploration and experimental characterization, but VISTA-based training is limited by developmental-stage, experimental, and conservation biases.The approach relies heavily on VISTA because it is the largest available collection of validated mammalian enhancers, while its regions do not span the full range of enhancers.
- Data integration: Complementary data from less directly relevant contexts improve predictions by revealing negative and reinforcing patterns that distinguish developmental enhancers.p300 binding and H3K4me1 contribute substantially, but other contexts also provide useful information.
- Tissue specificity: Heart enhancers are easier to predict because they have higher GC content, lower evolutionary conservation, and closer proximity to the nearest TSS than other tissue enhancers.GC content alone was sufficient to explain part of the distinctive predictability of heart enhancers, and GC content correlates with enhancer activity.
- Extensions and applications: EnhancerFinder enables human regulatory-element annotation and non-coding mutation interpretation, with ZEB2 demonstrating rapid novel-enhancer identification and GWAS enrichment supporting its utility.The framework can also incorporate population variation, three-dimensional genome information, and predicted target genes.
METHODS
EnhancerFinder applies a two-step multiple-kernel learning framework to genomic features, first distinguishing enhancers from non-enhancers and then predicting tissue-specific activity. It integrates functional genomics data, evolutionary conservation, and DNA sequence motifs using VISTA-derived training examples.
- Training examples: The first prediction step trained on 711 VISTA enhancers and 711 length- and chromosome-matched random genomic negatives.Negatives were filtered to remove known VISTA enhancers and assembly gaps.
- Training examples: The second step used tissue-specific subsets of 1,447 VISTA regions to learn enhancer activity for individual tissues.For heart, training used 84 heart-expressing positives and 1,363 regions without heart expression at E11.5, regardless of activity in other tissues.
- Feature data: Functional-genomics features included histone modifications, transcription-factor and p300 associations, and open-chromatin measurements from hundreds of cell types.Heart p300 data were also included.
- Feature data: The method represented sequence motifs by counting all possible 4-mers and represented conservation by each region’s maximum overlapping mammalian phastCons score.Regions without overlapping phastCons elements received a conservation score of zero.
- Machine-learning algorithms: EnhancerFinder integrates three feature types—functional genomics, evolutionary conservation, and DNA sequence motifs—through multiple-kernel learning.Its kernels quantify similarity between genomic regions, while the MKL algorithm learns both training-example weights and kernel contributions.
Performance evaluations
EnhancerFinder was evaluated using cross-validation and multiple performance measures, then compared with established enhancer-prediction methods and genomic segmentation approaches.
- Evaluation framework: Performance was evaluated with 10-fold cross-validation using ROC AUC, precision-recall curves, and power estimates at fixed false-positive rates.Differences between classification methods were assessed with McNemar’s test.
- Comparison to existing methods: The evaluation compared EnhancerFinder with CLARE, ChromHMM, and Segway predictions, including ChromHMM models from individual and combined ENCODE cell lines.CLARE was evaluated on the Step 1 prediction task using the study’s positive and negative datasets.
Identification of tissue-specific enhancers across the human genome
EnhancerFinder was applied genome-wide to identify high-confidence tissue-specific developmental enhancers, yielding separate heart, brain, and limb enhancer sets. These predictions were further characterized through genomic, functional, motif, GWAS-overlap, and transgenic assays.
- Genome-wide enhancer prediction: Enhancer scores were filtered at a 5% false-positive rate estimated by cross-validation, using 1500-bp windows scanned every 500 bp across the human genome.The genome-wide scan used a trained Step 1 MKL classifier without conservation.
- Genome-wide enhancer prediction: 19,051 heart, 11,693 brain, and 7,400 limb enhancers were predicted after tissue-specific 5% FPR filtering and merging overlapping windows.Tissue specificity was assigned using trained brain, limb, and heart classifiers applied to 299,039 windows with positive Step 1 enhancer scores.
- Prediction characterization: Predicted enhancers were evaluated for nearby gene-expression patterns across 79 tissues using GNF Atlas 2 and paired t-tests.The analysis compared expression of genes nearest to predicted heart, brain, and limb enhancers.
- Prediction characterization: Genomic regions near predicted enhancers were tested for Gene Ontology annotations, phenotypes, and pathways with GREAT using hypergeometric tests and basal-plus-extension associations.The association rule included 5 kb upstream, 1 kb downstream, and distal sequence up to 100 kb.
- Prediction characterization: Motif occurrences in tissue-specific enhancer sets were identified with FIMO using TRANSFAC motifs and a score threshold of 10e-5.The motifs were summarized to identify prevalent transcription factors in each tissue-specific enhancer set.
- Prediction characterization: Predicted enhancer overlap with 9,687 GWAS SNPs was assessed using length- and chromosome-matched randomizations to calculate unadjusted permutation p-values.Randomized regions avoided assembly gaps before overlap testing.
FIGURE CAPTIONS
The figures describe EnhancerFinder’s two-step multiple-kernel pipeline, show that integrating diverse data improves developmental enhancer prediction, and demonstrate tissue-specific predictions culminating in validation of a cranial nerve enhancer near ZEB2.
- Prediction pipeline and performance: EnhancerFinder combines functional genomics, evolutionary conservation, and DNA sequence motifs, significantly improving enhancer identification over each data type alone (p<2.0E-7 for all).The two-step pipeline uses multiple-kernel learning with positive and negative training examples.
- Prediction pipeline and performance: Considering all or developmentally relevant functional-genomics features improves developmental-enhancer classification over embryonic functional-genomics features alone (p=9.2E-9 and p=2.7E-6, respectively).The comparison includes contexts and assays not initially expected to associate with developmental enhancer activity.
- Tissue specificity: Heart enhancers at E11.5 are dramatically easier to identify than enhancers active in other tissues, and heart enhancers have significantly higher GC content.Tissue-specific classifiers used active VISTA enhancers as positives and inactive regions as negatives.
- Tissue specificity: Of 84 validated heart enhancers, 71 are heart-specific, while five overlap brain, four overlap limb, and four overlap both tissues.The figure contrasts true tissue overlap with predictions from the two-step approach.
- Experimental validation: A predicted enhancer overlapping 2xHAR.240 in the ZEB2 locus drove consistent cranial-nerve expression in transgenic mice for both human and chimpanzee sequences.EnhancerFinder predicted a dense enhancer region around ZEB2, which was tested at E11.5.
SUPPLEMENTARY FIGURE CAPTIONS
Supplementary analyses show that EnhancerFinder benefits from integrating functional genomics data, while enhancer predictability varies with kernel choice, tissue activity, sequence conservation, and GC content. A cranial nerve enhancer near ZEB2 was supported by consistent expression in transient transgenic mouse embryos.
- Supplementary Figure 2: Spectrum kernels using k-mer lengths between 4 and 7 perform best, while combining k-spectrum kernels provides no significant improvement over the best individual kernels.The 4-spectrum kernel performs competitively with other k-spectrum kernels and their combination.
- Supplementary Figure 4: Linear SVMs using functional genomics data outperform classifying all regions overlapping each feature as positives.This comparison concerns distinguishing VISTA enhancers from genomic background in Step 1.
- Supplementary Figures 6–7: Heart enhancers are less conserved and closer to the nearest TSS than limb and brain enhancers, but remain easier to identify because of their high GC content.Restricting limb and brain enhancers to less-conserved regions near a TSS improves identification without matching heart-enhancer predictability.
- Supplementary Figure 7: Heart enhancers have 49% GC content versus approximately 40% for the genomic background, and their classification scores correlate with GC content at Pearson rho=0.95.The uniquely high GC content of VISTA heart enhancers enables accurate classification.
- Supplementary Figure 8: EnhancerFinder distinguishes enhancers active in multiple tissues better than single-tissue enhancers, with AUC=0.75 versus AUC=0.67.The VISTA database contains 312 multi-tissue enhancers and 399 single-tissue enhancers active at E11.5.
- Supplementary Figure 9: Seven transient transgenic mouse embryos showed LacZ expression for a novel cranial nerve enhancer near ZEB2, with staining in all human- and chimp-sequence embryos.Three embryos carried the human region and four carried the chimp sequence; no significant human–chimp difference appeared at this time point.
VISTA Positive Tissue Overlap
The section presents one-step and two-step prediction tissue-overlap views alongside associated biological-process annotations and genomic regulatory tracks.
- The section compares one-step and two-step prediction tissue overlap.
- Additional annotations cover signaling, fibroblast proliferation, and negative regulation of kinase, phosphate-metabolic, phosphorylation, and transferase activities.
- The annotations include neural development and differentiation terms, including primary neural tube formation and cell differentiation in spinal cord.
- The genomic context displays regulatory marks, DNaseI hypersensitivity, transcription-factor ChIP-seq, conservation, gene annotations, and human accelerated regions.