Source-linked AI summary

Predicting genome-wide DNA methylation using methylation marks, genomic position, and DNA regulatory elements

Weiwei Zhang, Tim D Spector, Panos Deloukas, Jordana T Bell, Barbara E Engelhardt

arXiv:1308.2134v1q-bio.GN

TL;DR

Site-specific genome-wide DNA methylation remains challenging to characterize and predict. The paper develops a random forest predictor using CpG methylation patterns, genomic context, and regulatory-element information, achieving improved prediction while identifying features most predictive of methylation levels.

  • Problem

    DNA methylation remains challenging to characterize and predict, motivating improved site-specific genome-wide prediction.

  • Method

    The study builds a random forest predictor for CpG-site methylation using neighboring methylation levels, genomic position, and co-localized genomic and regulatory features.

  • Results

    The method outperformed state-of-the-art methylation classifiers and identified neighboring CpG methylation, CpG-island location, DHS sites, and specific transcription-factor binding sites as most predictive.

  • Takeaways & Limitations

    The approach links genome-wide methylation prediction with genomic features that interact with methylation and may help elucidate epigenetic regulation.

  • Takeaways & Limitations

    Prediction accuracy remains below accuracies currently achieved for genome-wide CpG sites from WGBS assays.

Abstract

from arXiv · show

Background: Recent assays for individual-specific genome-wide DNA methylation profiles have enabled epigenome-wide association studies to identify specific CpG sites associated with a phenotype. Computational prediction of CpG site-specific methylation levels is important, but current approaches tackle average methylation within a genomic locus and are often limited to specific genomic regions. Results: We characterize genome-wide DNA methylation patterns, and show that correlation among CpG sites decays rapidly, making predictions solely based on neighboring sites challenging. We built a random forest classifier to predict CpG site methylation levels using as features neighboring CpG site methylation levels and genomic distance, and co-localization with coding regions, CGIs, and regulatory elements from the ENCODE project, among others. Our approach achieves 91% -- 94% prediction accuracy of genome-wide methylation levels at single CpG site precision. The accuracy increases to 98% when restricted to CpG sites within CGIs. Our classifier outperforms state-of-the-art methylation classifiers and identifies features that contribute to prediction accuracy: neighboring CpG site methylation status, CpG island status, co-localized DNase I hypersensitive sites, and specific transcription factor binding sites were found to be most predictive of methylation levels. Conclusions: Our observations of DNA methylation patterns led us to develop a classifier to predict site-specific methylation levels that achieves the best DNA methylation predictive accuracy to date. Furthermore, our method identified genomic features that interact with DNA methylation, elucidating mechanisms involved in DNA methylation modification and regulation, and linking different epigenetic processes.

Keywords

The paper focuses on DNA methylation and its genomic and regulatory context, including CpG regions, regulatory elements, and EWAS.

  • DNA methylation is examined in relation to CpG islands, shores, and shelves.
  • Random forest classifiers are used to predict methylation patterns at CpG sites.
  • DNase I hypersensitive sites and transcription factor binding sites are considered regulatory features.

Background

The study addresses the challenge of predicting methylation at individual CpG sites across the genome, where existing methods often use regional averages or restricted genomic regions. It combines methylation measurements with genomic and regulatory features to develop a genome-wide random forest predictor and identify features associated with prediction accuracy.

  • Background: DNA methylation is an epigenetic modification involved in DNA replication, gene transcription, cell differentiation, development, and tumorigenesis.
  • Background: CpG islands are mostly unmethylated, GC-rich genomic regions that often occur in promoters and exons.
  • Measurement challenges: Whole-genome bisulfite sequencing measures methylation at approximately 26 million CpG sites but is expensive, subject to conversion bias, and difficult in some regions.
  • Proposed approach: The proposed random forest predicts methylation levels at single CpG sites genome-wide using neighboring methylation, genomic location, local features, and ENCODE regulatory marks.
  • Prior prediction methods: Existing predictors commonly estimate binary or average methylation across genomic regions, often restricting predictions to CpG islands.

Results

Genome-wide methylation patterns were highly context-specific, with rapid local correlation decay and distinct CGI-associated profiles. A random forest using neighboring methylation, genomic context, and regulatory features achieved strong site-level prediction, with performance improving under short-distance restrictions and within CGIs.

  • 81.2% of CpG sites in CGIs were hypomethylated, whereas 73.2% of non-CGI sites were hypermethylated.
  • CGI shores showed variable U-shaped methylation, while 78.2% of CGI shelf sites were hypermethylated.
  • Neighboring-site correlation declined to approximately 0.4 within ∼400 bp and depended strongly on genomic context.
  • CGI-neighbor MSE increased slowly with distance, whereas CGI shore and shelf MSE exceeded the background value of 0.30.
  • Upstream and downstream neighboring methylation statuses and co-localized DHS sites correlated most strongly with methylation PC1, at approximately 0.57.
  • Using all 97 features, the classifier achieved 91.6% accuracy and 0.96 AUC; restricting neighbors to 1 kb increased these values to 94.3% and 0.98.
  • Methylation-level prediction reached r = 0.90 with unrestricted neighbors and r = 0.94 within 1 kb, while RMSE decreased from 0.19 to 0.15.
  • Within CGIs, methylation-status prediction achieved 98.3% accuracy and 0.99 AUC, while methylation-level prediction reached correlation 0.94 and RMSE 0.09.

Discussion

The study characterizes genome-wide methylation patterns and develops a CpG-site-resolution random forest predictor using methylation, genomic, and regulatory features. Its discussion highlights predictive features, potential imputation applications, biological interpretation, and important limits involving reference panels, rare variants, and unresolved causal mechanisms.

  • The study characterized genome-wide and region-specific DNA methylation patterns, raising questions about relationships between methylation and other genomic and epigenomic processes.
  • A random forest classifier predicts DNA methylation levels at single-CpG resolution using neighboring methylation, genomic location, local features, and regulatory-element colocalization.
  • Random forests improved on SVMs for sparsely sampled genomic regions and offered biological interpretability through readily extractable feature importance.
  • The approach could support imputing unassayed genome-wide CpG methylation from array data, but its accuracy remains below that of current DNA-imputation methods.The discussion specifically considers imputing missing genome-wide CpG sites from Illumina 450K measurements using WGBS references.
  • Precise methylation imputation may require larger reference panels because biological, batch, and environmental effects influence methylation, and rare or unexpected variants remain difficult to predict.
  • Neighboring CpG methylation status and CpG-site features were among the most important predictors, while additional genomic features improved prediction.

Conclusion

Using methylation profiles from 100 individuals, the study characterized region-specific CpG correlation patterns and developed a random forest classifier for single-CpG methylation prediction. The model outperformed state-of-the-art classifiers and highlighted neighboring methylation, CpG-island location, DHS sites, and transcription-factor binding sites as predictive features.

  • 100 individuals were analyzed to characterize methylation patterns and CpG correlation structure across genomic regions.
  • CpG methylation correlation showed distinct patterns across CpG islands, CGI shores, and non-CGI regions, including distance-dependent shelf-region correlation.
  • The random forest predicted binary methylation status and continuous methylation levels at single-CpG precision using neighboring methylation, genomic position, sequence properties, and regulatory-element co-location.
  • The method outperformed state-of-the-art methylation classifiers, including an SVM-based classifier.
  • Neighboring methylation levels, CpG-island location, co-localized DHS sites, and specific transcription-factor binding sites were among the most predictive features.Identified TFBSs included Elf1, MAZ, Mxi1, and Runx3, whose predictive features may relate to methylation regulation or downstream cellular phenotypes.

Materials and methods

The study measured and quality-controlled genome-wide CpG methylation in 100 individuals, then evaluated correlations and trained random forest and SVM classifiers using genomic, neighboring-site, and regulatory features.

  • Data processing: Methylation level β was defined as the methylated-probe signal divided by the combined methylated and unmethylated probe signals, ranging from 0 to 1.The calculation used nonnegative signal values and included a quantity α in the denominator.
  • Data processing: Quality control removed multiply mapped, missing, or low-detection-quality probes and sex-chromosome sites, leaving 394,354 CpG sites for analysis.Residualization controlled for array number, array position, age, and sex; PCA identified no obvious outliers.
  • Correlation analysis: Methylation correlation and mean square error were assessed for adjacent CpG pairs across genomic-distance windows and for randomly sampled non-adjacent pairs.The analysis used Pearson correlation and compared neighboring sites within distances extending to 6,000 bp.
  • Prediction models: A random forest classifier used neighboring methylation status, genomic distance, and genomic or regulatory annotations among a comprehensive set of 99 features.Features included CpG-island, coding-region, GC-content, recombination, conservation, DNase hypersensitivity, and transcription-factor binding annotations; regional analyses excluded the corresponding region feature.

Author’s contributions

The listed contributors conceived, designed, performed, analyzed, and wrote the study, with additional responsibility for quality control, preprocessing, and data contribution.

  • Contributions: WZ, JTB, and BEE conceived, designed, performed, analyzed, and wrote the study across the listed contribution categories.The contribution record assigns conception to JTB and BEE, design to WZ and BEE, performance to WZ, analysis to WZ and BEE, and writing to WZ, JTB, and BEE.
  • Contributions: WZ, JTB, and BEE handled quality control and preprocessing, while TDS, PD, and JTB contributed valuable data.
Loading 1308.2134v1…