Source-linked AI summary

DNA Methylation Profiling in Melanoma: From Lesion Classification to Therapeutic Stratification

Jana T. Winterstein, Lukas Heinlein, Günter Raddatz, Carina Nogueira Garcia, Sarah Haggenmüller, Christoph Wies, Lucas Schneider, Annemarie Hoffsommer, Tim J. Zeuner, Friedegund Meier, Sarah Hobelsberger, Frank F. Gellrich, Mildred Sergon, Axel Hauschild, Lucie Heinzerling, Justin G. Schlager, Kamran Ghoreschi, Max Schlaak, Franz J. Hilke, Carola Berking, Markus V. Heppt, Michael Erdmann, Sebastian Haferkamp, Konstantin Drexler, Dirk Schadendorf, Wiebke Sondermann, Matthias Goebeler, Bastian Schilling, Daniel B. Lipka, Stefan Fröhling, Felix Sahm, Jakob N. Kather, Yuri Tolkach, Jochen S. Utikal, Benjamin Izar, Yevgeniy R. Semenov, Titus J. Brinker

arXiv:2608.21448v1q-bio.GNcs.LG

TL;DR

Melanoma lesion methylomes may contain information useful for diagnosis and disease progression, but their value across clinical categories requires evaluation. This study compared CpG-based and biology-guided machine-learning models, achieving AUROC 0.919 for lesion classification and mean absolute error 0.627 for treatment-group prediction.

  • Problem

    The study asks whether DNA methylation profiles can support biologically and clinically relevant stratification of melanocytic lesions.

  • Method

    The study compared CpG-based with biology-guided machine-learning representations of prospectively collected methylation profiles from melanocytic lesions.

  • Results

    AUROC 0.919 (95% CI: 0.878 to 0.952) was achieved for classifying nevi, noninvasive melanoma, and invasive melanoma, while treatment-group prediction had mean absolute error below one group.

  • Takeaways & Limitations

    DNA methylation profiling can resolve clinically relevant lesion categories and support therapeutic stratification within melanoma.

  • Takeaways & Limitations

    The clinical value of methylation-based therapeutic-group prediction for therapy selection remains to be assessed.

Abstract

from arXiv · show

DNA methylation provides a stable record of cellular identity, capturing epigenetic programs that distinguish specialized cell states despite a shared genome. Because malignant transformation and tumour progression are accompanied by extensive epigenetic remodeling, we hypothesized that the methylome of melanocytic lesions contains biologically and clinically relevant information for both diagnosis and disease progression. In a cohort of 1,001 tissue samples prospectively collected across eight German university hospitals profiled using Illumina Infinium MethylationEPIC arrays, we compared machine-learning models based on selected Cytosine phosphate Guanine (CpG) methylation sites with models incorporating biology-guided features, including epigenetic age acceleration, cell type composition and copy-number variation burden. In an external test set, the best diagnostic classifier was CpG-based and distinguished melanocytic nevi, noninvasive melanoma and invasive melanoma with a macro-averaged area under the receiver operating characteristic curve of 0.919 (95% CI: 0.878 to 0.952). Notably, across CpGs most strongly hyper- and hypomethylated between NV and IM, NIM showed an intermediate methylation profile, providing a molecular correlate of its diagnostic complexity. The best model for clinically relevant treatment group prediction, with AJCC stages grouped according to guideline-based management recommendations, relied on biology-guided features and achieved a macro-averaged mean absolute error of 0.627 (95% CI: 0.477 to 0.808). Together, these findings demonstrate that methylation-based models can capture both diagnostic identity and clinically relevant disease stratification, supporting DNA methylation as a promising biomarker for further validation and potential clinical translation.

Results

Methylation profiling revealed substantial epigenetic differences across melanocytic lesions, with noninvasive melanoma occupying an intermediate molecular position between nevi and invasive melanoma. The study evaluated CpG-based and biology-guided representations for diagnostic classification and therapeutic stratification in a large multicentre cohort.

  • Study framework: Models addressed three-class classification of invasive melanoma, noninvasive melanoma and nevi, alongside therapeutic-group prediction based on guideline-aligned AJCC stage groupings.Three model types were developed for each task: CpG-based, methylation-derived marker-based and stacked models combining both predictions.
  • Study framework: The analysis used 1,001 tissue samples from 923 patients prospectively and consecutively collected across eight German university hospitals.Methylation profiles generated genome-wide data for melanocytic lesions.
  • Study framework: External-test-set evaluation assessed whether CpG-based and biology-guided methylation representations captured complementary information for diagnosis and therapeutic stratification.The biology-guided features included epigenetic age acceleration, cell type composition and copy-number variation burden.
  • Methylation differences: 2,267 high-effect differentially methylated regions separated invasive melanoma from nevi, compared with 1,146 between invasive and noninvasive melanoma and 193 between noninvasive melanoma and nevi.The greatest number of high-effect DMRs occurred between invasive melanoma and benign nevi.
  • Methylation differences: Noninvasive melanoma showed intermediate hyper- and hypomethylation patterns between invasive melanoma and nevi, consistent with a molecular continuum across diagnostic categories.Its overlap with both nevi and invasive melanoma may underlie the diagnostic challenges associated with noninvasive melanoma.

Diagnostic Classification of Melanocytic Lesions

CpG methylation models accurately distinguished melanocytic nevi, noninvasive melanoma, and invasive melanoma, outperforming marker-based models particularly for noninvasive melanoma. Noninvasive melanoma showed diffuse methylation profiles overlapping nevus- and invasive-melanoma-associated signatures, while stacking provided no significant overall gain.

  • Diagnostic Classification Results: 0.886 AUROC for NIM was lower than for NV (0.946) and IM (0.926), indicating greater diagnostic ambiguity for NIM.These are class-wise AUROC values from the CpG-based model.
  • Diagnostic Classification Results: 0.076 ΔAUROC favored the CpG-based model over the marker-based model on the primary endpoint, with 95% CI 0.033 to 0.119 and p < 0.001.The advantage was primarily attributable to NIM, with DeLong ΔAUROC = 0.143 and Holm-adjusted p = 0.007.
  • Complementarity of the Diagnostic Classification Models: 129 of 187 external test cases (69.0%) were correctly classified by both CpG-based and marker-based models, whereas stacking achieved 0.899 AUROC and did not significantly improve on the CpG-based model.The stacked-versus-CpG comparison was ΔAUROC = 0.020, 95% CI: -0.005 to 0.054, Holm-adjusted p = 0.124.
  • Molecular characterization: 205 selected diagnostic CpGs mapped to high-effect DMRs whose heatmap profiles showed NIM overlapping both NV- and IM-associated signatures rather than forming a distinct pattern.Predictive CpGs were enriched for transcriptional regulation and DNA-binding terms, including DNA-binding transcription factor activity and chromatin (both FDR = 0.025).

Therapeutic Group Prediction for Melanoma

Methylation-based models predicted melanoma therapeutic groups with mean absolute error below one group, with marker-based biology-guided features numerically outperforming CpG features but without significant improvement. Copy-number burden was the dominant marker-based predictor, while combining model outputs did not improve performance.

  • Model comparison: 0.043 MAE difference favored the marker-based model numerically, but was not statistically significant (95% CI: -0.141 to 0.224, p = 0.65).Across 58 external test cases, both models were correct in 16 cases (27.6%) and equally wrong in 24 cases (41.4%); their errors were largely shared.
  • Model comparison: 0.758 macro-averaged MAE (95% CI: 0.603 to 0.947) was achieved by the best stacked model, which did not significantly improve on either base model.Versus the CpG-based model, Δmacro-MAE = -0.087 (95% CI: -0.205 to 0.037, Holm-adjusted p = 0.28); versus the marker-based model, Δmacro-MAE = -0.130 (95% CI: -0.303 to 0.046, Holm-adjusted p = 0.28).
  • Biology-guided features: 1.16 ± 0.17 therapeutic groups was the marker-based model’s permutation-importance increase after jointly permuting FGA and total CNV burden, exceeding EpiScore estimates and EAA.Total CNV burden correlated most strongly with therapeutic groups (ρ = +0.68), followed by FGA (ρ = +0.60), both Holm-adjusted p < 0.001.
  • Stacked model: 0.76 ± 0.15 therapeutic groups was the macro-MAE increase after permuting the stacked model’s CpG predictions, versus 0.13 ± 0.07 for marker predictions.The meta-regressor therefore relied more heavily on CpG-based predictions, although cross-validated CpG-based performance was 0.79 ± 0.19 versus 0.94 ± 0.25 for the marker-based model.

Ethics Statement and Reporting Standards

The multicentre study received ethics approval from the reported German university and hospital committees, with a stated exception for Charité Berlin. Patients provided informed written consent, and the study followed the Declaration of Helsinki.

  • Ethics approval: Ethics approval was obtained from committees at the Technical University of Dresden, Friedrich-Alexander University Erlangen-Nuremberg, LMU Munich, University of Regensburg, University of Würzburg, and University Hospitals Mannheim and Essen.The reported approval identifiers were BO-EK-53012021, 69_21 Bc, 21-0182, 20-2190-103, 293/20_z, 2010-318N-MA, 2014-835R-MA, and 20-9784-BO.
  • Ethics approval: Separate ethics approval was not required from Charité Berlin because approval from another German university or medical association ethics board covered the multicentre study under §15(2) of Berlin’s professional code.The passage states that an additional approval is unnecessary in this circumstance.
  • Consent and standards: Patients provided informed written consent, and the study was performed in accordance with the Declaration of Helsinki.These statements describe the study’s consent and international ethical framework.

Data collection

The study prospectively collected 1,001 melanocytic lesions from 923 patients across eight German university hospitals and profiled tumour DNA methylation using Illumina Infinium MethylationEPIC arrays. Samples were processed with pathologist-guided tumour dissection, quality-controlled preprocessing, patient-stratified cross-validation, and a hospital-held-out external test set.

  • Cohort: 1,001 lesions from 923 patients were prospectively and consecutively collected across eight German university hospitals between April 2021 and February 2023.The cohort comprised IMs (n = 349), NIMs (n = 117), and NV (n = 535).
  • Therapeutic stratification: Melanoma samples with available AJCC 8th edition stages were assigned to six ordinal therapeutic groups according to stage-specific management recommendations, excluding stage IV for insufficient sample size.Groups ranged from stage 0 to stages IIIA–IIID; no stage IIID patients were represented.
  • Tissue processing: Tumour regions were annotated by a lead senior dermatopathologist and manually dissected from unstained serial sections before DNA extraction from FFPE tissue.Each FFPE block yielded six serial 3–4 µm sections, and DNA was extracted using the QIAamp DNA FFPE Advanced Kit.
  • Data splitting: Hospital 8 was held out for external testing, while hospitals 1 to 7 supplied training and validation data with patient-stratified 5-fold cross-validation.Slides from the same patient never split across cross-validation folds, and stratification used the task target directly.

Statistical Analysis

Diagnostic models were evaluated primarily by macro-averaged one-versus-rest AUROC, while therapeutic prediction used macro-averaged rank-based MAE. Secondary metrics, patient-level bootstrap confidence intervals, pairwise model comparisons, and multiple-comparison correction were prespecified.

  • Diagnostic classification used one-versus-rest macro-averaged AUROC as the primary endpoint, with balanced accuracy and macro-averaged F1 score as secondary endpoints.
  • Therapeutic group prediction used macro-averaged, rank based MAE as the primary endpoint, with quadratic weighted Cohen's κ and Spearman's correlation ρ as secondary endpoints.
  • Test-set metrics were reported as point estimates with 95% bootstrap CIs from 1,000 patient-level resamples.
  • Models were compared pairwise on the external cohort using paired, patient-clustered bootstrap differences, while class-specific AUROCs were assessed with DeLong's test.
  • Marker associations were analyzed descriptively in training data, using Spearman's ρ for therapeutic monotonicity and Kruskal-Wallis followed by post-hoc Mann-Whitney U tests for group differences.
  • Related p-values were Holm-Bonferroni corrected, and all tests were two-sided with α = 0.05.

AJCC 8th Edition Stage Therapeutic Group

The therapeutic-group analysis characterized development and testing datasets and evaluated best-performing prediction models based on CpGs, methylation-derived markers, and stacked features. Model performance was further examined with confusion matrices and paired, patient-clustered bootstrap comparisons of macro-averaged mean absolute error.

  • Dataset characteristics: Supplementary Table 3 reports characteristics of the datasets used for therapeutic-group model development and testing.
  • Model evaluation: Confusion matrices were provided for the best-performing therapeutic-group prediction models.
  • Model evaluation: Best-performing therapeutic-group prediction models were evaluated using CpGs, methylation-derived markers, and stacked feature sets.
  • Model evaluation: Pairwise differences in macro-averaged mean absolute error were assessed using a paired, patient-clustered bootstrap with 1,000 resamples.
Loading 2608.21448v1…