Source-linked AI summary
GENEB: Why Genomic Models Are Hard to Compare
Daria Ledneva, Mikhail Nuridinov, Denis Kuznetsov
TL;DR
Genomic foundation models lack directly comparable evaluation because benchmarks and protocols are fragmented. GENEB addresses this problem by probing frozen representations from 40 models across 100 tasks and 13 categories, revealing unstable rankings and task-dependent trade-offs. The benchmark supports category-aware comparison while leaving important long-range and noneukaryotic regimes underrepresented.
Problem
Fragmented benchmarks and incompatible evaluation protocols prevent direct comparison of genomic foundation-model superiority and generality claims.
Method
GENEB evaluates frozen sequence representations from 40 genomic foundation models on 100 tasks across 13 functional categories under unified probing, including few-shot regimes.
Results
Model rankings vary across categories and supervision regimes; scale correlates with aggregate performance at ρ = 0.573, while architecture and pretraining alignment frequently offset scale differences.
Takeaways & Limitations
Category-aware controlled evaluation is preferable to aggregate leaderboards for principled genomic model comparison and selection.
Takeaways & Limitations
GENEB underrepresents very long-range regulatory tasks and lacks broad prokaryotic and viral genomics coverage, limiting conclusions in those regimes.
Abstract
from arXiv · showhide
Progress in genomic foundation models is difficult to assess due to fragmented benchmarks, incompatible evaluation protocols, and task-specific reporting. As a result, claims of superiority or generality across models are often not directly comparable. We introduce GENEB, a large-scale diagnostic benchmark that evaluates frozen representations from 40 genomic foundation models across 100 tasks spanning 13 functional categories under a unified probing-based protocol, including few-shot regimes. GENEB enables controlled comparison across model scale, architecture, tokenization, and pretraining data while explicitly exposing task-level trade-offs. Our analysis shows that aggregate leaderboards are unstable: model rankings vary sharply across task categories, scale provides only modest and inconsistent gains, and architectural and pretraining alignment frequently outweigh parameter count. These results highlight limitations of current evaluation practices and position GENEB as a reference framework for principled comparison and category-aware model selection in genomic machine learning.
1. Introduction
Genomic foundation models are difficult to compare because they are evaluated on disjoint benchmarks with incompatible protocols, weakening claims about progress. GENEB addresses this gap with a unified benchmark for controlled cross-model evaluation.
- Disjoint benchmarks and incompatible protocols make it unclear whether reported improvements reflect genuine progress.
- DNA-GPT, GENOMEOCEAN, and EVO cannot be compared directly because they use different tasks, preprocessing pipelines, and evaluation protocols.
- Growing claims of model superiority and generality outpace reproducible evidence from cross-model evaluation.
- GENEB evaluates 40 genomic foundation models on 100 tasks spanning 13 functional categories under a unified probing protocol.
- GENEB provides directly comparable results across models and tasks, functioning as a shared reference framework rather than a single-task leaderboard.
2. Related Work
Prior genomic benchmarks and comparative studies cover important tasks but use differing protocols or limited model sets. GENEB extends this work with matched evaluation of 40 models across 100 tasks and 13 functional categories.
- The field spans diverse architectures, tokenization schemes, and pretraining strategies.
- Architectures: Architectural development includes Transformer encoders, decoder-only models, generative systems, long convolutions, state-space models, and hybrids.
- Tokenization and Pretraining: Tokenization ranges from nucleotides and k-mers to learned BPE vocabularies, while pretraining data varies from species-specific to multi-species corpora.
- Benchmarks: Existing benchmarks cover regulatory, epigenetic, and cross-species tasks but differ in design and protocols and usually evaluate limited model subsets.
- Positioning of GENEB: GENEB evaluates 40 models on 100 DNA classification tasks across 13 functional categories under one probing protocol, enabling matched comparisons across model properties.
3. Methodology
GENEB evaluates frozen genomic representations with lightweight probes, using standardized data regimes and metrics to isolate representation quality and support controlled comparisons.
- Frozen sequence representations are assessed with lightweight classifiers, isolating representation quality across architectures and training regimes.
- Probing setup: Logistic regression probes are evaluated in 1-shot, 10-shot, and full-data regimes using five fixed random seeds.
- Metric and data: Matthews Correlation Coefficient is the reported metric because it is robust to class imbalance and standard in genomic evaluation.
- Metric and data: Tasks with more than 10^5 sequences are subsampled, based on evidence that MCC stabilizes beyond this dataset size.
4. Aggregate Performance Analysis Across 100 Genomic Tasks
GENEB shows that aggregate genomic-model performance reflects scale, architecture, tokenization, and pretraining choices, with substantial task- and data-regime-specific variation. Controlled comparisons reveal that architecture and pretraining alignment can outweigh parameter count, while few-shot performance degrades sharply and unevenly across models and categories.
- Benchmark scope: GENEB evaluates 40 DNA foundation models across 100 tasks and 13 functional categories to analyze scale, architecture, tokenization, and pretraining under one protocol.Statistics use MCC aggregated within GENEB.
- Scale and architecture: 25 in-domain cases show a model at least 5× smaller outperforming a larger counterpart, including MUTBERT exceeding ECCDNAMAMBA by +0.110 macro-MCC despite a 6.2-fold size difference.The 25-case count is identical under micro- and macro-averaging.
- Scale and architecture: Transformer models exceed Mamba by +0.131 macro-MCC and +0.149 macro-MCC in matched multi-species/BPE comparisons, while encoder–decoder rankings remain task- and setting-dependent.The reported aggregate gaps are 0.550 vs. 0.419 and 0.568 vs. 0.419, respectively.
- Scale and architecture: Architecture gaps become larger on cross-species tasks: Transformer–Mamba differences reach +0.355 macro-MCC on virus/phage and +0.305 on mouse enhancers.These gaps exceed the +0.075 macro-MCC difference between models above 1B and below 200M parameters.
- Pretraining and tokenization: Multi-species pretraining improves chromatin accessibility by +0.076 macro-MCC across 6/6 controlled pairs, while microbial-focused pretraining trails general multi-species pretraining by +0.084 macro-MCC on average.The largest category gaps favoring general multi-species pretraining occur on splice sites (+0.222), species classification (+0.130), lncRNA (+0.116), and DNA methylation (+0.108).
- Few-shot evaluation: Few-shot mean macro-MCC falls from 0.488 full-shot to 0.253 at 10-shot and 0.106 at 1-shot, with category collapses near random performance for virus/phage, DNA methylation, and lncRNA.Relative reductions are 48.2% and 78.2%, while promoter and species-classification retain more full-shot signal at 1-shot.
5. Conclusion
GENEB provides a unified framework for comparing 40 genomic foundation models across tasks and categories. Its results show that scale alone is an imperfect predictor, while category-aware evaluation is needed because rankings change with supervision and task domain.
- GENEB evaluates 40 genomic foundation models across 100 tasks from 13 functional categories under a unified linear-probing protocol.
- Model scale correlates with aggregate performance, but architecture and pretraining alignment frequently offset substantial parameter differences.
- Few-shot evaluation changes the leading model in 8 of 13 categories, showing that full-supervision winners are not universally optimal.
- Transformer models generally outperform the evaluated state-space alternative, with domain-specific exceptions such as chromatin accessibility.
- GENEB supports category-aware, controlled evaluation rather than relying on aggregate leaderboards for model comparison and selection.
6. Limitations
GENEB’s conclusions are bounded by its metric aggregation, task and model coverage, domain representation, and frozen-representation evaluation protocol. These constraints limit how broadly aggregate rankings and probing results should be interpreted.
- Long-range tasks: GENEB underrepresents tasks requiring explicit modeling of regulatory interactions longer than 10 kb.
- Task selection and curation: Task selection depends on available datasets and existing benchmarks, while some hard-regime tasks may have noisy or weakly defined labels.
- Model coverage: Some models are excluded because of unavailable weights, incompatible pipelines, or computational constraints.
- Prokaryotic and viral task gap: Only virus/phage classification represents noneukaryotic domains, making aggregate rankings unreliable for broader prokaryotic or viral genomics.
- Frozen representations and pooling-tokenization interactions: Linear probing may underestimate task-specific fine-tuning performance, and mean pooling does not fully disentangle pooling–tokenization interactions.
- Aggregate metric: Micro-averaging is biased toward histone-modification and promoter tasks, which together comprise over half of the per-task evaluations.
7. Use of Large Language Models
Large language models assisted with manuscript writing, editing, organization, phrasing, and LaTeX formatting. The authors performed the experimental design, data processing, analysis, interpretation, and verification of reported findings.
- Large language models were used for writing and editing assistance, including clarity, organization, phrasing, and LaTeX formatting.
- The authors performed all experimental design, data processing, analysis, interpretation, and verification against benchmark outputs.
Impact Statement
GENEB is intended to improve rigor in genomic model comparison through unified, category-aware evaluation. Its controlled comparisons may inform model selection and future design, while dataset biases require careful use beyond represented domains.
- GENEB replaces heterogeneous single-paper evaluations with a unified protocol across 40 models and 100 tasks.
- Category-aware evaluation reduces the risk of selecting models from aggregate leaderboards that mask heterogeneity across biological task types.
- GENEB is skewed toward eukaryotic, human, and well-studied model-organism settings, which may under-represent other biologically and clinically important domains.
- Users should consult per-task and per-category results when applying findings beyond domains directly represented in the benchmark.
A. Related Work
Genomic foundation models span diverse architectures, tokenization schemes, pretraining strategies, and application-specific designs, but existing benchmarks evaluate limited and incompatible model subsets. GENEB addresses this fragmentation through exhaustive, controlled comparison across a broad model and task space.
- Architectures: Genomic models span Transformer encoders, autoregressive decoders, state space and convolutional architectures, hybrids, multimodal systems, and long-range designs.Prior work varies in architectural design, tokenization, pretraining strategy, and benchmark scope.
- Tokenization: Tokenization choices trade fine-grained biological resolution against sequence length, vocabulary size, mutation sensitivity, and computational efficiency.Single-nucleotide, k-mer, and BPE-based approaches make different efficiency and biological-relevance compromises.
- Benchmarking gaps: Existing benchmarks cover important biological domains but remain fragmented in task scope, preprocessing assumptions, evaluation protocols, and model coverage.Prior resources often target limited task families or biological regimes and assess only small model subsets.
- GENEB: GENEB aggregates 100 tasks from multiple established benchmarks across 13 functional categories and evaluates all included models under matched conditions.The benchmark produces a complete performance matrix over 40 genomic foundation models, including recent architectures absent from many existing evaluations.
- GENEB: GENEB is positioned as a unified, community-facing reference framework analogous to MTEB rather than a single-task leaderboard.Planned public release and hosted evaluations aim to support transparent comparison and reproducible assessment of future models.
B. Task Taxonomy and Benchmark Composition
GENEB organizes 100 classification tasks into 13 functional categories spanning regulatory, epigenomic, evolutionary, and cross-taxonomic genomic challenges. It evaluates 40 models while documenting exclusions caused by reproducibility, execution, computational, and design constraints.
- Task organization: 100 classification tasks are organized into 13 functional categories spanning regulatory prediction, epigenomic detection, and evolutionary sequence analysis.Category sizes range from single-task groups to Histone Modifications with 30 tasks and Promoters with 22.
- Epigenomic and Chromatin Tasks: Histone Modifications is the largest category with 30 tasks, while Chromatin Accessibility contributes one complementary DNase-I task.The histone tasks cover multiple marks and draw from NT, NT-rev, and GUE sources.
- Regulatory and splicing tasks: Promoter recognition contains 22 tasks, enhancer prediction contains 8, TF Binding contains 5, and Splice Site detection contains 7.These tasks combine sources including NT, NT-rev, GUE, GB, and iPro-WAEL.
- Additional Tasks: Additional tasks cover mouse enhancers, virus/phage detection, coding discrimination, and transfer across human, plant, prokaryotic, and eukaryotic sequences.The predominance of eukaryotic tasks creates systematic disadvantages for prokaryotic-focused models.
D.3. Excluded Long-Range Regulatory Tasks
GENEB excludes tasks requiring explicit very long-range regulatory modeling because most evaluated models cannot process the necessary genomic context fairly. This omission limits evaluation of models whose architectural priors target long sequences.
- Excluded task regimes: GENEB excludes regulatory tasks requiring explicit modeling of interactions or effects beyond 10 kb.Excluded regimes include enhancer–promoter contacts, three-dimensional chromatin contacts, distal eQTL effects, and whole-locus expression.
- Fair-comparison constraint: Most GENEB models accept context windows below 6 kb, making fair evaluation of 50–500 kb enhancer–promoter interactions impossible without arbitrary cropping.Megabase-scale Hi-C prediction similarly exceeds the input limits of nearly all evaluated models.
- Affected models: Long-context models such as HYENADNA-LARGE-1M, CADUCEUS-PS-131K, EVO-1-131K, and JANUSDNA-72-W are not tested where their architectural priors might yield differentiating gains.The authors identify extending GENEB with unified long-range regulatory tasks as future work.
- Evaluation scope: Frozen-representation evaluation uses linear probing, which may obscure information accessible only through nonlinear readouts.Complementary MLP-probe and regularization analyses are used to assess whether conclusions depend on this choice.
E.1. Probe Stability Analysis
Probe-stability analyses indicate that GENEB’s linear-probe rankings closely track rankings from nonlinear MLP probes. Absolute performance changes are small, supporting the robustness of the main comparative conclusions within the tested subset.
- Rank stability: ρ = 0.964 across 143 model–task pairs and ρ = 0.973 for per-model average MCC show highly stable rankings between linear and MLP probes.All other evaluation components, including feature extraction, normalization, splits, and random seeds, were held identical.
- Rank stability: The top-3 and top-5 models are identical under both probes, while 12 of 13 tasks have positive per-task rank correlations.The median per-task correlation is ρ = 0.855.
- Exception: The sole exception, GB ENSEMBL REGULATORY, has tightly clustered MCC values, making rank ordering sensitive to stochastic noise.The reported negative correlation is attributed to noise rather than substantive representation differences.
- Absolute performance: The mean signed MCC difference is +0.011, and HYENADNA-LARGE-1M has the largest shift at +0.052 without changing its relative rank.These shifts indicate only a marginal aggregate benefit from nonlinear readouts.
- Conclusion: Within the representative subset, linear-probe orderings are a reliable proxy for nonlinear-probe orderings.The authors conclude that the main GENEB findings are unlikely to be artifacts of the linear-probe choice.
E.2. Few-Shot Protocol Sensitivity
Few-shot model rankings are highly stable under regularization changes in the most data-constrained regime and remain stable for typical settings at higher shot counts. The severe degradation from full-data to 1-shot persists across all tested regularization strengths.
- 1-shot rankings are nearly invariant across regularization strengths, with mean pairwise ρ = 0.993 and minimum ρ = 0.982.This stability indicates that rankings in the most data-constrained regime are dominated by representation quality rather than regularization choice.
- At 10-shot, mean pairwise ranking correlation is ρ = 0.805, while adjacent regularization values yield ρ ≥0.9.Larger divergences occur between extreme settings, such as C = 0.01 versus C = 100, where ρ = 0.582.
- Full-data rankings show greater sensitivity, with mean pairwise ρ = 0.766 and minimum ρ = 0.436 across regularization strengths.Absolute MCC sensitivity remains below 0.10 for most models, but is largest for LUCAONE, HYENADNA-LARGE-1M, and NT-V2-50M-MS.
- The sharp mean-MCC degradation from full-data to 1-shot is replicated at every tested regularization strength.Both the direction and severity of this low-data degradation are preserved across the sweep.
F.4. Architecture Comparison via Controlled Experiments
Controlled comparisons show that architecture, tokenization, pretraining data, and scale interact, so no single design choice consistently determines genomic-model performance. Across the benchmark, alignment and configuration often outweigh parameter count, while few-shot behavior and aggregate rankings expose additional trade-offs.
- Architecture: Transformer models outperform the evaluated Mamba alternative under matched multi-species and BPE conditions.GENOMEOCEAN-500M exceeds ECCDNAMAMBA by +0.131 macro-MCC (0.550 vs. 0.419), while OMNI-DNA-1B exceeds it by +0.149 (0.568 vs. 0.419).
- Architecture: Encoder models exceed decoders across all six matched Transformer pairs, with a strongest gap of +0.127 macro-MCC.GENA-LM-LARGE-T2T scores 0.552 versus OMNINA-220M at 0.425 under matched multi-species/BPE conditions.
- Tokenization: BPE outperforms alternative tokenization schemes in several matched comparisons, but tokenization has no single global ordering across designs.GPT2-GENE-V1 exceeds BIOFM-265M by +0.134 macro-MCC (0.476 vs. 0.342), whereas the broader controlled evidence shows interactions with architecture and pretraining data.
- Scale and alignment: 25 of 36 in-domain comparisons show a model at least 5× smaller outperforming its larger counterpart.MUTBERT (86M, 0.529) exceeds EVO-1-131K (7B, 0.298) by +0.231 macro-MCC, although this comparison is confounded by EVO-1's prokaryotic-only pretraining against mostly eukaryotic tasks.
- Few-shot evaluation: Few-shot rankings and robustness do not track full-shot quality consistently across capacity tiers.Capacity-tier degradation ranges from 75.4% for tiny models to 80.9% for small models, with no monotonic size-robustness trend; low absolute drops can reflect a low full-shot ceiling rather than useful signal retention.
- Pretraining data: Multi-species pretraining is favored in 11 of 13 task categories in the GENA-LM comparison.Across six human-vs. multi-species controlled comparisons, multi-species models are favored for Chromatin Accessibility in 6/6 pairs and lncRNA in 6/6 pairs, while human-only pretraining retains a small Virus/Phage advantage.
F.13.1. EASY TASKS: APPROACHING CEILING PERFORMANCE
Eighteen tasks achieve mean MCC above 0.70 across all 40 models, indicating robust performance on several promoter, species-classification, and coding-status problems.
- 18 tasks achieve mean MCC exceeding 0.70 across all 40 models.
- Promoter recognition: Human cell-type-specific promoter recognition is especially strong, with mean MCC values of 0.890, 0.875, and 0.855 across three tasks.SPACE leads all three tasks, while BIOFM-265M ranks lowest.
- Promoter classification: General promoter classification tasks cluster at 0.853–0.855 mean MCC, with GENA-LM-LARGE-T2T and OMNI-DNA-1B among the top performers.
- Species classification: GB Human-or-worm species classification reaches 0.857 mean MCC, while MUTBERT leads at MCC = 0.948 and EVO-1-131K performs worst.
- Coding status: GB Coding/Non-coding achieves 0.803 mean MCC, with GENERATOR-EUKARYOTE-3B leading and CADUCEUS-PH-1K performing weakest.Sequence composition differences provide strong discriminative signal for this task.
F.13.2. HARD TASKS: PERSISTENT CHALLENGES
Twenty-eight tasks have mean MCC below 0.35, with DNA methylation, plant lncRNA, and viral sequence classification remaining difficult for current models.
- 28 tasks have mean MCC below 0.35, representing problems where current models achieve limited success.
- DNA methylation: DNA methylation is particularly challenging, with mean MCC values of 0.061, 0.103, and 0.107 across three tasks.The maximum reported mean MCC is 0.206 by GENERATOR-EUKARYOTE-3B on 4mC G. subterraneus.
- Plant lncRNA: Plant lncRNA tasks reach only 0.221–0.238 mean MCC, although LUCAONE reaches 0.539 on PGB lncRNA S. lycopersicum.LUCAONE achieves the best performance across all six plant lncRNA tasks.
- Viral sequences: GUE COVID variants achieves only 0.200 mean MCC, while GENA-LM-LARGE-T2T leads at 0.472.The result suggests some transferable signal from multi-species genomic pretraining despite evolutionary distance.
F.13.3. HIGH-VARIANCE TASKS: DISCRIMINATING MODEL CAPABILITIES
Thirteen high-variance tasks distinguish model capabilities, revealing strong differences associated with architecture and pretraining alignment rather than uniform performance across models.
- 13 tasks have inter-model standard deviation above 0.12 and serve as stress tests for model capabilities.
- Fungi classification: GUE Fungi-20 has the highest variance, with std = 0.216 and range = 0.764 MCC; GENOMEOCEAN-4B reaches 0.939 MCC versus 0.175 for CADUCEUS-PH-1K.Top performers use multi-species or eukaryotic-gene pretraining, whereas bottom performers use narrower data.
- Splice detection: Splice tasks show high variance, with standard deviations of 0.195, 0.186, and 0.170 across donor, acceptor, and reconstruction tasks.EVO-1-131K consistently ranks last and achieves negative MCC on GUE Splice reconstr.
- Enhancer prediction: Mouse enhancer tasks have std = 0.151–0.175, with ENFORMER and SPACE dominating while human-only JANUSDNA and CADUCEUS variants perform poorly.The pattern supports taxonomically aligned pretraining for cross-species regulatory-element prediction.
- Architecture patterns: Among top performers on high-variance tasks, Transformer-decoder appears 18 times and Transformer-encoder 15 times, while Mamba appears 17 times among bottom performers.CNN-Transformer architectures appear 6 times among top performers; Hybrid-Mamba-MoE and StripedHyena appear 7 and 6 times among bottom performers.
- Pretraining patterns: Multi-species training appears 20 times and eukaryotic-gene training 12 times among top performers, whereas human-only training appears 29 times among bottom performers.Prokaryotic training appears 6 times in bottom positions despite comprising only one model.