Source-linked AI summary

The Geometric Alignment Tax: Tokenization vs. Continuous Geometry in Scientific Foundation Models

Prashant C. Raju

arXiv:2604.04155v1cs.LGcs.ITq-bio.QMstat.ML

TL;DR

Foundation models for biology and physics are evaluated mainly by predictive metrics that do not establish whether their representations preserve continuous system geometry. This paper combines controlled synthetic ablations, rate-distortion theory, MINE, and biological tests to characterize the Geometric Alignment Tax. It finds that discrete tokenization drives geometric instability, with continuous heads reducing distortion by up to 8.5×, while biological robustness signals can reflect sequence composition rather than learned symmetry.

  • Problem

    Predictive metrics do not establish whether foundation-model representations preserve the continuous geometry of biological and physical systems.

  • Method

    The paper isolates tokenization through controlled synthetic ablations and analyzes 14 biological foundation models using rate-distortion theory, MINE, and symmetry controls.

  • Results

    Discrete tokenization is the dominant source of geometric instability: continuous heads reduce distortion by up to 8.5×, while finer codebooks worsen geometry despite improving reconstruction.

  • Takeaways & Limitations

    Geometric evaluation must consider local geometry, global coherence, and information jointly because no studied model achieves all three simultaneously.

  • Takeaways & Limitations

    The empirical analysis covers AI-for-Science models up to 15B parameters and 1M context length, so mitigation beyond this horizon cannot be ruled out.

Abstract

from arXiv · show

Foundation models for biology and physics optimize predictive accuracy, but their internal representations systematically fail to preserve the continuous geometry of the systems they model. We identify the root cause: the Geometric Alignment Tax, an intrinsic cost of forcing continuous manifolds through discrete categorical bottlenecks. Controlled ablations on synthetic dynamical systems demonstrate that replacing cross-entropy with a continuous head on an identical encoder reduces geometric distortion by up to 8.5x, while learned codebooks exhibit a non-monotonic double bind where finer quantization worsens geometry despite improving reconstruction. Under continuous objectives, three architectures differ by 1.3x; under discrete tokenization, they diverge by 3,000x. Evaluating 14 biological foundation models with rate-distortion theory and MINE, we identify three failure regimes: Local-Global Decoupling, Representational Compression, and Geometric Vacuity. A controlled experiment confirms that Evo 2's reverse-complement robustness on real DNA reflects conserved sequence composition, not learned symmetry. No model achieves simultaneously low distortion, high mutual information, and global coherence.

1 Introduction

The paper argues that discrete tokenization imposes a Geometric Alignment Tax by distorting continuous manifolds, even when resolution and model scale improve predictive metrics. Its analyses isolate tokenization as the bottleneck, identify distinct failure regimes, and show that apparent reverse-complement robustness can reflect composition rather than learned symmetry.

  • Problem: The Geometric Alignment Tax is the intrinsic distortion incurred when continuous physical manifolds are forced through discrete categorical bottlenecks.The paper distinguishes this geometric failure from standard predictive metrics such as perplexity, AUC, and benchmark rankings.
  • Motivation: Resolution does not restore continuity: shrinking discrete bins reduces macroscopic error while leaving fractured manifolds governed by impractically slow convergence.The paper compares this behavior to constructing a ramp from discrete blocks, where cumulative directional error persists as block size shrinks.
  • Core claim: 1.3× versus 3,000×: continuous objectives keep Transformer, SSM, and hybrid stability within a narrow range, whereas discrete tokenization produces extreme divergence.The comparison spans synthetic continuous systems and a biological mutation walk, with tokenization changing while the broader architectural mechanisms differ across tracks.
  • Core claim: Finer learned codebooks improve reconstruction but worsen geometric stability, while ESM-2 scaling shows declining stability and a misleading 15B recovery caused by global manifold drift.The paper attributes the codebook double bind to denser decision boundaries and interprets the ESM-2 recovery as drift rather than genuine geometric improvement.
  • Biological validation: Evo 2’s apparent reverse-complement robustness on real DNA reflects conserved k-mer composition rather than learned symmetry.A controlled four-condition experiment supports this interpretation.
  • Information-theoretic analysis: Across 14 foundation models, MINE identifies Local-Global Decoupling, Representational Compression, and Geometric Vacuity as three failure regimes.These regimes distinguish shallow local encoding, information concentration with reduced fidelity, and smooth embeddings lacking meaningful biological structure.

2 Ground Truth: The Controlled Experiments

Controlled experiments isolate tokenization and output objectives as major determinants of geometric stability. Continuous heads preserve geometry across architectures, whereas discrete tokenization, codebook refinement, and biological mutation walks reveal fracture, divergence, and misleading smoothness.

  • Controlled setup: Three small architectures are trained from scratch on synthetic dynamical systems with known geometry using 256-bin causal language modeling.The systems are SmallBERT, SmallMamba, and SmallStripedHyena, evaluated on WAVEFORM, OSCILLATOR, and LORENZ datasets.
  • Evaluation protocol: The evaluation harness embeds clean and perturbed sequences and measures RDM similarity, perturbation stability, feature split, sample split, and their composite mean.It uses mean-pooled center windows and cosine-distance RDMs to compare representational geometry under perturbation.
  • Baseline discrete CE: 0.036, 0.038, and 0.038: discrete CE preserves Lorenz LLE estimates for SmallBERT, SmallMamba, and SmallStripedHyena near the ground truth 0.037.All estimates are within 3% of ground truth, and butterfly tests confirm attractor preservation across five seeds.
  • Continuous versus discrete: 8.5×: replacing CE with a continuous MSE head reduces SmallStripedHyena Lorenz Procrustes D from 0.072 to 0.0085.The encoder, training data, and perturbation protocol remain unchanged; the cross-architecture spread also contracts from 0.072-0.157 to 0.0085-0.034.
  • VQ bottleneck: VQ codebooks exhibit a double bind: distortion reaches a shallow optimum at K = 64 with D = 0.073, then increases as codebooks become finer.The experiment re-encodes continuous perturbations through learned k-means codebooks rather than perturbing unordered code indices.
  • Additional ablations: Jacobian regularization improves composite stability from 0.467 to 0.491 for SmallBERT but worsens mean validation CE from 0.81 to 1.07.The corresponding SmallStripedHyena sweep shows the same Pareto trade-off, with no setting achieving simultaneously low distortion and low CE.
  • Additional ablations: The LSTM’s Lorenz mean composite score is 0.480 versus SmallBERT’s 0.455, whereas SmallMamba reaches 0.369, implicating continuous ODE parameterization rather than recurrence alone.The MSE-head SmallMamba positive control supports the role of its continuous prior.
  • Track A and Track B: 1.3× versus ∼3,000×: architecture gaps remain small on continuous oscillator interpolation but become enormous on the discrete BRCA1 mutation walk.Track A uses MSE-trained models and Track B uses genomic foundation models, so the comparison illustrates magnitude rather than serving as a controlled causal test.

3 Model Performance at Scale

Scaling experiments show that larger biological foundation models can lose geometric stability, while Evo 2’s apparent reverse-complement robustness on real DNA is explained by conserved sequence composition rather than learned symmetry.

  • Parameter Scale vs. Stability: Composite stability declines monotonically across ESM-2 checkpoints from 0.463 at 8M to 0.391 at 3B parameters.The 15B checkpoint shows an apparent recovery, but Procrustes analysis identifies coherent global drift rather than resolved local stability.
  • Cross-Architecture Results: Across architectures, NT v2 and SaProt degrade with scale, while ProtMamba combines low Procrustes distortion with reasonable perplexity but informationally empty embeddings.The cross-architecture results indicate that structure-aware tokens do not rescue the geometric tax.
  • Parameter Scale vs. Stability: ESM-2 separates into Brittle Glass at 8M–650M and Untethered Gel at 3B–15B under 1% substitution.Brittle Glass has low Procrustes reduction, whereas Untethered Gel has higher reduction indicating coherent global drift.
  • Context Length and the RC Dissociation: Evo 2 gains only modest synthetic-DNA stability from 0.747 at 8K context to 0.817 at 1M, while real-chr22 similarity changes from 0.990 to 0.993.Frozen Head context-tax accuracy remains high across checkpoints: 0.988, 0.980, and 0.993.
  • Context Length and the RC Dissociation: Evo 2’s synthetic-DNA RC RDM similarity is low, with values of 0.139, 0.156, and 0.208 across 8K, 262K, and 1M contexts.This indicates failure of the A↔T / C↔G bijection under the synthetic test.
  • Context Length and the RC Dissociation: Dinucleotide shuffling recovers 97% of the real-random RC gap because it preserves per-sequence k-mer counts, showing that real-DNA robustness is a histogram artifact.Texture-matched Markov sequences recover only 3%, and the controlled result indicates short-subsequence counting rather than double-stranded symmetry.

4 The Information Theory of the Tax

The paper frames geometric distortion as a rate–distortion consequence of discrete tokenization and uses mutual-information analysis to distinguish three ways biological models trade geometry against information.

  • Rate–Distortion Framing: Cross-entropy over discrete vocabularies creates piecewise-constant embedding geometry because classification boundaries provide no gradient toward manifold preservation.The discrete channel has capacity log2 K bits per token, linking tokenization to a rate constraint.
  • Rate–Distortion Framing: Finer codebooks improve reconstruction but can worsen perturbation stability because denser cell boundaries increase boundary-crossing probability.The observed optimum is K=64, followed by increasing Procrustes distortion through K=1024; Dproc ∝ 1/log K with R2 = 0.98 and p < 0.001 on Lorenz.
  • Scope and Open Questions: The paper treats cross-entropy as a sufficient, not exclusive, source of distortion and leaves architectural equivariance as an open alternative to post-hoc regularization.Embedding-level RCCR achieves perfect per-sequence RC consistency but degrades population-level geometry, suggesting redistribution rather than elimination of the tax.
  • Three Pathologies of the Distortion Bound: Evo 2 exhibits Local-Global Decoupling: global MI exceeds local MI by only 14% despite a 64× context increase.Its positive excess MI reflects biological signal, but the signal remains shallow, scale-invariant local composition rather than long-range structure.
  • Three Pathologies of the Distortion Bound: OpenFold and ESM-1b exemplify Representational Compression, while ProtMamba exemplifies Geometric Vacuity with negative excess MI at every sequence length.The Evoformer adds +2.3 to +2.5 nats of structural context while warping geometry; ProtMamba has low distortion but less biological information than a matched random baseline.
  • Three Pathologies of the Distortion Bound: Discrete-token models do not simultaneously achieve low distortion, high mutual information, and global coherence under the Gaussian rate–distortion constraint.The paper describes the three regimes as different allocations of finite representational capacity.

5 Discussion

The discussion presents discrete tokenization as the dominant source of geometric instability and argues that refining codebooks or scaling models does not remove the problem. It frames geometric auditing and continuous-geometric integration as necessary directions while limiting generalization beyond AI for Science and tested scales.

  • Findings: 1/log K scaling makes continuous-head performance unreachable through codebook refinement alone, while finer quantization can worsen geometric stability.The VQ double bind emerges when cell boundaries become denser than the perturbation scale.
  • Findings: The three MINE failure regimes represent different allocations of finite capacity: local geometry trades off with global coherence, information concentration with geometric fidelity, and geometry with information.
  • Implications: Predictive metrics can miss ungrounded global geometry or smooth manifolds that encode no biological signal.The discussion therefore proposes Physical Alignment and geometric stability auditing alongside task performance.
  • Implications: The authors argue that scaling larger discrete models cannot simply eliminate the penalty, while continuous priors may erase biological signal.
  • Limitations: The findings are scoped to AI for Science and models up to 15B parameters with 1M context length, leaving emergent mitigation beyond that horizon unresolved.The biological benefit of continuous objectives remains to be demonstrated while retaining task performance.
  • Future directions: Future work likely requires architectures that natively unify continuous geometric priors with high-fidelity discrete encoding or jointly optimize prediction and manifold preservation.

Code

This section provides reproducibility resources and situates the paper among work on representation geometry, quantization, scientific foundation models, and symmetry enforcement.

  • Code: The full code, experiments, benchmarks, and analyses are publicly available in the paper’s repository.Infrastructure, software versions, and configurations are documented in Appendix J.
  • Related work: RSA and CKA compare neural representations, whereas linear probing measures downstream discriminative utility without directly testing manifold topology.
  • Related work: Classical quantization theory predicts reconstruction distortion scaling as K−2/d, while this paper distinguishes geometric distortion under perturbation from reconstruction error.
  • Related work: Scientific foundation models span protein language models, genomic sequence models, and chaotic-system time-series forecasters across evolving architectures.
  • Related work: Physical symmetries in discrete-token models can be enforced architecturally through equivariance or during training through post-hoc regularization.

B Extraction Details and Evaluation Metrics

The evaluation harness measures whether embeddings preserve representational geometry under perturbations using split-based, RDM, anchor, and composite stability metrics.

  • Extraction protocol: All evaluated models use cosine distance for RDM computation and mean-pooling over a center window unless otherwise noted.
  • Harness: Shesha extracts clean and perturbed representation geometries over equivalent context windows, with optional stratification and bootstrapping.
  • Metrics: Sample Split tests whether sample identities preserve proximity metrics across perturbations.
  • Metrics: Feature Split tests whether variance directions remain consistent across independent feature subspaces.
  • Metrics: RDM Similarity correlates clean and perturbed representational topologies, while Anchor Stability measures distance consistency from fixed anchor points.The anchor metric uses rank correlations between anchor-to-subset distance structures.
  • Metrics: The Composite Stability Score pools the component results into a unified model-comparison metric.

C Track A + Track B Extended Methods and Details

The appendices detail two tracks: controlled synthetic-physics experiments with discretized oscillator trajectories and a biological BRCA1 mutation walk across four genomic models. They report extraction procedures, architecture specifications, and geometric-stability outcomes.

  • Track A: Track A uses damped harmonic oscillator trajectories sampled at 512 points, globally discretized into 256 bins, with 50,000 training and 2,000 validation sequences.
  • Track A: The three Track A architectures share dmodel = 256, vocabulary size 258, and sequence length 512, with distinct Mamba and Hyena configuration details.
  • Track A: Models are trained under identical causal-language-modeling conditions with cross-entropy on discretized physical signals.
  • Track A: Ten trajectory pairs generate 101 interpolated sequences spanning distinct dynamical regimes, whose hidden states are mean-pooled for geometric analysis.
  • Track A: PCA arcs, cosine distance from the start, and L2 Lipschitz profiles assess trajectory smoothness, cumulative drift, and local sensitivity.
  • Track A: Mean Lipschitz values are 65.3 for SmallBERT, 79.3 for SmallStripedHyena, and 84.6 for SmallMamba, a 1.3× spread.
  • Track B: Track B constructs a 122-sequence BRCA1 mutation walk containing one pathogenic C61G substitution and 120 additional random SNPs.

D Complete Ablation Battery

The ablation battery identifies tokenization and discrete optimization as the dominant sources of geometric instability, while continuous objectives and continuous-time structure preserve geometry more reliably across architectures and scales.

  • D.1 Variant A: Continuous MSE Head: Replacing discrete CE with continuous MSE eliminates manifold fracture and sharply reduces Lorenz Procrustes distortion across architectures.SmallBERT improves 2.8× and SmallStripedHyena 8.5× at 1% noise; continuous-head distortion spans 0.0085–0.034 versus 0.072–0.157 under CE.
  • D.2 Variant B: Jacobian Norm Penalty: A Jacobian penalty improves geometric stability but worsens validation CE, producing a Pareto frontier rather than simultaneously low distortion and low prediction error.For SmallBERT, stability rises from 0.467 to 0.491 while validation CE increases from 0.81 to 1.07 as λ increases.
  • D.4 Variant D: LSTM Baseline: The LSTM fractures like the Transformer, whereas SmallMamba preserves attractor geometry, ruling out recurrence alone as the source of SSM stability.LSTM RDM similarity falls from 0.900 at 1% waveform noise to 0.600 at 10%, while SmallMamba remains above 0.99.
  • D.6 Variant F: Hyena Filter Order Sweep: Hyena filter order has little effect on geometric stability, indicating that the presence of continuous convolution matters more than deeper filter order.Lorenz composite stability remains near 0.471–0.477 across orders 1–8, with no systematic improvement after parameter normalization.
  • E.11 Cross-Architecture Scaling: Across biological models, Transformer stability generally declines with scale, while Caduceus remains nearly constant and structure-aware tokens do not rescue SaProt from degradation.ESM-2, Nucleotide Transformer, and SaProt show progressive tax patterns; Caduceus is near-constant across scale with near-perfect reverse-complement preservation.
  • E.11 Failure Regimes: Apparent large-model rebounds can reflect global manifold drift rather than genuine geometric recovery, while ProtMamba exhibits geometric vacuity and OpenFold representational compression.ProtMamba has negative excess mutual information, whereas OpenFold has high excess mutual information but geometric warping.

E.11 Cross-Architecture Summary

The cross-architecture summary compares geometric stability across model families and evaluation protocols, emphasizing progressive Transformer degradation, stable Caduceus behavior, and diagnostics beyond raw stability.

  • Cross-Architecture Patterns: Mean composite stability across synthetic sequences shows progressive degradation in ESM-2, Nucleotide Transformer, and SaProt as model scale increases.The summary orders models by parameter count within each domain and reports declining composite stability for these Transformer families.
  • Cross-Architecture Patterns: Caduceus remains nearly constant across scale with perfect reverse-complement preservation, making it the tax-exempt baseline in the comparison.Its composite stability is reported as 0.459–0.458 across scale.
  • Cross-Architecture Patterns: The continuous regime has a 1.3× cross-architecture spread, whereas the discrete regime spans three orders of magnitude in Lipschitz constants.This comparison supports tokenization as more consequential than architectural class for geometric stability.
  • Cross-Architecture Patterns: SaProt’s structure-aware tokens do not prevent the geometric tax, degrading at the same rate as ESM-2 at matched sizes.Foldseek structural tokens are therefore insufficient to remove the observed scale-related degradation.
  • Evaluation Protocol: The evaluation framework requires internal embeddings to compute Procrustes stability, RDM similarity, and MINE mutual information.The Ghost Detection protocol additionally tests whether apparent geometric stability remains connected to downstream utility.
  • Evaluation Boundary: Evo 2 40B and AlphaGenome were excluded because their APIs did not provide usable intermediate representations for geometric analysis.The analysis therefore restricts Evo 2 to the locally run 7B family; future direct weight access could extend coverage.
  • Evaluation Protocol: Local pooling isolates signal-region geometry, whereas global pooling tests whether random padding corrupts the full representation.The protocol uses these two pooling strategies for controlled species-classification tests.

F.2 Per-Model Ghost Detection Results

Ghost Detection reveals distinct interactions between sequence scaling, geometry, and biological information across Evo 2, ProtMamba, and OpenFold. Longer context does not uniformly improve frozen-head performance and can expose geometric or representational failure.

  • Evo 2: 0.963, 0.964, and 0.963: Evo 2’s SNP Procrustes ratio remains stable across 8K, 262K, and 1M contexts.Reverse-complement ratios show no systematic context-length trend, ranging from 0.744 to 0.732.
  • Evo 2: 128× more context yields only ∆= +0.005 accuracy for Evo 2, effectively no classification gain from 8K to 1M.On real chr22 DNA, SNP Procrustes ratios likewise remain stable at 0.984–0.985.
  • Evo 2: 85.3%, 84.9%, and 85.2%: Evo 2’s clean-versus-perturbed top-1 agreement is flat across context windows, while reverse-complement agreement falls to ∼28–30%.KL divergence also remains flat for SNPs but is much higher under reverse complement.
  • ProtMamba: ∼50% accuracy at every sequence length: ProtMamba contains no extractable species information despite stable geometry and perplexity of 40.1.Both linear and MLP probes remain at chance under local and global pooling, confirming Geometric Vacuity.
  • OpenFold: 7.0 percentage points: OpenFold accuracy drops as length increases eightfold, from 0.843 at L=100 to 0.773 at L=800.Its embeddings remain linearly informative, but the biological signal degrades with padding length.
  • Cross-model interpretation: 22–27% reduction: Evo 2 occupies a stable intermediate geometric position that does not vary with context length.Across architectures, the Brittle Glass / Untethered Gel transition is reported as a general phenomenon of discrete Transformer scaling.

G RCCR Experiment Details

The RCCR experiment tests whether embedding-level reverse-complement consistency can repair geometry on a discrete DNABERT-2 backbone. It produces perfect per-sequence consistency but worsens population-level geometric structure.

  • Training outcome: 99.4% reduction: RCCR training loss falls from 1.56 × 10^-3 to 9 × 10^-6 over 10 epochs.The experiment fine-tunes DNABERT-2 (117M) on 2,000 random DNA sequences.
  • Per-sequence consistency: 0.041 to 0.000: the per-sequence reverse-complement cosine gap collapses to zero under RCCR.Every sequence maps to the same point as its reverse complement.
  • Interpretation: RCCR achieves perfect per-sequence consistency but degrades population-level geometric structure across perturbation conditions.The result separates pointwise agreement from preservation of relational embedding geometry.
  • Population geometry: 91% increase: Procrustes disparity between forward and reverse-complement embedding matrices rises from 0.761 to 1.454.The composite score increases slightly because feature-split performance improves, while geometric-fidelity metrics degrade.

H MINE Extended Methods and Details

The MINE analysis estimates biological information in model embeddings using PCA, bias-corrected mutual information, and repeated neural estimation. It compares local and global signals across protein and DNA models while calibrating estimator behavior.

  • Replication: 5 independent runs: each experiment reports mean ± standard deviation across fixed seeds.The same seed sequence controls network initialization, NumPy randomness, and minibatch permutation.
  • Pre-processing: 50 principal components: raw embeddings are reduced before MINE to stabilize the Donsker–Varadhan bound in high dimensions.The retained components explain more than 90% of variance for all tested models.
  • Bias correction: MIexcess = MImodel − MIrandom: reported information is measured relative to a matched random baseline.The correction addresses finite-sample upward bias that scales with embedding dimensionality.
  • Local versus global information: 64× more context yields only a 14% increase in excess MI for Evo 2, supporting a local–global information gap.Local embeddings are compared with 128-bp features, while global embeddings use the full 8,192-token context.
  • Robustness: ProtMamba remains at or below the noise floor, while ESM-1b, OpenFold, and Evo 2 preserve positive or local–global information patterns across PCA settings.The qualitative regime ordering is stable at 30, 50, and 100 components.
  • Estimator validation: 0.15 nats: run-to-run standard deviation remains below this value across conditions after convergence.All models converge within 300 epochs, and the final 50 epochs form stable plateaus.

I Texture Hypothesis Test (Evo 2 Reverse Complement Mechanism)

The Texture Hypothesis Test separates learned reverse-complement symmetry from conserved sequence composition in Evo 2. Dinucleotide shuffling preserves per-sequence composition and nearly reproduces real-DNA reverse-complement stability, while Markov matching of population statistics does not.

  • Experimental design: The four-condition experiment uses 10,000 sequences per condition, each 1,000 bp, evaluated with Evo 2 7B at 8K context.Conditions include real chr22, pure random, texture-matched Markov, and dinucleotide-shuffled real DNA.
  • Control interpretation: Dinucleotide shuffling preserves each parent sequence’s compositional fingerprint while destroying positional organization, retaining pairwise diversity.It preserves exact per-sequence dinucleotide counts and lower-order frequencies.
  • Mechanism: Reverse-complement k-mer frequencies are related by a fixed complement permutation, so exact per-sequence counts can generate symmetric forward–RC representations without positional understanding.The test compares paired k-mer similarities before the embedding experiment.
  • Key comparison: 97% of the real–random gap: dinucleotide-shuffled DNA recovers nearly all real-DNA reverse-complement RDM stability, 0.858 versus 0.873.Its composite recovery is also 97%, at 0.464 versus 0.469.
  • Key comparison: 3% of the real–random gap: texture-matched Markov sequences recover little reverse-complement stability, with RDM values 0.167 versus 0.139.Matching population-level dinucleotide frequencies is insufficient when individual sequence compositions converge toward the population mean.
  • Mechanism: The Markov condition’s apparent stability is uninformative because nearly identical k-mer profiles collapse the pairwise RDM.The model cannot distinguish sequences from one another, making forward and reverse-complement versions look alike through representational collapse.
  • Caveat: Per-sequence dinucleotide composition is the dominant signal; higher-order texture may contribute only marginally under this test.A second-order Markov chain is suggested as a follow-up for stronger evidence about trinucleotide sensitivity.

J Reproducibility and Computational Infrastructure

The study documents its statistical procedures, computational safeguards, deterministic execution practices, hardware, software stack, and implementation patches to support reproducible geometric evaluations.

  • Bootstrap evaluations used five stratified resampling rounds and reported bootstrap means for core geometric metrics.Stratification followed sequence class proportions when labels were available and otherwise used uniform random sampling.
  • Datasets exceeding 2,500 samples were subsampled for pairwise RDM computations because of O(n^2) memory complexity.Dedicated Procrustes evaluations used 5,000 or 10,000 samples depending on the target.
  • Fixed random seeds and deterministic filename hashing made generated sequences, embeddings, and perturbation traces reproducible across architectures.The default seed was 320 using numpy.random.default_rng.
  • Experiments ran on Google Colab Pro instances with NVIDIA A100 and T4 GPUs, while AlphaGenome scale tests used API-based gRPC inference.The AlphaGenome tests required no local GPU allocation.
  • Full-scale evaluations over 10,000 genomic sequences required approximately 15 minutes to 20 hours per architecture, depending on context length.
  • The software environment included PyTorch, Transformers, mamba-ssm, shesha-geometry, SciPy, and scikit-learn, with compatibility patches for two model codebases.The patches addressed embedding extraction issues caused by Transformers 4.x API changes.
Loading 2604.04155v1…