Source-linked AI summary
Systematic Evaluation of Single-Cell Foundation Model Interpretability Reveals Attention Captures Co-Expression Rather Than Unique Regulatory Signal
Ihor Kendiukhov
TL;DR
The paper addresses whether attention-derived edges in single-cell foundation models capture causal regulatory information beyond gene-level statistics. It evaluates this question systematically across architectures, cell types, perturbation modalities, and causal interventions, finding that gene-level features dominate prediction and pairwise edges add no predictive value. CSSI improves curated GRN recovery under heterogeneous cell states, but the paper’s causal and predictive conclusions remain bounded by its tested architectures, biological settings, and lack of a decisive positive control.
Problem
Current interpretability practice assumes that attention patterns reflect causal regulation and transfer across biological contexts, but systematic validation beyond expression statistics and causal interventions has been lacking.
Method
The study applies a two-tier framework of 37 analyses and 153 tests to scGPT and Geneformer across four cell types and two perturbation modalities, including CSSI and causal ablation.
Results
Gene-level features dominate perturbation prediction: pairwise edges add zero incremental value, and ablating regulatory-ranked attention heads does not degrade performance.
Takeaways & Limitations
Attention can encode structured biological information without providing validated causal regulatory signal, while CSSI improves curated GRN recovery by controlling cell-state heterogeneity.
Takeaways & Limitations
Causal ablation covers only Geneformer V2-316M, additional tissues and architectures remain untested, and no decisive positive control establishes recovery of known causal regulatory structure.
Abstract
from arXiv · showhide
We present a systematic evaluation framework - thirty-seven analyses, 153 statistical tests, four cell types, two perturbation modalities - for assessing mechanistic interpretability in single-cell foundation models. Applying this framework to scGPT and Geneformer, we find that attention patterns encode structured biological information with layer-specific organisation - protein-protein interactions in early layers, transcriptional regulation in late layers - but this structure provides no incremental value for perturbation prediction: trivial gene-level baselines outperform both attention and correlation edges (AUROC 0.81-0.88 versus 0.70), pairwise edge scores add zero predictive contribution, and causal ablation of regulatory heads produces no degradation. These findings generalise from K562 to RPE1 cells; the attention-correlation relationship is context-dependent, but gene-level dominance is universal. Cell-State Stratified Interpretability (CSSI) addresses an attention-specific scaling failure, improving GRN recovery up to 1.85x. The framework establishes reusable quality-control standards for the field.
1 Introduction
The paper systematically tests whether attention-derived edges provide mechanistic regulatory information beyond gene-level expression statistics. Across models, contexts, and interventions, attention encodes structured but non-incremental information, while CSSI improves attention-derived GRN recovery under heterogeneous cell states.
- Evaluation framework: Attention-derived edges are evaluated through 37 analyses and 153 statistical tests spanning two architectures, four cell types, and two perturbation modalities.The framework includes baseline comparisons, incremental-value tests, residualisation, propensity matching, and causal ablation.
- Main finding: Attention patterns show layer-specific structure, but convergent tests find no regulatory information beyond gene-level features.The evidence includes residualisation, null models, incremental-value testing, and causal interventions.
- CSSI and causal tests: CSSI improves top-K TRRUST recovery up to 1.85× at 5–7 strata, but this improves curated-edge recovery rather than perturbation-outcome validity.Causal ablation likewise finds no degradation from masking regulatory-ranked heads, while interventions materially perturb internal representations.
- Perturbation prediction: AUROC 0.81–0.88 for gene-level baselines exceeds 0.70 for attention and correlation edges in CRISPRi target prediction.Variance, mean expression, and dropout rate each outperform both pairwise edge types.
- Incremental value: ∆AUROC = −0.0004 with attention and −0.002 with correlation shows that adding pairwise edges to gene-level features provides no incremental value.This null persists across harder gene and perturbation splits and linear and nonlinear models.
- Confounding: 76% of attention’s above-chance TRRUST signal disappears after expression residualisation, whereas correlation retains 91%.The asymmetry indicates stronger expression confounding for attention-derived regulatory scores.
3 Discussion
The study finds that attention encodes structured biological information but does not provide causal regulatory signal or incremental perturbation-prediction value beyond gene-level features. CSSI improves edge recoverability, while the conclusions remain bounded by component, model, and dataset limitations.
- Interpretation: Attention and MLP analyses indicate structured representations, but pairwise edge scores add no regulatory information beyond gene-level features.Convergent residualisation, incremental-value tests, and causal interventions support this conclusion.
- Context dependence: The attention–correlation relationship varies by context, yet gene-level features dominate perturbation prediction in both K562 and RPE1.Raw edge-score comparisons can therefore be misleading without confound controls.
- Limitations: The findings concern attention weights, heads, and MLP blocks rather than the predictive capacity of the complete foundation models.Additional tissues, species, architectures, and decisive positive controls remain limited or untested.
- Constructive contribution: CSSI controls heterogeneity-driven dilution and improves recoverability against curated references, but it does not establish perturbation-outcome predictive validity.Its scope is edge recovery rather than causal perturbation prediction.
- Implications: The authors recommend trivial-baseline and incremental-value tests, dual thresholded and continuous metrics, stratification for heterogeneity, and intervention-fidelity diagnostics.These recommendations are presented as quality-control practices for causal interpretability claims.
- Evaluation framework: The evaluation covers scGPT and Geneformer with multiple datasets, stratification procedures, perturbation analyses, cross-validation protocols, and causal interventions.CSSI partitions cells into embedding-based strata and aggregates within-stratum TF–target correlations.
Extended Data
Extended analyses show metric- and context-dependent scaling, near-random unstratified GRN recovery, and null incremental-value or ablation results. The figures also document cross-context differences and study metadata.
- Scaling: Top-K F1 degraded in all 9 tier×seed runs as cell count increased, while continuous AUROC improved monotonically.The scaling pattern is metric-dependent rather than uniformly negative.
- Multi-model recovery: Both scGPT and Geneformer achieved near-random unstratified AUROC of approximately 0.5 against TRRUST/DoRothEA.CSSI revealed recoverable signal within individual heads.
- Perturbation comparison: Geneformer attention at L13 matched correlation on K562 CRISPRi, with AUROC 0.704 versus 0.703 and p = 0.73.L13 was pre-specified from independent DLPFC data.
- Confounding: Attention edges lost approximately 76% of TRRUST signal after residualisation, compared with approximately 9% for correlation.Degree-preserving null models were shown for both edge types.
- Layer selection: Nested cross-validation selected L15 in every fold, with pooled held-out delta +0.040, but CRISPRa replication reversed the effect with d = −0.56.The figure reports both the pooled selection result and the replication reversal.
- Incremental value: Hard-generalisation tests found no incremental value across cross-gene, cross-perturbation, or joint splits for linear and nonlinear models.The null held across AUROC, AUPRC, and top-k recall.
- Sensitivity: Across 27 sensitivity combinations, AUROC remained above chance with p < 0.005 and ranged from 0.62 to 0.76.The combinations varied LFC thresholds, control-cell counts, and highly variable gene counts.
- Matched evaluation: Propensity matching reduced edge AUROCs to near chance and produced exactly zero logistic-regression incremental value.The decomposition matched DE-positive and DE-negative targets on gene-level covariates.
Supplementary Table 1: Analysis and Dataset Overview
The supplementary analyses examine scaling, baseline performance, metric sensitivity, mediation non-additivity, and ranking robustness across attention-derived and conventional gene-regulatory signals. Scaling can reduce attention-based recovery and stability, while metric choice, compositional heterogeneity, and interaction effects materially shape interpretation.
- Analysis and Dataset Overview: Top-100 attention-derived GRNs were evaluated across scGPT tiers, seeds, cell counts, reference databases, and complementary GRN baselines.The scaling analyses used TRRUST and DoRothEA reference edges, while DLPFC comparisons included correlation, mutual information, GENIE3, GRNBoost2, and attention-based scores.
- Scaling behavior: TRRUST F1 decreased from 200 to 1,000 cells in all 9 tier×seed pairs (p = 0.00195), with the same pattern for DoRothEA.The 1,000→3,000 step continued degradation in 7/9 pairs, but was weaker statistically (p = 0.09).
- Scaling behavior: Recovered true positives moved toward or below random expectation as cell count increased, while edge-set Jaccard overlap fell 46.6–47.9% from 200 to 1,000 cells.The robustness decline was significant across tiers and seeds (one-sided Mann–Whitney p = 2.1 × 10−4).
- Scaling behavior: Cell-type richness increased with sample size and was anti-correlated with TRRUST F1 (Spearman ρ = −0.76, p = 4.3 × 10−5).Controlled-composition experiments separately tested sample size, fixed composition, and increasing heterogeneity.
- Metric sensitivity: Continuous AUROC improved with cell count from 0.858 to 0.925 to 0.934, showing that the scaling result depends on the recovery metric.Top-K F1 was near zero because only 51 TRRUST edges occurred among approximately 3.7 million candidate pairs.
- Baseline comparisons: Dedicated GRN methods and attention achieved similar DLPFC AUROC values of 0.50–0.53, despite requiring 89–127 seconds versus 0.1 seconds for attention extraction.The compared methods included Spearman correlation, mutual information, GENIE3, GRNBoost2, and attention-based scores.
- Mediation and ranking robustness: Non-additivity was positive in 10 of 16 run-pairs, with median Alb/|TE| = 0.725, and certified pair coverage dropped from 0.0669 at λ = 1 to 0.0032 by λ ≥3.These results motivate reporting residual non-additivity and interaction-aware alternatives such as Shapley-value decomposition.
Detectability Phase Diagrams
The detectability analyses compare attention-like and intervention-like signals under varying noise and tissue contexts. Intervention-like signals can require fewer cells under sub-Gaussian conditions, but tail inflation, tissue variation, and technical confounding limit generalization.
- Detectability Phase Diagrams: Under sub-Gaussian baseline conditions, intervention-like signals required 44.4% as many cells as attention-like signals for equivalent detectability.The relative cell requirement approached unity when the tail inflation factor exceeded 3.
- Robust detection: Robust median-based or Huber estimation expanded the feasible detection region by 37% under 10% contamination.Real-data calibration generally produced relative cell ratios below one, but confidence intervals remained wide.
- Cross-tissue consistency: Cross-tissue edge effects ranged from Spearman ρ = −0.44 to 0.71, with only two of six pair-granularity comparisons surviving FDR control at α = 0.05.The observed variability is consistent with tissue-specific regulation but may also reflect tissue-specific confounds.
- Cross-tissue consistency: Technical covariates were recoverable from edge features, so distinguishing regulatory rewiring from artifacts requires matched protocols across tissues.This caveat follows the batch-leakage analysis described alongside the cross-tissue comparison.
Perturbation Validation Details
Perturbation validation finds weak-to-moderate, parameter-sensitive alignment between inferred edges and perturbation outcomes, with attention indistinguishable from correlation and weaker than trivial gene-level baselines. Additional transfer and pseudotime tests show strong context dependence and limited temporal validation.
- Perturbation validation: Cross-dataset perturbation alignment was weak and condition-specific, with the strongest positive adjusted consistency in Dixit 13-day (ρ = 0.199, p = 0.020).Other datasets showed weaker, non-significant, marginal, or adjustment-sensitive relationships.
- Perturbation-first validation: Correlation-based edge scores achieved mean per-gene AUROC 0.696 across 151 evaluable perturbations under the primary parameterization.All 27 sensitivity conditions remained above chance, with AUROC ranging from 0.619 to 0.756.
- Attention versus correlation: Geneformer attention was statistically indistinguishable from correlation at layers 6, 13, and 18, with AUROC values of 0.705, 0.704, and 0.708 versus 0.703.The corresponding p-values were 0.76, 0.73, and 0.75.
- Cross-species transfer: Human–mouse edge scores showed global conservation (ρ = 0.743), but per-TF conservation ranged from near-perfect lineage-factor transfer to poor signaling-responsive transfer.Lineage-specifying factors such as XBP1 and EPAS1 were highly conserved, whereas CTNNB1 and HIF1A were weakly conserved.
- Pseudotime validation: Only 12 of 56 curated TF–target pairs (21.4%) showed directionally consistent pseudotime ordering.The mean directionality score exceeded the shuffled-pseudotime null only marginally (p = 0.068), and pseudotime is recommended as a qualitative sanity check.
Supplementary Note 9 Batch and Donor Leakage Audit
The leakage and calibration audit shows that edge features can encode donor, assay, and dataset artifacts, while calibration improves score reliability but does not transfer across datasets. CSSI gains remain dependent on meaningful cell-state labels and structured heterogeneity.
- Leakage audit: Donor identity was recoverable above chance with AUC 0.85–0.87 in immune tissue and 0.94–0.96 in lung, while assay method reached AUC 0.96–0.99.These results indicate substantial technical signal in edge-product features.
- Dataset stability: The imbalanced immune dataset showed lower resampling stability (r = 0.929), 54.6% blacklisted edges, and 17.1% sign-flipped edges.The balanced lung dataset was more stable, with r = 0.997 and 10.1% blacklisted edges.
- Leakage controls: Donor-stratified splits are required because random cross-validation can obscure a 6.6-percentage-point generalization gap.The donor-balanced resampling results motivate reporting this gap as a quality check.
- Calibration: Isotonic regression reduced raw ECE from 0.269–0.469 to 0.062–0.079, a 4–7× reduction without changing discrimination.All six edge-scoring methods were severely miscalibrated before post-hoc calibration.
- Calibration transfer: K562-trained calibrators failed to transfer to Shifrut T cells, yielding ECE 0.320–0.424 versus 0.002–0.031 for locally trained calibrators.Split conformal prediction sets achieved valid marginal coverage of at least 95% for selected methods at α = 0.05.
- CSSI validation: CSSI-max with oracle labels maintained F1 ≥0.900 across synthetic configurations, whereas shuffled or random labels removed the advantage.The null tests support a label-dependent CSSI effect rather than generic inflation.
- CSSI validation: On structured PBMC data, CSSI-max recovered 22/22 known edges versus 19/22 for pooled inference, especially for cell-type-specific regulatory edges.The tested examples included BCL6⊣PRDM1, IRF8→IL12B, and RORC→IL17A.
N Pooled F1 CSSI-max F1 AUROCpool AUROCCSSI
Geneformer attention shows layer-specific GRN recovery and structured association with expression, but regulatory alignment and causal usefulness remain limited. Attention survives several controls, while perturbation and ablation tests provide little evidence of incremental mechanistic value.
- Layer-specific recovery: 0.694 AUROC is the best pooled Geneformer layer result, reached at L13, while several early layers remain near chance.The all-layer pooled baseline is 0.543; late layers perform better than early layers.
- Attention associations: Attention correlates with expression co-occurrence (ρ = 0.31–0.42) but not regulatory ground truth (ρ = −0.01–0.02).The attention–expression relationship weakens across tissues, with cross-tissue R2 < 0.02.
- Confound control: 0.66 → 0.54 AUROC after residualization shows attention loses ∼76% of above-chance TRRUST signal, whereas correlation retains ∼91%.This comparison uses Geneformer L13 attention edges on K562.
- Causal intervention: Top-5 and top-10 attention replacement leave AUROC at 0.704 and 0.703, while targeted MLP ablations also remain at 0.704.Random-layer MLP ablation at L8 produces a small but significant −0.005 drop, confirming that the intervention can disrupt computation.
- Perturbation prediction: 0.55 versus 0.65 AUROC shows attention underperforms correlation in K562 CRISPRa, while the two methods are indistinguishable in primary T cells.The K562 comparison is statistically significant at p < 10−6.
- Incremental predictive value: ∆AUROC = −0.000 shows attention edges add zero incremental value after matching DE-positive and DE-negative targets by expression profile.The matched analysis used k = 5 controls and 59,153 pairs.
4. Cross-Context Consistency
Cross-context analyses show that interpretability signals vary across datasets and biological settings, with calibration transfer failing and attention–correlation relationships depending on context. Several structured signals remain detectable, but their stability and biological interpretation are conditional.
- Donor effects: AUC 0.85–0.87 and 0.94–0.96 show that edges encode donor identity in immune and lung data, respectively.The donor-prediction analyses use logistic regression with 20k edges.
- Generalization and calibration: 6.6 percentage points is the cross-donor generalization gap, while isotonic calibration reaches ECE 0.06–0.08 across six methods.These findings indicate donor variation and calibration performance must be evaluated separately.
- Synthetic validation: r = 0.887 indicates that empirical detectability matches theoretical predictions in the synthetic validation.The result is reported for the phase-space detection analysis.
13. Controlled-Composition Scaling
Controlled-composition analyses show that attention-derived recovery is sensitive to cell-state composition and heterogeneity, while simple gene-level features often outperform correlation-based edges. Layer selection can improve attention performance, but pairwise scores provide limited additional value.
- Heterogeneity: ρ = +0.63 shows that greater diversity improves Spearman performance across 35 runs.The analysis is part of the perturbation-first validation.
- Edge interpretation: ρ = 0.842 shows that edge scores closely track co-expression across 75,962 pairs.This association is statistically significant at p < 10−50.
- Residual signal: R2 < 0.02 indicates that ordinary least squares explains little residual edge variation after the controlled analysis.The reported analysis covers 21k–26k edges.
- Baselines: Variance and dropout each outperform correlation as gene-level predictors in the controlled comparisons.These are model-free univariate baselines.
22. Bootstrap Per-TF CIs
Bootstrap and nested cross-validation analyses quantify uncertainty around layer and perturbation performance. The strongest attention layer is consistently selected, but performance varies substantially across layers and conditions.
- Split-sample validation: AUROC = 0.750 is obtained in split-sample discovery validation for the selected layer.The validation uses n = 140 and reports p = 0.017.
- Layer variability: AUROC 0.47–0.74 spans the per-layer results, demonstrating substantial heterogeneity across the 18-layer model.Only 1 of 18 layer comparisons is reported as significant in the table summary.
- Residual prediction: AUROC = 0.538, 0.554, and 0.622 are reported for the three residual prediction analyses.Each analysis uses 174 positive examples.
- Subgroup analysis: TF-stratified AUROC is 0.913 for TFs and 0.895 for non-TFs, indicating similar subgroup performance.The comparison is reported as a subgroup analysis.
- Nested validation: ∆ = +0.040 [0.018, 0.062] is the pooled held-out advantage for the selected best layer over correlation.Strict nested cross-validation selects L15 in all five folds.
31. Hard-Generalization Incremental Value
Attention and correlation edges show some signal, but adding them to gene-level features provides essentially no incremental predictive value. Across matched evaluations, pairwise scores do not improve AUROC, while ablation effects remain small and inconsistent.
- Causal ablation comparisons produce mostly small effects, including ∆=+0.0003, d=0.02 and ∆=+0.0021, d=0.08.Other tested conditions show decreases of ∆=−0.0025, d=−0.16 and ∆=−0.0030, d=−0.14.
- The expanded intervention design tests ranked, composite, full-layer, inverse, random, uniform-attention, and MLP-ablation conditions.
- AUROC=0.609± 0.121 for attention and AUROC=0.574± 0.169 for correlation edges exceed chance in the evaluated setting.
- ∆AUROC=−0.000 [−0.000, +0.000] shows that attention adds no matched predictive value beyond gene-only features.
- ∆AUROC=−0.000 [−0.001, +0.001] shows that correlation edges likewise add no matched predictive value beyond gene-only features.
- Attention adds only ∆AUPRC=+0.001 [+0.000, +0.002] when evaluated with gene features.
36. Intervention-Fidelity Diagnostics
The framework assessed intervention effects through hidden-state shifts, statistical testing, and cross-condition diagnostics. Its registry records 153 tests across 37 analyses, with 63 of 95 confirmatory tests significant after correction.
- 153 statistical tests covered 37 complementary analyses, including 95 confirmatory and 58 descriptive tests.
- 63 of 95 confirmatory tests were significant after Benjamini–Hochberg FDR correction.
- The framework used a framework-level α of 0.05 with Benjamini–Hochberg correction.
- Sample sizes varied from individual run-pairs to tens of thousands of cells and bootstrap resamples, depending on the analysis.
- The mediation analysis used six frozen cross-tissue runs spanning immune, kidney, and lung tissues, with 16 run-pairs across head and MLP granularities.
Detectability theory
The analyses show that attention contains layer-specific biological structure, but its detectability and predictive value depend strongly on how expression and pairwise information are controlled. Gene-level features remain stronger for perturbation prediction, while simple attention-derived edge scores add little or no incremental value.
- Protein–protein interaction signal peaks at L0, whereas transcriptional regulatory signal increases with depth and peaks at L15.STRING ≥700 reaches AUROC = 0.640 at L0, while TRRUST reaches AUROC = 0.750 at L15.
- 97% of the TRRUST attention–membership correlation remains after controlling for expression correlation.The partial correlation is r = 0.353 with p = 1.1 × 10^-11, compared with raw r = 0.363.
- Value-weighted cosine similarity underperforms raw attention and correlation because it collapses pairwise structure into similarity between blended context representations.Its edges alone achieve AUROC = 0.587, below raw attention at 0.707 and correlation at 0.637.