Source-linked AI summary
Efficient Auto-Interpretability of AI Models in Biology
Piotr Jedryszek, Oliver M. Crook
TL;DR
Biological model latents need separate tests for coherence, describability, and predictive validity, but evaluating thousands of latents is costly. The paper combines stability prioritisation, blind intruder detection, and falsifiable description checks; on the Boltz-1 Pairformer trunk, stability reduced evaluations and measured cost while recovering over half of interpretable latents. The approach also revealed that stability preferentially surfaces structure-related features over functional and compositional ones.
Problem
AI models can make powerful biological predictions, but latent usefulness requires separately establishing coherence, describability, and predictive power.
Method
The pipeline ranks latents by cross-seed stability, tests activating-example coherence with intruder detection, and independently turns descriptions into falsifiable annotation-based predictions.
Results
4.4× fewer latent evaluations and 5.2× lower measured cost per interpretable latent were achieved with just over half recall in Boltz-1 deployment.
Takeaways & Limitations
Stability prioritisation can reduce auto-interpretability costs while surfacing motifs enriched for claimed annotations, but it disproportionately favors structure-related features.
Takeaways & Limitations
The method’s current evidence suggests stability filtering loses many functional-site and compositional-region findings, and subspace results for complex motifs come from a small sample.
Abstract
from arXiv · showhide
Sparse autoencoders (SAEs), and other interpretability methods could turn AI models in Biology and other fields into engines of scientific discovery by explaining the superhuman capabilities of those models. However, a latent is only useful if we know three things: whether it is coherent, whether it can be described, and whether that description has predictive power. These questions are routinely conflated. We assemble them into a single pipeline and report the practical innovations each stage required. First, cross-seed dictionary stability prioritises which latents are worth spending resources to investigate. Second, an intruder-detection task asks whether a latents activating examples share a recognizable pattern. Third, a separate pass proposes a candidate biological description which we convert into falsifiable predictions which can be tested in silico. Deployed on the Boltz-1 Pairformer trunk, stability prioritisation finds interpretable latents using about 4.4 times fewer latent evaluations each, and at 5.2 times lower measured cost, while recovering over half of them, and the external check shows the surfaced motifs are significantly enriched for their claimed annotations. The results also suggest a possible tension: the cross- seed stability might be selecting for some types of features, like structure-related ones, much more than others, such as function-related features.
1 Introduction
AI biology models can make highly accurate predictions, but their internal representations remain difficult to understand. The paper frames interpretability as separate questions of latent coherence, description quality, and whether descriptions yield testable biological predictions.
- Motivation: SAEs decompose neural activations into simpler features that can be explored through examples where each feature is active or inactive.Their self-supervised setup may help discover concepts not specified by predefined labels.
- Problem: Three problems complicate biological auto-interpretability: evaluation cost, conflation of coherence with explanation quality, and weak validation of descriptions.The paper addresses these as distinct stages rather than treating a plausible explanation as sufficient evidence.
- Pipeline: Cross-seed stability ranks latents before LLM evaluation, intruder detection tests coherence, and a separate describe pass converts labels into falsifiable predictions.The pipeline separates prioritisation, pattern recognition, and biological claim validation.
- Contribution: 4.4× fewer latent evaluations and 5.2× lower measured cost per interpretable latent were achieved with just over half recall in Boltz-1 deployment.Stability prioritisation was presented as the pipeline’s main contribution.
- Scope: The stability filter predominantly retains residue-in-structural-position detectors while discarding many functional-site and compositional-region findings.This indicates that stability-based prioritisation may favor some feature types over others.
2 Methods
The methods separate stability-based prioritisation, calibrated intruder evaluation, independent description generation, and annotation-based checks. They apply these procedures to SAE latents from the Boltz-1 Pairformer trunk while also measuring subspace-level stability.
- Data and models: SAEs with TopK k = 256, dictionary size 2048, and input dimension 384 were trained on Boltz-1 Pairformer trunk activations from layer 20 at recycle 1.The trunk was targeted because concurrent work indicated sequence-level biochemistry was linearly accessible there.
- Intruder detection: An LLM judge identifies one intruder among four genuine activating windows, testing whether a latent’s examples share a recognisable pattern.The calibrated judge uses Claude Sonnet 4.6 at temperature 0.
- Intruder detection: 16 to 512 output tokens increased positive-control accuracy from 0.17 to 0.97, while prompt phrasing produced accuracies from 0.85 to 0.97.The restricted-budget version failed as a reliable harness on ground-truth controls.
- Description generation: An independent LLM describe pass proposes one-sentence biological labels without exposing those descriptions to the intruder judge.Descriptions attach hypotheses only after the blind coherence test.
- Validation: Ground-truth checks map each functional description to SwissProt residue categories and report fold-enrichment with by-protein bootstrap 95% confidence intervals.The residue-class conditional background avoids over-crediting detectors that already fire on the relevant residue.
- Stability: Cross-seed stability matches dictionary directions across independently trained seeds and averages matched absolute decoder cosine values.The continuous stability score ranges from 0 to 1 and is thresholded at 0.5; a separate shared flag uses |cos θ| > 0.7 in both held-out seeds.
- Stability: Subspace stability compares spans of up to eight co-activating latent directions across seeds using the mean cos2 of their principal angles.This addresses concepts that may be represented by rotated multi-dimensional bases rather than single matched axes.
3 Results
Cross-seed stability predicts which Boltz-1 Pairformer latents are interpretable, reducing evaluation cost while recovering over half of interpretable latents. However, it favors some feature types more than others, while separately generated descriptions generally validate residue claims but less reliably identify contextual restrictions.
- Cross-seed stability predicts interpretability: 4.4× fewer latent evaluations per interpretable latent and 5.2× lower measured cost were achieved by filtering on cross-seed stability, with just over half recall.In deployment, stable-latent filtering recovered 53% of interpretable latents.
- Cross-seed stability predicts interpretability: 1.97× was the supervised interpretable-rate lift from applying stability ≥0.5 in the discovery band.The analysis covered 4,881 latents in activation-frequency percentiles 20 to 60.
- Cross-seed stability predicts interpretability: Stability correlated with concept-recovery F1 overall, but within-concept correlations were only 0.09 to 0.19, indicating stronger separation of concept types than ranking within a type.The overall correlation was ρ = 0.51, while amino-acid identity was more reproducible than secondary structure or SwissProt function.
- Cross-seed stability predicts interpretability: 53% precision and 53% recall resulted when nine of 17 stable latents passed the intruder-accuracy interpretability bar.The stable arm required judging only 12% of the screened panel and produced a 4.4× lift over the panel base rate.
- Auto-interp descriptions as falsifiable hypotheses: All 17 surfaced descriptions validated their claimed residue classes, whereas named structural or functional contexts validated for 12 of 16 testable cases.Residue precision was at least 0.98 in 13 of 17 cases but recall was mostly low, consistent with context-restricted detectors.
- Complex motifs and cross-seed stability: Five metal-coordination or catalytic motifs fell below the stability threshold, and subspace stability ranked them below even uninterpretable latents.The authors caution against over-interpreting this result because the motif sample is small.
4 Discussion
The discussion presents stability filtering as a substantial efficiency gain while warning that it disproportionately excludes functional and compositional motifs. It also argues for falsifiable ground-truth checks, which currently validate annotated biology but not genuinely novel claims.
- 4 Discussion: Stability filtering keeps eight of nine structure-related finds but loses four of five functional-site finds and all three compositional-region finds.The authors therefore recommend ranking latents by stability rather than excluding unstable ones outright.
- 4 Discussion: 4.4× fewer evaluations per interpretable latent comes with a disproportionate loss of metal-coordination, catalytic, and compositional features.The paper treats this class-specific pattern as a possible tension requiring larger-scale causal intervention.
- 4 Discussion: Ground-truth checks separate trustworthy residue claims from unsupported context claims in auto-interpretability descriptions.Across the panel, residue claims held, while four of 16 testable context claims did not.
- 4 Discussion: The check is not automated, covers only already-annotated biology, and may require wet-lab experiments for genuinely novel claims.The authors envision a digital lab-in-the-loop to draft and score predictions before human inspection.
Responsible Use Statement
The study analyzes public Boltz-1 representations and SwissProt annotations without generating novel biological designs. It reports low direct misuse risk for the characterized motifs but notes that cheaper auto-interpretability could aid harmful latent discovery elsewhere.
- Responsible Use Statement: The study uses public Boltz-1 representations and SwissProt annotations rather than generating novel sequences, structures, or designs.The identified motifs are described as well characterized.
- Responsible Use Statement: Cheaper auto-interpretability could support harmful latent discovery in other models.The authors screened discovered descriptions for pathogenicity-related concepts and found none.
A Supplementary Methods
The supplementary methods describe how protein exemplars are extracted, filtered, and evaluated for interpretable latent discovery. The sweep uses activation-based screening and intruder detection across sampled latents and activation deciles.
- A Supplementary Methods: Exemplars use a dual-pass script that finds global maxima and localized protein peaks, then extracts 33-residue windows around peaks.Windows are partitioned into 10 activation deciles, with up to eight windows per decile.
- A Supplementary Methods: The supervised benchmark scores 15 concepts comprising three AlphaFold-DSSP secondary-structure classes and 12 SwissProt categories.Amino-acid identities are excluded from this benchmark passage.
- A Supplementary Methods: Latents with cumulative activation density above 0.85 are discarded, and sweep eligibility requires firing in at least 10 proteins.This screen is separate from the supervised proxy’s alive criterion of encoder-decoder cosine above 0.1.
- A Supplementary Methods: The discovery sweep samples 150 deduplicated layer-20, recycle-1 latents across three equal-count activation-frequency bins.After screening, 141 latents remain; each receives 24 to 54 intruder trials, with median 42.
A.5 Cost accounting
Cost accounting measures discovery-sweep requests, tokens, and dollar costs directly from the batches API. Stability filtering reduces evaluations and measured cost, but its savings depend on dictionary availability, threshold choice, and uncertain sample-level estimates.
- A.5 Cost accounting: 1.9 evaluations per interpretable latent versus 8.3 unfiltered yields a 4.4× saving.Request usage shows the same ratio because stable and other latents receive similar trial counts.
- A.5 Cost accounting: 5.2× lower measured cost corresponds to $0.33 per find versus $1.69.Stable latents also average 351 output tokens per trial versus 473 elsewhere, making money savings exceed the request-count savings.
- A.5 Cost accounting: All token counts and dollar figures are measured per request from the Message Batches API rather than estimated from prompt lengths.An earlier estimate understated output by 2.2× because it was calibrated on high-interpretability latents.
- A.5 Cost accounting: Table 1 treats cost per find as the efficiency metric, not total-spend ratio, because filtering returns fewer finds.The describe pass was not rerun because descriptions are independent of intruder construction.
- A.5 Cost accounting: A bootstrap gives a 95% CI of 2.2 to 8.5 for request-cost saving, while stability itself requires independently trained dictionaries.At the evaluated scale, judging 141 latents costs $28.76, and dictionary retraining does not amortise.
- A.5 Cost accounting: A stability cut of 0.60 gives precision 0.64 and 1.6 evaluations per interpretable latent at recall 0.41.The paper reports that precision and recall trade off smoothly across the swept thresholds.
A.7 How an intruder trial is constructed
Each intruder trial pairs four genuine activation windows with one biologically mismatched window while neutralising highlight-pattern cues. A contamination check shows occasional target firing in intruder windows but negligible impact on accuracy.
- Each trial uses five 33-residue windows: four from one latent’s activation decile and one intruder centred on another latent’s activation peak.Genuine windows come from one protein, while the intruder comes from a protein contributing none of the target’s selected windows.
- Matched highlight masks prevent an adversary from identifying the intruder using mark count, run length, or positional offset.The mask transplantation makes these statistics match between genuine and intruder windows by construction.
- 0.204 adversarial accuracy stays near the 0.200 chance level when highlight patterns are fully matched.Randomising highlighted positions instead creates a centre-marking cue and raises adversarial accuracy to 0.804.
- Target contamination occurs in 5.3% of trials, concentrates in denser latents, and is uncorrelated with per-latent judge accuracy.Removing contaminated trials changes the panel mean only from 0.343 to 0.338.
- Contaminated trials are slightly easier for judges, at 0.433 versus 0.338, but the authors report no confident mechanism.The difference is treated as a residual oddity rather than a causal explanation.
B Harness Calibration Ladder and Robustness
Harness calibration shows that the judge’s output budget, rather than prompt wording, determines whether intruder evaluation works. Expanding the budget restores positive-control accuracy and yields graded control-panel performance.
- 0.17 accuracy at 16 tokens rises to 0.97 at 512 tokens on 20 amino-acid positive controls.The restricted generation setting is therefore unsuitable for this control task, whereas the expanded budget recovers all 20 controls.
- 0.29 random-latent accuracy, 0.48 contextual-concept accuracy, and 0.17 scrambled-control accuracy form a graded control continuum.The calibrated results distinguish contextual concepts from random and scrambled controls rather than producing a flat near-chance panel.
- The calibrated pass-rate test uses chance = 0.20 and a one-sided binomial criterion at α = 0.05, weaker than the discovery bar of accuracy ≥0.5.The scrambled control is unaffected by the intruder construction because its windows come from different features.
- 0.85–0.97 accuracy across four prompt phrasings shows wording has less effect than the token allowance.The three calibrated production variants differ by less than 0.01, while the independent paraphrase reaches 0.85.
- 0.96–0.97 positive-control accuracy is reproduced across Claude Sonnet 4.6, Haiku 4.5, and Opus 4.8 under the expanded budget.The broken 16-token gate is model-dependent, collapsing to chance on Sonnet and Haiku but reaching 0.86 on Opus.
C Per-Feature Ground-Truth Validation
Per-feature validation scores both residue-class and contextual claims in each description against external annotation categories. Residue claims validate broadly, while contextual claims validate less consistently and face annotation and power limits.
- Table 3 separately scores each description’s named residue class and its additional structural or functional context.This preserves distinct evaluation targets rather than treating a description as a single indivisible claim.
- Four context failures include two underpowered cases based on only four and seven firing residues.Features 534 and 1237 are described as underpowered rather than clearly refuted.
- Five description clauses across four features are unscored because their categories are absent from the available ground truth.Transmembrane environment is the most common unavailable category.
- Secondary-structure categories come from DSSP on Boltz-predicted structures and are therefore not fully model-external.SwissProt categories and residue identities provide the more external annotation sources referenced in the passage.
D Subspace Stability: Full Method and Positive-Control Ladder
Subspace stability extends cross-seed comparison to groups of co-activating latents, addressing concepts whose bases may rotate across seeds. Positive controls establish the metric’s range, while Figure 6 shows the tested motifs near the low-stability floor.
- D Subspace Stability: Full Method and Positive-Control Ladder: For each seed, the method spans up to eight latents whose per-residue firing correlates most strongly with the target latent.Subspace stability is computed from the mean cos2 of principal angles between these spans across seeds.
- D Subspace Stability: Full Method and Positive-Control Ladder: Seed-1 auto-interpretation is transferred to seeds 2 and 3 using firing-signature matches rather than repeated interpretation.The subspace analyses use the same 196,175-residue SwissProt subset across all three seeds.
- D Subspace Stability: Full Method and Positive-Control Ladder: 0.85 decoder-cosine overlap is recovered for the 20 amino-acid stable-atom controls under the positive-control ladder.The metric is additionally tested for rotation invariance and against synthetic rotated bases.
- D Subspace Stability: Full Method and Positive-Control Ladder: 0.13 subspace stability for five metal-coordination and catalytic motifs falls below uninterpretable latents at 0.23 and amino-acid controls at 0.36.All groups are processed through the identical pipeline.
- D Subspace Stability: Full Method and Positive-Control Ladder: 1.00 for a planted rotated subspace and 0.85 for a genuinely stable atom group demonstrate that the metric has usable range.The tested motifs therefore sit near the metric’s floor rather than near either positive-control level.