Source-linked AI summary
Recovery Rates Are Not Comparable Across Transcription Factors: Chance Correction for Attribution Evaluation
Hyunkyung Han, Min Jung Kim
TL;DR
Genomic attribution scores are often reported without factor-specific chance baselines, making raw motif recovery and perturbation curves difficult to interpret or compare. The paper derives chance levels and corrected scores, adds positional and pre-attribution validity screens, and shows that correction changes published classifications while some perturbation curves fail their own precondition.
Problem
Raw motif-recovery and perturbation scores lack routinely reported chance baselines, even though their chance values vary with motif geometry and evaluation conditions.
Method
The paper derives closed-form and positional nulls, applies a chance-corrected overlap score, and proposes composition and fully masked-input screens before attribution.
Results
Across 268 factors, uniform chance levels range from 0.0118 to 0.0427, correction changes a published five-factor classification, and one perturbation curve is non-monotone with a fully masked score above threshold.
Takeaways & Limitations
Raw recovery rates should be compared against factor-specific chance levels, and perturbation-based faithfulness should be screened for valid baseline behavior before attribution is interpreted.
Abstract
from arXiv · showhide
Attribution methods for genomic sequence models are commonly evaluated by how much of a known motif they recover, or by how a prediction degrades as evidence is deleted. Neither score is interpretable without the value it would take by chance, and neither is routinely reported against one. We show that this omission is not a matter of precision but of validity. The uniform chance level for contiguous motif overlap is \(L/(N-L+1)\); across 268 transcription factors in UniBind it ranges from 0.0118 to 0.0427, a 3.6-fold spread determined by motif length and window size alone. For two factors the bootstrap intervals of the chance levels themselves do not overlap, so their raw recovery rates are not comparable quantities. Correcting for this dissolves a published three-way classification of five factors: a factor reported as a resolution failure attains the second-highest corrected value, ahead of one of the two positive controls, and two reported as complete failures fall at or below chance. We further show that perturbation-based evaluation can fail its own precondition: for one factor a fully masked input still scores above the decision boundary, and the curve is not monotone in the number of masked positions, so the area under it is not a measure of faithfulness. We provide chance levels in closed form, a chance-corrected score, and two screens that run before any attribution is computed.
A. 1. Introduction
Raw attribution recovery and perturbation scores are not interpretable or comparable without factor-specific chance baselines and valid null behavior. The paper introduces closed-form and pre-attribution checks to address these problems.
- Motivation: A raw motif-overlap score is not meaningful without knowing the overlap an uninformative attribution would achieve, which varies with motif length and window width.The same comparability issue applies to inter-method agreement scores when their chance midpoint is not zero or constant.
- Motivation: The uniform chance level spans a 3.6-fold range across 268 transcription factors, driven entirely by motif geometry rather than model information.For CTCF and MAX, bootstrap intervals of the chance levels do not overlap, so their raw recovery rates are not comparable.
- Motivation: Perturbation curves require a fully evidence-deleted prediction at or below the decision boundary; this precondition can be checked with one forward pass.If the precondition fails, the curve is bounded below by a quantity unrelated to the attribution.
- Contribution: The paper frames chance correction as an application of an established overlap-statistic correction, while emphasizing that its contribution is demonstrating changed conclusions in attribution evaluation.The approach subtracts the expected overlap under independence and rescales the result.
- Method: For a motif of length L placed uniformly in a window of length N, the contiguous-overlap chance level is L/(N-L+1).The paper verifies this closed-form expectation against simulation within 2.6% for L = 6...25 at N = 512 over 200,000 draws.
2) 2.2 Positional null:
The positional null accounts for motif-location bias that the uniform null can miss, while chance correction puts recovery and agreement measures on interpretable scales. This matters because below-chance adjusted values are treated as uninformative rather than anti-correlated.
- 2) 2.2 Positional null:: When motifs concentrate near the window centre, the uniform null is too loose, so the positional null simulates chance from the empirical motif-start distribution.For CTCF, the positional null is 0.0697 versus 0.0385 under the uniform null.
- 2) 2.2 Positional null:: Position alone produces a 1.59–2.02-fold enrichment over the uniform null across the five factors, showing that location can materially change the baseline.The CTCF comparison is 0.0697 against 0.0385.
- 4) 2.4 Adjusted score:: The chance value for inter-method Jaccard agreement is 0.020 for k = 20 in a 512-position window when both selected sets have fixed cardinality.The fixed cardinality prevents the size-driven bias described for variable-size sets.
- 4) 2.4 Adjusted score:: The corrected score is zero at chance, one at perfect recovery, and negative below chance, using the adjusted-Rand-index correction form.The correction subtracts the chance expectation and rescales the overlap statistic.
- 4) 2.4 Adjusted score:: Negative adjusted values near −0.042 and −0.041 are treated as uninformative because they are indistinguishable from attribution that ignores the sequence.They are not interpreted as evidence that the attribution avoids the motif.
1) 3.1 Composition screen:
The composition screen tests whether a model’s task requires sequence order before attribution is interpreted. If shuffled positives perform at least as well as originals, motif use is not required by the task.
- 1) 3.1 Composition screen:: If dinucleotide-preserving shuffled positives score at or above the originals, the task does not require sequence order and attribution cannot report motif use the model did not make.The screen compares classifier performance on original and shuffled positive sequences using the same positive and negative sets.
- 1) 3.1 Composition screen:: The evaluation uses UniBind data with 268 transcription factors, 512 bp positive windows centered on binding sites, and GC-matched genomic negatives.The detailed benchmark includes CTCF, GATA1, MAX, SP1, and TBP, with factor-specific fine-tuned models and integrated gradients.
- 1) 3.1 Composition screen:: GC matching uses positive-set histograms in bins of width 0.02, and six factors are excluded before matching because their positive sets are too small for the screen.The remaining factors are matched by drawing candidate windows until the histogram is filled.
E. 5. Results
Chance correction changes the interpretation of motif-recovery results: geometry-driven baselines vary substantially, and the published factor classes collapse into recovered-above-chance versus uninformative cases.
- 5.1 Chance levels are not comparable across factors:: 1.93-fold: CTCF’s chance level exceeds MAX’s, showing that raw recovery rates compare factors against different geometry-dependent baselines.SP1 and TBP share a chance level exactly because they share motif length.
- 5.2 A published classification does not survive correction:: GATA1 becomes the second-best recovered factor, while SP1 and TBP are at or below chance, replacing three published classes with two outcome groups.The same re-partition appears under uniform and centre-biased corrections, with GATA1 indistinguishable from the positive controls.
- 5.2 A published classification does not survive correction:: 0.094 and 0.095: SP1 and TBP have the smallest full-model versus best-k-mer AUC gaps, compared with 0.319 for CTCF, and their adjusted scores are at or below zero.When composition nearly solves the task, attribution is not necessarily failing because little sequence-specific signal remains to find.
- 5.2 A published classification does not survive correction:: 0.983 versus 0.745 for TBP and 0.946 versus 0.917 for SP1: uniformly random inputs can score above true positives, while accuracy remains high and nondiscriminative.These factors are also the ones whose adjusted recovery is at or below zero.
4) 5.4 A perturbation curve can fail its own precondition:
Perturbation curves require a fully masked prediction at or below the decision boundary, but CTCF violates this precondition and produces a non-faithful, non-monotone curve.
- 5.4 A perturbation curve can fail its own precondition:: 0.596 is CTCF’s fully masked prediction, above the decision boundary, so its deletion curve converges to a baseline unrelated to attribution.The curve therefore retains only 0.165 of AUC dynamic range.
- 5.4 A perturbation curve can fail its own precondition:: A perturbation curve is interpretable only when the fully masked score is at or below the decision boundary; across five factors, that score ranges from 0.032 to 0.596.Thus MAX passes the screen, whereas CTCF does not.
- 5.4 A perturbation curve can fail its own precondition:: CTCF’s deletion curve is non-monotone: it falls at 32 masked positions, rises to 0.767 at 128, then declines to 0.596.Over most of its extent, the curve exceeds the unmasked prediction, indicating that masking substitutes a positive signal rather than simply removing evidence.
- 5.4 A perturbation curve can fail its own precondition:: Random deletion initially lowers CTCF’s prediction more than attribution-ordered deletion, with scores of 0.401 and 0.555, respectively.This agrees with the negative deletion effect reported for CTCF.
5) 5.5 Deletion curves are not comparable across tokenizers:
Deletion curves are not comparable across tokenizers when masking changes input segmentation, because the resulting area under the curve measures different quantities.
- 5.5 Deletion curves are not comparable across tokenizers:: DNABERT-2 token counts expand from 111.5 unmasked tokens to 142.3, 186.2, 253.1, and 445.9 after masking 20, 50, 100, and 300 positions.At full masking, DNABERT-2 reaches 514.0 tokens, whereas HyenaDNA remains at 513 throughout.
- 5.5 Deletion curves are not comparable across tokenizers:: A 1.28- to 1.67-fold token expansion already occurs across the 20–50 masked positions used to compare attribution methods.This means the reported separation range is affected by segmentation changes.
- 5.5 Deletion curves are not comparable across tokenizers:: Deletion AUC measures evidence removal for HyenaDNA but evidence removal confounded with input segmentation for DNABERT-2.Because the confound is independent of masking order, it compresses attribution-method differences and is consistent with near-ceiling SP1 and TBP values.
6) 5.6 Faithfulness measured against the right baseline:
CTCF is the factor whose masking baseline fails and the only factor whose deletion effect is indistinguishable from zero, consistent with having no usable curve range.
- 5.6 Faithfulness measured against the right baseline:: CTCF’s deletion effect is indistinguishable from zero, consistent with its fully masked baseline failing the required condition and leaving no range for the curve to move.The masking baseline and deletion effect therefore point to the same interpretability failure.
7) 5.7 Two hypotheses we rejected:
The authors reject training-data bias and negative-set compositional confounding as explanations for the observed recovery pattern, while limiting the uniform null to expectation-based chance correction.
- 5.7 Two hypotheses we rejected:: Held-out and matched evaluation sets differ by at most 0.023 across factors, rejecting training-data bias as the explanation for recovery.The test compares recovery when binding sites may or may not have appeared in training.
- 5.7 Two hypotheses we rejected:: A GC-only classifier achieves AUC 0.522–0.569, so GC-matched negatives are not separable by GC composition alone.This rejects compositional confounding of the negative set under the reported matching procedure.
- 5.7 Two hypotheses we rejected:: The contribution is demonstrating that chance correction re-partitions published conclusions rather than merely refining them.The paper distinguishes this application from the novelty of the individual correction ingredients.
- 5.7 Two hypotheses we rejected:: The uniform null supplies an effect-size origin through its expectation, not a variance for significance testing or confidence intervals.The authors use dinucleotide-preserving shuffled controls for those inferential judgments.
- 5.7 Two hypotheses we rejected:: Motif clustering can increase sequence-to-sequence variation but does not alter the origin used by the adjusted score.Accordingly, the uniform null is used for correction rather than as a hypothesis test.
3) 6.3 Where the failure is located:
The paper locates attribution-evaluation failures upstream of model architecture: chance correction and null-model choice determine whether motif-recovery conclusions are interpretable.
- Compositionally separable tasks give attribution methods nothing to find, independently of whether architecture aligns filters with whole motifs.
- Chance correction follows established overlap-normalization ideas, while genomic null models show that constrained randomization can change conclusions.
- The audit already required random-baseline comparisons, yet its motif-overlap scores reported neither chance levels nor shuffled-sequence controls.
- The paper presents its analysis as a correction that makes the audit’s own recommended guidelines visible in a demonstration case.
6) 6.6 Reporting checklist:
The reporting checklist requires four preconditions or controls before interpreting attribution scores: masked-input behavior, chance correction, composition screening, and perturbation-unit disclosure.
- Report the fully masked prediction before showing any perturbation curve, because an above-boundary score makes the curve’s baseline attribution-independent.
- For every overlap metric, report both its chance level and adjusted value, using a closed-form correction where applicable.
- Run a composition screen to determine whether classification remains separable without sequence order; the paper does not propose constructing a non-separable task.
- Disclose the perturbation unit when comparing models with different tokenizers, while recognizing that in-distribution perturbation design remains outside this paper’s scope.
- Interpret single-run comparisons cautiously: adjusted recovery spreads across runs, and three runs are insufficient to characterize that variation.
- Treat SP1’s per-sequence results cautiously because its held-out set contains only 65 annotated positives.
9) 6.9 What follows:
The paper concludes that attribution recovery is partly determined by task design, so benchmark validity requires controlled negatives and explicit chance and masking checks.
- Without negatives controlled so motif use is necessary, recovery-rate benchmarks cannot be said to measure the attribution.
- The paper supplies code and implementation details, including N-token masking and a GC histogram rule in 0.02 bins, with six factors failing to fill.
- The conclusion’s re-partition is unaffected by seed variation, but smaller comparisons are unsafe from a single run because recovery is less stable than validation AUC.