Source-linked AI summary
PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning
Hao Ye, Gaopeng Zhang
TL;DR
Structured pruning commonly evaluates surrogates on broad mask samples, leaving unclear whether the selected mask is reliable. PruneShift separates broad fidelity, selector-neighborhood fidelity, and decision quality, then combines theoretical conditions with four empirical studies. The results are heterogeneous: local or predictive improvements do not consistently establish superior selected-mask task performance.
Problem
Broad surrogate error and rank correlation do not directly establish whether the mask selected by pruning search is reliable.
Method
PruneShift separates broad prediction, selector-neighborhood prediction, and finite comparison regret, and derives coverage, margin, and finite-pool conditions.
Results
The four studies yield heterogeneous outcomes, including controlled coverage effects and improved local OSSCAR fidelity without consistent independent fixed-mask confirmation.
Takeaways & Limitations
Predictive fit, decision reliability, and pruning method quality require separate evidence.
Takeaways & Limitations
The sufficient conditions are operationally limited because finite-sample certification additionally requires independent evaluation, valid upper confidence bounds, and known or conservatively bounded parameters.
Abstract
from arXiv · showhide
Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. Most evaluations report average surrogate error or rank correlation on broadly sampled masks. These summaries do not directly test the mask chosen by the surrogate. We introduce PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision. We first prove that Spearman and Kendall agreement can approach one while normalized selection regret remains maximal. We then derive sufficient conditions based on uniform error, selector suboptimality, decision margin, density ratio, and comparison mass. The analysis also yields a finite pool certificate with an explicit excess cost bound. Four studies test different links in this argument. External TextbookQA confirmation is heterogeneous: 7 of 20 simultaneous intervals favor the surrogate-selected mask, 6 favor its fixed comparator, and 7 cross zero. On a fixed Natural Questions pool, strict improvement holds in one of four settings. A controlled QQP experiment supports the proposed coverage mechanism in all 16 prespecified endpoints, although the sufficient bounds are conservative. Finally, a restricted OSSCAR reconstruction study on OPT-125M shows better local than broad fidelity in 68 of 75 primary endpoints. Independent fixed-mask confirmation is inconclusive in 24 of 25 endpoints and favors the comparator in one. These results show why predictive fit, decision reliability, and pruning method quality require separate evidence.
1 Introduction
Structured pruning relies on surrogate objectives because exhaustive task evaluation is intractable, but broad surrogate validation may not reflect the masks selected by search. PruneShift separates predictive fidelity, selector-neighborhood fidelity, and decision quality, with theory and studies designed to test their links.
- Structured pruning preserves regular tensor structure but requires surrogate objectives because exhaustive mask search is combinatorial.Surrogates may use weights, activations, gradients, curvature, or reconstruction error under a pruning budget.
- Broad average error and rank correlation can remain strong even when errors are common among masks competing for selection.The selector searches for unusually favorable surrogate values, creating a mismatch between broad sampling and selection-relevant regions.
- PruneShift evaluates broad-mask accuracy, selector-neighborhood accuracy, and selected-decision performance against declared alternatives.It is an evaluation framework rather than a new pruning score.
- The framework separates estimator error from selector error and derives guarantees using coverage, comparison mass, uniform error, selector suboptimality, and decision margin.
- Two operational routes provide stronger evidence: a finite-pool certificate bounds excess cost within a fixed pool, while independent confirmation tests a prespecified comparison.The paper presents these routes as answering different questions rather than as substitutes.
- Four studies combine external transport, finite-pool selection, controlled coverage intervention, and OSSCAR reconstruction, producing both positive and null findings.The studies are used to assess which claims transfer across settings.
2 Related Work
Prior work develops pruning rules and decision-focused evaluation, but PruneShift focuses on the evidence needed to trust a selected structured-pruning mask. It makes the transfer from broad predictive fit to selection-relevant and decision-level evidence explicit.
- Structured-pruning research has advanced local sensitivity, curvature, learned masks, initialization signals, global objectives, reconstruction, and discrete optimization.
- Structured units interact across layers and groups, while greedy and local-exchange search can produce different masks under the same objective.The experiments therefore cross matched estimators and selectors.
- Decision-focused and offline optimization research evaluates predictions through induced decisions and shows that selection can exploit errors in weakly supported regions.PruneShift applies this perspective to structured pruning.
- PruneShift adds explicit separation between broad masks, selector neighborhoods, and prespecified comparisons, including tests of transfer between them.
3 Problem Formulation and Evaluation Domains
PruneShift formalizes structured pruning as surrogate optimization over feasible masks and distinguishes three evaluation domains. These domains require different evidence because averages under one distribution do not automatically transfer to selector-relevant or decision-level guarantees.
- 3.1 Structured pruning as surrogate optimization: A feasible structured mask z belongs to F, which fixes pruning cardinality and layer constraints.
- 3.1 Structured pruning as surrogate optimization: Direct evaluation of L(z) over every feasible mask is usually intractable, so a pruning method constructs a surrogate and approximately minimizes it.
- 3.1 Structured pruning as surrogate optimization: The estimator defines the surrogate, whereas the selector determines how its constrained minimum is approximated; these are separate components.
- 3.1 Structured pruning as surrogate optimization: Regret is global when A = F, but experiments use a finite prespecified comparison domain whose regret is estimable from held-out data.This regret remains deliberately scoped to the declared comparison domain.
- 3.2 Three non-interchangeable domains: Broad validation distribution QU samples generic feasible masks and supports mean error, rank correlation, tail recall, and pairwise ordering.
- 3.2 Three non-interchangeable domains: Selector-neighborhood distribution QN balances neighborhoods around fixed selector outputs to measure errors search is more likely to expose.
- 3.2 Three non-interchangeable domains: Decision comparison domain CS contains a selected mask and declared alternatives, so held-out task loss measures the actual decision in a known finite domain.For fixed-cardinality masks, neighborhoods use half Hamming distance and shells of radii 1, 2, and 3 with equal center-radius weighting.
- 3.2 Three non-interchangeable domains: Agreement under QU does not automatically transfer to QN or to a uniform error bound on CS; Proposition 4.1 shows rank agreement can approach one while normalized regret remains one.Transfer requires explicit coverage and comparison mass.
4 Theory: From Fidelity to a Decision Guarantee
The theory shows that broad fidelity and rank agreement do not by themselves guarantee a good pruning decision, then identifies coverage and error conditions that support decision guarantees. It also provides finite-pool certificates whose claims remain limited to fixed, declared comparison domains.
- Rank agreement can miss regret: For every M > 0 and N ≥3, surrogate minimization can incur regret M while Spearman correlation approaches one.The construction swaps only the two best decisions, leaving the remaining ranks unchanged.
- Rank agreement can miss regret: For fixed M, rank disagreement and RMSE vanish as N grows, whereas selected-decision regret remains M.Thus aggregate rank and score fidelity can become arbitrarily strong without improving the selected decision.
- Sufficient conditions: The theory transfers broad error to decision comparisons through a density ratio, uniform comparison error, selector suboptimality, and comparison-domain mass.The density ratio measures concentration where broad validation has little mass, while qmin measures how little evaluation weight a decisive candidate receives.
- Sufficient conditions: A δ-approximate surrogate minimizer can receive a decision guarantee when uniform error and the comparison-domain conditions hold.The bound distinguishes estimator error from selector optimization error and applies to the declared finite comparison domain.
- Sufficient conditions: When λ = 0, the comparison domain has no validation-distribution support and the density ratio is unbounded, so the sufficient condition fails by design.This limiting case is not a claim about every possible estimator.
- Finite-pool certificates: The finite-pool certificate applies only to a fixed declared pool and requires the candidate set, surrogate, selector rule, and certificate to be fixed independently of evaluation observations.Positive pool regret disproves global optimality, but zero pool regret neither establishes global optimality nor bounds the unobserved global gap.
5 Surrogates and Selectors Under Evaluation
The studies evaluate measured finite differences, gradient second moments, deterministic local selectors, and OSSCAR reconstruction as distinct surrogate–decision links. The designs separate estimator behavior from selector behavior and avoid treating local or reconstruction objectives as global optimality guarantees.
- Measured finite differences: Measured singleton and pair finite differences are exact on calibration data, while the quadratic surrogate approximates arbitrary multiunit masks.The approximation arises when the exact singleton and pair terms are extended to larger masks.
- Gradient second moments: The QQP surrogates use gate-gradient first-order, diagonal second-moment, and full quadratic information.The matrix F is treated as a raw empirical second moment, not as population Fisher information or an exact Hessian.
- Selectors: Forward greedy adds the feasible unit with the smallest increment, whereas the local selector searches feasible one-for-one exchanges from the greedy set.Both implementations enumerate exchanges deterministically and apply study-specific numerical tolerances.
- Selectors: The one-swap selector terminates finitely and ends with no feasible exchange improving the objective beyond the terminal tolerance.Every accepted exchange preserves cardinality and strictly decreases the objective over a finite feasible set.
- OSSCAR reconstruction: OSSCAR evaluates restricted supports after refitting retained weights with reference damping, but its returned support is not certified globally optimal.The reconstruction equations define the evaluation quantity, while the search remains algorithmic.
- Study scope: The experiments use these surrogates and selectors to test four distinct transfers from surrogate evidence to a pruning decision.The studies span measured finite differences, gradient second moments, deterministic local search, and a public reconstruction method.
6 Experimental Design
The experimental protocol fixes the decision pipeline before independent confirmation and assigns distinct units, pools, partitions, and multiplicity families to four studies. This design separates selection, surrogate evaluation, coverage intervention, and fixed-mask comparison.
- Common protocol: PruneShift fixes the feasible mask set, surrogate, selector, comparison set, and statistical analysis before confirmation.Confirmation data cannot change the selected decision, and model outputs are reduced to prespecified experimental units.
- TextbookQA: TextbookQA uses 1,503 extractive QA examples in 389 passage contexts disjoint from SQuAD construction data.The primary analysis excludes one prespecified near-duplicate context and uses 388 clusters.
- TextbookQA: TextbookQA compares five fixed method roles against one comparator selected without using TextbookQA outcomes.The 20 contrasts form one multiplicity family with context-cluster bootstrap intervals.
- Natural Questions: Natural Questions assigns disjoint passage contexts to selection and confirmation, each containing 1,024 contexts.Each setting uses a fixed pool of K = 128 candidates and J = 4 cumulative looks.
- Natural Questions: The selected Natural Questions role is reproduced independently and compared with one fixed comparator using finite-population intervals.All four settings are considered jointly for the familywise analysis.
- QQP: The QQP coverage intervention crosses six checkpoints, two unit families, and two budgets, producing 24 cells and 73,728 theory-calibration records.The arms vary comparison-set mass while keeping estimator and selector definitions unchanged.
- OSSCAR: The OPT-125M OSSCAR study uses six layers, two unit families, two removal rates, and nonoverlapping calibration, primary, and confirmation documents.The document is the inferential unit, and 64 broad supports are matched to 64 selector-neighborhood supports per cell.
7 Results
Results separate broad or local surrogate fidelity from selected-decision performance. External confirmation is heterogeneous, the finite-pool certificate is scoped, coverage effects are consistently supported, and OSSCAR’s local fidelity does not establish fixed-mask task superiority.
- TextbookQA: 7 of 20 TextbookQA intervals favor the method, 6 favor the comparator, and 7 cross zero.The pattern varies by architecture and unit family, so the study rejects a universal transport claim.
- Natural Questions: 0.1392 is the per-setting 95% Natural Questions excess-cost bound, while 0.1486 is the four-setting familywise bound.These values concern the selected candidate relative to the best member of the 128-candidate pool.
- Natural Questions: Only BERT FFN shows strict improvement in one of four Natural Questions settings.The fixed-pool certificate bounds excess cost within the candidate pool but does not compare its winner with an external pruning strategy.
- QQP: All 16 QQP confirmation intervals support the predicted coverage effect across architectures, unit families, maximum error, and selected regret.Both the uniform and coverage-transfer inequalities hold in 73,728 of 73,728 records.
- QQP: 18.1% of QQP records show exact surrogate selection, while the sufficient margin condition holds in none.The intervention supports the mechanism, but the sufficient bounds are too conservative to explain most exact selections.
- OSSCAR: 68 of 75 OPT-125M primary fidelity intervals are negative, indicating better local than broad reconstruction fidelity.Aggregate local-minus-broad effects are −0.0567 for mean mismatch, −0.0801 for the 90th percentile, and −0.1302 for the maximum.
- OSSCAR: 24 of 25 OSSCAR fixed-mask confirmation intervals cross zero, and one favors the comparator; none favors OSSCAR.The aggregate OSSCAR-minus-comparator NLL is 0.000774 with interval [−0.000300, 0.001848].
- Interpretation: Local reconstruction fidelity and fixed-mask NLL answer different questions, so better local fidelity does not prove lower task loss for the selected support.The cross-study directional summary keeps endpoint families separate rather than treating metrics as commensurate effect sizes.
8 Discussion and Limitations
PruneShift’s evidence is deliberately scoped: coverage, finite-pool guarantees, reconstruction fidelity, and independent task comparisons answer different questions. The studies therefore produce heterogeneous, sometimes null findings that distinguish estimator, coverage, selector, and confirmation issues.
- Coverage determines whether broad evaluation reaches the thin region explored by a selector, while the QQP intervention supports this mechanism despite conservative bounds.The intervention changes comparison coverage while holding the main estimator and selector definitions fixed.
- A finite-pool certificate guarantees performance only within its declared pool and can coexist with an inconclusive comparison against an external strategy.Finite-set regret is a lower bound on global regret and need not approximate the global gap.
- OSSCAR reconstruction improves local fidelity after retained weights are refitted, but independent task superiority remains unresolved.This separates proxy fidelity from task superiority rather than implicating the reconstruction objective itself.
- The framework’s null results narrow claims: TextbookQA rejects a simple universal transport story, Natural Questions limits the finite-pool claim, and OSSCAR does not establish task superiority.These distinctions identify whether future work needs a better estimator, coverage, selector, or direct confirmation design.
- The conclusions are scoped to specific models, tasks, datasets, and pruning settings, so larger models, other tasks, and other methods may show different transfer patterns.The evidence uses BERT and RoBERTa, an answerable text-only TextbookQA subset, and six layers of OPT-125M.
- The intervals and checkpoints have limited interpretations because they are not distribution-free population intervals or random samples of training seeds.TextbookQA intervals describe stability across observed context clusters, while QQP checkpoints are fixed design instances.
- The theoretical guarantees are sufficient and potentially conservative, as the QQP study confirms both their validity and looseness.Improved data-dependent bounds must preserve independence and account for selector adaptation.
- All experiments use logical masks and do not measure exported latency, throughput, energy, recovery training, or hardware speedup.The work evaluates pruning decisions rather than deployment efficiency.
9 Conclusion
PruneShift argues that structured-pruning surrogate evaluation must reach the selected decision, not stop at broad predictive fit. Across four studies, transfer is heterogeneous, coverage effects are supported but bounds are loose, and local proxy gains do not establish task superiority.
- PruneShift separates broad predictive fidelity, selector-neighborhood fidelity, and fixed comparison quality for structured-pruning decisions.It also separates estimator error from selector error.
- Aggregate rank agreement cannot certify an argmin without additional assumptions, so the framework makes coverage, margin, finite-comparison, and finite-pool conditions explicit.
- The four studies show heterogeneous external transport, valid but non-universal finite-pool guarantees, loose coverage bounds, and improved local fidelity without fixed-mask task superiority.Together, these findings provide a clearer standard for evaluating decisions made by structured-pruning surrogates.