Source-linked AI summary

Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

Huiyuan Tian, Bonan Xu, Shijian Li

arXiv:2607.28308v1cs.LG

TL;DR

The paper asks whether multi-expert benefit requires experts to cover disjoint representation directions, a question relevant to explaining quality and guiding pruning. Using geometric diagnostics, factorial analyses, and functional interventions, it finds coherent overlap: later experts remain useful within a shared token-relevant neighborhood despite negative contextual narrowing and non-disjoint linear coverage.

  • Problem

    The paper examines whether multi-expert quality requires selected experts to contribute disjoint representation directions, an intuition motivating explanations, pruning, and compression.

  • Method

    The paper combines ESSI, matched-route residuals, a prefix-controlled 2 × 2 factorial, frozen-route interventions, and controlled Top-k training comparisons.

  • Results

    Across six MoEs, experts overlap while routes remain coherent; across 39 cells, selected candidates outperform rivals, but actual prefixes narrow their advantage and later experts remain useful.

  • Takeaways & Limitations

    Multi-expert benefit can arise from distinct computations within a shared token-relevant neighborhood without disjoint linear coverage.

  • Takeaways & Limitations

    The conclusions are limited to a linear, rank-128 router-input metric, three architectures for factorial analyses, and one small-scale matched-compute setting.

Abstract

from arXiv · show

Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled $2\times2$ factorial; frozen-route interventions and a controlled Top-$k$ study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.

1. Introduction

The introduction argues that MoE routing benefits cannot be explained by geometric complementarity alone because route coherence, candidate quality, and candidate-by-context interaction are distinct. It presents ESSI and a prefix-controlled 2 × 2 factorial to separate these quantities and reports overlapping expert subspaces alongside coherent, functionally useful routing.

  • MoE models scale capacity by activating only a small parameter fraction per token through input-dependent routing and specialized local computations.
  • Route coherence, candidate quality, and positive geometric complementarity are distinct properties rather than interchangeable evidence for multi-expert benefit.A coherent route may contain individually strong candidates, while a selected candidate may add little after preceding experts have covered similar directions.
  • ESSI and a prefix-controlled 2 × 2 factorial separately test expert geometry, candidate quality, contextual opportunity, and candidate-by-context interaction.The factorial crosses the selected candidate and strongest unselected rival with actual and matched alternative prefixes.
  • Across six architectures, expert subspaces overlap while actual routes fit tokens better than matched alternatives.The introduction frames this as a cross-architecture empirical pattern rather than evidence that experts occupy disjoint representational directions.
  • Selected candidates explain more residual representation than rivals, but actual prefixes narrow this advantage rather than increasing it.This result directly tests whether a candidate’s advantage depends positively on its co-selection context.

2. Method

The method separates expert-subspace geometry, route-conditioned candidate novelty, and functional contribution using ESSI, matched alternatives, a prefix-controlled 2×2 factorial, and frozen-route interventions. It evaluates these quantities across six MoE architectures and 39 factorial cells.

  • Subspace geometry: The route-coherence analysis compares actual routes with matched alternative contexts using residual ratios and token-representation fit.The geometry survey covers 18 model–layer cells across six open MoEs.
  • Subspace geometry: ESSI compares inter-expert separation with local within-expert dispersion using global and local rank-128 subspace bases.The index uses unweighted eligible expert pairs in its numerator and routing-load weighting in its denominator.
  • Prefix-controlled factorial: The candidate’s fractional novelty measures how much residual router-input energy an added expert explains after a route prefix, using separate uncentered rank-128 bases.Rows with zero input energy are excluded, and q = 0.2 means the candidate explains 20% of what remained after the prefix.
  • Prefix-controlled factorial: The factorial crosses the selected candidate and strongest unselected rival with actual and legal alternative contexts to estimate candidate quality, context effects, and interaction.The selected candidate is the actual expert at route position j; the rival is the highest-scoring eligible expert outside the complete actual Top-k route.
  • Evaluation design: The study analyzes 24 OLMoE/Mixtral factorial cells and 15 DeepSeek shared-expert cells, with frozen-route NLL interventions covering all 39 cells.Each cell uses up to 2,048 held-out tokens per layer and 1,000 paired source-context bootstrap resamples for 95% intervals.

3. Geometric Overlap and Route Coherence

Across six MoE architectures, expert subspaces overlap substantially rather than forming hard directional partitions, yet actual routes fit token representations better than matched alternatives. This pattern supports structured overlap with route coherence, while leaving candidate quality and interaction effects for Section 4.

  • Interpretation: Geometric route coherence alone cannot distinguish positive interaction from selecting individually strong candidates, so Section 4 separates these explanations.The comparison isolates route fit from global expert-subspace separation but does not by itself identify interaction effects.
  • Global geometric overlap: ESSI ranged from 0.776 to 1.060 across six models and 18 layers, with median 0.969, indicating substantial overlap rather than hard directional separation.Between-expert separation was comparable to local within-expert variation under the centered directional metric.
  • Interpretation: Together, ESSI and residual ratios reject hard directional partitions and interchangeability, supporting structured overlap with selective routing.Hard partitions predict ESSI well above one, whereas interchangeability predicts residual ratios near one; neither prediction holds.

4. Strong Candidates, Negative Interaction

Across 39 factorial cells, selected experts are stronger candidates than unselected rivals, but the actual prefix consistently narrows that advantage. This negative interaction persists across robustness checks, supporting coherent overlap rather than disjoint geometric coverage.

  • Candidate quality: All 39 comparisons favor the selected expert when the prefix is fixed, with mean lift from 0.001382 to 0.038642 and confidence intervals above zero.Changing candidate and context together instead reverses the result, producing negative later-expert comparisons.
  • Candidate quality: 0.02616 is DeepSeek’s macro candidate advantage, with 95% CI [0.02412, 0.02800]; Aactual > 0 in all 39 point estimates.The selected candidate explains more residual representation than the highest-scoring eligible rival under the actual prefix.
  • Context effects: Every Ts and Tr estimate is negative, including DeepSeek Ts = −0.26864 and Tr = −0.21094 under equal-length prefixes.For OLMoE/Mixtral, Ts ranges from −0.2284 to −0.0573 and Tr from −0.1531 to −0.0379.
  • Interaction: The interaction is negative in all 39 cells, and every 95% interval lies below zero; macro D is −0.05548 for OLMoE/Mixtral and −0.05770 for DeepSeek.The actual prefix removes more novelty from the selected candidate than from the rival.
  • Robustness: Negative interactions persist across raw-gain, nearest-match, and strict-caliper variants, while DeepSeek’s macro interaction remains negative under all three variants.Under the strict caliper, every point estimate remains negative and 20 of 24 intervals exclude zero at 7.88% coverage.

5. Functional Value Within Overlapping Geometry

Functional value persists within overlapping expert geometry: later experts often improve prediction, and controlled Top-2 routing outperforms Top-1. Although the leader is usually most important individually, collective later-expert value can be substantial, showing that geometric overlap is not functional redundancy.

  • Interpretation: Negative geometric interactions measure residual linear coverage, not the value of nonlinear expert computation.Additions measure marginal value, replacement tests leader concentration, and matched training permits adaptation.
  • Later-expert additions: 24 of 39 frozen-route additions reduce next-token NLL, while the remaining 15 intervals include zero.OLMoE is positive in 17 of 21 cases, Mixtral at all three layers, and DeepSeek has four positive additions at early ranks.
  • Marginal value: 0.0906 [0.0836, 0.0975] NLL is recovered by the first OLMoE layer-16 addition, versus 0.0030 [0.0021, 0.0039] for the last.Effect sizes decay with router rank, and later estimates become smaller and less precise.
  • Route concentration: Seven of nine configurations show greater NLL damage from replacing the leader than the aggregate later set.At OLMoE layer 16, replacing seven later experts instead causes 0.0848 [0.0724, 0.0971] NLL damage.
  • Controlled training: Top-2 has lower validation loss than Top-1 in all three controlled-study seeds, with ∆= 0.1016 ± 0.0025.The configurations have the same active intermediate capacity, while total parameters and active compute differ by less than 0.04%.

6. Potential Impact on MoE Analysis and Design

MoE analysis should distinguish route coherence, candidate quality, context-specific geometric synergy, and functional necessity rather than infer one from another. Claims about overlap and compression should specify the representation and intervention, then test retained-route effects directly.

  • Evidence levels: The framework separates route coherence, fixed-prefix candidate quality, context-specific geometric synergy, and output-level functional necessity.Only the interaction establishes context-specific geometric synergy; functional necessity requires an output-level counterfactual.
  • Geometric interpretation: Negative interaction can reflect geometric saturation: the actual prefix may remove directions available to both candidates, reducing the stronger candidate’s marginal novelty.The selected candidate can remain better under a fixed prefix even while losing more marginal novelty, because changing context and candidate mixes residual opportunity.
  • Measurement scope: Complementarity claims must identify the representation, rank, metric, and intervention because rank-128 linear router-input overlap may coexist with nonlinear, output, or logit specialization.Semantic routing and expert collaboration may reveal specialization that this subspace metric does not resolve.
  • Design implications: Similarity can screen compression candidates, but pruning, merging, or skipping should be tested under the retained route because removal changes other experts’ context.Separating subspaces need not improve prediction; adaptive-k routing could instead estimate another expert’s output gain under load and compute constraints.
  • Scope: The conclusions are limited to a linear, rank-128 router-input metric, factorial analyses of three architectures, and one small-scale matched-compute setting.Appendix S7 provides additional scope and interpretation details.

7. Conclusion

The paper identifies coherent overlap: experts share token-relevant geometric neighborhoods, yet routes remain coherent and later experts can provide functional value. Separating overlap, candidate quality, contextual interaction, and function shows that multi-expert benefit does not require disjoint linear coverage.

  • Coherent overlap: 39 factorial cells show selected candidates explain more residual representation than rivals, while negative interactions mean actual context narrows geometric advantage.The prefix-controlled 2 × 2 factorial separates candidate, context, and interaction effects.
  • Coherent overlap: Across six MoEs, eligible experts’ centered directional subspaces overlap under ESSI, while actual routes remain token-coherent.ESSI calibrates expert separation against within-expert variation.
  • Functional value: Frozen-route interventions and a controlled training study show that later experts can remain useful despite overlapping expert subspaces and non-disjoint linear coverage.This reframes multi-expert benefit as potentially arising from distinct computations within a shared token-relevant neighborhood.

Appendix · S1. Scope, Notation, and Experimental Assets

This appendix defines the factorial terminology, candidate constructions, interval notation, and layer-index convention used throughout the analyses. It also specifies the corpora supporting pretrained-model geometry, factorial, intervention, and matched-compute studies.

  • S1. Scope, Notation, and Experimental Assets: “Rival” denotes the highest-scoring eligible unselected routed expert in the factorial experiment.
  • S1. Scope, Notation, and Experimental Assets: Figure 3 uses load-near control candidates, which remain distinct from the factorial experiment’s rival construction.
  • S1. Scope, Notation, and Experimental Assets: Functional intervals overlapping zero are classified as inconclusive.
  • S1. Scope, Notation, and Experimental Assets: Unless otherwise specified in a table header, θ̂ [l, u] reports a point estimate with its 95% confidence interval.
  • S1. Scope, Notation, and Experimental Assets: The three-seed matched-training summary uses mean ± standard deviation, while L denotes the analyzed transformer layer index rather than total model depth.
  • S1. Scope, Notation, and Experimental Assets: The geometry, factorial, and intervention analyses use a fixed 8,192-record corpus spanning general, code, summarization, and dialogue sources.The corpus contains 2,048 records each from C4, CodeSearchNet-Python, ccdv/arxiv-summarization, and UltraChat 200k.
  • S1. Scope, Notation, and Experimental Assets: The matched-compute study uses a separate 8,192-record corpus with the same general, code, and dialogue sources plus ARC science data.Its ARC component comprises 1,024 ARC-Easy and 1,024 ARC-Challenge train records.

S2. ESSI and Route-Coherence Protocol · S2.1. Centered global and local tangent subspaces

The ESSI protocol measures expert overlap using routing-weighted global subspaces and local tangent bases built from sampled token representations. It standardizes expert eligibility, sampling, and normalization across the analyzed MoE architectures.

  • S2.1. Centered global and local tangent subspaces: ESSI uses native Top-k router weights, renormalized within each token, to weight expert assignments.These weights define the routing contribution used in the survey.
  • S2.1. Centered global and local tangent subspaces: An expert becomes eligible after receiving at least 2,048 selected tokens.Eligibility is defined by a minimum routing count.
  • S2.1. Centered global and local tangent subspaces: Each eligible expert receives a global rank-128 basis from centered, routing-weighted PCA.The basis is denoted Ge in the protocol.
  • S2.1. Centered global and local tangent subspaces: Local rank-128 tangent bases are fitted after centering the 256 nearest representations around each anchor.Anchors and candidate pools are sampled without replacement in proportion to expert routing weights.
  • S2.1. Centered global and local tangent subspaces: 2,048 anchors and up to 8,192 candidate tokens are used for OLMoE, Mixtral, and DeepSeek.Qwen3, Gemma4, and Qwen3.6 instead use 512 anchors and up to 4,096 candidate tokens.
  • S2.1. Centered global and local tangent subspaces: The ESSI denominator is the routing-load-weighted mean expert-level local tangent distance, with numerical floor ϵ = 10−12.The base random seed is zero, with the layer index added for layer-specific sampling.

S2.2. Route coherence … S7.2. Reproducibility

Across the analyzed MoE models, expert subspaces overlap, yet selected routes better explain token representations than matched alternatives. Candidate advantages persist across all 39 factorial cells and later experts often improve prediction, supporting coherent overlap rather than geometric redundancy.

  • S2.2. Route coherence; S2.3. Shared-core overlap; S3. Factorial Residual-Attribution Protocol; S3.1. Uncentered bases for the factorial analysis: ESSI and shared-core diagnostics distinguish global expert separation, local within-expert dispersion, and the fraction of expert bases captured by a common layer-wide core.The route-coherence analysis uses centered bases and matched alternative sets, while the factorial uses distinct uncentered bases.
  • S3.2. Load-near control diagnostic in Figure 3: All nine Figure 3 leader intervals exceed zero under changing context, whereas all 39 later-expert intervals are below zero; under fixed context, all 39 intervals exceed zero.Point estimates for the fixed-context intervals range from 0.001382 to 0.038642, with aggregate coverage of 79,852/79,872 token-cells.
  • S3.3. Rivals and alternative contexts: The factorial compares the selected candidate with the highest-scoring eligible routed expert outside the actual Top-k route across matched alternative contexts and source-context bootstrap intervals.DeepSeek keeps its two shared experts fixed and excludes them from routed IDs, candidate pools, rivals, alternative routes, and prefix lengths.
  • S4. Complete Factorial Estimates: 39 cells show positive selected-minus-rival advantage, while every candidate-by-context interaction interval is below zero.The equal-cell macro interaction is −0.05548 [−0.05780, −0.05340], and the DeepSeek macro candidate advantage is 0.02616 [0.02412, 0.02800].
  • S5. Sensitivity Analyses and Numerical Validation: Sensitivity analyses preserve negative interactions across raw-gain and nearest-residual measures, while strict-caliper estimates remain negative but have low coverage.For OLMoE/Mixtral, raw-gain and nearest-residual intervals are negative in every cell; the caliper retains 3,875 of 49,152 token-cells (7.88%).
  • S6. Functional Interventions and Matched Training; S6.1. Leader-versus-later replacement; S6.2. Adjacent ordered-prefix NLL recovery: 24 of 39 frozen-route additions produce positive NLL recovery, while 15 are inconclusive, showing that later experts can add functional value despite geometric narrowing.OLMoE contributes 17 positive and four inconclusive additions; Mixtral contributes three positive; DeepSeek contributes four positive and 11 inconclusive.
  • S6.3. Matched-compute Top-1 versus Top-2: Across three matched-compute seeds, Top-2 achieves lower validation loss than Top-1, with paired difference 0.101627 ± 0.002547.The comparison matches active intermediate capacity exactly and keeps parameter and active-FLOP differences below 0.04%.
  • S7. Scope, Reproducibility, and Interpretation; S7.1. Scope of the evidence; S7.2. Reproducibility: The evidence separates router-input residual geometry from functional consequences, with analytical score-ordered prefixes, controlled alternatives, fixed shared experts, and reproducible train-only fits and bootstrap procedures.NLL interventions and matched training provide complementary functional evidence to ESSI and candidate-by-context interaction.
Loading 2607.28308v1…