Source-linked AI summary
Demystifying Oversmoothing in Sheaf Neural Networks: An Index-Theoretic Criterion
Junwen Dong, Yuhan Peng, Hao Li, Huitao Feng, Kelin Xia
TL;DR
Absolute harmonic-space dimension can misrepresent anti-oversmoothing capacity in sheaf neural networks because larger kernels may contain only constant signals. The paper introduces a relative index-theoretic criterion, extends it to nonlinear stalks through local linearization, and finds that compliant models maintain depth-stable representations while violating models collapse.
Problem
Absolute harmonic-space dimension does not universally measure anti-oversmoothing capacity because enlarged harmonic spaces may contain only channel-wise constant signals.
Method
The paper combines index jumps with a degree-one heat-trace correction and natural holonomy and cohomological conditions, extending the criterion to nonlinear stalks by tangent-space linearization.
Results
Experiments across ten models confirm that the criterion identifies genuine harmonic enlargement beyond the embedded baseline and distinguishes depth-stable models from collapsing ones.
Takeaways & Limitations
Anti-oversmoothing capacity should be assessed through relative harmonic inclusion rather than kernel dimension alone.
Takeaways & Limitations
The simplified certificate has limited attainability, motivating refinement with global topological invariants intrinsic to graph geometry.
Abstract
from arXiv · showhide
To combat oversmoothing in Graph Convolutional Networks, Sheaf Neural Networks (SNNs) were proposed as a generalization by equipping the graph with a sheaf structure and replacing the graph Laplacian with a sheaf Laplacian $\mathcal{L}$. Existing analyses connect sheaf diffusion to oversmoothing via the harmonic space ($\ker\mathcal{L}$), taking its absolute dimension as an indicator of anti-oversmoothing capacity. However, absolute dimension alone is not a reliable measure: certain sheaf configurations inflate $\dim \ker \mathcal{L}$ while their harmonic sections remain entirely constant, without enriching discriminative capacity. We instead introduce the first relative, geometric approach, yielding a precise characterisation of anti-oversmoothing capacity. Under natural conditions on stalk transportation and global sheaf structure, we establish an index-theoretic comparison criterion showing that one sheaf's harmonic space genuinely contains another's beyond trivial inflation. We illustrate this with a concrete instance and further introduce \textit{GyroSheaf}, a sheaf with curved gyrovector-space stalks, extending the criterion to the non-linear setting via local tangent-space linearization. Experiments across ten models confirm the theoretical criterion: sheaf models violating the criterion collapse despite possessing index jumps, while compliant models maintain depth-stable representations.
1 Introduction
The introduction reframes anti-oversmoothing capacity as a relative geometric property of sheaf harmonic spaces rather than their absolute dimension. It proposes combining index jumps with structural corrections to certify strict harmonic-space enlargement and validates the criterion across ten models and three datasets.
- Background: Repeated graph diffusion projects signals onto constant harmonic modes, making node representations increasingly similar and causing classical oversmoothing.For connected graphs, the harmonic space consists of constant signals.
- Background: Sheaf Neural Networks generalize graph diffusion through stalk spaces and restriction maps, whose harmonic sections can preserve node-level differences in deep limits.Setting all stalks to R with identity maps recovers the graph Laplacian used by GCNs.
- Limitation: A positive raw index jump measures harmonic-capacity difference but cannot distinguish meaningful enlargement from trivial channel-wise constant-section inflation.IdentitySheaf and InverseSheaf exemplify configurations that increase dim H0 without improving anti-oversmoothing capacity.
- Our approach: Theorem 5 combines the index jump with a degree-one heat trace correction and non-trivial holonomy to prove strict geometric inclusion of harmonic spaces.Under the capacity conditions, H0(F)/ι*H0(ξ) ≠ 0, so H0(F) strictly contains the embedded H0(ξ).
- Contributions: The work presents an algebraic-topological comparison criterion for arbitrary sheaf frameworks and validates it through oversmoothing experiments across ten models on three datasets.The contribution is described as a precise mathematical characterization of anti-oversmoothing capacity.
2 Related Work
Prior work characterizes oversmoothing in message-passing GNNs through Laplacian smoothing and convergence to low-dimensional subspaces, while sheaf neural networks replace graph diffusion with sheaf diffusion to preserve richer harmonic signals. Existing extensions broaden sheaf methods to hypergraphs and cell complexes, but their analyses rely on absolute harmonic-space dimension.
- Oversmoothing in GNNs: Message-passing GNNs exhibit oversmoothing as Laplacian smoothing and exponential convergence to a low-dimensional subspace determined by graph structure.The cited characterization involves connected components and node degrees.
- Oversmoothing in GNNs: Existing oversmoothing remedies are organized into three broad families, including normalization and regularization.The supplied passage introduces these families but truncates before listing all three.
- Sheaf neural networks and geometric extensions: Sheaf neural networks replace graph diffusion with sheaf diffusion, enabling richer harmonic spaces that can preserve non-constant signals at the diffusion limit.Subsequent work extends sheaf methods to hypergraphs and cell complexes.
- Sheaf neural networks and geometric extensions: Existing sheaf-theoretic analyses measure anti-oversmoothing capacity through the absolute dimension of the harmonic space.The supplied passage ends while introducing this limitation.
3 A Sheaf-Theoretic Diffusion Framework on Graphs
This section develops a sheaf diffusion framework on graphs by defining stalks, restriction maps, coboundaries, pairings, and a sheaf Laplacian. Diffusion converges to the harmonic space, identified with global sections, making its structure central to understanding oversmoothing.
- Sheaf Structure: A discrete sheaf assigns stalk spaces to nodes and edges, with restriction maps encoding compatibility between adjacent stalks.Stalks may be vector spaces or Riemannian manifolds, while restriction maps are linear or smooth accordingly.
- Coboundary and Pairings: The coboundary maps node assignments to edge-wise discrepancies by comparing transported endpoint features in each common edge stalk.In vector spaces, this discrepancy is a signed difference; general stalk geometries require geometry-specific constructions.
- Laplacian Diffusion: The sheaf Laplacian L = δ*δ defines Dirichlet energy and a heat equation whose positive semidefiniteness dissipates energy toward harmonic equilibria.An explicit Euler step gives the diffusion layer X_k+1 = σ((I − L)X_k).
- Nonlinear Extensions: The framework extends nonlinearly through Lie-group operations or local tangent-space linearization when no suitable group structure exists.These alternatives support nonlinear stalk spaces while preserving the framework’s operator-based construction.
- Harmonic Equilibria and Oversmoothing: Diffusion converges to H0 = ker L = ker δ, the globally edge-consistent sections, and the projection determines whether node distinctions collapse or persist.Minimal harmonic capacity yields flat collapse, whereas richer H0 can preserve node-level distinctions.
4 An Index-Theoretic Analysis of Over-Smoothing
The section replaces raw index jumps with a corrected criterion combining degree-one heat-trace information and relative harmonic quotients. Under embedding, holonomy, and capacity conditions, the criterion certifies genuine harmonic enlargement beyond trivial channel replication.
- Corrected index criterion: The corrected capacity difference combines the raw index jump with the degree-one correction Δh1, because Euler balancing prevents the raw index from measuring H0 alone.The relative excess is measured by H0(F)/ι*H0(ξ), separating genuine capacity from trivial inflation.
- Corrected index criterion: Trivial Euclidean harmonic sections can scale with feature dimension while remaining channel-wise constant and preserving no node-level distinctions.Both IdentitySheaf and InverseSheaf exhibit this trivial inflation.
- Genuine harmonic inclusion: Theorem 5 requires a cohomology-injective map ι: ξ → F together with dimensional and capacity controls based on non-trivial holonomy and index-related certificates.The holonomy-fixed subspace must be smaller than the full stalk, and the theorem concludes dim H0(F) > dim H0(ξ) with a nontrivial quotient.
- Genuine harmonic inclusion: The resulting quotient H0(F)/ι*H0(ξ) measures excess non-dissipative capacity originating from holonomy rather than higher-rank channel enlargement.For small β1(G), one simplified certificate is Δind(F, ξ) ≥ (n − k)β1(G) + 1.
- SPD illustration: The SPD sheaf provides a concrete non-Euclidean comparison: its Euclidean embedding preserves global sections, congruence transports generate non-trivial holonomy, and heat-trace comparison supplies capacity control.Its finite-dimensional Hodge complex is justified through logarithmic and congruence-based structure after linearization.
5 Nonlinear Generalization
For nonlinear stalks, global linear index and spectral tools fail, so the theory is reformulated by tangent-space linearization at harmonic states. GyroSheaf realizes this framework on curved gyrovector-space stalks while preserving local Hodge-theoretic and index structure.
- Nonlinear Generalization: Nonlinear stalks produce a nonlinear harmonic set and gradient flow, invalidating the global kernel-dimension and spectral tools used by the linear index theorem.The compatibility map is nonlinear, so harmonic sections form H rather than ker δ, and diffusion is no longer generated by a trace-class operator.
- Tangent Linearization: At each harmonic state f, linearizing δ yields dδ_f on tangent cochains, making the nonlinear theory a local version of the preceding index-jump framework.The differential maps T_fC0 to T_0C1 and serves as the tangent-space analogue of the coboundary operator.
- Tangent Linearization: Nonlinear harmonic capacity is measured by dim T_fH = dim ker dδ_f, so enlargement becomes a local tangent-dimension increase rather than a global vector-space dimension jump.The local analytic index is ind_f(dδ) := dim ker dδ_f − dim coker dδ_f, and the index–heat trace criterion applies after tangent linearization.
- GyroSheaf Construction: GyroSheaf maps SPD features to a bounded Poincaré-ball gyrovector domain and defines nonlinear edge inconsistencies, tangent operators, and a Laplacian from Dirichlet-energy gradients.Its nonlinear coboundary is replaced by the tangent map Dδ(X), while the Laplacian is computed as the gradient of total inconsistency energy.
- GyroSheaf Construction: GyroSheaf supplies the tangent complex, Hodge decomposition, and local harmonic space, while O(n)-congruence restriction maps preserve the linearised SPD complex’s holonomy and index data.The tangent cochain spaces use the Frobenius pairing, making the linearised coboundary an operator between Hilbert spaces.
6 Experiments
Experiments progressively test Theorem 5 from trivial-sheaf GNN baselines through violating controls to compliant Euclidean and geometric sheaves. Results show that index jumps alone do not prevent oversmoothing, whereas satisfying the holonomy conditions yields depth-stable representations.
- Experimental design: The experiments compare GCN baselines, Tier 1 violating controls, and compliant Tier 2–3 sheaves with increasing index jumps.This progression directly tests whether Theorem 5 governs anti-oversmoothing capacity.
- Baseline: Rank-1 GCN, GAT, and GraphSAGE use the graph Laplacian as a trivial sheaf, with no index jump or sheaf structure.Oversmoothing is treated as the default baseline behavior.
- Tier 1: violating controls: IdentitySheaf and the second Tier 1 control have nonzero index jumps but trivial holonomy, reducing their harmonic spaces to channel-wise constant copies of the scalar baseline.Their theorem conditions are violated, so enlarged stalks do not provide genuine harmonic enlargement.
- Tier 2: Euclidean stalks: Tier 2 DiagSheaf, BundleSheaf, and GeneralSheaf use diagonal, orthogonal, and general maps whose learned holonomy satisfies the theorem’s conditions.They share Tier 1’s stalk dimension and index jump but can preserve enough fixed directions.
- Tier 3: geometric stalks: Tier 3 SPDSheaf and GyroSheaf satisfy both conditions, with GyroSheaf’s effective dimension r = d(d+1)/2 > d enabling strict enlargement of the relevant harmonic quantity.On cyclic graphs, the enlargement is expressed as Δh0 = Δind + Δ(1)Str.
- Untrained regime: Baseline and Tier 1 controls collapse by six to eight orders of magnitude within 128 layers, while Tier 2 delays decay and Tier 3 keeps both metrics essentially flat.This ordering appears across Cora, Citeseer, and Texas and confirms that index jumps without holonomy conditions are insufficient.
- Trained regime: On trained Cora, baseline accuracy drops sharply beyond L=8, whereas all sheaf models maintain ∼80% accuracy despite structural distinctions being masked by finite-depth optimization.This includes the Tier 1 controls and shows no finite-depth accuracy penalty from enlarged harmonic capacity.
7 Conclusion
The paper introduces a relative criterion for anti-oversmoothing capacity in Sheaf Neural Networks, extending it to nonlinear stalks and validating it across ten models.
- 7 Conclusion: Theorem 5 shows that a positive index jump, together with natural holonomy and cohomological conditions, guarantees genuine harmonic enlargement beyond the Euclidean baseline.This provides the paper’s relative criterion for anti-oversmoothing capacity.
- 7 Conclusion: The framework extends to nonlinear stalks through local linearisation and is instantiated by GyroSheaf.
- 7 Conclusion: Experiments across ten models confirm the proposed criterion.
A Limitations … B.2 Details of 3.2 (Coboundary Operator)
The paper limits its criterion to standard graphs while formalizing sheaves and coboundaries through discrete stalks, cochains, transport maps, and intrinsic differences. For non-linear stalks, the coboundary requires identifying suitable structure or locally linearizing tangent spaces, recovering the classical operator in the Euclidean case.
- A Limitations: The theoretical framework and experiments focus on standard graph structures, leaving extensions to hypergraphs and cell complexes for future work.The limitation concerns generalizing the index-theoretic criterion beyond standard graphs, where the coboundary operator may require further development.
- B.1 Details of 3.1 (Sheaves on Graphs): The graph is treated as a one-dimensional cell complex sampled from a manifold, connecting discrete graph data to a continuous fiber bundle.This construction bridges the continuous bundle π : E →M and the graph-based sheaf representation.
- B.1 Details of 3.1 (Sheaves on Graphs): Node stalks are manifold fibers, edge stalks represent midpoint or shared spaces, and restriction maps discretize connection-induced parallel transport.For an edge e = (u, v), Fu→e and Fv→e transport node features into the edge stalk Fe.
- B.1 Details of 3.1 (Sheaves on Graphs): A 0-cochain assigns local features to vertices, while 1-cochains encode feature differences across edges.The spaces C0(G, F) and C1(G, F) discretize smooth sections and corresponding differential forms, respectively.
- B.2 Details of 3.2 (Coboundary Operator): The coboundary operator δ : C0 →C1 measures whether a 0-cochain fails to form a globally parallel section.It directly simulates the continuous operator d0s = s∗ω.
- B.2 Details of 3.2 (Coboundary Operator): Along each edge, δ compares transported node features using an intrinsic geometric difference in the common edge stalk.The restriction maps carry both endpoint features into Fe before the difference is formulated.
- B.2 Details of 3.2 (Coboundary Operator): Because non-linear output spaces lack natural algebraic structure, δ can use inherent space structure or local linearization on tangent spaces.The latter approach identifies an algebraic structure locally for defining the required difference.
- B.2 Details of 3.2 (Coboundary Operator): For Euclidean stalks with linear restriction maps, the geometric difference becomes the standard vector difference and δ exactly recovers the classical degree-0 coboundary operator.The resulting formulation agrees with prior Euclidean sheaf treatments.
B.3 Proof of Proposition 1 (Energy-Laplacian Correspondence) … C.3 Details of Remark 3 (Index Jump Alone is Insufficient)
The appendices establish that sheaf diffusion dissipates Dirichlet inconsistency toward harmonic states, which coincide with global sections and determine long-time representation capacity. They further show that index increases must be interpreted relatively: oversmoothing resistance depends on additional, geometrically meaningful harmonic modes and frustration growth, not index jumps alone.
- B.3 Proof of Proposition 1 (Energy-Laplacian Correspondence): The sheaf Laplacian is the energy gradient, its quadratic form measures accumulated edgewise geometric discrepancies, and diffusion monotonically dissipates this inconsistency toward a harmonic state.The proof uses L = δ*δ and shows ∇E(x) = Lx, ⟨Lx, x⟩ = ∥δx∥2 ≥ 0, and energy decreases under ∂tx = −Lxt.
- B.4 Proof of Theorem 2 (0th-order Discrete Hodge Decomposition): The harmonic space satisfies ker L = ker δ and is isomorphic to H0(G, F), so diffusion steady states are exactly the sheaf’s global sections.Global sections are precisely 0-cochains compatible with all restriction maps, equivalently satisfying δf = 0.
- B.5 Technical Details for Proposition 3 (Harmonic Characterisation of OverSmoothing): As diffusion time tends to infinity, positive-eigenvalue components vanish and the representation converges to the projection of the input onto ker L.For centered inputs, the expected preservation ratio is approximately 1/N, so standard diffusion loses discriminative information at rate O(1/N).
- B.5 Technical Details for Proposition 3 (Harmonic Characterisation of OverSmoothing): Strong enrichment, H0_1 ⊊ H0_2, guarantees pointwise nondecrease of retained signal under projection, unlike weak enrichment based only on harmonic dimension.The decomposition H0_2 = H0_1 ⊕ W yields ∥P_H0_2 f(0)∥2 ≥ ∥P_H0_1 f(0)∥2 for every input.
- C.1 Algebraic Foundations of the Discrete Index 4.1: Discrete Hodge theory links harmonic capacity to the sheaf’s topological index, whose supertrace remains invariant under diffusion because positive spectral contributions cancel.The deep-limit capacity is tied to the kernel of the Laplacian, while the Hodge–Dirac heat supertrace equals Ind(F).
- C.2 Proof of Proposition 4 (Topological Index and Index Jump): For rank r sheaves, the index contribution is rχ(G), so comparing rank-r and rank-n systems produces a rank term (r − n)χ(G).The proposition derives the identity from Euler–Poincaré and applies the same argument to the rank-n Euclidean sheaf before subtraction.
- C.3 Details of Remark 3 (Index Jump Alone is Insufficient): When χ(G) < 0 and r > n, increasing stalk dimension contributes a negative rank term, so greater harmonic capacity requires sufficient growth of the frustration space H1(F).For connected cyclic graphs, |E| ≥ |V| and χ(G) ≤ 0; H1(F) is interpreted as the divergence-free space.
- C.3 Details of Remark 3 (Index Jump Alone is Insufficient): The relevant certificate is relative excess harmonic capacity, H0(F)/ι*H0(ξ), whose node-separating modes can resist oversmoothing; in the matched-rank case, this excess reflects H1(F) growth.Thus an index jump alone is not a pointwise oversmoothing theorem, while geometric frustration mediates resistance beyond the Euclidean baseline.
C.4 Details of Theorem 5 (Genuine Harmonic Inclusion) · C.5 Details of the SPD Sheaf Case Study 4.3 · D Nonlinear Generalization
The theorem distinguishes genuine harmonic inclusion from trivial channel replication by requiring holonomy-fixed structure and relative comparison with the Euclidean baseline. The SPD sheaf case study shows that this criterion can hold in a concrete non-Euclidean model through baseline-preserving embedding, geometric transport, and excess harmonic sections.
- C.4 Details of Theorem 5 (Genuine Harmonic Inclusion): Trivial holonomy only replicates scalar Euclidean cohomology across independent channels, creating strict capacity gain only when r > n.For connected G, the dimension difference is r − n; when r = n, there is no excess over the Euclidean baseline.
- C.4 Details of Theorem 5 (Genuine Harmonic Inclusion): Genuine inclusion instead requires a nontrivial holonomy-fixed subspace W that extends a Euclidean subcomplex while complementary directions retain nontrivial holonomy.This condition is weaker than global flatness and supports Euclidean cohomology without reducing the entire sheaf to replicated identity channels.
- C.4 Details of Theorem 5 (Genuine Harmonic Inclusion): For connected G, the holonomy-fixed dimension must satisfy k > n for strict enlargement relative to the rank-n Euclidean baseline, or k > 1 relative to the scalar baseline.Under the local-system condition, the excess is dim H0(F) − dim H0(ξ) = k − n.
- C.4 Details of Theorem 5 (Genuine Harmonic Inclusion): The corrected index–heat trace balance supplements the raw index with the degree-one heat-trace contribution needed to recover zeroth harmonic capacity.On cyclic graphs, the raw index jump can be negative, while the correction records cycle-level geometric frustration and restores the correct H0 comparison.
- C.5 Details of the SPD Sheaf Case Study 4.3: The SPD sheaf supplies a concrete non-Euclidean setting with SPD stalks, congruence-type transports, logarithmic linearization, admissible pairing, and Hodge-compatible operators.After logarithmic identification, congruence transports become linear isometric transports on symmetric matrices, enabling the finite-dimensional index–heat trace framework.
- C.5 Details of the SPD Sheaf Case Study 4.3: The Euclidean-to-SPD embedding preserves Euclidean global sections inside the SPD harmonic space, establishing the required baseline comparison map.The SPD sheaf therefore contains the Euclidean harmonic baseline as a comparable subspace rather than merely having a larger ambient stalk.
- C.5 Details of the SPD Sheaf Case Study 4.3: Because SPD global sections extend beyond the lifted Euclidean image, the relative quotient is nontrivial and detects genuine harmonic capacity beyond constant-channel replication.The comparison is explicitly relative to the Euclidean baseline rather than inferred from the absolute size of a single kernel.
- C.5 Details of the SPD Sheaf Case Study 4.3: The SPD prototype separates pairing-dependent diffusion energy from index-based harmonic comparison, illustrating the framework’s relative anti-oversmoothing principle.The cited discussion presents the construction as a concrete realization of the comparison theorem, not as a complete prior SPD oversmoothing theory.
D.1 Details of Non-Linear Index Jump 5 … E.1 Datasets
The nonlinear analysis replaces global linear index machinery with local tangent-space operators and characterizes harmonic capacity through local geometry. The appendix then specifies GyroSheaf differentiation, Hodge-type decomposition, experimental regimes, and datasets spanning homophilic and heterophilic graphs.
- D.1 Details of Non-Linear Index Jump 5: Nonlinear stalks turn cochains into a product manifold, replace the linear coboundary with a nonlinear constraint map, and replace ker δ with a nonlinear harmonic set.The global-section condition remains δf = 0, while harmonic sections form a nonlinear set.
- D.1 Details of Non-Linear Index Jump 5: Linearizing at a compatible state produces a self-adjoint tangent Laplacian whose kernel contains precisely the directions surviving nonlinear diffusion.The local tangent heat operator is well-defined even though global linear heat-kernel and supertrace machinery fails.
- D.1 Details of Non-Linear Index Jump 5: The nonlinear global-section capacity is determined by the local dimension of the harmonic submanifold rather than a single global kernel dimension.When the harmonic set is a smooth submanifold, this local dimension provides the nonlinear analogue of global section capacity.
- D.1 Details of Non-Linear Index Jump 5: Under surjectivity of dδf, the local harmonic dimension equals the local analytic index, while McKean–Singer cancellation survives after tangent-space linearization.The local index is ind_f(dδ) := dim ker dδf − dim coker dδf, and surjectivity yields dim T_fH = ind_f(dδ).
- D.2 The Calculation of Dδ and (Dδ)∗of 5.2: GyroSheaf defines edgewise gyro-coboundaries and computes their tangent maps using Fréchet derivatives of principal matrix inverse square roots.The associated adjoints use the Frobenius inner product, while experiments obtain equivalent gradients by automatic differentiation through the gyro-Laplacian.
- D.3 Proof of Proposition 6 (Hodge-type Decomposition for the Nonlinear GyroSheaf): The nonlinear GyroSheaf admits a Hodge-type decomposition because finite-dimensional tangent spaces provide inner products, a linearized boundary map, and its adjoint.These ingredients yield the corresponding orthogonal-complement relations and direct-sum decomposition.
- E Experimental Setup: Experiments separately evaluate untrained propagation dynamics and trained end-to-end behavior, with model implementations and compute details documented in the appendix.The two regimes are configured separately to test intrinsic propagation and whether the properties survive optimization.
- E.1 Datasets: Cora, Citeseer, and Texas evaluate layer-wise propagation across homophilic citation graphs and a heterophilic web graph.The datasets therefore span qualitatively different connectivity regimes.
E.2 Model Implementations … E.6 Compute Environment
The paper implements Euclidean, linear-sheaf, SPD, and GyroSheaf models, evaluates oversmoothing with geometry-appropriate but dimension-matched measures, and compares untrained and trained trajectories under shared experimental settings. All reported experiments use common seeds and depths and run on a single consumer-grade GPU.
- E.2 Model Implementations: Linear sheaf models use the original NSD diffusion implementation, while IdentitySheaf shares its backbone and Laplacian builder.DiagSheaf, BundleSheaf, GeneralSheaf, and IdentitySheaf are implemented within the NSD framework; edge transports use invertible LU-parameterized maps shared in both directions.
- E.2 Model Implementations: SPDSheaf is a from-scratch PyTorch implementation using SPD-valued node states, core SPD geometry, and a sheaf Laplacian.The benchmark omits the molecule-specific dual-stream architecture and coordinate-to-SPD lifting.
- E.2 Model Implementations: GyroSheaf maps SPD states into the symmetric unit ball through the Cayley transform and differs from SPDSheaf only in edgewise propagation geometry.This design isolates the effect of intrinsic curved propagation while retaining identical restriction-map parameterization.
- E.3 Oversmoothing Measures: Oversmoothing is measured with Dirichlet energy and mean-average-distance on the backbone representation immediately after the final propagation layer.The trained evaluations exclude classifier logits, and normalized Dirichlet energy divides by ||X||_F to ensure invariance to global feature rescaling.
- E.3 Oversmoothing Measures: Euclidean models use the ℓ2 inner product, SPD models use Log-Euclidean distances, and GyroSheaf is evaluated after inverse-Cayley recovery under the SPD measures.All four measures share functional form and normalization; matched feature dimensions, 120 = d, make trajectory shapes directly comparable across ten models.
- E.4 Untrained Regime: The untrained regime records purely forward-propagated representations across depths L ∈ {1, 2, 4, 8, 16, 32, 64, 128} and five seeds.No optimizer, loss, or training step is used, and Euclidean and linear-sheaf models have hidden width 120.
- E.5 Trained Regime: The trained regime covers Cora, depths L ∈ {2, 4, 8, 16, 32}, and five seeds using one shared end-to-end training recipe across model families.Oversmoothing curves again use trained backbone hidden representations rather than classifier logits.
- E.6 Compute Environment: All experiments ran on one NVIDIA GeForce RTX 3090 (24GB) with Intel Xeon Gold 5218R CPUs and 376GB system memory.The reported workload spans ten models, three datasets, five seeds, and eight depth configurations in both regimes, remaining reproducible on a single consumer-grade GPU.
F Additional Experiments · F.1 Within-Tier Discrimination: SPDSheaf vs. GyroSheaf under Nonlinear Distortion
Within Tier 3, SPDSheaf and GyroSheaf share the same index-theoretic quantities, so their distinction emerges only under nonlinear observation distortion. As AffErr increases, GyroSheaf preserves accuracy while SPDSheaf degrades.
- F.1 Within-Tier Discrimination: SPDSheaf vs. GyroSheaf under Nonlinear Distortion: SPDSheaf and GyroSheaf share stalk dimension, restriction maps, index jump ∆ind, and holonomy fixed subspace W, so the index criterion cannot separate them.The experiment instead compares how each model realizes the tangent complex.
- F.1 Within-Tier Discrimination: SPDSheaf vs. GyroSheaf under Nonlinear Distortion: SPDSheaf uses one global matrix-logarithm chart, whereas GyroSheaf uses a state-dependent tangent-complex construction.The global-chart shortcut is faithful when observations respect the chart but loses fidelity outside the globally linearizable regime.
- F.1 Within-Tier Discrimination: SPDSheaf vs. GyroSheaf under Nonlinear Distortion: The experiment uses a 3-arm pinwheel with angular labels, a fixed symmetrized 8-nearest-neighbor latent graph, and nonlinear projected 16-dimensional observations.The setup includes stratified splits with 20 train and 20 validation nodes.
- F.1 Within-Tier Discrimination: SPDSheaf vs. GyroSheaf under Nonlinear Distortion: Isotropic Gaussian noise progressively distorts observed features while latent coordinates, labels, and the graph remain fixed.Re-standardization keeps feature scales comparable across noise levels, so only η varies.
- F.1 Within-Tier Discrimination: SPDSheaf vs. GyroSheaf under Nonlinear Distortion: AffErr measures departure from local affinity by fitting a least-squares affine map to each node’s 8 latent-space nearest neighbors.AffErr equals 0 for globally affine observation maps, and larger values indicate increasing global-chart stress.
- F.1 Within-Tier Discrimination: SPDSheaf vs. GyroSheaf under Nonlinear Distortion: At AffErr = 0.002, both models perform similarly because the global chart remains faithful.This establishes the low-distortion regime in which the two linearization strategies both succeed.
- F.1 Within-Tier Discrimination: SPDSheaf vs. GyroSheaf under Nonlinear Distortion: At AffErr = 0.273, SPDSheaf trends downward to 0.73 while GyroSheaf remains essentially flat at 0.83.The results show GyroSheaf tolerating nonlinear distortion that defeats SPDSheaf’s global-chart linearization.