Source-linked AI summary

Multi-Domain Riemannian Graph Gluing for Building Graph Foundation Models

Li Sun, Zhenhao Huang, Silei Chen, Lanxu Yang, Junda Ye, Sen Su, Philip S. Yu

arXiv:2603.00618v1cs.LG

TL;DR

Multi-domain graph pre-training lacks a principled account of how knowledge is integrated and transferred across semantically heterogeneous graph domains. GraphGlue uses neural manifold gluing to unify graphs on a smooth Riemannian manifold, and it improves cross-domain few-shot transfer while exhibiting geometric scaling with more pre-training datasets.

  • Problem

    Existing methods lack a principled framework for integrating and transferring knowledge across heterogeneous graph domains, while general manifolds underlying multi-domain graphs remain unexplored.

  • Method

    GraphGlue uses neural manifold gluing to unify multi-domain graphs into a smooth Riemannian manifold, with EMA prototyping and geometric consistency for transferability.

  • Results

    GraphGlue outperforms strongest baselines in cross-domain few-shot transfer, including by 4.9% on Computers and 2.3% on Reddit in the 1-shot setting.

  • Takeaways & Limitations

    Adding semantically distinct datasets steadily improves GraphGlue’s transfer performance, supporting its geometric scaling law and knowledge integration across domains.

  • Takeaways & Limitations

    The theoretical construction assumes ε>0 is small and imposes conditions on the added edge weights.

Abstract

from arXiv · show

Multi-domain graph pre-training integrates knowledge from diverse domains to enhance performance in the target domains, which is crucial for building graph foundation models. Despite initial success, existing solutions often fall short of answering a fundamental question: how is knowledge integrated or transferred across domains? This theoretical limitation motivates us to rethink the consistency and transferability between model pre-training and domain adaptation. In this paper, we propose a fresh Riemannian geometry perspective, whose core idea is to merge any graph dataset into a unified, smooth Riemannian manifold, enabling a systematic understanding of knowledge integration and transfer. To achieve this, our key contribution is the theoretical establishment of neural manifold gluing, which first characterizes local geometry using an adaptive orthogonal frame and then "glues" the local pieces together into a coherent whole. Building on this theory, we present the GraphGlue framework, which supports batched pre-training with EMA prototyping and provides a transferability measure based on geometric consistence. Extensive experiments demonstrate its superior performance across diverse graph domains. Moreover, we empirically validated GraphGlue's geometric scaling law, showing that larger quantities of datasets improve model transferability by producing a smoother manifold. Codes are available at https://github.com/RiemannGraph/GraphGlue.

1 INTRODUCTION

Multi-domain graph pre-training faces semantic heterogeneity and leaves knowledge integration and transfer insufficiently explained. The paper addresses this gap through neural manifold gluing, which unifies graph datasets as a smooth Riemannian manifold and supports the GRAPHGLUE pre-training-adaptation framework.

  • Motivation: Semantic heterogeneity across domains makes multi-domain graph pre-training challenging, while LLM-based methods remain limited to text-attributed graphs.Many real-world graphs lack explicit textual information.
  • Motivation: Existing methods learn shared or invariant knowledge with graph structures and use adaptation techniques, but do not adequately explain how knowledge is integrated or transferred across domains.Examples include graph codebooks, motifs, computation trees, domain tokens, and in-context learning.
  • Theory: Neural Manifold Gluing integrates graph datasets into a unified, smooth Riemannian manifold by characterizing local geometry and coherently gluing local pieces.The theory provides a differential-geometry foundation for analyzing knowledge integration and transfer.
  • Framework: GRAPHGLUE extends local geometry globally with EMA prototyping for batched pre-training, learnable prompts, and a Riemannian Mixture-of-Experts for adaptation.EMA prototyping distinguishes domain semantics through different manifold locations and handles large-scale graphs in batches.
  • Contributions: The paper studies the theoretical foundations of multi-domain graph pre-training and proposes a geometry-based account of cross-domain knowledge transfer.Its contributions include a fresh differential-geometry perspective and neural manifold gluing for unified multi-domain graphs.

2 RELATED WORK

Related work spans general-purpose and specialized graph foundation models, multi-domain graph pre-training, graph adaptation, and Riemannian representation learning. Existing multi-domain methods address semantic heterogeneity, while the paper distinguishes its focus on constructing a general manifold for multi-domain pre-training.

  • Graph Foundation Models: Graph Foundation Models target pre-trainable, general-purpose deep learning architectures for graphs and have expanded to text-attributed graphs.
  • Graph Foundation Models: Specialized graph foundation models cover knowledge graphs, recommender systems, and molecular graphs, while multi-domain pre-training targets general-purpose modeling for text-free graphs.
  • Multi-domain Graph Pre-training: Multi-domain graph pre-training uses generative or contrastive self-supervised learning and seeks shared or invariant knowledge across semantically heterogeneous domains.
  • Graph Fine-tuning and Prompt Learning: Graph adaptation includes fine-tuning with limited target-domain data and prompting that keeps pre-trained parameters frozen while adding learnable components.
  • Riemannian Graph Representation Learning: Most Riemannian graph models target specific tasks, whereas the paper focuses on multi-domain pre-training and constructing a general rather than task-specific manifold.

3 NOTATIONS AND PRELIMINARIES

This section introduces the Riemannian-geometric concepts and moving-frame perspective underlying the paper, then formalizes multi-domain graph pre-training and transfer to target domains. The setup permits targets seen during pre-training or entirely unseen, motivating a principled interpretation of transferability.

  • Riemannian Geometry: A Riemannian manifold is a smooth manifold equipped with a metric tensor, with each point associated with a tangent space and volume element.The volume element is denoted |G(p)|, and Ricci curvature governs changes in volume elements.
  • Cartan’s Method of Moving Frame: Cartan’s Method of Moving Frame provides a principled framework for studying manifold geometry, while its deep-learning methodology remains largely unexplored.The paper positions its work as bridging this gap.
  • Multi-domain Graph Pre-training: Multi-domain graph pre-training first trains a model on graph datasets from multiple source domains and then adapts it to a target domain.Graphs are represented as G = (V, E) with feature matrix X ∈ R|V|×F, and the source collection contains K graphs from L domains.
  • Multi-domain Graph Pre-training: The encoder is frozen during target adaptation, and the target domain may be either included in or absent from the pre-training domains.The pre-trained model generates informative target-graph representations with slight adaptation, supporting the goal of principled transferability.

4 THEORY: CONSTRUCTING A UNIFIED, SMOOTH MANIFOLD

The theory constructs a unified, smooth Riemannian manifold by characterizing local graph geometry and gluing local pieces through metric-compatible translations, holonomy, and curvature-based smoothing. It establishes conditions for smooth gluing and motivates a geometric scaling law linking more datasets to smoother manifolds and improved transferability.

  • Neural Manifold Gluing: Neural manifold gluing characterizes local geometry before connecting local pieces into a unified, smooth Riemannian manifold.The framework is intended to provide a principled foundation for analyzing knowledge integration and transfer across graph domains.
  • Local Geometry: A (k, M)-sparse perturbation generates tangent vectors, and an adaptive orthogonal frame forms the tangent-space basis for each representation.The perturbation is attached to a parametric fGNN, while the frame is obtained through QR-decomposition with length recovery.
  • Local Geometry: Length recovery is important because tangent-vector lengths describe space deformation, while basis-vector angles and lengths reflect twisting and stretching.The theory provides an upper bound on tangent-vector length in terms of the perturbation.
  • Gluing: Tangent edge translations preserve inner products across adjacent local metrics, induce boundary isometries, and yield a unique global continuous metric.Higher-order motifs may create offsets, so holonomy enforces consistency around cycles before smoothing the connected manifold.
  • Smoothing: Log-determinant smoothness uses the logarithmic volume density and graph Dirichlet energy to formulate a computationally efficient curvature loss.The resulting smoothing targets C2 continuity and eliminates folds that hinder knowledge transport along the manifold.
  • Smooth Manifold and Scaling Law: Theorem 4.11 states that log-determinant ∞-order smoothness and trivial holonomy with metric-preserving diffeomorphisms are sufficient for smooth manifold gluing.The theory further deduces that increasing dataset quantities produces a smoother manifold and improves model transferability.

5 GRAPHGLUE: GEOMETRIC MULTI-DOMAIN GRAPH PRE-TRAINING

GraphGlue pre-trains a unified Riemannian manifold by learning local geometry, EMA-prototyping domain semantics, and gluing local pieces together. Its adaptation and GTM mechanisms align target graphs geometrically and quantify transfer difficulty through deformation, while holonomy and curvature losses promote smoothness.

  • Pre-training: GraphGlue learns local geometry, applies EMA prototyping to distinguish domain semantics, and glues local pieces into a unified manifold during batched pre-training.EMA prototyping supports efficient handling of large-scale graphs and stabilizes Riemannian prototype averages during pre-training.
  • Domain Adaptation: During adaptation, prompt matrices adjust target coordinates and tangent-frame metrics, while nearest-prototype connections and holonomy-curvature losses enforce geometric consistency.The transfer graph connects each target sample to its k-nearest prototypes; Lholo penalizes non-trivial holonomy and Lcurv penalizes abrupt volume changes.
  • Transferability: GTM measures the minimal geometric deformation required to merge a target graph into the pre-trained manifold without disrupting learned local geometry.GTM is defined as GTM(GT; S) = ΔH + ΔC, combining holonomy and curvature disagreement.
  • Transferability: Low GTM indicates seamless integration and high transferability, whereas high GTM indicates geometric alienness and greater transfer effort.Unlike source-target similarity measures, GTM assesses geometric consistency within GraphGlue and provides an interpretable transfer-difficulty measure.
  • Further Insight: Lholo preserves topological continuity at gluing boundaries, while Lcurv induces k-order smoothness through log-determinant smoothness, controlling the global metric’s smoothness.The framework connects objective smoothness with generalization error and provides complexity analysis in the appendices.

6 EXPERIMENTS

Experiments evaluate GraphGlue across six graph domains using few-shot cross-domain transfer, showing strong performance and support for its transferability measure, semantic-data integration, and geometric scaling law. Ablations further indicate that holonomy-based gluing and Ricci-curvature smoothing contribute to downstream performance.

  • Evaluation Protocol: Evaluation uses leave-one-out transfer, pre-training on five source datasets and fine-tuning on one held-out target with 1 or 5 labeled samples per class.The remaining target data is split into 10% validation and 90% testing, with ACC and other task-specific metrics used for evaluation.
  • Main Results on Cross-domain Transfer Learning: In the 1-shot setting, GraphGlue outperforms the strongest baselines by 4.9% on Computers and 2.3% on Reddit.In the 5-shot Reddit setting, GraphGlue achieves 85.0% ACC.
  • Ablation Study: Ablations show that both holonomy-based gluing and Ricci-curvature smoothing are important for downstream tasks.These components correspond to the proposed Lholo and Lcurv losses.
  • On Transferability Measure: The transferability measure GMT tracks transfer effort because curvature loss and cross-entropy test-task loss decrease and converge with the same pattern.Convergence of curvature-loss oscillation amplitude also implies convergence of test-task loss.
  • Case Study: Adding semantically distinct PROTEINS and HIV datasets to Reddit-only pre-training steadily improves GraphGlue on Reddit in the 1-shot setting, whereas GCOPE can suffer negative transfer.The case study incrementally adds the distinct datasets and evaluates on Reddit.
  • On Geometric Scaling Law: Increasing the number of pre-training datasets steadily raises average 1-shot accuracy on Computers and Reddit, with observed logarithmic scaling supporting the geometric scaling law.More labeled samples restrain the scaling effect, while larger dataset quantities increase the learned manifold’s expressive power.

7 CONCLUSION … B.2 PROOF OF THEOREM 4.5

The paper concludes that GraphGlue unifies arbitrary graph datasets on a smooth Riemannian manifold through neural manifold gluing, while its appendices formalize notation and prove key geometric results. In particular, Theorem 4.5 establishes the tangent edge translation as a minimum-norm isometry that induces a metric-preserving local map.

  • 7 CONCLUSION: The conclusion presents neural manifold gluing and GraphGlue as a Riemannian framework for unifying arbitrary graph datasets and understanding cross-domain knowledge transfer.The framework merges local pieces into a coherent, smooth Riemannian manifold.
  • APPENDIX: TABLE OF CONTENT: The appendix table of contents lists an Ethics Statement on page 41.
  • A NOTATIONS: The notation appendix defines symbols for graph datasets, Riemannian prototypes, tangent edge translation, holonomy, and curvature losses.It also introduces notation for positive-definite matrices and prototype-level contrastive learning.
  • B PROOFS: The proofs appendix develops theoretical guarantees for perturbation effects and geometric consistency in GraphGlue.Its highlighted results include an upper bound on tangent vector length and an isometric tangent translation theorem.
  • B.1 PROOF OF THEOREM 4.3: Theorem 4.3 analyzes heat diffusion under (k, M)-sparse graph perturbations and isolates the perturbation-induced component affecting other nodes.The affected node support has size at most kM, and the proof concludes using finiteness of heat-kernel elements.
  • B.2 PROOF OF THEOREM 4.5: Theorem 4.5 states that tangent edge translation is the optimal solution to the associated geometric optimization problem.The construction uses a general-linear-group transformation satisfying metric preservation between local frames.
  • B.2 PROOF OF THEOREM 4.5: The metric geometric mean is used to ensure geometric consistency and symmetry when the local metrics do not commute.The direct minimum-norm construction is asymmetric in Gi and Gj unless they commute.
  • B.2 PROOF OF THEOREM 4.5: P (i,j) is a minimum-norm isometry and global minimizer, and under smooth chart compatibility it lifts to a local diffeomorphism preserving the metric.The isometry satisfies Gj(P (i,j)u, P (i,j)v) = Gi(u, v).

B.3 PROOF OF THEOREM 4.6 … C.5 SMOOTHNESS AND HARMONIC FUNCTIONS

The appendix establishes conditions for gluing local graph charts into a unique continuous or smooth Riemannian manifold, then reviews the geometric notions underlying curvature, holonomy, volume change, and smoothness. It also connects harmonic log-volume density to regularity of the learned metric.

  • B.3 PROOF OF THEOREM 4.6: Local metrics connected by isometric tangent translations glue into a unique continuous global Riemannian metric extending every chart metric.The metric is well-defined on overlaps, continuous across overlaps, and uniquely determined by the local metrics.
  • B.4 PROOF OF THEOREM 4.8 AND CLARIFICATION: If every edge lies in a triangle and triangular holonomies are identity, holonomy is identity on every graph cycle; triangle coverage is sufficient but not necessary for gluing.The result follows because triangular cycles generate the cycle space and holonomy is multiplicative.
  • B.6 PROOF OF THEOREM 4.11: Log-determinant ∞-order smoothness and trivial, metric-preserving gluing maps yield a smooth Riemannian manifold from the graph data.Trivial holonomy ensures path-independent transport and compatible overlaps, while log-determinant smoothness supports smooth metric variation.
  • C.1 RIEMANNIAN MANIFOLD: THE CONTINUOUS SETTING: A Riemannian manifold is a smooth topological manifold equipped with a positive-definite metric tensor that determines tangent-vector lengths, angles, and volumes.In local coordinates, the metric is represented by the matrix G(p) = [gij(p)], and f(p) = 1 2 log det G(p) is the logarithmic volume density.
  • C.2 LEVI-CIVITA CONNECTION AND PARALLEL TRANSPORT: The Levi-Civita connection is metric-compatible and torsion-free, and parallel transport along a smooth curve is a linear isometry preserving inner products.Parallel transport is defined by ∇˙γ(t)V(t) = 0.
  • C.3 CURVATURE AND HOLONOMY: Curvature measures failure of path-independent parallel transport: zero curvature makes transport depend only on endpoints, while loop transport is represented by holonomy H(C).If H(C) = id for all loops, curvature vanishes and the manifold is flat.

C.6 CARTAN’S METHOD OF MOVING FRAME … D.1 MULTI-DOMAIN PRE-TRAINING

The paper connects Cartan’s moving-frame geometry to GraphGlue’s learned manifold framework, which constructs local geometry, aligns sampled tangent spaces, and enforces cycle consistency. Its multi-domain training processes graph-level samples through uniformly mixed batches, contrastive locality recognition, EMA prototypes, and staged manifold construction.

  • C.6 CARTAN’S METHOD OF MOVING FRAME: Cartan’s moving-frame method characterizes local geometry with frames and extends it to a global manifold, but its deep-learning theory remains largely unexplored.
  • C.7 CONNECTION TO OUR FRAMEWORK: GraphGlue assumes a smooth GNN embedding contains a low-dimensional submanifold whose intrinsic geometry encodes generalizable graph-data rules.
  • C.7 CONNECTION TO OUR FRAMEWORK: The Adaptive Frame Bank samples local tangent spaces, while optimal isometric alignment approximates Levi-Civita transport between sampled points.The alignment is identified with Theorem 5.6 in the supplied passage.
  • C.7 CONNECTION TO OUR FRAMEWORK: Cycle-consistency loss enforces trivial holonomy, thereby mimicking flatness in the learned manifold.
  • D ALGORITHMS: Algorithm 1 uses staged training: mixed batches first construct local representations with contrastive and prototype losses, then build a global manifold skeleton using cross-dataset KNN.Prototype loss begins after warm-up epochs, and prototypes are updated with the learned representations.
  • D.1 MULTI-DOMAIN PRE-TRAINING: The data loader converts each sample to a graph, using 2-hop neighborhood ego-subgraphs per Reddit node while retaining the global edge index.
  • D.1 MULTI-DOMAIN PRE-TRAINING: A mixture loader uniformly samples each batch from all source graph datasets to support multi-domain pre-training.
  • D.1 MULTI-DOMAIN PRE-TRAINING: Training combines graph contrastive learning for dataset-specific locality recognition with EMA-updated Riemannian prototypes and post-warm-up sample–prototype contrastive learning.For efficient implementation, adjacent-edge pairs replace triangle sampling because triangle sampling is costly, especially on large graphs.

D.2 COMPLEXITY ANALYSIS … E.1 GRAPH FOUNDATION MODELS

GraphGlue has explicit pretraining and adaptation complexity that scales linearly with graph size, while its computational and memory costs are compared with other graph few-shot methods. The related-work discussion situates these methods within graph foundation models and cross-domain transfer via text-attributed graphs.

  • D.2 COMPLEXITY ANALYSIS: GraphGlue’s pretraining cost is O(B(|V| + |E| + M^2 + K)d + TsM), while per-graph adaptation costs O((|V| + |E|)d + K(d + M) + TsM).B is batch size; |V| and |E| are average nodes and edges; d is hidden dimension; M is manifold dimension; K and Ts are module-specific quantities.
  • D.2 COMPLEXITY ANALYSIS: GraphGlue scales linearly with graph size and was pretrained on large-scale datasets including ogbn-arxiv and Reddit.The stated complexity concerns both pretraining and per-graph adaptation costs.
  • D.3 COMPLEXITY COMPARISON WITH OTHER GFMS: GraphGlue’s total computational cost is compared with other graph few-shot learning methods across pretraining and adaptation phases.The comparison results are summarized in Table 4.
  • D.3 COMPLEXITY COMPARISON WITH OTHER GFMS: The compared methods introduce distinct complexity factors, including full attention, tree construction, retrieved edges, structure and prompt tokens, virtual coordinators, and dense adjacency refinement.These factors correspond respectively to PRODIGY, GFT, RAGraph, SAMGPT, GCOPE, and MDGFM.
  • D.3 COMPLEXITY COMPARISON WITH OTHER GFMS: Memory costs are compared between GraphGlue, GCOPE, and MDGFM on six datasets under a 512 batch size, [10, 10] neighbor sampler size, and d = 512.The six incrementally included datasets are ogbn-arxiv, computers, FB15k-237, Reddit, PROTEINS, and HIV; GPU memory is reported in GB.
  • E.1 GRAPH FOUNDATION MODELS: Graph foundation models target pre-trainable, general-purpose deep learning architectures for graph-structured data.The related-work discussion also describes extending LLMs to text-attributed graphs for cross-domain transfer learning through textual descriptions.

E.2 MULTI-DOMAIN GRAPH PRE-TRAINING … F.2 BASELINES

The paper motivates multi-domain graph pre-training as a way to integrate knowledge across semantically heterogeneous domains, while framing adaptation, Riemannian representation learning, datasets, and baseline models as key empirical components. Its experiments use 12 diverse benchmark datasets and compare against supervised GNNs, self-supervised GNNs, and graph foundation models.

  • E.2 MULTI-DOMAIN GRAPH PRE-TRAINING: Multi-domain graph pre-training seeks shared or invariant knowledge across diverse domains, but knowledge integration and transfer remain theoretically challenging.GNN pre-training uses generative or contrastive self-supervised learning to capture intrinsic semantics from unlabeled data.
  • E.3 GRAPH FINE-TUNING AND PROMPT LEARNING: Downstream adaptation of pre-trained graph models is commonly organized around graph fine-tuning and prompt learning.The supplied passage describes fine-tuning as adapting model behavior with limited target-domain data, including strategies that update the entire model or preserve most parameters.
  • E.4 RIEMANNIAN GRAPH REPRESENTATION LEARNING: Riemannian manifolds provide an alternative to Euclidean spaces for graph representation learning, with existing models often tailored to specific tasks or manifolds.Examples include hyperbolic, spherical, and symmetric positive definite manifolds.
  • F.1 DATASET DESCRIPTION: The experiments evaluate the model on 12 benchmark datasets spanning citation, co-purchase, and social-network domains.Examples include PubMed and Arxiv for academic-paper classification, and Computers and Photo for product-category prediction.
  • F.2 BASELINES: Baselines are grouped into Supervised GNNs, Self-Supervised GNNs, and Graph Foundation Models.These categories cover models trained from scratch, models pre-trained on unlabeled graphs and fine-tuned, and large-scale models pre-trained on diverse datasets.
  • F.2 BASELINES: Supervised baselines include GCN, GraphSAGE, and GIN, representing neighborhood aggregation, scalable inductive learning, and a powerful graph-classification model.GraphSAGE uses neighborhood sampling, while GIN is commonly used as a supervised baseline for graph classification tasks.
  • F.2 BASELINES: Self-supervised baselines include DGI, GraphMAE, and GCC, which use mutual-information maximization, masked-feature reconstruction, and subgraph instance discrimination, respectively.GCC is designed to capture transferable structural representations across multiple networks.
  • F.2 BASELINES: Graph foundation-model baselines include PRODIGY, GFT, RAGraph, SAMGPT, GCOPE, and MDGFM, covering prompting, computation-tree reconstruction, retrieval augmentation, structure tokens, coordinator nodes, and topology alignment.SAMGPT targets multi-domain pretraining and cross-domain adaptation, while GCOPE and MDGFM address unified representation and domain-invariant transfer.

F.3 IMPLEMENTATION NOTES · G ADDITIONAL RESULTS · G.1 SUPPLEMENTARY RESULTS

The paper specifies a leave-one-out, few-shot cross-domain evaluation and concrete pre-training and transfer configurations. Supplementary material adds cross-domain and intra-domain results, ablations, and manifold visualization.

  • F.3 IMPLEMENTATION NOTES: Models are pre-trained on five source datasets and evaluated on one held-out target using leave-one-out cross-domain evaluation.Self-Supervised GNNs and Graph Foundation Models are pre-trained, whereas supervised GNNs train from scratch on the target task.
  • F.3 IMPLEMENTATION NOTES: All downstream evaluations use few-shot fine-tuning with k labeled samples per class from the target dataset.
  • F.3 IMPLEMENTATION NOTES: Pre-training uses 2-hop ego-graphs with 10 neighbors per hop, a 2-layer GCN, M = 32 virtual nodes, k = 15, dropout 0.1, and learning rate 1e−4.The model input dimension is 128, and k = 15 is also used for KNN construction in mixed-data and multi-graph training.
  • F.3 IMPLEMENTATION NOTES: Few-shot transfer tunes learning rate, dropout rate, prototype-target KNN number k, and balance coefficient λ across the reported datasets.Node and graph classification use a linear classifier head, while link classification uses a bilinear layer.
  • G ADDITIONAL RESULTS: Supplementary experiments report comprehensive few-shot results for both cross-domain and intra-domain transfer.These results are presented in Tables 18–21.
  • G.1 SUPPLEMENTARY RESULTS: An ablation study evaluates the effectiveness of the model’s key components.The ablation results are reported in Table 22.
  • G.1 SUPPLEMENTARY RESULTS: An additional visualization projects the 512-dimensional pre-trained embeddings into 3-D using t-SNE.The visualization appears in Figure 8.

G.2 COMPREHENSIVE ABLATION STUDY … K BROADER IMPACT

The paper evaluates GraphGlue through component ablations, hyperparameter and heterophilic-graph studies, manifold-gluing visualization, reproducibility checks, ethics compliance, and broader-impact discussion. These sections report component effectiveness, methodological transparency, and the framework’s societal implications.

  • G.2 COMPREHENSIVE ABLATION STUDY: GraphGlue ablations remove EMA, prototype loss, or Riemannian MoE to test each component, with both 1-shot and 5-shot results demonstrating their effectiveness.The variants replace EMA with batch averaging, omit prototype loss, or replace Riemannian MoE with typical prompting.
  • G.3 HYPERPARAMETER SENSITIVITY ANALYSIS: The AOF sensitivity analysis varies neighborhood size k and node count M in (k, M)-sparse perturbation across 1-shot and 5-shot settings.The reported k analysis fixes M = 32.
  • G.4 RESULTS ON HETEROPHILIC GRAPHS: GraphGlue is evaluated on Amazon-ratings, Roman-empire, Texas, and Wisconsin as benchmarking heterophilic graphs under multiple shot settings.The experiments use pretraining on six datasets in one table and eight datasets, including Amazon-Ratings and Roman-Empire, in another.
  • G.5 VISUALIZATION OF MANIFOLD GLUING: Manifold gluing constructs local geometry with (k, M)-sparse perturbation, joins similar graph grids through overlapping regions, and smooths curvature into a unified manifold.The joining process uses the isometry of Definition 4.4 and holonomy of Equation 5 to avoid stretching or twisting.
  • H REPRODUCIBILITY STATEMENT: The reproducibility statement says theoretical assumptions and complete proofs are provided, empirical protocols and details are disclosed, and code and data are openly available.It also states that experiments use 10 independent runs with means and standard deviations, while evaluation resources are described in Appendix F.
  • I ETHICS STATEMENT: The paper confirms conformity with the ICLR Code of Ethics and reports no human-subject or crowdsourcing research requiring such procedures.It states that original assets are cited, new assets are documented, and LLMs were used only to polish writing.
  • K BROADER IMPACT: The work unifies multi-domain graph pre-training with differential geometry through neural manifold gluing, supporting systematic knowledge-transfer analysis and broader graph applicability.The authors identify transferability and generality as positive societal impacts and report no negative impacts requiring specific emphasis.
Loading 2603.00618v1…