Source-linked AI summary
A Generative Adaptive Replay Continual Learning Model for Temporal Knowledge Graph Reasoning
Zhiyu Zhang, Wei Chen, Youfang Lin, Huaiyu Wan
TL;DR
Continual-learning methods for temporal knowledge graph reasoning must preserve historical semantics while handling conflicts between past and current data distributions. DGAR uses context-based replay and diffusion-guided distribution generation with adaptive integration. It outperforms baselines across evaluated datasets and reports gains on current and historical tasks.
Problem
Existing CL-based TKGR methods replay individual facts without fully preserving historical context and overlook conflicts between historical and current distributions.
Method
DGAR samples historical context prompts, uses a pre-trained diffusion model guided by the TKGR model to generate historical entity distributions, and adaptively replays them layer by layer.
Results
DGAR consistently outperforms all baselines, improving current-task MRR by 4.01% on average and historical-task MRR by 8.23% and Hits@10 by 9.79%.
Takeaways & Limitations
Across various datasets, DGAR’s combined techniques achieve strong performance while preserving historical knowledge.
Takeaways & Limitations
New entities and relations use Xavier initialization without dedicated modeling, which may constrain learning when their interrelations with previously learned knowledge are complex.
Abstract
from arXiv · showhide
Recent Continual Learning (CL)-based Temporal Knowledge Graph Reasoning (TKGR) methods focus on significantly reducing computational cost and mitigating catastrophic forgetting caused by fine-tuning models with new data. However, existing CL-based TKGR methods still face two key limitations: (1) They usually one-sidedly reorganize individual historical facts, while overlooking the historical context essential for accurately understanding the historical semantics of these facts; (2) They preserve historical knowledge by simply replaying historical facts, while ignoring the potential conflicts between historical and emerging facts. In this paper, we propose a Deep Generative Adaptive Replay (DGAR) method, which can generate and adaptively replay historical entity distribution representations from the whole historical context. To address the first challenge, historical context prompts as sampling units are built to preserve the whole historical context information. To overcome the second challenge, a pre-trained diffusion model is adopted to generate the historical distribution. During the generation process, the common features between the historical and current distributions are enhanced under the guidance of the TKGR model. In addition, a layer-by-layer adaptive replay mechanism is designed to effectively integrate historical and current distributions. Experimental results demonstrate that DGAR significantly outperforms baselines in reasoning and mitigating forgetting.
1 Introduction
Existing continual-learning methods for temporal knowledge graph reasoning struggle to preserve historical semantics and handle conflicts between historical and current distributions. DGAR replays historical entity distributions generated from full historical context, and the paper reports stronger reasoning and forgetting-mitigation results than baselines.
- 1 Introduction: Existing CL-based TKGR methods omit historical context when replaying individual facts and overlook conflicts between historical and current distributions.The paper links changing entity neighbors over time to semantic differences and distribution conflicts.
- 1 Introduction: DGAR samples Historical Context Prompts to retain historical context and generate entity-distribution representations from that context.The paper presents these prompts as an alternative to using individual facts as replay sampling units.
- 1 Introduction: Diff-HDG generates historical distribution representations while enhancing features shared with current distributions to mitigate conflicts.The approach uses a diffusion model for generation and emphasizes common features during the process.
- 1 Introduction: Deep Adaptive Replay injects historical entity distributions into current representations layer by layer.The paper describes this mechanism as integrating historical and current distributions.
- 1 Introduction: Extensive experiments on widely used TKGR datasets demonstrate the approach’s superiority, as it consistently outperforms all baseline methods across various evaluation metrics.
2 Related Work
The related-work section reviews temporal knowledge graph reasoning, continual learning for knowledge graphs, and diffusion models. DGAR builds on graph-neural-network reasoning and generative diffusion approaches.
- 2.1 Reasoning on TKGs: TKGR methods include distribution-based, GNN-based, and rule-based approaches; DGAR builds on GNN-based methods because of their strong TKGR performance.Rule-based methods mine temporal logical rules, while some approaches require retraining when new data arrives.
- 2.2 CL for Knowledge Graphs: CL approaches for TKGR use replay and regularization, while TIE’s restrictive regularization can lower overall performance and DEWC is limited to a small number of tasks.Other cited work applies regularization constraints to retain historical knowledge in KGE.
- 2.3 Diffusion Models: Diffusion models generate structured data through reverse denoising, and KG methods apply them by mapping discrete knowledge into continuous space.The paper notes that Gaussian noise is not directly compatible with discrete structures.
3 Preliminaries
The preliminaries define temporal knowledge graph reasoning and its continual-learning setting, then introduce the diffusion process used to model representations. In this setting, snapshots arrive sequentially, and the diffusion model transforms embedded facts through noise and denoising steps.
- 3.1 The Task of TKGR: A TKG is represented as timestamped graph snapshots, and TKGR predicts a missing subject or object entity from a query.Each fact is a subject-relation-object tuple with a timestamp.
- 3.2 Continual Learning for TKGR: In continual TKGR, snapshot tasks arrive as a stream and the model updates parameters sequentially from the previous task’s parameters.The paper focuses on entity representations because entity semantics tend to evolve more than relation semantics.
- 3.3 Denoising Diffusion Probabilistic Model: The diffusion model embeds discrete facts, adds Gaussian noise through a forward Markov chain, then iteratively denoises toward the original representation.The reverse process is parameterized by a model that predicts the previous latent representation from the current one and step.
- 3.3 Denoising Diffusion Probabilistic Model: The paper uses a Transformer for the diffusion model and states that testing data remains unseen during pretraining.The reverse-process parameters are implemented by a deep neural model.
4 The DGAR Method
DGAR builds historical context prompts, generates historical entity distributions, then injects them into current representations for reasoning. Its method covers prompt building, diffusion-enhanced generation, adaptive replay, and model training.
- 4.1 Historical Context Prompt Building: For each query entity, DGAR samples historical context prompts—sets of associated triples—from k distinct timestamps to preserve context while reducing computation and storage.The sampled prompts guide generation of the entity’s historical distribution.
- 4.2 Diffusion-enhanced Historical Distribution Generation: Diff-HDG generates historical distributions from prompt-conditioned neighbor and relation information, enhancing features shared with current distributions and weakening differing features.The current TKGR model scores historical facts to guide generation; θt−1 approximates θt, and entity representations are generated in parallel.
- 4.3 Deep Adaptive Replay: DAR injects historical distributions into current representations at each layer, adaptively balancing old and new knowledge across L evolution units.The method addresses the stated trade-off between complex injection that burdens learning and simplistic injection that loses knowledge.
- 4.4 Model Training: DGAR trains with current-task loss plus a replay loss computed from historical facts in the sampled prompts.The replay loss is included as a regularization term to address historical information loss during current-data optimization; μ is typically set to 1.
5 Experiments
DGAR is evaluated on four TKGR benchmarks against continual-learning baselines, with results covering current and historical tasks, ablations, knowledge retention, and entity-distribution visualizations.
- 5.1 Experimental Setup: The evaluation uses ICE14, ICE18, ICE05-15, and GDELT, measuring MRR and Hits@1/10 on current and average historical test sets.RE-GCN is the base model, and results are averaged over five runs.
- 5.2 Main Results: DGAR improves current-task MRR by 4.01% and historical-task MRR by 8.23% and Hits@10 by 9.79% against the strongest baseline, averaged across evaluated datasets.It also raises historical-task MRR by an average of 11.34% compared with fine-tuning.
- 5.3 Ablation Study: DGAR also outperforms direct fine-tuning on current and historical tasks when extended to TiRGN, supporting scalability across GNN-based TKGR models.Figure 3 presents performance for different base TKGR models.
- 5.3 Ablation Study: Removing the guider or adaptive parameter α lowers performance, while removing the combined replay components causes a clear drop attributed to historical-current distribution conflicts.The GDELT drops are smaller, attributed to shorter temporal gaps and less pronounced distribution shifts.
- 5.4 Effect of Memorizing in CL: DGAR outperforms the best baseline on the mean difference between final-task and earlier-task MRR; positive values indicate reverse transfer, while negative values indicate forgetting.The paper notes reverse transfer for DGAR and ER on ICE14 when the number of tasks is small, consistent with high data correlation in TKGs.
- 5.5 Case Study: Across four stages on ICE14 and ICE18, DGAR’s entity distributions are more general and consistent across time than fine-tuning’s, supporting knowledge retention and reduced forgetting.The case study samples entity feature distributions at successive times and analyzes them with U-MAP.
6 Conclusion
DGAR generates historical entity distributions from contextual prompts using a pretrained diffusion model guided by current model parameters. Its adaptive replay strategy is reported to perform well across datasets.
- Historical Context Prompts integrate contextual information to support generation of historical entity distribution representations.The prompts are designed as sampling units for historical information.
- A pretrained diffusion model generates historical distributions, guided by current model parameters to reinforce common features and minimize distribution conflicts.
- Deep Adaptive Replay derives entity distribution representations with historical knowledge, and the combined techniques achieve outstanding performance across datasets.
7 Limitations
DGAR has limitations in handling newly emerging entities and relations, and its added parameters create risks and complexity despite its forgetting mitigation.
- 7 Limitations: Xavier initialization handles new entities and relations without dedicated modeling, which may constrain learning when they have complex interrelations with previously learned knowledge.The authors identify more sophisticated strategies for new entities and relations in continual-learning scenarios as needed.
- 7 Limitations: Additional learnable parameters may lead DGAR to prioritize new knowledge and compromise retention of previously learned information.The authors describe this as a potential risk, despite DGAR's strong performance in reducing catastrophic forgetting.
- 7 Limitations: The additional parameters increase model complexity, making training and reasoning more cumbersome and requiring a balance between retention and complexity.
8 Ethics Statement
The paper reports its ethics and dataset practices, details adaptive replay and diffusion-model pretraining, and describes its dataset split and continual-learning baselines.
- 8 Ethics Statement: The study reports compliance with the ACL Code of Ethics, use of previously studied datasets without individual privacy data, and possible need for manual inspection of toxic or erroneous TKGR outputs.
- A.1 Details about Deep Adaptive Replay: Adaptive replay balances historical and current entity distributions with a parameter after direct fusion degraded performance.The weighting produces a final representation combining recent and historical information; the method also injects historical distributions into current representations across evolution units without additional parameters.
- A.2 Pre-train for DM: The diffusion model is continually trained: the prior timestep's model initializes the next, and training on the new task produces the model used for subsequent learning.
- B.1 Datasets Details: TKGR facts are split into train, validation, and test sets in an 8:1:1 ratio, with dataset statistics provided in Table 4.
- B.2 Baselines Details: The baseline descriptions cover fine-tuning, event replay, temporal regularization, historical-weight preservation, and IncDE-based ordering and distillation strategies.The paper also incorporates reconstruction loss and embedding regularization from LKGE and adopts its new-entity embedding transfer strategies.
B.3 Implementation Details
The implementation uses fixed embedding, optimization, and model settings, with dataset-specific choices for HCP sample counts.
- B.3 Implementation Details: Across datasets, embedding size is 200, learning rate is 0.001, batch size depends on task facts, Transformer depth is 2, and temperature τ is 0.5.
- B.3 Implementation Details: DGAR uses Adam, with γ=1 for Diff-HDG, L=3 for DAR, and loss coefficient µ=1.
- B.3 Implementation Details: The best HCP sample counts are 35, 25, 40, and 32 for ICE14, ICE18, ICE05-15, and GDELT, respectively; experiments ran on NVIDIA A40.
B.4 Sensitivity Analysis
DGAR outperforms retraining while requiring less time, and sensitivity results indicate that additional recall time slices offer limited gains at added computational cost.
- B.4 Sensitivity Analysis: On ICE14, performance initially improves and then plateaus as k increases; more recall time slices do not significantly improve final performance but increase computational cost.Figure 6 reports historical-task MRR and Hits@3; even five recall time slices outperform all baseline models.
- B.5 Inference Efficiency: The efficiency analysis reports average time per task and historical-task MRR for DGAR at different k values.Here, k is the number of distinct times represented by replay HCPs.
- B.5 Inference Efficiency: DGAR outperforms retraining and requires less time, although it takes more time than baselines because of its more complex operations for retaining historical knowledge.Table 5 reports average task time and historical-task MRR across DGAR k settings, baselines, and retention settings.
B.6 Effect Analysis of Lr
The effect analyses find that DGAR enhances LogCL under continual learning, random HCP selection outperforms nearest-slice selection on historical tasks, and removing Lr significantly harms historical-task performance.
- B.6 Effect Analysis of Lr: Without Lr, historical-task performance drops significantly, while Lr,t alleviates historical-information loss from current-data optimization and Diff-HDG.
- B.7 Random Selection of HCPs: Across different k values, random HCP selection outperforms nearest-time-slice selection on historical tasks.The authors say random selection provides more generalized data, improving test-set performance.
- B.8 Compare Base on LogCL: DGAR significantly enhances LogCL’s reasoning under CL; its full-training performance does not guarantee better reasoning in the CL setting.The authors attribute LogCL’s weaker ICE14 result under CL to overfitting and difficulty maintaining stable learned features as new data arrives.