Reliable Rule-Guided Augmentation for Knowledge Graph Completion
Abstract
1. Introduction
- 1.
- A model-agnostic interface transfers symbolic evidence into KGE training without replacing the embedding predictor with a rule reasoner.
- 2.
- Target-relation-guided sampling and candidate-level reliability distinguish groundings by support, type validity, path consistency, and redundancy.
- 3.
- Reliability-dependent weights control the contribution of each inferred triple to the embedding objective.
- 4.
- Experiments with five scoring functions evaluate link prediction, parameter sensitivity, component ablation, and computational cost on WN18RR and FB15k-237.
2. Related Work
2.1. Knowledge Graph Completion
2.2. Graph Sampling and Augmentation
2.3. Rule Learning and Reasoning
3. Method
3.1. Preliminaries
3.1.1. Knowledge Graph
3.1.2. Task
3.1.3. Multi-Hop Neighbors
3.1.4. Logic Rules
3.2. Rule-Based Augmentation
3.2.1. Rule Induction
| Algorithm 1 Path sampling |
|
| Algorithm 2 Rule induction |
|
3.2.2. Rule Inference
| Algorithm 3 Rule inference |
|
3.3. Encoder and Objective
4. Experiments
4.1. Setup
4.1.1. Datasets
4.1.2. Augmentation Protocol
4.1.3. Baselines
- ComplEx: ComplEx [12] extends bilinear scoring to complex embeddings, allowing both symmetric and antisymmetric relations,where and denote the real and imaginary components.
- ConvE: ConvE [13] applies a multilayer convolutional network to entity and relation embeddings,where ⋆ is convolution, is the convolutional kernel, W is a parameter matrix, and vec flattens its input.
- AutoBLM: AutoBLM [37] uses neural architecture search to select a bilinear scoring function.
4.1.4. Metrics
4.1.5. Implementation
4.2. Link Prediction
4.3. Additional Comparison Analysis
4.4. Parameter Sensitivity
4.4.1. Rule Threshold
4.4.2. Augmentation Weight
4.5. Ablation
4.6. Efficiency
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AnyBURL | Anytime bottom-up rule learning |
| AutoBLM | Automated bilinear scoring function search model |
| CompGCN | Composition-based graph convolutional network |
| ERMLP | Entity–relation multilayer perceptron |
| FB15k | Freebase 15k dataset |
| FB15k-237 | Freebase 15k-237 dataset |
| GAT | Graph attention network |
| GNN | Graph neural network |
| GPU | Graphics processing unit |
| HRAN | Heterogeneous relation attention network |
| KG | Knowledge graph |
| KGC | Knowledge graph completion |
| KGE | Knowledge graph embedding |
| LINE | Large-scale information network embedding |
| LR | LeakyReLU |
| LTE | Linearly transformed entity embedding |
| MR | Mean rank |
| MRR | Mean reciprocal rank |
| NCE | Noise contrastive estimation |
| NTN | Neural tensor network |
| R-GCN | Relational graph convolutional network |
| RLvLR | Rule learning via learning representation |
| SampledNCE | Sampled noise contrastive estimation |
| WN18 | WordNet 18 dataset |
| WN18RR | WordNet 18RR dataset |
References
- Zhu, G.; Iglesias, C.A. Exploiting semantic similarity for named entity disambiguation in knowledge graphs. Expert Syst. Appl. 2018, 101, 8–24. [Google Scholar] [CrossRef] [Scilit]
- Hu, S.; Zou, L.; Yu, J.X.; Wang, H.; Zhao, D. Answering natural language questions by subgraph matching over knowledge graphs. IEEE Trans. Knowl. Data Eng. 2017, 30, 824–837. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Zhang, F.; Wang, J.; Zhao, M.; Li, W.; Xie, X.; Guo, M. Exploring high-order user preference on the knowledge graph for recommender systems. ACM Trans. Inf. Syst. (TOIS) 2019, 37, 1–26. [Google Scholar] [CrossRef] [Scilit]
- Rosa, R.L.; Schwartz, G.M.; Ruggiero, W.V.; Rodríguez, D.Z. A knowledge-based recommendation system that includes sentiment analysis and deep learning. IEEE Trans. Ind. Inform. 2018, 15, 2124–2135. [Google Scholar] [CrossRef] [Scilit]
- Marino, K.; Salakhutdinov, R.; Gupta, A. The more you know: Using knowledge graphs for image classification. arXiv 2016, arXiv:1612.04844. [Google Scholar]
- Bollacker, K.; Evans, C.; Paritosh, P.; Sturge, T.; Taylor, J. Freebase: A collaboratively created graph database for structuring human knowledge. In Proceedings of the 2008 ACM SIGMOD International Conference on Management of Data; Association for Computing Machinery: New York, NY, USA, 2008; pp. 1247–1250. [Google Scholar]
- Miller, G.A. WordNet: An Electronic Lexical Database; MIT Press: Cambridge, MA, USA, 1998. [Google Scholar]
- Vrandečić, D.; Krötzsch, M. Wikidata: A free collaborative knowledgebase. Commun. ACM 2014, 57, 78–85. [Google Scholar]
- Mahdisoltani, F.; Biega, J.; Suchanek, F. Yago3: A knowledge base from multilingual wikipedias. In Proceedings of the 7th Biennial Conference on Innovative Data Systems Research, CIDR Conference, Asilomar, CA, USA, 4–7 January 2014. [Google Scholar]
- Wang, Q.; Mao, Z.; Wang, B.; Guo, L. Knowledge graph embedding: A survey of approaches and applications. IEEE Trans. Knowl. Data Eng. 2017, 29, 2724–2743. [Google Scholar] [CrossRef] [Scilit]
- Bordes, A.; Usunier, N.; Garcia-Duran, A.; Weston, J.; Yakhnenko, O. Translating embeddings for modeling multi-relational data. Adv. Neural Inf. Process. Syst. 2013, 26, 2787–2795. [Google Scholar]
- Trouillon, T.; Welbl, J.; Riedel, S.; Gaussier, É.; Bouchard, G. Complex embeddings for simple link prediction. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2016; pp. 2071–2080. [Google Scholar]
- Dettmers, T.; Minervini, P.; Stenetorp, P.; Riedel, S. Convolutional 2d knowledge graph embeddings. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Palo Alto, CA, USA, 2018; Volume 32. [Google Scholar]
- Schlichtkrull, M.; Kipf, T.N.; Bloem, P.; Berg, R.v.d.; Titov, I.; Welling, M. Modeling relational data with graph convolutional networks. In Proceedings of the European Semantic Web Conference; Springer: Berlin/Heidelberg, Germany, 2018; pp. 593–607. [Google Scholar]
- Li, Z.; Liu, H.; Zhang, Z.; Liu, T.; Xiong, N.N. Learning Knowledge Graph Embedding with Heterogeneous Relation Attention Networks. IEEE Trans. Neural Netw. Learn. Syst. 2022, 33, 3961–3973. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, J.; Shi, W.; Cao, X.; Chen, J.; Lei, W.; Zhang, F.; Wu, W.; He, X. DisenKGAT: Knowledge graph embedding with disentangled graph attention network. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management; Association for Computing Machinery: New York, NY, USA, 2021; pp. 2140–2149. [Google Scholar]
- Yang, Z.; Ding, M.; Zhou, C.; Yang, H.; Zhou, J.; Tang, J. Understanding negative sampling in graph representation learning. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining; Association for Computing Machinery: New York, NY, USA, 2020; pp. 1666–1676. [Google Scholar]
- Meilicke, C.; Chekol, M.W.; Ruffinelli, D.; Stuckenschmidt, H. Anytime Bottom-Up Rule Learning for Knowledge Graph Completion. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence; AAAI Press: Palo Alto, CA, USA, 2019; pp. 3137–3143. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, F.; Yang, Z.; Cohen, W.W. Differentiable Learning of Logical Rules for Knowledge Base Reasoning. In Proceedings of the Advances in Neural Information Processing Systems; Neural Information Processing Systems Foundation, Inc.: San Diego, CA, USA, 2017; Volume 30. [Google Scholar]
- Li, G.; Sun, Z.; Qian, L.; Guo, Q.; Hu, W. Rule-Based Data Augmentation for Knowledge Graph Embedding. AI Open 2021, 2, 186–196. [Google Scholar] [CrossRef] [Scilit]
- Shomer, H.; Jin, W.; Wang, W.; Tang, J. Toward Degree Bias in Embedding-Based Knowledge Graph Completion. In Proceedings of the ACM Web Conference 2023; Association for Computing Machinery: New York, NY, USA, 2023; pp. 705–715. [Google Scholar] [CrossRef] [Scilit]
- Yang, B.; Yih, W.t.; He, X.; Gao, J.; Deng, L. Embedding entities and relations for learning and inference in knowledge bases. arXiv 2014, arXiv:1412.6575. [Google Scholar]
- Vashishth, S.; Sanyal, S.; Nitin, V.; Talukdar, P. Composition-based multi-relational graph convolutional networks. arXiv 2019, arXiv:1911.03082. [Google Scholar]
- Zhang, Z.; Wang, J.; Ye, J.; Wu, F. Rethinking graph convolutional networks in knowledge graph completion. In Proceedings of the ACM Web Conference 2022; Association for Computing Machinery: New York, NY, USA, 2022; pp. 798–807. [Google Scholar]
- Mikolov, T.; Chen, K.; Corrado, G.; Dean, J. Efficient estimation of word representations in vector space. arXiv 2013, arXiv:1301.3781. [Google Scholar]
- Tang, J.; Qu, M.; Wang, M.; Zhang, M.; Yan, J.; Mei, Q. Line: Large-scale information network embedding. In Proceedings of the 24th International Conference on World Wide Web; International World Wide Web Conferences Steering Committee: Geneva, Switzerland, 2015; pp. 1067–1077. [Google Scholar]
- Perozzi, B.; Al-Rfou, R.; Skiena, S. Deepwalk: Online learning of social representations. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Association for Computing Machinery: New York, NY, USA, 2014; pp. 701–710. [Google Scholar]
- Grover, A.; Leskovec, J. node2vec: Scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Association for Computing Machinery: New York, NY, USA, 2016; pp. 855–864. [Google Scholar]
- Gilmer, J.; Schoenholz, S.S.; Riley, P.F.; Vinyals, O.; Dahl, G.E. Neural message passing for quantum chemistry. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2017; pp. 1263–1272. [Google Scholar]
- Veličković, P.; Fedus, W.; Hamilton, W.L.; Liò, P.; Bengio, Y.; Hjelm, R.D. Deep graph infomax. arXiv 2018, arXiv:1809.10341. [Google Scholar]
- Cai, C.; Wang, D.; Wang, Y. Graph coarsening with neural networks. arXiv 2021, arXiv:2102.01350. [Google Scholar]
- Jin, W.; Zhao, L.; Zhang, S.; Liu, Y.; Tang, J.; Shah, N. Graph condensation for graph neural networks. arXiv 2021, arXiv:2110.07580. [Google Scholar]
- Liu, X.; Sun, D.; Wei, W. Alleviating the over-smoothing of graph neural computing by a data augmentation strategy with entropy preservation. Pattern Recognit. 2022, 132, 108951. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Wang, Z.; Wang, K.; Omran, P.G.; Li, J. Rule Learning over Knowledge Graphs: A Review. Trans. Graph Data Knowl. 2023, 1, 7:1–7:23. [Google Scholar] [CrossRef]
- Liu, H.; Wang, Z.; Wang, K.; Zhang, X.; Feng, Z. Transfer Rule Learning over Large Knowledge Graphs. In Proceedings of the ACM on Web Conference 2025, New York, NY, USA, 28 April–2 May 2025; WWW ’25; Association for Computing Machinery: New York, NY, USA, 2025; pp. 2135–2143. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Zhang, J.; Feng, J.; Chen, Z. Knowledge graph embedding by translating on hyperplanes. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Palo Alto, CA, USA, 2014; Volume 28. [Google Scholar]
- Zhang, Y.; Yao, Q.; Kwok, J.T. Bilinear scoring function search for knowledge graph learning. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 1458–1473. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Meilicke, C.; Fink, M.; Wang, Y.; Ruffinelli, D.; Gemulla, R.; Stuckenschmidt, H. Fine-Grained Evaluation of Rule- and Embedding-Based Systems for Knowledge Graph Completion. In Proceedings of the International Semantic Web Conference; Springer International Publishing: Berlin/Heidelberg, Germany, 2018; pp. 3–20. [Google Scholar]
- Omran, P.G.; Wang, K.; Wang, Z. Scalable Rule Learning via Learning Representation. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence; AAAI Press: Palo Alto, CA, USA, 2018; pp. 2149–2155. [Google Scholar] [CrossRef] [Scilit] [PubMed]






| Datasets | WN18RR | FB15k-237 | |
|---|---|---|---|
| # Entities | 40,493 | 14,541 | |
| # Relations | 11 | 237 | |
| # Edges | Train | 86,835 | 272,115 |
| Valid | 3034 | 17,535 | |
| Test | 3134 | 20,466 | |
| Total | 93,003 | 310,116 | |
| # Mean Degree | 2.12 | 18.71 | |
| Methods | WN18RR | FB15k-237 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| MR | MRR | Hits | MR | MRR | Hits | ||||||
| @1 | @3 | @10 | @1 | @3 | @10 | ||||||
| Embeddings | TransE [11] | 2879 | 0.198 | 0.047 | 0.306 | 0.476 | 189 | 0.329 | 0.240 | 0.364 | 0.507 |
| DistMult [22] | 6024 | 0.434 | 0.402 | 0.451 | 0.498 | 269 | 0.331 | 0.244 | 0.363 | 0.504 | |
| ConvE [13] | 2520 | 0.476 | 0.447 | 0.492 | 0.540 | 180 | 0.358 | 0.266 | 0.393 | 0.545 | |
| ComplEx [12] | 4303 | 0.454 | 0.418 | 0.469 | 0.526 | 250 | 0.342 | 0.252 | 0.376 | 0.522 | |
| AutoBLM [37] | 3198 | 0.461 | 0.423 | 0.476 | 0.536 | 171 | 0.362 | 0.269 | 0.399 | 0.550 | |
| GNNs | CompGCN+TransE [23] | 3182 | 0.206 | 0.064 | 0.281 | 0.502 | 205 | 0.335 | 0.247 | 0.369 | 0.511 |
| CompGCN+DistMult [23] | 4559 | 0.430 | 0.395 | 0.439 | 0.513 | 200 | 0.342 | 0.252 | 0.372 | 0.520 | |
| CompGCN+ConvE [23] | 3065 | 0.469 | 0.433 | 0.480 | 0.543 | 245 | 0.351 | 0.254 | 0.386 | 0.535 | |
| LTE+TransE [24] | 3290 | 0.211 | 0.022 | 0.362 | 0.521 | 182 | 0.334 | 0.241 | 0.347 | 0.519 | |
| LTE+DistMult [24] | 4485 | 0.437 | 0.403 | 0.447 | 0.517 | 238 | 0.335 | 0.246 | 0.360 | 0.517 | |
| LTE+ConvE [24] | 3434 | 0.472 | 0.436 | 0.485 | 0.544 | 249 | 0.352 | 0.262 | 0.385 | 0.533 | |
| Rules | AnyBURL [18] | – | ≥0.470 | 0.441 | – | 0.552 | – | ≥0.310 | 0.233 | – | 0.486 |
| RuleN [38] | – | – | 0.427 | – | 0.536 | – | – | 0.182 | – | 0.420 | |
| RLvLR [39] | – | – | – | – | – | – | 0.240 | – | – | 0.393 | |
| Augmentation Baselines | KnowAug (ComplEx) [20] | – | 0.453 | 0.414 | – | 0.535 | – | 0.331 | 0.239 | – | 0.516 |
| KG-Mixup (ConvE) [21] | – | – | – | – | – | – | 0.343 | 0.250 | – | 0.531 | |
| Augmentations (Ours) | TransE–Aug | 2706 | 0.208 | 0.043 | 0.316 | 0.489 | 177 | 0.331 | 0.242 | 0.375 | 0.525 |
| DistMult–Aug | 5879 | 0.435 | 0.412 | 0.469 | 0.510 | 232 | 0.334 | 0.247 | 0.369 | 0.517 | |
| ConvE–Aug | 2032 | 0.491 | 0.457 | 0.503 | 0.553 | 173 | 0.362 | 0.271 | 0.401 | 0.552 | |
| ComplEx–Aug | 3042 | 0.461 | 0.423 | 0.472 | 0.541 | 245 | 0.346 | 0.256 | 0.379 | 0.527 | |
| AutoBLM–Aug | 2441 | 0.481 | 0.451 | 0.496 | 0.545 | 156 | 0.372 | 0.278 | 0.412 | 0.557 | |
| Decoder | WN18RR | FB15k-237 | ||||
|---|---|---|---|---|---|---|
| Base | Aug | MRR | Base | Aug | MRR | |
| TransE | ||||||
| DistMult | ||||||
| ConvE | ||||||
| ComplEx | ||||||
| AutoBLM | ||||||
| Setting | Dataset | Method | MRR | Hits@1 | Hits@10 |
|---|---|---|---|---|---|
| KnowAug | WN18RR | ComplEx † | 0.437 | 0.393 | 0.526 |
| Local Base | 0.436 | 0.392 | 0.526 | ||
| KnowAug–ComplEx † | 0.453 | 0.414 | 0.535 | ||
| Ours–ComplEx | 0.456 | 0.424 | 0.543 | ||
| FB15k-237 | ComplEx † | 0.317 | 0.224 | 0.504 | |
| Local Base | 0.319 | 0.227 | 0.506 | ||
| KnowAug–ComplEx † | 0.331 | 0.239 | 0.516 | ||
| Ours–ComplEx | 0.340 | 0.252 | 0.522 | ||
| KG-Mixup | FB15k-237 | ConvE † | 0.330 | 0.240 | 0.512 |
| Local Base | 0.331 | 0.242 | 0.513 | ||
| KG-Mixup–ConvE † | 0.343 | 0.250 | 0.531 | ||
| Ours–ConvE | 0.351 | 0.258 | 0.537 | ||
| Additional ConvE | WN18RR | Local Base | 0.436 | 0.421 | 0.523 |
| Ours–ConvE | 0.477 | 0.448 | 0.541 |
| Variant | MR | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|---|
| Base KGE | 171 | 0.362 | 0.269 | 0.399 | 0.550 |
| Random multi-hop augmentation | 197 | 0.328 | 0.236 | 0.360 | 0.510 |
| Rule-induced augmentation without candidate control | 174 | 0.349 | 0.255 | 0.384 | 0.538 |
| Complete framework | 156 | 0.372 | 0.278 | 0.412 | 0.557 |
| Dataset | Variant | Offline/Base | Training/Base | Total/Base | MRR |
|---|---|---|---|---|---|
| WN18RR | Base KGE | – | 1.00 | 1.00 | – |
| Uniform augmentation | 0.37 | 0.46 | 0.83 | ||
| Ours | 0.37 | 0.49 | 0.86 | ||
| FB15k-237 | Base KGE | – | 1.00 | 1.00 | – |
| Uniform augmentation | 0.31 | 0.82 | 1.14 | ||
| Ours | 0.31 | 0.82 | 1.14 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, Q.; Lv, Y.; Feng, X.; Wei, W.; Tang, S. Reliable Rule-Guided Augmentation for Knowledge Graph Completion. Mathematics 2026, 14, 2984. https://doi.org/10.3390/math14162984
Li Q, Lv Y, Feng X, Wei W, Tang S. Reliable Rule-Guided Augmentation for Knowledge Graph Completion. Mathematics. 2026; 14(16):2984. https://doi.org/10.3390/math14162984
Chicago/Turabian StyleLi, Qingsong, You Lv, Xiangnan Feng, Wei Wei, and Shaoting Tang. 2026. "Reliable Rule-Guided Augmentation for Knowledge Graph Completion" Mathematics 14, no. 16: 2984. https://doi.org/10.3390/math14162984
APA StyleLi, Q., Lv, Y., Feng, X., Wei, W., & Tang, S. (2026). Reliable Rule-Guided Augmentation for Knowledge Graph Completion. Mathematics, 14(16), 2984. https://doi.org/10.3390/math14162984

