Retrieval-Augmented Generation for Curated Thematic Corpora: A Critical Survey, Bibliometric Evidence, and the ThemePath-RAG Framework
Abstract
1. Introduction
- We provide a critical scoping survey of RAG research supported by a bibliometric analysis of 2815 Scopus-indexed records.
- We analyze the dominant bibliometric signals in the RAG literature, including document type composition, keyword distribution, thematic evolution, and country–keyword relationships.
- We introduce the concept of curated thematic corpora, where knowledge is already organized through human-authored thematic structures.
- We identify a limitation in current flat and graph-based RAG methods: they do not fully exploit curated thematic paths and often retrieve redundant evidence.
- We propose ThemePath-RAG, a query-aware evidence pruning framework for grounded question answering over pre-structured thematic corpora.
- We position Qur’anic question answering as a primary case study and discuss generalization to legal, medical, educational, and library knowledge systems.
2. Methodology
2.1. Research Design
2.2. Data Source and Search Strategy
TITLE-ABS-KEY(“retrieval augmented generation” OR “retrieval-augmented generation” OR “RAG”) AND PUBYEAR > 2022 AND PUBYEAR < 2027
2.3. Eligibility Criteria
- 1.
- The record was indexed in Scopus.
- 2.
- The record was published between 2023 and 2026.
- 3.
- The title, abstract, author keywords, or indexed keywords indicated relevance to retrieval-augmented generation, retrieval-enhanced language models, graph-based RAG, knowledge-graph RAG, sparse or hybrid retrieval for LLMs, domain-specific RAG, RAG evaluation, or hallucination mitigation in RAG systems.
- 4.
- The record contained sufficient bibliographic metadata for analysis, including title, publication year, document type, source title, and either abstract or keyword information.
- 5.
- The record belonged to scholarly document categories such as journal articles, conference papers, or review papers.
2.4. PRISMA-Informed Workflow
2.5. Bibliometric Analysis Procedure
- Dataset and metadata quality analysis: This analysis summarized the number of records, publication years, countries, institutions, sources, authors, document types, total citations, average citations per document, and metadata completeness.
- Document type analysis: This analysis examined whether the field was dominated by conference papers, journal articles, or review papers.
- Keyword analysis: Author keywords, indexed keywords, and Keywords Plus were analyzed to identify the dominant research vocabulary in RAG studies.
- Thematic evolution analysis: This analysis examined changes in dominant topics from 2023 to 2026.
- Country–keyword relationship analysis: This analysis examined whether major RAG topics were concentrated in particular countries or shared across leading research communities.
- Bibliometric gap synthesis: The bibliometric results were interpreted to assess whether curated thematic paths, canonical evidence units, and thematic path-guided evidence pruning appeared as consolidated research themes.
2.6. Critical Synthesis Procedure
How do existing RAG methods represent knowledge before retrieval, and what limitations emerge when the corpus already contains expert-authored thematic paths and canonical evidence units?
2.7. Methodological Limitations
3. Bibliometric Evidence of RAG Research Trends
3.1. Dataset and Metadata Quality
3.2. Publication Maturity and Document Type Composition
3.3. Dominant Research Vocabulary
3.4. Thematic Evolution
3.5. Geographical Distribution of Dominant Keywords
3.6. Bibliometric Gap Synthesis
4. Critical Taxonomy of RAG Methods
4.1. Representative Scopus Literature Coverage
4.2. Flat Chunk-Based RAG
4.3. Sparse and Hybrid RAG
4.4. Graph-Based RAG
4.5. Agentic and Multi-Hop RAG
4.6. Memory-Inspired and Path-Based RAG
4.7. Domain-Specific RAG
5. Structural Gap: Curated Thematic Corpora
5.1. Definition
5.2. Examples
5.3. Why Existing RAG Assumptions Are Insufficient
6. Comparison with Existing Methods
7. ThemePath-RAG Framework
7.1. Conceptual Overview
7.2. Stage 1: Thematic Path Retrieval
7.3. Stage 2: Candidate Evidence Expansion
7.4. Stage 3: Query-Aware Evidence Scoring
7.5. Stage 4: Evidence Pruning
7.6. Stage 5: Citation-Grounded Answer Generation
7.7. Proof-of-Concept Instantiation
| Algorithm 1 ThemePath-RAG Retrieval and Context Relevance Evaluation |
|
8. Proof-of-Concept Evaluation in Qur’anic Question Answering
8.1. Evaluation Objective
How does the ThemePath-RAG prototype compare with a Vector RAG baseline in retrieving contexts relevant to Qur’anic question-answering queries?
8.2. Evaluation Protocol
8.3. Context Relevance Results
8.4. Qualitative Retrieval Examples
Morals → Destructive Morals → Tyranny → On the Day of Judgment, the Wrongdoers Will Be in Fear.
The Qur’an states that wrongdoers will be fearful of what they have earned and that the consequence of their deeds will inevitably befall them (Qur’an 42:22). It also warns that those who commit crimes against believing men and women and do not repent face the punishment of Hell and the Burning Fire (Qur’an 85:10). A further warning of woe is given to those who cheat in measure (Qur’an 83:1).
8.5. Interpretation of the Proof-of-Concept Results
8.6. Implications for Further Evaluation
- Vector RAG using dense verse-level retrieval;
- BM25 RAG using lexical verse-level retrieval;
- thematic-path retrieval without query-aware evidence scoring; and
- ThemePath-RAG with weighted lexical, semantic, and path-based scoring followed by global top-k pruning.
9. Promising Cross-Domain Applications of ThemePath-RAG Beyond Qur’anic Question Answering
9.1. Case 1: Library Classification and Digital Library Retrieval
9.2. Case 2: Indonesian Legal Codes and Regulatory Question Answering
9.3. Case 3: Medical Guidelines and Clinical Pathways
9.4. Case 4: Educational Taxonomies and Curriculum-Based QA
9.5. Case 5: Policy, Government, and Compliance Documents
10. Evaluation Agenda and Benchmark Design
10.1. Research Questions
10.2. Dataset Construction
10.3. Baselines
10.4. Metrics
11. Discussion
11.1. Why ThemePath-RAG Is Not Just Hybrid Retrieval
11.2. Relation to HippoRAG and PathRAG
11.3. Generalization Potential
11.4. Practical Efficiency
12. Limitations and Open Challenges
13. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.t.; Rocktäschel, T.; et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Red Hook, NY, USA, Online, 6–12 December 2020. [Google Scholar]
- Karpukhin, V.; Oguz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; Yih, W.T. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online; Webber, B., Cohn, T., He, Y., Liu, Y., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, November 2020; pp. 6769–6781. [Google Scholar] [CrossRef]
- Izacard, G.; Caron, M.; Hosseini, L.; Riedel, S.; Bojanowski, P.; Joulin, A.; Grave, E. Unsupervised Dense Information Retrieval with Contrastive Learning. arXiv 2022, arXiv:2112.09118. [Google Scholar]
- Santhanam, K.; Khattab, O.; Saad-Falcon, J.; Potts, C.; Zaharia, M. ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, WA, USA; Carpuat, M., de Marneffe, M.C., Meza Ruiz, I.V., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, July 2022; pp. 3715–3734. [Google Scholar] [CrossRef]
- Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; Metropolitansky, D.; Ness, R.O.; Larson, J. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv 2025, arXiv:2404.16130. [Google Scholar]
- Guo, Z.; Xia, L.; Yu, Y.; Ao, T.; Huang, C. LightRAG: Simple and Fast Retrieval-Augmented Generation. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China; Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 4–9 November 2025; pp. 10746–10761. [Google Scholar] [CrossRef]
- Pan, S.; Luo, L.; Wang, Y.; Chen, C.; Wang, J.; Wu, X. Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Trans. Knowl. Data Eng. 2024, 36, 3580–3599. [Google Scholar] [CrossRef]
- Nabila, D.; Nasution, A.H.; Murakami, Y.; Koos, S.; Ergun, A.E. Modeling and Benchmarking GraphRAG for Indonesian Legal Question Answering. Artif. Intell. Lang. Model. 2026, 1, 1–12. [Google Scholar] [CrossRef]
- Chen, B.; Guo, Z.; Yang, Z.; Chen, Y.; Chen, J.; Liu, Z.; Shi, C.; Yang, C. PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational Paths. In Proceedings of the Association for the Advancement of Artificial Intelligence (AAAI), Singapore, 20–27 January 2026; Volume 40, pp. 30183–30191. [Google Scholar] [CrossRef]
- Gutiérrez, B.J.; Shu, Y.; Gu, Y.; Yasunaga, M.; Su, Y. HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems, 10–15 December 2024; Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 37, pp. 59532–59569. [Google Scholar] [CrossRef]
- Haveliwala, T.H. Topic-sensitive PageRank. In Proceedings of the 11th International Conference on World Wide Web, New York, NY, USA, 7–11 May 2002; WWW ’02. pp. 517–526. [Google Scholar] [CrossRef]
- Gutiérrez, B.J.; Shu, Y.; Qi, W.; Zhou, S.; Su, Y. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models. In Proceedings of the 42nd International Conference on Machine Learning. PMLR, 13–19 July 2025; Proceedings of Machine Learning Research: New York, NY, USA, 2025; Volume 267, pp. 21497–21515. [Google Scholar]
- Khalila, Z.; Nasution, A.H.; Monika, W.; Onan, A.; Murakami, Y.; Radi, Y.B.I.; Osmani, N.M. Investigating Retrieval-Augmented Generation in Quranic Studies: A Study of 13 Open-Source Large Language Models. Int. J. Adv. Comput. Sci. Appl. 2025, 16, 1361–1371. [Google Scholar] [CrossRef]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [PubMed]
- Elsevier. Scopus: A Comprehensive Abstract and Citation Database for Impact Makers. 2026. Available online: https://www.elsevier.com/products/scopus (accessed on 13 June 2026).
- Baas, J.; Schotten, M.; Plume, A.; Côté, G.; Karimi, R. Scopus as a curated, high-quality bibliometric data source for academic research in quantitative science studies. Quant. Sci. Stud. 2020, 1, 377–386. [Google Scholar] [CrossRef]
- Pereira, V.; Basilio, M.P.; Santos, C.H.T. PyBibX—A Python library for bibliometric and scientometric analysis powered with artificial intelligence tools. Data Technol. Appl. 2025, 59, 302–337. [Google Scholar] [CrossRef]
- Saeed, T.; Wang, B. Large Language Models for Generative Recommendation: A Systematic Review of Data-Centric Taxonomy, Evaluation, and Human-Centric Analytics. Int. J. Data Sci. Anal. 2026, 22, 130. [Google Scholar] [CrossRef]
- Billah, S.M.; Yusof, R.J.B.R. Exploring the Role of RAG in Cloud-Based Monolithic Chatbots: A Systematic Review. In Proceedings of the 2025 IEEE 23rd Student Conference on Research and Development (SCOReD), Kuala Lumpur, Malaysia, 25–26 November 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Tsallis, C.; Papageorgas, P.; Munteanu, R.A.; Dellagi, S. Large Language and Foundation Models for Machinery Health Monitoring: A Systematic Review. Appl. Sci. 2026, 16, 2493. [Google Scholar] [CrossRef]
- Choi, M.; Ahsan, I.; Yu, H.; Choe, T.; Kim, M. The Semantic Design Space of Retrieval-Augmented Recommender Systems: A Systematic Review of LLM-Based Approaches. Comput. Mater. Contin. 2026, 88, 3. [Google Scholar] [CrossRef]
- Fatawi, I.; Bilad, M.; Asy’ari, M. The role of prompt engineering in enhancing LLMs: A systematic review of applications and ethical implications. IAES Int. J. Artif. Intell. (IJ-AI) 2026, 15, 1071–1086. [Google Scholar] [CrossRef]
- Trad, F.; Yammine, R.; Charafeddine, J.; Chakhtoura, M.; Rahme, M.; El-Hajj Fuleihan, G.; Chehab, A. Streamlining Systematic Reviews with Large Language Models Using Prompt Engineering and Retrieval Augmented Generation. BMC Med. Res. Methodol. 2025, 25, 130. [Google Scholar] [CrossRef] [PubMed]
- Shah, J.; Gade, S.R.; Ponna, D.; Patel, K.A. Cloud-Native AI and Generative AI on AWS: A Systematic Review and a Proposed Unified Model. In Proceedings of the 2026 7th International Conference on Mobile Computing and Sustainable Informatics (ICMCSI), Goathgaun, Morang, Nepal, 7–9 January 2026; pp. 1710–1717. [Google Scholar] [CrossRef]
- Mikulić, I.; Vlaić, M.; Delač, G.; Šilić, M.; Vladimir, K. Integrating External Knowledge with LLMs: A Systematic Review of RAG Approaches. In Proceedings of the 2025 MIPRO 48th ICT and Electronics Convention, Opatija, Croatia, 2–6 June 2025; pp. 93–98. [Google Scholar] [CrossRef]
- Smajić, A.; Karlović, R.; Bobanović Dasko, M.; Lorencin, I. Large Language Models for Structured and Semi-Structured Data, Recommender Systems and Knowledge Base Engineering: A Survey of Recent Techniques and Architectures. Electronics 2025, 14, 3153. [Google Scholar] [CrossRef]
- Murugaraj, K.; Lamsiyah, S.; Theobald, M. ExpandFuse: A Hybrid Retrieval Framework with Query Expansion and Topic-Aware Reranking for Multi-Hop Question Answering. In Proceedings of the 2025 IEEE International Conference on Big Data (BigData), Macau, China, 8–11 December 2025; pp. 722–729. [Google Scholar] [CrossRef]
- Umadevi, K.S.; Rajeshwari, S.R.; Sai Likhitha, P. Medical RAG: Hybrid Retrieval & Re-Ranking. In Proceedings of the 2025 International Conference on Responsible, Generative and Explainable AI (ResGenXAI), Bhubaneswar, India, 10–12 September 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Sun, Y. Adaptive Hybrid Retrieval-Augmented Generation with Context-Aware Confidence Control. In Proceedings of the 2025 4th International Conference on Electronic Information Technology (EIT), Chengdu, China, 22–24 August 2025; pp. 891–895. [Google Scholar] [CrossRef]
- Abirami, S.; Joshma, M.J.; Jayavarthini, A.; Joshi, A.; Gajendran, M.K. Evaluating Lexical, Dense, and Hybrid Retrieval Pipelines for RAG. In Proceedings of the 2025 IEEE Pune Section International Conference (PuneCon), Pune, India, 12–14 December 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Hu, C.; Huang, Y.; Kuang, J.; Dai, B.; Peng, Y.; Xiao, Y.; Su, Y. Mitigating Hallucinations in Discipline Inspection QA: A Two-Stage RAG Framework with Late Interaction and Reranking. Electronics 2026, 15, 541. [Google Scholar] [CrossRef]
- Aljohani, B.; Alsanoosy, T. Enhancing Medical Question Answering with LLMs via a Hybrid Retrieval-Augmented Generation Framework. Information 2026, 17, 133. [Google Scholar] [CrossRef]
- Ahmad, S.; Nezami, Z.; Hafeez, M.; Raza Zaidi, S.A. Benchmarking Vector, Graph and Hybrid Retrieval Augmented Generation (RAG) Pipelines for Open Radio Access Networks (ORAN). In Proceedings of the 2025 IEEE 36th International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), İstanbul, Türkiye, 1–4 September 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Achyar, M.K.U.; Widyawan; Pratama, A.R. RAG Chatbot Architecture for Law & Crime News Using Hybrid Retrieval and Small Language Model. In Proceedings of the 2025 5th International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA), Yogyakarta, Indonesia, 17–19 December 2025; pp. 588–593. [Google Scholar] [CrossRef]
- Altınok, H.; Tekgoz, H.; Erdoğan, H.; Uz, H. A Comparative Analysis of Retrieval-Augmented Generation Architectures with Semantic Hashing for Enterprise Knowledge Systems. In Proceedings of the 2026 5th International Informatics and Software Engineering Conference (IISEC), Ankara, Türkiye, 5–6 February 2026; pp. 343–347. [Google Scholar] [CrossRef]
- Hamza, R.M.; Ajam, G.E. Context-Aware Intelligent Data Integration Approach: A practical Hybrid Retrieval Agent. In Proceedings of the 2026 2nd International Conference on Computing and Emerging Sciences (ICCES ’26), Erbil, Iraq, 4–5 February 2026; pp. 121–127. [Google Scholar] [CrossRef]
- Li, R.; Mao, S.; Zhu, C.; Yang, Y.; Tan, C.; Li, L.; Mu, X.; Liu, H.; Yang, Y. Enhancing Pulmonary Disease Prediction Using Large Language Models With Feature Summarization and Hybrid Retrieval-Augmented Generation: Multicenter Methodological Study Based on Radiology Report. J. Med. Internet Res. 2025, 27, e72638. [Google Scholar] [CrossRef] [PubMed]
- Shang, Y.; Ke, Z.; Lin, P.; Ren, Q.; Zhang, W.; Wang, X.; Li, X.; Gong, F.; Wang, S.; Wang, B.; et al. Empowering knowledge graphs with hybrid retrieval-augmented generation for the intelligent mix scheme of mass concrete. Case Stud. Constr. Mater. 2025, 23, e04979. [Google Scholar] [CrossRef]
- Abdolinejad, F.; Eftekhari, M. Augmenting RAG with Nonnegative Matrix Factorization-Driven Semantic Chunking in Embedding Space. J. Supercomput. 2026, 82, 224. [Google Scholar] [CrossRef]
- Lavarec, E.; Du, Y. Evaluating document chunking approaches for retrieval augmented generation in editorial content. IAES Int. J. Artif. Intell. (IJ-AI) 2026, 15, 1909–1918. [Google Scholar] [CrossRef]
- Koay, X.K.; Ong, L.Y.; Goh, P.Y. Structure-Aware Chunking for Complex Tables in Retrieval-Augmented Generation Systems. Emerg. Sci. J. 2026, 10, 184–205. [Google Scholar] [CrossRef]
- Zhang, L.; Ning, Y. Improving construction contract question answering through embedding optimization and semantic chunking in large language models. Adv. Eng. Inform. 2026, 69, 104027. [Google Scholar] [CrossRef]
- Moreno-Cediel, A.; Garcia-Lopez, E.; Garcia-Cabot, A.; De-Fitero-Dominguez, D. Optimising retrieval performance in RAG systems: A new growing window semantic chunking strategy to address weak semantic boundaries. Knowl.-Based Syst. 2026, 331, 114896. [Google Scholar] [CrossRef]
- Niu, Y.; Rong, X. Improving Retrieval-Augmented Generation for Educational Policy Understanding via Structure-Aware Text Chunking. In Proceedings of the 2026 3rd International Conference on Informatics Education and Computer Technology Applications (IECA 2026), Shanghai, China, 16–18 January 2026; pp. 1093–1100. [Google Scholar] [CrossRef]
- Li, X.; Xue, T. HTR-GEN: A Structure-Aware Framework for Hierarchical Table Retrieval and Generation. In Proceedings of the 2025 5th International Conference on Computer Systems (ICCS), Xi’an, China, 26–28 September 2025; pp. 77–81. [Google Scholar] [CrossRef]
- Lee, S.; Kim, N.; Lee, J. Structural Chunking: A Semantic-Structural Integrated Method for Retrieval-Augmented Generation. In Proceedings of the 2026 International Conference on Electronics, Information, and Communication (ICEIC), Macau, China, 18–21 January 2026; pp. 1–6. [Google Scholar] [CrossRef]
- Jaiswal, S.; Bisht, P.; Kansara, K.; Datta, M.S. Comparison of Chunking Techniques Across Diverse Document Types in NLP Retrieval Tasks. In Proceedings of the 2025 International Conference on Responsible, Generative and Explainable AI (ResGenXAI), Bhubaneswar, India, 10–12 September 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Widiiswa, R.A.N.; Henry, M.M.; Imani, A.; Pardamean, B. Retrieval-Augmented Generation Chunking Strategies for Financial Document Analysis: A Systematic Literature Review. In Proceedings of the 2026 30th International Conference on Information Technology (IT), Žabljak, Montenegro, 24–28 February 2026; pp. 1–4. [Google Scholar] [CrossRef]
- Qu, L.; Zhao, X.; Zhang, C.; Li, G. GraphRAG-Vet: A Knowledge Graph-Augmented Large Language Model for Precision Bovine Disease Diagnosis. Computers 2026, 15, 203. [Google Scholar] [CrossRef]
- Wang, X.; Fang, J. Claim Knowledge Graph Construction and GraphRAG-Based Question-Answering System. Buildings 2026, 16, 845. [Google Scholar] [CrossRef]
- Li, M.; Qin, R. DualGraphRAG: A Dual-View Graph-Enhanced Retrieval-Augmented Generation Framework for Reliable and Efficient Question Answering. Appl. Sci. 2026, 16, 2221. [Google Scholar] [CrossRef]
- Chen, S.; Chen, T. Co-MedGraphRAG: A Collaborative Large–Small Model Medical Question-Answering Framework Enhanced by Knowledge Graph Reasoning. Information 2026, 17, 247. [Google Scholar] [CrossRef]
- Jiang, B.; Liu, Z.; Wang, N.; Li, Z.; Shi, Y.; Lin, B. Process-Oriented Dual-Layer Knowledge GraphRAG for Reservoir Engineering Decision Support. Processes 2025, 13, 3230. [Google Scholar] [CrossRef]
- Knollmeyer, S.; Caymazer, O.; Grossmann, D. Document GraphRAG: Knowledge Graph Enhanced Retrieval Augmented Generation for Document Question Answering Within the Manufacturing Domain. Electronics 2025, 14, 2102. [Google Scholar] [CrossRef]
- Wang, C.; Fu, Y.; Wang, C.; Wang, L.; Zheng, Q.; Zhang, H.; Li, M.; Meng, F. CRAG-IKG: A Conflict- and Reliability-Aware GraphRAG Framework for Noisy Industrial Knowledge Graphs. In Proceedings of the 2025 5th International Conference on Electronic Communication, Computer Science and Technology (ECCST), Qinhuangdao, China, 26–28 December 2025; pp. 373–377. [Google Scholar] [CrossRef]
- Lv, X.; Feng, Y.; Zheng, J. SL-MERK: Synthetic Lethality Mechanism Explainer based on GraphRAG and Knowledge Graph. In Proceedings of the 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Copenhagen, Denmark, 14–18 July 2025; pp. 1–6. [Google Scholar] [CrossRef] [PubMed]
- Xu, J.; Zhang, H.; Zhang, H.; Lu, J.; Xiao, G. ChatTf: A Knowledge Graph-Enhanced Intelligent Q&A System for Mitigating Factuality Hallucinations in Traditional Folklore. IEEE Access 2024, 12, 162638–162650. [Google Scholar] [CrossRef]
- Utami, L.; Rachmi, H.; Hidayatulloh, S. A Hybrid TF-IDF and Knowledge Graph-Enhanced Retrieval-Augmented Generation Framework with Large Language Models for Domain-Aware Question Answering. J. Appl. Data Sci. 2026, 7, 866–881. [Google Scholar] [CrossRef]
- Lin, S.; Shao, S.; Liu, X.; Su, H. FusionGraphRAG: An Adaptive Retrieval-Augmented Generation Framework for Complex Disease Management in the Elderly. Information 2026, 17, 138. [Google Scholar] [CrossRef]
- Jovanovski, D.; Stojcheva, M.; Dodevska, M.; Lameski, P.; Mishkovski, I.; Gjorgjevikj, D. An Empirical Study of Knowledge Graph-Enhanced RAG for Information Security Compliance. Information 2026, 17, 389. [Google Scholar] [CrossRef]
- Lecu, A.; Groza, A.; Hawizy, L. Reducing Hallucinations in Medical AI: A Knowledge Graph-Augmented Retrieval System for Evidence-Based Age-Related Macular Degeneration Information. IEEE Access 2025, 13, 210624–210639. [Google Scholar] [CrossRef]
- Dong, C.; Yuan, Y.; Chen, K.; Cheng, S.; Wen, C. How to Build an Adaptive AI Tutor for Any Course Using Knowledge Graph-Enhanced Retrieval-Augmented Generation (KG-RAG). In Proceedings of the 2025 14th International Conference on Educational and Information Technology (ICEIT), Guangzhou, China, 14–16 March 2025; pp. 152–157. [Google Scholar] [CrossRef]
- Yadav, V.; Gaurav; Rana, A.; Sharma, S. MedRAG-Agent: Medical Query Resolution By Employing A Multi-Agent, Knowledge Graph-Enhanced RAG-Based AI Framework. In Proceedings of the 2025 IEEE 6th Global Conference for Advancement in Technology (GCAT), Bangalore, India, 24–26 October 2025; pp. 1–4. [Google Scholar] [CrossRef]
- Chen, W.; Zhou, Y.; Yan, J.; He, Y.; Lu, H.; Tian, Z. APT-ArgusQA: Knowledge Graph-Enhanced Large Language Model-based Question Answering for Advanced Persistent Threats. In Proceedings of the 2025 IEEE 6th International Conference on Computer, Big Data, Artificial Intelligence (ICCBD+AI), Xiamen, China, 21–23 November 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Flüh, M.; Kim, S.Y.; Schneider, C.V.; Geisler, S. FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis. In Proceedings of the 2025 IEEE International Conference on Knowledge Graph (ICKG), Limassol, Cyprus, 13–14 November 2025; pp. 90–97. [Google Scholar] [CrossRef]
- Huang, S.; Cheng, J. StructReason: A Multi-hop Retrieval-Augmented Generation Framework via Iterative Reasoning Chains and PCST Structural Refinement. In Proceedings of the 2025 5th International Conference on Communication Technology and Information Technology (ICCTIT), Guangzhou, China, 26–28 December 2025; pp. 201–206. [Google Scholar] [CrossRef]
- Ding, L.; Ding, N.; Tao, Q.; Shi, P. Enhancing graph multi-hop reasoning for question answering with LLMs: An approach based on adaptive path generation. J. Intell. Inf. Syst. 2025, 63, 1455–1485. [Google Scholar] [CrossRef]
- Jian, Y.; Lan, L.; Luo, Z. CogNav: Cognitive Navigation for Complex Multi-Hop Retrieval via Chain-of-Thought Reasoning. In Proceedings of the 2nd International Conference on Digital Society, Information Science and Risk Management (ICDIR 2026), Shenyang, China, 23–25 January 2026; pp. 114–117. [Google Scholar] [CrossRef]
- Huang, H. H2RAG: A Hub-Aware Hypergraph Retrieval-Augmented Generation Framework for Multi-Hop Reasoning. In Proceedings of the 2026 5th International Conference on Electronics Technology and Artificial Intelligence (ETAI), Harbin, China, 6–8 March 2026; pp. 362–365. [Google Scholar] [CrossRef]
- Huang, Y.; Yang, L.; Yang, X.H.; Xu, X. Retrieval-Augmented Generation for Multi-Hop Question Answering Based on Structured Planning. ACM Trans. Knowl. Discov. Data 2026, 20, 1–20. [Google Scholar] [CrossRef]
- Zhang, X.; Zhao, F.; Liu, Y.; Chen, P.; Wang, Y.; Wang, X.; Ma, D.; Xu, H.; Chen, M.; Li, H. TreeQA: Enhanced LLM-RAG with logic tree reasoning for reliable and interpretable multi-hop question answering. Knowl.-Based Syst. 2025, 330, 114526. [Google Scholar] [CrossRef]
- Wang, H.; Wang, T.; Sun, Z.; Li, H.; Cao, Z.; Feng, L.; Wang, D. A Query-Driven Graph Retrieval Framework with Adaptive Pruning for Multi-Hop Question Answering. Electronics 2026, 15, 1263. [Google Scholar] [CrossRef]
- Wang, S.; Wang, Z.; Qu, C.; Yin, Z. Multi-level retrieval with representation alignment enhances cross-document evidence synthesis for scientific knowledge generation. Appl. Soft Comput. 2026, 191, 114649. [Google Scholar] [CrossRef]
- Zai, X.; Tan, X.; Wang, X.; Liu, Q.; Xu, X.; Zhang, W. PRoH: Dynamic Planning and Reasoning over Knowledge Hypergraphs for Retrieval-Augmented Generation. In Proceedings of the ACM Web Conference 2026, Dubai, United Arab Emirates, 13–17 April 2026; pp. 4256–4267. [Google Scholar]
- Hamid, M.R.A.; El-Regaily, S.A.; Aref, M.M. Can Large Language Models Perform Retrieval-Augmented Generation as Multi-Hop Reasoning Over Knowledge Graphs? In Proceedings of the 2025 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), Cairo, Egypt, 17–18 September 2025; pp. 360–365. [Google Scholar] [CrossRef]
- Wang, J.; Shen, H.; Xie, B.; Chen, Y.; Zhao, W.; Wang, X.; Hong, Y.; Fu, C.; Pan, Z.; Sun, L.; et al. Agentic Graph-RAG: A Multi-Agent Framework for Robust, Decomposed Multi-Hop Reasoning. In Proceedings of the 2025 11th International Conference on Computer and Communications (ICCC), Chengdu, China, 12–15 December 2025; pp. 574–578. [Google Scholar] [CrossRef]
- Fu, R.; Wang, Y.; Xu, T.; Liu, Y.; Tang, W.; Wu, W.; Ma, X.; Fong, S. S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering. In Proceedings of the ACM Web Conference 2026, Dubai, United Arab Emirates, 29 June–3 July 2026; pp. 4057–4068. [Google Scholar] [CrossRef]
- Kamlakshya, T. Agentic RAG Systems for Real-Time Financial Decision Making: A Multi-Agent Framework with Model Context Protocol Integration. In Proceedings of the 2025 IEEE 2nd International Conference on Information Technology, Electronics and Intelligent Communication Systems (ICITEICS), Bangalore, India, 29–30 August 2025; pp. 1–3. [Google Scholar] [CrossRef]
- Hariharan, M.; Barma, S.; Arvapalli, S.; Sheela, E. Agentic RAG for Software Testing with Hybrid Vector-Graph and Multi-Agent Orchestration. In Proceedings of the 2025 IEEE International Conference on Data and Software Engineering (ICoDSE), Batam, Indonesia, 28–29 October 2025; pp. 54–59. [Google Scholar] [CrossRef]
- Hajaghaie, A.; Thulasiram, R.K. Agentic Portfolio Construction: A Multi-Agent Architecture for LLM-Driven Financial Asset Allocation. In Proceedings of the 2025 3rd International Conference on Foundation and Large Language Models (FLLM), Vienna, Austria, 25–28 November 2025; pp. 192–201. [Google Scholar] [CrossRef]
- Liang, C.; Cui, Y.; Shi, R.; Zha, G.; Yin, X.; Xiao, M.; Xu, D.; Duan, X.; Huang, B. GeoAgentic-RAG: A Multi-Agent framework for autonomous geospatial reasoning and visual insight generation with LLM. Int. J. Appl. Earth Obs. Geoinf. 2026, 147, 105195. [Google Scholar] [CrossRef]
- Tanasă, A.M.; Oprea, S.V.; Bâra, A. Designing an Architecture of a Multi-Agentic AI-Powered Virtual Assistant Using LLMs and RAG for a Medical Clinic. Electronics 2026, 15, 334. [Google Scholar] [CrossRef]
- Chondamrongkul, N.; Kyaw, M.P.P.; Ko, S.M.; Paing, P.P.; Than Swe, M.K.; Hongthong, T. RepoAI: Automated code refactoring through multi-agent LLM orchestration and retrieval-augmented generation. Sci. Comput. Program. 2026, 253, 103477. [Google Scholar] [CrossRef]
- Roumeliotis, K.I.; Sapkota, R.; Karkee, M.; Tselikas, N.D. Agentic AI With Orchestrator-Agent Trust: A Modular Visual Classification Framework With Trust-Aware Orchestration and RAG-Based Reasoning. IEEE Access 2026, 14, 26965–26982. [Google Scholar] [CrossRef]
- Ahi, K.; Hsieh, C.H.; Fenger, G. LLMs and LVMs for agentic AI: A GPU-accelerated multimodal system architecture for RAG-grounded, explainable, and adaptive intelligence. In Proceedings of the Photomask Technology 2025, Monterey, CA, USA, 22–25 September 2025; Melvin, L.S., III, Philipsen, V., Eds.; Proceedings of SPIE; International Society for Optics and Photonics: Bellingham, WA, USA, 2025; Volume 13687, p. 136871R. [Google Scholar] [CrossRef]
- Hadee, A.N.A.; Riznee, M. Code Prism: A Multi-Agent, Multi-LLM, Semantic Indexing Artifact for Regulatory Code Audits—A Design Science Research Study. In Proceedings of the 2025 1st International Conference on Emerging Innovation and Digital Technology (ICEIDT), Malé, Maldives, 24–25 November 2025; pp. 23–27. [Google Scholar] [CrossRef]
- Salehi, S.; Singh, Y.; Horst, K.K.; Hathaway, Q.A.; Erickson, B.J. Agentic AI and Large Language Models in Radiology: Opportunities and Hallucination Challenges. Bioengineering 2025, 12, 1303. [Google Scholar] [CrossRef] [PubMed]
- Liu, Z.; Kou, J.; Zhang, W.; Gu, C.; Fang, X.; Huang, Z.; Yuan, H.; Li, H.; Lu, X.; Yin, A.; et al. Comprehensive Evaluation of AI Hallucination and Novel UV-Oriented Framework toward Safe and Trustworthy AI. In Proceedings of the 2024 7th International Conference on Universal Village (UV), Boston, MA, USA, 19–22 October 2024; pp. 1–136. [Google Scholar] [CrossRef]
- Gu, C.; Zhang, W.; Huang, Z.; Kou, J.; Liu, Z.; Zhao, C.; Liu, C.; Zhang, L.; Lin, W.; Wang, Z.; et al. LENS: Layers of Evaluation of Hallucination in GenAI Systems. In Proceedings of the 2024 7th International Conference on Universal Village (UV), Boston, MA, USA, 19–22 October 2024; pp. 1–85. [Google Scholar] [CrossRef]
- Pawlik, L.; Deniziak, S. Reducing Hallucinations in Medical AI Through Citation Enforced Prompting in RAG Systems. Appl. Sci. 2026, 16, 3013. [Google Scholar] [CrossRef]
- Pai, K.C.; Hsu, W.C. Fine-tuning small language models for industrial retrieval-augmented generation: Efficiency, factuality, and hallucination analysis. Comput. Stand. Interfaces 2026, 98, 104163. [Google Scholar] [CrossRef]
- Tamilselvi, P.; Ahmed Amin, K.; Vigneshvaran, K.; Rohith, S. Benchmarking Quantized LLMS for Faithfulness in Resource-Constrained RAG. In Proceedings of the 2026 Second International Conference on Multi-Agent Systems for Collaborative Intelligence (ICMSCI), Erode, India, 2–4 March 2026; pp. 762–765. [Google Scholar] [CrossRef]
- Krainovskikh, V.; Samigulin, T. Domain-Adapted Retrieval-Augmented Generation for Technical Documentation: Enhancing Reliability and Faithfulness in Technical QA. IEEE Access 2026, 14, 36016–36024. [Google Scholar] [CrossRef]
- Olariu, M.E.; Buinceanu, V.G.; Simionescu, C.; Dospinescu, O.; Georgescu, R.; Tudor, C.; Iftene, A.; Bores, A.M. RO-FIN-LLM: A Benchmark with LLM-as-a-Judge and Human Evaluators for Romanian Tax and Accounting. Systems 2026, 14, 244. [Google Scholar] [CrossRef]
- Papageorgiou, G.; Sarlis, V.; Maragoudakis, M.; Magnisalis, I.; Tjortjis, C. Evaluating Faithfulness in Agentic RAG Systems for e-Governance Applications Using LLM-Based Judging Frameworks. Big Data Cogn. Comput. 2025, 9, 309. [Google Scholar] [CrossRef]
- Wang, Y.; Zhang, Y.; Xu, K.; Xu, Y.; Wang, Y.; Feng, Q. CORB-RAG: A Comprehensive Evaluation Benchmark for Retrieval-Augmented Generation Systems in the Chinese Telecommunications Operator Domain. In Proceedings of the 2025 IEEE 5th International Conference on Computer Communication and Artificial Intelligence (CCAI), Haikou, China, 23–25 May 2025; pp. 503–508. [Google Scholar] [CrossRef]
- More, R. A Unified Evaluation Framework for Grounded LLM Architectures: Comparative Analysis of RAG, Self-RAG, and Agentic RAG. In Proceedings of the 2025 5th International Conference on Artificial Intelligence and Signal Processing (AISP), Vijayawada, India, 22–24 November 2025; pp. 1–5. [Google Scholar] [CrossRef]
- Nishisako, S.; Higashi, T.; Wakao, F. Reducing Hallucinations and Trade-Offs in Responses in Generative AI Chatbots for Cancer Information: Development and Evaluation Study. JMIR Cancer 2025, 11, e70176. [Google Scholar] [CrossRef] [PubMed]
- Wysocka, M.; Wysocki, O.; Delmas, M.; Mutel, V.; Freitas, A. Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation. J. Biomed. Inform. 2024, 158, 104724. [Google Scholar] [CrossRef] [PubMed]
- Zhang, H.; Li, D.; Nie, X. Mitigating Execution Hallucinations and Computational Inflation in Agentic RAG via Strict Protocol Boundaries. Electronics 2026, 15, 1805. [Google Scholar] [CrossRef]
- Wallat, J.; Heuss, M.; Rijke, M.d.; Anand, A. Correctness is not Faithfulness in Retrieval Augmented Generation Attributions. In Proceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR), Padua, Italy, 18 July 2025; Zamani, H., Dietz, L., Piwowarski, B., Bruch, S., Eds.; ACM: New York, NY, USA, 2025; pp. 22–32. [Google Scholar] [CrossRef]
- Noguera, A.; Mogollón-Benavides, A.L.; Niño-Mojica, M.D.; Rua, S.; Sanin-Villa, D.; Tejada, J.C. Applications and Challenges of Retrieval-Augmented Generation (RAG) in Maternal Health: A Multi-Axial Review of the State of the Art in Biomedical QA with LLMs. Sci 2025, 7, 148. [Google Scholar] [CrossRef]
- Shukla, D.; Shirote, S.; Soma, O.; Sorte, A.; Banchhor, S.; Takale, D. MediSense: AI-Based Dual Summarization of Clinical Reports for Healthcare Professionals and Patients. In Proceedings of the 2025 3rd DMIHER International Conference on Artificial Intelligence in Healthcare, Education and Industry (IDICAIHEI), Wardha, India, 28–29 November 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Feng, Y.; Wang, J.; He, R.; Zhou, L.; Li, Y. A retrieval-augmented knowledge mining method with deep thinking LLMs for biomedical research and clinical support. GigaScience 2025, 14, giaf109. [Google Scholar] [CrossRef] [PubMed]
- Tang, W.; Chen, R.; Long, X.; Yu, D.; Zhao, S.; Chen, B. Medical large language models and systems in the clinical application of spinal diseases: Current status, challenges, and future prospects. J. Orthop. Transl. 2026, 57, 101050. [Google Scholar] [CrossRef] [PubMed]
- Rahulprasath, S.; Pranav Harshan, S.; Kabilash, P.V.; Lakshithraj, A.; Sreemathy, J. AI in Healthcare: Simplifying Medical Reports for Enhanced Patient Comprehension. In Proceedings of the 2025 International Conference on Emerging Technologies in Computing and Communication (ETCC), Bangalore, India, 26–27 June 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Thio, S.; Lewis, M.; Denaxas, S.; Dobson, R.J.B. Unlocking electronic health records: A hybrid graph RAG approach to safe clinical AI for patient QA. Front. Digit. Health 2026, 8. [Google Scholar] [CrossRef] [PubMed]
- Kulshreshtha, A.; Choudhary, A.; Taneja, T.; Verma, S. Enhancing Healthcare Accessibility: A RAG- Based Medical Chatbot Using Transformer Models. In Proceedings of the 2024 International Conference on IT Innovation and Knowledge Discovery (ITIKD), Manama, Bahrain, 13–15 April 2025; pp. 1–4. [Google Scholar] [CrossRef]
- Raputri, E.; Teguh, A.J.; Hidayah, S.N.; Anom, A.K.; Setiawan, F.; Qomariyah, N.N. Retrieval-Augmented LLMs with Indonesian Clinical Trials Guidelines: A Comparative Study. In Proceedings of the 2025 8th International Seminar on Research of Information Technology and Intelligent Systems (ISRITI), Yogyakarta, Indonesia, 11 December 2025; pp. 459–464. [Google Scholar] [CrossRef]
- Li, W.; Zhang, Y.; Wang, C.; Li, Y.; He, X.; Wang, A.L.; Xu, M.; Zhang, F.; Sun, H.; Wang, K.; et al. CARE: A clinical agentic reasoning engine to enhance real-World diagnostic accuracy via structured medical reasoning. Expert Syst. Appl. 2026, 313, 131476. [Google Scholar] [CrossRef]
- Wang, J.F.; Chang, C.C.; Chiang, T.M.; Yeh, T.C.; Cheng, E.; Lee, Y.T.; Chen, H.I. An innovative X-RAG technique combined with GPT-4o for summarizing medical information from EHR and EMR to assist doctors in clinical decision-making effectively and efficiently. Health Inform. J. 2025, 31, 1–16. [Google Scholar] [CrossRef] [PubMed]
- Guiducci, L.; Saulle, C.; Dimitri, G.M.; Valli, B.; Alpini, S.; Tenti, C.; Rizzo, A. Dialogical AI for Cognitive Bias Mitigation in Medical Diagnosis. Appl. Sci. 2026, 16, 710. [Google Scholar] [CrossRef]
- Berkowitz, J.S.; Srinivasan, A.; Acitores Cortina, J.M.; Fatapour, Y.; Tatonetti, N.P. Biomedical text normalization through generative modeling. J. Biomed. Inform. 2025, 167, 104850. [Google Scholar] [CrossRef] [PubMed]
- Kalyanasundaram, T.; Bussari, S.; Sundaram, D.; Sharma, K.K. Engineering Reliable Retrieval-Augmented Generation for Regulatory and Compliance Systems. In Proceedings of the SoutheastCon 2026, Huntsville, AL, USA, 13–15 March 2026; pp. 1–5. [Google Scholar] [CrossRef]
- Bharucha, D.; Palanisamy, N.; Reddy, M. ReguQuery: An Agentic Framework For Automated Regulatory Compliance Analysis Using Retrieval-Augmented Generation. In Proceedings of the 2026 6th Biennial International Conference on Nascent Technologies in Engineering (ICNTE), Navi Mumbai, India, 16–17 January 2026; pp. 1–6. [Google Scholar] [CrossRef]
- Arshad, U.; Corsar, D.; Nkisi-Orji, I. Integrating KGs and ontologies with RAG for personalised summarisation in regulatory compliance. In Proceedings of the SICSA REALLM Workshop 2024, Aberdeen, UK, 17 October 2024; CEUR Workshop Proceedings. Volume 3822, pp. 56–61.
- Xu, Y.; Xue, C.; Zhu, T.; Tian, Y.; Hu, J.N.; Tong, L. Research on the Technical Framework of Legal Compliance Review of Major Decision-Making based on Large Language Model. In Proceedings of the 2025 IEEE 5th International Conference on Applied Mathematics, Modeling and Computer Simulation (AMMCS), Wuhan, China, 23–24 August 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Zhai, H. Law GraphRAG: An Advanced Legal Question-Answering System. In Proceedings of the 2025 5th International Conference on Artificial Intelligence and Industrial Technology Applications (AIITA), Xi’an, China, 28–30 March 2025; pp. 1407–1410. [Google Scholar]
- Kim, J.; Hur, M.; Min, M. From RAG to QA-RAG: Integrating Generative AI for Pharmaceutical Regulatory Compliance Process. In Proceedings of the 40th ACM/SIGAPP Symposium on Applied Computing, Catania, Italy, 31 March 2025–4 April 2025; pp. 1293–1295. [Google Scholar] [CrossRef]
- Yao, S.; Ke, Q.; Wang, Q.; Li, K.; Hu, J. Lawyer GPT: A Legal Large Language Model with Enhanced Domain Knowledge and Reasoning Capabilities. In Proceedings of the 2024 3rd International Symposium on Robotics, Artificial Intelligence and Information Engineering, Singapore, 5–7 July 2024; pp. 108–112. [Google Scholar] [CrossRef]
- Kalyanasundaram, T.; Jadav, V. A Compliance-Focused Retrieval-Augmented AI System for Pharmacy Policy Assistance. In Proceedings of the 2025 International Conference on Computer and Applications (ICCA), Bahrain, Bahrain, 22–24 December 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Huang, S.; Sun, C.; Ning, M.; Yang, Y.; Ma, C.; Zhong, J.; Shu, K.; Shi, F.; Khajepour, A. DriveLegal: Toward legally compliant driving via trustworthy hybrid retrieval-augmented LLMs. Expert Syst. Appl. 2026, 314, 131593. [Google Scholar] [CrossRef]
- Amazou, Y.; Tayalati, F.; Mensouri, H.; Azmani, A.; Azmani, M. Accurate AI Assistance in Contract Law Using Retrieval-Augmented Generation to Advance Legal Technology. Int. J. Adv. Comput. Sci. Appl. 2025, 16, 1141–1150. [Google Scholar] [CrossRef]
- Xu, B.; Tong, R.J.; Li, Y.; Chen, P.; Li, H.; Liang, J.; Fan, X.; Tong, J. An Architectural Framework for Educational Knowledge Graphs (IEEE P2807.6): Ontology Design, Llm Integration, and Adaptive Learning Applications. In Proceedings of the 2025 IEEE Conference on Artificial Intelligence (CAI), Santa Clara, CA, USA, 5–7 May 2025; pp. 1610–1616. [Google Scholar] [CrossRef]
- Ho, T.L.; Lam, T.P. EduMSRA: A Multi-Source Educational Research Agent Integrating Retrieval-Augmented Generation and Model Context Protocol for Adaptive Intelligent Tutoring Systems. Appl. Sci. 2026, 16, 4400. [Google Scholar] [CrossRef]
- Golla, F. Enhancing Student Engagement Through AI-Powered Educational Chatbots: A Retrieval-Augmented Generation Approach. In Proceedings of the 2024 21st International Conference on Information Technology Based Higher Education and Training (ITHET), Paris, France, 6–8 November 2024; pp. 1–6. [Google Scholar] [CrossRef]
- Maity, S.; Deroy, A.; Sarkar, S. Leveraging In-Context Learning and Retrieval-Augmented Generation for Automatic Question Generation in Educational Domains. In Proceedings of the 16th Annual Meeting of the Forum for Information Retrieval Evaluation, Gandhinagar, India, 12–15 December 2024; pp. 40–47. [Google Scholar] [CrossRef]
- Sahana, S.; Ghosh, P.; Anutariya, C. Reimagining Education: An Architectural Blueprint for a Synergistic AI-Driven Learning Environment. In Proceedings of the 2025 IEEE 16th Annual Information Technology, Electronics and Mobile Communication Conference (IEMCON), Berkeley, CA, USA, 29-31 October 2025; pp. 0123–0130. [Google Scholar] [CrossRef]
- Németh, R.; Tátrai, A.; Szabó, M.; Zaletnyik, P.T.; Tamási, Á. Exploring the use of retrieval-augmented generation models in higher education: A pilot study on artificial intelligence-based tutoring. Soc. Sci. Humanit. Open 2025, 12, 101751. [Google Scholar] [CrossRef]
- Khan, A.; Chandekar, S.; Baikar, A.; Kanade, A.; Kadam, S.; Kadam, V. Educational Application of RAG-based LLMs. In Proceedings of the 2025 5th International Conference on Evolutionary Computing and Mobile Sustainable Networks (ICECMSN), Coimbatore, India, 24–26 November 2025; pp. 720–725. [Google Scholar] [CrossRef]
- Tawfik, M.K.; Ali, M.O.; Mohamed, S.K.; Ezz Elregal, F.S.; Edward, P.E.; Elsawaf, M.; Khattab, D. Mufakkir: RAG-Based Arabic Educational Chatbot for University Students. In Proceedings of the 2025 Twelfth International Conference on Intelligent Computing and Information Systems (ICICIS), Cairo, Egypt, 25–27 November 2025; pp. 706–712. [Google Scholar] [CrossRef]
- Tan, A.; Dorneich, M.C.; Cotos, E. Aligning Pedagogy with Generative AI: An Approach to Customizing Educational GPTs. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting; SAGE Publications: Los Angeles, CA, USA, 2025; Volume 69, pp. 894–899. [Google Scholar] [CrossRef]
- Hu, W.; Gong, R.; Wu, S.; Li, X. A conversational agent based on contingent teaching model to support collaborative learning activities: Impacts on students’ learning performance, self-efficacy and perceptions. Educ. Technol. Res. Dev. 2025, 73, 3341–3372. [Google Scholar] [CrossRef]
- Aljohani, B.; Aljuhani, A. Pioneering agentic retrieval-augmented generation in software quality: A novel framework for code smell detection dynamic retrieval. PeerJ Comput. Sci. 2026, 12, e3642. [Google Scholar] [CrossRef]
- İçöz, B.; Biricik, G. Context-Aware Code Review Automation: A Retrieval-Augmented Approach. Appl. Sci. 2026, 16, 1875. [Google Scholar] [CrossRef]
- Sharanarthi, T.; Polineni, S. Multi-Agent LLM Collaboration for Adaptive Code Review, Debugging, and Security Analysis. In Proceedings of the 2025 International Conference on Mechatronics, Robotics, and Artificial Intelligence (MRAI), Jinan, China, 19–21 June 2025; pp. 541–546. [Google Scholar] [CrossRef]
- Elezi, A.; Cico, B.; Hyseni, D. Tuning DeepSeek-Coder-V2-Lite-Base for C# Code Smell Detection: Advancing Towards Task Versatility in Software Maintenance. In Proceedings of the 2025 14th Mediterranean Conference on Embedded Computing (MECO), Budva, Montenegro, 10–14 June 2025; pp. 1–4. [Google Scholar] [CrossRef]
- Liao, D.; Pan, S.; Sun, X.; Ren, X.; Huang, Q.; Xing, Z.; Jin, H.; Li, Q. A3A3-CodGen: A Repository-Level Code Generation Framework for Code Reuse With Local-Aware, Global-Aware, and Third-Party-Library-Aware. IEEE Trans. Softw. Eng. 2024, 50, 3369–3384. [Google Scholar] [CrossRef]
- Jaoua, I.; Sghaier, O.B.; Sahraoui, H. Combining Large Language Models with Static Analyzers for Code Review Generation. In Proceedings of the 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR), Ottawa, ON, Canada, 28–29 April 2025; pp. 174–186. [Google Scholar] [CrossRef]
- Yang, L.; Chen, Y. Integrating RAG and LLM for Automated Code Review in Practice. In Proceedings of the 2025 10th International Conference on Computer and Information Processing Technology (ISCIPT), Fushun, China, 12–14 September 2025; pp. 591–596. [Google Scholar] [CrossRef]
- Denuri, T.T.L.; Ragukanthan, S.; Wijesekara, W.A.D.N.; Prageeth, M.D.; Silva, A.P.; Attanayaka, L. Archelon AI: Programming Assistant for Legacy Codebases. In Proceedings of the 2025 10th International Conference on Information Technology Research (ICITR), Colombo, Sri Lanka, 8–11 December 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Jiang, C.; Zhang, P.; Ni, Y.; Wang, X.; Peng, H.; Liu, S.; Fei, M.; He, Y.; Xiao, Y.; Huang, J.; et al. Multimodal retrieval-augmented generation for financial documents: Image-centric analysis of charts and tables with large language models. Vis. Comput. 2025, 41, 7657–7670. [Google Scholar] [CrossRef]
- Loo, C.W.; Qian Leong, Z.; Ong, J.K.; Thong Chong, H.; Yu, Y.P.; Ming Lim, T. Retrieval-Augmented Generation with GPT-4o-mini: Integrating Configurable Chunking, Hybrid Search, and Multimodal Image Retrieval. In Proceedings of the 2025 IEEE International Conference on Computation, Big-Data and Engineering (ICCBE), Penang, Malaysia, 27–29 June 2025; pp. 862–866. [Google Scholar] [CrossRef]
- Chang, C.Y.; Khanum, A.; Su, S.G.; Tsai, M.H.; Hsu, H.H.; Chen, W.T.; Lu, H.W. Optimizing Construction Safety: Multimodal Prompts For Automated Image Descriptions In Daily Construction Reports. In Proceedings of the 2025 IEEE International Conference on Image Processing Workshops (ICIPW), Anchorage, AK, USA, 14–17 September 2025; pp. 263–268. [Google Scholar] [CrossRef]
- Zhang, C.; Lin, K.; Yang, Z.; Wang, J.; Li, L.; Lin, C.C.; Liu, Z.; Wang, L. MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 13647–13657. [Google Scholar] [CrossRef]
- Ni, T.; Yuan, X.; Li, S.; Ni, W. Privacy-Preserving Multimodal Reasoning for Internet of Things: A Retrieval-Augmented Large Language and Vision Assistant Framework. IEEE Internet Things Mag. 2026, 9, 113–122. [Google Scholar] [CrossRef]
- Suram, K.; A J, H.; Kurian, A.; Augustine, K.; Majeed A, R. Vision-Language Transformers for Medical Report Synthesis: A Multimodal Approach. In Proceedings of the 2025 IEEE 4th International Conference for Advancement in Technology (ICONAT), Goa, India, 19–21 September 2025; pp. 1–5. [Google Scholar] [CrossRef]
- Markin, E.I.; Zuparova, V.V.; Martyshkin, A.I. Integration of Large Language Models and Computer Vision Algorithms in LMS: A Methodology for Automated Verification of Software Tasks and Multimodal Analysis of Educational Data. In Proceedings of the 2025 International Russian Smart Industry Conference (SmartIndustryCon), Sochi, Russian Federation, 24–28 March 2025; pp. 777–781. [Google Scholar] [CrossRef]
- Miao, J.; Lu, D.; Wang, Z. A Multimodal RAG Framework for Housing Damage Assessment: Collaborative Optimization of Image Encoding and Policy Vector Retrieval. In Proceedings of the 2025 International Conference on Artificial Intelligence and Product Design, New York, NY, USA, 18–20 July 2025; AIPD ’25. pp. 281–285. [Google Scholar] [CrossRef]
- He, H.; Yuan, X.; Wu, K.; Ni, W. Federated Retrieval-Augmented Generation for Cybersecurity in Resource-Constrained IoT and Edge Environments: A Deployment-Oriented Scoping Review. Electronics 2026, 15, 1409. [Google Scholar] [CrossRef]
- Zaw, H.M.; Khant Naing, H.; Myat, A.K.; Paing Linn, H.; Lin, T.; Khaing, T.T. Browser-Embedded, Legal-Aware Cybersecurity Co-Pilot for Myanmar with RAG, Multilingual Defense and Privacy. In Proceedings of the 2025 6th International Conference on Advanced Information Technologies (ICAIT), Yangon, Myanmar, 3 November 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Belmahjoub, M.; Benhiba, L. Causal-Aware Knowledge Graph Enhanced RAG for Predictive Cybersecurity Intelligence: A Framework for Attack Progression Analysis and Consequence Prediction. In Proceedings of the 2025 12th International Conference on Soft Computing & Machine Intelligence (ISCMI), Rio de Janeiro, Brazil, 21–23 November 2025; pp. 136–141. [Google Scholar] [CrossRef]
- Gregory, J.; Liao, Q. Autonomous Cyberattack with Security-Augmented Generative Artificial Intelligence. In Proceedings of the 2024 IEEE International Conference on Cyber Security and Resilience (CSR), London, United Kingdom, 2–4 September 2024; pp. 270–275. [Google Scholar] [CrossRef]
- Gokcimen, T.; Das, B. A novel system for strengthening security in large language models against hallucination and injection attacks with effective strategies. Alex. Eng. J. 2025, 123, 71–90. [Google Scholar] [CrossRef]
- Pu, X.; Zhang, Y. Cultivating Cybersecurity Talent: Localized RAG Approach Integrating Pattern Recognition and Document Analysis. Int. J. Pattern Recognit. Artif. Intell. 2026, 40, 2550034. [Google Scholar] [CrossRef]
- Cheng, Z.; Sun, J.; Gao, A.; Quan, Y.; Liu, Z.; Hu, X.; Fang, M. Secure Retrieval-Augmented Generation Against Poisoning Attacks. In Proceedings of the 2025 IEEE International Conference on Big Data (BigData), Macau, China, 8–11 December 2025; pp. 1799–1806. [Google Scholar] [CrossRef]
- Abd Elhakeem, M.G.; Abdullah, A.O.; Elhag, S.W.; Mohamed, E.H. Aqrag: Advanced quranic retrieval-augmented generation for low-resource question answering. Neural Comput. Appl. 2026, 38, 282. [Google Scholar] [CrossRef]
- Khanduja, N.; Kumar, D.N.; Arun Chauhan, D. A Retrieval-Augmented Generation Model for Faith-Aligned QA in Bhagvat Gita. In Proceedings of the 2025 International Conference on Intelligent and Secure Engineering Solutions (CISES), Greater Noida Gautam Budh Nagar, India, 11–13 August 2025; pp. 1524–1528. [Google Scholar] [CrossRef]
- Obaid, S.; Bawany, N.Z. SeerahGPT: Retrieval Augmented Generation based Large Language Model. In Proceedings of the 2024 18th International Conference on Open Source Systems and Technologies (ICOSST), Lahore, Pakistan, 26–27 December 2024; pp. 1–7. [Google Scholar] [CrossRef]
- Singh, A.; Shrivastava, P.; Misra, R. Chatbots for Religious Punjabi Texts Using Retrieval-Augmented Generation (RAG) for Factual Accuracy: A Review. In Proceedings of the 2025 IEEE 4th International Conference on Technology, Engineering, Management for Societal impact using Marketing, Entrepreneurship and Talent (TEMSMET), New Delhi, India, 8–10 October 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Sun, S. CIRR: Causal-Invariant Retrieval-Augmented Recommendation with Faithful Explanations under Distribution Shift. In Proceedings of the 2025 6th International Conference on Computer Science and Management Technology, Xiamen, China, 26–28 December 2025; pp. 888–894. [Google Scholar] [CrossRef]
- Bara, A.; Oprea, S.V. AI-Augmented Bibliometric Framework: A Paradigm Shift with Agentic AI for Dynamic, Snippet-Based Research Analysis. arXiv 2026, arXiv:2511.21745. [Google Scholar]
- Robertson, S.E.; Walker, S. Some Simple Effective Approximations to the 2-Poisson Model for Probabilistic Weighted Retrieval. In Proceedings of the SIGIR’94: Proceedings of the Seventeenth Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval, Organised by Dublin City University; Springer: London, UK, 1994; pp. 232–241. [Google Scholar] [CrossRef]
- Trivedi, H.; Balasubramanian, N.; Khot, T.; Sabharwal, A. Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, Canada; Rogers, A., Boyd-Graber, J., Okazaki, N., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 10014–10037. [Google Scholar] [CrossRef]
- Bosma, M.; Chi, E.; Ichter, B.; Le, Q.V.; Schuurmans, D.; Wang, X.; Wei, J.; Xia, F.; Zhou, D. Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems 35, New Orleans, LA, USA, 28 November–9 December 2022; pp. 24824–24837. [Google Scholar] [CrossRef]
- Namratha, N.; Naidu, S.S.S.; Sahoo, K.S.; Singla, M.; Praneeth, B.V.S.S.S.R. GeetaVani: A Retrieval-Augmented LLM Framework for Contextual Dialogue from the Bhagavad Geeta. In Proceedings of the 2025 13th International Conference on Intelligent Systems and Embedded Design (ISED), Raipur, India, 17–19 December 2025; pp. 654–661. [Google Scholar] [CrossRef]
- Es, S.; James, J.; Espinosa Anke, L.; Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, St. Julians, Malta, 17–22 March 2024; pp. 150–158. [Google Scholar] [CrossRef]
- Li, Z. Equipping a Virtual Reading Promoter for Digital Libraries with Retrieval-Augmented Generation. In Proceedings of the 2025 ACM/IEEE Joint Conference on Digital Libraries (JCDL), Dekalb, IL, USA, 15–19 December 2025; pp. 313–315. [Google Scholar] [CrossRef]
- Kelly, P.; Schild, J.; Jafari, A.H. FolkRAG: A retrieval-augmented generation system for cultural heritage materials. Neural Comput. Appl. 2025, 37, 20281–20297. [Google Scholar] [CrossRef]
- Júnior, E.F.P.D.S.; Baptista, C.d.S.; Alves, A.L.F.; Mendes, F.I.d.L. Using Retrieval-Augmented Generation to Improve Access to Legal Documents. In Proceedings of the 2026 International Conference on Semantic Computing (ICSC), Laguna Hills, CA, USA, 2–4 February 2026; pp. 328–335. [Google Scholar] [CrossRef]
- Anhein, M.A.G.; Wiharja, K.R.S. Integrating Knowledge Graphs and Semantic Retrieval for Indonesian Legal Question Answering. In Proceedings of the 2025 5th International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA), Yogyakarta, Indonesia, 17–19 December 2025; pp. 389–394. [Google Scholar] [CrossRef]
- Ongris, J.G.; Darari, F.; Tobing, B.C.; Faisal, D.R.; Lee, O. Benchmarking KG-based RAG Systems: A Case Study of Legal Documents. In Proceedings of the Second International Workshop on Retrieval-Augmented Generation Enabled by Knowledge Graphs, Nara, Japan, 2–6 November, 2025; Volume 4079, pp. 69–82. [Google Scholar]
- Wilsen, W.; Dewandaru, A.; Candra, M.Z.C. Generation and Visualization of BPMN from Legal Documents. In Proceedings of the 2025 IEEE International Conference on Data and Software Engineering (ICoDSE), Batam, Indonesia, 28–29 October 2025; pp. 232–237. [Google Scholar] [CrossRef]
- Rumambi, F.; Prasetya, D.D.; Widiyaningtyas, T. Groundedness-Aware Retrieval in Government Document Chatbots: Systematic Literature Review and Semantic Alignment Score (SAS) Formulation. ITEGAM-JETIA 2026, 12, 242–257. [Google Scholar]
- Rizki, A.; Panjaitan, G.P.H.; Azizy, F.N.; Purwarianti, A.; Utama, N.P. Retrieval-Augmented Question Answering for Dukcapil SOPs: Synthetic-Context Fine-Tuning and Chunking Design. In Proceedings of the 2026 International Seminar on Intelligent Business and Edge-Computing Research (ISIBER), Jakarta, Indonesia, 26–26 February 2026; pp. 745–750. [Google Scholar] [CrossRef]
- Liu, Z.; Agrawal, P.; Singhal, S.; Madaan, V.; Kumar, M.; Verma, P.K. LPITutor: An LLM based personalized intelligent tutoring system using RAG and prompt engineering. PeerJ Comput. Sci. 2025, 11, e2991. [Google Scholar] [CrossRef] [PubMed]







| Indicator | Value |
|---|---|
| Initial Scopus records | 2815 |
| Records after duplicate removal | 2809 |
| Timespan | 2023–2026 |
| Countries | 102 |
| Institutions | 4241 |
| Sources | 1267 |
| Authors | 11,532 |
| Conference papers | 1893 |
| Journal articles | 877 |
| Review papers | 39 |
| Total citations | 6816 |
| Average citations per document | 2.43 |
| Bibliometric Finding | Interpretation | Implication for ThemePath-RAG |
|---|---|---|
| Conference papers dominate the corpus (67.4%) | RAG remains a rapidly evolving technical field | Conceptual consolidation and critical surveys are still needed |
| Review papers account for only 1.39% | Survey literature is limited relative to technical production | A critical survey with a new research agenda is justified |
| Top keywords emphasize RAG, LLMs, information retrieval, search engines, semantics, and question answering | The field is organized around generic LLM-retrieval integration | Existing vocabulary does not foreground curated thematic paths |
| Knowledge graph appears among major keywords but with lower frequency than general RAG and LLM terms | Graph RAG is visible but not dominant | Graph-based retrieval needs more specialized analysis |
| Medical, legal, education, and compliance-oriented terms appear in the bibliometric landscape | RAG is increasingly applied to high-value structured domains | These domains may benefit from retrieval over curated thematic structures |
| Theme evolution shows expansion from knowledge graphs to broader RAG and LLM themes | The field has broadened quickly but remains generic | Pre-existing thematic organization remains under-theorized |
| No dominant cluster for curated thematic paths or canonical evidence units | Current RAG research has not consolidated around this problem | ThemePath-RAG addresses this missing retrieval setting |
| Method | Retrieval Unit | Structure Source | Main Limitation | ThemePath-RAG Difference |
|---|---|---|---|---|
| Vector RAG | Text chunks | Chunking and embedding | Ignores thematic hierarchy | Retrieves curated thematic paths first |
| BM25 RAG | Terms and documents | Lexical index | Weak semantic structure | Uses lexical scoring for evidence pruning |
| Hybrid RAG | Chunks | Dense and sparse indexes | Still mostly flat | Adds path-level retrieval and pruning |
| GraphRAG | Communities or subgraphs | Auto-generated graph | Broad context and redundancy | Uses curated paths and pruned canonical evidence |
| LightRAG | Local/global graph retrieval | Auto-generated graph | May retrieve noisy neighbors | Avoids all-neighbor retrieval |
| HippoRAG | KG nodes and passages | OpenIE graph | Built for constructed KGs | Uses curated thematic paths instead of rebuilding graph from scratch |
| PathRAG | Relational paths | Indexing graph | General graph paths, not domain-authored thematic paths | Treats expert-authored thematic paths as retrieval routes |
| ThemePath-RAG | Thematic paths and evidence units | Curated thematic corpus | Requires curated structure | Designed for pre-structured canonical corpora |
| Parameter | Vector RAG | ThemePath-RAG |
|---|---|---|
| Evaluation instances | 150 | 150 |
| Canonical evidence unit | Ayat | Ayat |
| Query representation | Indonesian | English translation |
| Ayat representation | Indonesian translation | English translation |
| Thematic-path representation | Not applicable | English thematic paths |
| Path-selection rule | Not applicable | Top- paths |
| Lexical similarity weight | Not applicable | 0.30 |
| Semantic similarity weight | Not applicable | 0.50 |
| Path-similarity weight | Not applicable | 0.20 |
| Initial path-candidate generation | Not applicable | English thematic-path vector search |
| Candidate evidence expansion | Verse-level retrieval | Cypher-based retrieval from selected thematic paths |
| Lexical scoring corpus | Not applicable | Unique English candidate ayat texts per query |
| Semantic and path embedding model | text-embedding- 3-small | text-embedding-3-small |
| Similarity function | Cosine similarity | Cosine similarity |
| Component-score normalization | Not applicable | Per-query min–max normalization across expanded pairs |
| Duplicate handling | Not applicable | Maximum combined score retained per ayat reference |
| Final evidence selection | Top- ayat | Global top- unique ayat |
| Actual retrieved contexts | 3 for all questions | 3 for 149 questions; 2 for 1 question |
| Mean retrieved contexts | 3.00 | 2.99 |
| Mean context words | 77.96 | 103.13 |
| Evaluation framework | RAGAS | RAGAS |
| LLM evaluator | gpt-4o-mini | gpt-4o-mini |
| Primary metric | Context relevance | Context relevance |
| Measure | Vector RAG | ThemePath-RAG |
|---|---|---|
| Number of paired questions | 150 | 150 |
| Mean context relevance score | 0.920 | 0.798 |
| Standard deviation | 0.201 | 0.285 |
| Median context relevance score | 1.000 | 1.000 |
| First quartile | 1.000 | 0.500 |
| Third quartile | 1.000 | 1.000 |
| Minimum score | 0.000 | 0.000 |
| Maximum score | 1.000 | 1.000 |
| Perfect relevance score, | 128 (85.3%) | 94 (62.7%) |
| Zero relevance score, | 1 (0.7%) | 3 (2.0%) |
| Mean retrieved contexts | 3.00 | 2.99 |
| Mean context words | 77.96 | 103.13 |
| Question | Vector RAG | ThemePath-RAG | Interpretation |
|---|---|---|---|
| Bagaimana Musa menjelaskan siapa Tuhan itu kepada Fir’aun? (How did Moses explain who the Lord is to Pharaoh?) | 1.00 | 1.00 | Both methods retrieved evidence directly related to Fir’aun’s question and Musa’s explanation of the Lord. This case shows that thematic-path retrieval can preserve direct verse-level relevance when the selected paths align closely with the question. |
| Bagaimana jawaban Hud kepada kaumnya ketika ia dituduh kurang waras?(How did Hud respond to his people when they accused him of lacking reason?) | 1.00 | 0.50 | Vector RAG retrieved the verse in which Hud explicitly rejects the accusation. ThemePath-RAG retrieved related evidence concerning Hud’s call and his people’s rejection but did not retrieve the most direct verse containing his response. |
| Apa akibat yang disebutkan bagi orang-orang yang zalim?(What consequence is mentioned for the wrongdoers?) | 0.50 | 1.00 | ThemePath-RAG retrieved a more consistently consequence-oriented evidence set. The selected thematic route and final evidence ranking are detailed in Table 7. |
| Panel A. Top- Thematic-Path Retrieval | ||||||
| Rank | Curated Thematic Path | Raw Path Similarity | ||||
| P1 | Story → People → Ashhab al-Ukhdud → The Bad Consequences of Their Crime | 0.539828 | ||||
| P2 | Morals → Destructive Morals → Tyranny → On the Day of Judgment, the Wrongdoers Will Be in Fear | 0.444480 | ||||
| P3 | Sharia → Fiqh Mu’amalah → Condemnation for Those Who Cheat | 0.414062 | ||||
| Panel B. Global Top- Evidence Ranking After Candidate Expansion and Deduplication | ||||||
| Rank | Selected Ayat and Abbreviated Evidence | Path | Score | |||
| 1 | Qur’an 42:22: wrongdoers are fearful of what they have earned, and the consequence will inevitably befall them | P2 | 1.000000 | 1.000000 | 0.241862 | 0.848372 |
| 2 | Qur’an 85:10: those who commit crimes against believing men and women and do not repent face the punishment of Hell and the Burning Fire | P1 | 0.000000 | 0.933415 | 1.000000 | 0.666707 |
| 3 | Qur’an 83:1: warning of woe for those who give less than due | P3 | 0.000000 | 0.902207 | 0.000000 | 0.451103 |
| Panel C. Vector RAG Top- Retrieval | ||||||
| Rank | Retrieved Ayat and Abbreviated Evidence | Vector Score | ||||
| 1 | Qur’an 6:21: wrongdoers who fabricate falsehood or deny Allah’s signs do not prosper | 0.684167 | ||||
| 2 | Qur’an 24:50: a statement identifying a group as wrongdoers, without directly describing their consequence | 0.682845 | ||||
| 3 | Qur’an 10:54: wrongdoers witness punishment and conceal their regret | 0.682791 | ||||
| Domain | Curated Path Example | Canonical Evidence Unit | Pruning Criteria | Primary Risk |
|---|---|---|---|---|
| Library classification | Knowledge domain → class → subject heading → collection | Bibliographic record, abstract, subject heading, full-text passage | Subject overlap, metadata proximity, source type, publication context | Overbroad recommendation |
| Indonesian legal codes | Legal area → law → chapter → article → clause | Article, paragraph, clause, explanatory note, SOP step | Article specificity, jurisdiction, clause match, authority level | Wrong legal citation or advice |
| Medical guidelines | Disease category → diagnosis → treatment → contraindication | Guideline statement, evidence grade, clinical note | Patient constraints, guideline section, evidence grade, contraindication match | Unsafe medical suggestion |
| Education and curriculum | Competency → topic → subtopic → learning outcome | Learning outcome, module section, assessment item | Grade level, prerequisite, Bloom level, topic match | Mismatched difficulty or objective |
| Policy and compliance | Policy area → regulation → requirement → procedure | Regulation clause, audit item, required document | Actor, obligation type, effective date, procedure match | Misstated obligation |
| Dimension | Example Metrics |
|---|---|
| Current PoC metric | RAGAS context relevance |
| Retrieval quality | Precision@k, Recall@k, MRR, nDCG@k, gold evidence hit rate |
| Evidence efficiency | Context length, context reduction ratio, evidence redundancy ratio |
| Answer grounding | Faithfulness, groundedness, hallucination rate |
| Citation quality | Citation accuracy, source distinction accuracy |
| Domain sensitivity | Evidence sufficiency, unsupported-claim rate, interpretive caution |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Monika, W.; Dewi, D.A.; Nasution, A.H.; Onan, A.; Murakami, Y. Retrieval-Augmented Generation for Curated Thematic Corpora: A Critical Survey, Bibliometric Evidence, and the ThemePath-RAG Framework. Information 2026, 17, 660. https://doi.org/10.3390/info17070660
Monika W, Dewi DA, Nasution AH, Onan A, Murakami Y. Retrieval-Augmented Generation for Curated Thematic Corpora: A Critical Survey, Bibliometric Evidence, and the ThemePath-RAG Framework. Information. 2026; 17(7):660. https://doi.org/10.3390/info17070660
Chicago/Turabian StyleMonika, Winda, Deshinta Arrova Dewi, Arbi Haza Nasution, Aytuğ Onan, and Yohei Murakami. 2026. "Retrieval-Augmented Generation for Curated Thematic Corpora: A Critical Survey, Bibliometric Evidence, and the ThemePath-RAG Framework" Information 17, no. 7: 660. https://doi.org/10.3390/info17070660
APA StyleMonika, W., Dewi, D. A., Nasution, A. H., Onan, A., & Murakami, Y. (2026). Retrieval-Augmented Generation for Curated Thematic Corpora: A Critical Survey, Bibliometric Evidence, and the ThemePath-RAG Framework. Information, 17(7), 660. https://doi.org/10.3390/info17070660

