Knowledge Graphs vs. SQL over Structured EHR Data
Abstract
1. Introduction
2. Related Work
2.1. Text-to-SQL
2.2. Retrieval-Augmented and Full-Context Clinical QA
2.3. Knowledge Graph Approaches and Text-to-Cypher
2.4. Tool-Calling LLM Agents
2.5. LLM-as-Judge Scoring
3. Materials and Methods
3.1. Datasets and Data Representations
3.1.1. Knowledge Graph
3.1.2. Relational Database
3.1.3. Patient Snapshots
3.1.4. Question Bank
3.2. Question Answering
3.2.1. Retrieval Methodologies
3.2.2. Models
3.2.3. Generation and Agent Settings
3.3. Evaluation
3.3.1. LLM-as-Judge Agreement
3.3.2. Deterministic Reference Scoring
3.3.3. Statistical Analysis
4. Results
4.1. Overall Accuracy
4.2. Statistical Significance of Retrieval Approaches
4.3. Performance by Question Type
4.4. Performance Across Cohort Sizes
4.5. Retrieval Gains
4.6. Latency and Failures by Approach and Model
4.7. Cost by Approach and Model
5. Discussion
5.1. Retrieval Benefit Depends on the Model
5.2. Retrieval Approach Comparison
5.2.1. Simple Lookup
5.2.2. Multi-Hop
5.2.3. Temporal
5.2.4. Cohort
5.2.5. Reasoning
5.2.6. Unanswerable
5.3. Cohort-Size Effects
5.4. Latency and Deployment Tradeoffs
5.5. Cost and Accuracy Tradeoffs
5.6. Limitations
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| EHR | Electronic Health Record |
| LLM | Large Language Model |
| QA | Question Answering |
| SQL | Structured Query Language |
| RAG | Retrieval-Augmented Generation |
| FHIR | Fast Healthcare Interoperability Resources |
| MCP | Model Context Protocol |
| KG | Knowledge Graph |
| NLP | Natural Language Processing |
| CI | Confidence Interval |
References
- Du, X.; Zhou, Z.; Wang, Y.; Chuang, Y.-W.; Li, Y.; Yang, R.; Zhang, W.; Wang, X.; Chen, X.; Guan, H.; et al. Testing and Evaluation of Generative Large Language Models in Electronic Health Record Applications: A Systematic Review. medRxiv 2025. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shi, W.; Xu, R.; Zhuang, Y.; Yu, Y.; Zhang, J.; Wu, H.; Zhu, Y.; Ho, J.C.; Yang, C.; Wang, M.D. EHRAgent: Code Empowers Large Language Models for Few-shot Complex Tabular Reasoning on Electronic Health Records. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; Al-Onaizan, Y., Bansal, M., Chen, Y.-N., Eds.; Association for Computational Linguistics: Miami, FL, USA, 2024; pp. 22315–22339. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- De Maio, C.; Fenza, G.; Furno, D.; Grauso, T.; Loia, V. A Multi-Agent Architecture for Privacy-Preserving Natural Language Interaction with FHIR-Based Electronic Health Records. In Proceedings of the 2024 International Conference on Software, Telecommunications and Computer Networks (SoftCOM); IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- De Maio, C.; Fenza, G.; Furno, D.; Grauso, T.; Loia, V. Privacy-Preserving Healthcare Data Interactions: A Multi-Agent Approach Using LLMs. J. Commun. Softw. Syst. 2025, 21, 13–22. [Google Scholar] [CrossRef] [Scilit]
- Lee, G.; Hwang, H.; Bae, S.; Kwon, Y.; Shin, W.; Yang, S.; Seo, M.; Kim, J.-Y.; Choi, E. EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records. arXiv 2026, arXiv:2301.07695. [Google Scholar] [CrossRef] [Scilit]
- Pampari, A.; Raghavan, P.; Liang, J.; Peng, J. emrQA: A Large Corpus for Question Answering on Electronic Medical Records. arXiv 2018, arXiv:1809.00732. [Google Scholar] [CrossRef] [Scilit]
- Jiang, P.; Xiao, C.; Cross, A.; Sun, J. GraphCare: Enhancing Healthcare Predictions with Personalized Knowledge Graphs. arXiv 2024, arXiv:2305.12788. [Google Scholar] [CrossRef] [Scilit]
- Lee, G.; Bach, E.; Yang, E.; Pollard, T.; Johnson, A.; Choi, E.; Jia, Y.; Lee, J.H. FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering. arXiv 2025, arXiv:2509.19319. [Google Scholar] [CrossRef] [Scilit]
- Walonoski, J.; Kramer, M.; Nichols, J.; Quina, A.; Moesel, C.; Hall, D.; Duffett, C.; Dube, K.; Gallagher, T.; McLachlan, S. Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record. J. Am. Med. Inform. Assoc. 2018, 25, 230–238. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, P.; Shi, T.; Reddy, C.K. Text-to-SQL Generation for Question Answering on Electronic Medical Records. arXiv 2020, arXiv:1908.01839. [Google Scholar] [CrossRef] [Scilit]
- Johnson, A.E.W.; Pollard, T.J.; Shen, L.; Lehman, L.H.; Feng, M.; Ghassemi, M.; Moody, B.; Szolovits, P.; Anthony Celi, L.; Mark, R.G. MIMIC-III, a freely accessible critical care database. Sci. Data 2016, 3, 160035. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kweon, S.; Kim, J.; Kwak, H.; Cha, D.; Yoon, H.; Kim, K.; Yang, J.; Won, S.; Choi, E. EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries. arXiv 2024, arXiv:2402.16040. [Google Scholar] [CrossRef] [Scilit]
- Johnson, A.E.W.; Bulgarelli, L.; Shen, L.; Gayles, A.; Shammout, A.; Horng, S.; Pollard, T.J.; Hao, S.; Moody, B.; Gow, B.; et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data 2023, 10, 1. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; Metropolitansky, D.; Ness, R.O.; Larson, J. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv 2025, arXiv:2404.16130. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; Zhu, J.; Qi, Y.; Chen, J.; Xu, M.; Menolascina, F.; Jin, Y.; Grau, V. Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented Generation. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics: Vienna, Austria, 2025; pp. 28443–28467. [Google Scholar] [CrossRef] [Scilit]
- Ozsoy, M.G.; Messallem, L.; Besga, J.; Minneci, G. Text2Cypher: Bridging Natural Language and Graph Databases. arXiv 2024, arXiv:2412.10064. [Google Scholar] [CrossRef] [Scilit]
- Hu, N.; Wu, Y.; Qi, G.; Min, D.; Chen, J.; Pan, J.Z.; Ali, Z. An Empirical Study of Pre-trained Language Models in Simple Knowledge Graph Question Answering. arXiv 2023, arXiv:2303.10368. [Google Scholar] [CrossRef] [Scilit]
- Tan, Y.; Zhang, X.; Chen, Y.; Ali, Z.; Hua, Y.; Qi, G. CLRN: A reasoning network for multi-relation question answering over Cross-lingual Knowledge Graphs. Expert Syst. Appl. 2023, 231, 120721. [Google Scholar] [CrossRef] [Scilit]
- Khan, A.; Ali, Z.; Aziz, A.; Kefalas, P. Talk2Doc: A Patient Q&A system using Retrieval-Augmented Generation with Weighted Knowledge Graphs and LLMs. In Proceedings of the 21st International Conference on Intelligent Computing (ICIC 2025), Ningbo, China, 26–29 July 2025. [Google Scholar] [CrossRef]
- Ali, Z.; Huang, Y.; Khan, A.; Qi, G.; Zhang, Y.; Feng, J.; Deng, C.; Kefalas, P. Pythia-RAG: Retrieval-augmented generation over a unified multimodal knowledge graph for enhanced QA. Knowl.-Based Syst. 2026, 335, 115200. [Google Scholar] [CrossRef] [Scilit]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. arXiv 2023, arXiv:2210.03629. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Y.; Black, K.C.; Geng, G.; Park, D.; Zou, J.; Ng, A.Y.; Chen, J.H. MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents. arXiv 2025, arXiv:2501.14654. [Google Scholar] [CrossRef] [Scilit]
- Su, X.; Wang, Y.; Gao, S.; Liu, X.; Giunchiglia, V.; Clevert, D.-A.; Zitnik, M. KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA. arXiv 2025, arXiv:2410.04660. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; Zhu, C. G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; Bouamor, H., Pino, J., Bali, K., Eds.; Association for Computational Linguistics: Singapore, 2023; pp. 2511–2522. [Google Scholar] [CrossRef] [Scilit]
- Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv 2023, arXiv:2306.05685. [Google Scholar] [CrossRef] [Scilit]
- Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; Szolovits, P. What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams. arXiv 2020, arXiv:2009.13081. [Google Scholar] [CrossRef] [Scilit]
- Feng, X.; Jin, G.; Chen, Z.; Liu, C.; Salihoğlu, S. KÙZU Graph Database Management System. In Proceedings of the 13th Annual Conference on Innovative Data Systems Research (CIDR’23), Amsterdam, The Netherlands, 8–11 January 2023. [Google Scholar]
- Introducing Claude Haiku 4.5. Available online: https://www.anthropic.com/news/claude-haiku-4-5 (accessed on 22 May 2026).
- Yang, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Wei, H.; et al. Qwen2.5 Technical Report. arXiv 2025, arXiv:2412.15115. [Google Scholar] [CrossRef] [Scilit]
- Grattafiori, A.; Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Vaughan, A.; et al. The Llama 3 Herd of Models. arXiv 2024, arXiv:2407.21783. [Google Scholar] [CrossRef] [Scilit]
- Friedman, M. The Use of Ranks to Avoid the Assumption of Normality Implicit in the Analysis of Variance. J. Am. Stat. Assoc. 1937, 32, 675–701. [Google Scholar] [CrossRef]
- Cliff, N. Dominance statistics: Ordinal analyses to answer ordinal questions. Psychol. Bull. 1993, 114, 494–509. [Google Scholar] [CrossRef]
- Vezakis, I.A.; Lambrou, G.I.; Matsopoulos, G.K. Deep Learning Approaches to Osteosarcoma Diagnosis and Classification: A Comparative Methodological Approach. Cancers 2023, 15, 2290. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vezakis, A.; Vezakis, I.; Petropoulou, O.; Miloulis, S.T.; Anastasiou, A.; Kakkos, I.; Matsopoulos, G.K. Comparative Analysis of Deep Neural Networks for Automated Ulcerative Colitis Severity Assessment. Bioengineering 2025, 12, 413. [Google Scholar] [CrossRef] [Scilit] [PubMed]








| Resource | 200 Patients | 2000 Patients | 20,000 Patients |
|---|---|---|---|
| Patients | 200 | 2000 | 20,000 |
| Encounters | 9028 | 97,091 | 969,499 |
| Conditions | 6297 | 67,253 | 657,619 |
| Medications | 7196 | 74,788 | 696,078 |
| Observations | 110,349 | 1,250,228 | 11,977,238 |
| Procedures | 24,981 | 276,189 | 2,739,587 |
| Providers | 1183 | 1183 | 1183 |
| Organizations | 1183 | 1183 | 1183 |
| Model | Subset | Best Approach | Best Score | Pooled Score |
|---|---|---|---|---|
| Claude Haiku 4.5 | 200 | sql-t2s | 0.885 | 0.736 |
| Claude Haiku 4.5 | 2000 | sql-t2s | 0.806 | 0.690 |
| Claude Haiku 4.5 | 20,000 | graph-cypher | 0.837 | 0.715 |
| Qwen 2.5 72B | 200 | sql-t2s | 0.736 | 0.600 |
| Qwen 2.5 72B | 2000 | sql-t2s | 0.730 | 0.619 |
| Qwen 2.5 72B | 20,000 | sql-t2s | 0.766 | 0.655 |
| Llama 3.1 8B | 200 | sql-fts | 0.676 | 0.304 |
| Llama 3.1 8B | 2000 | sql-fts | 0.646 | 0.366 |
| Llama 3.1 8B | 20,000 | sql-fts | 0.592 | 0.314 |
| Llama 3.3 70B | 200 | llm-only | 0.502 | 0.436 |
| Llama 3.3 70B | 2000 | llm-only | 0.498 | 0.404 |
| Llama 3.3 70B | 20,000 | graph | 0.579 | 0.477 |
| Method | Haiku 4.5 | Qwen 72B | Llama 3.3 70B | Llama 3.1 8B |
|---|---|---|---|---|
| Graph | 16.85 | 1.05 | 2.17 | 0 |
| Graph-cypher | 14.68 | 1.45 | 2.35 | 0 |
| Sql-t2s | 13.31 | 1.41 | 1.26 | 0 |
| Sql-fts | 5.54 | 0.24 | 0.38 | 0 |
| llm-only | 4.86 | 0.48 | 0.75 | 0 |
| rag-dense (est.) | 4.50 | 0.40 | 0.60 | 0 |
| Method | Haiku 4.5 | Qwen 72B | Llama 3.3 70B | Llama 3.1 8B |
|---|---|---|---|---|
| Graph | 0.064 | 0.005 | 0.013 | 0 |
| Graph-cypher | 0.053 | 0.007 | 0.019 | 0 |
| Sql-t2s | 0.047 | 0.006 | 0.008 | 0 |
| Sql-fts | 0.036 | 0.001 | 0.003 | 0 |
| llm-only | 0.023 | 0.002 | 0.004 | 0 |
| rag-dense (est.) | 0.035 | 0.004 | 0.006 | 0 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Anagnou, L.; Vezakis, A.; Vezakis, I.; Kakkos, I.; Petropoulou, O.; Matsopoulos, G.K. Knowledge Graphs vs. SQL over Structured EHR Data. Future Internet 2026, 18, 365. https://doi.org/10.3390/fi18070365
Anagnou L, Vezakis A, Vezakis I, Kakkos I, Petropoulou O, Matsopoulos GK. Knowledge Graphs vs. SQL over Structured EHR Data. Future Internet. 2026; 18(7):365. https://doi.org/10.3390/fi18070365
Chicago/Turabian StyleAnagnou, Leonidas, Andreas Vezakis, Ioannis Vezakis, Ioannis Kakkos, Ourania Petropoulou, and George K. Matsopoulos. 2026. "Knowledge Graphs vs. SQL over Structured EHR Data" Future Internet 18, no. 7: 365. https://doi.org/10.3390/fi18070365
APA StyleAnagnou, L., Vezakis, A., Vezakis, I., Kakkos, I., Petropoulou, O., & Matsopoulos, G. K. (2026). Knowledge Graphs vs. SQL over Structured EHR Data. Future Internet, 18(7), 365. https://doi.org/10.3390/fi18070365

