From Fine-Tuning to Grounding: Retrieval-Augmented Generation for Biomedical LLMs in Research and Clinical Data Infrastructures
Abstract
1. Biomedical LLMs Need Grounding, Not Only Scale
1.1. Introduction and Scope
1.2. Terminology: Databases, Knowledge Bases, Vector Stores, Knowledge Graphs, and RAG Systems
1.3. Grounding Types in Biomedical RAG
1.4. Why RAG Instead of Training or Fine-Tuning
2. Two Anchor Scenarios: From Omics Interpretation to Clinical Documentation
2.1. Anchor Scenario I: RAG for Single-Cell Annotation and Omics Interpretation
2.2. Anchor Scenario II: RAG for EHRs, PDFs, and Clinical Free Text
2.3. Comparative Design Matrix: Why Biomedical RAG Is Not One-Size-Fits-All
3. Developing and Evaluating Biomedical RAG Systems
3.1. Design Choices Across the RAG Lifecycle
- Clinical task: What problem is the system intended to support, such as patient-history summarization, guideline-grounded question answering, trial eligibility screening, or tumor board preparation?
- Target users: Who is expected to use the system, such as physicians, molecular tumor board coordinators, data managers, researchers, or patients?
- Clinical setting: In which setting will the system be used, such as outpatient care, inpatient documentation review, molecular tumor board preparation, retrospective research, or patient communication?
- Source scope: Which source types are allowed, such as EHR notes, structured events, PDFs, laboratory values, guidelines, SOPs, consent documents, or institutional policies?
- Expected output: What should the system produce, such as a cited summary, candidate evidence list, eligibility explanation, uncertainty statement, or structured report?
- Decision influence: Is the output informational only, prioritizing findings, suggesting possible interpretations, or contributing to diagnostic or therapeutic decisions?
- Responsibility model: Who reviews the output, who may act on it, and who remains responsible for final interpretation?
- Escalation and abstention: Under which conditions should the system refuse to answer, request additional information, or escalate to expert review?
- Governance requirements: Which access-control, logging, validation, monitoring, and update procedures are required for the intended use?
3.2. Evaluation and Benchmarking: From Biological Plausibility to Clinical Safety
3.3. Semi-Automated and Agentic RAG Construction
4. From Prototypes to Governed and Future-Ready Biomedical Infrastructures
4.1. Embedding RAG into Biomedical Research and Clinical Data Infrastructures
4.2. Normative Grounding, Regulation, and Deployer-Side Governance
4.3. Open Challenges and Research Agenda
4.4. Future Outlook: RAG-Ready Knowledge Resources and Agent-Readable Reviews
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Gargari, O.K.; Habibi, G. Enhancing medical AI with retrieval-augmented generation: A mini narrative review. Digit. Health 2025, 11, 20552076251337177. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kumar, D.; Choudhary, N.; Gupta, V.; Tanwar, R.; Behmani, K.; Antal, S.; Jarora, S.; Gupta, S.; Shaikh, M.S.; Webster, T.J.; et al. Natural Language Processing in Healthcare: From Unstructured Data to Clinical Intelligence. Intell. Syst. Appl. 2026, 31, 200698. [Google Scholar] [CrossRef] [Scilit]
- Lin, X.; Deng, G.; Li, Y.; Ge, J.; Ho, J.W.K.; Liu, Y. GeneRAG: Enhancing Large Language Models with Gene-Related Task by Retrieval-Augmented Generation. bioRxiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Artsi, Y.; Sorin, V.; Glicksberg, B.S.; Korfiatis, P.; Nadkarni, G.N.; Klang, E. Large language models in real-world clinical workflows: A systematic review of applications and implementation. Front. Digit. Health 2025, 7, 1659134. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Li, X.; Peng, L.; Wang, Y.-P.; Zhang, W. Open challenges and opportunities in federated foundation models towards biomedical healthcare. BioData Min. 2025, 18, 2. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Anisuzzaman, D.M.; Malins, J.G.; Friedman, P.A.; Attia, Z.I. Fine-Tuning Large Language Models for Specialized Use Cases. Mayo Clin. Proc. Digit. Health 2025, 3, 100184. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Oche, A.J.; Folashade, A.G.; Ghosal, T.; Biswas, A. A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions. arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Amugongo, L.M.; Mascheroni, P.; Brooks, S.; Doering, S.; Seidel, J. Retrieval augmented generation for large language models in healthcare: A systematic review. PLoS Digit. Health 2025, 4, e0000877. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Neha, F.; Bhati, D.; Shukla, D.K. Retrieval-Augmented Generation (RAG) in Healthcare: A Comprehensive Review. AI 2025, 6, 226. [Google Scholar] [CrossRef] [Scilit]
- Zandigohar, M.; Dai, Y. RAGulate: Retrieval-Augmented Generation for Post-hoc Literature-Grounded Regulatory Assessment. bioRxiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Gupta, S.; Ranjan, R.; Singh, S.N. A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Irany, F.A.; Akwafuo, S. A Hybrid Retrieval and Reranking Framework for Evidence-Grounded Retrieval-Augmented Generation. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Thaker, N.G.; Liu, W.; Waddle, M.; Showalter, T.; Mastroleo, F.; Luh, J.; Levitt, C.; Ning, M.; Loaiza-Bonilla, A.; Hong, J. Retrieval-Augmented Generation in Oncology: Promises, Pitfalls, and Early Applications. AI Precis. Oncol. 2026, 3, 34–45. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Mikkelsen, Y. Clinical Context Variables Collectively Rival Model Choice in Embedding-Based Retrieval: Multi-Corpus Benchmark Study. JMIR Med. Inform. 2026, 14, e94241. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Niyonkuru, E.; Gomez, M.S.; Casarighi, E.; Antogiovanni, S.; Blau, H.; Reese, J.T.; Valentini, G.; Robinson, P.N. Replacing non-biomedical concepts improves embedding of biomedical concepts. PLoS ONE 2025, 20, e0322498. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Yang, C.; Zhang, X.; Chen, J. Large language model consensus substantially improves the cell type annotation accuracy for scRNA-seq data. Commun. Biol. 2026. Online ahead of print. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lopez, I.; Swaminathan, A.; Vedula, K.; Narayanan, S.; Nateghi Haredasht, F.; Ma, S.P.; Liang, A.S.; Tate, S.; Maddali, M.; Gallo, R.J.; et al. Clinical entity augmented retrieval for clinical information extraction. npj Digit. Med. 2025, 8, 45. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Berman, E.; Sundberg Malek, H.; Bitzer, M.; Malek, N.; Eickhoff, C. Retrieval Augmented Therapy Suggestion for Molecular Tumor Boards: Algorithmic Development and Validation Study. J. Med. Internet Res. 2025, 27, e64364. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Nouri, N.; Artzi, R.; Savova, V. An agentic AI framework for ingestion and standardization of single-cell RNA-seq data analysis. npj Artif. Intell. 2026, 2, 8. [Google Scholar] [CrossRef] [Scilit]
- Baek, S.; Song, K.; Lee, I. Single-cell foundation models: Bringing artificial intelligence into cell biology. Exp. Mol. Med. 2025, 57, 2169–2181. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chiu, H.-H.; Varghese, A.; Shao, K.; Lu, Y.-C.; Nahar, R.; Chen, H.; Deng, Q.; Bao, X.; Li, C. scChat: A large language model-powered co-pilot for contextualized single-cell RNA sequencing analysis. AIChE J. 2026, 72, e70285. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yu, Z.; Zheng, C.; Chen, C.; Hua, X.-S.; Luo, X. scRAG: Hybrid Retrieval-Augmented Generation for LLM-based Cross-Tissue Single-Cell Annotation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025; Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; Association for Computational Linguistics: Vienna, Austria, 2025; pp. 954–970. [Google Scholar]
- Daw, R.H.; Deijnen, H.R.; Rattray, M.; Grainger, J.R. CellTypeAI: Cell annotation for scRNA-seq using local generative-AI. Bioinformatics 2026, 42, btag425. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xie, E.; Cheng, L.; Shireman, J.; Cai, Y.; Liu, J.; Mohanty, C.; Dey, M.; Kendziorski, C. CASSIA: A multi-agent large language model for automated and interpretable cell annotation. Nat. Commun. 2025, 17, 389. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ou, J.; Huang, T.; Zhao, Y.; Yu, Z.; Lu, P.; Shen, Y.; Ying, R. Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QA. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Liakata, M., Moreira, V.P., Zhang, J., Jurgens, D., Eds.; Association for Computational Linguistics: San Diego, CA, USA, 2026; pp. 23406–23430. [Google Scholar]
- Bazzo, G.T.; Lorentz, G.A.; Suarez Vargas, D.; Moreira, V.P. Assessing the Impact of OCR Errors in Information Retrieval. Adv. Inf. Retr. 2020, 12036, 102–109. [Google Scholar] [CrossRef] [Scilit]
- Abo El-Enen, M.; Saad, S.; Nazmy, T. A survey on retrieval-augmentation generation (RAG) models for healthcare applications. Neural Comput. Appl. 2025, 37, 28191–28267. [Google Scholar] [CrossRef] [Scilit]
- Cao, L.; Chen, Q.; Guo, Y. EHR-RAG: Bridging Long-Horizon Structured Electronic Health Records and Large Language Models via Enhanced Retrieval-Augmented Generation. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Y.; Ren, C.; Wang, Z.; Zheng, X.; Xie, S.; Feng, J.; Zhu, X.; Li, Z.; Ma, L.; Pan, C. EMERGE: Enhancing Multimodal Electronic Health Records Predictive Modeling with Retrieval-Augmented Generation. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, Boise, ID, USA, 21–25 October 2024; pp. 3549–3559. [Google Scholar]
- Chouhan, A.; Gertz, M. heiDS at ArchEHR-QA 2025: From Fixed-k to Query-dependent-k for Retrieval Augmented Generation. In Proceedings of the 24th Workshop on Biomedical Language Processing (Shared Tasks); Soni, S., Demner-Fushman, D., Eds.; Association for Computational Linguistics: Vienna, Austria, 2025; pp. 50–61. [Google Scholar]
- Fahim, Y.A.; Hasani, I.W.; Kabba, S.; Ragab, W.M. Artificial intelligence in healthcare and medicine: Clinical applications, therapeutic advances, and future perspectives. Eur. J. Med. Res. 2025, 30, 848. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Liu, S.; McCoy, A.B.; Wright, A. Improving large language model applications in biomedicine with retrieval-augmented generation: A systematic review, meta-analysis, and clinical development guidelines. J. Am. Med. Inform. Assoc. JAMIA 2025, 32, 605–615. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Ayalew, A.M.; Hasan, M.R.; Seppänen, T.; Oussalah, M. Large Language Models for Explainable Medical Text Summarization: A Systematic Literature Review. WIREs Data Min. Knowl. Discov. 2026, 16, e70089. [Google Scholar] [CrossRef] [Scilit]
- Gomez-Cabello, C.A.; Prabha, S.; Haider, S.A.; Genovese, A.; Collaco, B.G.; Wood, N.G.; Bagaria, S.; Forte, A.J. Comparative Evaluation of Advanced Chunking for Retrieval-Augmented Generation in Large Language Models for Clinical Decision Support. Bioengineering 2025, 12, 1194. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Kabak, Y.; Erturkmen, G.B.L.; Gencturk, M.; Namli, T.; Sinaci, A.A.; Corcoles, R.A.; Ballesteros, C.G.; Abizanda, P.; Dogac, A. FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Support. AI 2026, 7, 246. [Google Scholar] [CrossRef] [Scilit]
- Wijesekara, Y.; Brahma, R.; Lotfi, M.; Vollmer, M.; Kaderali, L. VectorSage: Enhancing article retrieval with advanced semantic search. Bioinform. Adv. 2026, 6, vbag116. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wei, S.; Sasi, C.; Piepenbrock, J.; Huynen, M.A.; ’t Hoen, P.A.C. The use of knowledge graphs for drug repurposing: From classical machine learning algorithms to graph neural networks. Comput. Biol. Med. 2025, 196, 110873. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nanua, S.; Steward, R.; Neely, B.; Datto, M.; Youens, K. Retrieval-augmented generation for interpreting clinical laboratory regulations using large language models. J. Pathol. Inform. 2025, 19, 100520. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Li, D.; Jiang, N.; Huang, K.; Tu, R.; Ouyang, S.; Yu, H.; Qiao, L.; Yu, C.; Zhou, T.; Tong, D.; et al. Streamlining evidence based clinical recommendations with large language models. npj Digit. Med. 2025, 8, 793. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Xu, X.; Weytjens, H.; Zhang, D.; Lu, Q.; Weber, I.; Zhu, L. RAGOps: Operating and Managing Retrieval-Augmented Generation Pipelines. arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Yu, H.; Gan, A.; Zhang, K.; Tong, S.; Liu, Q.; Liu, Z. Evaluation of Retrieval-Augmented Generation: A Survey. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Es, S.; James, J.; Espinosa Anke, L.; Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations; Aletras, N., De Clercq, O., Eds.; Association for Computational Linguistics: St. Julians, Malta, 2024; pp. 150–158. [Google Scholar]
- Saad-Falcon, J.; Khattab, O.; Potts, C.; Zaharia, M. ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers); Duh, K., Gomez, H., Bethard, S., Eds.; Association for Computational Linguistics: Mexico City, Mexico, 2024; pp. 338–354. [Google Scholar]
- Donabauer, P.; Elsweiler, D. Exploring the Impact of Warnings on User Perception towards AI-Generated Content in Search Results. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management; Association for Computing Machinery: New York, NY, USA, 2025; pp. 605–614. [Google Scholar]
- Fletcher, E.; Burns, A.; Wiering, B.; Lavu, D.; Shephard, E.; Hamilton, W.; Campbell, J.L.; Abel, G. Workload and workflow implications associated with the use of electronic clinical decision support tools used by health professionals in general practice: A scoping review. BMC Prim. Care 2023, 24, 23. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Veiner, M.; Supek, F. The DNA dialect: A comprehensive guide to pretrained genomic language models. Mol. Syst. Biol. 2026, 22, 309–332. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Kim, H.; Sohn, J.; Gilson, A.; Cochran-Caggiano, N.; Applebaum, S.; Jin, H.; Park, S.; Park, Y.; Park, J.; Choi, S.; et al. Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights. arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Bressem, K.K.; Papaioannou, J.-M.; Grundmann, P.; Borchert, F.; Adams, L.C.; Liu, L.; Busch, F.; Xu, L.; Loyen, J.P.; Niehues, S.M.; et al. medBERT.de: A comprehensive German BERT model for the medical domain. Expert Syst. Appl. 2024, 237, 121598. [Google Scholar] [CrossRef] [Scilit]
- Arzideh, K.; Schäfer, H.; Idrissi-Yaghir, A.; Schmidt, C.S.; Eryilmaz, B.; Bahn, M.; Turki, A.T.; Pollok, O.B.; Hartmann, E.M.; Winnekens, P.; et al. Improving Retrieval Augmented Generation for Health Care by Fine-Tuning Clinical Embedding Models: Development and Evaluation Study. J. Med. Internet Res. 2026, 28, e82997. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tu, X.; Shi, C.; Qian, P.; Wang, L. Ethical Imperatives for Retrieval-Augmented Generation in Clinical Nursing: Viewpoint on Responsible AI Use. JMIR Med. Inform. 2026, 14, e79922. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Ammann, L.; Ott, S.; Landolt, C.R.; Lehmann, M.P. Securing RAG: A Risk Assessment and Mitigation Framework. In Proceedings of the 2025 IEEE Swiss Conference on Data Science (SDS), Zürich, Switzerland, 26–27 June 2025; pp. 127–134. [Google Scholar]
- Wang, Y.; Cheng, C.; Chang, J.-W. From Documents to Decisions: Enterprise-Grade LLM Systems for Zero-Hallucination, Attributed Generation, and Regulatory Alignment. Comput. Model. Eng. Sci. 2026, 147, 8. [Google Scholar] [CrossRef] [Scilit]
- Singh, A.; Ehtesham, A.; Kumar, S.; Khoei, T.T.; Vasilakos, A.V. Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Farooq, A.; Raza, S.; Karim, N.; Iqbal, H.; Vasilakos, A.V.; Emmanouilidis, C. Evaluating and regulating agentic AI: A study of benchmarks, metrics, and regulation. Inf. Fusion 2026, 136, 104444. [Google Scholar] [CrossRef] [Scilit]
- Tang, Y.; Zhou, Y.; Chen, H. Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Athni, T.S. Emerging Risks of AI-to-AI Interactions in Health Care: Lessons From Moltbook. J. Med. Internet Res. 2026, 28, e96199. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Prabha, S.; Gomez-Cabello, C.A.; Haider, S.A.; Genovese, A.; Trabilsy, M.; Wood, N.G.; Bagaria, S.; Tao, C.; Forte, A.J. Enhancing Clinical Decision Support with Adaptive Iterative Self-Query Retrieval for Retrieval-Augmented Large Language Models. Bioengineering 2025, 12, 895. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Semler, S.C.; Wissing, F.; Heyder, R. German Medical Informatics Initiative. Methods Inf. Med. 2018, 57, e50–e56. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Ahmadi, N.; Peng, Y.; Wolfien, M.; Zoch, M.; Sedlmayr, M. OMOP CDM Can Facilitate Data-Driven Studies for Cancer Prediction: A Systematic Review. Int. J. Mol. Sci. 2022, 23, 11834. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ferber, D.; Hilgers, L.; Höper, C.; Kinny-Köster, B.; Eckardt, J.-N.; Egger-Heidrich, K.; Bill, M.; Schneider, M.M.K.; Clusmann, J.; Kadric, L.; et al. Towards autonomous medical artificial intelligence agents. Nature 2026, 655, 1282–1291. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ardel, H.K.; Randmaa, R.; Bossenko, I.; Piho, G.; Ross, P. Toward bidirectional FHIR–OMOP CDM transformations using TermX to support the secondary use of real-world health data within a patient-centered digital health paradigm. Front. Med. 2026, 13, 1736785. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Saba, W.; Wendelken, S.; Shanahan, J. Question-Answering Based Summarization of Electronic Health Records using Retrieval Augmented Generation. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Ziletti, A.; D’Ambrosi, L. Generating Patient Cohorts from Electronic Health Records Using Two-Step Retrieval-Augmented Text-to-SQL Generation. In Proceedings of the Artificial Intelligence for Healthcare, and Hybrid Models for Coupling Deductive and Inductive Reasoning; Bruno, P., Calimeri, F., Cauteruccio, F., Dragoni, M., Stella, F., Terracina, G., Eds.; Springer Nature: Cham, Switzerland, 2026; pp. 231–242. [Google Scholar]
- Schmiedmayer, P.; Rao, A.; Zagar, P.; Aalami, L.; Ravi, V.; Zahedivash, A.; Yao, D.; Fereydooni, A.; Aalami, O. LLMonFHIR: A Physician-Validated, Large Language Model–Based Mobile Application for Querying Patient Electronic Health Data. JACC Adv. 2025, 4, 101780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lammert, J.; Dreyer, T.; Mathes, S.; Kuligin, L.; Borm, K.J.; Schatz, U.A.; Kiechle, M.; Lörsch, A.M.; Jung, J.; Lange, S.; et al. Expert-Guided Large Language Models for Clinical Decision Support in Precision Oncology. JCO Precis. Oncol. 2024, 8, e2400478. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Timilsina, M.; Buosi, S.; Razzaq, M.A.; Haque, R.; Judge, C.; Curry, E. Harmonizing foundation models in healthcare: A comprehensive survey of their roles, relationships, and impact in artificial intelligence’s advancing terrain. Comput. Biol. Med. 2025, 189, 109925. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rashid, M.T. From Recommender to Actor: The Normative Boundary When RAG Tools Become Tool-Calling Agents. Minds Mach. 2026, 36, 30. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Pei, J.; Huang, J.; Si, C.; Qu, A.; Tang, X.; Lu, R.; Chen, L.; Bai, X.; Zheng, H.; et al. The Last Human-Written Paper: Agent-Native Research Artifacts. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]


| Term | Meaning in This Review | Biomedical Examples | Role in a RAG System |
|---|---|---|---|
| Database | Organized collection of structured or semi-structured data | NCBI Gene, OMOP CDM tables, FHIR-based clinical data, marker-gene tables, medication databases | Provides structured or semi-structured source data |
| Knowledge base | Curated collection of facts, documents, rules, guidelines, metadata, or references | CellMarker, Cell Ontology, institutional SOP repository, clinical guideline collection, atlas resource | Provides curated domain knowledge for retrieval or reasoning |
| Vector store/vector database | Index of embedded chunks or objects for similarity-based retrieval | FAISS, ChromaDB, Weaviate, embedded guideline passages, embedded gene descriptions, embedded report sections | Enables semantic retrieval of relevant content |
| Knowledge graph | Structured representation of entities and relationships | Gene-cell type-tissue graph, disease-drug-phenotype network, PrimeKG, ontology graph | Enables relation-aware retrieval, structured reasoning, and entity linking |
| RAG system | Full retrieval-and-generation architecture that grounds an LLM in external sources | Single-cell annotation RAG, guideline-grounded QA, EHR summarization RAG, MTB preparation RAG | Combines source ingestion, retrieval, reranking, prompt construction, generation, citation, and evaluation |
| Grounding | Linking model outputs to external evidence, context, constraints, or workflow-specific information | Literature-grounded annotation, patient-specific EHR retrieval, SOP-based response constraint, citation to PDF page | Makes generated outputs more traceable, contextualized, and inspectable |
| Grounding Type | Guiding Question | Single-Cell/Omics Example | Clinical EHR/PDF Example |
|---|---|---|---|
| Factual grounding | What is known? | Marker genes, cell ontology terms, pathway literature, atlas entries | Clinical guidelines, medication databases, publications, disease-specific evidence |
| Contextual grounding | What is relevant here? | Tissue, species, disease state, assay platform, cluster markers | Patient, encounter, date, document type, clinical question, guideline version |
| Analytical grounding | What was computed? | Differential expression, enrichment results, trajectory analysis, cluster statistics | Laboratory trends, longitudinal events, risk scores, structured EHR features |
| Provenance grounding | Where does the answer come from? | Cited database entry, atlas version, publication, analysis output | Cited note passage, PDF page, report section, FHIR resource, guideline version |
| Normative grounding | What is allowed or appropriate? | Data-use limitations, lab annotation rules, sharing restrictions | Consent, SOPs, access rights, GDPR/AI Act constraints, clinical responsibility boundaries |
| Operational grounding | How is this used in practice? | Scanpy/Jupyter/Nextflow integration, lab repositories, annotation workflow | DIZ, FHIR/OMOP, EHR systems, document archives, MTB platforms, audit logs |
| Design Dimension | Single-Cell Annotation and Omics Interpretation | Clinical EHR, PDF, and Free-Text Integration |
|---|---|---|
| Primary purpose | Biological interpretation and annotation support | Patient-specific information synthesis and workflow support |
| Main data sources | Marker databases, cell ontologies, atlases, pathway databases, the literature, analysis outputs | EHR notes, structured events, PDFs, reports, guidelines, SOPs, PROs, institutional policies |
| Typical retrieval unit | Gene description, marker set, ontology term, cell type entry, publication passage, analysis result | Encounter, note passage, report section, PDF page, lab event, medication record, guideline passage |
| Main grounding types | Factual, contextual, analytical, provenance | Contextual, provenance, normative, operational |
| Evaluation focus | Expert annotation, hierarchical F1, marker consistency, ontology mapping, biological plausibility | Factuality, completeness, attribution, omission rate, temporal correctness, non-harmfulness |
| Infrastructure setting | Scanpy/Jupyter workflows, research repositories, lab knowledge bases, atlas resources | EHR systems, DIZ, FHIR/OMOP, document archives, MTB platforms, audit systems |
| Governance focus | Reproducibility, data reuse, source quality, benchmark transparency | Consent, access control, accountability, clinical responsibility, regulation |
| Human role | Expert biological validation and correction | Clinician review, responsibility, and workflow supervision |
| Lifecycle Step | Key Design Question | Single-Cell/Omics Example | Clinical EHR/PDF Example |
|---|---|---|---|
| Intended use | What task should the system support? | Cluster annotation, marker interpretation, pathway explanation | Patient-history summary, guideline QA, MTB preparation |
| Source curation | Which sources are trusted and current? | Cell Ontology, CellMarker, atlases, the literature | EHR notes, reports, guidelines, SOPs, PDFs |
| Preprocessing | How are sources transformed into retrievable units? | Marker sets, ontology terms, pathway tables | Note sections, PDF pages, lab events, report passages |
| Indexing | How are sources represented technically? | Vector store, marker table, knowledge graph | Vector store, FHIR/OMOP query layer, document index |
| Retrieval | How is relevant evidence selected? | Tissue-, species-, and gene-aware retrieval | Patient-, date-, encounter-, and access-aware retrieval |
| Reranking | Which evidence should be prioritized? | Most specific marker/cell-type evidence | Most recent and clinically relevant document passages |
| Context construction | How is retrieved evidence given to the LLM? | Cluster markers plus supporting ontology/literature | Patient context plus cited reports/guidelines |
| Generation | What should the model produce? | Candidate annotation with uncertainty | Traceable clinical summary or answer |
| Provenance | How can users inspect the evidence? | Marker/database/publication citations | Note/PDF/guideline citations |
| Monitoring | How is the system maintained? | Update atlases and marker sources | Update guidelines, documents, access rights, logs |
| Grounding Type | Evaluation Target | Possible Metrics or Checks | Example Failure Mode |
|---|---|---|---|
| Factual grounding | Correctness of retrieved and generated factual claims | Factuality score, FactScore, expert correctness rating, unsupported-claim rate | The answer states a marker-cell association or guideline recommendation not supported by the retrieved sources |
| Contextual grounding | Fit between retrieved evidence and biological or clinical context | Context precision, patient/encounter/date matching, tissue/species/disease-state match, metadata-filter accuracy | The system retrieves evidence from the wrong tissue, species, patient, encounter, or guideline version |
| Analytical grounding | Faithfulness to computed results or structured data | Recalculation checks, agreement with differential expression/pathway/lab/risk-score outputs, structured-data consistency | The response contradicts the underlying analysis output or laboratory trend |
| Provenance grounding | Traceability of claims to inspectable sources | Citation accuracy, source-support rate, page/section match, evidence attribution, retrieval recall | The citation points to an irrelevant note section, PDF page, or publication passage |
| Normative grounding | Consistency with rules, consent, policies, and allowed use | Policy/rule match, access-right verification, consent-status check, intended-use compliance, audit-log review | The system retrieves or summarizes information that the user is not authorized to access |
| Operational grounding | Fit with workflow and deployment environment | Usability, task completion time, workload, integration success, audit completeness, user-feedback review | Correct information is generated but appears too late, is hard to verify, or disrupts the clinical workflow |
| Challenge | Why It Matters | Priority Direction |
|---|---|---|
| Source curation | RAG can only retrieve what has been made available; missing or outdated sources can directly limit answer quality | Versioned source registries, source-quality criteria, update cycles, provenance standards |
| Context-aware retrieval | Biomedical similarity is not generic similarity; tissue, species, patient, encounter, date, and access rights may determine relevance | Metadata-aware, ontology-aware, time-aware, patient-aware, and hybrid retrieval strategies |
| Realistic benchmarking | Generic QA benchmarks do not capture biological ambiguity or clinical safety requirements | Task-specific benchmarks for rare cell types, disease states, longitudinal EHRs, PDFs, conflicting sources, and multilingual settings |
| Uncertainty and abstention | A fluent answer may be unsafe when retrieval is incomplete or sources conflict | Calibrated uncertainty, explicit insufficiency statements, non-answer behavior, escalation to expert review |
| Agentic RAG validation | Multiple calls in multi-agent systems may increase latency and computational cost | Evaluation of intermediate decisions, routing, tool calls, self-checks, reproducibility, and orchestration |
| Infrastructure integration | Prototype RAG systems do not automatically translate into real research or clinical environments | Integration with analysis workflows, DIZ, FHIR, OMOP, document archives, consent systems, and audit logs |
| Governance and intended use | The same architecture may be low risk in research support but high risk in clinical decision support | Intended-use definitions, medical-device boundary clarification, role-based access, post-deployment monitoring |
| Multilingual and local adaptation | English benchmarks and sources may not reflect German or European clinical documentation and regulation | Local-language corpora, German clinical benchmarks, local terminology, institutional SOPs, and deployer-specific evaluation |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Enayati, M.; Priya, V.; Prochaska, E.; Sobe, K.; Wolfien, M. From Fine-Tuning to Grounding: Retrieval-Augmented Generation for Biomedical LLMs in Research and Clinical Data Infrastructures. Sci 2026, 8, 266. https://doi.org/10.3390/sci8090266
Enayati M, Priya V, Prochaska E, Sobe K, Wolfien M. From Fine-Tuning to Grounding: Retrieval-Augmented Generation for Biomedical LLMs in Research and Clinical Data Infrastructures. Sci. 2026; 8(9):266. https://doi.org/10.3390/sci8090266
Chicago/Turabian StyleEnayati, Mahdi, Vishnu Priya, Eveline Prochaska, Kathrin Sobe, and Markus Wolfien. 2026. "From Fine-Tuning to Grounding: Retrieval-Augmented Generation for Biomedical LLMs in Research and Clinical Data Infrastructures" Sci 8, no. 9: 266. https://doi.org/10.3390/sci8090266
APA StyleEnayati, M., Priya, V., Prochaska, E., Sobe, K., & Wolfien, M. (2026). From Fine-Tuning to Grounding: Retrieval-Augmented Generation for Biomedical LLMs in Research and Clinical Data Infrastructures. Sci, 8(9), 266. https://doi.org/10.3390/sci8090266

