Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing
Abstract
1. Introduction
2. Identifying Relevant Literature
3. The Black Box Problem
4. Reference Hallucination Score
5. Persistent Citation Errors Despite Advancements
6. Understanding How LLMs Work
7. Consequences for Academic Integrity
8. Strategies to Tackle the Challenges
8.1. Strengthening Verification Protocols
8.2. Enhancing Transparency in LLMs
8.3. Embedding AI Ethics and Literacy into Education
8.4. Establishing Institutional and Editorial Policies
8.5. Recognizing and Responding to Broader Ethical Implications
9. Limitations
10. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| ChatGPT | Chat Generative Pre-trained Transformer |
| DOI | Digital Object Identifier |
| LLM | Large Language Model |
| NLP | Natural Language Processing |
| RHS | Reference Hallucination Score |
References
- Bajwa, J.; Munir, U.; Nori, A.; Williams, B. Artificial Intelligence in Healthcare: Transforming the Practice of Medicine. Future Healthc. J. 2021, 8, e188–e194. [Google Scholar] [CrossRef] [Scilit]
- Nayak, P.; Gogtay, N. Large Language Models and the Future of Academic Writing. J. Postgrad. Med. 2024, 70, 67–68. [Google Scholar] [CrossRef] [Scilit]
- Kobak, D.; González-Márquez, R.; Horvát, E.-Á.; Lause, J. Delving into LLM-Assisted Writing in Biomedical Publications through Excess Vocabulary. Sci. Adv. 2025, 11, eadt3813. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mishra, T.; Sutanto, E.; Rossanti, R.; Pant, N.; Ashraf, A.; Raut, A.; Uwabareze, G.; Oluwatomiwa, A.; Zeeshan, B. Use of Large Language Models as Artificial Intelligence Tools in Academic Research and Publishing among Global Clinical Researchers. Sci. Rep. 2024, 14, 31672. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fernández-Pichel, M.; Pichel, J.C.; Losada, D.E. Evaluating Search Engines and Large Language Models for Answering Health Questions. NPJ Digit. Med. 2025, 8, 153. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hou, W.; Ji, Z. Comparing Large Language Models and Human Programmers for Generating Programming Code. Adv. Sci. 2025, 12, e2412279. [Google Scholar] [CrossRef] [Scilit]
- Ghosh, A.; Li, H.; Trout, A.T. Large Language Models Can Help with Biostatistics and Coding Needed in Radiology Research. Acad. Radiol. 2025, 32, 604–611. [Google Scholar] [CrossRef] [Scilit]
- Cheng, A.; Nagesh, V.; Eller, S.; Grant, V.; Lin, Y. Exploring AI Hallucinations of ChatGPT: Reference Accuracy and Citation Relevance of ChatGPT Models and Training Conditions. Simul. Healthc. 2025, 20, 413–418. [Google Scholar] [CrossRef] [Scilit]
- Wu, K.; Wu, E.; Wei, K.; Zhang, A.; Casasola, A.; Nguyen, T.; Riantawan, S.; Shi, P.; Ho, D.; Zou, J. An Automated Framework for Assessing How Well LLMs Cite Relevant Medical References. Nat. Commun. 2025, 16, 3615. [Google Scholar] [CrossRef] [Scilit]
- Linardon, J.; Jarman, H.K.; McClure, Z.; Anderson, C.; Liu, C.; Messer, M. Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models: Experimental Study. JMIR Ment. Health 2025, 12, e80371. [Google Scholar] [CrossRef] [Scilit]
- Ngatuvai, M.; Autrey, C.; McKenny, M.; Elkbuli, A. Significance and Implications of Accurate and Proper Citations in Clinical Research Studies. Ann. Med. Surg. 2021, 72, 102841. [Google Scholar] [CrossRef] [Scilit]
- Masic, I. The Importance of Proper Citation of References in Biomedical Articles. Acta Inform. Med. 2013, 21, 148–155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Santini, A. The Importance of Referencing. J. Crit. Care Med. 2018, 4, 3–4. [Google Scholar] [CrossRef] [Scilit]
- Bhattacharyya, M.; Miller, V.M.; Bhattacharyya, D.; Miller, L.E. High Rates of Fabricated and Inaccurate References in ChatGPT-Generated Medical Content. Cureus 2023, 15, e39238. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ghanem, D.; Zhu, A.R.; Kagabo, W.; Osgood, G.; Shafiq, B. ChatGPT-4 Knows Its A B C D E but Cannot Cite Its Source. JBJS Open Access 2024, 9, e24.00099. [Google Scholar] [CrossRef] [Scilit]
- Gravel, J.; D’Amours-Gravel, M.; Osmanlliu, E. Learning to Fake It: Limited Responses and Fabricated References Provided by ChatGPT for Medical Questions. Mayo Clin. Proc. Digit. Health 2023, 1, 226–234. [Google Scholar] [CrossRef] [Scilit]
- Jaźwińska, K.; Chandrasekar, A. How ChatGPT Search (Mis)Represents Publisher Content. Available online: https://www.cjr.org/tow_center/how-chatgpt-misrepresents-publisher-content.php (accessed on 1 January 2025).
- Lechien, J.R.; Briganti, G.; Vaira, L.A. Accuracy of ChatGPT-3.5 and -4 in Providing Scientific References in Otolaryngology-Head and Neck Surgery. Eur. Arch. Otorhinolaryngol. 2024, 281, 2159–2165. [Google Scholar] [CrossRef] [Scilit]
- Sanchez-Ramos, L.; Lin, L.; Romero, R. Beware of References When Using ChatGPT as a Source of Information to Write Scientific Articles. Am. J. Obstet. Gynecol. 2023, 229, 356–357. [Google Scholar] [CrossRef] [Scilit]
- Walters, W.H.; Wilder, E.I. Fabrication and Errors in the Bibliographic Citations Generated by ChatGPT. Sci. Rep. 2023, 13, 14045. [Google Scholar] [CrossRef] [Scilit]
- Frosolini, A.; Franz, L.; Benedetti, S.; Vaira, L.A.; de Filippis, C.; Gennaro, P.; Marioni, G.; Gabriele, G. Assessing the Accuracy of ChatGPT References in Head and Neck and ENT Disciplines. Eur. Arch. Otorhinolaryngol. 2023, 280, 5129–5133. [Google Scholar] [CrossRef] [Scilit]
- Sallam, M. ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. Healthcare 2023, 11, 887. [Google Scholar] [CrossRef] [Scilit]
- Sawamura, S.; Bito, T.; Ando, T.; Masuda, K.; Kameyama, S.; Ishida, H. Evaluation of the Accuracy of ChatGPT’s Responses to and References for Clinical Questions in Physical Therapy. J. Phys. Ther. Sci. 2024, 36, 234–239. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ambrosio, L.; Schol, J.; La Pietra, V.A.; Russo, F.; Vadalà, G.; Sakai, D. Threats and Opportunities of Using ChatGPT in Scientific Writing—The Risk of Getting Spineless. JOR Spine 2023, 7, e1296. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Day, T. A Preliminary Investigation of Fake Peer-Reviewed Citations and References Generated by ChatGPT. Prof. Geogr. 2023, 75, 1024–1027. [Google Scholar] [CrossRef] [Scilit]
- Schwartzman, J.D.; Shaath, M.K.; Kerr, M.S.; Green, C.C.; Haidukewych, G.J. ChatGPT Is an Unreliable Source of Peer-Reviewed Information for Common Total Knee and Hip Arthroplasty Patient Questions. Adv. Orthop. 2025, 2025, 5534704. [Google Scholar] [CrossRef] [Scilit]
- Suppadungsuk, S.; Thongprayoon, C.; Krisanapan, P.; Tangpanithandee, S.; Garcia Valencia, O.; Miao, J.; Mekraksakit, P.; Kashani, K.; Cheungpasitporn, W. Examining the Validity of ChatGPT in Identifying Relevant Nephrology Literature: Findings and Implications. J. Clin. Med. 2023, 12, 5550. [Google Scholar] [CrossRef] [Scilit]
- Emsley, R. ChatGPT: These Are Not Hallucinations—They’re Fabrications and Falsifications. Schizophrenia 2023, 9, 52. [Google Scholar] [CrossRef] [Scilit]
- Aljamaan, F.; Temsah, M.-H.; Altamimi, I.; Al-Eyadhy, A.; Jamal, A.; Alhasan, K.; Mesallam, T.A.; Farahat, M.; Malki, K.H. Reference Hallucination Score for Medical Artificial Intelligence Chatbots: Development and Usability Study. JMIR Med. Inform. 2024, 12, e54345. [Google Scholar] [CrossRef] [Scilit]
- Şengül, H.B.; Akın, B.; Kayaalp, M.E.; Sezgin, E.A. Deep Research Capabilities in GPT-5 Thinking and Gemini 2.5 Pro Improve Citation Integrity and Concordance with American Academy of Orthopaedic Surgeons Anterior Cruciate Ligament and Rotator Cuff Guidelines. Knee Surg. Sports Traumatol. Arthrosc. 2026; early view. [CrossRef] [Scilit]
- Chinonso, O.E.; Theresa, A.M.-E.; Aduke, T.C. ChatGPT for Teaching, Learning and Research: Prospects and Challenges. Glob. Acad. J. Humanit. Soc. Sci. 2023, 5, 33–40. [Google Scholar] [CrossRef] [Scilit]
- Dergaa, I.; Chamari, K.; Zmijewski, P.; Ben Saad, H. From Human Writing to Artificial Intelligence Generated Text: Examining the Prospects and Potential Threats of ChatGPT in Academic Writing. Biol. Sport. 2023, 40, 615–622. [Google Scholar] [CrossRef] [Scilit]
- Chelli, M.; Descamps, J.; Lavoué, V.; Trojani, C.; Azar, M.; Deckert, M.; Raynier, J.-L.; Clowez, G.; Boileau, P.; Ruetsch-Chelli, C. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. J. Med. Internet Res. 2024, 26, e53164. [Google Scholar] [CrossRef] [Scilit]
- Aiumtrakul, N.; Thongprayoon, C.; Suppadungsuk, S.; Krisanapan, P.; Miao, J.; Qureshi, F.; Cheungpasitporn, W. Navigating the Landscape of Personalized Medicine: The Relevance of ChatGPT, BingChat, and Bard AI in Nephrology Literature Searches. J. Pers. Med. 2023, 13, 1457. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Athaluri, S.A.; Manthena, S.V.; Kesapragada, V.S.R.K.M.; Yarlagadda, V.; Dave, T.; Duddumpudi, R.T.S. Exploring the Boundaries of Reality: Investigating the Phenomenon of Artificial Intelligence Hallucination in Scientific Writing Through ChatGPT References. Cureus 2023, 15, e37432. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- McGowan, A.; Gui, Y.; Dobbs, M.; Shuster, S.; Cotter, M.; Selloni, A.; Goodman, M.; Srivastava, A.; Cecchi, G.A.; Corcoran, C.M. ChatGPT and Bard Exhibit Spontaneous Citation Fabrication during Psychiatry Literature Search. Psychiatry Res. 2023, 326, 115334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, R.T.; Dang, R.R. ChatGPT in Head and Neck Scientific Writing: A Precautionary Anecdote. Am. J. Otolaryngol. 2023, 44, 103980. [Google Scholar] [CrossRef] [Scilit]
- Mugaanyi, J.; Cai, L.; Cheng, S.; Lu, C.; Huang, J. Evaluation of Large Language Model Performance and Reliability for Citations and References in Scholarly Writing: Cross-Disciplinary Study. J. Med. Internet Res. 2024, 26, e52935. [Google Scholar] [CrossRef] [Scilit]
- Alkaissi, H.; McFarlane, S.I. Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus 2023, 15, e35179. [Google Scholar] [CrossRef] [Scilit]
- Arif, T.B.; Munaf, U.; Ul-Haque, I. The Future of Medical Education and Research: Is ChatGPT a Blessing or Blight in Disguise? Med. Educ. Online 2023, 28, 2181052. [Google Scholar] [CrossRef] [Scilit]
- Zhang, M.; Zhao, T. Citation Accuracy Challenges Posed by Large Language Models. JMIR Med. Educ. 2025, 11, e72998. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Łukasz, K.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Berman, D.S.; Stapleton, A.G. A Path to Natural Language through Tokenisation and Transformers. arXiv 2026, arXiv:2601.03368. [Google Scholar] [CrossRef] [Scilit]
- Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2024, arXiv:2312.10997. [Google Scholar] [CrossRef] [Scilit]
- Maheshwari, H.; Tenneti, S.; Nakkiran, A. CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track); Rehm, G., Li, Y., Eds.; Association for Computational Linguistics: Vienna, Austria, 2025; pp. 310–317. [Google Scholar]
- De Cassai, A.; Dost, B. Concerns Regarding the Uncritical Use of ChatGPT: A Critical Analysis of AI-Generated References in the Context of Regional Anesthesia. Reg. Anesth. Pain. Med. 2024, 49, 378–380. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Choudhury, A.; Chaudhry, Z. Large Language Models and User Trust: Consequence of Self-Referential Learning Loop and the Deskilling of Health Care Professionals. J. Med. Internet Res. 2024, 26, e56764. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alyasiri, O.M.; Salman, A.M.; Akhtom, D.; Salisu, S. ChatGPT Revisited: Using ChatGPT-4 for Finding References and Editing Language in Medical Scientific Articles. J. Stomatol. Oral. Maxillofac. Surg. 2024, 125, 101842. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, Y.; Ma, X.; Chen, W. Augmenting Black-Box LLMs with Medical Textbooks for Biomedical Question Answering. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024; Al-Onaizan, Y., Bansal, M., Chen, Y.-N., Eds.; Association for Computational Linguistics: Miami, FL, USA, 2024; pp. 1754–1770. [Google Scholar]
- Gibney, E. Can Researchers Stop AI Making up Citations? Nature 2025, 645, 569–570. [Google Scholar] [CrossRef] [Scilit]
- Huang, G.; Li, Y.; Jameel, S.; Long, Y.; Papanastasiou, G. From Explainable to Interpretable Deep Learning for Natural Language Processing in Healthcare: How Far from Reality? Comput. Struct. Biotechnol. J. 2024, 24, 362–373. [Google Scholar] [CrossRef] [Scilit]
- Chopra, A.; Rajput, D.S.; Patel, H. Enhancing Medical Decision-Making with ChatGPT and Explainable AI. Int. J. Surg. 2024, 110, 5167–5168. [Google Scholar] [CrossRef] [Scilit]
- Dabis, A.; Csáki, C. AI and Ethics: Investigating the First Policy Responses of Higher Education Institutions to the Challenge of Generative AI. Humanit. Soc. Sci. Commun. 2024, 11, 1006. [Google Scholar] [CrossRef] [Scilit]
- Masters, K. Ethical Use of Artificial Intelligence in Health Professions Education: AMEE Guide No. 158. Med. Teach. 2023, 45, 574–584. [Google Scholar] [CrossRef] [Scilit]
- Cheng, K.; Wu, H. Policy Framework for the Utilization of Generative AI. Crit. Care 2024, 28, 128. [Google Scholar] [CrossRef] [Scilit]
- Smith, G.R.; Bello, C.; Bialic-Murphy, L.; Clark, E.; Delavaux, C.S.; Fournier de Lauriere, C.; van den Hoogen, J.; Lauber, T.; Ma, H.; Maynard, D.S.; et al. Ten Simple Rules for Using Large Language Models in Science, Version 1.0. PLoS Comput. Biol. 2024, 20, e1011767. [Google Scholar] [CrossRef]
- Tools Such as ChatGPT Threaten Transparent Science; Here Are Our Ground Rules for Their Use. Nature 2023, 613, 612. Available online: https://www.nature.com/articles/d41586-023-00191-1 (accessed on 17 February 2026). [CrossRef] [Scilit]
- Resnik, D.B.; Hosseini, M. Disclosing Artificial Intelligence Use in Scientific Research and Publication: When Should Disclosure Be Mandatory, Optional, or Unnecessary? Account. Res. 2025, 33, 2481949. [Google Scholar] [CrossRef] [Scilit]
| Reference Identifier | Item Hallucination Score |
|---|---|
| Erroneous date of publication | 1 |
| Erroneous web link of the paper | 1 |
| Erroneous citation relevance | 1 |
| Erroneous title of the paper | 2 |
| Erroneous digital object identifier | 2 |
| Erroneous authors’ names | 2 |
| Erroneous name of the journal | 2 |
| Total Reference Hallucination Score (RHS) | 11 |
| Author, Year | Type of Article | Main Findings | Summary of Discussion | Recommendations |
|---|---|---|---|---|
| Sallam, 2023 [22] | Systematic review |
| The imminent dominant use of large language models like ChatGPT in healthcare education, research, and practice is inevitable, but appropriate guidelines and regulations are urgently needed to ensure their safe and responsible use, while also weighing their potential benefits against the possible risks. |
|
| Chinonso et al., 2023 [31] | Literature review |
| The paper discusses the prospects and challenges of using ChatGPT. Authors find the model to be verbose, often over-explaining or repeating certain phrases. It infers user intent instead of asking clarifying questions, leading to potential misunderstandings or incorrect answers. Occasionally, the model may react to damaging or biased instructions. | Users should avoid over-reliance on AI-generated responses, continue conducting independent research using established tools such as search engines, and ensure that all sources and literary works are properly cited and referenced. |
| Dergaa et al., 2023 [32] | Literature review |
| NLP tools like ChatGPT enhance efficiency through data analysis, summarization, and research support, and may improve accessibility. However, they raise concerns regarding factual inaccuracies, bias, plagiarism, transparency, and overreliance, with implications for research integrity. |
|
| Şengü et al., 2026 [30] | Research Study |
| GPT-5 demonstrated strong baseline citation integrity, whereas Gemini relied heavily on live web search to mitigate hallucinated references. Performance was optimized under structured, citation-enforced prompts, suggesting that real-world unstructured queries may underperform. Clinician oversight remains essential, and live web search introduces temporal variability in outputs. | Individualized training in prompt creation or interface-level systems is required for safe clinical adoption. |
| Linardon et al., 2025 [10] | Research Study |
| Citation fabrication and bibliographic errors were common, with lower fabrication and higher accuracy in widely studied and familiar conditions, and higher errors in less familiar or more specialized topics, showing that subject area and prompt granularity influence LLM citation reliability. |
|
| Cheng et al., 2025 [8] | Research Study |
| ChatGPT-4 and ChatGPT-o1 demonstrated poor performance in generating accurate and relevant references. There was no improvement with prompt training interventions. These limitations arise because LLMs generate responses based on learned text patterns without built-in factchecking. | The responsibility for ensuring accuracy rests with human authors, who must carefully verify all AI-generated content, citations, and references, and clearly disclose the use of generative AI in manuscript preparation. |
| Schwartzman et al., 2025 [26] | Research Study |
| ChatGPT-3.5 was an unreliable source of peer-reviewed literature for total knee and hip arthroplasty, with 64% of citations inaccurate or unverifiable, indicating it can generate convincing responses without consistently citing credible sources. | Medical researchers must externally validate all information and sources provided by ChatGPT. |
| Day, 2023 [25] | Research study |
| ChatGPT can generate fake academic citations and references, highlighting limitations in its use for research. However, it may have applications in tasks that do not require references, provided domain expertise is applied to identify and correct inaccurate information. | Researchers should adopt a cautious and critical approach to using ChatGPT in teaching and research. |
| Frosolini et al., 2023 [21] | Research study |
| ChatGPT 4.0 improved reference reliability compared to version 3.5; however, the persistent generation of erroneous or inexistent citations raises concerns for scientific credibility and evidence-based practice. | Journals and institutions should establish strategies and good-practice guidelines to mitigate bias in AI-assisted scientific writing, and further longitudinal research is needed. |
| Sawamura et al., 2024 [23] | Research Study |
| ChatGPT-4.0’s training dataset includes public information, research articles, and sources like Wikipedia and blogs, which may introduce biases. Citing references is more complex than answering questions, likely contributing to the inaccuracies in the reference output. | ChatGPT should be used cautiously as a support tool rather than a replacement for professional expertise, and further research is needed to improve reference accuracy and reliability. |
| Chelli et al., 2024 [33] | Research study |
| Current LLMs demonstrate insufficient precision and high hallucination rates, making them unreliable as standalone tools for systematic reviews and requiring careful human validation of generated references. |
|
| Bhattacharyya et al., 2023 [14] | Research study |
| The findings emphasize the need for caution when seeking medical information on ChatGPT, as most references provided were fabricated or inaccurate, and individuals should verify information from reliable sources and consult medical professionals for serious health concerns. |
|
| Walters & Wilder, 2023 [20] | Research study |
| Although GPT-4 shows marked improvement over GPT-3.5, persistent fabricated citations and substantive errors limit its reliability for scholarly use. |
|
| Aljamaan et al., 2024 [29] | Research study |
| AI chatbots demonstrated higher hallucination severity when prompted with complex medical scenarios. These findings underscore the need for standardized evaluation tools, such as the proposed Reference Hallucination Score (RHS), to systematically assess and monitor the reliability of AI-generated references. |
|
| Aiumtrakul et al., 2023 [34] | Research study |
| The discussion emphasizes the need for careful vetting of AI-generated references in the medical field, given the varying levels of citation accuracy across different AI chatbots and the potential consequences of inaccuracies on medical decision-making and patient care. |
|
| Suppadungsuk et al., 2023 [27] | Research study |
| The hallucination effect of ChatGPT may vary across studies due to factors like input data, prompts, or algorithm updates. Interestingly, when ChatGPT provided the correct link, all reference components were found to be authentic. This suggests a potential correlation between the presence of a correct link and the accuracy of the reference. | Rigorous validation and clinical critical review of AI-generated references are necessary, and ongoing model updates may help reduce hallucinations and improve reliability. |
| Athaluri et al., 2023 [35] | Research study |
| The discussion highlights AI hallucination as a significant limitation of ChatGPT-3, with potential ethical and legal implications, and emphasizes the need for caution and independent verification of AI-generated references. |
|
| Gravel et al., 2023 [16] | Research study |
| ChatGPT has significant limitations in providing accurate and valid information, especially in the references it provides, highlighting the need for caution, oversight, and human involvement when using such language models for scientific purposes. |
|
| McGowan et al., 2023 [36] | Research study |
| The authors attribute citation fabrication to the probabilistic generative nature of LLMs and raise concerns about implications for scientific integrity. | AI-generated references should be independently verified and not relied upon without scrutiny. |
| Wu & Dang, 2023 [37] | Research study |
| The discussion highlights concern that fabricated citations generated by ChatGPT may mislead researchers and undermine scientific integrity, particularly as AI-generated text becomes increasingly sophisticated. | AI-generated references should be independently verified before use in academic work due to the risk of fabrication. |
| Mugaanyi et al., 2024 [38] | Research study |
| ChatGPT’s performance in generating accurate citations and references varies across academic disciplines, with lower accuracy in the humanities compared to the natural sciences, particularly for generating valid DOIs. |
|
| Alkaissi & McFarlane, 2023 [39] | Editorial | ChatGPT was able to generate coherent scientific text from structured prompts and assist indirectly with reference management (e.g., identifying recurrent PMIDs via generated code). However, it produced fabricated references and confident but factually incorrect content, exemplifying what the authors describe as “artificial hallucination. | The paper discusses the implications of using ChatGPT in scientific writing, highlighting concerns about the integrity and accuracy of the content generated by large language models, and proposes that journal and conference policies be modified to maintain scientific standards and include AI output detectors in the editorial process. | Journals and medical conferences should modify their policies and practices for evaluating scientific manuscripts to maintain rigorous scientific standards, given the potential for ChatGPT to generate inaccurate or fabricated content. |
| Arif et al., 2023 [40] | Editorial |
| While ChatGPT can be useful for certain tasks in medical education and research, such as writing and rephrasing text, it also has significant limitations, including a lack of critical thinking, reasoning, and access to up-to-date medical databases, which raises concerns about its credibility and potential for misuse. | The use of ChatGPT in medical education and research requires careful supervision to prevent misuse and ensure academic integrity. |
| Sanchez-Ramos et al., 2023 [19] | Letter |
| Authors need to verify the accuracy of references provided by ChatGPT before submitting manuscripts, though the limitations of these tools are expected to improve over time. | Authors must ensure the accuracy of LLM-generated outputs before submitting manuscripts for publication. |
| Zhang and Zhao, 2025 [41] | Letter |
| Fictional citations mislead users, undermine academic rigor and credibility, and impede the dissemination of accurate scientific knowledge, requiring coordinated efforts from multiple stakeholders to address. |
|
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Rajaratnam, V.; Omar, U.F.; Kee, K.; Kaliya-Perumal, A.-K. Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing. Standards 2026, 6, 10. https://doi.org/10.3390/standards6010010
Rajaratnam V, Omar UF, Kee K, Kaliya-Perumal A-K. Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing. Standards. 2026; 6(1):10. https://doi.org/10.3390/standards6010010
Chicago/Turabian StyleRajaratnam, Vaikunthan, Usama Farghaly Omar, Kristen Kee, and Arun-Kumar Kaliya-Perumal. 2026. "Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing" Standards 6, no. 1: 10. https://doi.org/10.3390/standards6010010
APA StyleRajaratnam, V., Omar, U. F., Kee, K., & Kaliya-Perumal, A.-K. (2026). Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing. Standards, 6(1), 10. https://doi.org/10.3390/standards6010010

