Next Article in Journal
Intelligence for Regional Development in Maranhão, Brazil: Insights from Logistics Process Management
Previous Article in Journal
Assessment and Standards in Hygienic Design of Food Equipment: A Comprehensive Cross-Industry Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing

by
Vaikunthan Rajaratnam
1,
Usama Farghaly Omar
1,
Kristen Kee
2 and
Arun-Kumar Kaliya-Perumal
3,*
1
Department of Orthopaedic Surgery, Hand and Reconstructive Microsurgery Service, Khoo Teck Puat Hospital, Singapore 768828, Singapore
2
Yong Loo Lin School of Medicine, National University of Singapore, Singapore 117594, Singapore
3
Rehabilitation Research Institute of Singapore, Nanyang Technological University, Singapore 308232, Singapore
*
Author to whom correspondence should be addressed.
Standards 2026, 6(1), 10; https://doi.org/10.3390/standards6010010
Submission received: 18 February 2026 / Revised: 12 March 2026 / Accepted: 18 March 2026 / Published: 20 March 2026

Abstract

Generative artificial intelligence (AI)-based large language models (LLMs) are increasingly being used in medical writing to improve efficiency and broaden access to knowledge. However, concerns have emerged regarding the accuracy of the citations they generate. This review discusses the issue of citation inaccuracies in AI-assisted medical writing and its implications for scientific reliability and accountability in academic medicine. Published literature describing citation errors in AI-generated content, particularly in medical and academic contexts, was examined to understand the nature and persistence of this problem and to consider potential safeguards. Reports consistently describe citation inaccuracies, including fabricated references, incorrect bibliographic details, and incomplete source information such as missing authors, journal titles, publication years, or digital object identifiers. Although these tools continue to evolve, such errors remain reported and highlight limitations in their reliability. While LLMs offer clear benefits in supporting medical writing, their outputs require careful verification. As developers continue to address these challenges, responsible use will depend on continued human oversight, improved transparency, greater user awareness, and institutional and policy-level guidance to ensure accurate and trustworthy use of generative AI in medical writing.

1. Introduction

Artificial intelligence (AI) is increasingly becoming a valuable tool in medicine [1], offering significant benefits in clinical decision support, research, data analysis, and academic writing. In the realm of scientific and clinical writing, generative AI-based large language models (LLMs), such as Chat Generative Pre-trained Transformer (ChatGPT), have shown the potential to enhance efficiency, support the writing process, and broaden access to knowledge [2]. These tools are already widely used in clinical documentation and academic preparation across the globe [3,4]. Since their widespread adoption across disciplines, LLMs have been used in diverse applications, ranging from content generation and question answering to code writing and problem-solving [5,6,7]. However, concerns persist regarding the reliability of AI-generated content, particularly the accuracy and authenticity of citations, including the risk of fabricated references [8,9,10].
In medicine, a field that is inherently evidence-based and highly sensitive to accuracy, precise referencing is not merely academic; it can influence clinical decisions and patient outcomes [11,12,13]. Although some LLMs have recently begun attempting to cite sources alongside their responses, their ability to accurately reference peer-reviewed articles from indexed journals for the purpose of academic writing remains limited, particularly in providing complete and correctly formatted details such as all author names, publication year, journal name, volume, issue, page numbers, and digital object identifiers (DOIs). These AI-generated references are often incomplete, incorrectly formatted, or fabricated.
Several studies across various fields have highlighted this issue since the introduction of LLM tools, reporting that LLMs frequently generate inaccurate or fabricated citations and exhibit inconsistent referencing practices [14,15,16,17,18,19,20]. Despite ongoing improvements in AI models, these problems persist and continue to challenge the credibility and reliability of AI-assisted writing. Therefore, human verification of AI-generated citations remains essential, especially in medical and clinical contexts.
Although existing literature often highlights the functional capabilities of LLMs in academic settings, citation accuracy receives comparatively limited attention [15,21,22,23]. Empirical research on the subject is relatively scarce, with many articles identifying citation errors without exploring their broader implications or providing individual, institutional or policy-level recommendations. The prevailing consensus is to approach AI-generated references with caution and verify them manually.

2. Identifying Relevant Literature

Given the limited and heterogeneous nature of existing evidence, a narrative approach is adopted to synthesize current findings, assess the scope of the problem, and explore its ethical and practical implications. We considered a broad range of articles including original studies, reviews, editorials, and commentaries published in academic journals that explicitly discuss the accuracy of LLM-generated citations. Only articles published in English up to early 2026 were reviewed. The search strategy involved querying databases such as PubMed, Scopus, Web of Science, and Google Scholar using a search string tailored to each platform. Keywords included terms related to artificial intelligence and large language models, combined with terms related to citation inaccuracies, such as fabricated citations, bibliography errors, or reference hallucinations. The extracted data were summarized in a tabular format. Thematic analysis was used to synthesize the findings from the discussions and policy recommendations. The review also aims to offer actionable recommendations to help researchers, clinicians, and educators responsibly integrate AI tools into academic writing while maintaining the standards of credibility and integrity.

3. The Black Box Problem

Although there is growing interest in integrating generative AI-based LLMs into academic and medical writing, a significant challenge remains: the ‘black box’ problem, which refers to the lack of transparency in how these models generate content, particularly when they fail to provide verifiable sources such as peer-reviewed journal articles [24]. When prompted for such citations, studies indicate a high prevalence of fabricated, inaccurate, or incomplete references generated by LLMs across various disciplines [21,22,23,25]. Fabricated citations, which refer to non-existent articles, authors, or journals, have been reported in AI-generated texts at rates as high as 60%, posing serious risks to academic integrity by misleading readers and undermining the credibility of scholarly work [14,26,27]. In addition to fabrications, AI-generated content often includes errors in authentic references, such as misspelled author names, incorrect publication dates, and inaccurate article titles. One of the most pervasive issues is erroneous DOIs, which impede the verification process fundamental to scholarly research and discourse [20,28].

4. Reference Hallucination Score

To quantify the problems mentioned above, a reference hallucination score (RHS) was developed by Aljamaan et al. (Table 1), testing six chatbots with 10 medical prompts, each requesting 10 references [29]. RHS evaluates citation errors, with a total score of 11, covering inaccuracies in publication date (1 point), web link (1 point), relevance (1 point), title (2 points), DOI (2 points), authors’ names (2 points), and journal name (2 points). In practice, an RHS close to zero indicates accurate references with no or minimal bibliographic errors, whereas higher scores indicate increasing levels of reference hallucination and unreliable referencing. Unlike simple binary assessments of whether a citation exists, the RHS provides a structured evaluation of the type and severity of bibliographic inaccuracies, allowing comparison of hallucination patterns across different systems. In their 2024 study, Aljamaan et al. reported that Elicit and SciSpace had negligible hallucination levels, while ChatGPT and Bing had critical hallucination levels [29].

5. Persistent Citation Errors Despite Advancements

Despite advancements in newer AI models, significant problems with citation accuracy persist, as consistently reported in the literature. We have tabulated 25 relevant articles, highlighting their main findings, key discussion points, and recommendations regarding citation inaccuracies and responsible use (Table 2). Most of these articles evaluated GPT-3.5, GPT-4, and other earlier models, while the most recent study by Şengül et al. examined GPT-5 Thinking and Gemini 2.5 Pro and reported improved citation integrity [30]. While newer models show measurable improvements in citation accuracy, available evidence remains limited in scope, and reliability across diverse topics cannot yet be assumed. Prior studies demonstrate that citation performance varies widely depending on subject matter, with more frequently represented topics tending to yield higher accuracy than less common or specialized areas [10]. Collectively, these findings stress the need for caution in utilizing AI-generated citations in academic and medical contexts, with a strong emphasis on human oversight to verify accuracy and ensure reliability [22].

6. Understanding How LLMs Work

LLMs do not function like traditional search engines that retrieve verified information. Instead, they generate content through probabilistic prediction. The input is first broken into small units called tokens, and the model processes relationships between these tokens using a transformer neural network and a mechanism known as self-attention [42,43]. Based on patterns learned during training, the model calculates probabilities for possible next tokens, selects the most likely option, and repeats this process sequentially until the response is complete.
Retrieval-Augmented Generation (RAG) has emerged as an important advancement, combining external database retrieval with LLM generation to improve factual grounding and citation accuracy [44]. In RAG systems, relevant documents are retrieved and provided as context before the model generates a response. However, the final output remains a generated synthesis rather than a direct retrieval, and the internal decision pathways are not fully explainable. Even with RAG, LLMs remain probabilistic generative systems, and citation inaccuracies may still occur. Ongoing research is focused on improving citation reliability through post-processing and verification methods [45]. Therefore, external validation of AI-generated citations remains necessary.

7. Consequences for Academic Integrity

The impact of this issue is multifaceted. Firstly, inaccurate citations undermine the validation of arguments, acknowledgment of prior work, and accessibility of sources, thereby compromising academic credibility [25]. Secondly, the fabrication of sources raises severe ethical concerns by introducing non-existent studies into the academic ecosystem, risking the propagation of false information and distorting the scholarly record [14]. Then, hybrid papers that mix accurate and fabricated citations make plagiarism detection and source verification challenging, thereby complicating the maintenance of academic integrity [20]. Finally, careless citation practices, including incomplete references and inconsistent formatting, exacerbate these challenges by failing to meet the foundational standards of scholarly work [46]. The implications of these shortcomings are significant.
Trust in academic research relies on the credibility of sources and the accuracy of citations, and AI-generated inaccuracies threaten this trust [22]. Fabricated or incorrect citations can also lead to legal and ethical violations, exposing institutions and individuals to significant risks [25,47]. Furthermore, complete reliance on AI tools by students and early-career researchers for their writing tasks could instill poor citation habits, compromising their academic development and the quality of future scholarship [46].

8. Strategies to Tackle the Challenges

Effectively addressing the citation accuracy issue and related concerns surrounding LLM-generated academic content requires a structured, multi-pronged response. The following strategies outline key areas of focus:

8.1. Strengthening Verification Protocols

Academics must be encouraged to critically evaluate AI-generated citations and verify them against legitimate, peer-reviewed sources in terms of both accuracy and relevance [48]. The consistently high rates of citation hallucinations and inaccuracies reported across multiple studies underscore the urgent need for rigorous verification at the user level. Cheng et al. further emphasize that the responsibility for ensuring accuracy rests with human authors, who must carefully verify all AI-generated content, citations, and references, and transparently disclose the use of generative AI in the manuscript [8]. Such practices are essential to maintain the credibility of academic outputs and uphold the integrity of research.

8.2. Enhancing Transparency in LLMs

Enhancing transparency in LLMs remains challenging due to their probabilistic generative nature. Although recent approaches such as RAG improve source attribution and factual grounding, complete interpretability remains limited due to the black-box nature of these models [44,49,50]. To improve, developers can prioritize reliable sourcing, leverage verified academic databases, and implement safeguards to identify and flag unverifiable content [14,51]. Such approaches contribute to explainability, enabling users to better assess the reliability of AI-generated information [52]. Ultimately, these measures are essential to support responsible academic use while acknowledging the limitations of generative AI systems.

8.3. Embedding AI Ethics and Literacy into Education

Institutions should embed ethical AI usage into the curriculum, offering training programs that promote academic integrity, citation best practices, and awareness of AI’s limitations [20,53,54]. It is equally important for users to understand that LLMs generate responses based on learned patterns, including their probabilistic token-based generation, rather than directly retrieving verified references. As a result, the information provided may be correct, but the cited reference may be incomplete, inaccurate, or improperly formatted. Awareness of this distinction helps users critically evaluate AI-generated outputs and verify references appropriately. This educational approach is essential to developing a culture of critical engagement with AI tools, rather than blind reliance, and will help reduce the risk of academic dishonesty or the unintentional propagation of errors.

8.4. Establishing Institutional and Editorial Policies

Academic institutions and journals must implement clear, enforceable policies on the use of generative AI in research and publishing. These should include guidelines on citation standards, authorship, data transparency, and consequences for ethical violations [46,55,56]. Such policies are particularly important given the documented risk of inaccurate or fabricated citations generated by AI tools, which may compromise scholarly integrity if left unchecked. Establishing formal frameworks will help ensure appropriate verification of AI-generated content and support responsible integration of these technologies into academic practice.

8.5. Recognizing and Responding to Broader Ethical Implications

In the broader context, the widespread use of generative AI raises important ethical questions regarding authorship, originality, and the nature of scholarly contribution. When AI systems generate text and references, the boundary between human authorship and machine assistance becomes less distinct, raising concerns about academic authenticity. While journals increasingly allow and encourage the use of LLMs as tools to support authors, they also emphasize the need for transparent disclosure and affirm that authors remain fully responsible for the accuracy and integrity of the content [57,58]. This shift is already altering how knowledge is produced, evaluated, and trusted within the scientific community. Therefore, institutions and researchers must foster critical awareness and ethical responsibility to ensure that AI serves as a supportive tool rather than undermining scholarly rigor and academic integrity.

9. Limitations

This narrative review has several inherent limitations. Article selection was guided by relevance to the topic, narrative coherence, and the key messages the author intended to convey. Although a search was carried out across multiple databases and rigorous screening was applied, the review is not systematic, and selection bias may still be present. The included studies vary in design, sample size, methodological rigor, and the version of the LLM used, which can affect the generalizability of the findings. The exclusion of non-English literature may also have affected the comprehensiveness of the review by omitting potentially valuable insights. Furthermore, LLMs are rapidly evolving, with frequent updates improving their accuracy and functionality over time. As such, some earlier findings may become outdated. Newer LLM versions, including ChatGPT-5, remain to be characterized in peer-reviewed literature in terms of citation accuracy; while users may encounter errors or assess performance in practice, systematic and independent evaluations have not yet been sufficiently documented. Nonetheless, this article captures a moment in the ongoing evolution of LLMs, and highlighting these issues is crucial, as critical evaluations contribute to continued improvements in LLMs’ performance and promote responsible development and use.

10. Conclusions

Generative AI-based LLMs offer promising capabilities, but their use in generating academic citations and references continues to present significant challenges, especially in the medical field. This review highlights three key insights. First, multiple studies consistently report substantial rates of fabricated or inaccurate citations, indicating that citation hallucination remains a persistent limitation of current LLMs. Second, although newer models demonstrate improving citation accuracy, reliability remains highly dependent on prompt design, model configuration, and external information retrieval mechanisms. Third, the literature consistently emphasizes the need for human verification and institutional safeguards when AI tools are used in academic writing. Future research should focus on improving verification mechanisms, increasing transparency, and advancing explainability in reference generation. While integrating AI into academic workflows holds great promise for innovation and efficiency, it must be guided by clear ethical standards and rigorous oversight. By addressing these challenges through a proactive and collaborative approach, the academic community can harness the benefits of AI while upholding the principles of credible and ethical scholarship.

Author Contributions

Conceptualization, V.R., U.F.O., K.K. and A.-K.K.-P.; methodology, V.R.; validation, U.F.O. and A.-K.K.-P.; formal analysis, V.R., K.K. and A.-K.K.-P.; resources, V.R., U.F.O. and A.-K.K.-P.; data curation, V.R., K.K. and A.-K.K.-P.; writing—original draft preparation, V.R.; writing—review and editing, U.F.O. and A.-K.K.-P.; visualization, A.-K.K.-P.; supervision, V.R. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

All data generated is contained within the article.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT Version 5.2 (OpenAI) for the purposes of grammar and syntax. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
ChatGPTChat Generative Pre-trained Transformer
DOIDigital Object Identifier
LLMLarge Language Model
NLPNatural Language Processing
RHSReference Hallucination Score

References

  1. Bajwa, J.; Munir, U.; Nori, A.; Williams, B. Artificial Intelligence in Healthcare: Transforming the Practice of Medicine. Future Healthc. J. 2021, 8, e188–e194. [Google Scholar] [CrossRef] [Scilit]
  2. Nayak, P.; Gogtay, N. Large Language Models and the Future of Academic Writing. J. Postgrad. Med. 2024, 70, 67–68. [Google Scholar] [CrossRef] [Scilit]
  3. Kobak, D.; González-Márquez, R.; Horvát, E.-Á.; Lause, J. Delving into LLM-Assisted Writing in Biomedical Publications through Excess Vocabulary. Sci. Adv. 2025, 11, eadt3813. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Mishra, T.; Sutanto, E.; Rossanti, R.; Pant, N.; Ashraf, A.; Raut, A.; Uwabareze, G.; Oluwatomiwa, A.; Zeeshan, B. Use of Large Language Models as Artificial Intelligence Tools in Academic Research and Publishing among Global Clinical Researchers. Sci. Rep. 2024, 14, 31672. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Fernández-Pichel, M.; Pichel, J.C.; Losada, D.E. Evaluating Search Engines and Large Language Models for Answering Health Questions. NPJ Digit. Med. 2025, 8, 153. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Hou, W.; Ji, Z. Comparing Large Language Models and Human Programmers for Generating Programming Code. Adv. Sci. 2025, 12, e2412279. [Google Scholar] [CrossRef] [Scilit]
  7. Ghosh, A.; Li, H.; Trout, A.T. Large Language Models Can Help with Biostatistics and Coding Needed in Radiology Research. Acad. Radiol. 2025, 32, 604–611. [Google Scholar] [CrossRef] [Scilit]
  8. Cheng, A.; Nagesh, V.; Eller, S.; Grant, V.; Lin, Y. Exploring AI Hallucinations of ChatGPT: Reference Accuracy and Citation Relevance of ChatGPT Models and Training Conditions. Simul. Healthc. 2025, 20, 413–418. [Google Scholar] [CrossRef] [Scilit]
  9. Wu, K.; Wu, E.; Wei, K.; Zhang, A.; Casasola, A.; Nguyen, T.; Riantawan, S.; Shi, P.; Ho, D.; Zou, J. An Automated Framework for Assessing How Well LLMs Cite Relevant Medical References. Nat. Commun. 2025, 16, 3615. [Google Scholar] [CrossRef] [Scilit]
  10. Linardon, J.; Jarman, H.K.; McClure, Z.; Anderson, C.; Liu, C.; Messer, M. Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models: Experimental Study. JMIR Ment. Health 2025, 12, e80371. [Google Scholar] [CrossRef] [Scilit]
  11. Ngatuvai, M.; Autrey, C.; McKenny, M.; Elkbuli, A. Significance and Implications of Accurate and Proper Citations in Clinical Research Studies. Ann. Med. Surg. 2021, 72, 102841. [Google Scholar] [CrossRef] [Scilit]
  12. Masic, I. The Importance of Proper Citation of References in Biomedical Articles. Acta Inform. Med. 2013, 21, 148–155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Santini, A. The Importance of Referencing. J. Crit. Care Med. 2018, 4, 3–4. [Google Scholar] [CrossRef] [Scilit]
  14. Bhattacharyya, M.; Miller, V.M.; Bhattacharyya, D.; Miller, L.E. High Rates of Fabricated and Inaccurate References in ChatGPT-Generated Medical Content. Cureus 2023, 15, e39238. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Ghanem, D.; Zhu, A.R.; Kagabo, W.; Osgood, G.; Shafiq, B. ChatGPT-4 Knows Its A B C D E but Cannot Cite Its Source. JBJS Open Access 2024, 9, e24.00099. [Google Scholar] [CrossRef] [Scilit]
  16. Gravel, J.; D’Amours-Gravel, M.; Osmanlliu, E. Learning to Fake It: Limited Responses and Fabricated References Provided by ChatGPT for Medical Questions. Mayo Clin. Proc. Digit. Health 2023, 1, 226–234. [Google Scholar] [CrossRef] [Scilit]
  17. Jaźwińska, K.; Chandrasekar, A. How ChatGPT Search (Mis)Represents Publisher Content. Available online: https://www.cjr.org/tow_center/how-chatgpt-misrepresents-publisher-content.php (accessed on 1 January 2025).
  18. Lechien, J.R.; Briganti, G.; Vaira, L.A. Accuracy of ChatGPT-3.5 and -4 in Providing Scientific References in Otolaryngology-Head and Neck Surgery. Eur. Arch. Otorhinolaryngol. 2024, 281, 2159–2165. [Google Scholar] [CrossRef] [Scilit]
  19. Sanchez-Ramos, L.; Lin, L.; Romero, R. Beware of References When Using ChatGPT as a Source of Information to Write Scientific Articles. Am. J. Obstet. Gynecol. 2023, 229, 356–357. [Google Scholar] [CrossRef] [Scilit]
  20. Walters, W.H.; Wilder, E.I. Fabrication and Errors in the Bibliographic Citations Generated by ChatGPT. Sci. Rep. 2023, 13, 14045. [Google Scholar] [CrossRef] [Scilit]
  21. Frosolini, A.; Franz, L.; Benedetti, S.; Vaira, L.A.; de Filippis, C.; Gennaro, P.; Marioni, G.; Gabriele, G. Assessing the Accuracy of ChatGPT References in Head and Neck and ENT Disciplines. Eur. Arch. Otorhinolaryngol. 2023, 280, 5129–5133. [Google Scholar] [CrossRef] [Scilit]
  22. Sallam, M. ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. Healthcare 2023, 11, 887. [Google Scholar] [CrossRef] [Scilit]
  23. Sawamura, S.; Bito, T.; Ando, T.; Masuda, K.; Kameyama, S.; Ishida, H. Evaluation of the Accuracy of ChatGPT’s Responses to and References for Clinical Questions in Physical Therapy. J. Phys. Ther. Sci. 2024, 36, 234–239. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Ambrosio, L.; Schol, J.; La Pietra, V.A.; Russo, F.; Vadalà, G.; Sakai, D. Threats and Opportunities of Using ChatGPT in Scientific Writing—The Risk of Getting Spineless. JOR Spine 2023, 7, e1296. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Day, T. A Preliminary Investigation of Fake Peer-Reviewed Citations and References Generated by ChatGPT. Prof. Geogr. 2023, 75, 1024–1027. [Google Scholar] [CrossRef] [Scilit]
  26. Schwartzman, J.D.; Shaath, M.K.; Kerr, M.S.; Green, C.C.; Haidukewych, G.J. ChatGPT Is an Unreliable Source of Peer-Reviewed Information for Common Total Knee and Hip Arthroplasty Patient Questions. Adv. Orthop. 2025, 2025, 5534704. [Google Scholar] [CrossRef] [Scilit]
  27. Suppadungsuk, S.; Thongprayoon, C.; Krisanapan, P.; Tangpanithandee, S.; Garcia Valencia, O.; Miao, J.; Mekraksakit, P.; Kashani, K.; Cheungpasitporn, W. Examining the Validity of ChatGPT in Identifying Relevant Nephrology Literature: Findings and Implications. J. Clin. Med. 2023, 12, 5550. [Google Scholar] [CrossRef] [Scilit]
  28. Emsley, R. ChatGPT: These Are Not Hallucinations—They’re Fabrications and Falsifications. Schizophrenia 2023, 9, 52. [Google Scholar] [CrossRef] [Scilit]
  29. Aljamaan, F.; Temsah, M.-H.; Altamimi, I.; Al-Eyadhy, A.; Jamal, A.; Alhasan, K.; Mesallam, T.A.; Farahat, M.; Malki, K.H. Reference Hallucination Score for Medical Artificial Intelligence Chatbots: Development and Usability Study. JMIR Med. Inform. 2024, 12, e54345. [Google Scholar] [CrossRef] [Scilit]
  30. Şengül, H.B.; Akın, B.; Kayaalp, M.E.; Sezgin, E.A. Deep Research Capabilities in GPT-5 Thinking and Gemini 2.5 Pro Improve Citation Integrity and Concordance with American Academy of Orthopaedic Surgeons Anterior Cruciate Ligament and Rotator Cuff Guidelines. Knee Surg. Sports Traumatol. Arthrosc. 2026; early view. [CrossRef] [Scilit]
  31. Chinonso, O.E.; Theresa, A.M.-E.; Aduke, T.C. ChatGPT for Teaching, Learning and Research: Prospects and Challenges. Glob. Acad. J. Humanit. Soc. Sci. 2023, 5, 33–40. [Google Scholar] [CrossRef] [Scilit]
  32. Dergaa, I.; Chamari, K.; Zmijewski, P.; Ben Saad, H. From Human Writing to Artificial Intelligence Generated Text: Examining the Prospects and Potential Threats of ChatGPT in Academic Writing. Biol. Sport. 2023, 40, 615–622. [Google Scholar] [CrossRef] [Scilit]
  33. Chelli, M.; Descamps, J.; Lavoué, V.; Trojani, C.; Azar, M.; Deckert, M.; Raynier, J.-L.; Clowez, G.; Boileau, P.; Ruetsch-Chelli, C. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. J. Med. Internet Res. 2024, 26, e53164. [Google Scholar] [CrossRef] [Scilit]
  34. Aiumtrakul, N.; Thongprayoon, C.; Suppadungsuk, S.; Krisanapan, P.; Miao, J.; Qureshi, F.; Cheungpasitporn, W. Navigating the Landscape of Personalized Medicine: The Relevance of ChatGPT, BingChat, and Bard AI in Nephrology Literature Searches. J. Pers. Med. 2023, 13, 1457. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Athaluri, S.A.; Manthena, S.V.; Kesapragada, V.S.R.K.M.; Yarlagadda, V.; Dave, T.; Duddumpudi, R.T.S. Exploring the Boundaries of Reality: Investigating the Phenomenon of Artificial Intelligence Hallucination in Scientific Writing Through ChatGPT References. Cureus 2023, 15, e37432. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. McGowan, A.; Gui, Y.; Dobbs, M.; Shuster, S.; Cotter, M.; Selloni, A.; Goodman, M.; Srivastava, A.; Cecchi, G.A.; Corcoran, C.M. ChatGPT and Bard Exhibit Spontaneous Citation Fabrication during Psychiatry Literature Search. Psychiatry Res. 2023, 326, 115334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Wu, R.T.; Dang, R.R. ChatGPT in Head and Neck Scientific Writing: A Precautionary Anecdote. Am. J. Otolaryngol. 2023, 44, 103980. [Google Scholar] [CrossRef] [Scilit]
  38. Mugaanyi, J.; Cai, L.; Cheng, S.; Lu, C.; Huang, J. Evaluation of Large Language Model Performance and Reliability for Citations and References in Scholarly Writing: Cross-Disciplinary Study. J. Med. Internet Res. 2024, 26, e52935. [Google Scholar] [CrossRef] [Scilit]
  39. Alkaissi, H.; McFarlane, S.I. Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus 2023, 15, e35179. [Google Scholar] [CrossRef] [Scilit]
  40. Arif, T.B.; Munaf, U.; Ul-Haque, I. The Future of Medical Education and Research: Is ChatGPT a Blessing or Blight in Disguise? Med. Educ. Online 2023, 28, 2181052. [Google Scholar] [CrossRef] [Scilit]
  41. Zhang, M.; Zhao, T. Citation Accuracy Challenges Posed by Large Language Models. JMIR Med. Educ. 2025, 11, e72998. [Google Scholar] [CrossRef] [Scilit]
  42. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Łukasz, K.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  43. Berman, D.S.; Stapleton, A.G. A Path to Natural Language through Tokenisation and Transformers. arXiv 2026, arXiv:2601.03368. [Google Scholar] [CrossRef] [Scilit]
  44. Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2024, arXiv:2312.10997. [Google Scholar] [CrossRef] [Scilit]
  45. Maheshwari, H.; Tenneti, S.; Nakkiran, A. CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track); Rehm, G., Li, Y., Eds.; Association for Computational Linguistics: Vienna, Austria, 2025; pp. 310–317. [Google Scholar]
  46. De Cassai, A.; Dost, B. Concerns Regarding the Uncritical Use of ChatGPT: A Critical Analysis of AI-Generated References in the Context of Regional Anesthesia. Reg. Anesth. Pain. Med. 2024, 49, 378–380. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Choudhury, A.; Chaudhry, Z. Large Language Models and User Trust: Consequence of Self-Referential Learning Loop and the Deskilling of Health Care Professionals. J. Med. Internet Res. 2024, 26, e56764. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Alyasiri, O.M.; Salman, A.M.; Akhtom, D.; Salisu, S. ChatGPT Revisited: Using ChatGPT-4 for Finding References and Editing Language in Medical Scientific Articles. J. Stomatol. Oral. Maxillofac. Surg. 2024, 125, 101842. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  49. Wang, Y.; Ma, X.; Chen, W. Augmenting Black-Box LLMs with Medical Textbooks for Biomedical Question Answering. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024; Al-Onaizan, Y., Bansal, M., Chen, Y.-N., Eds.; Association for Computational Linguistics: Miami, FL, USA, 2024; pp. 1754–1770. [Google Scholar]
  50. Gibney, E. Can Researchers Stop AI Making up Citations? Nature 2025, 645, 569–570. [Google Scholar] [CrossRef] [Scilit]
  51. Huang, G.; Li, Y.; Jameel, S.; Long, Y.; Papanastasiou, G. From Explainable to Interpretable Deep Learning for Natural Language Processing in Healthcare: How Far from Reality? Comput. Struct. Biotechnol. J. 2024, 24, 362–373. [Google Scholar] [CrossRef] [Scilit]
  52. Chopra, A.; Rajput, D.S.; Patel, H. Enhancing Medical Decision-Making with ChatGPT and Explainable AI. Int. J. Surg. 2024, 110, 5167–5168. [Google Scholar] [CrossRef] [Scilit]
  53. Dabis, A.; Csáki, C. AI and Ethics: Investigating the First Policy Responses of Higher Education Institutions to the Challenge of Generative AI. Humanit. Soc. Sci. Commun. 2024, 11, 1006. [Google Scholar] [CrossRef] [Scilit]
  54. Masters, K. Ethical Use of Artificial Intelligence in Health Professions Education: AMEE Guide No. 158. Med. Teach. 2023, 45, 574–584. [Google Scholar] [CrossRef] [Scilit]
  55. Cheng, K.; Wu, H. Policy Framework for the Utilization of Generative AI. Crit. Care 2024, 28, 128. [Google Scholar] [CrossRef] [Scilit]
  56. Smith, G.R.; Bello, C.; Bialic-Murphy, L.; Clark, E.; Delavaux, C.S.; Fournier de Lauriere, C.; van den Hoogen, J.; Lauber, T.; Ma, H.; Maynard, D.S.; et al. Ten Simple Rules for Using Large Language Models in Science, Version 1.0. PLoS Comput. Biol. 2024, 20, e1011767. [Google Scholar] [CrossRef]
  57. Tools Such as ChatGPT Threaten Transparent Science; Here Are Our Ground Rules for Their Use. Nature 2023, 613, 612. Available online: https://www.nature.com/articles/d41586-023-00191-1 (accessed on 17 February 2026). [CrossRef] [Scilit]
  58. Resnik, D.B.; Hosseini, M. Disclosing Artificial Intelligence Use in Scientific Research and Publication: When Should Disclosure Be Mandatory, Optional, or Unnecessary? Account. Res. 2025, 33, 2481949. [Google Scholar] [CrossRef] [Scilit]
Table 1. Reference Hallucination Score Framework.
Table 1. Reference Hallucination Score Framework.
Reference IdentifierItem Hallucination Score
Erroneous date of publication1
Erroneous web link of the paper1
Erroneous citation relevance1
Erroneous title of the paper2
Erroneous digital object identifier2
Erroneous authors’ names2
Erroneous name of the journal2
Total Reference Hallucination Score (RHS)11
Reproduced from Aljamaan et al. 2024 [29].
Table 2. Various articles reporting inaccuracies in LLM generated citations (2023–2026).
Table 2. Various articles reporting inaccuracies in LLM generated citations (2023–2026).
Author, YearType of Article Main FindingsSummary of DiscussionRecommendations
Sallam, 2023 [22]Systematic review
  • ChatGPT has promising applications in healthcare education, research, and practice, but also raises valid concerns that need to be addressed.
  • Key concerns include ethical issues, risk of hallucination, transparency issues, legal issues, and potential decline in human expertise and critical thinking.
  • An initiative involving all stakeholders is urgently needed to establish a code of ethics for the responsible use of ChatGPT and other LLMs in healthcare and academia.
The imminent dominant use of large language models like ChatGPT in healthcare education, research, and practice is inevitable, but appropriate guidelines and regulations are urgently needed to ensure their safe and responsible use, while also weighing their potential benefits against the possible risks.
  • Researchers should use ChatGPT with caution due to risks of inaccuracy, bias, hallucination, plagiarism, and ethical or legal concerns.
  • The impact of ChatGPT and other LLMs on healthcare education, research, and practice should be systematically evaluated to prevent misuse.
  • Clear policies and ethical frameworks are urgently needed to ensure transparent and responsible use of ChatGPT.
Chinonso et al., 2023 [31]Literature review
  • ChatGPT provides rapid, conversational responses and supports self-paced learning and research activities.
  • It has notable limitations, including lack of citation and referencing, risk of plagiarism, inaccurate responses, and over-reliance concerns.
  • ChatGPT should enhance, not replace, traditional research methods, and users must conduct proper research and provide appropriate citations.
  • Despite its advantages, search engines and independent verification remain necessary.
The paper discusses the prospects and challenges of using ChatGPT. Authors find the model to be verbose, often over-explaining or repeating certain phrases. It infers user intent instead of asking clarifying questions, leading to potential misunderstandings or incorrect answers. Occasionally, the model may react to damaging or biased instructions.Users should avoid over-reliance on AI-generated responses, continue conducting independent research using established tools such as search engines, and ensure that all sources and literary works are properly cited and referenced.
Dergaa et al., 2023 [32]Literature review
  • ChatGPT and other NLP technologies have the potential to enhance academic writing and research efficiency.
  • However, their use also raises concerns about the impact on the authenticity and credibility of academic work.
  • The review calls for comprehensive discussion of the uses and risks of NLP tools, emphasizing ethical principles and human critical thinking.
NLP tools like ChatGPT enhance efficiency through data analysis, summarization, and research support, and may improve accessibility. However, they raise concerns regarding factual inaccuracies, bias, plagiarism, transparency, and overreliance, with implications for research integrity.
  • Researchers, reviewers, editors, and publishers should understand the capabilities and potential issues of LLMs, acting as guardians of honest science.
  • Educators should discuss the ethics and use of AI tools with students, while group leaders and teachers need to establish guidelines for responsible use.
  • Contributors to research must take full responsibility for verifying the accuracy of their work, regardless of AI support.
Şengü et al., 2026 [30]Research Study
  • In a blind evaluation of 46 orthopedic prompts, four LLM configurations (GPT-5 Thinking, GPT-5 Thinking Deep Research, Gemini 2.5 Pro, Gemini 2.5 Pro Deep Research) were assessed.
  • GPT-5 Thinking achieved 96.8% citation integrity, improving to 100% with Deep Research.
  • Gemini 2.5 Pro showed low baseline integrity (64.6%), with frequent hallucinated references, improving to 98.6% with Deep Research.
GPT-5 demonstrated strong baseline citation integrity, whereas Gemini relied heavily on live web search to mitigate hallucinated references. Performance was optimized under structured, citation-enforced prompts, suggesting that real-world unstructured queries may underperform. Clinician oversight remains essential, and live web search introduces temporal variability in outputs.Individualized training in prompt creation or interface-level systems is required for safe clinical adoption.
Linardon et al., 2025 [10]Research Study
  • Across six reviews generated by ChatGPT-4, 19.9% of citations were fabricated, and 45.4% of real citations contained errors, most commonly incorrect digital object identifiers.
  • Fabrication and accuracy rates varied significantly by disorder and review type, with higher fabrication and lower accuracy in certain conditions and specialized reviews.
Citation fabrication and bibliographic errors were common, with lower fabrication and higher accuracy in widely studied and familiar conditions, and higher errors in less familiar or more specialized topics, showing that subject area and prompt granularity influence LLM citation reliability.
  • Researchers and students should systematically check, validate, and verify the accuracy and authenticity of citations generated by LLMs.
  • Journal editors and publishers can use citation matching through plagiarism detection tools to confirm whether references exist in published sources.
  • Academic institutions and organizations should establish clear policies and training to help users identify hallucinated citations, verify references, and appropriately disclose AI assistance in scientific writing.
Cheng et al., 2025 [8]Research Study
  • Fifteen articles were generated in total, including nine by ChatGPT-4 and six by ChatGPT-o1 across five training conditions, and citation accuracy and relevance were assessed.
  • Overall, 23.4% of references were completely fabricated and 16.9% were partially fabricated, with no differences across conditions.
  • Among valid citations, 12.9% were irrelevant and 25.3% were only moderately relevant, with no differences across conditions.
ChatGPT-4 and ChatGPT-o1 demonstrated poor performance in generating accurate and relevant references. There was no improvement with prompt training interventions. These limitations arise because LLMs generate responses based on learned text patterns without built-in factchecking.The responsibility for ensuring accuracy rests with human authors, who must carefully verify all AI-generated content, citations, and references, and clearly disclose the use of generative AI in manuscript preparation.
Schwartzman et al., 2025 [26]Research Study
  • This study evaluated whether ChatGPT-3.5 provides reliable peer-reviewed references for knee and hip arthroplasty related questions and found that only 36% of cited sources were accurate.
  • There was no clear bias toward US-based or open-access journals.
ChatGPT-3.5 was an unreliable source of peer-reviewed literature for total knee and hip arthroplasty, with 64% of citations inaccurate or unverifiable, indicating it can generate convincing responses without consistently citing credible sources.Medical researchers must externally validate all information and sources provided by ChatGPT.
Day, 2023 [25]Research study
  • ChatGPT-generated citations for five geography prompts were systematically checked against journal databases and online records.
  • None of the peer-reviewed citations generated by ChatGPT could be verified as genuine.
  • The fabricated references appeared plausible, using real journal names, authors, and formatting, making them difficult to detect.
  • Some references were partially derived from real papers but contained altered metadata (authors, journal, volume, or pages).
ChatGPT can generate fake academic citations and references, highlighting limitations in its use for research. However, it may have applications in tasks that do not require references, provided domain expertise is applied to identify and correct inaccurate information.Researchers should adopt a cautious and critical approach to using ChatGPT in teaching and research.
Frosolini et al., 2023 [21]Research study
  • Twenty clinical questions across five head and neck disciplines were posed to ChatGPT 3.5 and 4.0, and generated references were verified against scientific databases and categorized as true, erroneous, or inexistent.
  • ChatGPT 4.0 produced 74.29% true references, significantly outperforming ChatGPT 3.5 (16.66%).
  • However, 25.71% of ChatGPT 4.0 references were erroneous or inexistent, indicating persistent reliability concerns.
ChatGPT 4.0 improved reference reliability compared to version 3.5; however, the persistent generation of erroneous or inexistent citations raises concerns for scientific credibility and evidence-based practice.Journals and institutions should establish strategies and good-practice guidelines to mitigate bias in AI-assisted scientific writing, and further longitudinal research is needed.
Sawamura et al., 2024 [23]Research Study
  • Five clinical questions from three sections of the Japanese Physical Therapy Guidelines were submitted to the free version of ChatGPT in Japanese, and response content and references were evaluated by expert raters.
  • ChatGPT generated generally accurate clinical responses, but reference accuracy was poor, with 40.5% of citations being fictitious and only 18.9% correctly matching PMIDs/DOIs.
ChatGPT-4.0’s training dataset includes public information, research articles, and sources like Wikipedia and blogs, which may introduce biases. Citing references is more complex than answering questions, likely contributing to the inaccuracies in the reference output.ChatGPT should be used cautiously as a support tool rather than a replacement for professional expertise, and further research is needed to improve reference accuracy and reliability.
Chelli et al., 2024 [33]Research study
  • The performance of GPT-3.5, GPT-4, and Bard in replicating human-conducted systematic reviews was evaluated across 11 reviews in 4 fields, involving 33 prompts and 471 references.
  • Precision rates were 9.4% for GPT-3.5, 13.4% for GPT-4, and 0% for Bard. Recall rates were 11.9% for GPT-3.5 and 13.7% for GPT-4, with Bard retrieving no relevant papers.
  • Hallucination rates were 39.6% for GPT-3.5, 28.6% for GPT-4, and 91.4% for Bard.
Current LLMs demonstrate insufficient precision and high hallucination rates, making them unreliable as standalone tools for systematic reviews and requiring careful human validation of generated references.
  • A statement or scholarly usage guideline should be clearly presented before using the tool or integrated into the software itself, outlining its lack of liability for any inaccuracies in citation.
  • The use of GPT-based chatbots for tasks such as spelling correction, proofreading, or text restructuring be explicitly mentioned in the materials and methods section of academic writings.
Bhattacharyya et al., 2023 [14]Research study
  • ChatGPT-3.5 generated 30 short biomedical papers using standardized prompts across multiple therapeutic areas, and 115 references were verified for authenticity and accuracy.
  • 47% of the references generated by ChatGPT were fabricated, 46% were authentic but inaccurate, and only 7% were authentic and accurate.
  • The likelihood of fabricated references significantly differed based on prompt variations, with the highest rates in papers on healthcare disparity, pulmonology, dermatology, and gastroenterology.
  • The most common errors in the references were incorrect PMID numbers (93%), volume (64%), page numbers (64%), and year of publication (60%).
The findings emphasize the need for caution when seeking medical information on ChatGPT, as most references provided were fabricated or inaccurate, and individuals should verify information from reliable sources and consult medical professionals for serious health concerns.
  • Users should exercise caution when using ChatGPT for biomedical information and verify references against trusted scientific sources.
  • Full reliance on AI-generated content should be avoided, with rigorous human verification required.
Walters & Wilder, 2023 [20]Research study
  • GPT-3.5 and GPT-4 generated 84 short literature reviews across 42 multidisciplinary topics, producing 636 citations that were evaluated for fabrication and substantive citation errors.
  • 55% of the citations generated by GPT-3.5 were fabricated, compared to only 18% for GPT-4.
  • 43% of the real (non-fabricated) citations generated by GPT-3.5 contained substantive errors, compared to 24% for GPT-4.
Although GPT-4 shows marked improvement over GPT-3.5, persistent fabricated citations and substantive errors limit its reliability for scholarly use.
  • Users should verify AI-generated citations and critically assess reference quality.
  • Journal editors and publishers should ensure fabricated or inaccurate citations are not incorporated into the academic literature.
  • Educators should carefully examine citations in student assignments to detect AI-generated errors.
Aljamaan et al., 2024 [29]Research study
  • Six AI chatbots were tested with 10 medical prompts (10 references each), and references were evaluated using a Reference Hallucination Score (RHS).
  • ChatGPT 3.5 and Bing exhibited the highest hallucination scores (median RHS = 11), whereas Elicit and SciSpace showed the lowest (median RHS = 1).
  • Bard generated no references, and Perplexity demonstrated intermediate hallucination levels (median RHS = 8).
AI chatbots demonstrated higher hallucination severity when prompted with complex medical scenarios. These findings underscore the need for standardized evaluation tools, such as the proposed Reference Hallucination Score (RHS), to systematically assess and monitor the reliability of AI-generated references.
  • Scholars should rigorously assess the accuracy of AI-generated citations, and structured tools such as the Reference Hallucination Score (RHS) may assist in evaluating reference reliability.
  • The RHS can support the triaging of hallucinated references and promote greater scrutiny of AI-generated content.
Aiumtrakul et al., 2023 [34]Research study
  • ChatGPT, Bing Chat, and Bard were prompted with 20 medical questions, and generated references were verified and classified for accuracy and fabrication.
  • ChatGPT provided the highest proportion of accurate references (38%) compared to Bing Chat (30%) and Bard (3%).
  • Bard had the highest proportion of fabricated (63%) and incomplete (11%) references.
  • The most common source of inaccuracy was incorrect DOIs for ChatGPT and Bing Chat, while it was incorrect author names for Bard.
The discussion emphasizes the need for careful vetting of AI-generated references in the medical field, given the varying levels of citation accuracy across different AI chatbots and the potential consequences of inaccuracies on medical decision-making and patient care.
  • AI-generated references should be rigorously verified before use in clinical or research contexts.
  • Greater transparency regarding the use of AI-generated content is needed to allow critical appraisal of citation reliability.
  • Clinicians, researchers, and journals should exercise caution and scrutiny when incorporating AI-generated references into academic work.
Suppadungsuk et al., 2023 [27]Research study
  • ChatGPT-3.5 was prompted to provide Vancouver-style references with links across 12 nephrology subdomains.
  • Out of the 610 nephrology references provided by ChatGPT, 62% were valid, while 31% were fabricated and 7% were incomplete.
  • Furthermore, 68% of the reference links were incorrect, and 54% of the DOIs were inaccurate.
The hallucination effect of ChatGPT may vary across studies due to factors like input data, prompts, or algorithm updates. Interestingly, when ChatGPT provided the correct link, all reference components were found to be authentic. This suggests a potential correlation between the presence of a correct link and the accuracy of the reference.Rigorous validation and clinical critical review of AI-generated references are necessary, and ongoing model updates may help reduce hallucinations and improve reliability.
Athaluri et al., 2023 [35]Research study
  • ChatGPT-3 generated 50 medical research proposals, yielding 178 references that were independently verified for DOI validity and existence by five reviewers.
  • Of 178 references, 69 lacked valid DOI; 28 could not be found via Google search and had no existing DOI; and 3 were books rather than research articles.
The discussion highlights AI hallucination as a significant limitation of ChatGPT-3, with potential ethical and legal implications, and emphasizes the need for caution and independent verification of AI-generated references.
  • Researchers should exercise caution and avoid full reliance on ChatGPT-generated citations in research proposals.
  • AI-generated references should be independently verified before use, and continued model improvement is necessary to mitigate hallucination risks.
Gravel et al., 2023 [16]Research study
  • ChatGPT was prompted with 20 medical questions derived from high-impact journals, and 59 references were evaluated for validity by domain experts.
  • Of the 59 references, 69% were fabricated despite appearing credible. Domain experts identified major factual errors in 29% of responses, and the median response quality score was 60%.
ChatGPT has significant limitations in providing accurate and valid information, especially in the references it provides, highlighting the need for caution, oversight, and human involvement when using such language models for scientific purposes.
  • Journals should implement clear policies regarding the use and disclosure of LLM-generated content in scientific publications.
  • Developers should improve model reliability to reduce fabricated and inaccurate citations.
McGowan et al., 2023 [36]Research study
  • ChatGPT-3.5 was prompted for peer-reviewed references on suicidal behavior in psychosis, and citations were verified against scientific databases; updated ChatGPT-3.5 and Bard 2.0 were also tested.
  • Of 35 ChatGPT citations, only 2 were real; 12 contained incorrect details, and 21 were fabricated. Bard generated eight citations, none were accurate.
The authors attribute citation fabrication to the probabilistic generative nature of LLMs and raise concerns about implications for scientific integrity.AI-generated references should be independently verified and not relied upon without scrutiny.
Wu & Dang, 2023 [37]Research study
  • ChatGPT was asked to generate 10 complete citations for each of five head and neck surgery topics (50 total), and references were independently verified for title, authors, journal, year, and DOI accuracy.
  • Only 10% of references were completely correct, and just 14% had the correct DOI; references on free flap reconstruction were the least accurate, with no correct DOI.
The discussion highlights concern that fabricated citations generated by ChatGPT may mislead researchers and undermine scientific integrity, particularly as AI-generated text becomes increasingly sophisticated.AI-generated references should be independently verified before use in academic work due to the risk of fabrication.
Mugaanyi et al., 2024 [38]Research study
  • ChatGPT (GPT-3.5) was prompted by two independent researchers to generate introduction sections with citations for 10 topics (5 natural sciences, 5 humanities), and 102 citations were evaluated for existence, accuracy, DOI presence, and DOI hallucination.
  • Of 102 citations, 72.7% (40/55) in the natural sciences and 76.6% (36/47) in the humanities were confirmed to exist (p = 0.42). However, DOI presence (70.9% vs. 38.3%) and DOI accuracy (32.7% vs. 8.5%) were significantly lower in the humanities, with DOI hallucination more frequent in the humanities (89.4%).
ChatGPT’s performance in generating accurate citations and references varies across academic disciplines, with lower accuracy in the humanities compared to the natural sciences, particularly for generating valid DOIs.
  • There is a need to evaluate discipline-specific limitations of ChatGPT.
  • Rigorous validation steps must be in place to ensure the accuracy and credibility of generated content.
Alkaissi & McFarlane, 2023 [39]EditorialChatGPT was able to generate coherent scientific text from structured prompts and assist indirectly with reference management (e.g., identifying recurrent PMIDs via generated code). However, it produced fabricated references and confident but factually incorrect content, exemplifying what the authors describe as “artificial hallucination.The paper discusses the implications of using ChatGPT in scientific writing, highlighting concerns about the integrity and accuracy of the content generated by large language models, and proposes that journal and conference policies be modified to maintain scientific standards and include AI output detectors in the editorial process.Journals and medical conferences should modify their policies and practices for evaluating scientific manuscripts to maintain rigorous scientific standards, given the potential for ChatGPT to generate inaccurate or fabricated content.
Arif et al., 2023 [40]Editorial
  • ChatGPT can assist in drafting scientific content using information from online sources but lacked the capacity to perform comprehensive literature searches or critically analyze articles.
  • Its limited access to updated databases (e.g., PubMed and Cochrane) restricted its ability to provide credible, fully reliable academic output.
  • The authors raise concerns about accuracy, accountability, ethical implications, and the potential misuse of AI in medical education and research.
While ChatGPT can be useful for certain tasks in medical education and research, such as writing and rephrasing text, it also has significant limitations, including a lack of critical thinking, reasoning, and access to up-to-date medical databases, which raises concerns about its credibility and potential for misuse.The use of ChatGPT in medical education and research requires careful supervision to prevent misuse and ensure academic integrity.
Sanchez-Ramos et al., 2023 [19]Letter
  • ChatGPT frequently provides erroneous and fictitious references, including incorrect authors, journals, titles, years, and PubMed identifiers.
  • Errors in references undermine the credibility of the authors and trustworthiness of the scientific report.
Authors need to verify the accuracy of references provided by ChatGPT before submitting manuscripts, though the limitations of these tools are expected to improve over time.Authors must ensure the accuracy of LLM-generated outputs before submitting manuscripts for publication.
Zhang and Zhao, 2025 [41]Letter
  • The authors argue that LLMs such as DeepSeek, ChatGPT, and ChatGLM generate correctly formatted but fictional citations, undermining academic rigor.
  • They attribute these citation inaccuracies to restricted access to subscription databases, pattern-based text generation without true understanding, and opaque underlying algorithms
Fictional citations mislead users, undermine academic rigor and credibility, and impede the dissemination of accurate scientific knowledge, requiring coordinated efforts from multiple stakeholders to address.
  • Coordinated efforts from developers, users, and academic institutions are needed to address citation fabrication and improve reliability.
  • Users should critically evaluate AI-generated citations before incorporating them into academic work.
Considering the pace of development of this field, earlier studies might soon lose relevance as LLMs continue to improve.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Rajaratnam, V.; Omar, U.F.; Kee, K.; Kaliya-Perumal, A.-K. Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing. Standards 2026, 6, 10. https://doi.org/10.3390/standards6010010

AMA Style

Rajaratnam V, Omar UF, Kee K, Kaliya-Perumal A-K. Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing. Standards. 2026; 6(1):10. https://doi.org/10.3390/standards6010010

Chicago/Turabian Style

Rajaratnam, Vaikunthan, Usama Farghaly Omar, Kristen Kee, and Arun-Kumar Kaliya-Perumal. 2026. "Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing" Standards 6, no. 1: 10. https://doi.org/10.3390/standards6010010

APA Style

Rajaratnam, V., Omar, U. F., Kee, K., & Kaliya-Perumal, A.-K. (2026). Citation Inaccuracies and the Need for Multi-Level Oversight in AI-Assisted Medical Writing. Standards, 6(1), 10. https://doi.org/10.3390/standards6010010

Article Metrics

Back to TopTop