Artificial Intelligence in Peer Review: A Bibliometric-Guided Thematic Review and a Task-Contingent Legitimacy Framework
Abstract
1. Introduction
2. Methods
2.1. Search Strategy
TITLE ((“artificial intelligence” OR “generative AI” OR ChatGPT OR “large language model*” OR LLM*) AND (“peer review” OR reviewer* OR referee*))
2.2. Screening and Eligibility
2.3. Coding Procedure
2.4. Analytical Approach
3. Results
3.1. Bibliometric Results
3.2. Content Analysis
3.2.1. Thematic Distribution
3.2.2. Editorial Stance
3.2.3. Stated Position on AI Autonomy
3.3. Empirical Evidence Synthesis
3.4. Thematic Literature Review
4. A Task-Contingent Legitimacy Framework of AI in Peer Review
5. Discussion
6. Limitations and Future Research Directions
6.1. Limitations
6.2. Future Research Directions
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Arabzadeh, N., Ebrahimi, S., Sadeghian, S., Hosseini, S. M., Daqiq, A., Le, H. S., Bashari, M., & Bagheri, E. (2026). Can LLMs uphold research integrity? Evaluating the role of LLMs in peer review quality. In WSDM 2026—Proceedings of the 19th ACM international conference on web search and data mining (pp. 1341–1342). ACM Digital Library. [Google Scholar] [CrossRef] [Scilit]
- Bauchner, H., & Rivara, F. P. (2024). Use of artificial intelligence and the future of peer review. Health Affairs Scholar, 2(5), qxae058. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chang, Y., Li, Z., Zhang, H., Kong, Y., Wu, Y., So, H. K.-H., Guo, Z., Zhu, L., & Wong, N. (2025). TreeReview: A dynamic tree of questions framework for deep and efficient LLM-based scientific peer review. In EMNLP 2025—2025 conference on empirical methods in natural language processing, proceedings of the conference (pp. 15651–15682). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
- Cheng, K., Sun, Z., Liu, X., Wu, H., & Li, C. (2024). Generative artificial intelligence is infiltrating peer review process. Critical Care, 28(1), 149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chu, Z., Ai, Q., Tu, Y., Li, H., & Liu, Y. (2024). Automatic large language model evaluation via peer review. In International conference on information and knowledge management, proceedings (pp. 384–393). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
- Crawford, J., Allen, K.-A., & Lodge, J. (2024). Humanising peer review with artificial intelligence: Paradox or panacea? Journal of University Teaching and Learning Practice, 21(1), 9–16. [Google Scholar] [CrossRef] [Scilit]
- Donker, T. (2023). The dangers of using large language models for peer review. The Lancet Infectious Diseases, 23(7), 781. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Doskaliuk, B., Zimba, O., Yessirkepov, M., Klishch, I., & Yatsyshyn, R. (2025). Artificial intelligence in peer review: Enhancing efficiency while preserving integrity. Journal of Korean Medical Science, 40(7), e92. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Farber, S. (2024). Enhancing peer review efficiency: A mixed-methods analysis of artificial intelligence-assisted reviewer selection across academic disciplines. Learned Publishing, 37(4), e1638, (Erratum in 2025, Learned Publishing, 38, e1663). [Google Scholar] [CrossRef] [Scilit]
- Felländer-Tsai, L., & Overgaard, S. (2023). Adapting to the rapidly moving target artificial intelligence (AI) in scholarly publishing. Acta Orthopaedica, 94, 625. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Garcia, M. B. (2024). Using AI tools in writing peer review reports: Should academic journals embrace the use of ChatGPT? Annals of Biomedical Engineering, 52(2), 139–140. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gatrell, C., Muzio, D., Post, C., & Wickert, C. (2024). Here, there and everywhere: On the responsible use of artificial intelligence (AI) in management research and the peer-review process. Journal of Management Studies, 61(3), 739–751. [Google Scholar] [CrossRef] [Scilit]
- Hopkins, B. S., Shah, I., Dallas, J., Borja, A. J., Gomez, D., Briggs, R. G., Cote, D. J., Chung, L., Shasby, G., Sisti, J., Rutka, J. T., & Zada, G. (2026). Application of large language and artificial intelligence modeling in the prediction of peer-review outcomes. Journal of Neurosurgery, 144(3), 720–728. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huang, X., Dou, S., & Yin, Z. (2025). The dual-edged sword: Artificial intelligence’s evolving role in academic peer review. Science China Information Sciences, 68(11), 216101:1–216101:3. [Google Scholar] [CrossRef] [Scilit]
- Jin, Y., Zhao, Q., Wang, Y., Chen, H., Zhu, K., Xiao, Y., & Wang, J. (2024). AgentReview: Exploring peer review dynamics with LLM agents. In EMNLP 2024—2024 conference on empirical methods in natural language processing, proceedings of the conference (pp. 1208–1226). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
- Joachim, M. V., Dodson, T. B., & Laviv, A. (2025). How artificial intelligence differs from humans in peer review. Journal of Oral and Maxillofacial Surgery, 83(8), 1040–1050. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kadi, G., & Aslaner, M. A. (2024). Exploring ChatGPT’s abilities in medical article writing and peer review. Croatian Medical Journal, 65(2), 93–100. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kankanhalli, A. (2024). Peer review in the age of generative AI. Journal of the Association for Information Systems, 25(1), 76–84. [Google Scholar] [CrossRef] [Scilit]
- Kousha, K., & Thelwall, M. (2024). Artificial intelligence to support publishing and peer review: A summary and review. Learned Publishing, 37(1), 4–12. [Google Scholar] [CrossRef] [Scilit]
- Levin, J. M., Oprea, T. I., Davidovich, S., Clozel, T., Overington, J. P., Vanhaelen, Q., Cantor, C. R., Bischof, E., & Zhavoronkov, A. (2020). Artificial intelligence, drug repurposing and peer review. Nature Biotechnology, 38(10), 1127–1131. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, J., Li, Y., Hu, X., Gao, M., & Wan, X. (2025). Where do LLMs go wrong? Diagnosing automated peer review via aspect-guided multi-level perturbation. In CIKM 2025—Proceedings of the 34th ACM international conference on information and knowledge management (pp. 1572–1581). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.-Q., Xu, H.-L., Cao, H.-J., Liu, Z.-L., Fei, Y.-T., & Liu, J.-P. (2024). Use of artificial intelligence in peer review among top 100 medical journals. JAMA Network Open, 7(12), e2448609. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., Chen, L., Ye, H., Liu, S., Huang, Z., McFarland, D., & Zou, J. Y. (2024). Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews. In Proceedings of the 41st international conference on machine learning (ICML 2024), PMLR 235 (pp. 29575–29620). PMLR. [Google Scholar]
- Lin, T.-L., Chen, W.-C., Hsiao, T.-F., Liu, H.-I., Yeh, Y.-H., Chan, Y.-K., Lien, W.-S., Kuo, P.-Y., Yu, P. S., & Shuai, H.-H. (2025). Breaking the reviewer: Assessing the vulnerability of large language models in automated peer review under textual adversarial attacks. In EMNLP 2025—2025 conference on empirical methods in natural language processing, findings of EMNLP 2025 (pp. 4819–4839). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
- Mohamed, A. A., Colome, D., Yang, J., Sargent, E. C., Flores-Milan, G., Sorrentino, Z., Sharma, A., Adogwa, O., Pirris, S., & Lucke-Wold, B. (2025a). Leveraging artificial intelligence in the peer review of neurosurgical research articles. Neurosurgical Review, 48(1), 631. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mohamed, A. A., Rajendran, S., Colome, D., Sargent, E. C., Schirmer, C. M., Vessell, M., Lucke-Wold, B., Sharma, A., Adogwa, O., & Pirris, S. (2025b). Neurosurgical journals’ policies on artificial intelligence use in manuscript preparation and peer review. Neurosurgical Review, 48(1), 670. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mrowinski, M. J., Fronczak, P., Fronczak, A., Ausloos, M., & Nedic, O. (2017). Artificial intelligence in peer review: How can evolutionary computation support journal editors? PLoS ONE, 12(9), e0184711. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Munafò, M. (2024). A policy on the use of artificial intelligence and large language models in peer review. Nicotine and Tobacco Research, 26(5), 519. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nabavi, A., Safari, F., Shmoury, A. H., Tabet, S., Perdomo-Luna, C., & Celi, L. A. (2026). Artificial intelligence in scholarly peer review: A scoping review of applications, risks, and governance challenges. International Journal of Medical Informatics, 214, 106418. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ning, K.-P., Yang, S., Liu, Y.-Y., Yao, J.-Y., Liu, Z.-H., Tian, Y.-H., Song, Y., & Yuan, L. (2025). PiCO: Peer review in LLMs based on consistency optimization. In Proceedings of the thirteenth international conference on learning representations (ICLR 2025). ICLR. [Google Scholar]
- (2026). Peer review in the time of artificial intelligence. Nature Nanotechnology, 21(4), 479. [CrossRef] [Scilit] [PubMed]
- Perlis, R. H., Christakis, D. A., Bressler, N. M., Öngür, D., Kendall-Taylor, J., Flanagin, A., & Bibbins-Domingo, K. (2025). Artificial intelligence in peer review. JAMA, 334(17), 1520–1522. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Price, S., & Flach, P. A. (2017). Computational support for academic peer review: A perspective from artificial intelligence. Communications of the ACM, 60(3), 70–79. [Google Scholar] [CrossRef] [Scilit]
- Rajakumar, H. K., Sankaran, K. A., Ashok, M. P., & Rachoori, S. (2026). Peer review in the age of artificial intelligence: A comparative study of human and AI-generated review reports. Postgraduate Medical Journal, 102(1208), 551–559. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rozencwajg, S., & Benhamou, D. (2026). Declare the use or non-use of artificial intelligence in peer-review: A relic of the past? Anaesthesia Critical Care and Pain Medicine, 45(4), 101812. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Saad, A., Jenko, N., Ariyaratne, S., Birch, N., Iyengar, K. P., Davies, A. M., Vaishya, R., & Botchu, R. (2024). Exploring the potential of ChatGPT in the peer review process: An observational study. Diabetes and Metabolic Syndrome: Clinical Research and Reviews, 18(2), 102946. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sebo, P. (2024). Use of ChatGPT to explore gender and geographic disparities in scientific peer review. Journal of Medical Internet Research, 26, e57667. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. In FAT* 2019—Proceedings of the 2019 conference on fairness, accountability, and transparency (pp. 59–68). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
- Shen, S. M., Wang, Z., Paul, K., Li, M.-H., Huang, X., & Koizumi, N. (2026). Evaluation of large language models for peer review in transplantation research: Algorithm validation study. JMIR AI, 5, e84322. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sheridan, G. A., Howard, L. C., Neufeld, M. E., Doyle, T. R., Hughes, A. J., Sculco, P. K., Beverland, D. E., Garbuz, D. S., & Masri, B. A. (2025). Can artificial intelligence generate scientific discussion that passes peer review for publication in a high-impact orthopaedic journal? Irish Journal of Medical Science, 194(4), 1191–1198. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Singh Chawla, D. (2024). Is ChatGPT corrupting peer review? Telltale words hint at AI use. Nature, 628(8008), 483–484. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Suchman, M. C. (1995). Managing legitimacy: Strategic and institutional approaches. Academy of Management Review, 20(3), 571–610. [Google Scholar] [CrossRef] [Scilit]
- Sun, Z. (2026). Relationship between peer review quality and scientific impact: Insights from LLMs-assessed reviews. Journal of Informetrics, 20(2), 101801. [Google Scholar] [CrossRef] [Scilit]
- Tang, Y., Kang, Y., Wu, S., Zhang, R., & Sun, Z. (2026). Can large language models assess the quality of peer review? An empirical study. Scientometrics, 131(4), 2237–2259. [Google Scholar] [CrossRef] [Scilit]
- Thakkar, N., Yuksekgonul, M., Silberg, J., Garg, A., Peng, N., Sha, F., Yu, R., Vondrick, C., & Zou, J. (2026). A large-scale randomized study of large language model feedback in peer review. Nature Machine Intelligence, 8(3), 326–336. [Google Scholar] [CrossRef] [Scilit]
- Thelwall, M., & Yaghi, A. (2025). Evaluating the predictive capacity of ChatGPT for academic peer review outcomes across multiple platforms. Scientometrics, 130(10), 5285–5307. [Google Scholar] [CrossRef] [Scilit]
- Verharen, J. P. H. (2023). ChatGPT identifies gender disparities in scientific peer review. eLife, 12, RP90230. [Google Scholar] [CrossRef] [PubMed]
- Vincent-Lamarre, P., & Larivière, V. (2021). Textual analysis of artificial intelligence manuscripts reveals features associated with peer review outcome. Quantitative Science Studies, 2(2), 662–677. [Google Scholar] [CrossRef] [Scilit]
- Von Wedel, D., Schmitt, R. A., Thiele, M., Leuner, R., Shay, D., Redaelli, S., & Schaefer, M. S. (2024). Affiliation bias in peer review of abstracts by a large language model. JAMA, 331(3), 252–253. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yin, S., Huang, S., Xue, P., Xu, Z., Lian, Z., Ye, C., Ma, S., Liu, M., Hu, Y., Lu, P., & Li, C. (2025). Generative artificial intelligence (GAI) usage guidelines for scholarly publishing: A cross-sectional study of medical journals. BMC Medicine, 23(1), 77. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Q., Ning, M., Liu, Z., Huang, Y., Yang, S., Wang, Y., Ye, J., Chen, X., Song, Y., & Yuan, L. (2025). UPME: An unsupervised peer review framework for multimodal large language model evaluation. In Proceedings of the IEEE computer society conference on computer vision and pattern recognition (pp. 9165–9174). IEEE. [Google Scholar] [CrossRef] [Scilit]
- Zou, J. (2024). ChatGPT is transforming peer review—How can we use it responsibly? Nature, 635(8037), 10. [Google Scholar] [CrossRef] [Scilit] [PubMed]







| Rank | Title | Year | Source | Citations |
|---|---|---|---|---|
| 1 | Computational support for academic peer review: A perspective from artificial intelligence (Price & Flach, 2017) | 2017 | Communications of the ACM | 151 |
| 2 | Artificial intelligence to support publishing and peer review: A summary and review (Kousha & Thelwall, 2024) | 2024 | Learned Publishing | 121 |
| 3 | Artificial intelligence, drug repurposing and peer review (Levin et al., 2020) | 2020 | Nature Biotechnology | 64 |
| 4 | Here, there and everywhere: On the responsible use of artificial intelligence (AI) in management research and the peer-review process (Gatrell et al., 2024) | 2024 | Journal of Management Studies | 61 |
| 5 | The dangers of using large language models for peer review (Donker, 2023) | 2023 | The Lancet Infectious Diseases | 60 |
| 6 | Exploring the potential of ChatGPT in the peer review process: An observational study (Saad et al., 2024) | 2024 | Diabetes and Metabolic Syndrome: Clinical Research and Reviews | 56 |
| 7 | Use of artificial intelligence and the future of peer review (Bauchner & Rivara, 2024) | 2024 | Health Affairs Scholar | 52 |
| 8 | Use of artificial intelligence in peer review among top 100 medical journals (Z.-Q. Li et al., 2024) | 2024 | JAMA Network Open | 51 |
| 9 | Artificial intelligence in peer review: How can evolutionary computation support journal editors? (Mrowinski et al., 2017) | 2017 | PLoS ONE | 51 |
| 10 | Using AI tools in writing peer review reports: Should academic journals embrace the use of ChatGPT? (Garcia, 2024) | 2024 | Annals of Biomedical Engineering | 47 |
| 11 | Is ChatGPT corrupting peer review? Telltale words hint at AI use (Singh Chawla, 2024) | 2024 | Nature | 45 |
| 12 | Peer review in the age of generative AI (Kankanhalli, 2024) | 2024 | Journal of the Association for Information Systems | 42 |
| 13 | Artificial intelligence in peer review: Enhancing efficiency while preserving integrity (Doskaliuk et al., 2025) | 2025 | Journal of Korean Medical Science | 41 |
| 14 | Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews (Liang et al., 2024) | 2024 | Proceedings of Machine Learning Research | 35 |
| 15 | Generative artificial intelligence is infiltrating peer review process (Cheng et al., 2024) | 2024 | Critical Care | 35 |
| Category | Study | Model(s) | Sample | Finding |
|---|---|---|---|---|
| Predicting outcomes | AI in neurosurgical manuscript prediction | ChatGPT-4o, Gemini, Copilot | 51 preprints | Accuracy 40–66.67% (p < 0.001) |
| Predicting outcomes | Predicting outcomes from reviewer text | Fine-tuned GPT-4-mini, GPT-3, BERT, GPT-2, GRU | JNS journals, 2021–2023 | Best AUC 0.91 (fine-tuned); untrained 0.67–0.75 |
| Predicting outcomes | ChatGPT predictive capacity across platforms | ChatGPT (30-run avg.) | F1000Research, SciPost, ICLR | ρ = 0.00 to 0.46, platform-dependent |
| Predicting outcomes | LLM evaluation, transplantation research (Shen et al., 2026) | 5 open-source LLMs | 200 papers | Best: exact-match 0.35, loose-match 0.78 |
| Generating content | AI-generated discussion passes peer review | ChatGPT-4 | 1 article, 6 reviewers | 83% recommended acceptance after revision |
| Generating content | ChatGPT’s abilities in writing and review (Kadi & Aslaner, 2024) | ChatGPT 3.0/4.0 | 15 reports + 15 articles | Merit score 4.9/10; review weaker than generation |
| Generating content | ChatGPT in peer review, observational (Saad et al., 2024) | ChatGPT 3.5/4.0 vs. 2 humans | 21 articles | Agreement 3.6–3.76/5 |
| Generating content | Human vs. AI-generated review reports | ChatGPT | 398 reports, 119 articles | AI reviews more routine; human reviews more diverse |
| Generating content | Review Feedback Agent (randomized) | Multi-LLM agent | 20,000+ reviews, ICLR 2025 | 27% of reviewers revised after AI feedback |
| Generating content | TreeReview, question-tree generation (Chang et al., 2025) | LLM-based | ICLR/NeurIPS benchmark | Outperforms baselines; up to 80% reduction in token usage |
| Evaluating quality | Can LLMs assess peer review quality? | GPT-4o, Claude 3.5, Gemini | 180 eLife papers | ~70% agreement; Kappa poor-to-moderate |
| Evaluating quality | UPME (unsupervised MLLM evaluation) | Multimodal LLMs | MMstar, ScienceQA | Pearson r = 0.944 and 0.814 |
| Evaluating quality | PRE, LLM evaluation via peer review (Chu et al., 2024) | 11 LLMs incl. GPT-4 | Summarization, QA tasks | Outperforms baselines; confirms single-LLM bias |
| Evaluating quality | Can LLMs uphold research integrity? (Arabzadeh et al., 2026) | Zero-shot/fine-tuned LLMs vs. ML | Expert-annotated benchmarks | Adequate for structure; hybrid needed for facts |
| Diagnosing weaknesses | Where Do LLMs Go Wrong? (perturbation) | GPT-4o, Gemini 2.0, LLaMA 3 | Multi-conference data | Misclassifies flaws; overweights rejection tone |
| Diagnosing weaknesses | Breaking the Reviewer (adversarial) | Multiple LLMs | Adversarial test set | Text manipulation distorts assessments |
| Diagnosing weaknesses | AgentReview (bias simulation) | LLM agents | Simulated review process | 37.1% of decision variance from reviewer bias |
| Diagnosing weaknesses | Textual analysis of AI-conference manuscripts (Vincent-Lamarre & Larivière, 2021) | N/A (linguistic) | AI conference submissions | Word choice predicts acceptance independent of quality |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kim, E.; Singh, V. Artificial Intelligence in Peer Review: A Bibliometric-Guided Thematic Review and a Task-Contingent Legitimacy Framework. Publications 2026, 14, 63. https://doi.org/10.3390/publications14040063
Kim E, Singh V. Artificial Intelligence in Peer Review: A Bibliometric-Guided Thematic Review and a Task-Contingent Legitimacy Framework. Publications. 2026; 14(4):63. https://doi.org/10.3390/publications14040063
Chicago/Turabian StyleKim, Eungi, and Vaishali Singh. 2026. "Artificial Intelligence in Peer Review: A Bibliometric-Guided Thematic Review and a Task-Contingent Legitimacy Framework" Publications 14, no. 4: 63. https://doi.org/10.3390/publications14040063
APA StyleKim, E., & Singh, V. (2026). Artificial Intelligence in Peer Review: A Bibliometric-Guided Thematic Review and a Task-Contingent Legitimacy Framework. Publications, 14(4), 63. https://doi.org/10.3390/publications14040063

