On the Vulnerability of Citation Metrics in the Era of Generative Artificial Intelligence †
Abstract
1. Introduction
2. Background on Large Language Models and Generative Artificial Intelligence
3. A Systematic Literature Review of Citation Metrics in the Era of Generative Artificial Intelligence
- RQ 1: What methods of citation metric manipulation are reported?
- RQ 2: To what extent are citation metrics used in formal research evaluation procedures?
- ∘
- RQ 2.1: …specifically in the civil engineering domain?
- RQ 3: How widely are large language models being used to generate scientific papers?
3.1. Identification
- Definition of core concepts and initial search terms: For each research question, logical blocks of concepts are defined, for example, a “core concepts” block and a “general manipulation terms” block are established for RQ 1. Within each block, initial keywords (e.g., “citation metrics”, “citation indicators”) are expanded into specific search terms by considering synonyms, alternative expressions, and related terminology (e.g., “citation metric*”, “citation indicator*”). Truncation techniques and Boolean operators are applied to combine the search terms and to capture various linguistic variations. The sets of search terms form the basis of initial search strings defined for each research question.
- Conducting the initial search: Using the initial search strings specified for each research question, an initial search is carried out. The objective of this step is to obtain an initial overview of the retrieved records and to assess the completeness, specificity, and precision of the search strategy. The findings obtained from the initial search form the basis for subsequent refinement of the search strings.
- Refinement of the search strings: Following the initial search, the search strings for each research question are systematically revised. The revision includes verifying whether predefined seminal studies are retrieved by the current strings. If key studies are not captured, the strings are adjusted and extended to improve coverage. Conversely, overly broad or irrelevant terms that produce large numbers of non-relevant records are removed. The use of Boolean operators and truncation is further optimized to increase overall precision.
- Execution of the final search: Upon refinement, a final search string is established for each research question, targeting an appropriate balance between comprehensive coverage and high precision. The final search strings for each research question, organized into logical blocks, are illustrated in Figure 2 and reported in full in the Supplementary Materials (Table S1). The final search is conducted in the “title”, “abstract”, and “keywords” fields of the Scopus database. Additional filters are applied for language (English) and document type (article, review, and conference paper) to further increase precision and relevance.
- Removal of duplicates: All records returned by the final search are examined for duplicate entries, which are then removed manually.
3.2. Screening
3.3. Eligibility
3.4. Snowballing and Inclusion
3.5. Results
3.5.1. Results Regarding Research Question 1 (“What Methods of Citation Metric Manipulation Are Reported?”)
3.5.2. Results Regarding Research Question 2 (“To What Extent Are Citation Metrics Used in Formal Research Evaluation Procedures?”)
3.5.3. Results Regarding Research Question 2.1 (“To What Extent Are Citation Metrics Used in Formal Research Evaluation Procedures Specifically in the Civil Engineering Domain?”)
3.5.4. Results Regarding Research Question 3 (“How Widely Are Large Language Models Being Used to Generate Scientific Papers?”)
4. A Google Scholar Proof-of-Concept Case Study
- RQ 4: How easily can papers generated with the assistance of large language models influence author-level citation metrics on Google Scholar?
4.1. Experimental Design and Methodology
4.2. Results
4.3. Summary
5. Discussion, Implications, and Recommendations
5.1. Discussion and Implications
5.1.1. Research Question 1 (“What Methods of Citation Metric Manipulation Are Reported?”)
- Citation metric manipulation undermines the validity and accuracy of research evaluations, affecting critical academic decisions, such as faculty promotion, hiring procedures, performance-based remuneration, institutional benchmarking, and the allocation of research funding.
- Persistent manipulation reduces overall trust in scholarly publications, complicating the identification of genuinely impactful research and diminishing confidence in research outcomes.
- The rise of citation stacking and citation cartels reveals a more organized and collaborative form of metric manipulation that disadvantages ethical researchers.
- Reliance on manipulated metrics encourages a culture where quantity supersedes quality, thereby undermining robust scholarship and hindering the pursuit of innovative research.
5.1.2. Research Question 2 (“To What Extent Are Citation Metrics Used in Formal Research Evaluation Procedures?”)
- Dependence on citation metrics encourages strategic publishing behaviors aimed at maximizing citations, thereby reducing incentives for high-risk, innovative, or interdisciplinary research.
- Younger researchers and/or researchers from emerging disciplines face disadvantages due to the inherent preference of citation-based metrics for senior scholars and/or scholars situated in established disciplines, which traditionally generate higher citation frequencies.
- Uniform application of citation metrics across diverse disciplines and geographical contexts aggravates inequalities, disadvantaging institutions and researchers in regions with lower citation frequencies or less journal indexing coverage.
- Most critically, reliance on non-curated platforms, such as Google Scholar, allows researchers to manipulate their own citation metrics directly. Since Google Scholar indexes all types of academic output without strict quality control, researchers can artificially inflate their citation counts, potentially influencing key evaluation processes, tenure decisions, funding allocations, and even salaries if remuneration is tied to metric-based performance agreements.
5.1.3. Research Question 2.1 (“To What Extent Are Citation Metrics Used in Formal Research Evaluation Procedures Specifically in the Civil Engineering Domain?”)
- As in other disciplines, overemphasis on citation metrics directs researchers toward citation-rich topics rather than innovative, interdisciplinary, or practically impactful civil engineering research, limiting overall research diversity and quality.
- Researchers with substantial industry experience, crucial in civil engineering, may experience disadvantages if citation metrics are weighted too heavily in metric-centric evaluations, as periods spent in professional practice negatively impact the publication output.
- Citation-based evaluations systematically undervalue practically relevant research fields and niche areas within civil engineering, potentially impacting funding decisions, career opportunities, and the attractiveness of practice-oriented careers.
- A comparatively higher weighting of multi-authorship metrics may disadvantage individuals or smaller collaborative groups, particularly in less frequently cited but practically significant research.
5.1.4. Research Question 3 (“How Widely Are Large Language Models Being Used to Generate Scientific Papers?”)
- Despite the evident advantages, increased reliance on LLMs in scholarly communication risks diluting research quality and eroding trust in published academic literature, exacerbating existing problems, such as predatory publishing.
- Although democratizing academic productivity and accessibility, particularly for researchers from non-native English-speaking regions, widespread use of LLMs without proper disclosure compromises transparency, accountability, and intellectual rigor in scholarly discourse.
- Established norms of authorship and attribution face disruption due to the emerging practice of explicitly acknowledging generative AI, thereby creating ambiguity regarding responsibility and genuine scholarly contribution in published research.
5.1.5. Research Question 4 (“How Easily Can Papers Generated with the Assistance of Large Language Models Influence Author-Level Citation Metrics on Google Scholar?”)
- The simplicity of artificially enhancing citation metrics through AI-generated content poses fundamental risks to the integrity and fairness of citation-based academic evaluations.
- Bibliographic databases operating without or with limited quality control, such as Google Scholar, are highly susceptible to citation manipulations, a vulnerability further aggravated by the free online access to LLM-based chatbots.
- The minimal effort and limited expertise required to produce apparently legitimate scholarly articles significantly lower the barriers for citation metric manipulation.
5.2. Recommendations
- Researchers
- Adhering to ethical citation practices: Researchers should strictly comply with ethical citation standards, transparently disclose self-citations, and clearly acknowledge the contributions of each author. Additionally, researchers should actively avoid participation in citation cartels or other manipulative citation schemes, documented in Section 3.
- Disclosing and verifying AI usage: Any use of large language models in research or manuscript preparation should be explicitly disclosed. Researchers should thoroughly verify all AI-generated content—particularly references, empirical data, and methodological details—to ensure accuracy and prevent dissemination of fabricated and/or false information that may propagate into bibliographic indexing systems.
- Prioritizing quality over metrics: Research quality and real-world relevance should take precedence over strategies aimed solely at maximizing citation counts. Particularly in applied research disciplines, such as civil engineering, researchers are encouraged to document practical industry engagements and collaborative efforts, demonstrating genuine impact beyond purely bibliometric indicators used in formal evaluation frameworks (Section 3.5.2).
- Academic institutions and committees
- 4.
- Diversifying research assessment: Academic institutions and committees should reduce their reliance on purely quantitative bibliometric indicators, particularly on indicators derived from non-curated bibliographic databases without or with limited quality control, such as Google Scholar, whose platform-specific sensitivities are demonstrated in Section 4. Evaluation procedures should shift toward more diverse assessments, emphasizing long-term research impact, innovation, interdisciplinary collaboration, and tangible societal benefits.
- 5.
- Recognizing industry and interdisciplinary work: Evaluation frameworks should explicitly recognize and fairly assess industry experience, practical achievements, and interdisciplinary engagement of researchers, which is particularly relevant in applied disciplines, such as civil engineering, where industry collaborations significantly contribute to research impact yet may be inadequately represented by traditional citation metrics.
- 6.
- Establishing ethics and transparency policies: Institutions should implement clear policies addressing unethical citation practices, with defined sanctions for citation manipulation, such as excessive self-citation or organized citation exchanges, identified in the systematic review (Section 3.5.1). Policies should mandate transparency regarding generative AI usage in scholarly work and evaluation materials, ensuring ethical and openly disclosed use.
- Publishers and editors
- 7.
- Strengthening editorial guidelines: Stringent peer-review and editorial guidelines should be established to prevent unethical citation behavior and the dissemination of AI-generated content. Measures should identify and mitigate coercive citation suggestions and irregular reference structures consistent with manipulation mechanisms summarized in Section 3.
- 8.
- Implementing clear policies on AI and authorship: Publishers should define explicit guidelines concerning the appropriate use of generative AI tools in research and manuscript preparation. To uphold trust in authorship attribution and scholarly content integrity, guidelines should specify methods for acknowledging AI assistance (e.g., in an acknowledgments section) and clearly reaffirm that all listed authors must be accountable human contributors.
- 9.
- Enhancing AI-awareness in the review process: Editors and reviewers should receive training and develop expertise to better recognize and critically assess AI-generated content during manuscript review. Editorial workflows should integrate verification procedures proportionate to documented risks of reference hallucination and citation inflation.
- Providers of bibliographic databases
- 10.
- Detecting citation anomalies: Providers of bibliographic databases and indexing platforms should (further) develop and refine automated systems to identify manipulated citation patterns, counterfeit references, and papers potentially produced by paper mills or AI-assisted pipelines. Algorithms should monitor abnormal citation clustering and sequential indexing effects similar to those demonstrated in Section 4.
- 11.
- Flagging questionable content: Upon detecting manipulative practices or AI-generated content, platforms should clearly flag affected records, which is particularly pertinent for platforms without rigorous editorial oversight (e.g., Google Scholar or certain preprint servers). Explicit markers or warnings on suspicious articles should alert users about potential reliability issues.
- 12.
- Ensuring transparency in indexing: Bibliographic databases should adopt enhanced transparency measures, introducing clear labeling or markers indicating AI-generated content or records with irregular citation patterns. Providing explicit metadata regarding indexing processes strengthens the transparency and integrity of citation data and supports user trust in bibliometric evaluations.
- Funding institutions and policymakers
- 13.
- Rewarding long-term impact over short-term metrics: Funding institutions and research policymakers should revise evaluation criteria to reward genuine innovation, long-term impact, collaboration, and societal value, rather than emphasizing high bibliometric scores; greater emphasis should be placed on sustained contributions of research projects, rather than on immediate citation counts or journal impact factors.
- 14.
- Establishing policies against metric gaming: Explicit guidelines should discourage strategic manipulation of citation metrics in grant applications and evaluations. Funding institutions should revise assessment frameworks to prioritize qualitative indicators of research excellence (including originality, rigor, and societal significance) over publication quantity or citation volume, to signal that attempts to “game the system” will not confer funding advantages.
- 15.
- Encouraging AI transparency and ethical standards in proposals: Funding institutions should mandate transparent disclosure concerning the use of generative AI tools throughout the research lifecycle, including proposal preparation, peer-review processes, and research execution. Applicants should be required to openly report AI involvement in preparing proposals or analyzing data and to explain how such use complies with ethical standards and rigorous research methodologies. In addition, funding institutions should establish mechanisms for monitoring unethical practices in evaluation-relevant submissions.
6. Summary and Conclusions
Supplementary Materials
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial intelligence |
| LLM | Large language model |
| PRISMA | Preferred reporting items for systematic reviews and meta-analyses |
| RQ | Research question |
| XAICE | Explainable artificial intelligence in civil engineering |
References
- Abramo, G., Cicero, T., & D’Angelo, C. A. (2012). Revisiting the scaling of citations for research assessment. Journal of Informetrics, 6(4), 470–479. [Google Scholar] [CrossRef] [Scilit]
- Abramo, G., D’Angelo, C. A., & Grilli, L. (2021). The effects of citation-based research evaluation schemes on self-citation behavior. Journal of Informetrics, 15(4), 101204. [Google Scholar] [CrossRef] [Scilit]
- Aksnes, D. W., Langfeldt, L., & Wouters, P. (2019). Citations, citation indicators, and research quality: An overview of basic concepts and theories. SAGE Open, 9(1), 1–17. [Google Scholar] [CrossRef] [Scilit]
- Ali, I., Sultan, P., & Aboelmaged, M. (2021). A bibliometric analysis of academic misconduct research in higher education: Current status and future research opportunities. Accountability in Research, 28(6), 372–393. [Google Scholar] [CrossRef] [Scilit]
- Ali, N., Halim, Z., & Hussain, S. F. (2023). An artificial intelligence-based framework for data-driven categorization of computer scientists: A case study of world’s top 10 computing departments. Scientometrics, 128(3), 1513–1545. [Google Scholar] [CrossRef] [Scilit]
- Altanopoulou, P., Dontsidou, M., & Tselios, N. (2012). Evaluation of ninety-three major Greek university departments using Google Scholar. Quality in Higher Education, 18(1), 111–137. [Google Scholar] [CrossRef] [Scilit]
- Asfour, O. S., & Al-Qawasmi, J. (2024). Research metrics in architecture: An analysis of the current challenges compared to engineering disciplines. Publications, 12(4), 50. [Google Scholar] [CrossRef] [Scilit]
- Besançon, L., Cabanac, G., Labbé, C., & Magazinov, A. (2023). Sneaked references: Cooked reference metadata inflate citation counts. Journal of the Association for Information Science and Technology, 74(7), 699–712. [Google Scholar]
- Cabezas-Clavijo, Á., Magadán-Díaz, M., Rivas-García, J. I., & Sidorenko-Bautista, P. (2024). This book is written by ChatGPT: A quantitative analysis of ChatGPT authorships through Amazon.com. Publishing Research Quarterly, 40(2), 147–163. [Google Scholar] [CrossRef] [Scilit]
- Camp, N. T., Bengtson, J. A., & Sandstrom, J. C. (2025). The citation catastrophe: Propagation of AI-generated counterfeit citations in scholarship. The Journal of Academic Librarianship, 51(4), 103065. [Google Scholar] [CrossRef] [Scilit]
- Chomsky, N. (1956). Three models for the description of language. IRE Transactions on Information Theory, 2(3), 113–124. [Google Scholar] [CrossRef] [Scilit]
- CoARA (Coalition for Advancing Research Assessment). (2022). Agreement on reforming research assessment. Available online: https://coara.eu/agreement/ (accessed on 27 March 2026).
- Dehnad, A., Abdekhoda, M., & Fallah Atatalab, F. (2019). H-index and promotion decisions. Annals of Library and Information Studies, 66(4), 171–175. [Google Scholar]
- DORA (San Francisco Declaration on Research Assessment). (2012). San Francisco declaration on research assessment. Available online: https://sfdora.org/read/ (accessed on 27 March 2026).
- Dragos, K. (2025a, May 5). The impact of generative artificial intelligence on a metric-driven academic system—A review. 7th International Workshop on Explainable Artificial Intelligence in Civil Engineering (XAICE), Hamburg, Germany. Available online: https://smarsly.wordpress.com/wp-content/uploads/2025/06/dragos2025b.pdf (accessed on 27 March 2026).
- Dragos, K. (2025b, May 5). A review of the h-index in the era of generative artificial intelligence. 7th International Workshop on Explainable Artificial Intelligence in Civil Engineering (XAICE), Hamburg, Germany. Available online: https://smarsly.wordpress.com/wp-content/uploads/2025/06/dragos2025c.pdf (accessed on 27 March 2026).
- El-Adaway, A. G., Assaad, R., Elsayegh, A., & Abotaleb, I. S. (2019). Analytic overview of citation metrics in the civil engineering domain with focus on construction engineering and management specialty area and its subdisciplines. Journal of Construction Engineering and Management, 145(10), 04019060. [Google Scholar] [CrossRef] [Scilit]
- European Commission. (2024). Living guidelines on the responsible use of generative AI in research. Available online: https://research-and-innovation.ec.europa.eu (accessed on 27 March 2026).
- Finkel-Gates, A. (2025). ChatGPT in academic assessments: Upholding integrity. Journal of Learning Development in Higher Education, (36). [Google Scholar] [CrossRef] [Scilit]
- Fire, M., & Guestrin, C. (2019). Over-optimization of academic publishing metrics: Observing Goodhart’s Law in action. GigaScience, 8(6), giz053. [Google Scholar] [CrossRef] [Scilit]
- Fister, I., Jr., Fister, I., & Perc, M. (2016). Toward the discovery of citation cartels in citation networks. Frontiers in Physics, 4, 49. [Google Scholar] [CrossRef] [Scilit]
- Floridi, L., & Chiriatti, M. (2020). GPT-3: Its nature, scope, limits, and consequences. Minds and Machines, 30(4), 681–694. [Google Scholar] [CrossRef] [Scilit]
- Fong, E. A., & Wilhite, A. W. (2017). Authorship and citation manipulation in academic research. PLoS ONE, 12(12), e0187394. [Google Scholar] [CrossRef] [Scilit]
- Goyanes, M., Lopezosa, C., & Piñeiro-Naval, V. (2025). The use of artificial intelligence (AI) in research: A review of author guidelines in leading journals across eight social science disciplines. Scientometrics, 130(7), 3725–3741. [Google Scholar] [CrossRef] [Scilit]
- Guraya, S. Y., Norman, R. I., Khoshhal, K. I., Guraya, S. S., & Forgione, A. (2016). Publish or Perish mantra in the medical field: A systematic review of the reasons, consequences and remedies. Pakistan Journal of Medical Sciences, 32(6), 1562–1567. [Google Scholar] [CrossRef] [Scilit]
- Haddow, G., & Hammarfelt, B. (2019). Quality, impact, and quantification: Indicators and metrics use by social scientists. Journal of the Association for Information Science and Technology, 70(1), 16–26. [Google Scholar] [CrossRef] [Scilit]
- Haider, J., Söderström, K. R., Ekström, B., & Rödl, M. (2024). GPT-fabricated scientific papers on Google Scholar: Key features, spread, and implications for preempting evidence manipulation. Harvard Kennedy School Misinformation Review, 5(5). [Google Scholar] [CrossRef] [Scilit]
- Hicks, D., Wouters, P., Waltman, L., de Rijcke, S., & Rafols, I. (2015). Bibliometrics: The Leiden Manifesto for research metrics. Nature, 520, 429–431. [Google Scholar] [CrossRef] [Scilit]
- Hosseini, M., & Horbach, S. P. J. M. (2023). Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer review. Research Integrity and Peer Review, 8(1), 4. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hosseini, M., Resnik, D. B., & Holmes, K. (2023). The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts. Research Ethics, 19(4), 449–465. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ibrahim, H., Liu, F., Zaki, Y., & Rahwan, T. (2025). Citation manipulation through citation mills and pre-print servers. Scientific Reports, 15(1), 5480. [Google Scholar] [CrossRef] [Scilit]
- Kazakis, N. A. (2014). Bibliometric evaluation of the research performance of the Greek civil engineering departments in National and European context. Scientometrics, 101(1), 505–525. [Google Scholar] [CrossRef] [Scilit]
- Kendall, G. (2024). More transparency is needed when citing h-indexes, journal impact factors and citescores. Publishing Research Quarterly, 40(1), 80–99. [Google Scholar] [CrossRef] [Scilit]
- Kendall, G., & Teixeira da Silva, J. A. (2024). Risks of abuse of large language models, like ChatGPT, in scientific publishing: Authorship, predatory publishing, and paper mills. Learned Publishing, 37(1), 55–62. [Google Scholar] [CrossRef] [Scilit]
- Kobak, D., González-Márquez, R., Horvát, E.-A., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances, 11(27), eadt3813. [Google Scholar] [CrossRef] [Scilit]
- Kojaku, S., Livan, G., & Masuda, N. (2021). Detecting anomalous citation groups in journal networks. Scientific Reports, 11(1), 14524. [Google Scholar] [CrossRef] [Scilit]
- Labbé, C. (2010). Ike Antkare one of the great stars in the scientific firmament. International Society for Scientometrics and Informetrics Newsletter, 6(2), 48–52. [Google Scholar]
- Labbé, C., & Labbé, D. (2013). Duplicate and fake publications in the scientific literature: How many SCIgen papers in computer science? Scientometrics, 94, 379–396. [Google Scholar] [CrossRef] [Scilit]
- Lendvai, G. F. (2025). ChatGPT in academic writing: A scientometric analysis of literature published between 2022 and 2023. Journal of Empirical Research on Human Research Ethics, 20(3), 131–148. [Google Scholar] [CrossRef] [Scilit]
- Lim, B. H., D’Ippoliti, C., Dominik, M., Hernández-Mondragón, A. C., Vermeir, K., Chong, K. K., Hussein, H., Morales-Salgado, V. S., Cloete, K. J., Kimengsi, J. N., Balboa, L., Mondello, S., dela Cruz, T. E., Lopez-Verges, S., Sidi Zakari, I., Simonyan, A., Palomo, I., Režek Jambrak, A., Germo Nzweundji, J., … Bueso, F. (2025). Regional and institutional trends in assessment for academic promotion. Nature, 638(8050), 459–468. [Google Scholar] [CrossRef] [Scilit]
- Lippi, G., & Mattiuzzi, C. (2017). Scientist impact factor (SIF): A new metric for improving scientists’ evaluation? Annals of Translational Medicine, 5(15), 303. [Google Scholar] [CrossRef] [Scilit]
- López-Cózar, E. D., Robinson-García, N., & Torres-Salinas, D. (2014). The Google Scholar experiment: How to index false papers and manipulate bibliometric indicators. Journal of the Association for Information Science and Technology, 65(3), 446–454. [Google Scholar] [CrossRef] [Scilit]
- Lund, B. D., & Naheem, K. T. (2024). Can ChatGPT be an author? A study of artificial intelligence authorship policies in top academic journals. Learned Publishing, 37(1), 13–21. [Google Scholar] [CrossRef] [Scilit]
- Lund, B. D., Wang, T., Mannuru, N. R., Nie, B., Shimray, S., & Wang, Z. (2023). ChatGPT and a new academic reality: Artificial Intelligence-written research papers and the ethics of large language models in scholarly publishing. Journal of the Association for Information Science and Technology, 74(5), 570–581. [Google Scholar] [CrossRef] [Scilit]
- Marsicano, C. R., Braxton, J. M., & Nichols, A. R. K. (2022). The use of Google Scholar for tenure and promotion decisions. Scientometrics, 47(4), 639–660. [Google Scholar] [CrossRef] [Scilit]
- Mazov, N. A., & Gureev, V. N. (2019, September 2). Detection of inappropriate types of authorship using bibliometric approaches. 17th Conference of the International Society for Scientometrics and Informetrics, Rome, Italy. [Google Scholar]
- Mehregan, M., & Moghiman, M. (2024). The unnoticed issue of coercive citation behavior for authors. Publishing Research Quarterly, 40(2), 164–168. [Google Scholar] [CrossRef] [Scilit]
- Min, B., Ross, H., Sulem, E., Veyseh, A. P. B., Nguyen, T. H., Sainz, O., Agirre, E., Heinz, I., & Roth, D. (2023). Recent advances in natural language processing via large pre-trained language models: A survey. ACM Computing Surveys, 56(2), 30. [Google Scholar] [CrossRef] [Scilit]
- Mingers, J., O’Hanley, J. R., & Okunola, M. (2023). Using Google Scholar institutional level data to evaluate the quality of university research. Scientometrics, 113(3), 1627–1643. [Google Scholar] [CrossRef] [Scilit]
- Moustafa, K. (2016). Aberration of the citation. Accountability in Research, 23(4), 230–244. [Google Scholar] [CrossRef] [Scilit]
- Mustafa, G., Afzal, M. T., Rauf, A., & Khan, M. A. (2025). Beyond publication numbers: A novel approach to academic ranking using evolutionary programming. Evolutionary Intelligence, 18(3), 62. [Google Scholar] [CrossRef] [Scilit]
- Nature. (2023). Tools such as ChatGPT threaten transparent science; here are our ground rules for their use (Editorial). Nature, 613, 612. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nicholas, D., Herman, E., Jamali, H. R., Abrizah, A., Boukacem-Zeghmouri, C., Xu, J., Rodríguez-Bravo, B., Watkinson, A., Polezhaeva, T., & Świgon, M. (2020). Millennial researchers in a metric-driven scholarly world: An international study. Research Evaluation, 29(3), 263–274. [Google Scholar] [CrossRef] [Scilit]
- Ortega, J.-L., & Delgado-Quirós, L. (2023). How do journals deal with problematic articles. Editorial. Profesional de la información, 32(1), e320118. [Google Scholar] [CrossRef] [Scilit]
- Öztürk, O., & Taşkın, Z. (2024). How metric-based performance evaluation systems fuel the growth of questionable publications? Scientometrics, 129(5), 2729–2748. [Google Scholar] [CrossRef] [Scilit]
- Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. [Google Scholar] [CrossRef] [Scilit]
- Raheel, M., Ayaz, S., & Afzal, M. T. (2018). Evaluation of h-index, its variants and extensions based on publication age and citation intensity in civil engineering. Scientometrics, 114(3), 1107–1127. [Google Scholar] [CrossRef] [Scilit]
- Ramadhan, M. A., Sutarto, S., Widodo, S., & Anisah, M. (2024). Bibliometric analysis to reveal research evolution and educational technology trends in civil engineering education. International Journal of Learning, Teaching and Educational Research, 23(3), 87–110. [Google Scholar] [CrossRef] [Scilit]
- Ramoni, D., Sgura, C., Liberale, L., Montecucco, F., Ioannidis, J. P. A., & Carbone, F. (2024). Artificial intelligence in scientific medical writing: Legitimate and deceptive uses and ethical concerns. European Journal of Internal Medicine, 127, 31–35. [Google Scholar] [CrossRef] [Scilit]
- Salman, M., Ahmed, M. M., & Afzal, M. T. (2021). Assessment of author ranking indices based on multi-authorship. Scientometrics, 126(5), 4153–4172. [Google Scholar] [CrossRef] [Scilit]
- Smarsly, K. (2026, March 23–26). On the reliability of citation metrics in civil and building engineering in the age of generative artificial intelligence. The International Conference on Computing in Civil and Building Engineering (ICCCBE), Taipei, Taiwan. [Google Scholar]
- Tang, A., Li, K.-K., Kwok, K. O., Cao, L., Luong, S., & Tam, W. (2024). The importance of transparency: Declaring the use of generative artificial intelligence (AI) in academic writing. Journal of Nursing Scholarship, 56(2), 314–318. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Thelwall, M., & Kurt, Z. (2025). Research evaluation with ChatGPT: Is it age, country, length, or field biased? Scientometrics, 130(10), 5323–5343. [Google Scholar] [CrossRef] [Scilit]
- Tripathi, M., Sonkar, S. K., & Kumar, S. (2019). A cross sectional study of retraction notices of scholarly journals of science. Journal of Library and Information Technology, 39(2), 74–81. [Google Scholar] [CrossRef] [Scilit]
- Usman, M., Mustafa, G., & Afzal, M. T. (2021). Ranking of author assessment parameters using logistic regression. Scientometrics, 126(1), 335–353. [Google Scholar] [CrossRef] [Scilit]
- Vincent, J. (2023). Top academic publisher bans AI-generated text in scientific papers. The Register. Available online: https://www.theregister.com/2023/01/27/top_academic_publisher_science_bans/ (accessed on 30 June 2025).
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13(1), 14045. [Google Scholar] [CrossRef] [Scilit]
- Wang, R., Zhou, Y., & Zeng, A. (2022). Evaluating scientists by citation and disruption of their representative works. Scientometrics, 128(3), 1689–1710. [Google Scholar] [CrossRef] [Scilit]
- Wilsdon, J., Allen, L., Belfiore, E., Campbell, P., Curry, S., Hill, S., Jones, R., Kain, R., Kerridge, S., Thelwall, M., Tinkler, J., Viney, I., Wouters, P., Hill, J., & Johnson, B. (2015). The metric tide: Report of the independent review of the role of metrics in research assessment and management. Higher Education Funding Council for England (HEFCE). [Google Scholar]
- Wohlin, C. (2014, May 4). Guidelines for snowballing in systematic literature studies and a replication in software engineering. 18th International Conference on Evaluation and Assessment in Software Engineering, London, UK. [Google Scholar]
- Yoo, J.-H. (2025). Defining the boundaries of AI use in scientific writing: A comparative review of editorial policies. Journal of Korean Medical Science, 40(23), e187. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zerem, E., Kunosić, S., Imširović, B., & Kurtčehajić, A. (2021). Science metrics systems and academic promotion: Bosnian reality, 2021. Psychiatria Danubina, 33(Suppl. S3), S371–S377. [Google Scholar]
- Zhang, M., & Zhao, T. (2025). Citation accuracy challenges posed by large language models. JMIR Medical Education, 11, e72998. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Q., Abraham, J., & Fu, H. Z. (2020). Collaboration and its influence on retraction based on retracted publications during 1978–2017. Scientometrics, 125(1), 213–232. [Google Scholar] [CrossRef] [Scilit]
- Zubiaga, A. (2024). Natural language processing in the era of large language models. Frontiers in Artificial Intelligence, 6, 1350306. [Google Scholar] [CrossRef] [Scilit]





| Paper A (Dragos, 2025a) | Paper B (Dragos, 2025b) | |
|---|---|---|
| Author’s papers cited | 50 | 75 |
| External papers cited | 60 | 40 |
| Dissemination date | 5 June 2025 | 28 June 2025 |
| Indexing date | 19 June 2025 | 15 July 2025 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Smarsly, K. On the Vulnerability of Citation Metrics in the Era of Generative Artificial Intelligence. Publications 2026, 14, 23. https://doi.org/10.3390/publications14020023
Smarsly K. On the Vulnerability of Citation Metrics in the Era of Generative Artificial Intelligence. Publications. 2026; 14(2):23. https://doi.org/10.3390/publications14020023
Chicago/Turabian StyleSmarsly, Kay. 2026. "On the Vulnerability of Citation Metrics in the Era of Generative Artificial Intelligence" Publications 14, no. 2: 23. https://doi.org/10.3390/publications14020023
APA StyleSmarsly, K. (2026). On the Vulnerability of Citation Metrics in the Era of Generative Artificial Intelligence. Publications, 14(2), 23. https://doi.org/10.3390/publications14020023
