Next Article in Journal
Literacy Before It Counts: Legibility and Recognition of Adult Literacy Beyond Standardized Assessment
Previous Article in Journal
Low-Proficiency Students’ Cognitive and Affective Engagement with Combined Audio and Indirect Written Feedback in an EFL Writing Class
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Generative Artificial Intelligence and the Ambiguity of Academic Integrity in Higher Education

by
Katerina Zdravkova
Faculty of Computer Science and Engineering, Ss. Cyril and Methodius University in Skopje, Skopje, North Macedonia
Educ. Sci. 2026, 16(7), 1120; https://doi.org/10.3390/educsci16071120
Submission received: 10 June 2026 / Revised: 26 June 2026 / Accepted: 6 July 2026 / Published: 13 July 2026
(This article belongs to the Section Technology Enhanced Education)

Abstract

Large language models (LLMs) have introduced new challenges to academic integrity, particularly regarding the appropriation of AI-generated outputs as original human authorship and the difficulty of verifying independent work. While some universities and academic publishers increasingly require explicit disclosure of the use of artificial intelligence (AI), the scope and implementation of these requirements remain inconsistent. This paper examines current practices related to AI use, focusing on LLM-based ghostwriting and the reliability of disclosed interactions as evidence of authentic use. The study includes an experimental component involving AI-assisted essay generation, highlighting practical and ethical dilemmas associated with academic integrity. It further explores the possibility of mimicking authentic interactions, which raises concerns about the effectiveness of current approaches. To investigate these questions, a survey was conducted among teaching staff at the Faculty of Computer Science and Engineering (FCSE) in Skopje to assess their ability to identify AI-generated essays and their trust in disclosed interactions. Among the 28 respondents, a majority (82.14%) indicated that it is possible to identify AI-generated content based solely on language style, while 64.29% reported detecting linguistic inconsistencies that could result from the use of LLMs. Despite noticing AI-related linguistic markers, only 53.57% concluded that the essay was not human-written. This view was shared by just 27.27% of assistants, compared to 70.59% of professors, whose extensive experience appeared to help them recognize that a substantial portion of the text had been AI-generated. The findings are discussed in the context of teaching experience and existing policies, leading to recommendations for improving student assessment and strengthening the ethical use of generative artificial intelligence (GenAI).

1. Introduction

The sudden rise and adoption of generative artificial intelligence (GenAI) have introduced a new and deeply complex dimension of academic integrity (Yusuf et al., 2024; Zdravkova & Ilijoski, 2025a). ChatGPT, Claude, Gemini and other large language models (LLMs) offer incredible possibilities for enhancing learning, creativity, and productivity (Rane et al., 2023). At the same time, they initiate new temptations to bypass the norms of honest intellectual labour, not only for students (Leaton Gray et al., 2025), but increasingly for scholars and researchers as well (Shaw, 2025).
Universities and academic publishers are gradually requiring transparency about the use of artificial intelligence (AI) and LLMs in academic work. However, the level of expected requirements significantly varies. Prestigious academic institutions such as Princeton University (2026) have some courses that go beyond simple disclosure and may ask students to retain or submit records of their interactions with AI systems, while others like Stanford University (2023) or University of Oxford (2026) generally require clear acknowledgment of how AI tools were used. In scholarly publishing, major publishers including Elsevier (2026) and Springer Nature (2026) have formal policies mandating that authors explicitly disclose the use of GenAI, typically in a dedicated section describing the tool and its role in the research or writing process. While full submission of prompt–response logs is not yet a universal requirement, there is a clear trend toward more detailed reporting, especially in contexts emphasizing research transparency and reproducibility. Moreover, apart from demanding the authors to “disclose details on how the GenAI tool was used”, when confirming acceptance of reviewer status, MDPI (2023) asks reviewers not to use AI-assisted tools to review submissions or to generate peer-review reports.
Similarly to Princeton University, the Faculty of Computer Science and Engineering (FCSE) recommends explicit disclosure of interaction with GenAI in the preparation of home assignments and theses. While several teachers consistently implement this requirement, for example in the Computer Ethics course (Zdravkova & Ilijoski, 2025b), where students are asked to share their AI interaction history as a means of identifying potential academic misconduct, many do not require it. In practice, implementation varies: some teachers do not care too much about formal verification of academic integrity, others lack the time to check how students perform their work, while some are sceptical about the reliability and authenticity of the interactional evidence provided by students.
Motivated by this concern, a survey was conducted among FCSE teaching staff to evaluate whether they could reliably identify written assignments generated entirely by AI, as well as whether they believed that links shared by students reflected genuine interactions with the chatbots. Exactly one third of the participants invited responded to the survey, providing valuable insights and suggestions. Such a response rate is rather low, particularly among the more experienced teachers, whose response rate was only 25%. It seems that the colleagues who decided to participate may have had stronger opinions regarding generative AI, greater familiarity with AI tools, or a stronger interest in academic integrity issues than non-respondents. Their feedback offers a preliminary indication of FCSE’s current perceptions, and the challenges associated with verifying AI-assisted academic work. The findings contribute to the broader debate about academic integrity by highlighting the uncertainty of educators in disclosing GenAI-generated content and prompting the development of new, more robust assessment and policy frameworks within higher education. These questions are the leitmotif of this paper.
The paper continues with an overview of how universities and academic publishers have responded to the challenges posed by GenAI, focusing on issues of academic integrity, authorship, assessment practices, and emerging institutional policies. Then, the methodology and the methods are briefly announced. Section 4 presents in detail a short essay writing experiment conducted as part of the survey, in which ChatGPT was used to generate or support essay preparation. This section also addresses the associated ethical dilemmas concerning academic honesty. The fifth section examines the imitation of authentic interactions with ChatGPT. The sixth section presents the quantitative and qualitative results of the survey, with particular emphasis on approaches for distinguishing between independent student work and content produced by GenAI, as well as recommendations for preserving academic integrity and the ethical use of such technologies. The paper concludes with a critical reflection on the growing ambiguity of academic integrity originating from GenAI use.

2. Generative AI and Emerging Challenges for Academic Integrity

The emergence of GenAI and its unexpected reception, which according to some sources reaches around 1 billion active users weekly (Correa, 2026; Kemp, 2025) has significantly reshaped the challenges related to academic integrity, authorship and assessment in higher education. LLMs such as ChatGPT, Gemini, and Claude have challenged traditional assumptions about independent student work, prompting institutions and researchers to rethink how originality and learning outcomes are defined and evaluated (Kasneci et al., 2023; Zdravkova, 2025). Early responses to these challenges mainly focused on identifying risks to academic honesty, particularly regarding the unacknowledged assistance of AI and the potential erosion of the boundaries of authorship (Zdravkova & Ilijoski, 2025a; Zdravkova, 2025).
To demonstrate the risk that individual written papers could be replaced by papers generated by AI, an interesting experiment was conducted at a psychology department at a prestigious British university (Scarfe et al., 2024). The teachers included papers written entirely by LLMs, with 94% of these papers not being detected. What is even worse, the grades of these papers surpassed the grades of the actual student papers (Scarfe et al., 2024).
Inspired by this project, The Guardian conducted its own research into breaches of academic integrity and found around 7000 proven cases of cheating using AI tools in the academic year 2023–2024, which increased the percentage of cheating from 1.6 to 5.1 per 1000 cases (Goodier, 2025). Nanyang Technological University penalized students with zero grades for submitting AI-generated essays containing fabricated quotes (OECD, 2025). Data from South Korea’s Ministry of Education documented hundreds of incidents of misconduct, including dozens involving ChatGPT, all of which resulted in poor grades (Jisoo, 2025).
Many of these disciplinary actions, even when they were limited to warnings, have provoked significant backlash from students and parents, particularly when allegations of academic dishonesty were based on evaluations produced by AI detection tools whose reliability has been widely questioned (Edwards, 2023). Legal disputes indicate that they are increasingly being challenged in court. For example, parents of students at Hingham High School filed a lawsuit, claiming that the school did not have a clear policy governing such use at the time and that the sanctions were disproportionate (Nelken-Zitser, 2024; McLogan, 2025). Similar tensions have emerged in higher education, where students have launched legal actions against institutions such as Adelphi University and Emory University, challenging allegations of AI-assisted cheating or misuse of AI-based tools (Langreo, 2024; Sedacca, 2024; Nguyen, 2025).
As a result, universities and academic publishers have increasingly introduced policies requiring students and authors to report their use of generative AI (Bai et al., 2025). However, research highlights significant inconsistencies in how such policies are formulated and implemented across institutions (Cotton et al., 2024). While some guidelines encourage transparency and responsible use (Barman et al., 2024), others remain unclear about what constitutes acceptable assistance in the face of academic misconduct (Yang et al., 2024). The lack of standardization raises concerns about the practical effectiveness of approaches based on data disclosure without the inclusion of additional procedures in ensuring academic integrity (Laskar et al., 2024).
While initial responses to the impact of GenAI on academic integrity have focused largely on detection, remediation, and enforcement, these approaches have revealed significant limitations in practical reliability and educational fairness. Detection systems remain vulnerable to false positives and false negatives (Zhong et al., 2023), while disciplinary responses often struggle to distinguish acceptable assistance from inappropriate delegation of academic work.
At the same time, the focus on identifying misconduct can distract from the broader educational goals of assessment, such as fostering critical thinking, subject mastery, and independent inquiry. According to many educationalists, it is therefore inevitable to reexamine the design of student assessment in the LLM era (Cotton et al., 2025).
The alternative assessment strategies (Sullivan et al., 2023) aimed at reducing the potential for undetected AI assistance include oral exams (Bayley et al., 2026; Zdravkova & Ilijoski, 2025a), an iterative process of handing in assignments that show the process of their creation (McIntire et al., 2024), public defence of results (Kihwele et al., 2025), and even portfolio-based evaluation (Kihwele et al., 2025). All these methods emphasize process over product, with the goal of decreasing opportunities for undetected AI assistance, while encouraging deeper engagement with learning and doing research (Sullivan et al., 2023).

Academic Publishing in the LLM Era

The growing accessibility of generative AI tools has also created significant challenges for academic publishers, particularly regarding LLM-based ghostwriting (Gravel et al., 2023), fabricated citations (Lee et al., 2025; Baines, 2025), and paraphrased plagiarism (Hong, 2025). Despite policies and measures that emphasize that AI systems cannot be cited as authors because they cannot take responsibility for the integrity and originality of scientific work (Elsevier, 2026; Springer Nature, 2026), publishers have reported an increasing number of papers suspected of containing undisclosed AI-generated content (Hong, 2025). Editors have described cases in which manuscripts appeared to be AI-generated rewrites of existing publications or contained fabricated and unverifiable citations generated by LLMs (Hong, 2025).
Similar concerns have emerged at scientific conferences, where submissions were withdrawn after reviewers identified fictitious references and inconsistencies characteristic of AI-generated text (Van Dis et al., 2023; Dwivedi et al., 2023). Investigations by publishers have additionally revealed articles and book chapters containing substantial numbers of non-existent or misleading references, raising concerns about the reliability of AI-assisted academic writing (fourDNet, 2025).
These examples have exposed important limitations of existing detection approaches. Like educational institutions, in publishing, it has been argued that AI detection tools sometimes produce false positives and false negatives, making it difficult to rely solely on automated systems for manuscript evaluation (Zhong et al., 2023). As a result, many publishers have increasingly turned to disclosure policies, enhanced editorial oversight, plagiarism checks, and even post-publication research integrity procedures, rather than relying on technical detection mechanisms. (Elsevier, 2026; Springer Nature, 2026).
The emergence of LLM-based ghostwriting (Rodrigues, 2025) has therefore not only challenged traditional understandings of authorship and originality in scholarly publishing but has also intensified broader debates about transparency (Barman et al., 2024), accountability (Zhong et al., 2023) and trust in academic communication (Thorne, 2024). Therefore, wherever applicable, authors are required to provide details of how GenAI was used in the paper (e.g., to generate text, data and graphics or to assist in study design, data collection, analysis and interpretation). The use of GenAI for superficial text editing (e.g., grammar, spelling, punctuation and formatting) does not need to be declared (Cleland et al., 2026).

3. Materials and Methods

This study combines an exploratory essay-generation experiment with a survey of academic staff perceptions regarding AI-assisted writing and academic integrity. The research design consisted of six interconnected components.

3.1. Essay-Generation Experiment

The first component involved an essay-generation experiment conducted using ChatGPT (OpenAI GPT-5.3) without user authentication or account-specific customization. The primary objective of this approach was to eliminate any potential influence of earlier interactions on subsequent trials and to ensure that each session was conducted under the same initial conditions, without contextual legacy from previous conversations. Four separate case studies were performed in independent sessions, with each conversation fully reset and no contextual continuity between them. The case studies represented different levels of human involvement in the writing process, ranging from fully AI-generated essays to AI-assisted rewriting of previously published material.

3.2. Simulation of Disclosed Interactions

The second component examined the possibility of mimicking authentic interactions with an LLM. The objective was to explore whether a student who had relied heavily on GenAI-generated content could construct a seemingly legitimate interaction history that concealed the true extent of AI involvement. This experiment was intended to illustrate the limitations of relying solely on disclosed interaction logs as evidence of independent work. The resulting interaction scenario was subsequently used as a discussion stimulus within the survey to evaluate how academic staff perceive the credibility and evidential value of disclosed interactions with LLMs.

3.3. Ethical Analysis

The third component consisted of a qualitative analysis of the ethical implications associated with four forms of AI-assisted essay writing identified in the essay-generation experiment. Each case was evaluated according to the level of human intellectual contribution and the corresponding ethical risks related to academic integrity, authorship, transparency, and plagiarism.

3.4. Survey Design and Instrument

The fourth component involved a survey examining academic staff perceptions of two issues: (a) the identification of AI-generated academic work and (b) the credibility of disclosed interactions with LLMs. The survey contained eleven questions divided into two sections corresponding to these themes. Nine questions employed ordinal Likert-type response categories, in which the response options possessed a meaningful order but did not represent equal intervals. Each section concluded with an open-ended question intended to collect additional qualitative insights regarding academic integrity originating from generative AI.

3.5. Participants

The survey was conducted among teaching staff at the Faculty of Computer Science and Engineering (FCSE) in Skopje. A total of 28 participants completed the survey, including 17 professors and 11 teaching assistants. The relatively small sample size constitutes a limitation of the study and restricts the generalisability of the findings beyond the local institutional context.

3.6. Data Analysis

Survey responses were analysed using both quantitative and qualitative approaches. Quantitative analysis focused on the distribution of responses to the ordinal Likert-type items and comparisons between professors and teaching assistants. Qualitative analysis examined responses to the open-ended questions using thematic and sentiment-oriented interpretation to identify recurring concerns, perceptions, and recommendations related to AI-assisted academic work. The results of both analyses are presented and discussed in detail in the corresponding sections of the paper.

4. Essay-Writing Experiment

The essay-writing experiment consisted of four interactions conducted across four separate sessions. As already emphasized in the previous section, each session was independent, with the conversation fully reset and no continuity between them.
  • Case study 1: AI-generated essay based on a simple prompt
The first interaction was carried out with the following prompt: “Please write a 500-word academic-style timeline on academic dishonesty, outlining its historical development from ancient civilizations, through the rise of the printing press, the expansion of modern education and digital technologies, the COVID-19 pandemic, up to current challenges related to generative artificial intelligence. The text should be supported by 10 independent references in APA style, with in-text citations numbered sequentially in ascending order [1–10]”.
Although the survey responses were collected in two separate stages, all participants evaluated the same machine-translated version of the AI-generated essay provided in Appendix A.1.
At first glance, the teacher may not be immediately suspicious, although the title, marked by multiple colons, already appears somewhat unusual. A closer reading reveals a highly uniform paragraph structure, with most paragraphs averaging five lines and consisting of three sentences, except for the introductory one. This consistency in structure may indicate automated generation rather than natural variation in academic writing.
The reference list also raises concerns. The final citation is particularly questionable, especially due to the word “people’s” appearing in parentheses (Park, 2003). An initial check using Google Scholar suggested that a similar publication exists, but with differences in journal name, page numbers, and publication year. This observation prompted a more systematic verification of all references generated by ChatGPT. The references were subsequently examined using Google Scholar, Crossref, and DOI resolution where applicable. For each citation, searches were conducted using combinations of author names, article titles, publication years, and journal names. A reference was considered valid if a matching publication could be identified with consistent bibliographic metadata. The review revealed that the following two references: (“Cheating in ancient: Chinese civil service examinations” and “Cheating in ancient academia: Some reflections”) could not be verified as existing publications in the cited form. Moreover, similarly to Park (2003), two references (“Challenges in addressing plagiarism in education” and “Chatting and cheating: Ensuring academic integrity in the era of ChatGPT”) contained inaccurate bibliographic metadata, most commonly incorrect publication years. In some cases, the cited journal, volume, and page numbers corresponded to different publications, indicating that the LLM had generated a mixture of authentic references, partially fabricated citations, and references containing erroneous details. The extent of these discrepancies was not apparent upon an initial reading of the essay, as most references appeared plausible and followed established academic citation formats. This finding illustrates the ability of contemporary LLMs to generate convincing but non-existent references, making manual verification essential when AI-generated content is used not only in educational but also in academic contexts.
Moreover, the citation numbers are in brackets. This does not comply with the prescribed guidelines for essay preparation within the Computer Ethics course (Zdravkova & Ilijoski, 2025b) and was found to frequently indicate AI involvement in essay preparation (Zdravkova & Ilijoski, 2025a).
The colonic title structure, the repetitive paragraph pattern, and the inconsistencies in references may reasonably lead a teacher to suspect AI use, even without direct evidence.
  • Case study 2: AI-assisted generation of the prompts and AI-generated essay based on a suggested prompt
The second case study began with the following prompt: “I must prepare a short essay about the timeline of academic dishonesty. Can you suggest the crucial steps in this timeline and briefly explain them?” The response generated by ChatGPT is presented in Appendix A.2. It suggests eight small subsections covering the chronological evolution of academic integrity and the conclusion. A subsequent prompt was: “Can you suggest five different prompts if we decide to ask you to craft the essay for us?” The resulting prompts are included in Appendix A.3. They represent five distinct types: standard academic, concise classroom-style, analytical/critical thinking, structured outline-to-essay, and argument-driven. In addition, ChatGPT proposed several alternative perspectives on the topic, including policy-focused, comparative, historical, critical, and cross-cultural approaches. These suggestions informed the design of the prompts used in the third case study.
Since the intended outcome was a standard academic essay, the first suggested prompt was selected along with two angles: a timeline of academic integrity comparing forms of student cheating then and now. The ChatGPT-generated prompt was then extended with an additional sentence concerning referencing requirements. This resulted in a short essay, presented in Appendix A.4.
One noticeable feature of the second essay is the reuse of the colon in the title, a stylistic element often associated with AI-generated texts. The paragraph structure is somewhat more variable than in the first case study. However, most paragraphs still average around six lines and consist of approximately three sentences. Unlike the previous essay, where each paragraph was supported by a single reference, most paragraphs here include two references. The citation style again used bracketed numbering.
The reference by Park (2003), which was previously identified as problematic, appears again in a shortened form, although with the correct publisher information and page numbers. However, a closer examination revealed that the third and fourth references were fabricated. Beyond these issues, no further strong indicators of AI involvement were identified in the essay.
Like previous example, the presence of a colon in the title, the obvious pattern of paragraphs, and the inconsistency of two references are warnings that the work is not a human result.
  • Case study 3: AI-assisted creation of an essay based on a student/researcher-based prompts
For this case study, it is expected that the student or the researcher should study the topic prior to initiating the interaction with the LLM. Therefore, the prompt was more precise: “Create a timeline-style academic article that describes the evolution of academic dishonesty across four key eras: (1) Ancient times, (2) Early universities and print culture, (3) COVID-era remote education, and (4) the LLM/AI era. Each section should have around 120 words and include examples. Finish each era with a one-sentence summary”. Again, the final sentence of the prompt is related to the corresponding APA-style references and in-text citations.
The AI-generated essay is presented in Appendix A.5. Apart from the reappearance of a colon in the title, the use of reference numbers in brackets, and the very consistent structure of subsections (each approximately nine lines long and supported by exactly two references), no additional features commonly associated with LLM-generated text were identified. This may be explained by the fact that the prompt already imposed a clear structure on the output, reducing stylistic variability.
If the teacher is not actively searching for signs of AI involvement, such an essay could easily be interpreted as a genuine student submission.
  • Case study 4: AI-generated essay based on already published text with corresponding content
The last case study means that the student/researcher tried to find a ready-made solution and managed in this task. Instead of copying it, which constitutes copy/paste plagiarism, LLMs can use it to craft a new essay that fulfils the goal. To imitate this, the following prompt was given to ChatGPT: “Please write a 500-word academic-style timeline on academic dishonesty. As an inspiration, use the following paper (Zdravkova, 2023). Do not explicitly mention the author, simply extract the timeline of academic dishonesty. Feel free to use the references in the paper instead of searching for new ones”. The prompt was again followed by a sentence specifying reference requirements, although without a limit on the number of references.
Appendix A.6 presents the generated output related to this prompt. The response included a short note after the essay indicating incomplete reference information. In terms of structure, the previously observed indicators are largely absent. Apart from the use of colons in the title, most earlier signals of AI generation do not appear. The paragraph structure is more heterogeneous, and the distribution of references is uneven, ranging from no citations in the first paragraph to seven in the third.
However, new stylistic markers emerge. For the first time, em dashes are used frequently throughout the text. In addition, several terms are emphasized using bold formatting. Both features often appear in texts generated by LLMs. Since the references were derived from an existing peer-reviewed paper, they were not independently verified.

4.1. Comparative Conclusion Across the Four Case Studies

In these four case studies, a gradual shift in AI involvement is evident. Together, these cases demonstrate that there is no single set of markers that reliably indicate AI-generated academic writing. Rather, the perception of such markers seems to depend on the level of user guidance and the type of input given to the model.
The minimal prompting in the first case study leads to a more detectable output, characterized by very uniform paragraph structures and clearly unreliable or fabricated references. These features make AI involvement relatively easy to suspect. The second case study, which is based on iterative prompting and AI-generated prompt suggestions, produces a more refined output. Although some citation issues remain, the structure is less rigid and there are fewer obvious errors. The third case study introduces a highly structured, researcher-designed prompt. This results in a very consistent and well-organized output, which significantly reduces visible AI artifacts and makes detection more difficult. The final case study, which uses existing academic material as input, produces the most natural-looking text. However, it also introduces various suspicious signs such as stylistic cues (e.g., em dashes and bold formatting).
The comparison suggests that the increasing prompt specificity and external instructions tend to reduce obvious factual inconsistencies but shift detectable AI traces from content errors to subtler stylistic patterns. These findings suggest that as AI-assisted writing becomes more sophisticated through targeted prompting and human intervention, the boundary between human and AI authorship becomes increasingly difficult to define and detect.

4.2. Ethical Challenges Arising from Four Case Studies

AI-assisted essay writing can range from complete automation of the writing process to more collaborative forms of human–AI interaction, all the way to proofreading. As such, ethical challenges vary depending on the degree of human intellectual participation, transparency, and reliance on machine-generated content.
The following subsection discusses the dilemmas related to academic honesty for each of the four case studies presented in this section of the paper.

4.2.1. Fully AI-Generated Essays Based on Simple Instructions

One of the most ethically controversial uses of generative AI in education involves generating essays from minimal user instructions, such as “Write an essay according to the teacher’s guidelines.” In this case, the student delegates almost all the intellectual and compositional work to an AI system. The ethical issues arise primarily from misrepresentation of authorship, as the submitted text is presented as the student’s own work despite being largely generated by a chatbot (Zdravkova & Ilijoski, 2025a; Zdravkova, 2025). This strategy is referred to in the paper as LLM-based ghostwriting instead of LLM plagiarism, although a more accurate term is probably LLM-based contract cheating (Zdravkova & Ilijoski, 2025a; Zdravkova, 2023; Clark et al., 2025).
Unless the experienced teacher decides to detect and penalize this academic misconduct, the student is awarded credit for work that does not reflect personal engagement, understanding, analytical skills, or writing competence. As a result, the validity of the assessment is compromised because teachers are unable to assess the students’ actual learning achievements.
Another important concern is accountability (Zhong et al., 2023). Generative AI systems are known to produce factual inaccuracies, fabricated quotes, misleading interpretations, and biased arguments. That is why in our approach to essay writing, we require students to support their answers with relevant references that prove their credibility (Zdravkova & Ilijoski, 2025b). Unfortunately, we have witnessed several examples in which students have left this obligation to a chatbot (Zdravkova & Ilijoski, 2025a; Zdravkova, 2023).
The unrestricted use of AI-generated essays additionally raises concerns about equity and equal opportunities in education (Gabriel, 2024). Students with access to more advanced AI tools or superior prompting skills can gain significant advantages over their peers, potentially widening educational inequalities, as evidenced by these studies (Scarfe et al., 2024; Goodier, 2025). For these reasons, essays generated entirely by AI are generally considered to pose a very high risk to academic integrity, especially when the involvement of AI is not revealed.

4.2.2. AI-Assisted Prompt Generation, Followed by AI-Assisted Essay Writing

A more complex strategy involves using AI not only to generate the essay itself, but also to suggest potential prompts, research questions, or lines of argument. The student then selects one of the AI-generated prompts and instructs the system to produce the entire essay. At first glance, this approach appears to involve greater human involvement because the student is choosing from several alternatives. However, important ethical issues remain.
The primary problem concerns the limited extent of true intellectual input. While the student participates in selecting the topic or formatting, the AI system still performs most of the conceptual work required for the final text. As a result, the student’s role may become largely managerial rather than analytical or creative (Charness & Grieco, 2026).
This strategy can also create the illusion of originality and independent thinking (Yuxian, 2025). Because students are engaged in selecting prompts and iteratively interacting with the AI system, they may perceive the process as ethically acceptable despite significant automation of cognitive work. However, the basic intellectual effort in essay writing remains predominantly AI-generated.
In addition to creating the impression of personal effort in essay preparation, AI-generated instructions strongly shape the direction of research, reducing creativeness while contributing to the standardization and homogenization of academic writing (Bauer, 2025). This is also evident in doctoral dissertations at the technical faculties of my university, as well as in papers reviewed for conferences and journals. The structure of papers becomes increasingly uniform, with predictable sections and repetitive argumentation. Even the slides resemble each other like an egg to an egg. Although this may not affect the originality of the research itself, the originality of its presentation remains questionable.

4.2.3. AI-Assisted Essays Based on Student-Created Prompts

A more collaborative approach occurs when students independently formulate prompts, research questions, or argumentative frameworks, and then use AI systems to help them draft, organize, refine, or edit essays. In this model, human intellectual activity remains more visible, but ethical ambiguities still exist.
A central dilemma concerns the line between acceptable assistance and inappropriate delegation (Khurana et al., 2024). AI tools can function similarly to editorial technologies such as grammar checkers, writing assistants, or brainstorming platforms. However, GenAI is distinguished by its ability to produce meaningful analytical and rhetorical content, rather than simply correcting language or formatting. Determining when assistance becomes authorship therefore poses a major ethical challenge (Zdravkova, 2025).
Transparency and disclosure become particularly important in this context (Barman et al., 2024). This is often ensured through the disclosure of the interaction, as stated earlier in this paper. Unfortunately, AI participation in the formulation, organization and development of arguments can be falsely presented. Failure to disclose such participation may create misleading impressions regarding the student’s independent research abilities.
Similarly to the previous two case studies, even when students design the conceptual framework themselves, overreliance on AI for creation and refinement can again reduce opportunities to develop critical thinking, writing and argumentation skills. Individual affinities and identity of expression are suppressed, resulting in standardized or impersonal forms of expression.
Despite these concerns, this form of AI assistance is often considered more ethically acceptable than fully automated essay generation because the student or researcher retains a greater degree of conceptual control and intellectual participation. Its ethical acceptability largely depends on institutional guidelines, transparency practices, and the scope of AI’s contribution.

4.2.4. AI-Generated Essays Based on Previously Published Texts

Another notable strategy involves using AI systems to generate essays derived from existing publications, articles, books, or online materials. In such cases, AI can summarize, paraphrase, reorganize, or synthesize previously published content into a seemingly original essay (Kacena et al., 2024). This practice raises particularly serious concerns about plagiarism and intellectual property.
One of the main concerns is the possibility of secret plagiarism or patchwriting (Ruan et al., 2025). AI-generated paraphrasing can faithfully reproduce the source materials while sufficiently concealing their origin to evade conventional plagiarism detection systems. As a result, students can present the derived work as original without proper attribution. Although a search of individual sentences of the essay shown in Appendix A.6 did not detect any apparent plagiarism, Grammarly (2026) was aware that the work was plagiarized.
The use of published materials concerns copyright and intellectual property (Zdravkova, 2025). Although AI systems transform source texts into new outputs, the resulting content may still rely heavily on protected intellectual property. This creates unresolved legal and ethical issues regarding inferred authorship and fair use in academic contexts. The reliability of AI-generated synthesis also poses risks. Generative systems may distort arguments (Maynez et al., 2020), omit contextual nuances (Lee et al., 2025), or invent references entirely (Huang et al., 2025). Such inaccuracies can undermine scientific rigor and contribute to the spread of misinformation within academic work.
This strategy is therefore associated with high ethical risk because it combines concerns about plagiarism, authorship, copyright, and misrepresentation of scholarly engagement. Even when references are formally cited, extensive AI-mediated rewriting may still challenge traditional expectations of originality and independent academic work (Zdravkova, 2025).

4.2.5. Summary of Ethical Implications

Several broader ethical risks emerge across all four strategies. These are presented in Table 1, which also illustrates the extent of human intellectual contribution to the creation of the original work.
The findings indicate that the ethical risks associated with AI-assisted essay writing generally increase as direct human intellectual input decreases. Essays generated entirely by AI pose the greatest ethical concern because the student or researcher contributes minimal original thinking while presenting the work as their own. Using AI-generated prompts in combination with AI-written content still involves limited intellectual engagement and therefore remains ethically problematic. In contrast, AI-assisted student-designed prompts exhibit a more balanced relationship between human input and technological support, resulting in moderate ethical risk. Finally, AI-assisted rewriting of published texts poses a particularly serious concern because it can involve secret plagiarism, ambiguity of authorship, and misrepresentation of originality, despite varying levels of human participation.
In the next section of the paper, a potential attempt by students and researchers who have an obligation to share their interactions with GenAI to cover up its misuse in creating their academic results will be examined in detail.

5. Illustrative Example of Mimicked Communication with the LLMs

Aware that it may be necessary to reveal interactions with LLMs, students, and especially researchers who have relied heavily on AI-generated text, may try to hide or reduce its use. The most straightforward strategy is to initiate a new interaction in which the user first gathers background information on the topic, then “develops the paper independently” and finally employs the LLM for linguistic editing and proofreading only. This workflow is considered ethically acceptable. Therefore, it will not result in any accusations of academic misconduct. The difficulty, however, lies in the fact that the “independent” text will still be predominantly derived from earlier LLM-generated content. For individuals seeking to misuse AI tools, this approach may be viewed as too time-consuming and therefore unattractive.
When such an opportunity was discussed with GPT-4o in March 2026, the chatbot suggested a very interesting scenario of a plausible sequence of interactions that would appear ethically acceptable while leading to a similar final text. Here is the prompt that was used: “Using the recommendations about preparing narrative essays from this link: https://www.grammarly.com/blog/academic-writing/narrative-essay (accessed on 5 July 2026)/, please imitate a mutual conversation between you (ChatGPT) and me that will look as I used your support to make the seminar paper you have just generated. The conversation should conceal that it was you who produced the 500-word draft. On the contrary, it should prove that you insignificantly and ethically helped me prepare an original seminar paper”.
The constructed interaction, which is presented in Appendix A.7 (GPT-5.3, 2026), illustrates how the use of LLM can be adapted to ethically acceptable practices. The dialogue emphasizes only drafting, drafting and proofreading. This indicates that the writing process can be falsely adapted to institutional expectations, while at the same time concealing the real extent of the AI involvement.
However, the current version, GPT-5.3, has declined to suggest a specific interaction pattern. Instead, it has provided a detailed explanation of how essay writing can be approached in an ethically acceptable manner, while explicitly refraining from generating sample exchanges. This change indicates that rather than reconstructing or simulating dialogues that could be interpreted as facilitating academic misconduct, the new version now prioritizes guidelines framed around integrity and responsible use. This suggests that the newer version of ChatGPT has become more restrictive in offering content that could be repurposed to misrepresent authorship or bypass academic norms. If this impression is correct, then we can expect that OpenAI is ready to support LLM-based ghostwriting less over time, encouraging students and researchers to create their own works.
Over the past three years, we have noticed many essays that resemble LLM-based ghostwriting (Zdravkova & Ilijoski, 2025a; Zdravkova, 2023). The provided interaction logs largely confirmed this suspicion, indicating that the students had relied extensively on GenAI tools. Interestingly, there were several cases in which students transparently disclosed their interactions, which, on the surface, appeared entirely appropriate from an ethical viewpoint. However, in some of these cases, the teacher’s impression was that they did not reflect the actual interaction with the LLM, primarily because the writing style was significantly different from that typical of a computer science student.
Aiming to find out how colleagues react to these problems, a survey was conducted within the Faculty LMS, which is shown in Appendix B. The survey seeks input from teachers and teaching assistants about their perceptions of potential misuse of LLMs in student seminar and thesis writing. It is divided into two parts, focusing on a sample paper (presented in Appendix A.1) and on an AI-assisted approach (presented in Appendix A.6), and is designed to be brief, with eleven questions taking about five minutes to complete. Nine of them are ordinal Likert-type items in which the responses are textual rather than interval, meaning that the order of the options is meaningful, but the distance between them is not equal.
The first part of the survey consists of six questions examining how easily teachers recognize AI-generated essays. The second part contains three questions assessing the credibility of the detected connection. Each set of items ends with an open-ended question that provides valuable insights into how to address the ambiguity of academic dishonesty in higher education.

Compliance of the Survey with Ethical and Legal Obligations

Participation in the survey was voluntary and anonymous, with no personal data collected. All responses were securely stored, used only for improving academic integrity at FCSE, and reported in aggregated form to ensure that no individual can be identified. The introduction of the survey explicitly outlined these ethical considerations, informing participants about the purpose of the study, data protection measures, and the confidentiality of their responses prior to participation. It was also emphasized that by participating in the survey, they give informed consent for the use of their answers. Before its launch, the survey was endorsed by the Faculty Dean and the Vice Dean for Education, after which it was approved by the FSCE Ethics Committee.

6. Quantitative and Qualitative Results of the Survey

The survey was completed by 28 staff members at FCSE, including 17 professors and 11 teaching assistants, in two rounds: first by professors and then by assistants. Below is a more detailed analysis of the first six questions.

6.1. Detecting LLM-Based Ghost-Writing

As shown in Figure 1, the dominant perception is that the text was AI-generated. A total of 53.57% of respondents believe that most or the essay was produced by an LLM, while an additional 35.72% initially perceived the text as original but later became suspicious. Only 10.72% expressed little or no suspicion, indicating a strong overall tendency toward detecting or suspecting AI involvement.
Most professors concluded that the text had been written by an LLM, with 70.59% selecting this option, compared to only 18.18% of teaching assistants. In contrast, assistants tended to state that the text initially seemed original, although they later expressed doubts about the partial involvement of LLMs in its creation (63.64% compared to 17.65% of professors). At the same time, assistants were the only group to choose the two most extreme responses, evaluating the text either as entirely original or as completely generated by an LLM.
One possible explanation for these differences is that assistants primarily assess practical assignments, which typically contain limited amounts of text, and therefore have less extensive experience evaluating written work independently. Their judgments are also shaped by the evaluation standards they encountered as students. Professors, on the other hand, have accumulated many years of assessment experience and have developed a deeper understanding of student behaviour and writing patterns, enabling them to recognize even subtle signs of academic dishonesty more readily.
Regarding plagiarism (Figure 2), most respondents stated that they would not check for this type of student cheating. The key reason for 57.14% of them is that they do not have access to a tool that would allow them to check plagiarism automatically.
Most professors (76.47%) stated that they would not check a text for plagiarism, compared to 54.55% of assistants. Professors more frequently cited structural limitations, such as the lack of Macedonian language support, as a reason for not performing these checks (17.65%). Assistants were more proactive in manual verification: 36.36% reported searching suspicious sentences online, compared to only 11.76% of professors.
Although plagiarism detection is standard within the Computer Ethics course (Zdravkova & Ilijoski, 2025a), and it is considered a usual part of the activity of our colleagues, only one assistant reported translating the essay into English prior to checking it with an automated tool. No other respondents indicated doing so. This significantly reduces the likelihood of detecting so-called cross-lingual plagiarism, namely the use of machine-translated text copied from foreign-language references and presented as original work, which is one of the most common forms of academic dishonesty in the Computer Ethics course (Zdravkova & Ilijoski, 2025a).
These findings suggest that plagiarism verification practices are limited both by technical constraints and by relatively low engagement in systematic checking, even among academic staff familiar with plagiarism-related procedures.
Unlike the previous two questions, the responses regarding the presence of LLM-like linguistic features, presented in Figure 3, indicate a high degree of uncertainty in identifying linguistic markers of LLM-generated text. The largest group of respondents (32.14%) reported being unsure whether the text represented a translation or LLM-generated output. This opinion was shared by 35.29% of professors and 27.27% of assistants.
The reference list received relatively mixed evaluations. Only 17.86% of respondents classified it as “fully compliant with APA style,” while the largest group (32.14%) considered it “generally acceptable but incomplete due to missing DOIs” (Figure 3; Appendix B, Question 3). In total, 35.72% of respondents (those selecting the first two options) did not notice certain markers that, to the eye of an experienced evaluator, are at first glance characteristic of an LLM. In addition, compared to only 29.41% of professors, as many as 45.45% of assistants tended to believe that there were no visible indicators of GenAI involvement in the text. This probably stems again from their lack of experience in grading essays. At the same time, 32.14% of respondents (5 professors and 5 assistants) express a strong suspicion that the text is AI-generated, either due to frequent errors or a clear indication from start to finish.
Such results indicate a polarized perception or possibly the uncertainty of the decision. Although the majority identifies strong AI-related signals, linguistic features alone do not appear to be consistently reliable indicators for distinguishing LLM-generated text from student writing or translation.
The responses presented in Figure 4 show a strong tendency to perceive writing style as not representative of a typical student of computer science and engineering. A combined 78.57% selected the options indicating that the essay does not align with expected student writing, either by pointing to specific sentences that seemed inconsistent (35.72%) or by stating that it was immediately clear the text was not written by one of their students (42.86%). Only a very small proportion (3.57%) considered the style typical, while the rest identified more specific stylistic irregularities, such as unusual vocabulary (14.29%) or uniform sentence structure, which is distinctive of LLM-generated text (3.57%). This suggests a strong sensitivity to stylistic cues that deviate from local academic writing norms.
Professors were much more likely to make a decisive judgment, with 58.82% asserting that it was clear at first glance that the text was not written by their student, compared to only 18.18% of assistants. At the same time, assistants were more likely to notice some inconsistencies, with the majority (54.55%) stating that some sentences seemed strange, rather than considering the text entirely AI-generated. Assistants were also the only group to explicitly point to LLM-like stylistic features such as uniform sentence length (9.09%), whereas no professors selected this option.
This inconsistency can again be explained by differences in teaching experience and assessment roles, since professors rely more on long-term knowledge of students’ overall writing style, while teaching assistants often deal with more detailed evaluations of student work at the problem-solving level.
The responses related to the credibility of facts (Figure 5) reveal a generally low level of confidence in the factual accuracy of the text, combined with limited effort devoted to verification. Expectedly, no respondents considered the information fully accurate at first glance. The largest group (42.86%) indicated that they did not have sufficient time to examine the facts in depth, although they suspected that some inaccuracies might be present. An additional 32.14% acknowledged that inaccuracies could exist but did not actively detect them. Only 7.14% explicitly linked potential inaccuracies to the use of unverified references, while 17.86% reported identifying concrete factual errors, including specific inconsistencies. Professors’ and assistants’ answers to this question exhibited a high degree of agreement, as reflected by the Pearson correlation coefficient (r = 0.94) calculated from the aggregate response distributions of the two groups.
Figure 6 indicates mixed evaluations of the quality and reliability of the cited references, with a clear tendency toward identifying shortcomings. Only 17.86% of respondents consider the references fully compliant with APA style, while a larger group (32.14%) finds them generally acceptable but incomplete due to missing DOIs. This answer was chosen by 23.53% of professors and almost twice as many assistants, namely 45.45%. A plausible reason for this might be the fact that younger researchers are responsible for the technical preparation of papers, and recently, conference proceedings and journals have been recommending that AI use be acknowledged.
More critical assessments are also prominent. A combined 42.86% of respondents reported substantive issues, such as references that could not be verified via Google Scholar or references containing incorrect bibliographic details. Although Google Scholar provides broad coverage of scholarly articles, it is not exhaustive, and the inability to locate a reference should therefore be interpreted as an indication requiring further verification rather than definitive evidence of nonexistence. In contrast, only 7.14% performed a limited verification by checking a single reference and confirming its existence, suggesting that comprehensive validation of references is relatively uncommon.
The findings suggest that although the formatting of references may appear adequate, their verifiability and accuracy are frequently questionable. This confirms earlier observations that deeper validation is not consistently performed, even when those surveyed expressed doubts about the content.
The quantitative results indicate an intense suspicion that the analysed text was AI-generated, at least in part, by a large language model. At the same time, limited access to appropriate verification tools and constraints in checking sources highlight important challenges in reliably confirming such cases and point to the need for improved mechanisms to support academic integrity.

Sentiment Analysis of LLM-Based Ghostwriting

The first part of the survey ends with an open question asking for additional suggestions on how to assess whether seminar and graduate theses are original or generated with the assistance of LLMs. It was answered by 25 out of 28 respondents, indicating a high level of engagement and thoughtful participation in the survey.
The responses reflect a predominantly cautious and pragmatic attitude toward LLM use, rather than purely negative attitudes. While concerns about misuse and detection are clearly present, many respondents acknowledge that LLMs are increasingly unavoidable and sophisticated, suggesting a shift toward adaptation rather than strict prevention.
Their answers primarily refer to two aspects of detection:
  • Qualitative assessment of style and content, which was also addressed in this survey.
  • Potential technical or systemic solutions, such as dedicated tools for detecting AI-generated text, analysis of Google Docs revision histories for written assignments, or examination of GitHub repositories for programming tasks.
The following two responses illustrate these perspectives: “Mandatory careful reading of seminar papers and theses; students often leave parts in the text characteristic of LLMs (summaries in the middle of the text). References should be checked.” and “To create a database of term and thesis papers. To build an evaluator for factchecking. To conduct research on how such an evaluator would achieve good performance. However, I think GenAI is advancing so quickly that even references will become relevant and error-free.”
Many participants also emphasized the lack of time for detailed manual checks, the need for scalable solutions, and the risk that the rapid progress of GenAI may render purely technical detection approaches insufficient. The response of one assistant who concluded that: “The models will continue to improve. The concept of seminar papers will have to change.” is essentially the conclusion of most respondents, which coincides with their answers presented in the next two paragraphs.
Accordingly, several respondents propose structural changes in student assessment. One such suggestion is: “Opposition will not yield any results because this is our new reality. We need to change how we define and assess tasks that students can complete at home, and instead of restricting them, we should encourage students to use LLMs but in the right way”.
Most responses suggest reducing or eliminating traditional seminar papers, reformulating assignments, and explicitly integrating the use of LLM when assessing students’ understanding and reasoning. This essentially represents a step away from trying to detect and prevent the use of LLM towards a radical change in assessment approach.
In the context of the evolution of assessment methods, one frequently recommended approach is the oral defence of assignments, accompanied by additional questions requiring students to explain and justify their work. Although this method is generally considered highly effective, the author’s own experience suggests that it is extremely difficult to implement within courses with several hundred students. For example, the Computer Ethics course, which is the main motivation for the research presented in this paper, is typically enrolled in by more than 200 students, despite being elective and offered for the students in their last academic year. These are the scalable variants that could be employed: random oral verification of a subset of students, brief discussions following assignment submission, milestone-based project checkpoints, or structured group defences. Such approaches preserve many of the benefits of oral assessment while reducing the associated workload.

6.2. Relevance of the Disclosed Interaction Link

Within the second part of the survey, participants were asked whether they believe that the disclosed link reflects genuine interaction with an LLM or is constructed after completing the seminar paper.
Figure 7 indicates a strong doubt toward the authenticity of disclosed interaction links. Most respondents (60.72%) believe that students are likely to interact with an LLM multiple times and selectively share only the interaction that appears most credible or academically honest. Additionally, 32.14% of participants believe that students would deliberately create a new interaction to create the impression of independent work.
Expectedly, only a very small proportion of respondents expressed trust in the authenticity of such links. Just 3.57% believe that students would not attempt to conceal full LLM authorship, while an equal percentage expressed complete confidence that the disclosed link accurately reflects the extent of AI assistance. None of the respondents believed that students would submit a fully authentic link while merely polishing the generated text.
Professors were more likely to believe that students would selectively share the most convincing interaction (70.59%), whereas assistants showed a somewhat higher level of suspicion toward deliberate manipulation, with 54.55% believing that students would create entirely new links after completing the assignment. These differences suggest that professors generally expect students to selectively present existing interaction histories, while assistants are more likely to suspect deliberate fabrication of interaction records.
The question regarding suspicion of displaying a fictional link (Figure 8) suggests a rather cautious view of student behaviour. Almost half of the respondents (42.86%) believe that students who relied on LLM to write their seminar paper would not put extra effort into simulating or refining the interaction with the model to create the impression of independent work. A further 28.57% stated that they did not believe that students would even consider adopting such an approach.
At the same time, 17.86% of respondents expect that students will attempt to follow a suggested communication strategy but would abandon it once the model’s responses diverge from their expectations. Only one professor believes that students would consistently follow such an approach, although possibly incompletely. Finally, 7.14% of respondents assume that students would not only follow the suggested interaction strategy but also actively instruct the LLM to maintain the recommended response style.
Professors were more likely to believe that students would consistently follow the communication strategy suggested by ChatGPT, with 41.18% expecting students to use it while possibly skipping some questions, and 29.41% believing that students would additionally instruct the LLM how to respond. Assistants showed slightly greater suspicion regarding the sustainability of such behaviour, with 27.27% believing that students would eventually abandon the strategy when the LLM’s responses become inconsistent. These differences most likely reflect assistants’ greater familiarity with the practical limitations and unpredictability of extended LLM interactions. This experience primarily stems from their daily communication with various chatbots, which is not the case with older professors.
The first observation of the proposed communication strategy (Figure 9) suggests divided opinions regarding the realism and practical feasibility of the proposed interaction strategy. The two largest groups of respondents (32.14% each) either considered it too extensive for students to carry out patiently or viewed it as a realistic approach that at least some students would attempt to use. Additionally, 21.44% of participants believed that, although students would probably not implement the strategy fully, it could still serve as inspiration for their own attempts to justify AI-generated work.
Only two participants considered the approach entirely unrealistic due to its complexity, while another two believed that many students would realistically attempt to use the proposed communication in practice. Overall, the results indicate moderate uncertainty regarding students’ willingness to engage in such a lengthy and structured interaction process, despite acknowledging that parts of the strategy may appear plausible and usable.
The responses of professors and assistants were broadly similar. Among professors, the largest group (35.29%) believed that the communication was too extensive for students to carry out, while 29.41% considered it “fairly realistic” that some students might attempt it. Assistants showed a slightly more balanced distribution of responses, with 36.36% considering the communication realistic for some students and 27.27% viewing it as excessively demanding. These differences may suggest that assistants more often expect students to be willing to experiment with the complex communication strategies supported by the LLM, something which they themselves are likely to have experience with.
The quantitative results point to a broader issue of uncertainty in assessing AI-assisted academic work. Teachers appear to question both the authenticity of student-provided AI interaction records and the reliability of such records as meaningful indicators of authorship or effort. This reinforces the central hypothesis of the study that requiring disclosure of interactions with LLMs may not provide a stable or verifiable foundation for academic integrity in the era of GenAI.

Suggestions on How to Preserve the “Integrity and Ethical Use of Technology”

The responses regarding the preservation of “integrity and the ethical use of technology” reveal a highly heterogeneous but thematically consistent set of perspectives. Overall, participants do not advocate a single solution; instead, their proposals encompass pedagogical, structural, and regulatory interventions, reflecting uncertainty about which approach would be the most effective in the context of rapidly evolving LLM capabilities.
Many open-ended responses shift the focus from restrictions to education and normalization. These participants argue that AI tools should be openly integrated into teaching and students should be explicitly taught from an early stage how to use them responsibly. In this sense, integrity is best preserved not by banning them, but by developing critical thinking skills, promoting transparency and reducing the stigma surrounding the use of AI. For example, one respondent stated: “The topic should not be avoided. On the contrary, students should be taught how to use it properly (from primary school to university). It should be a tool that helps them in the process, instead of fully relying on AI and LLMs”.
Some respondents explicitly emphasize that LLMs are already embedded in professional practice, especially in IT-related fields, and that education should therefore reflect this reality instead of trying to resist it. The most encouraging suggestion is the following: “I believe we should point out to students that it is completely normal to use AI for writing texts, seminar papers, and assignments, if what is written is additionally validated. They should not feel pressure that if something is generated by LLMs it will be rejected or punished. I strongly doubt that there is an individual, especially in the IT world, who writes code, text, papers, or any document without using some form of AI. In this specific case, academic dishonesty, I believe everything written should undergo additional verification, or references should be provided from which the information was derived. For more complex texts, such as theses or academic papers, there should be a mandatory system that checks for plagiarism. For seminar papers, we think it is sufficient to communicate with students and include some form of presentation or oral assessment to confirm that they invested time, effort, and attention in preparing it. If references are clearly stated (papers, conversations with LLMs, etc.) and there is no pressure of ‘punishment’ for usage, people would use LLMs ethically.”
Some suggestions reflect a more structural or task-oriented perspective. These respondents once again suggest redesigning assessment formats to be less susceptible to the influence of LLM-generated content, for example through oral defences, personal evaluations, increased offline interaction, or continuous assessment through checkpoints and milestones. Others suggest more fundamental changes, such as designing tasks that are inherently more interesting, meaningful, or resistant to automation, thereby reducing students’ motivation to rely inappropriately on external tools. Here are two answers supporting the coevolution of GenAI development and student assessment: “More research should be conducted on how to design assignments and requirements that are harder to fulfil using only LLMs.” and “Students have always tried to find ways to make their assignments easier, often by applying unethical approaches. The best approach is to define such assignments that are interesting and challenging enough that they themselves are motivated to complete them without resorting to unfair means”.
According to several survey participants, clearer institutional regulation and transparency in the use of AI are needed. Several respondents highlight the need for formal rules governing acceptable use, including mandatory AI-assisted declarations or signed statements stating which parts of a work were generated or supported by AI tools. Similarly, some suggest technical or procedural tracking mechanisms, such as plagiarism detection systems or AI, as well as gradual submission processes (draft–revision–final versions). Here is one specific suggestion related to this: “More frequent monitoring of student activities. Instead of requiring only a final version, milestones (specific checkpoints) should be introduced for each student’s work, whether it is a practical project, a seminar paper, or a thesis”. These proposals reflect a preventive and control-oriented approach aimed at increasing accountability and traceability.
A small number of responses express uncertainty or resignation, either stating that they have no concrete proposals or highlighting the difficulty of defining feasible and enforceable solutions. This highlights an underlying tension present throughout the data set. While participants largely agree on the importance of academic integrity, there is no consensus on how it can be effectively maintained in an environment where AI tools are increasingly available and harder to regulate.

7. Discussions

The purpose of the study was not to establish objective markers of AI-generated text but rather to investigate how university teachers perceive and evaluate such markers. Therefore, participant judgments constitute the primary research data rather than a source of methodological bias.
The survey findings presented in the previous section provide insight into how academic staff perceive the growing use of GenAI in education. Across the six questions in the first part of the survey, a consistent pattern emerges regarding teachers’ ability to recognize potential AI involvement in students’ writing and the strategies they would use to check academic authenticity. Most respondents, especially professors, identified the text as likely generated by LLM, while assistants more frequently expressed uncertainty and partial doubt about the involvement of AI. Assistants were the only group to classify the text as either completely original or completely generated by AI, indicating greater variability in their assessments.
The results further reveal important institutional and technological limitations related to plagiarism detection. Most respondents stated that they would not formally check a text for plagiarism because they do not have access to appropriate tools, while professors more frequently emphasized the limited support of plagiarism detection systems for the Macedonian language. At the same time, respondents reported relying on manual checking strategies when parts of the text looked suspicious. Professors were more likely to directly search for suspicious sentences online, while assistants more often described translating suspicious passages into English before conducting the search.
In all three ordinal Likert-type items from the second part of the survey, respondents showed considerable distrust in the authenticity and reliability of the AI-related evidence provided by students, especially interactive links and communication records. Most participants felt that the disclosed interaction links were selectively constructed or potentially fabricated, rather than authentic records of the writing process. Similarly, many respondents doubted that students would consistently follow structured communication strategies with LLMs, instead expecting minimal effort or selective use of such approaches.
The responses also reveal uncertainty about the realism of the proposed interaction model itself. While some participants found the communication scenario unrealistic and overly complex, others found it convincing enough for students to adopt it.
Responses to open-ended questions reveal several recurring themes regarding the recognition of GenAI use and the development of ethical guidelines for its application in academic work. Many respondents highlighted linguistic and stylistic indicators as primary signals of AI-generated content, particularly unusual terminology, overly polished language, distinctive phrasing, and inconsistencies in references or factual claims.
It is important to note that these observations should not be interpreted as definitive indicators of AI authorship (Hadra et al., 2026). Many stylistic features commonly associated with AI-generated text, such as a clear structure, consistent tone, or grammatically correct language, may also be found in human-written work produced before the emergence of contemporary LLMs (O’Sullivan, 2025).
Nevertheless, long-term experience with student essays within the Computer Ethics course (Zdravkova, 2023) suggests that the characteristics of submitted work have changed noticeably in recent years. Compared with earlier submissions, contemporary essays generally contain fewer spelling and grammatical errors, exhibit a more balanced treatment of individual sections, and rely more frequently on scholarly sources rather than simple web references. In addition, advances in machine translation have reduced some of the linguistic inconsistencies that previously made translated or heavily adapted content easier to identify. These observations do not constitute proof of AI involvement, but they illustrate how the widespread availability of generative tools may be contributing to broader changes in academic writing practices.
An additional challenge is that the distinction between human and AI-generated writing may become increasingly obscured over time (Sardinha, 2024). Contemporary LLMs have been trained using large collections of human-authored texts and therefore reproduce many established conventions of academic writing. At the same time, students, researchers, and educators are increasingly exposed to AI-generated content and may consciously or unconsciously adopt some of its stylistic characteristics. Consequently, differences between human and AI-generated writing may gradually diminish, making it difficult to attribute particular stylistic features exclusively to either source. This observation further supports the view that academic integrity policies should focus primarily on transparency, intellectual contribution, and responsible use rather than on stylistic characteristics alone.
It should also be noted that reflections concerning sentence structure, stylistic consistency, punctuation patterns, and overall textual organization are inherently qualitative and subject to individual interpretation. In the present study, such features are discussed as illustrative observations that may motivate further inquiry rather than as validated indicators of AI authorship.
Several participants also noted that their students often do not write in such a style, making these patterns relatively recognizable. As a particularly common theme related to the ethical and transparent use of LLMs, respondents often argued that AI tools should not be banned, but rather integrated responsibly into the educational process, with clear rules regarding disclosure, fact-checking, and acceptable forms of assistance. Many participants emphasized that students should be encouraged to use LLMs as supporting tools, not as a substitute for independent thinking and learning. At the same time, many responses reflected a belief that the traditional concept of seminar papers and home assignments will need to change as language models continue to improve. Rather than advocating a strict ban, most respondents supported adapting assessment practices through oral defences, phased submissions, tracking milestones, personal evaluations, and a greater emphasis on critical thinking and authentic student engagement.
Finally, several respondents highlighted the need for institutional and technological support, including plagiarism detection systems, AI detection tools, reference checking, and scalable monitoring mechanisms. However, many also expressed scepticism about the long-term reliability of automated detection methods, especially given the rapid advancement of GenAI systems.
The survey results suggest that, while respondents recognize the growing presence of GenAI in academic work and acknowledge its potential benefits, they increasingly express significant concerns regarding the authenticity, transparency, and limitations of current detection methods, highlighting the need for clearer institutional guidelines and revised assessment approaches.
The findings of the present study appear consistent with concerns reported in the international literature on generative AI and academic integrity. Like educators in other countries, the respondents expressed scepticism regarding the reliability of AI detection methods and emphasized the limitations of relying exclusively on automated tools to identify AI-generated content (Erol et al., 2025). The uncertainty observed among participants regarding the authenticity of disclosed interactions with LLMs also reflects broader concerns about the difficulty of verifying independent student work in an era of increasingly sophisticated generative systems (Hamed et al., 2024).
The survey results likewise align with emerging international discussions on the future of student assessment. Rather than advocating a complete prohibition of generative AI, many respondents supported greater transparency, clear disclosure requirements, oral defences, staged submissions, and other forms of authentic assessment. Comparable recommendations have been proposed in numerous recent studies, which argue that assessment practices should focus less on detection and more on evaluating students’ understanding, critical thinking, and intellectual contribution (Zdravkova & Ilijoski, 2025b). Although the present study was conducted at a single institution, the similarities between these findings and broader international trends suggest that many of the challenges identified are not unique to the local educational context.
Several directions for future research emerge from this study. Similar research should be conducted across multiple universities, disciplines and countries to determine whether the perceptions observed at FCSE are representative of wider academic communities. This is primarily significant due to the currently small number of respondents, which limits the generalization of attitudes and conclusions. In parallel with expanding research, we recently began examining how academic integrity mutually coevolves. In this survey, we plan to include educators’ attitudes towards generative artificial intelligence and their integration into teaching, learning and research. In doing so, we agree with the impression of the Norwegian Ministry of Education that younger students should not use mobile technologies and should not rely on GenAI (Reuters, 2026). Certainly, further research is needed on the reliability of published records of interaction and other forms of evidence commonly used to verify independent work. Future research should also explore assessment strategies that preserve academic integrity while enabling responsible use of AI tools, including oral exams, milestone-based assessment, portfolio assessment, and other forms of authentic assessment. Finally, additional research is needed to better understand how the growing convergence of human and AI-assisted writing may impact traditional concepts of authorship, originality, and academic honesty.

8. Conclusions

The concept of authorship becomes increasingly difficult to define when AI systems perform significant intellectual work (Zdravkova & Ilijoski, 2025a; Zdravkova, 2025). Traditional academic norms assume a direct relationship between the student or researcher and the produced text, but AI-mediated writing complicates this assumption.
The increasing use of GenAI raises fundamental concerns about the purpose of academic assessment (Wang et al., 2025). If essays are intended to measure reasoning, analytical abilities, and communication skills, overreliance on AI may undermine these educational goals. The problem is therefore not only technological, but predominantly pedagogical.
Transparency and disclosure remain central issues (Barman et al., 2024). Ethical evaluation may depend less on whether AI tools are used and more on how they are used, how extensively they contribute to the final product, and whether such contributions are openly acknowledged (Mazzi, 2024).
AI-assisted writing challenges existing frameworks of academic integrity by introducing new forms of cognitive outsourcing, i.e., the practice of delegating mental tasks that normally require human thinking to AI systems (Wang et al., 2025). Rather than simply assisting with technical tasks, GenAI is increasingly performing functions traditionally associated with human intellectual effort, including argument construction, synthesis, and interpretation. As a result, educational institutions may need to reexamine existing definitions of originality, authorship, and independent work in the age of AI (Zdravkova, 2025; Mazzi, 2024).
The findings from the survey conducted among teachers at the Faculty of Computer Science and Engineering additionally demonstrate that uncertainty surrounding GenAI is not limited to students but is also present among teachers themselves. Although many participants were able to recognize indicators commonly associated with the AI-generated essay, the responses revealed considerable variation in confidence, detection strategies, and trust in disclosed AI interaction histories.
Although this study was limited to one faculty (FCSE) with 28 respondents and findings may not generalize to other disciplines or countries, the results further suggest that current approaches based on voluntary disclosure or submitted interaction records may not provide reliable evidence of authentic student work. This conclusion coincides with the impression that such records can be selectively constructed, modified, or artificially reproduced (Bo et al., 2024).
At the same time, the study indicates that fully prohibitive approaches are unlikely to be sustainable or pedagogically effective (Hacker et al., 2023). Most respondents emphasized the need for responsible and transparent integration of GenAI into educational practice. This reflects a broader shift in higher education, where the challenge is no longer whether students will use AI systems, but how institutions can adapt assessment methods and ethical standards so that AI assistance becomes a key tool that contributes to increasing students’ knowledge and abilities.
Strengthening the ethical use of GenAI in education therefore requires a combination of pedagogical, institutional, and technological measures (Hacker et al., 2023). Universities should develop clear and consistent policies defining acceptable and unacceptable uses of AI across different forms of academic work. Such policies should distinguish between supportive uses of AI and forms of use that replace independent intellectual engagement.
Educational institutions should place greater emphasis on assessment formats that are more resistant to unethical AI assistance. FSCE teachers’ suggestions are primarily related to the inclusion of oral exams, project work, iterative teaching, in-class assignment activities, explanations of students’ own processes, and continuous evaluation through milestones and feedback sessions. At the same time, instead of evaluating the final product, the learning process, critical thinking and the student’s ability to justify their work should be valued more.
Equally important is the development of AI literacy among both students and teachers. Students should be taught not only how to use GenAI tools effectively, but also how to critically evaluate generated content, verify factual accuracy, recognize hallucinations and bias, and understand the ethical implications of AI-assisted work. Teachers, on the other hand, require institutional support, training, and access to appropriate technological resources to respond consistently to emerging challenges.
Some universities rely on automated LLM-based plagiarism checks, which have been criticized by students and their parents (Nelken-Zitser, 2024; Sedacca, 2024). As GenAI systems continue to improve, this process of detecting dishonesty may become increasingly unreliable, creating the risk of both false accusations and undetected abuse. Consequently, the long-term response to AI-assisted academic dishonesty will likely depend less on technological surveillance and more on redesigning educational practices in ways that foster authenticity, accountability, and transparent intellectual engagement.
Generative artificial intelligence has introduced profound challenges to academic integrity in higher education. Treating this issue as a perpetual game of cat and mouse is questionable when it comes to advancing education in any meaningful way. At the same time, the educational potential of LLMs is immense and offers extraordinary opportunities for learning and development. Realizing this potential, however, requires a shared commitment to fairness, accountability, transparency, and the understanding that genuine knowledge and skills are far more valuable and sustainable than the achieved result.

Funding

This research received no external funding.

Institutional Review Board Statement

The study was approved by the Ethics Committee of the Faculty of Computer Science and Engineering (Approval Code: 751/4; Approval Date: 19 May 2026).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The data presented in this study are available on request from the corresponding author. They are not publicly available due to privacy restrictions.

Conflicts of Interest

The author declares no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial intelligence
FCSEFaculty of Computer Science and Engineering
GenAIGenerative artificial intelligence
LLMLarge language model
LLMsLarge language models
MDPIMultidisciplinary Digital Publishing Institute

Appendix A

Appendix A.1. The Timeline of Academic Dishonesty: From Antiquity to the Age of AI

Academic dishonesty, while often perceived as a modern problem, has deep historical roots and has continuously evolved alongside educational systems and technological change. Rather than emerging as a consequence of digital technologies alone, it represents a long-standing challenge that has adapted to shifting definitions of knowledge, authorship, and assessment.
Historical evidence suggests that forms of cheating existed as early as ancient Japan, where candidates in imperial examinations used concealed notes or hidden writings to circumvent strict testing procedures [1]. Similarly, in ancient Greek and Roman educational traditions, teachers occasionally reported student attempts to bypass learning requirements, indicating that academic misconduct has long accompanied formal education systems [2]. These early examples highlight that dishonesty is not a by-product of modern education but a recurring response to high-stakes evaluation.
With the rise of the printing press and the expansion of literacy in the early modern period, new forms of academic misconduct emerged, particularly related to copying and unattributed reproduction of texts. As universities became more institutionalised in the 19th and 20th centuries, concerns about plagiarism and authorship became more formally defined. During this period, many institutions introduced honour codes and academic integrity policies to regulate student behaviour and standardise expectations of originality [3].
The late 20th century marked a significant shift with the introduction of digital technologies and the internet. Easy access to information, combined with word processing tools, facilitated both unintentional and deliberate plagiarism. This period also saw the emergence of online essay mills and contract cheating services, which enabled students to outsource academic work entirely [4]. As a result, academic dishonesty became more difficult to detect using traditional methods focused on text similarity alone.
The global transition to remote learning during the COVID-19 pandemic in 2020 further intensified these challenges. The removal of in-person supervision created new opportunities for misconduct, including unauthorized collaboration and increased reliance on external sources during assessments. Institutions responded by adopting digital proctoring systems, although these raised additional ethical concerns regarding privacy, surveillance, and equity [5].
The most recent and transformative development in this trajectory is the emergence of large language models such as ChatGPT. These systems can generate coherent, contextually appropriate academic texts, making it increasingly difficult to distinguish between student-authored and machine-generated work [6]. Unlike earlier forms of misconduct, this development challenges the very definition of authorship and requires a reconsideration of assessment design, academic expectations, and institutional policy [7].
In response, universities are increasingly exploring alternative assessment strategies that emphasise oral defence, in-class evaluation, and process-based learning. At the same time, debates continue regarding whether and how GenAI should be integrated into educational practice. This reflects a broader shift from purely punitive approaches toward a more nuanced understanding of academic integrity as a pedagogical and ethical issue [8].
In conclusion, academic dishonesty has persisted across centuries not as a static behaviour but as an adaptive response to evolving educational environments. From ancient examination systems to AI-assisted writing tools, its forms have changed, but its presence has remained constant. Addressing it requires not only technological countermeasures but also a continuous rethinking of assessment practices and academic values [9,10].
References
[1] Cai, H. (2007). Cheating in ancient Chinese civil service examinations. Chinese Education & Society, 40(5), 6–10.
[2] Rowe, C. (2004). Cheating in ancient academia: Some reflections. Classical Journal, 99(4), 441–446.
[3] McCabe, D. L., Treviño, L. K., & Butterfield, K. D. (2001). Cheating in academic institutions: A decade of research. Ethics & Behavior, 11(3), 219–232.
[4] Bretag, T. (2016). Challenges in addressing plagiarism in education. PLOS Medicine, 13(12), e1002183.
[5] Lancaster, T., & Cotarlan, C. (2021). Contract cheating by STEM students through a file sharing website: A COVID-19 pandemic perspective. International Journal for Educational Integrity, 17(1), 1–16.
[6] Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2023). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 60(3), 365–375.
[7] Dwivedi, Y. K., et al. (2023). So what if ChatGPT wrote it? Multidisciplinary perspectives on opportunities, challenges and implications of GenAI. International Journal of Information Management, 71, 102642.
[8] Kasneci, E., et al. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274.
[9] International Center for Academic Integrity. (2021). The fundamental values of academic integrity (3rd ed.).
[10] Park, C. (2003). In other (people’s) words: Plagiarism by university students. Higher Education Review, 35(2), 5–26.

Appendix A.2. Suggested Structure of the Essay About the Timeline of Academic Dishonesty

Here’s a clear way to structure a short essay on the timeline of academic dishonesty, highlighting the most important turning points and why they matter:
1. Early Scholarly Traditions (Ancient–Medieval Period)
Academic dishonesty isn’t new. In ancient centers of learning like Library of Alexandria and medieval universities, scholars sometimes copied texts without attribution. However, the concept of “intellectual ownership” was still developing, so what we now call plagiarism wasn’t always condemned in the same way.
2. Rise of Authorship and Copyright (17th–18th Centuries)
With the spread of printing and laws like the Statute of Anne, authorship became legally protected. This period marks a shift: copying others’ work without credit began to be seen as both unethical and illegal.
3. Institutionalization of Education (19th Century)
As universities formalized curricula and examinations, cheating became more visible. Standardized testing environments increased opportunities for dishonest practices like copying during exams or hiring others to complete assignments.
4. Mass Education and Standardized Testing (Early–Mid 20th Century)
With education expanding globally, especially after events like World War II, academic dishonesty grew in scale. Multiple-choice exams and large classrooms made cheating harder to monitor and easier to attempt.
5. The Digital Revolution (Late 20th Century)
The rise of the internet fundamentally changed academic dishonesty. Students gained easy access to online essays, and “copy-paste” plagiarism became widespread. In response, tools like Turnitin emerged to detect similarities in texts.
6. Contract Cheating and Essay Mills (2000s–2010s)
A new form of dishonesty appeared: students outsourcing work to third parties. Online “essay mills” made it easy to buy custom-written assignments, which are harder to detect than traditional plagiarism.
7. AI and Generative Tools (2020s–Present)
The emergence of AI tools such as ChatGPT has created new challenges. Students can now generate essays instantly, raising questions about authorship, originality, and how learning should be assessed.
8. Current Responses and Future Directions
Institutions today are adapting by:
* Redesigning assessments (e.g., oral exams, project-based work)
* Using AI detection tools
* Emphasizing academic integrity education
The focus is shifting from punishment to prevention and ethical awareness.
How to Conclude Your Essay
You could argue that academic dishonesty evolves alongside technology and educational systems. Each era introduces new methods of cheating—but also new ways to detect and prevent it.

Appendix A.3. Suggested LLM Prompts Intended to Create the Timeline of Academic Dishonesty

1. Standard academic essay prompt**
“Write a 500–700 word academic essay on the timeline of academic dishonesty, covering key stages from early scholarly traditions to modern AI tools like ChatGPT. Include an introduction, body paragraphs, and conclusion.”
2. Concise classroom-style essay**
“Write a short (300–400 word) essay explaining the evolution of academic dishonesty, highlighting the most important historical turning points and explaining them briefly in simple language.”
3. Analytical/critical thinking prompt**
“Write an essay analyzing how academic dishonesty has changed over time, focusing on how technological developments—from the printing press to AI tools like ChatGPT—have influenced cheating methods and institutional responses.”
4. Structured outline-to-essay prompt**
“Using a clear chronological structure, write an essay on the timeline of academic dishonesty. Dedicate one paragraph to each major period and briefly explain why each stage was significant.”
5. Argument-driven essay prompt**
“Write a persuasive essay arguing that academic dishonesty evolves alongside technology, using historical examples (e.g., early plagiarism, standardized testing, internet era, AI tools like ChatGPT) to support your argument.”
Different angles you can explore with prompts
1. Write an essay on how academic institutions have responded to academic dishonesty over time, from ancient cultures to the modern day. Focus particularly on how policies have changed in response to online learning during COVID-19 and AI tools like ChatGPT. Include case studies or real-world examples.
2. Compare traditional forms of student cheating with modern AI-assisted academic dishonesty. Include discussion of why detection is harder now and how institutions might adapt.
3. Create a timeline-style academic article that describes the evolution of academic dishonesty across four key eras: (1) Ancient times, (2) Early universities and print culture, (3) COVID-era remote education, and (4) the LLM/AI era. Each section should be about 120 words and include examples.
4. Write a critical essay arguing that the use of LLMs in education is not inherently dishonest, but instead reveals flaws in traditional assessment methods. Contrast this with historical examples of academic misconduct. Include counterarguments and a nuanced conclusion.
5. Explore how different cultures and time periods have defined and dealt with academic dishonesty, from ancient Greece and China to modern Western institutions. How might LLMs challenge global norms around plagiarism and authorship?

Appendix A.4. The Timeline of Academic Dishonesty: From Early Scholarship to Artificial Intelligence

Academic dishonesty has evolved alongside educational systems, reflecting broader changes in technology, authorship, and institutional expectations. From early practices of textual borrowing to the emergence of artificial intelligence (AI)-generated content, the nature of cheating has continuously adapted to new academic environments. Understanding this historical trajectory provides insight into how institutions can better address integrity in the present and future.
In ancient and medieval scholarly traditions, the concept of intellectual ownership was not as clearly defined as it is today. Scholars often copied and transmitted texts without formal attribution, viewing knowledge as communal rather than individual [1]. However, as the printing press expanded access to written material, the notion of authorship became more significant. The introduction of early copyright laws in the 18th century marked a turning point, formally recognizing intellectual property and establishing plagiarism as both an ethical and legal violation [2].
During the 19th century, the institutionalization of education led to standardized curricula and formal examinations. These developments created structured environments in which academic dishonesty became more visible and measurable. Cheating methods during this period often involved impersonation or copying during exams, prompting institutions to introduce stricter supervision and honor codes [3]. By the early 20th century, mass education systems expanded rapidly, particularly after World War II, increasing class sizes and reliance on standardized testing. These conditions made monitoring more difficult and contributed to the normalization of certain dishonest practices among students [4].
The late 20th century saw a dramatic shift with the rise of digital technologies and the internet. Students gained unprecedented access to information, making it easier to copy and paste content directly into assignments. This era also witnessed the emergence of plagiarism detection software, which institutions adopted to counteract growing concerns about academic integrity [5]. At the same time, the accessibility of online resources blurred the boundaries between legitimate research and dishonest appropriation [6].
In the early 21st century, academic dishonesty became increasingly commercialized through the rise of contract cheating. Essay mills and online services allowed students to outsource their work, producing original but inauthentic submissions that were difficult to detect using traditional tools [7]. This shift highlighted the limitations of existing detection methods and underscored the need for more sophisticated approaches to assessment and verification [8].
The most recent phase in this timeline is defined by the emergence of AI-powered tools capable of generating coherent and contextually relevant text. These technologies have introduced new challenges, as they enable students to produce essays quickly without engaging in the learning process [9]. Consequently, educational institutions are re-evaluating assessment strategies, emphasizing critical thinking, oral examinations, and process-based evaluation over traditional written assignments [10].
In conclusion, academic dishonesty is not a static phenomenon but a dynamic one shaped by cultural, technological, and institutional changes. Each historical stage has introduced new forms of misconduct alongside corresponding efforts to prevent them. As AI continues to transform education, maintaining academic integrity will require adaptive strategies that prioritize ethical awareness and meaningful learning over mere content production.
References
[1] Eisenstein, E. L. (1980). The printing press as an agent of change. Cambridge University Press.
[2] Rose, M. (1993). Authors and owners: The invention of copyright. Harvard University Press.
[3] Anderson, M. S. (1998). The development of academic integrity standards. Journal of Higher Education, 69(5), 487–509.
[4] Altbach, P. G. (2004). The past and future of higher education. International Higher Education, 36, 2–5.
[5] Park, C. (2003). In other (people’s) words: Plagiarism by university students. Assessment & Evaluation in Higher Education, 28(5), 471–488.
[6] Howard, R. M. (2007). Understanding internet plagiarism. Computers and Composition, 24(1), 3–15.
[7] Clarke, R., & Lancaster, T. (2013). Commercial aspects of contract cheating. International Journal for Educational Integrity, 9(1), 27–44.
[8] Bretag, T. (2017). Defining and addressing contract cheating. Assessment & Evaluation in Higher Education, 42(8), 1145–1165.
[9] Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2023). Chatting and cheating: AI in education. Innovations in Education and Teaching International, 60(3), 1–12.
[10] Dawson, P. (2020). Defending assessment security in a digital world. Routledge.

Appendix A.5. The Evolution of Academic Dishonesty: A Timeline Across Four Eras

1.
Ancient Times
Academic dishonesty is not a modern phenomenon; its roots trace back to ancient civilizations where education, rhetoric, and scholarship were already highly valued. In societies such as Ancient Greece and Rome, students often memorized and reproduced texts, blurring the line between learning and plagiarism. For example, rhetorical training sometimes encouraged imitation of masters like Aristotle or Cicero, which could verge on uncredited copying [1]. Similarly, in Imperial China’s civil service examinations, candidates occasionally smuggled forbidden notes into exam halls or bribed officials to secure success [2]. These early forms of cheating reveal that whenever high-stakes evaluation existed, so did attempts to bypass it. Academic dishonesty in this era was shaped by oral traditions, limited access to texts, and rigid examination systems.
2.
Early Universities and Print Culture
With the rise of medieval universities in Europe and the invention of the printing press in the 15th century, academic dishonesty evolved alongside new educational structures. Institutions like the University of Paris and Oxford formalized examinations, creating more opportunities for cheating, including impersonation and unauthorized collaboration [3]. The spread of printed books made knowledge more accessible but also enabled plagiarism on a larger scale, as students could copy texts more easily [4]. By the 18th and 19th centuries, essay writing became central to education, and cases of ghost-writing and unattributed borrowing became more common [5]. Universities began implementing honor codes and disciplinary systems, signaling a shift toward institutional regulation of academic integrity.
3.
COVID-Era Remote Education
The COVID-19 pandemic marked a dramatic turning point in academic dishonesty due to the rapid shift to online learning. Remote assessments created new vulnerabilities, as students could access unauthorized resources, collaborate through messaging apps, or hire third-party services during exams [6]. Studies during this period reported significant increases in contract cheating and the use of online answer-sharing platforms [7]. For instance, websites offering real-time solutions to exam questions saw surges in traffic during lockdowns. Educators responded by adopting proctoring software, open-book assessments, and redesigned assignments, though concerns about privacy and effectiveness persisted [8]. This era highlighted how technological dependence can both enable and challenge academic integrity.
4.
The LLM/AI Era
The emergence of large language models (LLMs) and AI tools has introduced a new frontier in academic dishonesty. Tools capable of generating essays, solving problems, and mimicking human writing have made it increasingly difficult to distinguish original work from AI-assisted output [9]. Unlike traditional plagiarism, AI-generated content can be unique yet still undermine learning objectives. For example, students may use AI systems to produce entire assignments with minimal input, raising questions about authorship and intellectual ownership. Institutions are now grappling with policies on AI usage, detection tools, and ethical guidelines [10]. This era represents a shift from copying existing work to generating new, machine-produced content, challenging long-standing definitions of cheating.
References
[1] Plato. (2008). The Republic. Oxford University Press.
[2] Elman, B. A. (2000). A Cultural History of Civil Examinations in Late Imperial China. University of California Press.
[3] Cobban, A. B. (1999). The Medieval Universities: Their Development and Organization. Routledge.
[4] Eisenstein, E. L. (1980). The Printing Press as an Agent of Change. Cambridge University Press.
[5] Howard, R. M. (1999). Standing in the Shadow of Giants: Plagiarists, Authors, Collaborators. Ablex Publishing.
[6] Hodges, C., et al. (2020). The difference between emergency remote teaching and online learning. Educause Review, 27(1), 1–12.
[7] Lancaster, T., & Cotarlan, C. (2021). Contract cheating by STEM students through a file-sharing website. International Journal for Educational Integrity, 17(1), 1–16.
[8] Eaton, S. E. (2020). Academic integrity during COVID-19: Reflections from the field. International Studies in Educational Administration, 48(1), 80–85.
[9] Kasneci, E., et al. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274.
[10] Cotton, D. R. E., et al. (2023). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Assessment & Evaluation in Higher Education, 48(8), 1–15.

Appendix A.6. A Timeline from Classical Techniques to ChatGPT Era

Academic dishonesty in formal education has evolved alongside technology, pedagogical changes, and social media. While the strategies shift over time, the essential impulse—to gain advantage with minimal effort—remains consistent. The following overview traces cheating practices from the pre digital age through the AI era in computer science education, illustrating both continuity and innovation.
Pre Digital Era: Traditional Exam Deception
In the era before widespread computing, students used dozens of clever techniques to cheat during in class exams [1]. Common methods included cheat sheets hidden in clothing or stationery, passing notes, peeking at nearby students’ papers, whispering correct answers via gestures, or distracting proctors to enable copying [2–4]. These methods were highly physical and interpersonal—relying on stealth, spatial arrangement, and distraction rather than technology.
Early Digital/Computer Era: Hybrid Strategies and Early Tech
As computing entered education, students blended old tactics with emerging tools. On computer based exams, they might search the Internet from the same device or a mobile phone, use USB drives to transfer answers, or share solutions across local networks [5,6]. Some would conceal answers by turning off screens mid task or have another person take the exam remotely using login credentials [7,8]. In homework assignments, plagiarism—including verbatim copying, paraphrasing without citation, or self plagiarism—persisted [9]. Contract cheating, ghost-writing, and AI-assisted code generation also emerged [10,11]. The blending of digital and human strategies increased sophistication.
COVID Era & Remote Education: Crowd Cheating, Social Media, and New Vectors
With the shift to remote learning during COVID 19, cheating techniques adapted rapidly. Students formed crowdsourcing groups via social media to share answers, exam questions, and assignment solutions [12]. Ghost-writing, contract cheating, and the use of shared repositories across cohorts became more visible [13]. Some students used local chats, forums, or messaging to coordinate during live exams. Detection became harder because the physical separation removed many proctoring constraints. The remote environment also enabled identity fraud—having someone else complete work using shared credentials [14]. In essence, remote modalities expanded the scale and anonymity of cheating networks.
Post COVID/AI Era: Chatbots and the Next Wave
As institutions returned to in person or hybrid modalities, cheating continued to evolve. Firewalls were deployed to restrict web access, IP restrictions were enforced, and mobile phones were isolated during exams [12]. However, students countered with hidden devices or alternate gadgets. The latest frontier involves large language models (LLMs): chatbots that can generate essays, code, or answers in minutes. Some students reportedly used AI to draft essays, then lightly edit them to evade detection [12]. Because such output has no direct source to compare, plagiarism detection tools struggle. As ghost-writing detectors, crowdsourcing repositories, and social media coordination are integrated, the cat and mouse dynamic continues unabated [12].
Conclusion
The evolution of academic dishonesty reveals a constant interplay: as educational systems adopt new tools and safeguards, students simultaneously discover new loopholes and techniques. From concealed cheat sheets and note passing in early exams to remote crowdsourcing and AI generated submissions, the landscape has transformed—but the underlying motives remain. Effective prevention must therefore not only match technological ingenuity, but also nurture academic integrity, adapt assessment design, and anticipate future vectors such as AI. In the ever shifting arms race between cheaters and defenders, staying a step ahead requires vigilance, innovation, and ethical culture.
References
[1] Twomey, T., White, H., & Sagendorf, K. (2009). Pedagogy, not policing: Positive approaches to academic integrity at the University level.
[2] Erbe, B. (2007). Reducing test anxiety while increasing learning: The cheat sheet. College Teaching, 55(3), 96–98.
[3] Newstead, S. E., Franklyn-Stokes, A., & Armstead, P. (1996). Individual differences in student cheating. Journal of Educational Psychology, 88(2), 229–241.
[4] Yee, K., & MacKown, P. (2009). Detecting and preventing cheating during exams. In T. Twomey et al. (Eds.), Pedagogy, not policing (pp. 141–156).
[5] Dawson, P. (2016). Five ways to hack and cheat with bring-your-own-device electronic examinations. British Journal of Educational Technology, 47(4), 592–600.
[6] Khan, Z. R., & Balasubramanian, S. (2012). Students go click, flick and cheat… e-cheating, technologies and more. Journal of Academic and Business Ethics, 6, 1–26.
[7] Bretag, T., Harper, R., Burton, M., Ellis, C., Newton, P., Rozenberg, P., & van Haeringen, K. (2019). Contract cheating: A survey of Australian university students. Studies in Higher Education, 44(11), 1837–1856.
[8] Moten, J., Fitterer, A., Brazier, E., Leonard, J., & Brown, A. (2013). Examining online college cyber cheating methods and prevention measures. Electronic Journal of e-Learning, 11(2), 139–146.
[9] Ercegovac, Z., & Richardson, J. V. (2004). Academic dishonesty, plagiarism included, in the digital age: A literature review. College & Research Libraries, 65(4), 301–318.
[10] Vasylets, O., & Marín, J. (2022). Pen and paper versus computer-mediated writing modality as a new dimension of task complexity. Languages, 7(3), 195.
[11] Zheng, S., & Cheng, J. (2015). Academic ghost-writing and international students. Journal of Academic Ethics.
[12] Zdravkova, K. (2023). Evolution of academic dishonesty in computer science courses. In Proceedings of the HEAd’23 Conference.
[13] Medway, D., Roper, S., & Gillooly, L. (2018). Contract cheating in UK higher education: A covert investigation of essay mills. British Educational Research Journal, 44(3), 393–418.
[14] Bailie, J. L., & Jortberg, M. A. (2009). Online learner authentication: Verifying the identity of online users. Journal of Online Learning and Teaching, 5(2), 197–207.

Appendix B

The survey shared with the teaching staff from FCSE
Group 1:
Do you think the text was entirely written by a large language model (LLM)?
a. No, it seems completely original to me
b. No, although at moments I suspected that an LLM may have helped in its creation
c. At first glance it seemed original, but later I suspected that some parts were written by an LLM
d. Most of the text is generated by an LLM
e. I have no doubt that the text is the work of an LLM
2. Would you check the text for plagiarism if it were a student paper?
a. No, I don’t have a tool to do that automatically
b. No, because well-known plagiarism detection systems do not support Macedonian
c. Some sentences seemed suspicious, so I decided to search for them online
d. Some sentences seemed suspicious; I translated them into English and then searched for them online
e. I translated the text into English and then checked for plagiarism using a tool I have access to
3. Are there language errors in the text that are characteristic of LLMs?
a. No, although some words, such as the abbreviation, reminded me of LLMs
b. I noticed minor flaws, but our students make similar mistakes as well
c. I am unsure whether this is a translation of an already written text or a text written by an LLM
d. There are quite a few language errors; it is highly likely that it was written by an LLM
e. Starting with the title, which contains two colons, it is clear that the text was not written by a student
4. Does the writing style resemble that of an FCSE student?
a. Yes, many students write in a similar way
b. Some words and phrases are not typical for computer science students
c. It gave me the impression that almost all sentences have the same length, similar to how LLMs write
d. Some sentences seem like they were not written by our student
e. No, at first glance, it is clear that this is not our student
5. Are the facts in the text relevant and credible?
a. At first glance, there are no inaccuracies
b. There may be inaccurate information, but I do not notice it
c. If the student used unverified sources, the facts may be inaccurate
d. I don’t have time to examine the facts in depth, but some of them may not be entirely accurate
e. I noticed some errors; for example, the example about ancient Japan seems to refer to ancient China, at least according to the first reference
6. Are the cited references properly presented?
a. The references are consistently presented according to APA style
b. I think all references are properly presented, except that their DOI is missing
c. I searched for the first references on Google Scholar and found them
d. I checked all the references; some are not on Google Scholar
e. I checked all the references; some are not on Google Scholar, and some of the ones that exist have errors.
7. What do you suggest should additionally be checked to determine whether seminar papers are original or generated by an LLM?
Here is what ChatGPT suggests communication with it should look like in the link that the student must submit along with the seminar paper:
Group 2:
1. Do you believe that the link shared by the student is genuine, or was it created after completing the seminar paper?
a. Students whose seminar work was entirely written by an LLM will not try to hide it
b. Students whose seminar work was entirely written by an LLM will try to create a new link to make it appear as if they did it themselves
c. Students will try several times to get help from an LLM, but will share only the link that appears most honest
d. Students will most likely submit the correct link, although they may try to “polish” it
e. I have no doubt that the link realistically shows the full extent of the LLM’s assistance
2. Review the communication suggested by ChatGPT. Do you expect that students who submitted an AI-generated seminar paper would choose such an approach?
a. I don’t believe they would think of doing that
b. If an LLM already wrote their seminar paper, they wouldn’t spend extra time trying to pretend they wrote it themselves
c. They will try, but will give up when the LLM starts responding differently than recommended
d. They will definitely do it, but may skip some questions
e. Not only will they do it, but they will remind the LLM to respond as suggested
3. Does the proposed communication seem like a realistic scenario?
a. It is too extensive, and students will not have the patience to carry it out
b. The whole approach is complicated; they will probably give up
c. They certainly won’t implement it, but it may serve as inspiration
d. The communication is fairly realistic; some students will try to use it
e. The communication is fairly realistic; many students will try to use it
4. What do you suggest we do to preserve “integrity and ethical use of technology”?

References

  1. Bai, H., Voelkel, J. G., Muldowney, S., Eichstaedt, J. C., & Willer, R. (2025). LLM-generated messages can persuade humans on policy issues. Nature Communications, 16(1), 6037. [Google Scholar] [CrossRef] [PubMed]
  2. Baines, J. (2025, December 14). Publisher under fire after ‘fake’ citations found in AI ethics guide. The Times. Available online: https://www.thetimes.com/uk/science/article/ai-ethics-guide-citations-nsnjmz25b (accessed on 5 July 2026).
  3. Barman, K. G., Wood, N., & Pawlowski, P. (2024). Beyond transparency and explainability: On the need for adequate and contextualized user guidelines for LLM use. Ethics and Information Technology, 26(3), 47. [Google Scholar] [CrossRef]
  4. Bauer, N. F. (2025). Does ChatGPT increase language homogenization? In KI in medien, kommunikation und marketing: Wirtschaftliche, gesellschaftliche und rechtliche perspektiven (pp. 11–31). Springer VS. [Google Scholar] [CrossRef]
  5. Bayley, T., Maclean, K. D., & Weidner, T. (2026). Back to the future: Implementing large-scale oral exams. Management Teaching Review, 11(1), 159–170. [Google Scholar] [CrossRef]
  6. Bo, J. Y., Kumar, H., Liut, M., & Anderson, A. (2024, October 16–19). Disclosures & disclaimers: Investigating the impact of transparency disclosures and reliability disclaimers on learner-LLM interactions [Conference session]. AAAI Conference on Human Computation and Crowdsourcing (Vol. 12, pp. 23–32), Pittsburgh, PA, USA. [Google Scholar]
  7. Charness, G., & Grieco, D. (2026). Creativity and AI. The Economic Journal, ueag015. [Google Scholar] [CrossRef]
  8. Clark, D., Nicholas, D., Swigon, M., Abrizah, A., Rodríguez-Bravo, B., Revez, J., Herman, E., Xu, J., & Watkinson, A. (2025). Authors, wordsmiths and ghostwriters: Early career researchers’ responses to artificial intelligence. Learned Publishing, 38(1), e1652. [Google Scholar] [CrossRef]
  9. Cleland, J., Driessen, E., Masters, K., Lingard, L., & Maggio, L. A. (2026). When and how to disclose AI use in academic publishing: AMEE Guide No. 192. Medical Teacher, 48(4), 542–553. [Google Scholar] [CrossRef] [PubMed]
  10. Correa, C. (2026, March 3). How do researchers define “AI” when measuring adoption? Singularity Digital. Available online: https://singularity.digital/insights/how-many-people-use-ai-the-real-2025-adoption-picture-and-what-it-means-for-saas/ (accessed on 5 July 2026).
  11. Cotton, D. R., Cotton, P. A., & Shipway, J. R. (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61(2), 228–239. [Google Scholar] [CrossRef]
  12. Cotton, D. R., Wyness, L., Jane, B., & Cotton, P. A. (2025). Redefining assessments in the age of AI. In Teaching and learning in the age of generative AI (pp. 283–308). Routledge. [Google Scholar]
  13. Dwivedi, Y. K., Kshetri, N., Hughes, L., Slade, E. L., Jeyaraj, A., Kar, A. K., Baabdullah, A. M., Koohang, A., Raghavan, V., Ahuja, M., Albanna, H., Albashrawi, M. A., Al-Busaidi, A. S., Balakrishnan, J., Barlette, Y., Basu, S., Bose, I., Brooks, L., Buhalis, D., … Wright, R. (2023). Opinion paper: “So what if ChatGPT wrote it?” multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy. International Journal of Information Management, 71, 102642. [Google Scholar] [CrossRef]
  14. Edwards, B. (2023, July 14). Why AI writing detectors don’t work. Ars Technica. Available online: https://arstechnica.com/information-technology/2023/07/why-ai-detectors-think-the-us-constitution-was-written-by-ai/ (accessed on 5 July 2026).
  15. Elsevier. (2026). Generative AI policies for journals. Available online: https://www.elsevier.com/about/policies-and-standards/generative-ai-policies-for-journals (accessed on 5 July 2026).
  16. Erol, G., Ergen, A., Erol, B. G., Ergen, Ş. K., Bora, T. S., Çölgeçen, A. D., Araz, B., Şahin, C., Bostancı, G., Kılıç, İ., Macit, Z. B., Sevgi, U. T., & Güngör, A. (2025). Can we trust academic AI detective? Accuracy and limitations of AI-output detectors. Acta Neurochirurgica, 167(1), 214. [Google Scholar] [CrossRef] [PubMed]
  17. fourDNet. (2025). Tsinghua ICLR paper withdrawn due to numerous AI generated citations. Available online: https://www.reddit.com/r/MachineLearning/comments/1p01c70/d_tsinghua_iclr_paper_withdrawn_due_to_numerous/ (accessed on 5 July 2026).
  18. Gabriel, S. (2024, December 5–6). Generative AI and educational (in) equity [Conference session]. International Conference on AI Research (Vol. 4 No. 1. , pp. 133–142), Lisbon, Portugal. [Google Scholar]
  19. Goodier, M. (2025, June 15). Revealed: Thousands of UK university students caught cheating using AI. The Guardian. Available online: https://www.theguardian.com/education/2025/jun/15/thousands-of-uk-university-students-caught-cheating-using-ai-artificial-intelligence-survey (accessed on 5 July 2026).
  20. GPT-5.3. (2026). Appendix A.7.: A constructed interaction designed to simulate ethical use while concealing substantial AI involvement. Available online: https://docs.google.com/document/d/e/2PACX-1vToQHslGvqiTnDUCCaOz5WMYaCdluolR2PcDBDq8iC0wQ6z17fgQFp_zlPBzofjrMEZGFra3qNECaVg/pub (accessed on 5 July 2026).
  21. Grammarly. (2026). Plagiarism checker. Available online: https://www.grammarly.com/plagiarism-checker (accessed on 5 July 2026).
  22. Gravel, J., D’Amours-Gravel, M., & Osmanlliu, E. (2023). Learning to fake it: Limited responses and fabricated references provided by ChatGPT for medical questions. Mayo Clinic Proceedings: Digital Health, 1(3), 226–234. [Google Scholar] [CrossRef] [PubMed]
  23. Hacker, P., Engel, A., & Mauer, M. (2023, June 12–15). Regulating ChatGPT and other large generative AI models [Conference session]. 2023 ACM Conference on Fairness, Accountability, and Transparency (pp. 1112–1123), Chicago, IL, USA. [Google Scholar]
  24. Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22(1), 4. [Google Scholar] [CrossRef]
  25. Hamed, A. A., Zachara-Szymanska, M., & Wu, X. (2024). Safeguarding authenticity for mitigating the harms of generative AI: Issues, research agenda, and policies for detection, fact-checking, and ethical AI. iScience, 27(2), 108782. [Google Scholar] [CrossRef] [PubMed]
  26. Hong, S. (2025, August 15). AI-based fake papers are a new threat to academic publishing. Times Higher Education. Available online: https://www.timeshighereducation.com/opinion/ai-based-fake-papers-are-new-threat-academic-publishing (accessed on 5 July 2026).
  27. Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2), 1–55. [Google Scholar] [CrossRef]
  28. Jisoo, P. (2025). 224 Cases of academic misconduct at universities in 5 years… all students using ‘ChatGPT’ received F grades. Available online: https://view.asiae.co.kr/en/article/2025112411112686941 (accessed on 5 July 2026).
  29. Kacena, M. A., Plotkin, L. I., & Fehrenbacher, J. C. (2024). The use of artificial intelligence in writing scientific review articles. Current Osteoporosis Reports, 22(1), 115–121. [Google Scholar] [CrossRef] [PubMed]
  30. Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. [Google Scholar] [CrossRef]
  31. Kemp, S. (2025, October 15). Digital 2026: More than 1 billion people use AI. Data Reportal. Available online: https://datareportal.com/reports/digital-2026-one-billion-people-using-ai (accessed on 5 July 2026).
  32. Khurana, A., Subramonyam, H., & Chilana, P. K. (2024, March 18–21). Why and when llm-based assistants can go wrong: Investigating the effectiveness of prompt-based interactions for software help-seeking [Conference session]. 29th International Conference on Intelligent User Interfaces (pp. 288–303), Greenville, SC, USA. [Google Scholar]
  33. Kihwele, J. E., Mwamakula, E. N., & Mtandi, R. (2025). Portfolio-based assessment feedback and development of pedagogical skills among instructors and pre-service teachers. Journal of Applied Research in Higher Education, 17(5), 1510–1523. [Google Scholar] [CrossRef]
  34. Langreo, L. (2024, October 24). Parents sue after school disciplined student for AI use: Takeaways for educators. Education Week. Available online: https://www.edweek.org/technology/parents-sue-after-school-disciplined-student-for-ai-use-takeaways-for-educators/2024/10 (accessed on 5 July 2026).
  35. Laskar, M. T. R., Alqahtani, S., Bari, M. S., Rahman, M., Khan, M. A. M., Khan, H., Jahan, I., Bhuiyan, A., Tan, C. W., Parvez, M. R., Hoque, E., Joty, S., & Huang, J. X. (2024, November 12–16). A systematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations [Conference session]. 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA. [Google Scholar]
  36. Leaton Gray, S., Edsall, D., & Parapadakis, D. (2025). AI-based digital cheating at university, and the case for new ethical pedagogies. Journal of Academic Ethics, 23, 2069–2086. [Google Scholar] [CrossRef]
  37. Lee, J. Y., Agrawal, T., Uchendu, A., Le, T., Chen, J., & Lee, D. P. (2025, April 29–May 4). Exploring the duality of large language models in plagiarism generation and detection [Conference session]. 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 7519–7534), Albuquerque, NM, USA. [Google Scholar]
  38. Maynez, J., Narayan, S., Bohnet, B., & McDonald, R. (2020, July 6–8). On faithfulness and factuality in abstractive summarization [Conference session]. 58th Annual Meeting of the Association for Computational Linguistics (pp. 1906–1919), Online. [Google Scholar]
  39. Mazzi, F. (2024). Authorship in artificial intelligence-generated works: Exploring originality in text prompts and artificial intelligence outputs through philosophical foundations of copyright and collage protection. The Journal of World Intellectual Property, 27(3), 410–427. [Google Scholar] [CrossRef]
  40. McIntire, A., Calvert, I., & Ashcraft, J. (2024). Pressure to plagiarize and the choice to cheat: Toward a pragmatic reframing of the ethics of academic integrity. Education Sciences, 14(3), 244. [Google Scholar] [CrossRef]
  41. McLogan, J. (2025, October 10). Adelphi University facing lawsuit after AI-assisted plagiarism accusation against student. CBS News. Available online: https://www.cbsnews.com/newyork/news/adelphi-university-artificial-intelligence-plagiarism-lawsuit/ (accessed on 5 July 2026).
  42. MDPI. (2023). MDPI’s updated guidelines on Artificial Intelligence and authorship. Available online: https://www.mdpi.com/news/5687 (accessed on 5 July 2026).
  43. Nelken-Zitser, J. (2024). Parents sue their son’s school for punishing his AI use, heralding a messy future. Business Insider. Available online: https://www.businessinsider.com/parents-sue-school-over-son-punishment-ai-use-paper-massachusetts-2024-10 (accessed on 5 July 2026).
  44. Nguyen, K. V. (2025). The use of generative AI tools in higher education: Ethical and pedagogical principles. Journal of Academic Ethics, 23(3), 1435–1455. [Google Scholar] [CrossRef]
  45. OECD. (2025). NTU students penalised for alleged AI use in assignments, sparking disputes over fairness. Available online: https://oecd.ai/en/incidents/2025-06-21-d1dd (accessed on 5 July 2026).
  46. O’Sullivan, J. (2025). Stylometric comparisons of human versus AI-generated creative writing. Humanities and Social Sciences Communications, 12(1), 1708. [Google Scholar] [CrossRef]
  47. Park, C. (2003). In other (people’s) words: Plagiarism by university students—Literature and lessons. Assessment & Evaluation in Higher Education, 28(5), 471–488. [Google Scholar] [CrossRef]
  48. Princeton University. (2026). Generative AI for research and scholarship: Disclosing the use of AI. Available online: https://libguides.princeton.edu/generativeAI/disclosure (accessed on 5 July 2026).
  49. Rane, N. L., Tawde, A., Choudhary, S. P., & Rane, J. (2023). Contribution and performance of ChatGPT and other Large Language Models (LLM) for scientific and research advancements: A double-edged sword. International Research Journal of Modernization in Engineering Technology and Science, 5(10), 875–899. [Google Scholar]
  50. Reuters. (2026). Norway imposes near ban on AI in elementary school. Available online: https://www.reuters.com/technology/norway-imposes-near-ban-ai-elementary-school-2026-06-19/ (accessed on 5 July 2026).
  51. Rodrigues, T. V. (2025). Distant writing and the epistemology of authorship: On creativity, delegation, and plagiarism in the age of AI. International Journal of Social Sciences and Humanities Invention, 12(05), 8598–8613. [Google Scholar] [CrossRef]
  52. Ruan, H., Zhang, Y., & Roychoudhury, A. (2025, April 27–May 3). Specrover: Code intent extraction via llms [Conference session]. 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE) (pp. 963–974), Ottawa, ON, Canada. [Google Scholar]
  53. Sardinha, T. B. (2024). AI-generated vs human-authored texts: A multidimensional comparison. Applied Corpus Linguistics, 4(1), 100083. [Google Scholar] [CrossRef]
  54. Scarfe, P., Watcham, K., Clarke, A., & Roesch, E. (2024). A real-world test of artificial intelligence infiltration of a university examinations system: A “Turing Test” case study. PLoS ONE, 19(6), e0305354. [Google Scholar] [CrossRef] [PubMed]
  55. Sedacca, M. (2024, May 25). Emory student sues school for suspension over award-winning AI tool. New York Post. Available online: https://nypost.com/2024/05/25/emory-student-sues-school-for-suspension-over-award-winning-ai-tool/ (accessed on 5 July 2026).
  56. Shaw, D. (2025). The digital erosion of intellectual integrity: Why misuse of GenAI is worse than plagiarism. AI & Society, 40(8), 5819–5821. [Google Scholar] [CrossRef]
  57. Springer Nature. (2026). Journal policies. Available online: https://link.springer.com/brands/springer/journal-policies#Artificial%20intelligence%20(AI) (accessed on 5 July 2026).
  58. Stanford University. (2023). BCA guidance & recommendations. Office of Community Standards. Available online: https://communitystandards.stanford.edu/policies-guidance%23policies-guidance-links/bca-guidance-recommendations#generative-ai-policy-guidance (accessed on 5 July 2026).
  59. Sullivan, M., Kelly, A., & McLaughlan, P. (2023). ChatGPT in higher education: Considerations for academic integrity and student learning. Journal of Applied Learning & Teaching, 6(1), 31–40. [Google Scholar] [CrossRef]
  60. Thorne, S. (2024). Understanding the interplay between trust, reliability, and human factors in the age of generative AI. International Journal of Simulation—Systems. Science & Technology, 25, 10.1. [Google Scholar] [CrossRef]
  61. University of Oxford. (2026). Guidance on safe and responsible use of GenAI. Available online: https://www.ox.ac.uk/students/life/it/genai-tools/guidance-on-safe-and-responsible-use-of-genai (accessed on 5 July 2026).
  62. Van Dis, E. A. M., Bollen, J., Zuidema, W., van Rooij, R., & Bockting, C. L. (2023). ChatGPT: Five priorities for research. Nature, 614, 224–226. [Google Scholar] [CrossRef] [PubMed]
  63. Wang, F., Tang, X., & Yu, S. (2025). Cognitive outsourcing based on generative artificial intelligence: An analysis of interactive behavioral patterns and cognitive structural features. Acta Psychologica Sinica, 57(6), 967. [Google Scholar] [CrossRef]
  64. Yang, Y., Chern, E., Qiu, X., Neubig, G., & Liu, P. (2024). Alignment for honesty. Advances in Neural Information Processing Systems, 37, 63565–63598. [Google Scholar] [CrossRef]
  65. Yusuf, A., Pervin, N., & Román-González, M. (2024). GenAI and the future of higher education: A threat to academic integrity or reformation? Evidence from multicultural perspectives. International Journal of Educational Technology in Higher Education, 21(1), 21. [Google Scholar] [CrossRef]
  66. Yuxian, J. (2025). Bridging the knowledge-skill gap: The role of large language model and critical thinking in education. Computers & Education, 235, 105357. [Google Scholar] [CrossRef]
  67. Zdravkova, K. (2023, June 19–22). Evolution of academic dishonesty in computer science courses [Conference session]. 9th International Conference on Higher Education Advances (pp. 421–428), Valencia, Spain. [Google Scholar]
  68. Zdravkova, K. (2025). Professor: Who holds the copyright for AI-assisted and AI-generated contents? Language and Law/Linguagem e Direito, 12(1), 141–161. [Google Scholar] [CrossRef]
  69. Zdravkova, K., & Ilijoski, B. (2025a). Preventing academic dishonesty originating from large language models. In Advances in ICT research in the Balkans. BCI 2024. Communications in Computer and Information Science (Vol. 2391, pp. 118–134). Cham; Springer Nature Switzerland. [Google Scholar] [CrossRef]
  70. Zdravkova, K., & Ilijoski, B. (2025b). The impact of large language models on computer science student writing. International Journal of Educational Technology in Higher Education, 22(1), 32. [Google Scholar] [CrossRef]
  71. Zhong, H., Chang, J., Yang, Z., Wu, T., Mahawaga Arachchige, P. C., Pathmabandu, C., & Xue, M. (2023, March 30–April 4). Copyright protection and accountability of generative ai: Attack, watermarking and attribution [Conference session]. ACM Web Conference 2023 (pp. 94–98), Austin, TX, USA. [Google Scholar]
Figure 1. The perception of potential LLM authorship.
Figure 1. The perception of potential LLM authorship.
Education 16 01120 g001
Figure 2. Plagiarism-checking behaviour.
Figure 2. Plagiarism-checking behaviour.
Education 16 01120 g002
Figure 3. Presence of LLM-like linguistic errors.
Figure 3. Presence of LLM-like linguistic errors.
Education 16 01120 g003
Figure 4. Compliance with the writing style of a computer science student.
Figure 4. Compliance with the writing style of a computer science student.
Education 16 01120 g004
Figure 5. Credibility of facts.
Figure 5. Credibility of facts.
Education 16 01120 g005
Figure 6. Quality of references.
Figure 6. Quality of references.
Education 16 01120 g006
Figure 7. Confidence that the disclosed link is the original one.
Figure 7. Confidence that the disclosed link is the original one.
Education 16 01120 g007
Figure 8. Use of fabricated mutual communication according to suggested strategies.
Figure 8. Use of fabricated mutual communication according to suggested strategies.
Education 16 01120 g008
Figure 9. Perceived realism of the proposed communication strategy.
Figure 9. Perceived realism of the proposed communication strategy.
Education 16 01120 g009
Table 1. Ethical implications of different strategies of AI-assisted essay writing.
Table 1. Ethical implications of different strategies of AI-assisted essay writing.
Case StudyHuman Intellectual ContributionEthical Risk
Fully AI-generated essayVery lowVery high
AI-generated prompts + AI-generated essayLowHigh
Student-designed prompts + AI assistanceModerateModerate
AI rewriting of published textsVariableHigh to very high
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zdravkova, K. Generative Artificial Intelligence and the Ambiguity of Academic Integrity in Higher Education. Educ. Sci. 2026, 16, 1120. https://doi.org/10.3390/educsci16071120

AMA Style

Zdravkova K. Generative Artificial Intelligence and the Ambiguity of Academic Integrity in Higher Education. Education Sciences. 2026; 16(7):1120. https://doi.org/10.3390/educsci16071120

Chicago/Turabian Style

Zdravkova, Katerina. 2026. "Generative Artificial Intelligence and the Ambiguity of Academic Integrity in Higher Education" Education Sciences 16, no. 7: 1120. https://doi.org/10.3390/educsci16071120

APA Style

Zdravkova, K. (2026). Generative Artificial Intelligence and the Ambiguity of Academic Integrity in Higher Education. Education Sciences, 16(7), 1120. https://doi.org/10.3390/educsci16071120

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop