2. Materials and Methods
To answer the research questions, we used written coursework requirements (academic texts) submitted by first-year student teachers enrolled in the introductory teaching course at a university for teacher education. The project was registered with The Norwegian Agency for Shared Services in Education and Research (SIKT).
2.1. The Participants
We shared information about the project with all first-year students enrolled in the Primary and Lower Secondary Teacher Education (grades 5–10) specialization in mathematics (N = 45). Written consent was obtained from those who wished to participate (41 out of 45). Consent was collected after the students had completed the semester to prevent the students from feeling compelled to participate out of fear of sanctions that might affect their final grade or their relationship with their teachers. A total of 79 texts were assessed, divided between two submissions. The first submission (Text 1) included 38 texts, and the second submission (Text 2) included 41 texts; three of the 41 students did not submit Text 1. The students had a median age of 22 years and had all completed upper secondary school in the Norwegian school system with average or above grades in the Norwegian language, which is a requirement for admission to teacher education. However, all students reported having little to no experience with written self-assessments. Approximately 70% were female and 30% male. The teacher educators involved included one female, age 45, with 17 years of teaching and assessment experience at lower secondary, upper secondary and university levels, and one male, age 57, with 26 years of teaching experience from upper secondary and university levels.
2.2. Pilot Study
After the students had submitted their assignments, including their self-assessment, we decided to do a pilot (see
Figure 1). We selected six texts, including self-assessments submitted by first-year student teachers. Based on the evaluations and ratings made by one teacher educator, we selected two very good, two good, and two moderate texts. The texts ranged from 800 to 1000 words and concerned theories of motivation.
Pilot_phase A. A second teacher educator read through the texts and assessed them based on the criteria to get a common baseline.
Pilot_phase B. Developing prompts in the LLM. The first step was to enter the criteria that had been drawn up for the LLM. The next step was to try out various prompts, such as “assess based on criteria”, “rank the texts from best to worst,” and assess academic writing skills”. We found that these prompts did not provide thorough feedback on how the text could be improved. Consequently, we tested prompts that were explicitly asking for suggestions for improvement.
Pilot_phase C. During the pilot, we iteratively refined prompts to identify those that enabled the LLM to perform comprehensive evaluations. evaluation before we concluded on which prompts to include in the main study. After each round of prompting, the researchers reviewed and discussed the LLM outputs to determine necessary adjustments before finalizing the prompts for the main study.
Based on the pilot, the following steps of prompts were included in the main study:
Step 1: Can you evaluate an academic text written by a first-year student teacher?
Step 2: Thanks! This is the assignment given to the students and the assessment criteria for the text.
Step 3: Can you evaluate the text on a scale from 1 to 3, where 1 is low, 2 is middle, and 3 is high?
Step 4: This is a new text for evaluation. Step 3_1 was used if there was a discrepancy between evaluations.
2.3. Main Study: Essays and Criteria
Our goal was to gain insight into whether LLMs could serve as useful partners in the assessment of academic texts, both for students and educators. We used interrater reliability to determine the degree of alignment between different assessments. Additionally, we compared feedback provided by the various evaluators to compare and evaluate accuracy. In the discussion section, we discuss whether LLMs represent a pedagogically useful tool for supporting students in developing academic writing skills and for teacher educators in their assessment tasks. We analyzed two different essays from the students, both based on selected learning outcomes from the program. These learning outcomes are critical because they specify what the assignments aim to measure. Yet, there are multiple learning goals, and each essay is related to a subset of learning goals. Interpretation of whether students meet the criteria for their essays can be used to discuss aspects of content validity and construct validity (
Messick, 1995).
However, an important question in this study is if we can examine what teacher educators, student teachers, and LLMs emphasize when applying assessment criteria for evaluating academic texts.
The first essay (E1) was submitted in the beginning of the semester on the topic of “The teacher’s role”, and the second essay (E2) was submitted at the end of the semester on the topic of Motivation. The criteria were:
A clear and precise problem statement.
You use relevant course literature (minimum 3 sources). It is not sufficient to use notes from lectures.
Clear professional language and good coherence between the different parts of the assignments. Use technical terms.
Reflect on the significance it can have for your work with the students.
Correct citation and reference list.
The essay must be presented in line with research ethic guidelines.
First, the students were asked to provide a self-assessment of their own essays and responded to questions regarding what they felt they had mastered and areas requiring improvement. This was voluntary, and most student teachers submitted a self-assessment together with their academic texts.
Second, the teacher educator evaluated each essay to determine whether the students passed the assignment or if they needed to resubmit it, and the educator provided feedback regarding the quality of the text and areas of improvement. Additionally, each essay was given a score from 1 to 3 (1 indicating low quality, 2 indicated middle, and 3 indicated high), but the students only got information about pass/resubmit (not the numeric score) alongside the qualitative feedback. This assessment was conducted without the use of LLMs.
When preparing data for analysis, the students did not score their essays numerically, but the researchers estimated their scores based on their response to the open-ended questions in the self-assessments. Overall, 38 essays (E1) and 41 essays (E2) were collected and could be used in the further evaluations by the LLMs. We used an LLM service supported and approved by the university, ensuring that all texts were anonymized and contained no personal information. Each researcher operated their own LLM.
All of the analyses were carried out in the same thread/chat. In some cases, the LLMs gave a score of 4, and if this happened repeatedly the prompt about the scale being from 1 to 3 was re-introduced.
The results from the evaluations were analyzed by looking at the agreement and interclass correlations between the LLMs, the teacher educator, and the student teachers.
2.4. Analytical Approaches
To answer RQ1, we analyzed how the teacher educator, the student teachers, and the LLMs evaluated the academic texts. We focus on how these evaluators used criteria operationalized from the learning outcomes. Potential challenges include the operationalization of criteria, as well as differences in interpreting and weighting them. Validity refers to whether the interpretations of scores and related decisions are well founded (
Messick, 1995). Ensuring good validity in essay assessment depends on raters’ thorough understanding of the criteria they apply. Both student teachers and teacher educators had access to the task descriptions of the two essays, the assessment criteria, and relevant literature. The LLMs were also provided with the task descriptions and criteria for both essays, in addition the LLMs had access to theories relevant for the essay. It appears rather clear that teacher educators and LLMs actively use these criteria in their feedback; for example, both LLMa and LLMb base their feedback on points in the criteria and include relevant theories. In contrast, student teachers tend to adopt a more precise but concise style, applying the criteria to a limited extent, which may challenge the validity of their evaluations and ratings. There may be a difference in that teacher educators emphasize the discussion aspect in the essay to a greater extent than LLMs and student teachers. Overall, assessments conducted by teacher educators and LLMs are well founded on the criteria, which supports the validity of content and conceptual understanding (
Bouwer et al., 2023).
Several measures could be used to examine the consistency of ratings. In this study, we use agreement in percentages, correlation, Cronbach’s alpha, and Fleiss’s kappa. As the latter is sensitive to cases with one dominating rating, we also used the prevalence-adjusted kappa (PBAK) (
Byrt et al., 1993).
To answer RQ2, we started off with an inductive approach to the data and performed a descriptive content analysis.
Coding procedure and analytical approach:
Both researchers read through the feedback provided by both LLMs, the teacher educator, and the student teachers. We coded the feedback first according to the assessment criteria, more precisely, which criteria were implicitly or explicitly featured in the feedback. We uploaded the content to an Excel spreadsheet to be able to compare the four different types of assessment, looking for patterns across the dataset. In
Section 3.1 we provide some examples to help visualize the differences and similarities in the evaluations. Thereafter, we proceeded with a more deductive approach and coded the feedback according to
Hattie and Timperley’s (
2007) framework (task, process, self-regulation and self) to identify which feedback level seemed to be salient in the different types of feedback. Regarding step 3_1, when disagreement among the graders was observed, we analyzed the differences in feedback and then prompted the LLMs to assess the quality of the text and suggest improvements. We compared the emphasis in the LLM-generated feedback with that of the educator’s feedback. This step was applied to three essays where the LLMs’ assessments were high but the students’ and teacher educator’s assessments were low, and to five essays where the LLMs’ assessments were high and the students’ and teacher educators’ assessments were middle. We concluded the review of feedback after these eight essays, as no new patterns or insights emerged from the LLMs’ responses.
3. Results
This section contains results related to the two research questions.
3.1. RQ1: Comparing Ratings from LLMs with Teacher Educator and Student Teachers
To answer RQ1, we analyzed how the LLMs evaluated the essays and compared their assessments with those of the teacher educator and the student teachers’ self-assessments.
Table 1 shows the frequency of ratings (1 low, 2 middle, and 3 high), the mean, the standard error, and the median from the evaluation of the texts carried out by LLMs and the teacher educator, along with the self-assessments made by the student teachers. Analysis conducted with IBM SPSS 29.
When comparing the scoring from the LLMs, the teacher educator, and student teachers, it seemed that both LLMa and LLMb were more positive than both the teacher educator and the student teachers. The LLMs assessed a larger proportion of the texts as high level (93% by LLMa and 84% by LLMb) compared to 32% assessed as high level by the teacher educator and 22% assessed as high level by the student teachers. Furthermore, we found differences between the LLMs, and LLMa seemed to be slightly more positive than LLMb (74 vs. 67 high). Finally, the teacher educator had higher ratings of the texts compared with the student teachers (28 vs. 17 high).
To compare the ratings, we analyzed the agreements between the LLMs, the teacher educator, and the student teachers (
Table 2). The table also contains measures of the strength of co-variation, namely the correlation (Corr) and Cronbach’s alpha (CA), and the strength of agreement (Cohen’s Kappa). Kappa can be low even though an observed agreement is high if one category is dominating the ratings. The ratings of academic texts from the LLMs are mostly as “high = 3”, a few as middle (2), and none as low (1) (see
Table 1). To compensate for skewness in ratings and low levels of kappa, the prevalence-adjusted kappa (PBAK) (
Byrt et al., 1993) was calculated for the cases with one dominating rating.
The LLMs had lower levels of agreement with both the teacher educator and the student teachers, and there seemed to be an increased level of agreement between the teacher educator and the student teachers from Text 1 (41.7%) to Text 2 (64.9%). There was strong agreement between the two LLMs regarding coding, an acceptable level of PBAK, and a low level of kappa. Yet the kappa is sensitive to one dominating rating, and the LLMs used high in rating of most texts. The results also show that there was a lack of co-variation and a lack of strength of agreement beyond what the LLMs agreed on (see example in
Table 3).
As can be seen from this example (
Table 3), the two LLMs comment on different aspects of the texts. LLM1 comments on language and understandability, while LLM2 calls for deeper analysis. It seems that LLM1 in this case is more in line with the criteria concerning academic tone and language, while LLM2 emphasizes the criteria that deal with practical implications for teachers.
3.2. RQ2: Evaluation of the Academic Texts
RQ2: What do the LLMs emphasize in the evaluation of the academic texts compared to the teacher educator and the student teachers?
As mentioned, we systematically reviewed and coded the written evaluations provided by the LLMs, the teacher educator, and the student teachers to compare the main topics and arguments, identifying similarities and differences across the assessments. Examples are presented below to illustrate these points. We begin by highlighting what the LLMs typically emphasized in their feedback.
Both LLMa and LLMb consistently provided longer and more systematic assessment texts compared to both the teacher educator and the student teachers. The LLMs’ evaluations aligned closely with the topics outlined in the assessment criteria; however, they were less precise in their analysis of the use of literature within the discussion section.
Summary of feedback from LLM. See
Appendix A for detailed feedback:
The student presents a well-justified problem statement and provides a clear overview of the assignment. Key concepts are well-defined and supported by relevant literature, demonstrating a solid understanding of the teacher’s role. The discussion is structured and coherent, with strong arguments and relevant examples. The conclusion effectively summarizes the main points and reflects on the implications for the student’s future role as a teacher. The bibliography is complete and correctly formatted, and the use of subheadings helps structure the assignment clearly. The academic language is clear, and there is good coherence throughout the text. The assignment is academically sound and ethically responsible. However, the discussion could benefit from more depth regarding the impact of societal changes on the teacher’s role. Overall, the text meets all the requirements and demonstrates a strong understanding of the subject, earning a high rating of 3.The LLMs addressed the central points outlined in the criteria when assessing the essays, but they encountered difficulties with both the formal and conceptual aspects. Formally, they struggled to accurately identify the number of correct references and to distinguish genuine references from incorrect ones. Conceptually, the LLMs had difficulty evaluating whether the literature was discussed in an academically sound manner. Additionally, the LLMs appeared to weigh all criteria equally, whereas the teacher educator and student teachers placed greater emphasis on certain aspects as more critical to producing quality academic writing.
Regarding the feedback from the teacher educator, it did not include a written evaluation explicitly addressing each criterion. Instead, the feedback focused on the extent to which the assignment was correctly understood, highlighting areas that required improvement as well as those that were well presented, although this was not always consistent. The teacher educator’s evaluations frequently emphasized the students’ understanding of literature and their ability to incorporate it into their discussions. Additionally, some feedback included motivational comments, such as “you impress me”. The teacher educator’s feedback must therefore be understood as a combination of feedback on the task level, process level, and self-level (
Hattie & Timperley, 2007).
This is an example of teacher feedback: You are nearly there. I see from your self-assessment that you find it challenging to know what should go into the clarification of terms and what belongs in the discussion. In this text, I think it would have been useful to first differentiate between traits and competencies. The course literature is not entirely clear on this (the terms are somewhat interchangeable), so I completely understand why you might not have considered it. Next, you could explain what the different competencies entail in the explanation and then discuss them in the discussion section. You end up with less space for discussion because you are explaining as you go along. These are things you should think about for your next assignment. Also note that you need to include page numbers for the references in the text.
Finally, the students provided the shortest feedback, typically consisting of one or two brief comments that highlighted one or two specific points. They often emphasized challenges related to organizing their writing (ranging from initiating the writing process to using the APA style for reference formatting). They also reported difficulties integrating research literature into their discussions and conclusions. Thus, their self-assessments reflected feedback on both the process level and the self-regulation level (
Hattie & Timperley, 2007). This is a representative example of self-assessment:
I feel that I to some extent manage to use the course literature in a meaningful way and write about my own experiences. I need to start writing my assignments at an earlier stage and understand where to stop clarifying concepts and where to start my discussion. I hope we can have writing workshops on campus to help me get started with my assignments.Regarding what the LLMs emphasize, the excerpts presented below are representative examples illustrating their focus on (a) theories, (b) analyses (including deeper and more critical examination), (c) discussions (emphasizing balanced and critical reflections), and (d) practical implications.
- (a)
Theories.
The LLMs commented on the need for more literature or citations. There were also examples of the LLMs advising the students to develop their academic texts through the connections between theories, for example: You can try to draw more connections between the different theories. How can the various theories complement or contradict each other? Is it possible that a mix of approaches could be most effective in this specific case?
- (b)
Analyses.
The LLMs suggested further work with the analysis, for example: Although the text does a good job of explaining each motivation theory and how they can be applied to the case, you can take it a step further by analyzing and discussing the potential strengths and weaknesses of each theory in relation to the case. This will provide a more critical analysis and can reveal insights that might not emerge by merely describing the application of the theories.
- (c)
Discussions.
The LLMs provided feedback about the discussions in the academic texts and about developing a more balanced discussion, for example: Try to give each theory roughly equal space in the discussion. In this text, there is a bit more focus on the behaviorist perspective, while the humanistic and cognitive perspectives could receive more attention.
The LLMs also suggested more critical reflections in the discussions, for example: Finally, you can spend more time on critical reflection regarding the theories. Are there aspects of the theories that do not fit with what you have observed in practice? Are there limitations in the theories that should be considered? This type of critical reflection can enrich the discussion and demonstrate a deeper understanding of the material.
- (d)
Implications.
The LLMs suggested improvements on the implications of the analysis and discussion, for example: You can more clearly discuss what the insights from the motivation theories mean for educational practice. How can teachers use this knowledge in the classroom? How can the insights be translated into concrete educational strategies?
4. Discussion
The first research question of this paper examined the agreement of two LLMs, one teacher educator, and the student teacher who wrote the text. The results revealed a relatively strong agreement between the two LLMs (80%) but a low Cronbach’s alpha (0.23).
This result could be partly due to the use of a three-point scale, where agreement by chance is likely to occur. If a more nuanced scale had been employed, the results might have differed. The low Cronbach’s alpha might also reflect poor consistency between the two LLMs. While some literature suggests that LLMs can contribute to providing consistent evaluation, other research indicates that different LLMs can give different assessments (
Pack et al., 2024). However, it is surprising that we got such results when we used the same LLMs and the same prompts and performed the assessments at the same time.
Further, the LLMs were more positive in their ratings of the academic texts compared to the teacher educator and the student teachers. The variation between the LLMs and the human assessors is to some extent in line with recent research (
Bui & Barrot, 2025;
Kostic et al., 2024;
Manning et al., 2025). Further, the agreement between the teacher educator and the student teachers was more consistent compared to their agreements with the LLMs. This difference may be explained by the student teachers’ prior knowledge of what the teacher educator valued and disvalued, as they had received instructions in advance of writing the essays and had received formative assessments on essays handed in at earlier stages in the course.
RQ2: What do the LLMs emphasize in the evaluation of the academic texts compared to the teacher educator and the student teachers?
Regarding the differences between the LLMs, the teacher educator, and the student teachers (RQ2), we found that the LLMs produced longer and more comprehensive evaluation texts. This is in line with findings from
Dai et al. (
2023) who found that LLMs can provide longer and more detailed comments compared to teachers/instructors. In our study, the evaluations from the LLMs followed the main points of the assessment criteria and included comments under each main point. Such feedback could be advantageous for the student, given that the evaluation is perceived as specific, detailed, and thorough (cf.
Dawson et al., 2019,
Boud & Dawson, 2023). This type of response was much more detailed compared to the teacher educator and the students’ own responses, and was typically on the process level.
The literature is inconclusive regarding the need for individualized and personalized feedback.
Dawson et al. (
2019) point out varying interpretations of what constitutes personalized and individualized feedback but emphasize that students generally prefer feedback directed specifically at their own texts rather than generic comments.
Boud and Dawson’s (
2023) framework underscores the importance of feedback that recognizes and responds to the students’ different needs. The teacher will, in most cases, have access to more information regarding the students’ individual needs, such as prior knowledge and background, specific learning needs (e.g., dyslexia), challenging home situations, and the students’ own preferences for feedback modes based on previous experiences. This contextual understanding enables teachers to provide more tailored feedback that considers the student’s broader lifeworld, something that current LLMs cannot fully replicate.
Another challenge with LLM-generated assessments in essay writing is the risk of students receiving disproportionately positive evaluations that can foster unrealistic expectations, leading to disappointment and confusion if they receive more critical feedback from their teachers. Finally, as noted above, LLM feedback can be quite comprehensive, which might overwhelm students who are still developing their academic skills, potentially undermining their motivation to take the necessary next steps.
Across the feedback on these texts, we found that the LLMs tended to emphasize areas such as theory, data analysis, discussion, and implications. Yet, it appears that the LLMs consistently rated the texts as better than what students or teacher educators did, and with longer and more detailed feedback. Regardless of this, the feedback seemed generic and seldom fine-tuned. This is in line with research (
Bui & Barrot, 2025;
Kostic et al., 2024;
Manning et al., 2025). For example,
Kostic et al. (
2024) show that it can be difficult for LLMs to evaluate complex texts. This may be an explanation for why the LLMs did not propose extensive changes to individual parts of the texts. Thus, even if the student chooses to improve the text in accordance with the feedback provided by the LLM, it is uncertain whether the suggested improvements are sufficient.
Another important aspect of feedback discussed in the literature is timeliness (
Dawson et al., 2019), emphasizing that feedback is most effective when delivered at a stage in the process where the student has a genuine opportunity to make necessary adjustments (
Gamlem & Smith, 2013). LLMs offer the advantage of providing feedback instantly at any point in the student’s writing process, unlike waiting for potentially delayed responses from busy teachers. However, this benefit depends on the student’s ability to recognize the appropriate moments to seek feedback, which involves both process understanding and self-regulation skills.
On the other hand, LLMs can support educators in providing students with detailed and comprehensive feedback in line with research on good feedback practices (
García-Varela et al., 2025;
Wisniewski et al., 2020), particularly at the task and process levels. Further, teacher educators can calibrate LLM outputs to align with their own standards, although there is no guarantee that the teachers’ assessments are more correct, true, or accurate compared to the LLM. Nonetheless, this AI–human collaboration can be beneficial for both educators and students (
Xiao et al., 2024). Further, the teacher educator can facilitate prompt-engineering classes or, in collaboration with the students, develop prompt cheat-sheets tailored for different stages of the writing process. In this way, students could use the LLM as scaffolding in their learning process.
4.1. Limitations
This study has some limitations that should be acknowledged. The number of texts was limited, and a larger sample would increase the generalizability of the findings; however, there were still significant variations within the current sample. The texts were evaluated by a single experienced teacher educator have led to bias in the ratings (e.g., disproportionately strict or permissive). The reported agreement reflects alignment between LLMs and one experienced teacher’s judgments (not with a group of human raters). There were two human raters in the pilot study to investigate consistency and inter-coder reliability, which yielded promising findings, but this was not continued in the main study.
Furthermore, the students’ self-assessment was voluntary, and they were aware that these self-assessments would not be graded as part of the assignment. This may have led some students to invest minimal effort in evaluating their own work, which could explain why their self-assessment texts were significantly shorter than the feedback provided by the teacher educator and the LLMs. A more structured and thorough self-assessment process might yield deeper insights into the students’ reflections. It is also possible that the students, all being first-year, lacked experience with self-assessment and were uncertain of the expectations of higher education. Moreover, the students’ scores were inferred indirectly by the researchers, potentially reducing the reliability. Therefore, student ratings should be interpreted cautiously and as indicative rather than definitive representations of their evaluative judgments.
Furthermore, the use of a limited three-point rating scale may increase the likelihood of coincidental agreement, potentially inflating interrater reliability coefficients. Therefore, caution is necessary when drawing conclusions about the meaning of findings derived from such comparisons.
In our study, we observed that the same LLM, operated independently by two researchers at the same time, provided differing feedback and ratings. While we do not know whether these results would differ if different LLMs had been used, as was the case in
Li and Liu’s (
2024) study, we cannot rule out the possibility that a higher-quality LLM might provide more consistent and higher-quality feedback than a lower-quality model.
Overall, these limitations suggest that the findings on interrater reliability should be viewed as preliminary evidence of agreement under limited conditions, rather than as strong claims about the validity assessments by humans and LLMs.
4.2. Practical Implications
LLMs can request improved argumentation and application of relevant theory; however, student teachers often lack access to model texts, making them reliant on prompts derived from task descriptions and assessment criteria. While the LLMs can support students in this regard, students must disclose their use of such tools in their submissions.
LLMs can generate extensive evaluations based on established criteria, but their limited focus on discussion complicates their use without developing tailored prompts. Educators must still manually review texts, which can potentially double their workload and reduce time efficiency.
Low and mid-level student texts may receive overly positive feedback from LLMs, which can lead to unrealistic expectations. Students need to critically evaluate the reliability of such feedback and refine their prompts to better align with the educator’s standards. While this process can foster learning and critical thinking. Such dynamics challenge the validity of assessment outcomes, particularly if top performers succeed primarily by mastering prompt engineering rather than academic writing. Access to different LLMs varies because higher-quality models often require paid subscriptions, while lower-quality or free options may offer limited features or accuracy. This variation in both quality and cost impacts students differently based on their financial resources.
5. Conclusions
By comparing the evaluations made by LLMs, the teacher educator, and the student teachers, we contribute to the expanding field of AI-supported assessment. We found agreement regarding the high-level ratings of the texts by the LLMs. Furthermore, some agreement was observed between the student teachers and the teacher educator; however, the LLMs did not demonstrate agreement with either the students or the teacher educator. These findings highlight the ongoing need to improve the evaluation capabilities of LLMs.
Second, we examined the characteristics of feedback provided by the LLMs, the students, and the teacher educator. The LLMs provided much longer feedback compared to that from the students, while the teacher educator’s feedback was longer than that of the students but shorter than that of the LLMs. The students typically provided a few sentences, whereas the teacher educator wrote one or two paragraphs. The LLMs’ evaluation included several paragraphs in length and explicitly followed the criteria. However, the teacher educator seemed to be selecting and using only some of the criteria in the written feedback, with particular emphasis on improving the integration of theory and research within the discussion section.
Overall, LLMs provided students with specific feedback on how to improve their academic texts. This feedback was generally more comprehensive than that offered by the teacher educator or the students themselves and was typically focused on the process level. Nevertheless, the LMMs’ feedback was not necessarily in alignment with that of the teacher educator or the aspects that were emphasized when evaluating the student teachers’ texts. Yet, there is reason to believe that LLMs can be valuable tools both for students developing academic writing skills and for educators who invest significant time in assessment, especially when used in a collaborative mode combining human and machine input. Based on this study and recent research, we are not certain that LLMs’ feedback helps to improve the academic texts in the right direction, and thus, further research is needed into how students and teachers can use and collaborate with LLMs in the writing and assessment of academic texts.