Next Article in Journal
Sustainable AI Integration in Teacher Education: From Personalised Learning to Signature Pedagogies
Previous Article in Journal
Mathematics as a Gateway, Not a Barrier: Reimagining Engineering Preparation for the 21st Century
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Seeking Help from Large Language Models: Exploring the Assessment and Feedback of Student Teachers’ Academic Texts

by
Astrid Gillespie
* and
Ove Edvard Hatlevik
Department of Primary and Secondary Teacher Education, Oslo Metropolitan University, 0130 Oslo, Norway
*
Author to whom correspondence should be addressed.
Educ. Sci. 2026, 16(5), 787; https://doi.org/10.3390/educsci16050787
Submission received: 20 March 2026 / Revised: 5 May 2026 / Accepted: 12 May 2026 / Published: 16 May 2026

Abstract

Student teachers can seek help from large language models when writing their academic texts during their initial teacher training, and teacher educators can use large language models to evaluate the submitted academic texts. However, existing research presents inconsistent evidence regarding whether large language models assess academic work in ways comparable to human evaluators. To our knowledge, few studies have examined evaluations made by both student teachers and teacher educators alongside those generated by large language models. This study addresses two research questions concerning how academic texts are evaluated by student teachers, teacher educators, and large language models. First, we found that the two large language models showed agreement with each other but did not consistently align with the evaluations provided by either the student teachers or the teacher educator. Second, the large language models produced substantially longer evaluation texts that closely followed the structure of the assessment criteria but struggled with evaluating the discussion sections. Although the large language models offered practical suggestions for improving academic texts, their feedback did not emphasize the same aspects highlighted by the teacher educator. Implications for practical use of generative AI-tools and needs for further research are discussed.

1. Introduction

1.1. Aim and Background

Generative artificial intelligence, including large language models (LLMs), presents both significant opportunities and complex challenges for students and educators. In recent years, there have been major developments in the field (Agostini & Picasso, 2024). Research shows that students use LLMs and that LLMs are, to varying degrees, able to provide consistent assessments and academic feedback to users (e.g., students and educators) (Jansen et al., 2024; Lindsay et al., 2026; Seo et al., 2025). In higher education, academic writing holds a central place through work requirements, assessments, and exams in fields such as the educational sciences, psychology, and the social sciences.
There are three important reasons for conducting this study. First, as teacher educators, we observe that most students and teachers frequently use LLMs. Second, research on the topic is inconsistent (Dai et al., 2023; Manning et al., 2025) and there is no clear evidence that LLMs evaluate academic texts in line with human assessors (Bui & Barrot, 2025; Kostic et al., 2024). Third, regarding human assessors, to our knowledge there is a lack of research including both students and teachers as human assessors compared to LLMs.
From the educator’s perspective, AI tools have the potential to ease the time-consuming task of evaluating longer texts and essays (García-Varela et al., 2025; Seo et al., 2025). However, research indicates that assessments produced by LLMs are not always aligned with those of human assessors (Bui & Barrot, 2025; Manning et al., 2025). From the students’ perspective, AI can be used not only to generate text and full essays but also as a personal tutor that guides them through their writing process (Jansen et al., 2024). The extent to which AI support improves academic outcomes depends on how well the AI-generated suggestions are recognized as valuable by the educator, who remains responsible for the final assessment, as well as on the students’ ability to identify their own strengths and weaknesses and tailor their prompts accordingly.
Students have access to LLMs and can use them to request additional feedback which can enhance both performance and motivation (Jansen et al., 2024; Meyer et al., 2024). Several studies have compared feedback from LLMs with that of human assessors (Bui & Barrot, 2025; Dai et al., 2023; Kostic et al., 2024; Manning et al., 2025), and investigated differences between feedback provided by different LLMs (Li & Liu, 2024) and between different human assessors (Manning et al., 2025; Shermis et al., 2002), However, to our knowledge, no studies have compared feedback from LLMs, teachers, and the students themselves. Yet, self-assessment can play an important role in student learning in higher education (Chan & Chen, 2023; Sharma et al., 2016). Therefore, it is valuable to compare the feedback provided by all three agents.
This study examines teacher educators’ use of LLMs in automated essay scoring and compares assessments generated by LLMs with both teachers’ grades and feedback, as well as student teachers’ self-assessments. The term student teacher (ST) refers to students enrolled in a five-year master’s teacher education program for primary- and lower secondary education. The term teacher educator (TE) refers to university instructors who teach in the same teacher education program. The primary aim of this paper is to examine how academic texts are assessed by student teachers, teacher educators, and LLMs by looking at the interrater reliability of the assessments. The second aim is to investigate the characteristics of the feedback provided by LLMs and explore how both students and educators may effectively utilize this feedback. Finally, our interest in this topic is also grounded in our background as educational practitioners, and we aim at contributing to the practitioner community through our research.

1.2. Theoretical Framework and Prior Research

We draw on well-established theory within the field of assessment in education. In particular, we use Hattie and Timperley’s (2007) levels of feedback to analyze the feedback provided. Furthermore, we also rely on Boud and Dawson’s (2023) framework for feedback literacy.

1.2.1. Feedback

Feedback has long been recognized as a key factor in promoting learning (Black et al., 2004; Black & Wiliam, 1998; Hattie & Timperley, 2007; Lindsay et al., 2026; Sadler, 1989; Wisniewski et al., 2020). Hattie and Timperley (2007) define feedback as information provided by an agent (e.g., teacher, peer, book, parent, self, experience) about an individual’s performance or understanding. They identify four levels of feedback:
-
Task Level: How well the task is accomplished, including corrective suggestions.
-
Process Level: Understanding of the learning process and deeper comprehension of the topic.
-
Self-Regulation Level: Encouraging self-evaluation and monitoring.
-
Self-Level: General praise like “well done” or “you are a good student”.
Wisniewski et al. (2020) found that feedback incorporating task, process, and self-regulation levels is most effective for learning outcomes. Dawson et al. (2019) highlighted that feedback has evolved to be student-driven, involving multiple contributors. Their study revealed that 84% of students valued detailed, specific comments, and 25% appreciated personalized feedback. While timeliness was not a major concern, Gamlem and Smith (2013) noted that students prefer feedback that allows them to revise their work before final submission.
Boud and Dawson (2023) argued that there has been a lack of investigation into the capabilities university teachers need to possess to give good feedback. In response, they developed a framework of competencies at the macro, meso and micro levels. These competencies include, but are not limited to, strategic planning of feedback, resource management, including managing work pressure, improvement of feedback processes and the application of technological aids. They also emphasize the ability and willingness to attend to different students’ needs. For a detailed description, see Boud and Dawson (2023).
Research reveals that raters can have their own “standards of what constitutes a good text, which influences how they rate text quality and how much weight they apply to certain aspects of writing” (Bouwer et al., 2023, p. 303). Although they recommend using a benchmarking procedure, developing the criteria and procedures can be time-consuming. Moreover, involving students in using these types of benchmarks as part of their own learning process may be rather complex.

1.2.2. Self-Assessment

Various terms are used to describe the process by which students reflect on, give feedback to, and evaluate themselves (Yan & Brown, 2017). Klenowski (1995) defines self-assessment as the ‘evaluation or judgement of “the worth” of one’s performance and the identification of one’s strengths and weaknesses with a view to improving one’s learning outcomes’ (p. 146). Paris and Paris (2001) state that self-assessment involves students’ ‘internalization of standards,’ which in turn enables them to regulate their own learning more effectively. Similarly, Andrade and Valtcheva (2009) understand self-assessment as an activity where students identify areas of strength and weakness in their work to make improvements and promote learning. In this way, students can take an active part in their learning processes in higher education (Chan & Chen, 2023).

1.2.3. Previous Research on AI-Assisted Feedback

A large body of research has investigated various aspects of AI-assisted teaching and learning, and the field is rapidly expanding (Agostini & Picasso, 2024, Bui & Barrot, 2025; Bond et al., 2024; Elmourad et al., 2026; Jansen et al., 2024; Manning et al., 2025; Seo et al., 2025). Kasneci et al. (2023) claim that using LLMs can reduce the time spent on grading and feedback—for example, scoring multiple-choice exams or screening for plagiarism. Several studies have investigated the use of LLMs in the assessment of student essays in higher education courses. Many of these studies have specifically focused on the use of LLMs in evaluations related to foreign language learning (Kostic et al., 2024; Li & Liu, 2024; Meyer et al., 2024; Mizumoto & Eguchi, 2023; Pack et al., 2024). Taken together, these studies suggest that the use of LLMs can offer valuable and time-saving contributions to the assessment process.
Dai et al. (2023) conducted a case study comparing ChatGPT (GPT-3.5 developed by Open AI) generated feedback with instructor-generated feedback. Their findings suggest that ChatGPT gave more detailed feedback compared to the instructors. Additionally, they observed a high degree of agreement between ChatGPT and the instructor. Kostic et al. (2024) compared human and AI-generated assessments of texts written by German language students, and they found that LLMs struggled to evaluate complex texts according to predefined criteria. They noted that LLM-generated assessments lacked the nuanced understanding needed to successfully evaluate students’ essays. Other studies have similarly found that LLMs did not consistently align with evaluations from human assessors (Bui & Barrot, 2025; Manning et al., 2025). Yet, Sun et al. (2024) suggest that AI can play a role in STEM education for pre-service teachers. Some studies have also compared different LLMs: Li and Liu (2024) compared Jess, jWriter and three LLMs (GPT, BERT, and a local Japanese LLM), concluding that GPT outperformed the others in terms of accuracy. Xiao et al. (2024) further explored AI–human collaboration in this context, finding that such collaboration improved human grading skills and elevated novice graders to more advanced levels. Jansen et al. (2024) reported that a higher proportion of students trusted feedback from experts (88%) compared to trusting LLMs (59%). Accordingly, the field of research on interrater reliability between LLMs and human raters is expanding, particularly in the context of foreign language learning. Current findings suggest that AI assistance can be valuable both in reducing time spent on assessment work and in further developing assessment skills.
The present study contributes to the field by comparing assessments conducted by a human grader and two LLMs to the students’ own assessment of their work, which to our knowledge has not been done previously.
The literature distinguishes between cross-prompt models and prompt-specific models used in automated essay scoring (e.g., Yang et al., 2024) depending on whether the model is trained to assess one specific essay or employs a more general approach to essay assessment. In our study, we applied a prompt-agnostic or hybrid approach. The method can be considered prompt-specific in that we gave the model the assessment criteria for that particular assignment and thereby provided specific scoring guidelines. These criteria were derived from selected learning outcomes for the program. Using these criteria allows us to assess whether the students’ academic texts demonstrate the intended knowledge and competencies as outlined in those learning outcomes. At the same time, we employed an LLM that had been trained on a broad range of language data, and thus not trained specifically on essay-scoring. By providing the criteria to the LLM, we ensured that the LLM, the teacher educator and the student teachers evaluated the essays according to the same standards. The criteria are listed in Section 2.

1.2.4. The Present Study

The Norwegian teacher education for primary and lower secondary school is divided into two integrated master’s programs: Years 1–7 and Years 5–10. In the Years 1–7 program, Norwegian, mathematics, and pedagogy are required, with a focus on teaching young children (ages 6–12). For the Years 5–10 program, the emphasis shifts to specialization in various subjects, with pedagogy as the only mandatory component. Applicants to these teacher education programs must meet general university entrance qualifications. Throughout their education, the student teachers work on developing their academic writing skills through compulsory coursework, workshops, and seminars. Teacher educators affiliated with the different courses are expected to assess and give feedback on numerous student essays—an often time-consuming task.
Both students and educators have access to various AI tools, including LLMs (Jansen et al., 2024). The guidelines for their use remain unclear, aside from informing students that submitting AI-generated texts is prohibited. Nonetheless, when used wisely, LLMs can assist students in many aspects of their writing (Agostini & Picasso, 2024). For example, students can ask the LLMs to assess drafts or texts based on specific criteria and prompt the LLM in the process to help improve their texts (Mills et al., 2025; Seo et al., 2025). Similarly, educators can use LLMs to assess and provide feedback on student submissions in accordance with predefined criteria. Research indicates that LLMs hold potential as evaluators (Agostini & Picasso, 2024; Bond et al., 2024; Dai et al., 2023; Jansen et al., 2024; Seo et al., 2025), although their assessments do not always align with those of human assessors (Bui & Barrot, 2025; Kostic et al., 2024; Manning et al., 2025).
Ultimately, in 2026, the educator has the decision-making authority to pass or fail students’ work. If the assistance provided by LLMs cannot be trusted, it holds minimal value. With this as our point of departure, we formulated the following research questions:
RQ1: What characterizes the agreement (interrater reliability) of two LLMs, the teacher educator, and the student teacher when rating student essays?
RQ2: What do the LLMs emphasize in the evaluation of the academic texts compared to the teacher educator and the student teachers?

2. Materials and Methods

To answer the research questions, we used written coursework requirements (academic texts) submitted by first-year student teachers enrolled in the introductory teaching course at a university for teacher education. The project was registered with The Norwegian Agency for Shared Services in Education and Research (SIKT).

2.1. The Participants

We shared information about the project with all first-year students enrolled in the Primary and Lower Secondary Teacher Education (grades 5–10) specialization in mathematics (N = 45). Written consent was obtained from those who wished to participate (41 out of 45). Consent was collected after the students had completed the semester to prevent the students from feeling compelled to participate out of fear of sanctions that might affect their final grade or their relationship with their teachers. A total of 79 texts were assessed, divided between two submissions. The first submission (Text 1) included 38 texts, and the second submission (Text 2) included 41 texts; three of the 41 students did not submit Text 1. The students had a median age of 22 years and had all completed upper secondary school in the Norwegian school system with average or above grades in the Norwegian language, which is a requirement for admission to teacher education. However, all students reported having little to no experience with written self-assessments. Approximately 70% were female and 30% male. The teacher educators involved included one female, age 45, with 17 years of teaching and assessment experience at lower secondary, upper secondary and university levels, and one male, age 57, with 26 years of teaching experience from upper secondary and university levels.

2.2. Pilot Study

After the students had submitted their assignments, including their self-assessment, we decided to do a pilot (see Figure 1). We selected six texts, including self-assessments submitted by first-year student teachers. Based on the evaluations and ratings made by one teacher educator, we selected two very good, two good, and two moderate texts. The texts ranged from 800 to 1000 words and concerned theories of motivation.
Pilot_phase A. A second teacher educator read through the texts and assessed them based on the criteria to get a common baseline.
Pilot_phase B. Developing prompts in the LLM. The first step was to enter the criteria that had been drawn up for the LLM. The next step was to try out various prompts, such as “assess based on criteria”, “rank the texts from best to worst,” and assess academic writing skills”. We found that these prompts did not provide thorough feedback on how the text could be improved. Consequently, we tested prompts that were explicitly asking for suggestions for improvement.
Pilot_phase C. During the pilot, we iteratively refined prompts to identify those that enabled the LLM to perform comprehensive evaluations. evaluation before we concluded on which prompts to include in the main study. After each round of prompting, the researchers reviewed and discussed the LLM outputs to determine necessary adjustments before finalizing the prompts for the main study.
Based on the pilot, the following steps of prompts were included in the main study:
  • Step 1: Can you evaluate an academic text written by a first-year student teacher?
  • Step 2: Thanks! This is the assignment given to the students and the assessment criteria for the text.
  • Step 3: Can you evaluate the text on a scale from 1 to 3, where 1 is low, 2 is middle, and 3 is high?
    • Step 3_1: If there is a discrepancy between the LLMs and the teacher educator or the student teachers in that they have chosen different extremes (low vs. high) of the scale: What is the quality of the academic text? How can the text be improved?
  • Step 4: This is a new text for evaluation. Step 3_1 was used if there was a discrepancy between evaluations.

2.3. Main Study: Essays and Criteria

Our goal was to gain insight into whether LLMs could serve as useful partners in the assessment of academic texts, both for students and educators. We used interrater reliability to determine the degree of alignment between different assessments. Additionally, we compared feedback provided by the various evaluators to compare and evaluate accuracy. In the discussion section, we discuss whether LLMs represent a pedagogically useful tool for supporting students in developing academic writing skills and for teacher educators in their assessment tasks. We analyzed two different essays from the students, both based on selected learning outcomes from the program. These learning outcomes are critical because they specify what the assignments aim to measure. Yet, there are multiple learning goals, and each essay is related to a subset of learning goals. Interpretation of whether students meet the criteria for their essays can be used to discuss aspects of content validity and construct validity (Messick, 1995).
However, an important question in this study is if we can examine what teacher educators, student teachers, and LLMs emphasize when applying assessment criteria for evaluating academic texts.
The first essay (E1) was submitted in the beginning of the semester on the topic of “The teacher’s role”, and the second essay (E2) was submitted at the end of the semester on the topic of Motivation. The criteria were:
  • A clear and precise problem statement.
  • You use relevant course literature (minimum 3 sources). It is not sufficient to use notes from lectures.
  • Clear professional language and good coherence between the different parts of the assignments. Use technical terms.
  • Reflect on the significance it can have for your work with the students.
  • Correct citation and reference list.
  • The essay must be presented in line with research ethic guidelines.
First, the students were asked to provide a self-assessment of their own essays and responded to questions regarding what they felt they had mastered and areas requiring improvement. This was voluntary, and most student teachers submitted a self-assessment together with their academic texts.
Second, the teacher educator evaluated each essay to determine whether the students passed the assignment or if they needed to resubmit it, and the educator provided feedback regarding the quality of the text and areas of improvement. Additionally, each essay was given a score from 1 to 3 (1 indicating low quality, 2 indicated middle, and 3 indicated high), but the students only got information about pass/resubmit (not the numeric score) alongside the qualitative feedback. This assessment was conducted without the use of LLMs.
When preparing data for analysis, the students did not score their essays numerically, but the researchers estimated their scores based on their response to the open-ended questions in the self-assessments. Overall, 38 essays (E1) and 41 essays (E2) were collected and could be used in the further evaluations by the LLMs. We used an LLM service supported and approved by the university, ensuring that all texts were anonymized and contained no personal information. Each researcher operated their own LLM.
All of the analyses were carried out in the same thread/chat. In some cases, the LLMs gave a score of 4, and if this happened repeatedly the prompt about the scale being from 1 to 3 was re-introduced.
The results from the evaluations were analyzed by looking at the agreement and interclass correlations between the LLMs, the teacher educator, and the student teachers.

2.4. Analytical Approaches

To answer RQ1, we analyzed how the teacher educator, the student teachers, and the LLMs evaluated the academic texts. We focus on how these evaluators used criteria operationalized from the learning outcomes. Potential challenges include the operationalization of criteria, as well as differences in interpreting and weighting them. Validity refers to whether the interpretations of scores and related decisions are well founded (Messick, 1995). Ensuring good validity in essay assessment depends on raters’ thorough understanding of the criteria they apply. Both student teachers and teacher educators had access to the task descriptions of the two essays, the assessment criteria, and relevant literature. The LLMs were also provided with the task descriptions and criteria for both essays, in addition the LLMs had access to theories relevant for the essay. It appears rather clear that teacher educators and LLMs actively use these criteria in their feedback; for example, both LLMa and LLMb base their feedback on points in the criteria and include relevant theories. In contrast, student teachers tend to adopt a more precise but concise style, applying the criteria to a limited extent, which may challenge the validity of their evaluations and ratings. There may be a difference in that teacher educators emphasize the discussion aspect in the essay to a greater extent than LLMs and student teachers. Overall, assessments conducted by teacher educators and LLMs are well founded on the criteria, which supports the validity of content and conceptual understanding (Bouwer et al., 2023).
Several measures could be used to examine the consistency of ratings. In this study, we use agreement in percentages, correlation, Cronbach’s alpha, and Fleiss’s kappa. As the latter is sensitive to cases with one dominating rating, we also used the prevalence-adjusted kappa (PBAK) (Byrt et al., 1993).
To answer RQ2, we started off with an inductive approach to the data and performed a descriptive content analysis.
Coding procedure and analytical approach:
Both researchers read through the feedback provided by both LLMs, the teacher educator, and the student teachers. We coded the feedback first according to the assessment criteria, more precisely, which criteria were implicitly or explicitly featured in the feedback. We uploaded the content to an Excel spreadsheet to be able to compare the four different types of assessment, looking for patterns across the dataset. In Section 3.1 we provide some examples to help visualize the differences and similarities in the evaluations. Thereafter, we proceeded with a more deductive approach and coded the feedback according to Hattie and Timperley’s (2007) framework (task, process, self-regulation and self) to identify which feedback level seemed to be salient in the different types of feedback. Regarding step 3_1, when disagreement among the graders was observed, we analyzed the differences in feedback and then prompted the LLMs to assess the quality of the text and suggest improvements. We compared the emphasis in the LLM-generated feedback with that of the educator’s feedback. This step was applied to three essays where the LLMs’ assessments were high but the students’ and teacher educator’s assessments were low, and to five essays where the LLMs’ assessments were high and the students’ and teacher educators’ assessments were middle. We concluded the review of feedback after these eight essays, as no new patterns or insights emerged from the LLMs’ responses.

3. Results

This section contains results related to the two research questions.

3.1. RQ1: Comparing Ratings from LLMs with Teacher Educator and Student Teachers

To answer RQ1, we analyzed how the LLMs evaluated the essays and compared their assessments with those of the teacher educator and the student teachers’ self-assessments.
Table 1 shows the frequency of ratings (1 low, 2 middle, and 3 high), the mean, the standard error, and the median from the evaluation of the texts carried out by LLMs and the teacher educator, along with the self-assessments made by the student teachers. Analysis conducted with IBM SPSS 29.
When comparing the scoring from the LLMs, the teacher educator, and student teachers, it seemed that both LLMa and LLMb were more positive than both the teacher educator and the student teachers. The LLMs assessed a larger proportion of the texts as high level (93% by LLMa and 84% by LLMb) compared to 32% assessed as high level by the teacher educator and 22% assessed as high level by the student teachers. Furthermore, we found differences between the LLMs, and LLMa seemed to be slightly more positive than LLMb (74 vs. 67 high). Finally, the teacher educator had higher ratings of the texts compared with the student teachers (28 vs. 17 high).
To compare the ratings, we analyzed the agreements between the LLMs, the teacher educator, and the student teachers (Table 2). The table also contains measures of the strength of co-variation, namely the correlation (Corr) and Cronbach’s alpha (CA), and the strength of agreement (Cohen’s Kappa). Kappa can be low even though an observed agreement is high if one category is dominating the ratings. The ratings of academic texts from the LLMs are mostly as “high = 3”, a few as middle (2), and none as low (1) (see Table 1). To compensate for skewness in ratings and low levels of kappa, the prevalence-adjusted kappa (PBAK) (Byrt et al., 1993) was calculated for the cases with one dominating rating.
The LLMs had lower levels of agreement with both the teacher educator and the student teachers, and there seemed to be an increased level of agreement between the teacher educator and the student teachers from Text 1 (41.7%) to Text 2 (64.9%). There was strong agreement between the two LLMs regarding coding, an acceptable level of PBAK, and a low level of kappa. Yet the kappa is sensitive to one dominating rating, and the LLMs used high in rating of most texts. The results also show that there was a lack of co-variation and a lack of strength of agreement beyond what the LLMs agreed on (see example in Table 3).
As can be seen from this example (Table 3), the two LLMs comment on different aspects of the texts. LLM1 comments on language and understandability, while LLM2 calls for deeper analysis. It seems that LLM1 in this case is more in line with the criteria concerning academic tone and language, while LLM2 emphasizes the criteria that deal with practical implications for teachers.

3.2. RQ2: Evaluation of the Academic Texts

RQ2: What do the LLMs emphasize in the evaluation of the academic texts compared to the teacher educator and the student teachers?
As mentioned, we systematically reviewed and coded the written evaluations provided by the LLMs, the teacher educator, and the student teachers to compare the main topics and arguments, identifying similarities and differences across the assessments. Examples are presented below to illustrate these points. We begin by highlighting what the LLMs typically emphasized in their feedback.
Both LLMa and LLMb consistently provided longer and more systematic assessment texts compared to both the teacher educator and the student teachers. The LLMs’ evaluations aligned closely with the topics outlined in the assessment criteria; however, they were less precise in their analysis of the use of literature within the discussion section.
Summary of feedback from LLM. See Appendix A for detailed feedback: The student presents a well-justified problem statement and provides a clear overview of the assignment. Key concepts are well-defined and supported by relevant literature, demonstrating a solid understanding of the teacher’s role. The discussion is structured and coherent, with strong arguments and relevant examples. The conclusion effectively summarizes the main points and reflects on the implications for the student’s future role as a teacher. The bibliography is complete and correctly formatted, and the use of subheadings helps structure the assignment clearly. The academic language is clear, and there is good coherence throughout the text. The assignment is academically sound and ethically responsible. However, the discussion could benefit from more depth regarding the impact of societal changes on the teacher’s role. Overall, the text meets all the requirements and demonstrates a strong understanding of the subject, earning a high rating of 3.
The LLMs addressed the central points outlined in the criteria when assessing the essays, but they encountered difficulties with both the formal and conceptual aspects. Formally, they struggled to accurately identify the number of correct references and to distinguish genuine references from incorrect ones. Conceptually, the LLMs had difficulty evaluating whether the literature was discussed in an academically sound manner. Additionally, the LLMs appeared to weigh all criteria equally, whereas the teacher educator and student teachers placed greater emphasis on certain aspects as more critical to producing quality academic writing.
Regarding the feedback from the teacher educator, it did not include a written evaluation explicitly addressing each criterion. Instead, the feedback focused on the extent to which the assignment was correctly understood, highlighting areas that required improvement as well as those that were well presented, although this was not always consistent. The teacher educator’s evaluations frequently emphasized the students’ understanding of literature and their ability to incorporate it into their discussions. Additionally, some feedback included motivational comments, such as “you impress me”. The teacher educator’s feedback must therefore be understood as a combination of feedback on the task level, process level, and self-level (Hattie & Timperley, 2007).
This is an example of teacher feedback: You are nearly there. I see from your self-assessment that you find it challenging to know what should go into the clarification of terms and what belongs in the discussion. In this text, I think it would have been useful to first differentiate between traits and competencies. The course literature is not entirely clear on this (the terms are somewhat interchangeable), so I completely understand why you might not have considered it. Next, you could explain what the different competencies entail in the explanation and then discuss them in the discussion section. You end up with less space for discussion because you are explaining as you go along. These are things you should think about for your next assignment. Also note that you need to include page numbers for the references in the text.
Finally, the students provided the shortest feedback, typically consisting of one or two brief comments that highlighted one or two specific points. They often emphasized challenges related to organizing their writing (ranging from initiating the writing process to using the APA style for reference formatting). They also reported difficulties integrating research literature into their discussions and conclusions. Thus, their self-assessments reflected feedback on both the process level and the self-regulation level (Hattie & Timperley, 2007). This is a representative example of self-assessment: I feel that I to some extent manage to use the course literature in a meaningful way and write about my own experiences. I need to start writing my assignments at an earlier stage and understand where to stop clarifying concepts and where to start my discussion. I hope we can have writing workshops on campus to help me get started with my assignments.
Regarding what the LLMs emphasize, the excerpts presented below are representative examples illustrating their focus on (a) theories, (b) analyses (including deeper and more critical examination), (c) discussions (emphasizing balanced and critical reflections), and (d) practical implications.
(a)
Theories.
The LLMs commented on the need for more literature or citations. There were also examples of the LLMs advising the students to develop their academic texts through the connections between theories, for example: You can try to draw more connections between the different theories. How can the various theories complement or contradict each other? Is it possible that a mix of approaches could be most effective in this specific case?
(b)
Analyses.
The LLMs suggested further work with the analysis, for example: Although the text does a good job of explaining each motivation theory and how they can be applied to the case, you can take it a step further by analyzing and discussing the potential strengths and weaknesses of each theory in relation to the case. This will provide a more critical analysis and can reveal insights that might not emerge by merely describing the application of the theories.
(c)
Discussions.
The LLMs provided feedback about the discussions in the academic texts and about developing a more balanced discussion, for example: Try to give each theory roughly equal space in the discussion. In this text, there is a bit more focus on the behaviorist perspective, while the humanistic and cognitive perspectives could receive more attention.
The LLMs also suggested more critical reflections in the discussions, for example: Finally, you can spend more time on critical reflection regarding the theories. Are there aspects of the theories that do not fit with what you have observed in practice? Are there limitations in the theories that should be considered? This type of critical reflection can enrich the discussion and demonstrate a deeper understanding of the material.
(d)
Implications.
The LLMs suggested improvements on the implications of the analysis and discussion, for example: You can more clearly discuss what the insights from the motivation theories mean for educational practice. How can teachers use this knowledge in the classroom? How can the insights be translated into concrete educational strategies?

4. Discussion

The first research question of this paper examined the agreement of two LLMs, one teacher educator, and the student teacher who wrote the text. The results revealed a relatively strong agreement between the two LLMs (80%) but a low Cronbach’s alpha (0.23).
This result could be partly due to the use of a three-point scale, where agreement by chance is likely to occur. If a more nuanced scale had been employed, the results might have differed. The low Cronbach’s alpha might also reflect poor consistency between the two LLMs. While some literature suggests that LLMs can contribute to providing consistent evaluation, other research indicates that different LLMs can give different assessments (Pack et al., 2024). However, it is surprising that we got such results when we used the same LLMs and the same prompts and performed the assessments at the same time.
Further, the LLMs were more positive in their ratings of the academic texts compared to the teacher educator and the student teachers. The variation between the LLMs and the human assessors is to some extent in line with recent research (Bui & Barrot, 2025; Kostic et al., 2024; Manning et al., 2025). Further, the agreement between the teacher educator and the student teachers was more consistent compared to their agreements with the LLMs. This difference may be explained by the student teachers’ prior knowledge of what the teacher educator valued and disvalued, as they had received instructions in advance of writing the essays and had received formative assessments on essays handed in at earlier stages in the course.
RQ2: What do the LLMs emphasize in the evaluation of the academic texts compared to the teacher educator and the student teachers?
Regarding the differences between the LLMs, the teacher educator, and the student teachers (RQ2), we found that the LLMs produced longer and more comprehensive evaluation texts. This is in line with findings from Dai et al. (2023) who found that LLMs can provide longer and more detailed comments compared to teachers/instructors. In our study, the evaluations from the LLMs followed the main points of the assessment criteria and included comments under each main point. Such feedback could be advantageous for the student, given that the evaluation is perceived as specific, detailed, and thorough (cf. Dawson et al., 2019, Boud & Dawson, 2023). This type of response was much more detailed compared to the teacher educator and the students’ own responses, and was typically on the process level.
The literature is inconclusive regarding the need for individualized and personalized feedback. Dawson et al. (2019) point out varying interpretations of what constitutes personalized and individualized feedback but emphasize that students generally prefer feedback directed specifically at their own texts rather than generic comments. Boud and Dawson’s (2023) framework underscores the importance of feedback that recognizes and responds to the students’ different needs. The teacher will, in most cases, have access to more information regarding the students’ individual needs, such as prior knowledge and background, specific learning needs (e.g., dyslexia), challenging home situations, and the students’ own preferences for feedback modes based on previous experiences. This contextual understanding enables teachers to provide more tailored feedback that considers the student’s broader lifeworld, something that current LLMs cannot fully replicate.
Another challenge with LLM-generated assessments in essay writing is the risk of students receiving disproportionately positive evaluations that can foster unrealistic expectations, leading to disappointment and confusion if they receive more critical feedback from their teachers. Finally, as noted above, LLM feedback can be quite comprehensive, which might overwhelm students who are still developing their academic skills, potentially undermining their motivation to take the necessary next steps.
Across the feedback on these texts, we found that the LLMs tended to emphasize areas such as theory, data analysis, discussion, and implications. Yet, it appears that the LLMs consistently rated the texts as better than what students or teacher educators did, and with longer and more detailed feedback. Regardless of this, the feedback seemed generic and seldom fine-tuned. This is in line with research (Bui & Barrot, 2025; Kostic et al., 2024; Manning et al., 2025). For example, Kostic et al. (2024) show that it can be difficult for LLMs to evaluate complex texts. This may be an explanation for why the LLMs did not propose extensive changes to individual parts of the texts. Thus, even if the student chooses to improve the text in accordance with the feedback provided by the LLM, it is uncertain whether the suggested improvements are sufficient.
Another important aspect of feedback discussed in the literature is timeliness (Dawson et al., 2019), emphasizing that feedback is most effective when delivered at a stage in the process where the student has a genuine opportunity to make necessary adjustments (Gamlem & Smith, 2013). LLMs offer the advantage of providing feedback instantly at any point in the student’s writing process, unlike waiting for potentially delayed responses from busy teachers. However, this benefit depends on the student’s ability to recognize the appropriate moments to seek feedback, which involves both process understanding and self-regulation skills.
Moreover, LLMs can function as tutors by providing explanations and offering real-time feedback. Research shows positive outcomes in disciplines where the correctness of answers is more clearly defined, such as foreign language education (Kostic et al., 2024; Li & Liu, 2024; Meyer et al., 2024; Pack et al., 2024) and STEM fields (Sun et al., 2024). Additionally, LLMs have the potential to provide feedback that promotes critical thinking and encourages consideration of ethical aspects in academic writing (Agostini & Picasso, 2024).
On the other hand, LLMs can support educators in providing students with detailed and comprehensive feedback in line with research on good feedback practices (García-Varela et al., 2025; Wisniewski et al., 2020), particularly at the task and process levels. Further, teacher educators can calibrate LLM outputs to align with their own standards, although there is no guarantee that the teachers’ assessments are more correct, true, or accurate compared to the LLM. Nonetheless, this AI–human collaboration can be beneficial for both educators and students (Xiao et al., 2024). Further, the teacher educator can facilitate prompt-engineering classes or, in collaboration with the students, develop prompt cheat-sheets tailored for different stages of the writing process. In this way, students could use the LLM as scaffolding in their learning process.

4.1. Limitations

This study has some limitations that should be acknowledged. The number of texts was limited, and a larger sample would increase the generalizability of the findings; however, there were still significant variations within the current sample. The texts were evaluated by a single experienced teacher educator have led to bias in the ratings (e.g., disproportionately strict or permissive). The reported agreement reflects alignment between LLMs and one experienced teacher’s judgments (not with a group of human raters). There were two human raters in the pilot study to investigate consistency and inter-coder reliability, which yielded promising findings, but this was not continued in the main study.
Furthermore, the students’ self-assessment was voluntary, and they were aware that these self-assessments would not be graded as part of the assignment. This may have led some students to invest minimal effort in evaluating their own work, which could explain why their self-assessment texts were significantly shorter than the feedback provided by the teacher educator and the LLMs. A more structured and thorough self-assessment process might yield deeper insights into the students’ reflections. It is also possible that the students, all being first-year, lacked experience with self-assessment and were uncertain of the expectations of higher education. Moreover, the students’ scores were inferred indirectly by the researchers, potentially reducing the reliability. Therefore, student ratings should be interpreted cautiously and as indicative rather than definitive representations of their evaluative judgments.
Furthermore, the use of a limited three-point rating scale may increase the likelihood of coincidental agreement, potentially inflating interrater reliability coefficients. Therefore, caution is necessary when drawing conclusions about the meaning of findings derived from such comparisons.
In our study, we observed that the same LLM, operated independently by two researchers at the same time, provided differing feedback and ratings. While we do not know whether these results would differ if different LLMs had been used, as was the case in Li and Liu’s (2024) study, we cannot rule out the possibility that a higher-quality LLM might provide more consistent and higher-quality feedback than a lower-quality model.
Overall, these limitations suggest that the findings on interrater reliability should be viewed as preliminary evidence of agreement under limited conditions, rather than as strong claims about the validity assessments by humans and LLMs.

4.2. Practical Implications

LLMs can request improved argumentation and application of relevant theory; however, student teachers often lack access to model texts, making them reliant on prompts derived from task descriptions and assessment criteria. While the LLMs can support students in this regard, students must disclose their use of such tools in their submissions.
LLMs can generate extensive evaluations based on established criteria, but their limited focus on discussion complicates their use without developing tailored prompts. Educators must still manually review texts, which can potentially double their workload and reduce time efficiency.
Low and mid-level student texts may receive overly positive feedback from LLMs, which can lead to unrealistic expectations. Students need to critically evaluate the reliability of such feedback and refine their prompts to better align with the educator’s standards. While this process can foster learning and critical thinking. Such dynamics challenge the validity of assessment outcomes, particularly if top performers succeed primarily by mastering prompt engineering rather than academic writing. Access to different LLMs varies because higher-quality models often require paid subscriptions, while lower-quality or free options may offer limited features or accuracy. This variation in both quality and cost impacts students differently based on their financial resources.

5. Conclusions

By comparing the evaluations made by LLMs, the teacher educator, and the student teachers, we contribute to the expanding field of AI-supported assessment. We found agreement regarding the high-level ratings of the texts by the LLMs. Furthermore, some agreement was observed between the student teachers and the teacher educator; however, the LLMs did not demonstrate agreement with either the students or the teacher educator. These findings highlight the ongoing need to improve the evaluation capabilities of LLMs.
Second, we examined the characteristics of feedback provided by the LLMs, the students, and the teacher educator. The LLMs provided much longer feedback compared to that from the students, while the teacher educator’s feedback was longer than that of the students but shorter than that of the LLMs. The students typically provided a few sentences, whereas the teacher educator wrote one or two paragraphs. The LLMs’ evaluation included several paragraphs in length and explicitly followed the criteria. However, the teacher educator seemed to be selecting and using only some of the criteria in the written feedback, with particular emphasis on improving the integration of theory and research within the discussion section.
Overall, LLMs provided students with specific feedback on how to improve their academic texts. This feedback was generally more comprehensive than that offered by the teacher educator or the students themselves and was typically focused on the process level. Nevertheless, the LMMs’ feedback was not necessarily in alignment with that of the teacher educator or the aspects that were emphasized when evaluating the student teachers’ texts. Yet, there is reason to believe that LLMs can be valuable tools both for students developing academic writing skills and for educators who invest significant time in assessment, especially when used in a collaborative mode combining human and machine input. Based on this study and recent research, we are not certain that LLMs’ feedback helps to improve the academic texts in the right direction, and thus, further research is needed into how students and teachers can use and collaborate with LLMs in the writing and assessment of academic texts.

Author Contributions

Conceptualization, A.G. and O.E.H.; methodology, A.G. and O.E.H.; software, O.E.H.; validation, A.G. and O.E.H.; formal analysis, A.G. and O.E.H.; investigation, A.G. and O.E.H.; resources, A.G. and O.E.H.; data curation, A.G. and O.E.H.; writing—original draft preparation, A.G. and O.E.H.; writing—review and editing, A.G. and O.E.H.; visualization, A.G. and O.E.H.; project administration, A.G. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Norwegian Agency for Shared Services in Education and Research (SIKT) (protocol code: 474695; date of approval: 8 June 2025).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

Data cannot be made available.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
LLMLarge Language Models
AIArtificial Intelligence
TETeacher Educator
STStudent Teacher

Appendix A

Some of the differences in evaluations undertaken by the student, the teacher educator, and the LLM can be illustrated in two examples:
  • Example A
One student was evaluated as low (1) by the teacher educator and by himself, but as high (3) by both LLMs.
The teacher educator approved the academic text with some reservations due to the variation in language throughout the text. It seemed like the text was written by at least two authors. Overall, the discussion was rated by the teacher educator as better than the previous text the student teacher had handed in, but the student was asked to “to support your arguments in the discussion with course literature”.
The student perceived that he was better at writing his way through the text without being disturbed in his work. But the student still thought that it was difficult to find and use current literature as part of his academic writing.
LLMa focused on a deep understanding and effective presentation of the topic, but did not comment on the bibliography, while LLMb focused on the text being an effective presentation of the topic and reported that the text had a formal tone with correct bibliography.
In example A, the teacher educator had doubts as to whether the text was entirely written by the student himself or if he had had some kind of assistance. The student reported that he had used ChatGPT for language editing and to correct spelling mistakes. The LLMs evaluated this text as very good due to its formal tone, correct bibliography, and effective presentation, but they did not assess the quality of the discussion or the use of appropriate literature to support the arguments in the text. The student’s assessment was to a larger degree concerning matters of self-regulation than matters regarding academic writing. These aspects were related to the student’s work habits, which neither the teacher educator nor the LLMs had any knowledge of.
  • Example B
Another student was evaluated as low (1) by himself, middle (2) by the teacher educator, and high (3) by both LLMs.
In example B, the teacher educator praised the student teacher’s progress and the use of relevant literature but called for a more coherent text and to include critical perspectives. The LLMs, in contrast, reported that the text was well written and that the student had “reflected deeply on the subject”.
The evaluation from the student was more focused on how to work further and improve their academic texts. The student stated, “I need to improve my academic writing and the discussion”.
When we found discrepancies between the evaluations from the LLMs and the students, we analyzed what the LLMs suggested as appropriate feedback to the students to improve their academic texts.
We prompted the LLMs to give feedback about the quality of the answers as academic texts. The LLMs emphasized structural aspects such as clear paragraphs, well-formed sentences, and appropriate literature use and included comments on the use of academic language and the understanding of subject terminologies.

References

  1. Agostini, D., & Picasso, F. (2024). Large language models for sustainable assessment and feedback in higher education. Intelligenza Artificiale, 18(1), 121–138. [Google Scholar] [CrossRef] [Scilit]
  2. Andrade, H., & Valtcheva, A. (2009). Promoting learning and achievement through self-assessment. Theory into Practice, 48(1), 12–19. [Google Scholar] [CrossRef] [Scilit]
  3. Black, P. J., Harrison, C., Lee, C., Marshall, B., & Wiliam, D. (2004). Working inside the black box: Assessment for learning in the classroom. Phi Delta Kappan, 86(1), 8–21. [Google Scholar] [CrossRef] [Scilit]
  4. Black, P. J., & Wiliam, D. (1998). Inside the black box: Raising standards through classroom assessment. Phi Delta Kappan, 80(2), 139. [Google Scholar] [CrossRef] [Scilit]
  5. Bond, M., Khosravi, H., De Laat, M., Bergdahl, N., Negrea, V., Oxley, E., Pham, P., Chong, S. W., & Siemens, G. (2024). A meta systematic review of artificial intelligence in higher education: A call for increased ethics, collaboration, and rigour. International Journal of Educational Technology in Higher Education, 21(1), 4. [Google Scholar] [CrossRef] [Scilit]
  6. Boud, D., & Dawson, P. (2023). What feedback literate teachers do: An empirically-derived competency framework. Assessment & Evaluation in Higher Education, 48(2), 158–171. [Google Scholar]
  7. Bouwer, R., Koster, M., & Van den Bergh, H. (2023). Benchmark rating procedure, best of both worlds? Comparing procedures to rate text quality in a reliable and valid manner. Assessment in Education: Principles, Policy & Practice, 30(3–4), 302–319. [Google Scholar] [CrossRef] [Scilit]
  8. Bui, N. M., & Barrot, J. S. (2025). ChatGPT as an automated essay scoring tool in the writing classrooms: How it compares with human scoring. Education and Information Technologies, 30, 2041–2058. [Google Scholar] [CrossRef] [Scilit]
  9. Byrt, T., Bishop, J., & Carlin, J. B. (1993). Bias, prevalence and kappa. Journal of Clinical Epidemiology, 46(5), 423–429. [Google Scholar] [CrossRef] [Scilit]
  10. Chan, C. K. Y., & Chen, S. W. (2023). Student partnership in assessment in higher education: A systematic review. Assessment & Evaluation in Higher Education, 48(8), 1402–1414. [Google Scholar] [CrossRef] [Scilit]
  11. Dai, W., Lin, J., Jin, H., Li, T., Tsai, Y.-S., Gašević, D., & Chen, G. (2023). Can large language models provide feedback to students? A case study on ChatGPT. In 2023 IEEE international conference on advanced learning technologies (ICALT) (pp. 323–325). IEEE. [Google Scholar]
  12. Dawson, P., Henderson, M., Mahoney, P., Phillips, M., Ryan, T., Boud, D., & Molloy, E. (2019). What makes for effective feedback: Staff and student perspectives. Assessment & Evaluation in Higher Education, 44(1), 25–36. [Google Scholar]
  13. Elmourad, T., Hadjiphanis, L., Christofi, K., Chourides, P., & Kythreotis, A. (2026). AI adoption in K–12 education: A model of skills transformation, productivity, and institutional readiness. Education Sciences, 16(2), 337. [Google Scholar] [CrossRef] [Scilit]
  14. Gamlem, S. M., & Smith, K. (2013). Student perceptions of classroom feedback. Assessment in Education: Principles, Policy & Practice, 20(2), 150–169. [Google Scholar] [CrossRef] [Scilit]
  15. García-Varela, F., Nussbaum, M., Mendoza, M., Martínez-Troncoso, C., & Bekerman, Z. (2025). ChatGPT as a stable and fair tool for automated essay scoring. Education Sciences, 15(8), 946. [Google Scholar] [CrossRef] [Scilit]
  16. Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. [Google Scholar] [CrossRef] [Scilit]
  17. Jansen, T., Höft, L., Bahr, L., Fleckenstein, J., Møller, J., Köller, O., & Meyer, J. (2024). Comparing generative AI and expert feedback to students’ writing: Insights from student teachers. Psychologie in Erziehung und Unterricht, 71(2), 80–92. [Google Scholar] [CrossRef] [Scilit]
  18. Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. [Google Scholar] [CrossRef] [Scilit]
  19. Klenowski, V. (1995). Student self-evaluation processes in student-centred teaching and learning contexts of Australia and England. Assessment in Education: Principles, Policy & Practice, 2(2), 145–163. [Google Scholar]
  20. Kostic, M., Witschel, H. F., Hinkelmann, K., & Spahic-Bogdanovic, M. (2024). LLMs in automated essay evaluation: A case study. In Proceedings of the AAAI symposium series (Vol. 3, pp. 143–147). IEEE. [Google Scholar]
  21. Li, W., & Liu, H. (2024). Applying large language models for automated essay scoring for non-native Japanese. Humanities and Social Sciences Communications, 11(1), 723. [Google Scholar] [CrossRef] [Scilit]
  22. Lindsay, E., Rodda, A., Lindqvist, A. L., Quince, Z., Lim, M., & Jiang, D. (2026). The dimensions of abundance in AI-generated feedback. Education Sciences, 16(3), 465. [Google Scholar] [CrossRef] [Scilit]
  23. Manning, J., Baldwin, J., & Powell, N. (2025). Human versus machine: The effectiveness of ChatGPT in automated essay scoring. Innovations in Education and Teaching International, 62, 1500–1513. [Google Scholar] [CrossRef] [Scilit]
  24. Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741–749. [Google Scholar] [CrossRef]
  25. Meyer, J., Jansen, T., Schiller, R., Liebenow, L. W., Steinbach, M., Horbach, A., & Fleckenstein, J. (2024). Using LLMs to bring evidence-based feedback into the classroom: AI-generated feedback increases secondary students’ text revision, motivation, and positive emotions. Computers and Education: Artificial Intelligence, 6, 100199. [Google Scholar] [CrossRef] [Scilit]
  26. Mills, E., Mizouri, A., & Peach, A. (2025). Prompting better feedback: A study of custom GPT for formative assessment in undergraduate physics. Education Sciences, 15(8), 1058. [Google Scholar] [CrossRef] [Scilit]
  27. Mizumoto, A., & Eguchi, M. (2023). Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics, 2(2), 100050. [Google Scholar] [CrossRef] [Scilit]
  28. Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education: Artificial Intelligence, 6, 100234. [Google Scholar] [CrossRef] [Scilit]
  29. Paris, S. G., & Paris, A. H. (2001). Classroom applications of research on self-regulated learning. Educational Psychologist, 36(2), 89–101. [Google Scholar] [CrossRef] [Scilit]
  30. Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Instructional Science, 18(2), 119–144. [Google Scholar] [CrossRef] [Scilit]
  31. Seo, H., Hwang, T., Jung, J., Kang, H., Namgoong, H., Lee, Y. H., & Jung, S. (2025). Large language models as evaluators in education: Verification of feedback consistency and accuracy. Applied Sciences, 15(2), 671. [Google Scholar] [CrossRef] [Scilit]
  32. Sharma, R., Jain, A., Gupta, N., Garg, S., Batta, M., & Dhir, S. K. (2016). Impact of self-assessment by students on their learning. International Journal of Applied and Basic Medical Research, 6(3), 226–229. [Google Scholar] [CrossRef] [Scilit]
  33. Shermis, M. D., Koch, C. M., Page, E. B., Keith, T. Z., & Harrington, S. (2002). Trait ratings for automated essay grading. Educational and Psychological Measurement, 62(1), 5–18. [Google Scholar] [CrossRef] [Scilit]
  34. Sun, F., Tian, P., Sun, D., Fan, Y., & Yang, Y. (2024). Pre-service teachers’ inclination to integrate AI into STEM education: Analysis of influencing factors. British Journal of Educational Technology, 55(6), 2574–2596. [Google Scholar] [CrossRef] [Scilit]
  35. Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 487662. [Google Scholar] [CrossRef] [Scilit]
  36. Xiao, C., Ma, W., Xu, S. X., Zhang, K., Wang, Y., & Fu, Q. (2024). From automation to augmentation: Large language models elevating essay scoring landscape. arXiv, arXiv:2401.06431. [Google Scholar]
  37. Yan, Z., & Brown, G. T. (2017). A cyclical self-assessment process: Towards a model of how students engage in self-assessment. Assessment & Evaluation in Higher Education, 42(8), 1247–1262. [Google Scholar]
  38. Yang, K., Raković, M., Li, Y., Guan, Q., Gašević, D., & Chen, G. (2024). Unveiling the tapestry of automated essay scoring: A comprehensive investigation of accuracy, fairness, and generalizability. In Proceedings of the AAAI conference on artificial intelligence (Vol. 38, pp. 22466–22474). IEEE. [Google Scholar]
Figure 1. Illustrating the process of data and data analysis in the study.
Figure 1. Illustrating the process of data and data analysis in the study.
Education 16 00787 g001
Table 1. Frequencies.
Table 1. Frequencies.
Groups123Mean (se)MedianOther Ratings or Missing
Both texts (n = 79)
LLMa03742.96 (0.03)3One 2.5 and one 4
LLMb06672.94 (0.04)3Three 2.5 and three 4
TE446282.31 (0.06)2One 2.5
ST945172.11 (0.07)2One 1.5, one 2.5 and six missing
Text 1 (n = 38)
LLMa00373.03 (0.03)3One 4
LLMb00333.05 (0.05)3Two 2.5 and three 4
TE222142.32 (0.09)2
ST72082.01 (0.11)2One 1.5 and two missing
Text 2 (n = 41)
LLMa03372.91 (0.04)3One 2.5
LLMb06342.84 (0.06)3One 2.5
TE224142.30 (0.09)2One 2.5
ST22592.20 (0.08)2One 2.5 and four missing
Table 2. Comparing the agreement, the correlation between the rating, Cronbach’s alpha (CA), and Cohen’s kappa of the evaluations between LLMs, teacher educator (TE), and student teacher (ST). The prevalence-adjusted kappa (PBAK) was calculated for LLMa vs. LLMb.
Table 2. Comparing the agreement, the correlation between the rating, Cronbach’s alpha (CA), and Cohen’s kappa of the evaluations between LLMs, teacher educator (TE), and student teacher (ST). The prevalence-adjusted kappa (PBAK) was calculated for LLMa vs. LLMb.
GroupsAgreementCorrelationCAKappaPBAK
All texts (n = 79)
LLMa vs. LLMb80.0%0.140.230.060.60
LLMa vs. TE36.7%0.130.130.02
LLMa vs. ST24.7%−0.07−0.120.00
LLMb vs. TE32.9%0.060.09−0.03
LLMb vs. ST26.0%−0.13−0.090.02
ST vs. TE52.1%0.220.190.16
Text 1 (n = 38)
LLMa vs. LLMb84.2%−0.03−0.01−0.040.68
LLMa vs. TE34.2%0.200.21−0.03
LLMa vs. ST22.2%0.000.000.00
LLMb vs. TE31.6%0.290.32−0.01
LLMb vs. ST22.2%−0.08−0.010.03
ST vs. TE41.7%0.100.070.02
Text 2 (n = 41)
LLMa vs. LLMb78.0%0.140.230.090.56
LLMa vs. TE39.0%0.090.110.06
LLMa vs. ST27.0%−0.06−0.140.00
LLMb vs. TE34.1%−0.12−0.10−0.04
LLMb vs. ST27.0%−0.10−0.07−0.02
ST vs. TE64.9%0.390.370.32
Table 3. Example of how the two LLMs evaluated the same text differently.
Table 3. Example of how the two LLMs evaluated the same text differently.
LLM1LLM2
I would grade this to a 3. It is a good balance between the use of every-day language and professional terminology, which contributes to making the text understandable and academic.I would grade this to a 2. The discussion could have included a deeper analysis and a critical evaluation of the theories. I addition the student could have given a more detailed discussion of the practical implications for teaching.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Gillespie, A.; Hatlevik, O.E. Seeking Help from Large Language Models: Exploring the Assessment and Feedback of Student Teachers’ Academic Texts. Educ. Sci. 2026, 16, 787. https://doi.org/10.3390/educsci16050787

AMA Style

Gillespie A, Hatlevik OE. Seeking Help from Large Language Models: Exploring the Assessment and Feedback of Student Teachers’ Academic Texts. Education Sciences. 2026; 16(5):787. https://doi.org/10.3390/educsci16050787

Chicago/Turabian Style

Gillespie, Astrid, and Ove Edvard Hatlevik. 2026. "Seeking Help from Large Language Models: Exploring the Assessment and Feedback of Student Teachers’ Academic Texts" Education Sciences 16, no. 5: 787. https://doi.org/10.3390/educsci16050787

APA Style

Gillespie, A., & Hatlevik, O. E. (2026). Seeking Help from Large Language Models: Exploring the Assessment and Feedback of Student Teachers’ Academic Texts. Education Sciences, 16(5), 787. https://doi.org/10.3390/educsci16050787

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop