Next Article in Journal
The Institutional Plateau in Microcredential Adoption: A Hybrid Review of Peer-Reviewed Evidence and Practitioner Reports in Higher Education
Next Article in Special Issue
Evaluating Competency Transfer from Factory I/O to Real PLC Systems in Industry 4.0 Engineering Education
Previous Article in Journal
Twice-Exceptional University Students: Perceptions of Effective Support
Previous Article in Special Issue
Early Childhood Educators’ AI Literacy: Validation of the Meta-AI Literacy Scale in a Greek Context
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Assessing Linguistic and Content Characteristics of First-Grade Students’ Texts Using Generative AI

1
Chair of Primary School Education, Julius-Maximilians-University Würzburg, 97070 Würzburg, Germany
2
College of Engineering and Information Technology, Adelaide University, Adelaide, SA 5005, Australia
*
Author to whom correspondence should be addressed.
Educ. Sci. 2026, 16(8), 1227; https://doi.org/10.3390/educsci16081227
Submission received: 16 May 2026 / Revised: 19 July 2026 / Accepted: 28 July 2026 / Published: 3 August 2026

Abstract

The present study examines how accurately a common large language model (ChatGPT 5.0) assesses linguistic and content characteristics of texts written by first-grade students. The focus is on the consistency of LLM assessments and their alignment with human assessments. The database consists of N = 539 texts written by elementary school students. The results show that, contrary to our expectations, LLM assessments of linguistic text characteristics are not generally more consistent than LLM assessments of content text characteristics. Furthermore, alignment between LLM and human assessments is higher for the assessment of single linguistic characteristics (such as the use of cohesive devices) than for the assessment of single content characteristics (such as originality), but not in general. Thus, LLM assessments of linguistic text characteristics in early elementary school are not generally more accurate than LLM assessments of content text characteristics. Based on these findings, we discuss implications for using LLMs as assessment tools in early stages of schooling.

1. Introduction

Writing is a key cultural skill that is essential for successful social participation (Trapman et al., 2018). Since writing skills develop in the early years of elementary school (Rohloff et al., 2023), their promotion is among the core tasks of elementary school teachers. The accurate assessment of students’ written products plays a key role in this. In a current model of writing research (revised writers-within-community model of writing), Graham (2018) postulates that writing development does not depend exclusively on cognitive factors. As writing is a social activity that is situated within specific contexts (writing communities), the exchange with interaction partners such as teachers is also important. Teachers’ assessments of written products and the resulting feedback are therefore influential for students’ writing development (see also Deane, 2018).
However, the individual assessment of student texts is time-consuming and complex. Teachers report that they frequently lack adequate preparation for this task (Brindle et al., 2016; Parr & Jesson, 2016). Consequently, the question arises of how teachers can be supported in assessing student texts in everyday school life. Given current technological developments in society and education (Chiu et al., 2023), technology-based solutions appear promising in this regard. Automated applications for text assessment—known as Automated Essay Scoring Systems (AES)—are considered to have particular potential here.
AES are computer-based systems designed for the automated assessment of written texts (Ramesh & Sanampudi, 2022). To date, a wide range of AES has been developed: Ke and Ng (2019) distinguish between three dimensions that may be used to systematize current AES (the task for which the application was developed, its approach, and the features it offers). The state of research on AES is also extensive and shows that the assessment of texts by AES is largely consistent with the assessments of human raters (Huawei & Aryadoust, 2023; Hussein et al., 2019), even when assessing student texts in elementary education (Wilson & Huang, 2024). Nevertheless, there is also criticism of AES. For instance, the development of AES is considered to be time-consuming and resource-intensive (Oğuz, 2025). In addition, AES are strongly tied to the writing tasks for which they were developed. Thus, an AES is not suitable to assess different types of texts written by students across school lessons. Instead, it is only able to assess the texts for the evaluation of which it was developed. AES are therefore of limited use in school practice (Xu et al., 2024).
An alternative to AES that has potential to address these criticisms is offered by generative AI and, in particular, large language models (LLMs). LLMs are statistical, pre-trained language models that generate texts by calculating probabilities about word sequences: texts are created by linking words that are most likely to follow each other (Minaee et al., 2024). One of the most prominent and powerful LLMs is ChatGPT. The GPT series is based on the principle of LLM decoder-only architecture. That is, the LLM predicts the next word in a sentence based on previous word(s) to reconstruct the pre-training data, which consists of hundreds of billions of parameters (Zhao et al., 2026). LLMs are therefore widely established tools for automated text generation (Shi et al., 2026).
At the same time, LLMs are a potentially promising aid for assessing texts (Bucol & Sangkawong, 2025). Several authors consider LLMs to be used as assessment tools and describe potential advantages of LLMs for evaluating texts in educational contexts: Kasneci et al. (2023), for instance, state that LLMs are quite easy to implement in school settings and can be used flexibly for different assessment tasks. Guo and Wang (2024) report that LLMs can provide balanced feedback on different text characteristics such as content or language. At the same time, various studies have focused on the accuracy of LLM assessments of student texts. Two indicators are often used to measure the accuracy of LLM assessments: (1) the intra-rater reliability of different LLM ratings, i.e., the consistency of AI-generated assessments; (2) the inter-rater reliability of LLM and human ratings, meaning the alignment between LLM and human assessments. For both indicators, research findings paint a mixed picture. The intra-rater reliability of LLM assessments is shown to be high in some studies (Altamimi, 2023), even higher than the intra-rater reliability of human raters (Tate et al., 2024). Oğuz (2025), for example, found evidence that ChatGPT, Gemini, and DeepSeek assess sixth- to 12th-grade students’ essays with excellent intra-rater reliability, with an intra-class correlation ranging between 0.753 and 0.807. Closed-source models such as ChatGPT are shown to be particularly promising in this regard (Seßler et al., 2025). There is also some evidence that high-quality texts are assessed with higher consistency than low-quality texts (Li et al., 2024). However, in other studies, the intra-rater reliability of LLM assessments ranges across a low level (Bui & Barrot, 2025; Manning et al., 2025). Bui and Barrot (2025), for example, found comparatively low intra-rater scores in ChatGPT assessments of college students’ essays, indicating low consistency in LLM scoring. Likewise, the alignment between LLM and human assessments is high in some studies (Oğuz, 2025; Tate et al., 2024; Yavuz et al., 2025) but low in others (Bui & Barrot, 2025; Manning et al., 2025). Pack et al. (2024), for example, found substantial reliability scores between ChatGPT and human assessments of English-language learners’ texts at the university level. Mathew et al. (2026), on the other hand, report on weak agreement between human and LLM assessments of secondary and undergraduate students’ essays.
The question arises how these mixed findings can be explained. One possible explanation is the complexity of text assessment. Evaluating texts is a multifaceted task that requires consideration of various aspects such as linguistic, formal, and content text characteristics (Kürzinger & Pohlmann-Rother, 2015). At the same time, research findings point to differences in the accuracy of LLM assessments when evaluating different text characteristics. On the one hand, LLM assessments of linguistic text characteristics appear to be quite adequate. Mizumoto and Eguchi (2023), for example, show that accuracy and reliability of LLM assessments increase when linguistic criteria are taken into account. Lexical (e.g., vocabulary/lexical diversity), syntactic, and cohesive criteria are particularly relevant here. The assessment of content text characteristics (e.g., fulfillment of the writing task, originality of the text), on the other hand, appears to be less appropriate. Seßler et al. (2025) report on low inter-rater scores between human and LLM ratings of seventh- and eighth-grade students’ essays when content text characteristics are assessed. Yavuz et al. (2025) show that the alignment between human and LLM ratings of university students’ essays is higher when it comes to assessing linguistic rather than content-related text qualities. Lan et al. (2025) demonstrate that the association between human and LLM assessments of university students’ texts is higher when linguistic rather than content characteristics are focused on. In summary, the findings on assessment accuracy seem to diverge depending on the criteria considered in a study. However, it should be noted that these studies focused exclusively on texts written by students at the secondary or university level. It is unclear whether and to what extent these findings apply to the assessment of texts from early elementary education. Since students in elementary education are in the early stages of writing development (Martin & Dockrell, 2024) and their texts are therefore significantly different from those in higher grades (e.g., shorter length, less syntactic complexity), a specific investigation of this is warranted. Nevertheless, research on this topic is still limited. One exception is the study by Theurer et al., which focused on the extent to which ChatGPT 4 can evaluate texts written by first-grade students. The results indicate that LLM assessments are inconsistent and show little alignment with human assessments. Even gradually specifying the prompt that initiated LLM assessments did not lead to increasing consistency or alignment. However, that study only considered assessments of holistic text quality. It did not evaluate different criteria such as linguistic or content text characteristics. This desideratum is addressed in the present study.

2. Objective and Hypotheses

The aim of this study is to investigate how accurately a common LLM can assess linguistic and content characteristics of first-graders’ texts and to what extent the accuracy of the assessments varies between the characteristics considered. Following previous studies on LLM text assessment, the accuracy of LLM assessments is measured using two indicators. The first indicator is the consistency, i.e., intra-rater reliability of LLM assessments. Based on results from secondary and higher education, we assume that the intra-rater reliability of LLM assessments is higher for linguistic text characteristics than for content text characteristics. Therefore, the first hypothesis is:
H1. 
LLM assessments of linguistic text characteristics are more consistent than LLM assessments of content text characteristics.
The second indicator is the alignment between LLM assessments and human assessments. Based on previous findings, we assume that LLM assessments and human assessments align more closely when assessing linguistic text characteristics than content text characteristics. We expect this pattern to be evident both in the association and in the agreement of LLM and human assessments. Thus, our second hypothesis is:
H2. 
Alignment between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.
H2.1. 
Association between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.
H2.2. 
Agreement between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.
Consequently, we consider both consistency of LLM assessments and alignment between human and LLM assessments to be indicators of LLM assessment accuracy. However, it is important to note that in this way, we do not capture assessment accuracy in an absolute sense. For example, human ratings may also be biased and do therefore not reflect an unquestionable ideal of assessment accuracy. Instead of measuring accuracy in absolute terms, we do thus merely approximate LLM assessment accuracy. For a critical discussion on this issue, please consider Section 6.

3. Methods

The present study is a secondary analysis of data collected in the context of the German NaSch1 project1 (Pohlmann-Rother et al., 2016). The NaSch1 project is therefore presented first (Section 3.1), followed by a description of the procedure used in the present study (Section 3.2).

3.1. Database: NaSch1

NaSch12 is a research project that was conducted with data from German elementary schools between 2012 and 2016. NaSch1 was part of the PERLE project, a multi-faceted and multi-level study, which collected data between 2006 and 2010 (Theurer et al., 2025). NaSch1 uses data from the video study “German” that was conducted as part of PERLE. In this video study, a 90 min lesson was videotaped in 38 first-grade classes. At the center of these lessons was the picture book ‘Lucy rettet Mama Kroko’ (Doucet & Wilsdorf, 2005; original title: ‘Alligator Sue’; Doucet, 2003). The book centers on a human girl named Lucy who grows up among crocodiles but runs away when the other crocodiles begin to tease her. The students’ task was to write a letter from Lucy’s perspective to Lucy’s mother among the crocodiles, explaining why she ran away from the crocodiles. Within PERLE, parents approved the production of the texts as well as their editing and analysis for scientific purposes. Parents were informed about the aim of the study, data protection proceedings, how they can withdraw their potential consent at any given time, and the fact that no child will ever experience disadvantages in case parents do not consent to participation in the PERLE study.
In NaSch1 and PERLE, these letters were analyzed in detail. In NaSch1, the focus was on assessing the letters’ holistic quality and various linguistic, formal, and content-related criteria (Pohlmann-Rother et al., 2016) in order to relate those individual data to other multi-level indicators (e.g., precursor skills, self-concept, instructional quality measures). Before the analysis, the handwritten texts were transformed into digital form, eliminating potential personal indicators (Kürzinger et al., 2013). For the sake of data protection within the present study, there is no further individual information available than age, sex, and class affiliation.
Sample. The sample of NaSch1 comprised 540 texts written by first-graders who were 7.1 years old on average (SD = 4.80). Approximately half of the students were female (55.92%; n = 301). 42.2% of the students attended private schools.
Human Scoring of Data Material. Human raters were trained to evaluate the texts using predefined assessment rubrics (Kürzinger et al., 2013). Specifically, two rubrics were used: one to assess the holistic quality of the texts (“holistic rating”; Pohlmann-Rother et al., 2016). The other rubric addressed different text characteristics and defined criteria for assessing these characteristics individually (“criterion-based rating”; Kürzinger & Pohlmann-Rother, 2015). The characteristics included linguistic and content-related criteria, among others. Linguistic criteria included, for example, the appropriateness of the language used in the letters, the vocabulary used in the letters, and the use of cohesive devices (e.g., “whereas”, “while”, “as”). Content-related criteria included, among other things, the originality of the letters, the correct adoption of Lucy’s perspective in the letters, and the addressing of the farewell. All texts were assessed by two independent raters on a four-point Likert scale (1 = low quality; 4 = high quality). Inter-rater reliability was calculated as a measure of agreement between the raters, leading to satisfactory scores (0.76 < κ < 0.97; 0.86 < g < 0.94; Kürzinger et al., 2013, Pohlmann-Rother et al., 2016). For building the human score in the present study, we used the mean of both human ratings of every text.

3.2. The Present Study

In the present study, the texts from NaSch1 were re-assessed by a common LLM, ChatGPT (see Section 3.2.1). Afterwards, these assessments were analyzed and compared to the human assessments from NaSch1 (see Section 3.2.2). The data for the present study were generated on 16 October 2025.

3.2.1. LLM Assessments of Data Material

In the first step, we uploaded the NaSch1 texts (N = 5393) to ChatGPT Version 5.0. We opted for ChatGPT due to its accessibility and popularity in educational contexts (Roumeliotis & Tselikas, 2023). We used a ChatGPT Team license, as ChatGPT does not store the inputs and use them for further training of the language model when using a Team license. In this way, we made sure that our study fulfills ethical requirements and is in line with applicable data protection regulations. Furthermore, this ensured that later GPT, i.e., LLM assessments, were not distorted by the results of earlier assessments. We left the settings at default with a temperature of 1.00, a top_p of 0.90, and a token limit of 128,000. We used the following prompt for initiating LLM assessments:
I will provide you with an Excel spreadsheet containing a column titled ‘Fair copies’ with texts written by first-grade students (text type: letter). In the appendix, you will also find a Word document containing a description of rating levels of a text assessment criterion. Your task is to assess the fair copies in the Excel spreadsheet based on this document and the rating guidelines it contains. Keep in mind that the language proficiency level of a first-grader is not comparable to that of an adult, so do not assess for absolute perfection but for realistic, age-appropriate performance. Enter your assessment in the column next to the fair copies and then send me your results as an .xlsx file for download.
We then uploaded the rating levels, i.e., assessment guidelines for the single criteria, to the LLM separately one after another. Both rating levels (four-point: 1 = lowest level; 4 = highest level) and rating descriptions were the same as rating levels and descriptions used by the human raters in NaSch1 (see Appendix A for a detailed description). In line with our research interest, we took into account both linguistic and content-related criteria:
The linguistic criteria included in further analyses were cohesive devices and vocabulary. Cohesive devices are understood as linguistic elements that signal relationships between sentences or parts of sentences, e.g., “because”, “whether”, “as”. The cohesive devices criterion captures the extent to which cohesive devices are used in a letter. The vocabulary criterion indicates how differentiated the vocabulary in a letter is, i.e., the diversity of the vocabulary (Kürzinger & Pohlmann-Rother, 2015). We chose these criteria because they have been shown to be particularly relevant for the accuracy of LLM assessments (Mizumoto & Eguchi, 2023).
The content-related criteria included in further analyses were originality and perspective taking. The originality criterion captures the degree of novelty and authenticity evident in a letter, e.g., when content is integrated in the letter that is not explicitly stated in the writing task but nevertheless enriches the content of the letter. The perspective taking criterion indicates whether and how consistently a letter reflects Lucy’s point of view (Kürzinger & Pohlmann-Rother, 2015). We chose both criteria because they are central content text characteristics considered in NaSch1.

3.2.2. Data Analysis

In the second step, we analyzed the LLM assessments considering the research interest of our study.
Preliminary Analysis. To be able to opt for parametric or non-parametric measures of data analysis, we first tested the normality of our data. For this, we used Q-Q-Plots, histograms, skewness, and kurtosis values in descriptive statistics. Non-normality was assumed if either skewness and kurtosis values pointed to non-normality (i.e., skewness value above 0.7 and kurtosis value above 3.5, Lei & Lomax, 2005) or if visual inspection of Q-Q-Plots and histograms indicated non-normality. All in all, tests of normality point to several non-normal distributions in our dataset. We thus opted for non-parametric measures for further analyses.
Hypothesis Testing. We proceeded with testing the hypotheses of our study. To test H1, we computed Intraclass-Correlation of all LLM ratings using a two-way random effects model (ICC2,1) with absolute agreement. Due to the non-normality of our data, we opted for calculating rank Intraclass-Correlations (rank ICCs) as rank ICCs are insensitive to skewed data and violations of normal distribution (Tu et al., 2023). According to Koo and Li (2016), we considered ICC scores of < 0.5 to be low, ICC scores of 0.5 < ICC < 0.75 to be moderate, ICC scores of 0.75 < ICC < 0.9 to be good, and ICC scores of > 0.9 to be excellent. To make reliable statements about ChatGPT’s intra-rater reliability, we decided to conduct several LLM assessment rounds for each criterion. Specifically, we conducted 25 LLM assessments of all texts and each criterion, i.e., for criterion cohesive devices: 25 assessments of all 539 texts (all in all 13,475 assessments); for criterion vocabulary: 25 assessments of all 539 texts (all in all 13,475 assessments), etc. To come to this number of iterations (25), our procedure was as follows: first, we conducted five assessment rounds, selected single text assessments, and determined the modes of these text assessments through visual inspection. We then conducted five additional assessment rounds, determined the modes of the text assessments again, and analyzed the extent to which the modes had changed. We repeated this process until there were no further changes in the modes of the selected text assessments. Since this was the case after 25 assessment rounds, we conducted a total of 25 assessment rounds for each criterion. All assessment rounds were conducted in independent chats using the same prompt (see Section 3.2.1). Subsequently, we calculated rank ICC for each criterion considering all 25 assessment rounds. We preferred a random-effects model to a mixed-effects model (ICC3,1) because we strived to generalize the 25 assessment results to any number of assessments conducted by ChatGPT.
As exploratory trials show, the LLM tends to rate the same text in several assessment rounds with different scores. Therefore, we decided to identify the LLM’s overall assessment tendency for further analysis, each per text and criterion. For this, we looked at the results of the 25 assessment rounds for a text and used the mode to identify the value that had been assigned most frequently to this text on the respective criterion (see previous paragraph). If two scores occurred with equal frequency, we opted for the lower score. In this way, we received one value for every text per criterion, i.e., 539 assessments for each criterion. Finally, we summarized the values of assessment tendencies of the four criteria in separate variables (e.g., for cohesive devices: LLM_Cohesive Devices; for vocabulary: LLM_Vocabulary, etc.). These assessment tendencies were used for analyses conducted to test H2.1 and H2.2.
To test H2.1, we computed rank correlations (Spearman correlations ρ) between human assessments and overall assessment tendency of the LLM for each criterion. We considered ρ ≈ 0.1 to be a small correlation, ρ ≈ 0.3 to be a moderate correlation, and ρ > 0.5 to be a large correlation (Cohen, 1988). We opted for rank correlations because LLM assessments are shown to be non-normally distributed.
To test H2.2, we computed Linear Weighted Kappas (κw), with κw < 0.00 indicating poor agreement, 0.00 < κw < 0.20 indicating slight agreement, 0.21 < κw < 0.40 indicating fair agreement, 0.41 < κw < 0.60 indicating moderate agreement, 0.61 < κw < 0.80 indicating substantial agreement, and 0.81 < κw < 1.0 indicating (almost) perfect agreement (Landis & Koch, 1977). We opted for Weighted Kappas because this measure provides the opportunity to consider proximity of assessments between human and LLM ratings. That is, a deviation of 1 between human and LLM assessments (e.g., human assessment: 1; LLM assessment: 2) is not considered the same as a deviation of 2 or 3 (e.g., human assessment: 1; LLM assessment: 4). Furthermore, we opted for Linear Weighted Kappas because this measure has recently been shown to be more appropriate to evaluate agreement between human and automated assessments of text characteristics than Quadratic Weighted Kappas (Doewes et al., 2023).

4. Results

4.1. Tests of Normality and Descriptive Statistics

First, we tested normal distribution and computed descriptive statistics of all criteria. Table 1 summarizes descriptive statistics for each criterion, both for human assessments and LLM assessment tendencies. We start with linguistic criteria (cohesive devices; vocabulary). Afterwards, we report on findings on content-related criteria (originality; perspective taking). As we conducted 25 LLM assessments per criterion and treated each assessment round as a single rater, we also report on normality and descriptive statistics of all 25 assessment rounds per criterion (Table A5, Table A6, Table A7 and Table A8, see Appendix A).
Mean scores of LLM assessments of the criterion cohesive devices seem to be quite similar, with a range of about one rating level (min.: LLM21 with M = 2.35; max.: LLM7 with M = 3.48, Table A5). Means of human assessments and LLM assessment tendency of the criterion cohesive devices (LLM_Cohesive Devices) are quite similar, too, with human assessments ranging slightly below the theoretical mean score (M = 2.28, Table 1) and LLM assessments ranging slightly above the theoretical mean (M = 2.59). Furthermore, human ratings are shown to have a higher standard deviation (SD = 1.02) than the LLM assessment tendency of the criterion cohesive devices (SD = 0.73), indicating more variance in human assessments.
Means of human assessments of the criterion vocabulary (M = 2.80, Table 1) and LLM assessment tendency (LLM_Vocabulary, M = 2.02, Table 1) are quite different, with LLM assessments being stricter and showing less variance. Furthermore, mean scores between LLM assessments differ considerably, ranging between very low (LLM23: M = 1.12) and very high scores (LLM6: M = 3.56, Table A6). Most LLM assessments do not cover the whole rating scale: assessments do often range between 1 and 3, sometimes even between 2 and 3.
Mean scores of human assessments of the criterion originality (M = 1.09) and LLM assessment tendency of the criterion originality (M = 1.08) are almost identical, reflecting a tendency toward low scoring (Table 1). Standard deviations of human assessments (SD = 0.42) and LLM assessment tendency of originality (SD = 0.38) are comparable as well, indicating similar variance in assessments. The majority of LLM assessment rounds do not contain a score of 4, emphasizing the tendency toward strict LLM assessments of this criterion (Table A7).
Means of LLM assessment tendency and human assessments of the criterion perspective taking both point to high rating scores, with the LLM scoring somewhat stricter (M = 3.23, Table 1) and showing higher variance (SD = 0.77) than humans (M = 3.86; SD = 0.55). Furthermore, single LLM assessment rounds are highly imbalanced, rating all texts with a score of 4 (e.g., LLM2, LLM5, Table A8).

4.2. Consistency of LLM Assessments

To test hypothesis H1 (“LLM assessments of linguistic text characteristics (cohesive devices, vocabulary) are more consistent than LLM assessments of content text characteristics (originality, perspective taking).”), we computed rank Intraclass-Correlations of LLM assessments of all criteria (Table 2).
Linguistic criteria. The rank ICC score of LLM assessments of cohesive devices points to significant Intraclass-Correlation at an acceptable level (rank ICC(2,1) = 0.697, p < 0.001, Table 2). That is, LLM assessments of cohesive devices are moderately consistent. In contrast, the rank ICC score of LLM assessments of vocabulary is notably low (rank ICC(2,1) = 0.099, p < 0.001), indicating inconsistency in LLM assessments of this criterion.
Content-related criteria. The rank ICC score of LLM assessments of perspective taking points to significant but rather weak Intraclass-Correlation (rank ICC(2,1) = 0.438, p < 0.001). At the same time, the rank ICC score of LLM assessments of originality indicates low Intraclass-Correlation (rank ICC(2,1) = 0.223; p < 0.001). All in all, LLM assessments of content-related criteria are not fully consistent.
Summary. H1 is partially supported: while LLM assessments of cohesive devices are more consistent than LLM assessments of both content-related criteria, LLM assessments of vocabulary are not.

4.3. Associations Between Human and LLM Assessments

To test hypothesis H2.1 (“Association between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.”), we computed Spearman correlations between human and LLM assessments for all criteria (Table 3).
Linguistic criteria. There is a significant and quite large correlation between human assessments and LLM assessments of cohesive devices (Spearman’s ρ = 0.555, p < 0.001, Table 3). Human assessments and LLM assessments of the criterion cohesive devices are thus positively associated. Additionally, there is a significant but small correlation between human assessments and LLM assessments of vocabulary (Spearman’s ρ = 0.133, p < 0.001). That is, human and LLM assessments of vocabulary are positively associated, too, but this association is rather weak.
Content-related criteria. The correlation between human assessments and LLM assessments of originality is non-significant and considerably low (Spearman’s ρ = 0.036, p = 0.398). Consequently, there is no association between human and LLM assessments of the criterion originality. At the same time, there is a significant and moderate association between human and LLM assessments of perspective taking (Spearman’s ρ = 0.307, p < 0.001).
Summary. H2.1. is partially supported: the correlation between human and LLM assessments of cohesive devices is stronger than correlations between human and LLM assessments of content-related criteria. The correlation between human and LLM assessments of vocabulary, on the other hand, is stronger than one (originality), but not all correlations between human and LLM assessments of content-related criteria.

4.4. Agreement Between Human and LLM Assessments

To test hypothesis H2.2 (“Agreement between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.”), we calculated Linear Weighted Kappas for all criteria (Table 4).
Linguistic Criteria. Kappa points to significant, fair agreement between human and LLM assessments of cohesive devicesw = 0.356; p < 0.001, 95% Confidence Interval with a minimum limit of >0.3 and a maximum limit of >0.4, Table 4). In contrast, agreement between human and LLM assessments of vocabulary is non-significant and ranges quite low (κw = 0.023; p = 0.051, 95% Confidence Interval with a minimum and maximum limit of <0.1).
Content-related criteria. Kappa points to significant but low agreement between human and LLM assessments of perspective takingw = 0.165; p < 0.001). The Confidence Interval (95%) shows a minimum limit of <0.2 and a maximum limit of <0.3, i.e., suggest insufficient agreement. At the same time, there is no significant agreement between human assessments and LLM assessments of originalityw = 0.048; p = 0.217). Consequently, human and LLM assessments of originality do not agree beyond chance.
Summary. H2.2 is partially supported: LLM and human assessments of cohesive devices show higher agreement than LLM and human assessments of content-related criteria. However, the agreement between LLM and human assessments of vocabulary is lower than agreement between human and LLM assessments of content-related criteria.

5. Discussion

The aim of this study was to investigate how accurately a common LLM (ChatGPT 5.0) assesses linguistic and content characteristics of first-grade students’ texts and to what extent these assessments vary between text characteristics. Descriptive statistics show that LLM assessments of linguistic text characteristics are not consistently similar to human assessments. On the one hand, mean scores of human and LLM assessments of cohesive devices are quite comparable (Mhuman = 2.28, SDhuman = 1.02; MLLM = 2.59, SDLLM = 0.73). On the other hand, mean scores of human and LLM assessments of vocabulary are clearly different, with humans scoring higher than the LLM (Mhuman = 2.80, SDhuman = 0.82; MLLM = 2.02, SDLLM = 0.32). At the same time, mean scores of human and LLM assessments of originality (content characteristic) appear to be highly comparable (Mhuman = 1.09, SDhuman = 0.42; MLLM = 1.08, SDLLM = 0.38). However, mean scores of human and LLM assessments of perspective taking are noticeably different, with humans scoring higher than the LLM (Mhuman = 3.86, SDhuman = 0.55; MLLM = 3.23, SDLLM = 0.77). Descriptive statistics thus indicate that human and LLM assessments of linguistic text characteristics are not generally more similar than human and LLM assessments of content text characteristics. This pattern is also evident in the results of our further analyses.
Accordingly, in some cases, LLM assessments of linguistic characteristics are more consistent than LLM assessments of content characteristics (i.e., linguistic characteristic cohesive devices; content characteristic originality). However, in other cases, they are not (i.e., linguistic characteristic vocabulary; content characteristic perspective taking; H1). This is also evident in the analyses conducted on alignment between human and LLM assessments. While alignment between human and LLM assessments of the linguistic characteristic cohesive devices is quite fair (Spearman’s ρ = 0.555, p < 0.001; κw = 0.356, p < 0.001), it is considerably low for the linguistic characteristic vocabulary (Spearman’s ρ = 0.133, p < 0.001; κw = 0.023, p = 0.051). Considering content characteristics, we find substantial differences in the alignment between human and LLM assessments of originality (Spearman’s ρ = 0.036, p = 0.398; κw = 0.048, p = 0.217) but rather moderate differences in human and LLM assessments of perspective taking (Spearman’s ρ = 0.307, p < 0.001; κw = 0.165, p < 0.001). All in all, LLM assessments of linguistic characteristics are shown to be partially more consistent and aligned to human ratings than assessments of content characteristics (for the criterion cohesive devices)—and partially not (for the criterion vocabulary). Thus, contrary to findings of previous research in secondary and higher education (e.g., Lan et al., 2025; Yavuz et al., 2025), the present study does not indicate that LLM assessments of linguistic text characteristics are systematically more accurate than LLM assessments of content characteristics. The accuracy of assessments seems to depend more on the specific characteristic than on the overall domain (linguistic, content). Accordingly, it is worth taking a differentiated look at the characteristics that the LLM was able to assess more accurately and less accurately.
The cohesive devices criterion was rated most accurately in terms of intra- and inter-rater reliability. One explanation for this finding could be the way LLMs work. Cohesive devices that are relevant to student texts in early elementary education comprise a comparatively small and well-defined group of terms such as “or”, “however”, “as” (Kürzinger & Pohlmann-Rother, 2015). It is among the ‘core competencies’ of decoder-only LLMs to assess texts for the presence of such terms (Minaee et al., 2024). Thus, LLMs may be particularly suitable to assess this criterion. Since cohesion and, consequently, the use of cohesive devices is a major resource for text construction (Halliday & Hasan, 1976), LLMs’ potential to capture this criterion is quite remarkable. At this point, the present findings thus indicate a potential area of application for LLM assessments in school practice. At the same time, it is important to note that the reliability of LLM assessments of the criterion cohesive devices is—although comparatively high—all in all (merely) fair. LLMs may therefore offer helpful assistance in determining the quality of elementary school students’ texts regarding this criterion. However, it is not appropriate to use LLMs as stand-alone assessment tools for young children’s writing here.
In contrast to LLM assessments of cohesive devices, low intra- and inter-rater reliability scores are evident in LLM assessments of originality and vocabulary. Here, the specific context of the study—elementary school and the specifics of elementary school students’ texts—could be relevant (see also Theurer et al., under review). That is, the low consistency and agreement with human assessments could be related to the brevity and lack of information in the texts from NaSch1:
Capturing the originality of an idea within a text seems challenging when the text is rather short and hardly provides enough space for this idea to unfold. Consequently, identifying originality in texts consisting of only a few sentences or words, such as the texts from NaSch1, seems particularly difficult. The present findings suggest that the LLM reaches its limits in this regard. At the same time, originality as an indicator of creativity is a complex construct (Corazza, 2016). The question arises as to what extent an LLM can be expected to capture this criterion at all—especially since the LLM itself seems to be capable of creativity only to a certain extent (Bellemare-Pepin et al., 2026; Cropley, 2025).
In NaSch1, the vocabulary criterion was defined as diversity of words used in a letter. Its aim was to capture how differentiated the vocabulary in a text was. However, short texts with comparatively few words offer little scope for achieving a high degree of vocabulary diversity. Accordingly, it could be difficult for the LLM to identify differences in first-grade students’ texts based on differences in vocabulary diversity. Texts written by students in higher grades offer a potentially richer basis for assessment in this regard. This could explain the more accurate LLM assessment of this characteristic in studies with older students (e.g., Mizumoto & Eguchi, 2023).
All in all, the present study indicates that a blanket distinction between ‘easily assessable’ linguistic text characteristics and ‘less easily assessable’ content text characteristics is not viable for LLM assessments of first-grade students’ texts. Instead, it seems more promising to consider the specific assessment criteria individually. In doing so, it also seems expedient to take greater account of the way LLMs work, as the way LLMs solve tasks makes them predestined for the assessment of certain text characteristics—but less for the assessment of others (see above). Accordingly, in the future, LLMs should be predominantly used for the assessment of text characteristics for which they are suitable to assess according to their functioning. In this way, LLMs may offer assistance to teachers in early elementary classrooms. LLMs may, for example, generate feedback on students’ texts, which could be tailored to individual students’ needs by the teacher (Xiao et al., 2025). As a result, LLMs can help reduce teachers’ workload (Nkoyo et al., 2025) and, at the same time, set the ground for providing individual feedback to students (Steiss et al., 2024). However, it is important that LLMs do not replace but supplement and enrich teachers’ practices in this regard (Jukiewicz & Wyrwa, 2026). At this point, one should always consider that the way LLMs come to their conclusions differs significantly from human cognitive processes (Wang et al., 2025)—even if LLMs may imitate human-like thinking—and thus should be carefully situated by teachers when used for educational purposes.

6. Limitations and Outlook

This study has several limitations that offer starting points for future research. For example, we used inter-rater reliability between human and LLM assessments as an indicator of assessment accuracy. Human assessments thus served as a baseline for LLM assessments. However, this can be viewed critically as human raters are shown to be affected by biases—some of which are even reproduced by LLMs (Gallegos et al., 2024). Consequently, we do not capture LLM assessment accuracy in an absolute sense, but in terms of its internal consistency and alignment with human raters. Therefore, future benchmarks should be developed that can serve as additional baselines for determining the accuracy of LLM assessments (Novikova et al., 2025). Moreover, we did not consider the validity of LLM and human assessments. It is thus conceivable that human raters and LLM ratings differ somewhat in their construct understanding of the criteria. In future studies, the validity of assessment should also be taken into account. Another limitation of the study concerns the operationalization of the assessment criteria. These criteria were drawn from the NaSch1 study, in which they were adapted and specified for the elementary school context. However, it should be noted that the theoretical constructs behind these criteria are more comprehensive than the facets reflected in the criteria from NaSch1. Originality, for example, encompasses not only the facets of novelty and authenticity, which are reflected in the respective criterion from NaSch1, but also the uniqueness of the idea (Corazza, 2016). Since first-grade students’ texts can only be expected to be unique to a very limited extent, this aspect is not reflected in the NaSch1 criterion of originality. The present results are therefore only valid for LLM assessments of originality of first-grade students’ texts, not for LLM assessments of originality in general. This also applies to other assessment criteria. To be able to make statements about the appropriateness of LLM assessments in other contexts, future studies should specifically focus on these contexts. Further limitations concern the skewness of LLM assessment data: descriptive statistics point to floor effects of LLM assessments of originality, which may have affected the results to some degree. A comparable pattern emerges from the LLM assessments of perspective taking. Here, descriptive statistics point to ceiling effects, which may have influenced the results as well. In this context, it is also important to note that the analyses of perspective taking are based on a slightly smaller sample than the analyses of the other criteria. Moreover, some methodological decisions must be considered when reading the findings of the present study. For example, we used the mode as a measure of assessment tendency. Consequently, we were only able to approach overall LLM assessment scores, not to determine LLM scores that claim absolute validity. Additionally, further analyses could yield results under different temperature settings. Possibly, ChatGPT rates differently when forced to more “logical” and conservative processing with low(er) temperature settings. Moreover, future research could consider other LLMs such as Gemini, DeepSeek, or Claude and examine the extent to which the present findings can be replicated using these LLMs. Finally, there are some technical limitations left to be considered, including the lack of full reproducibility and repeatability of LLM outputs (Shyr et al., 2026) and a general instability of LLM assessments over time (Pack et al., 2024). These limitations must be taken into account, too, when interpreting the findings of the present study.

Author Contributions

Conceptualization, D.T., C.T., D.C., and S.P.-R.; methodology, D.T. and C.T.; formal analysis, D.T.; investigation, D.T. and C.T.; data curation, C.T.; writing—original draft, D.T.; writing—review and editing, D.T., Caroline, Theurer, D.C., and S.P.-R.; visualization, D.T. and C.T.; supervision, D.T., C.T., D.C., and S.P.-R.; project administration, D.T., C.T., and S.P.-R. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Since the study is a secondary analysis without collecting new data, no institutional review board statement was obtained.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available from the authors upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1. Rating levels—criterion cohesive devices (linguistic) 1 (Kürzinger & Pohlmann-Rother, 2015).
Table A1. Rating levels—criterion cohesive devices (linguistic) 1 (Kürzinger & Pohlmann-Rother, 2015).
Rating LevelDescription
1The value “1” is assigned if no cohesive devices are used in a text. Single sentences or parts of the text are strung together without any connection. There is no discernible text structure or linguistic cohesion. The value “1” is also assigned if the letter contains only cohesive devices that are incorrect in terms of content (usually linking devices such as “and”, “or”).
2The value “2” is assigned if cohesive devices are used to some extent in a text. The single sentences or parts of the text are thus connected to each other to some extent. Text structure or linguistic cohesion is partially recognizable through referential devices. The value “2” is also assigned if the letter additionally contains cohesive devices that are incorrect in terms of content.
3The value “3” is assigned if cohesive devices are used extensively in a text. The single sentences or parts of the text are thus well connected linguistically in most cases. Text structuring or linguistic cohesion is usually recognizable. For a typical letter with the content “I’m leaving because you annoy me”, the value “3” is assigned due to the two cohesive devices (“because”, “you”). The use of cohesive devices must be correct without exception.
4The value “4” is assigned if cohesive devices are used consistently throughout a text. The single sentences or parts of the text are all connected in a linguistically correct way. There is complete text structuring or linguistic cohesion without exception and without incorrect links. In other words, the cohesive devices are particularly used to develop the content and argumentation of a letter, so the language conveys the content coherently and correctly. The use of cohesive devices must be correct without exception.
1 In addition to the rating levels, we uploaded indicators for each of the four criteria to ChatGPT, explaining specific technical terms (here, among others, the terms “cohesive device” or “linking device”). The indicators were taken from the NaSch1 study. They are not listed separately here but can be found in Kürzinger and Pohlmann-Rother (2015).
Table A2. Rating levels—criterion vocabulary (linguistic) (Kürzinger & Pohlmann-Rother, 2015).
Table A2. Rating levels—criterion vocabulary (linguistic) (Kürzinger & Pohlmann-Rother, 2015).
Rating LevelDescription
1The value “1” is assigned if the vocabulary is very simple throughout the letter. Colloquial language, spoken or dialect expressions, and a highly simplified style of expression dominate the text. In other words, written language is hardly present or absent in the text. Decorative words (e.g., verbs or adjectives) are lacking.
2The value “2” is assigned if the vocabulary in the letter is simple and functional. Colloquial language and expressions typical of spoken language are used but do not dominate the text. Decorative words (e.g., verbs or adjectives) are rarely used.
3The value “3” is assigned if the vocabulary in the letter is largely elaborate. The vocabulary reflects almost or completely written language and varies, for example, by partially using embellishments.
4The value “4” is assigned if the vocabulary in the letter is completely elaborate and thus fully meets the standard of written language. In addition, the esthetic value of the vocabulary is particularly used. Elaborate verbs or expressions such as “leave”, “say goodbye”, “travel”, or “attention” can be found. However, a single standard expression is not sufficient to rate a text with a score of “4”.
Table A3. Rating levels—criterion originality (content-related) (Kürzinger & Pohlmann-Rother, 2015).
Table A3. Rating levels—criterion originality (content-related) (Kürzinger & Pohlmann-Rother, 2015).
Rating LevelDescription
1The value “1” is assigned if a letter does not contain any special ideas, i.e., it does not go beyond the framework of the picture book.
2The value “2” is assigned if a letter contains an original idea but this idea does neither fit the picture book nor the letter and is not linked in a meaningful way.
3The value “3” is assigned if a letter contains a special idea and this idea is developed in a meaningful way.
4The value “4” is assigned if a letter contains several special ideas and these ideas are developed in a meaningful way.
Table A4. Rating levels—criterion perspective taking (content-related) (Kürzinger & Pohlmann-Rother, 2015).
Table A4. Rating levels—criterion perspective taking (content-related) (Kürzinger & Pohlmann-Rother, 2015).
Rating LevelDescription
1The value “1” is assigned if Lucy’s perspective is not adopted in a letter and/or the character Lucy is consistently referred to in the third person in a letter. The writer therefore writes about Lucy rather than as Lucy throughout the letter. If a letter consists of a single sentence and Lucy is referred to in the third person in that sentence, this is to be assessed as consistent use of Lucy in the third person.
2The value “2” is assigned if Lucy’s perspective is adopted in a letter and, at the same time, the character Lucy is referred to in the third person several times. There must be a (part of a) sentence (proposition) in which Lucy is not referred to in the third person. The value “2” is also assigned if Lucy is referred to in the third person in one sentence (proposition) and Lucy’s perspective is successfully adopted in a second sentence (proposition).
3The value “3” is assigned if Lucy’s perspective is adopted in a letter and the character Lucy is referred to in the third person one time. There must be at least two other sentences or propositions in which Lucy is not referred to in the third person.
4The value “4” is assigned if Lucy’s perspective is adopted in a letter and the character Lucy is not referred to in the third person. The letter writer therefore consistently takes Lucy’s perspective.
Table A5. Descriptive statistics and normality (criterion cohesive devices).
Table A5. Descriptive statistics and normality (criterion cohesive devices).
MSDSkewnessKurtosisMinMaxNormality
Human2.281.020.35−0.9914Yes
LLM12.470.56−0.43−0.8213Yes
LLM22.550.56−0.64−0.5914Yes
LLM32.550.52−0.47−1.1613Yes
LLM42.960.59−1.063.1814Yes
LLM53.250.82−0.980.4114No
LLM63.111.00−0.28−1.8114No
LLM73.480.79−1.431.2414No
LLM82.560.52−0.52−1.1213Yes
LLM92.961.01−0.06−1.7814No
LLM102.460.56−0.40−0.8413Yes
LLM112.660.860.32−0.9714Yes
LLM122.550.52−0.46−1.2413Yes
LLM132.770.700.06−0.4914Yes
LLM142.450.57−0.43−0.7813Yes
LLM152.420.59−0.40−0.7014Yes
LLM162.910.30−3.3611.0013No
LLM172.850.93−0.19−1.0414No
LLM183.200.86−0.64−0.7414No
LLM192.880.930.17−1.6914No
LLM202.550.56−0.64−0.5114Yes
LLM212.350.87−0.14−0.8414Yes
LLM222.730.860.09−0.9214Yes
LLM232.730.860.09−0.9214Yes
LLM242.710.730.37−0.8314Yes
LLM252.920.87−0.19−1.0014No
LLM_Cohesive Devices2.590.730.29−0.4214Yes
Notes: N = 539; LLM_Cohesive Devices: LLM assessment tendency considering all 25 rating rounds.
Table A6. Descriptive Statistics and Normality (criterion vocabulary).
Table A6. Descriptive Statistics and Normality (criterion vocabulary).
MSDSkewnessKurtosisMinMaxNormality
Human2.800.82−0.17−0.5914Yes
LLM11.760.580.08−0.4413Yes
LLM22.040.194.7820.9123No
LLM31.420.601.100.1913No
LLM41.800.960.41−1.7913No
LLM51.980.40−0.163.4413Yes
LLM63.560.55−0.77−0.4724No
LLM72.850.860.30−1.5824No
LLM82.140.361.752.3213No
LLM91.660.830.79−0.7914No
LLM102.000.35−0.035.3513No
LLM111.140.402.515.5913No
LLM122.850.36−1.981.9423No
LLM131.460.550.61−0.7713No
LLM141.880.35−1.783.0013Yes
LLM151.690.730.780.0714No
LLM162.040.340.775.5513No
LLM172.240.450.88−0.2513Yes
LLM182.570.52−0.61−0.9613No
LLM192.780.60−2.474.3513No
LLM201.890.45−0.471.5613Yes
LLM211.950.41−0.322.7613Yes
LLM222.040.223.5016.5613No
LLM231.120.453.7212.4013No
LLM242.090.344.1517.6124No
LLM252.040.223.4415.7613No
LLM_Vocabulary2.020.320.446.8213No
Notes: N = 539; LLM_Vocabulary: LLM assessment tendency considering all 25 rating rounds.
Table A7. Descriptive Statistics and Normality (criterion originality).
Table A7. Descriptive Statistics and Normality (criterion originality).
MSDSkewnessKurtosisMinMaxNormality
Human1.090.424.8723.4414No
LLM11.750.910.60−1.3314No
LLM21.590.850.98−0.6714No
LLM31.740.970.56−1,6714No
LLM41.050.245.7735.8013No
LLM51.070.334.8423.4213No
LLM61.010.0211.07131.2113No
LLM71.100.434.2315.9413No
LLM82.370.520.15−1.1713Yes
LLM91.180.572.876.2313No
LLM101.190.542.705.8913No
LLM111.020.1910.27103.7913No
LLM121.880.941.040.3014No
LLM131.030.198.1571.4213No
LLM141.030.269.7999.8814No
LLM151.120.433.7112.7913No
LLM161.120.433.6712.5113No
LLM171.340.771.972.2214No
LLM181.010.1611.43133.8513No
LLM191.050.255.1228.3013No
LLM201.470.771.300.1814No
LLM211.390.701.500.7113No
LLM221.640.931.02−0.4914No
LLM231.540.891.08−0.7914No
LLM242.240.510.30−0.2013Yes
LLM251.390.531.060.9714No
LLM_Originality 1.080.384.6420.1313No
Notes: N = 539; LLM_Originality: LLM assessment tendency considering all 25 rating rounds.
Table A8. Descriptive Statistics and Normality (criterion perspective taking).
Table A8. Descriptive Statistics and Normality (criterion perspective taking).
MSDSkewnessKurtosisMinMaxNormality
Human3.860.55−4.3118.0214No
LLM13.230.530.17−0.2024Yes
LLM24.000.00--44No
LLM33.140.70−1.072.3114Yes
LLM43.081.34−0.86−1.2114No
LLM54.000.00--44No
LLM64.000.00--44No
LLM72.490.930.88−0.8114Yes
LLM83.690.86−2.524.6414No
LLM92.720.91−0.01−0.9614Yes
LLM102.680.97−0.11−1.0014Yes
LLM114.000.00--44No
LLM123.130.70−1.092.3314Yes
LLM133.120.70−1.062.2414Yes
LLM143.140.70−1.072.3114Yes
LLM153.130.70−1.082.2814Yes
LLM162.960.87−1.030.7014Yes
LLM174.000.00--44No
LLM183.430.98−1.290.0314No
LLM192.831.04−0.73−0.6314Yes
LLM202.960.87−1.010.6314Yes
LLM213.400.99−1.611.2914No
LLM223.830.59−4.0115.7614No
LLM232.960.87−1.030.7014Yes
LLM242.960.87−1.030.7014Yes
LLM253.540.88−2.063.2014No
LLM_Perspective Taking 3.230.77−1.181.6714No
Notes: n = 533; LLM_Perspective Taking: LLM assessment tendency considering all 25 rating rounds.

Notes

1
NaSch1: Narrative Schreibkompetenz in Klasse 1 (Narrative Writing Proficiency in Grade 1).
2
NaSch1 was funded by the German Research Foundation (grant number: PO 1739/1-1). The PERLE-Project was funded by the German Federal Ministry of Education and Research (grant numbers: PLI 3026A and PLI 3026B).
3
Due to insufficient data, one text was excluded from further analysis.

References

  1. Altamimi, A. B. (2023). Effectiveness of ChatGPT in essay autograding. In 2023 International Conference on Computing, Electronics & Communications Engineering (iCCECE), Swansea, United Kingdom, 14–16 August (pp. 102–106). IEEE. [Google Scholar] [CrossRef] [Scilit]
  2. Bellemare-Pepin, A., Lespinasse, F., Thölke, P., Harel, Y., Mathewson, K., Olson, J. A., Bengio, Y., & Jerbi, K. (2026). Divergent creativity in humans and large language models. Scientific Reports, 16, 1279. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Brindle, M., Graham, S., Harris, K. R., & Hebert, M. (2016). Third and fourth grade teacher’s classroom practices in writing: A national survey. Reading and Writing, 29, 929–954. [Google Scholar] [CrossRef] [Scilit]
  4. Bucol, J. L., & Sangkawong, N. (2025). Exploring ChatGPT as a writing assessment tool. Innovations in Education and Teaching International, 62(3), 867–882. [Google Scholar] [CrossRef] [Scilit]
  5. Bui, N. M., & Barrot, J. S. (2025). ChatGPT as an automated essay scoring tool in the writing classrooms: How it compares with human scoring. Education and Information Technologies, 30, 2041–2058. [Google Scholar] [CrossRef] [Scilit]
  6. Chiu, T. K. F., Xia, Q., Zhou, X., Chai, C. S., & Cheng, M. (2023). Systematic literature review on opportunities, challenges, and future research recommendations of artificial intelligence in education. Computers and Education: Artificial Intelligence, 4, 100118. [Google Scholar] [CrossRef] [Scilit]
  7. Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates. [Google Scholar]
  8. Corazza, G. E. (2016). Potential originality and effectiveness: The dynamic definition of creativity. Creativity Research Journal, 28(3), 258–267. [Google Scholar] [CrossRef] [Scilit]
  9. Cropley, D. H. (2025). “The cat sat on the…?” Why generative AI has limited creativity. The Journal of Creative Behavior, 59, e70077. [Google Scholar] [CrossRef] [Scilit]
  10. Deane, P. (2018). The challenges of writing in school: Conceptualizing writing development within a sociocognitive framework. Educational Psychologist, 53(4), 280–300. [Google Scholar] [CrossRef] [Scilit]
  11. Doewes, A., Kurdhi, N. A., & Saxena, A. (2023). Evaluating quadratic weighted kappa as the standard performance metric for automated essay scoring. In Proceedings of the 16th International Conference on Educational Data Mining, Bengaluru, India, 5 July (pp. 103–113). International Educational Data Mining Society. [Google Scholar] [CrossRef]
  12. Doucet, S. A. (2003). Alligator Sue. Farrar Straus & Giroux. [Google Scholar]
  13. Doucet, S. A., & Wilsdorf, A. (2005). Lucy rettet Mama Kroko. Oetinger. [Google Scholar]
  14. Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., & Ahmed, N. K. (2024). Bias and fairness in large language models: A survey. Computational Linguistics, 50(3), 1097–1179. [Google Scholar] [CrossRef] [Scilit]
  15. Graham, S. (2018). A revised Writer(s)-Within-Community model of writing. Educational Psychologist, 53(4), 258–279. [Google Scholar] [CrossRef] [Scilit]
  16. Guo, K., & Wang, D. (2024). To resist it or to embrace it? Examining ChatGPT’s potential to support teacher feedback in EFL writing. Education and Information Technologies, 29, 8435–8463. [Google Scholar] [CrossRef] [Scilit]
  17. Halliday, M. A. K., & Hasan, R. (1976). Cohesion in English. Longman. [Google Scholar]
  18. Huawei, S., & Aryadoust, V. (2023). A systematic review of automated writing evaluation systems. Education and Information Technologies, 28, 771–795. [Google Scholar] [CrossRef] [Scilit]
  19. Hussein, M. A., Hassan, H., & Nassef, M. (2019). Automated language essay scoring systems: A literature review. PeerJ Computer Science, 5, e208. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Jukiewicz, M., & Wyrwa, M. (2026). Can ChatGPT replace the teacher in assessment? A review of research on the use of Large Language Models in grading and providing feedback. Applied Sciences, 16(2), 680. [Google Scholar] [CrossRef] [Scilit]
  21. Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. [Google Scholar] [CrossRef] [Scilit]
  22. Ke, Z., & Ng, V. (2019). Automated essay scoring: A survey of the state of the art. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, Macao, 10–16 August (pp. 6300–6308). International Joint Conferences on Artificial Intelligence. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Koo, T. K., & Li, M. Y. (2016). A Guideline of selecting and reporting intraclass correlation coefficients for reliability research. Journal of Chiropractic Medicine, 15, 155–163. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Kürzinger, A., Lotz, M., Gleich, A.-K., & Kempter, I. (2013). Auswertung der Lucybriefe: Perspektivenübernahme und Schreibkompetenz. In M. Lotz, F. Lipowsky, & G. Faust (Eds.), Technischer bericht zu den PERLE-videostudien (pp. 255–296). GFPF. [Google Scholar]
  25. Kürzinger, A., & Pohlmann-Rother, S. (2015). Kriterienkatalog textkorpus. Ein instrument zur bestimmung von textqualität in Klasse 1. University of Bamberg Press. [Google Scholar]
  26. Lan, G., Li, Y., Yang, J., & He, X. (2025). Investigating a customized generative AI chatbot for automated essay scoring in a disciplinary writing task. Assessing Writing, 66, 100959. [Google Scholar] [CrossRef] [Scilit]
  27. Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33, 159–174. [Google Scholar] [CrossRef] [Scilit]
  28. Lei, M., & Lomax, R. G. (2005). The effects of varying degrees of nonnormality in structural equation models. Structural Equation Modeling, 12(1), 1–27. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Li, J., Jangamreddy, N., Bhansali, R., Hisamoto, R., Zaphir, L., Dyda, A., & Glencross, M. (2024). AI-assisted marking: Functionality and limitations of ChatGPT in written assessment evaluation. Australasian Journal of Educational Technology, 40(4), 56–72. [Google Scholar] [CrossRef] [Scilit]
  30. Manning, J., Baldwin, J., & Powell, N. (2025). Human versus machine: The effectiveness of ChatGPT in automated essay scoring. Innovations in Education and Teaching International, 62(5), 1500–1513. [Google Scholar] [CrossRef] [Scilit]
  31. Martin, C., & Dockrell, J. E. (2024). Writing productivity development in elementary school: A systematic review. Assessing Writing, 60, 100834. [Google Scholar] [CrossRef] [Scilit]
  32. Mathew, J. G., Taher, S., Kundu, A., & Barbosa, D. (2026). LLMs do not grade essays like humans [Preprint]. arXiv. [Google Scholar] [CrossRef] [Scilit]
  33. Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., & Gao, J. (2024). Large language models: A survey [Preprint]. arXiv. [Google Scholar] [CrossRef] [Scilit]
  34. Mizumoto, A., & Eguchi, M. (2023). Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics, 2, 100050. [Google Scholar] [CrossRef] [Scilit]
  35. Nkoyo, F. E. T.-A., Ijezue, C. F., Amjad, M., Amjad, A. I., Butt, S., & Castañeda-Garza, G. (2025). Advances in auto-grading with Large Language Models: A cross-disciplinary survey. In Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications, Vienna, Austria, July 31–August (pp. 477–498). Association for Computational Linguistics. [Google Scholar] [CrossRef] [Scilit]
  36. Novikova, J., Anderson, C., Blili-Hamelin, B., Rosati, D., & Majumdar, S. (2025). Consistency in language models: Current landscape, challenges, and future directions [Preprint]. arXiv. [Google Scholar] [CrossRef] [Scilit]
  37. Oğuz, E. (2025). Can generative AI figure out figurative language? The influence of idioms on essay scoring by ChatGPT, Gemini, and Deepseek. Assessing Writing, 66, 100981. [Google Scholar] [CrossRef] [Scilit]
  38. Pack, A., Barrett, A., & Escalante, J. (2024). Large language models and automated essay scoring of English language learner writing: Insights into validity and reliability. Computers and Education: Artificial Intelligence, 6, 100234. [Google Scholar] [CrossRef] [Scilit]
  39. Parr, J. M., & Jesson, R. (2016). Mapping the landscape of writing instruction in New Zealand primary school classrooms. Reading and Writing, 29, 981–1011. [Google Scholar] [CrossRef] [Scilit]
  40. Pohlmann-Rother, S., Schoreit, E., & Kürzinger, A. (2016). Schreibkompetenzen von Erstklässlern quantitativ-empirisch erfassen—Herausforderungen und Zugewinn eines analytisch-kriterialen Vorgehens gegenüber einer holistischen Bewertung. Journal for Educational Research Online, 8(2), 107–135. [Google Scholar] [CrossRef] [PubMed]
  41. Ramesh, D., & Sanampudi, S. K. (2022). An automated essay scoring system: A systematic literature review. Artificial Intelligence Review, 55, 2495–2527. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Rohloff, R., Tortorelli, L., Gerde, H. K., & Bingham, G. E. (2023). Teaching early writing: Supporting early writers from preschool to elementary school. Early Childhood Education Journal, 51, 1227–1239. [Google Scholar] [CrossRef] [Scilit]
  43. Roumeliotis, K. I., & Tselikas, N. D. (2023). ChatGPT and Open-AI models: A preliminary review. Future Internet, 15(6), 192. [Google Scholar] [CrossRef] [Scilit]
  44. Seßler, K., Fürstenberg, M., Bühler, B., & Kasneci, E. (2025). Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring. In Proceedings of the 15th International Learning Analytics and Knowledge Conference, Dublin, Ireland, 3–7 March (pp. 462–472). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  45. Shi, Y., Yu, K., Dong, Y., & Chen, F. (2026). Large language models in education: A systematic review of empirical applications, benefits, and challenges. Computers and Education: Artificial Intelligence, 10, 100529. [Google Scholar] [CrossRef] [Scilit]
  46. Shyr, C., Ren, B., Hsu, C.-Y., Yan, C., Tinker, R. J., Cassini, T. A., Hamid, R., Wright, A., Bastarache, L., Peterson, J. F., Malin, B. A., & Xu, H. (2026). A statistical framework for evaluating the repeatability and reproducibility of large language model [Preprint]. medRxiv. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Steiss, J., Tate, T., Graham, S., Cruz, J., Hebert, M., Wang, J., Moon, Y., Tseng, W., Warschauer, M., & Booth Olson, C. (2024). Comparing the quality of human and ChatGPT feedback of students’ writing. Learning and Instruction, 91, 101894. [Google Scholar] [CrossRef] [Scilit]
  48. Tate, T. P., Steiss, J., Bailey, D., Graham, S., Moon, Y., Ritchie, D., Tseng, W., & Warschauer, M. (2024). Can AI provide useful holistic essay scoring? Computers and Education: Artificial Intelligence, 7, 100255. [Google Scholar] [CrossRef] [Scilit]
  49. Theurer, C., Hess, M., Denn, A.-K., & Lipowsky, F. (Eds.). (2025). Persönlichkeits- und Lernentwicklung in der Grundschule: Neue Ergebnisse der PERLE Studie. Springer. [Google Scholar]
  50. Theurer, C., Then, D., Cropley, D., & Pohlmann-Rother, S. (under review). AI-based assessment of text quality in early primary school using large language models: Not quite good enough (yet?). Computers and Education Open. [Google Scholar]
  51. Trapman, M., van Gelderen, A., van Schooten, E., & Hulstijn, J. (2018). Writing proficiency level and writing development of low-achieving adolescents: The roles of linguistic knowledge, fluency, and metacognitive knowledge. Reading and Writing, 31, 893–926. [Google Scholar] [CrossRef] [Scilit]
  52. Tu, S., Li, C., Zeng, D., & Shepherd, B. E. (2023). Rank intraclass correlation for clustered data. Statistics in Medicine, 42, 4333–4348. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Wang, Y., Deng, Y., Wang, G., Li, T., Xiao, H., & Zhang, Y. (2025). The fluency-based semantic network of LLMs differs from humans. Computers in Human Behavior: Artificial Humans, 3, 100103. [Google Scholar] [CrossRef] [Scilit]
  54. Wilson, J., & Huang, Y. (2024). Validity of automated essay scores for elementary-age English language learners: Evidence of bias? Assessing Writing, 60, 100815. [Google Scholar] [CrossRef] [Scilit]
  55. Xiao, C., Ma, W., Song, Q., Xu, S. X., Zhang, K., Wang, Y., & Fu, Q. (2025). Human-AI collaborative essay scoring: A dual-process framework with LLMs. In Proceedings of the 15th International Learning Analytics and Knowledge Conference, Dublin, Ireland, 3–7 March (pp. 293–305). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  56. Xu, W., Mahmud, R., & Hoo, W. L. (2024). A systematic literature review: Are automated essay scoring systems competent in real-life education scenarios? IEEE Access, 12, 77639–77657. [Google Scholar] [CrossRef] [Scilit]
  57. Yavuz, F., Çelik, Ö., & Yavaş Çelik, G. (2025). Utilizing large language models for EFL essay grading: An examination of reliability and validity in rubric-based assessments. British Journal of Educational Technology, 56(1), 150–166. [Google Scholar] [CrossRef] [Scilit]
  58. Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., … Wen, J.-R. (2026). A survey of large language models. Frontiers in Computer Science, 12(20), 2012627. [Google Scholar] [CrossRef] [Scilit]
Table 1. Descriptive statistics and normality—human assessment and LLM assessment tendencies.
Table 1. Descriptive statistics and normality—human assessment and LLM assessment tendencies.
MSDMedianIQRSkewnessKurtosisMinMaxNormality
Cohesive devicesHuman2.281.02220.35−0.9914Yes
LLM2.590.73310.29−0.4214Yes
VocabularyHuman2.800.8231−0.17−0.5914Yes
LLM2.020.32200.446.8213No
OriginalityHuman1.090.42104.8723.4414No
LLM1.080.38104.6420.1313No
Perspective takingHuman3.860.5540−4.3118.0214No
LLM3.230.7731−1.181.6714No
Notes: M = mean, SD = standard deviation, IQR = interquartile range.
Table 2. Rank ICC(2,1), absolute agreement, single measure.
Table 2. Rank ICC(2,1), absolute agreement, single measure.
ICC95% Confidence Intervalp
Min.Max.
Cohesive Devices aLLM1, LLM2, LLM3, …, LLM25 0.6970.6710.723<0.001
Vocabulary aLLM1, LLM2, LLM3, …, LLM25 0.0990.0850.115<0.001
Originality aLLM1, LLM2, LLM3, …, LLM250.2230.2000.249<0.001
Perspective Taking bLLM1, LLM2, LLM3, …, LLM25 0.4380.4080.471<0.001
Notes: a N = 539; b n = 533: six cases were excluded due to missing values (listwise deletion).
Table 3. Spearman correlations (rho; p) between human and LLM assessments.
Table 3. Spearman correlations (rho; p) between human and LLM assessments.
Spearman’s ρp
LLM_Cohesive Devices a0.555 <0.001
LLM_Vocabulary a0.133<0.001
LLM_Originality a0.036 0.398
LLM_Perspective Taking b0.307<0.001
Notes: a N = 539; b n = 533: six cases were excluded due to missing values (listwise deletion).
Table 4. Agreement between human and LLM assessments (Linear Weighted Kappas κw).
Table 4. Agreement between human and LLM assessments (Linear Weighted Kappas κw).
κw95% Confidence Intervalp
Min.Max.
LLM_Cohesive Devices a0.3560.3060.407<0.001
LLM_Vocabulary a0.0230.0010.0440.051
LLM_Originality a0.048−0.0670.1640.217
LLM_Perspective Taking b0.1650.1030.227<0.001
Notes: a N = 539; b n = 533: six cases were excluded due to missing values (listwise deletion).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Then, D.; Theurer, C.; Cropley, D.; Pohlmann-Rother, S. Assessing Linguistic and Content Characteristics of First-Grade Students’ Texts Using Generative AI. Educ. Sci. 2026, 16, 1227. https://doi.org/10.3390/educsci16081227

AMA Style

Then D, Theurer C, Cropley D, Pohlmann-Rother S. Assessing Linguistic and Content Characteristics of First-Grade Students’ Texts Using Generative AI. Education Sciences. 2026; 16(8):1227. https://doi.org/10.3390/educsci16081227

Chicago/Turabian Style

Then, Daniel, Caroline Theurer, David Cropley, and Sanna Pohlmann-Rother. 2026. "Assessing Linguistic and Content Characteristics of First-Grade Students’ Texts Using Generative AI" Education Sciences 16, no. 8: 1227. https://doi.org/10.3390/educsci16081227

APA Style

Then, D., Theurer, C., Cropley, D., & Pohlmann-Rother, S. (2026). Assessing Linguistic and Content Characteristics of First-Grade Students’ Texts Using Generative AI. Education Sciences, 16(8), 1227. https://doi.org/10.3390/educsci16081227

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop