1. Introduction
Writing is a key cultural skill that is essential for successful social participation (
Trapman et al., 2018). Since writing skills develop in the early years of elementary school (
Rohloff et al., 2023), their promotion is among the core tasks of elementary school teachers. The accurate assessment of students’ written products plays a key role in this. In a current model of writing research (
revised writers-within-community model of writing),
Graham (
2018) postulates that writing development does not depend exclusively on cognitive factors. As writing is a social activity that is situated within specific contexts (writing communities), the exchange with interaction partners such as teachers is also important. Teachers’ assessments of written products and the resulting feedback are therefore influential for students’ writing development (see also
Deane, 2018).
However, the individual assessment of student texts is time-consuming and complex. Teachers report that they frequently lack adequate preparation for this task (
Brindle et al., 2016;
Parr & Jesson, 2016). Consequently, the question arises of how teachers can be supported in assessing student texts in everyday school life. Given current technological developments in society and education (
Chiu et al., 2023), technology-based solutions appear promising in this regard. Automated applications for text assessment—known as Automated Essay Scoring Systems (AES)—are considered to have particular potential here.
AES are computer-based systems designed for the automated assessment of written texts (
Ramesh & Sanampudi, 2022). To date, a wide range of AES has been developed:
Ke and Ng (
2019) distinguish between three dimensions that may be used to systematize current AES (the
task for which the application was developed, its
approach, and the
features it offers). The state of research on AES is also extensive and shows that the assessment of texts by AES is largely consistent with the assessments of human raters (
Huawei & Aryadoust, 2023;
Hussein et al., 2019), even when assessing student texts in elementary education (
Wilson & Huang, 2024). Nevertheless, there is also criticism of AES. For instance, the development of AES is considered to be time-consuming and resource-intensive (
Oğuz, 2025). In addition, AES are strongly tied to the writing tasks for which they were developed. Thus, an AES is not suitable to assess different types of texts written by students across school lessons. Instead, it is only able to assess the texts for the evaluation of which it was developed. AES are therefore of limited use in school practice (
Xu et al., 2024).
An alternative to AES that has potential to address these criticisms is offered by generative AI and, in particular, large language models (LLMs). LLMs are statistical, pre-trained language models that generate texts by calculating probabilities about word sequences: texts are created by linking words that are most likely to follow each other (
Minaee et al., 2024). One of the most prominent and powerful LLMs is ChatGPT. The GPT series is based on the principle of LLM decoder-only architecture. That is, the LLM predicts the next word in a sentence based on previous word(s) to reconstruct the pre-training data, which consists of hundreds of billions of parameters (
Zhao et al., 2026). LLMs are therefore widely established tools for automated text generation (
Shi et al., 2026).
At the same time, LLMs are a potentially promising aid for assessing texts (
Bucol & Sangkawong, 2025). Several authors consider LLMs to be used as assessment tools and describe potential advantages of LLMs for evaluating texts in educational contexts:
Kasneci et al. (
2023), for instance, state that LLMs are quite easy to implement in school settings and can be used flexibly for different assessment tasks.
Guo and Wang (
2024) report that LLMs can provide balanced feedback on different text characteristics such as content or language. At the same time, various studies have focused on the accuracy of LLM assessments of student texts. Two indicators are often used to measure the accuracy of LLM assessments: (1) the intra-rater reliability of different LLM ratings, i.e., the consistency of AI-generated assessments; (2) the inter-rater reliability of LLM and human ratings, meaning the alignment between LLM and human assessments. For both indicators, research findings paint a mixed picture. The intra-rater reliability of LLM assessments is shown to be high in some studies (
Altamimi, 2023), even higher than the intra-rater reliability of human raters (
Tate et al., 2024).
Oğuz (
2025), for example, found evidence that ChatGPT, Gemini, and DeepSeek assess sixth- to 12th-grade students’ essays with excellent intra-rater reliability, with an intra-class correlation ranging between 0.753 and 0.807. Closed-source models such as ChatGPT are shown to be particularly promising in this regard (
Seßler et al., 2025). There is also some evidence that high-quality texts are assessed with higher consistency than low-quality texts (
Li et al., 2024). However, in other studies, the intra-rater reliability of LLM assessments ranges across a low level (
Bui & Barrot, 2025;
Manning et al., 2025).
Bui and Barrot (
2025), for example, found comparatively low intra-rater scores in ChatGPT assessments of college students’ essays, indicating low consistency in LLM scoring. Likewise, the alignment between LLM and human assessments is high in some studies (
Oğuz, 2025;
Tate et al., 2024;
Yavuz et al., 2025) but low in others (
Bui & Barrot, 2025;
Manning et al., 2025).
Pack et al. (
2024), for example, found substantial reliability scores between ChatGPT and human assessments of English-language learners’ texts at the university level.
Mathew et al. (
2026), on the other hand, report on weak agreement between human and LLM assessments of secondary and undergraduate students’ essays.
The question arises how these mixed findings can be explained. One possible explanation is the complexity of text assessment. Evaluating texts is a multifaceted task that requires consideration of various aspects such as linguistic, formal, and content text characteristics (
Kürzinger & Pohlmann-Rother, 2015). At the same time, research findings point to differences in the accuracy of LLM assessments when evaluating different text characteristics. On the one hand, LLM assessments of linguistic text characteristics appear to be quite adequate.
Mizumoto and Eguchi (
2023), for example, show that accuracy and reliability of LLM assessments increase when linguistic criteria are taken into account. Lexical (e.g., vocabulary/lexical diversity), syntactic, and cohesive criteria are particularly relevant here. The assessment of content text characteristics (e.g., fulfillment of the writing task, originality of the text), on the other hand, appears to be less appropriate.
Seßler et al. (
2025) report on low inter-rater scores between human and LLM ratings of seventh- and eighth-grade students’ essays when content text characteristics are assessed.
Yavuz et al. (
2025) show that the alignment between human and LLM ratings of university students’ essays is higher when it comes to assessing linguistic rather than content-related text qualities.
Lan et al. (
2025) demonstrate that the association between human and LLM assessments of university students’ texts is higher when linguistic rather than content characteristics are focused on. In summary, the findings on assessment accuracy seem to diverge depending on the criteria considered in a study. However, it should be noted that these studies focused exclusively on texts written by students at the secondary or university level. It is unclear whether and to what extent these findings apply to the assessment of texts from early elementary education. Since students in elementary education are in the early stages of writing development (
Martin & Dockrell, 2024) and their texts are therefore significantly different from those in higher grades (e.g., shorter length, less syntactic complexity), a specific investigation of this is warranted. Nevertheless, research on this topic is still limited. One exception is the study by
Theurer et al., which focused on the extent to which ChatGPT 4 can evaluate texts written by first-grade students. The results indicate that LLM assessments are inconsistent and show little alignment with human assessments. Even gradually specifying the prompt that initiated LLM assessments did not lead to increasing consistency or alignment. However, that study only considered assessments of holistic text quality. It did not evaluate different criteria such as linguistic or content text characteristics. This desideratum is addressed in the present study.
2. Objective and Hypotheses
The aim of this study is to investigate how accurately a common LLM can assess linguistic and content characteristics of first-graders’ texts and to what extent the accuracy of the assessments varies between the characteristics considered. Following previous studies on LLM text assessment, the accuracy of LLM assessments is measured using two indicators. The first indicator is the consistency, i.e., intra-rater reliability of LLM assessments. Based on results from secondary and higher education, we assume that the intra-rater reliability of LLM assessments is higher for linguistic text characteristics than for content text characteristics. Therefore, the first hypothesis is:
H1. LLM assessments of linguistic text characteristics are more consistent than LLM assessments of content text characteristics.
The second indicator is the alignment between LLM assessments and human assessments. Based on previous findings, we assume that LLM assessments and human assessments align more closely when assessing linguistic text characteristics than content text characteristics. We expect this pattern to be evident both in the association and in the agreement of LLM and human assessments. Thus, our second hypothesis is:
H2. Alignment between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.
H2.1. Association between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.
H2.2. Agreement between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.
Consequently, we consider both consistency of LLM assessments and alignment between human and LLM assessments to be indicators of LLM assessment accuracy. However, it is important to note that in this way, we do not capture assessment accuracy in an
absolute sense. For example, human ratings may also be biased and do therefore not reflect an unquestionable ideal of assessment accuracy. Instead of measuring accuracy in absolute terms, we do thus merely
approximate LLM assessment accuracy. For a critical discussion on this issue, please consider
Section 6.
4. Results
4.1. Tests of Normality and Descriptive Statistics
First, we tested normal distribution and computed descriptive statistics of all criteria.
Table 1 summarizes descriptive statistics for each criterion, both for human assessments and LLM assessment tendencies. We start with linguistic criteria (
cohesive devices;
vocabulary). Afterwards, we report on findings on content-related criteria (
originality;
perspective taking). As we conducted 25 LLM assessments per criterion and treated each assessment round as a single rater, we also report on normality and descriptive statistics of all 25 assessment rounds per criterion (
Table A5,
Table A6,
Table A7 and
Table A8, see
Appendix A).
Mean scores of LLM assessments of the criterion
cohesive devices seem to be quite similar, with a range of about one rating level (min.: LLM21 with
M = 2.35; max.: LLM7 with
M = 3.48,
Table A5). Means of human assessments and LLM assessment tendency of the criterion
cohesive devices (LLM_Cohesive Devices) are quite similar, too, with human assessments ranging slightly below the theoretical mean score (
M = 2.28,
Table 1) and LLM assessments ranging slightly above the theoretical mean (
M = 2.59). Furthermore, human ratings are shown to have a higher standard deviation (
SD = 1.02) than the LLM assessment tendency of the criterion
cohesive devices (
SD = 0.73), indicating more variance in human assessments.
Means of human assessments of the criterion
vocabulary (
M = 2.80,
Table 1) and LLM assessment tendency (LLM_Vocabulary,
M = 2.02,
Table 1) are quite different, with LLM assessments being stricter and showing less variance. Furthermore, mean scores between LLM assessments differ considerably, ranging between very low (LLM23:
M = 1.12) and very high scores (LLM6:
M = 3.56,
Table A6). Most LLM assessments do not cover the whole rating scale: assessments do often range between 1 and 3, sometimes even between 2 and 3.
Mean scores of human assessments of the criterion
originality (
M = 1.09) and LLM assessment tendency of the criterion
originality (
M = 1.08) are almost identical, reflecting a tendency toward low scoring (
Table 1). Standard deviations of human assessments (
SD = 0.42) and LLM assessment tendency of
originality (
SD = 0.38) are comparable as well, indicating similar variance in assessments. The majority of LLM assessment rounds do not contain a score of 4, emphasizing the tendency toward strict LLM assessments of this criterion (
Table A7).
Means of LLM assessment tendency and human assessments of the criterion
perspective taking both point to high rating scores, with the LLM scoring somewhat stricter (
M = 3.23,
Table 1) and showing higher variance (
SD = 0.77) than humans (
M = 3.86;
SD = 0.55). Furthermore, single LLM assessment rounds are highly imbalanced, rating all texts with a score of 4 (e.g., LLM2, LLM5,
Table A8).
4.2. Consistency of LLM Assessments
To test hypothesis H1 (“LLM assessments of linguistic text characteristics (
cohesive devices,
vocabulary) are more consistent than LLM assessments of content text characteristics (
originality, perspective taking).”), we computed rank Intraclass-Correlations of LLM assessments of all criteria (
Table 2).
Linguistic criteria. The rank ICC score of LLM assessments of
cohesive devices points to significant Intraclass-Correlation at an acceptable level (rank ICC(2,1) = 0.697,
p < 0.001,
Table 2). That is, LLM assessments of cohesive devices are moderately consistent. In contrast, the rank ICC score of LLM assessments of
vocabulary is notably low (rank ICC(2,1) = 0.099,
p < 0.001), indicating inconsistency in LLM assessments of this criterion.
Content-related criteria. The rank ICC score of LLM assessments of perspective taking points to significant but rather weak Intraclass-Correlation (rank ICC(2,1) = 0.438, p < 0.001). At the same time, the rank ICC score of LLM assessments of originality indicates low Intraclass-Correlation (rank ICC(2,1) = 0.223; p < 0.001). All in all, LLM assessments of content-related criteria are not fully consistent.
Summary. H1 is partially supported: while LLM assessments of cohesive devices are more consistent than LLM assessments of both content-related criteria, LLM assessments of vocabulary are not.
4.3. Associations Between Human and LLM Assessments
To test hypothesis H2.1 (“
Association between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.”), we computed Spearman correlations between human and LLM assessments for all criteria (
Table 3).
Linguistic criteria. There is a significant and quite large correlation between human assessments and LLM assessments of
cohesive devices (Spearman’s ρ = 0.555,
p < 0.001,
Table 3). Human assessments and LLM assessments of the criterion
cohesive devices are thus positively associated. Additionally, there is a significant but small correlation between human assessments and LLM assessments of
vocabulary (Spearman’s ρ = 0.133,
p < 0.001). That is, human and LLM assessments of
vocabulary are positively associated, too, but this association is rather weak.
Content-related criteria. The correlation between human assessments and LLM assessments of originality is non-significant and considerably low (Spearman’s ρ = 0.036, p = 0.398). Consequently, there is no association between human and LLM assessments of the criterion originality. At the same time, there is a significant and moderate association between human and LLM assessments of perspective taking (Spearman’s ρ = 0.307, p < 0.001).
Summary. H2.1. is partially supported: the correlation between human and LLM assessments of cohesive devices is stronger than correlations between human and LLM assessments of content-related criteria. The correlation between human and LLM assessments of vocabulary, on the other hand, is stronger than one (originality), but not all correlations between human and LLM assessments of content-related criteria.
4.4. Agreement Between Human and LLM Assessments
To test hypothesis H2.2 (“
Agreement between human and LLM assessments is higher for the assessment of linguistic text characteristics than for the assessment of content text characteristics.”), we calculated Linear Weighted Kappas for all criteria (
Table 4).
Linguistic Criteria. Kappa points to significant, fair agreement between human and LLM assessments of
cohesive devices (κ
w = 0.356;
p < 0.001, 95% Confidence Interval with a minimum limit of >0.3 and a maximum limit of >0.4,
Table 4). In contrast, agreement between human and LLM assessments of
vocabulary is non-significant and ranges quite low (κ
w = 0.023;
p = 0.051, 95% Confidence Interval with a minimum and maximum limit of <0.1).
Content-related criteria. Kappa points to significant but low agreement between human and LLM assessments of perspective taking (κw = 0.165; p < 0.001). The Confidence Interval (95%) shows a minimum limit of <0.2 and a maximum limit of <0.3, i.e., suggest insufficient agreement. At the same time, there is no significant agreement between human assessments and LLM assessments of originality (κw = 0.048; p = 0.217). Consequently, human and LLM assessments of originality do not agree beyond chance.
Summary. H2.2 is partially supported: LLM and human assessments of cohesive devices show higher agreement than LLM and human assessments of content-related criteria. However, the agreement between LLM and human assessments of vocabulary is lower than agreement between human and LLM assessments of content-related criteria.
5. Discussion
The aim of this study was to investigate how accurately a common LLM (ChatGPT 5.0) assesses linguistic and content characteristics of first-grade students’ texts and to what extent these assessments vary between text characteristics. Descriptive statistics show that LLM assessments of linguistic text characteristics are not consistently similar to human assessments. On the one hand, mean scores of human and LLM assessments of cohesive devices are quite comparable (Mhuman = 2.28, SDhuman = 1.02; MLLM = 2.59, SDLLM = 0.73). On the other hand, mean scores of human and LLM assessments of vocabulary are clearly different, with humans scoring higher than the LLM (Mhuman = 2.80, SDhuman = 0.82; MLLM = 2.02, SDLLM = 0.32). At the same time, mean scores of human and LLM assessments of originality (content characteristic) appear to be highly comparable (Mhuman = 1.09, SDhuman = 0.42; MLLM = 1.08, SDLLM = 0.38). However, mean scores of human and LLM assessments of perspective taking are noticeably different, with humans scoring higher than the LLM (Mhuman = 3.86, SDhuman = 0.55; MLLM = 3.23, SDLLM = 0.77). Descriptive statistics thus indicate that human and LLM assessments of linguistic text characteristics are not generally more similar than human and LLM assessments of content text characteristics. This pattern is also evident in the results of our further analyses.
Accordingly, in some cases, LLM assessments of linguistic characteristics are more consistent than LLM assessments of content characteristics (i.e., linguistic characteristic
cohesive devices; content characteristic
originality). However, in other cases, they are not (i.e., linguistic characteristic
vocabulary; content characteristic
perspective taking; H1). This is also evident in the analyses conducted on alignment between human and LLM assessments. While alignment between human and LLM assessments of the linguistic characteristic
cohesive devices is quite fair (Spearman’s ρ = 0.555,
p < 0.001; κ
w = 0.356,
p < 0.001), it is considerably low for the linguistic characteristic
vocabulary (Spearman’s ρ = 0.133,
p < 0.001; κ
w = 0.023,
p = 0.051). Considering content characteristics, we find substantial differences in the alignment between human and LLM assessments of
originality (Spearman’s ρ = 0.036,
p = 0.398; κ
w = 0.048,
p = 0.217) but rather moderate differences in human and LLM assessments of
perspective taking (Spearman’s ρ = 0.307,
p < 0.001; κ
w = 0.165,
p < 0.001). All in all, LLM assessments of linguistic characteristics are shown to be partially more consistent and aligned to human ratings than assessments of content characteristics (for the criterion
cohesive devices)—and partially not (for the criterion
vocabulary). Thus, contrary to findings of previous research in secondary and higher education (e.g.,
Lan et al., 2025;
Yavuz et al., 2025), the present study does not indicate that LLM assessments of linguistic text characteristics are systematically more accurate than LLM assessments of content characteristics. The accuracy of assessments seems to depend more on the specific characteristic than on the overall domain (linguistic, content). Accordingly, it is worth taking a differentiated look at the characteristics that the LLM was able to assess more accurately and less accurately.
The
cohesive devices criterion was rated most accurately in terms of intra- and inter-rater reliability. One explanation for this finding could be the way LLMs work. Cohesive devices that are relevant to student texts in early elementary education comprise a comparatively small and well-defined group of terms such as “or”, “however”, “as” (
Kürzinger & Pohlmann-Rother, 2015). It is among the ‘core competencies’ of decoder-only LLMs to assess texts for the presence of such terms (
Minaee et al., 2024). Thus, LLMs may be particularly suitable to assess this criterion. Since cohesion and, consequently, the use of cohesive devices is a major resource for text construction (
Halliday & Hasan, 1976), LLMs’ potential to capture this criterion is quite remarkable. At this point, the present findings thus indicate a potential area of application for LLM assessments in school practice. At the same time, it is important to note that the reliability of LLM assessments of the criterion
cohesive devices is—although comparatively high—all in all (merely) fair. LLMs may therefore offer helpful
assistance in determining the quality of elementary school students’ texts regarding this criterion. However, it is not appropriate to use LLMs as
stand-alone assessment tools for young children’s writing here.
In contrast to LLM assessments of
cohesive devices, low intra- and inter-rater reliability scores are evident in LLM assessments of
originality and
vocabulary. Here, the specific context of the study—elementary school and the specifics of elementary school students’ texts—could be relevant (see also
Theurer et al., under review). That is, the low consistency and agreement with human assessments could be related to the brevity and lack of information in the texts from NaSch1:
Capturing the
originality of an idea within a text seems challenging when the text is rather short and hardly provides enough space for this idea to unfold. Consequently, identifying originality in texts consisting of only a few sentences or words, such as the texts from NaSch1, seems particularly difficult. The present findings suggest that the LLM reaches its limits in this regard. At the same time, originality as an indicator of creativity is a complex construct (
Corazza, 2016). The question arises as to what extent an LLM can be expected to capture this criterion at all—especially since the LLM itself seems to be capable of creativity only to a certain extent (
Bellemare-Pepin et al., 2026;
Cropley, 2025).
In NaSch1, the
vocabulary criterion was defined as diversity of words used in a letter. Its aim was to capture how differentiated the vocabulary in a text was. However, short texts with comparatively few words offer little scope for achieving a high degree of vocabulary diversity. Accordingly, it could be difficult for the LLM to identify differences in first-grade students’ texts based on differences in vocabulary diversity. Texts written by students in higher grades offer a potentially richer basis for assessment in this regard. This could explain the more accurate LLM assessment of this characteristic in studies with older students (e.g.,
Mizumoto & Eguchi, 2023).
All in all, the present study indicates that a blanket distinction between ‘easily assessable’ linguistic text characteristics and ‘less easily assessable’ content text characteristics is not viable for LLM assessments of first-grade students’ texts. Instead, it seems more promising to consider the specific assessment criteria individually. In doing so, it also seems expedient to take greater account of the way LLMs work, as the way LLMs solve tasks makes them predestined for the assessment of certain text characteristics—but less for the assessment of others (see above). Accordingly, in the future, LLMs should be predominantly used for the assessment of text characteristics for which they are suitable to assess according to their functioning. In this way, LLMs may offer assistance to teachers in early elementary classrooms. LLMs may, for example, generate feedback on students’ texts, which could be tailored to individual students’ needs by the teacher (
Xiao et al., 2025). As a result, LLMs can help reduce teachers’ workload (
Nkoyo et al., 2025) and, at the same time, set the ground for providing individual feedback to students (
Steiss et al., 2024). However, it is important that LLMs do not replace but supplement and enrich teachers’ practices in this regard (
Jukiewicz & Wyrwa, 2026). At this point, one should always consider that the way LLMs come to their conclusions differs significantly from human cognitive processes (
Wang et al., 2025)—even if LLMs may imitate human-like thinking—and thus should be carefully situated by teachers when used for educational purposes.
6. Limitations and Outlook
This study has several limitations that offer starting points for future research. For example, we used inter-rater reliability between human and LLM assessments as an indicator of assessment accuracy. Human assessments thus served as a baseline for LLM assessments. However, this can be viewed critically as human raters are shown to be affected by biases—some of which are even reproduced by LLMs (
Gallegos et al., 2024). Consequently, we do not capture LLM assessment accuracy in an absolute sense, but in terms of its internal consistency and alignment with human raters. Therefore, future benchmarks should be developed that can serve as additional baselines for determining the accuracy of LLM assessments (
Novikova et al., 2025). Moreover, we did not consider the validity of LLM and human assessments. It is thus conceivable that human raters and LLM ratings differ somewhat in their construct understanding of the criteria. In future studies, the validity of assessment should also be taken into account. Another limitation of the study concerns the operationalization of the assessment criteria. These criteria were drawn from the NaSch1 study, in which they were adapted and specified for the elementary school context. However, it should be noted that the theoretical constructs behind these criteria are more comprehensive than the facets reflected in the criteria from NaSch1. Originality, for example, encompasses not only the facets of novelty and authenticity, which are reflected in the respective criterion from NaSch1, but also the uniqueness of the idea (
Corazza, 2016). Since first-grade students’ texts can only be expected to be unique to a very limited extent, this aspect is not reflected in the NaSch1 criterion of
originality. The present results are therefore only valid for LLM assessments of originality of first-grade students’ texts, not for LLM assessments of originality in general. This also applies to other assessment criteria. To be able to make statements about the appropriateness of LLM assessments in other contexts, future studies should specifically focus on these contexts. Further limitations concern the skewness of LLM assessment data: descriptive statistics point to floor effects of LLM assessments of
originality, which may have affected the results to some degree. A comparable pattern emerges from the LLM assessments of
perspective taking. Here, descriptive statistics point to ceiling effects, which may have influenced the results as well. In this context, it is also important to note that the analyses of
perspective taking are based on a slightly smaller sample than the analyses of the other criteria. Moreover, some methodological decisions must be considered when reading the findings of the present study. For example, we used the mode as a measure of assessment tendency. Consequently, we were only able to approach overall LLM assessment scores, not to determine LLM scores that claim absolute validity. Additionally, further analyses could yield results under different temperature settings. Possibly, ChatGPT rates differently when forced to more “logical” and conservative processing with low(er) temperature settings. Moreover, future research could consider other LLMs such as Gemini, DeepSeek, or Claude and examine the extent to which the present findings can be replicated using these LLMs. Finally, there are some technical limitations left to be considered, including the lack of full reproducibility and repeatability of LLM outputs (
Shyr et al., 2026) and a general instability of LLM assessments over time (
Pack et al., 2024). These limitations must be taken into account, too, when interpreting the findings of the present study.