Next Article in Journal
A Rolling Wavelet-Denoised TENET Framework for Analyzing Tail-Risk Spillovers in Crypto-Equity Complex Systems
Previous Article in Journal
Memory-Dependent Shifts of the Period-Doubling Cascade in the Finite-Memory Grünwald–Letnikov Fractional Logistic Map
Previous Article in Special Issue
Research on UAV 3D Airspace Signal Strength Prediction Based on Physical Perception Feature Engineering
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

LLM-Assisted Scoring for College English Writing Assessment: Statistical Calibration Against Teacher Standards

1
School of General Education, Xihang University, No. 259 Xierhuan Road, Xi’an 710077, China
2
Faculty of Applied Sciences, Macao Polytechnic University, Macao SAR, China
*
Author to whom correspondence should be addressed.
Mathematics 2026, 14(17), 3033; https://doi.org/10.3390/math14173033 (registering DOI)
Submission received: 16 July 2026 / Revised: 18 August 2026 / Accepted: 19 August 2026 / Published: 23 August 2026
(This article belongs to the Special Issue Applications of Machine Learning and Pattern Recognition)

Abstract

Large classes in Chinese College English programmes make frequent analytic assessment of student writing difficult. Large language models (LLMs) may support more frequent formative assessment, but their scores may vary across queries and be systematically harsher or more lenient than local teacher ratings. Using a corpus-based, five-fold cross-validated comparative rater-evaluation design, this study examined whether statistical calibration could make LLM-assisted scores more interpretable for College English writing assessment and where their use should remain limited. Data comprised 414 timed argumentative essays written by Chinese non-English majors at one applied undergraduate institution. Two trained College English teachers independently rated the essays on a seven-dimension analytic rubric informed by China’s Standards of English Language Ability, providing the local reference standard. Three LLMs rated each essay–dimension pair on five occasions. Under five-fold cross-validation, uncalibrated scores were compared with location–scale correction, isotonic calibration, and equipercentile linking, using quadratic weighted kappa, Spearman correlation, mean absolute error, signed bias, and half-point tolerance accuracy. Agreement between models did not imply agreement with teachers: two models showed inter-model kappa values of 0.70–0.78 but an average kappa of only 0.15 with teacher ratings while rating the essays about one band more severely. Calibration removed most of this severity difference and raised pooled kappa to 0.61–0.70 depending on the method (0.63–0.64 under equipercentile linking), compared with a teacher–teacher agreement benchmark of 0.747. The three methods differed little, and the improvement mainly reflected closer alignment of score distributions rather than better judgement of writing quality. Agreement was higher for vocabulary, syntax, and grammar but remained low for cohesion and conventions. The findings suggest that LLM-assisted scoring may support low-stakes formative feedback when calibrated to local teacher standards and used under teacher supervision, while teachers retain responsibility for judging content, coherence, argumentation, and communicative quality.

1. Introduction

Frequent analytic assessment of student writing is difficult to sustain in many Chinese College English programmes. College English instructors may teach two or three classes of more than 50 students each, so a writing task assigned across their classes can generate more than one hundred scripts for a single instructor. Such large-class conditions create substantial marking demands and reduce opportunities for frequent, individualised feedback [1,2]. Under these conditions, writing tasks may be assigned only a few times each semester, and feedback is often limited to a holistic score or brief comments. This workload constrains the formative assessment encouraged in the Guidelines on College English Teaching [3] and the timely and specific feedback that formative assessment research associates with learning improvement [4]. Automated writing evaluation (AWE) systems such as Pigai, a Chinese online writing evaluation platform, can provide immediate feedback on large numbers of essays, particularly on grammar, collocation, and mechanics. However, studies have reported uneven feedback accuracy and continued reliance on teacher judgement [5,6].
Large language models (LLMs) extend the range of automated support because they can assign scale- or rubric-based scores to second-language (L2) essays [7,8,9,10]. Their educational value, however, depends on how they are incorporated into the assessment process and how teachers interpret and use their outputs. Research on technology-supported instruction in other subjects similarly emphasises the role of teachers’ pedagogical decisions, motivational strategies, and interaction with students [11]. In L2 essay scoring, prompting strategies, calibration examples, and repeated scoring can affect score consistency and agreement with human ratings [8,9,10]. More general LLM-as-a-judge research also reports position, verbosity, and self-enhancement biases [12]. From a language-assessment perspective, such variation can be treated as a rater-management problem. Operational writing assessments address comparable problems through rater training with benchmark scripts [13]; that is, previously rated exemplar essays used to clarify score-level standards, through double rating and consistency monitoring [13,14] and through statistical modelling of rater severity [14,15]. The central question for College English programmes is therefore not whether an LLM can produce a score, but whether its scores can be aligned with local teacher standards and limited to the kinds of judgement they can support.
To address this question, the study treats each LLM as an auxiliary ordinal rater and evaluates a teacher-referenced procedure for making its scores interpretable on the local assessment scale. Two sources of benchmark essays are examined in separate stages: panels selected through agreement between two separate LLMs and panels based on exact agreement between the two College English teachers. Repeated ratings are used to examine within-model consistency. The resulting scores are then calibrated to the teacher reference scale using equipercentile linking [16,17], with location–scale correction and isotonic calibration included as comparison methods. An instability-based rule is also evaluated as a means of identifying ratings for teacher review. These procedures address differences in consistency, severity, score distribution, and scale use. They do not establish that an LLM understands writing quality in the same way as a teacher, and this boundary defines the scope of the claims made here.
The empirical analysis uses the Corpus of College English Writing (CCEW-414), which contains 414 timed classroom essays written by first- and second-year Chinese non-English majors at one applied undergraduate institution in Northwest China. Each essay was independently rated by two trained College English instructors using a seven-dimension analytic rubric informed by China’s Standards of English Language Ability (CSE). The two independent ratings provide the teacher data against which the LLM scores are evaluated; the construction of the reference score is described in Section 3.3. Three LLMs from different model families rated the same essays under repeated scoring conditions. The analysis examines whether agreement among LLMs corresponds to alignment with teacher judgement, whether statistical calibration reduces systematic severity differences, how agreement varies across writing dimensions, and whether rating instability can be used to identify cases requiring teacher review.
By placing local teacher judgement at the centre of model evaluation, the study makes three related contributions. First, it distinguishes inter-model agreement from alignment with teacher reference scores and compares how benchmark panels validated through machine agreement and teacher agreement are associated with raw scoring performance. Second, it compares alternative calibration methods and separates improvements in scoring severity and scale use from evidence of dimension-level alignment, while also examining whether instability across repeated rating occasions can support targeted teacher review. Third, for College English teaching, it provides an empirical basis for the selective formative use of LLM-assisted scores: relatively stable lexical and grammatical dimensions may inform diagnostic profiling and targeted feedback, whereas content, coherence, argumentation, communicative quality, and other discourse-level judgements remain under direct teacher interpretation. Together, these contributions position LLM-assisted scoring as a locally calibrated and teacher-governed resource for low-stakes writing assessment rather than as an autonomous replacement for teacher judgement.

2. Literature Review

2.1. Writing Assessment in College English: The Formative Gap

Formative assessment uses evidence about learning to adapt teaching and provide feedback that can guide subsequent work [4]. It is distinguished from summative assessment, which certifies attainment at the end of a course or unit, by the use to which the evidence is put rather than by the instrument itself. In the Chinese College English context, the Guidelines on College English Teaching call for programmes to combine formative and summative assessment and to develop students’ written communication ability [3]. The China’s Standards of English Language Ability (CSE) provide national, level-based descriptors of written expression that can inform local rating rubrics [18], while the College English Test (CET) syllabus provides a familiar reference for the writing construct used with non-English majors [19].
Within second-language (L2) writing instruction, written feedback connects assessment with revision and subsequent teaching decisions. Research reviews indicate that L2 writers value teacher written feedback and may use it to improve later writing [20]. Studies of Chinese university L2 writers further show that engagement with teacher and automated feedback depends on how learners understand the comments and apply them during revision [5,6]. Providing individualised feedback on both language and meaning, however, requires substantial teacher time, and case studies of large classes report that teachers adapt or reduce assessment-for-learning practices when marking loads become heavy [1].
One way to organise assessment evidence for formative use is through an analytic rating rubric. A rating rubric specifies the dimensions on which a performance is judged, the ordered score levels available for each dimension, and the descriptors associated with those levels. Whereas a holistic rubric assigns a single score to the text as a whole, an analytic rubric produces separate scores for different aspects of writing. This dimension-level information can help teachers and students identify areas requiring revision or further instruction, particularly when the scores are interpreted alongside the text and the instructional aims of the course [13,21]. Analytic scores do not replace written feedback or teacher judgement, but they can provide a structured basis for deciding where more detailed explanation and teaching attention are needed. Their formative value, however, also comes with a practical cost: rating several dimensions for every essay requires substantially more teacher time than assigning a single holistic score.
This creates a gap between the formative assessment encouraged in College English policy and the feedback that teachers can regularly provide. The problem is not only how to increase scoring capacity, but also how to ensure that any additional assessment reflects the writing construct, learner population, and instructional priorities of the local programme [21].

2.2. Automated Support for L2 Writing Assessment

Automated essay scoring (AES) refers to the assignment of scores to written responses by a computational system; automated writing evaluation (AWE) refers more broadly to systems that return feedback as well as, or instead of, scores. A machine rater, as the term is used in this article, is any such system considered in the role that a human rater would otherwise occupy: it receives a script and a rubric and returns an ordinal score on that rubric. Earlier AES and AWE systems commonly used engineered linguistic features to represent aspects of writing performance. More recent hybrid approaches have combined contextual embeddings with handcrafted linguistic features to capture both textual meaning and observable language characteristics [22]. In classroom-facing systems such as Pigai, automated feedback has often focused on language form, mechanics, and other features that can be identified directly from the text. The usefulness of such feedback depends on its accuracy, clarity, and relevance to the revisions students are expected to make [5,6].
Automated scoring cannot be evaluated only by whether it produces a score. Psychometric guidelines emphasise comparison with human ratings, examination of possible bias, clear definition of the intended score use, and continuing monitoring under the conditions in which the system is applied [23,24]. These requirements are particularly important in formative assessment, where scores may influence feedback, revision priorities, and later teaching decisions.
Large language models extend the range of automated support because they can apply analytic rubrics to several aspects of an essay without task-specific model training. Mizumoto and Eguchi [7] reported that a GPT-family model could score the essays of the TOEFL11 corpus with meaningful agreement, although its performance remained below human-rater levels. Yancey et al. [8] found that GPT-4 agreement with CEFR ratings of short L2 essays increased when calibration examples representing different score levels were included in the prompt. In an instructional assessment setting, Kim [9] found that prompting strategy and consistency across repeated ratings affected the alignment of GPT-4 placement scores with human ratings. Yavuz et al. [10] also reported high reliability in a small-scale comparison of rubric-based LLM and human grading of EFL essays.
Together, these studies indicate that LLMs may support more frequent analytic assessment. They also leave unresolved whether model scores reflect the construct and standards of a particular teaching programme. Agreement obtained with an external rubric or benchmark dataset does not necessarily indicate that the same model will apply a locally developed rubric in the way College English teachers intend.
Related concerns appear in the broader literature on LLMs used as evaluators. Model judgements may be influenced by response position and verbosity [12], while ratings across repeated occasions may not always remain consistent [9]. In predictive modelling more broadly, post hoc methods such as isotonic regression and temperature scaling have been used to recalibrate model confidence [25]. Conformal prediction has also been adapted to represent uncertainty in LLM-based ordinal judgements [26]. These methods may adjust model outputs or identify uncertain cases, but they do not by themselves establish alignment with a locally defined writing construct.
Public corpora such as ELLIPSE, PERSUADE 2.0, and ASAP 2.0 provide important benchmark resources for automated essay scoring research [27,28,29]. Evidence obtained from these corpora, however, does not directly establish how an LLM will evaluate shorter classroom essays written by Chinese non-English majors or how its scores relate to a local College English rubric. Differences in essay length, task conditions, learner population, scoring criteria, and score distributions may all affect agreement. The remaining issue is therefore not simply whether machine scores can be made more consistent, but how they can be anchored to locally defined teacher standards and governed as part of an educational assessment process.

2.3. Rater Effects, Benchmark Training, and Score Linking

Language assessment provides established procedures for addressing variation in writing ratings. Writing scores are mediated by raters, whose interpretations of a rubric and levels of severity may differ; rater severity denotes a rater’s general tendency to award lower or higher scores than other raters applying the same rubric to the same scripts. Norming, the training procedure in which raters discuss and score a common set of scripts before operational marking, can help raters develop a shared understanding of score levels. Benchmark scripts, also called exemplar or sample scripts, are previously rated essays whose agreed scores illustrate each level of the scale and are used during norming. Double rating makes disagreement visible, while adjudication provides a procedure for reviewing substantial differences, normally by a third, more senior rater. Measurement models such as many-facet Rasch measurement can also be used to estimate differences in rater severity [13,14,15].
When two score scales are intended to represent the same underlying construct, statistical linking can make their results more comparable. Equipercentile linking maps scores from one distribution to another by treating scores with the same percentile rank as corresponding values [16,17]. In this way, a systematically harsher or more lenient rating distribution can be placed on a reference scale. Such linking addresses differences in score location, spread, and distribution rather than differences in how writing quality is interpreted.
The LLM-scoring studies reviewed above provide evidence about agreement with human ratings, but less information about continuing processes of standard setting, consistency monitoring, statistical calibration, and score review. Similar quality-management principles can be adapted when an LLM is used as an auxiliary rater. Teacher-agreed benchmark essays can make the local standard visible in the scoring prompt. Repeated ratings can provide evidence about whether the model applies that standard consistently. Statistical calibration can adjust systematic differences between machine and teacher score distributions, while uncertain or disputed cases can be referred for teacher review [13,14,15,16,17]. The purpose of these procedures is to make machine scores more interpretable on a locally defined scale and to keep decisions about standards, score use, and disputed cases under teacher control.

2.4. Research Focus and Questions

Against this background, the study examined whether and under what conditions LLM-assisted scoring could support formative writing assessment in College English when referenced to independent ratings by two trained instructors. Four research questions guided the analysis. First, to what extent did raw LLM ratings agree with the teacher reference scores, both overall and across writing dimensions, and did agreement between models indicate alignment with teacher judgement? Second, how did benchmark referencing affect raw-score consistency and teacher-referenced alignment, and what differences were observed between benchmark panels validated through machine agreement and those validated through teacher agreement? Third, how did location–scale correction, isotonic calibration, and equipercentile linking differ in correcting systematic severity and score-distribution differences, both overall and by writing dimension? Fourth, did instability across repeated rating occasions identify scores with larger teacher-referenced errors, and what proportions of essay–dimension pairs and essays would require teacher review? As supplementary analyses, the study also examined whether the calibrated scores agreed equally well with each of the two teachers considered separately, and whether the main agreement patterns were consistent across gender and essay-length subgroups.

3. Materials and Methods

3.1. Corpus and Essay Profile

This study used a corpus-based, cross-validated comparative rater-evaluation design to examine the agreement, calibration, and repeatability of LLM-assisted scores against a local teacher reference. The design is comparative in that several machine raters and several calibration methods were evaluated against the same teacher reference on the same held-out data; it is a rater-evaluation design in that the object of study is the behaviour of the raters rather than the writing development of the students. No instructional intervention was administered, and no experimental manipulation was applied to the students.
The data were drawn from the Corpus of College English Writing (CCEW-414), which comprises 414 timed argumentative essays written in regular College English courses by first- and second-year non-English majors at an applied undergraduate institution in Northwest China. The essays were collected over two academic years, with each student contributing one essay. Students wrote by hand for 30 min under supervised classroom conditions and were not permitted to use dictionaries, online resources, electronic devices, or large language model tools. The five writing topics were drawn from course instruction: time management, technology and life, smartphone influence, digital media, and staying up late.
The essays were produced as part of regular course activities. Students were informed that their writing would be included in the corpus and used for research purposes, and they agreed to its use. The essays were submitted anonymously, and anonymity was maintained throughout the rating and analysis. The analytic corpus included all eligible essays completed under the stated conditions. No random sampling or score-based filtering was applied. Blank scripts, duplicate submissions, and scripts with substantial missing or illegible content were not eligible for inclusion. Essays were not filtered by length, writing quality, language accuracy, or subsequent teacher ratings. Of the 414 essays, 300 were written by male students and 114 by female students, reflecting the enrolment profile of the institution. All 414 essays entered every analysis reported below; under the cross-validation design each essay contributed to the training partition of four folds and to the held-out partition of exactly one fold.
Essay length averaged 133.1 words (SD = 28.7; range = 53–240), and the essays contained on average 3.6 paragraphs (SD = 1.0; range = 1–7). At the essay level, the mean sentence count was 10.7 (SD = 3.6), and the mean sentence length was 13.4 words (SD = 4.3). The mean type–token ratio was 0.604 (SD = 0.093), and the mean bidirectional measure of textual lexical diversity (MTLD) was 73.6 (SD = 40.4). MTLD has been shown to be relatively insensitive to text length [30]. These indices describe the observed corpus and do not constitute an independent classification of student proficiency.
Approximately 11.6% of the essays began with “Nowadays…” or “With the (rapid) development of…”, while 20.3% began with one of a broader set of textbook-style openings, including “As we all know…” and “In recent years…”. The relatively short texts and recurring textbook-style openings compress the range of observable quality differences that any rater, human or machine, must discriminate, and were taken into account when interpreting agreement statistics. Text-level indices recorded for each essay included token, type, sentence, and paragraph counts; type–token ratio and MTLD; mean sentence length; and counts of selected connectors, complex structures, and surface error patterns.

3.2. Analytic Rating Rubrics

Two analytic rubrics were used at different stages of the study. The teacher rubric provided the local reference framework for the main analyses, whereas a preliminary six-dimension machine rubric was used only in the initial machine-anchored comparison. The two rubrics were not treated as construct-equivalent; comparisons were limited to dimensions with sufficient overlap in their descriptors.
The teacher rubric adopted the overall proficiency score and six analytic dimensions/traits of the English Language Learner Insight, Proficiency and Skills Evaluation (ELLIPSE) framework [27]—overall writing quality, cohesion and coherence, syntactic ability, vocabulary, phraseology, grammatical accuracy, and conventions—together with its 1–5 scale in half-point increments. The ELLIPSE framework was selected because it is a documented, peer-reviewed analytic scheme developed for the rating of English language learner writing, with published rater training materials and level descriptors. The level descriptors used in the present study were rewritten with reference to the 2018 edition of China’s Standards of English Language Ability, which was the version in force during rubric development, together with the Guidelines on College English Teaching and the College English Test writing specifications [3,18,19], and were localised to the writing profile and instructional expectations of the study population. The structure of the revision followed established principles for analytic rating-scale development in second-language writing assessment [13,21]. For example, a score of 3 for overall writing quality indicated that the essay conveyed the basic message but showed limited development and noticeable structural and language problems.
The analytic design was intended to produce a dimension profile rather than only a single holistic mark. Such a profile can identify aspects of writing for which teacher explanation, student revision, or further instruction may be needed.
The preliminary machine rubric was a study-specific adaptation of the same ELLIPSE framework, and Table 1 documents each step of that adaptation. Three changes were made. First, pairs of closely related source dimensions were merged so that the preliminary rubric would remain short enough to be applied by a single model call per dimension: vocabulary and phraseology were combined into vocabulary and expression, and grammar and syntax were combined into grammar and syntax. Second, two dimension labels were renamed to match the terminology used in College English teaching documents in China: cohesion became organisation and coherence, and conventions became mechanics. Third, one dimension with no counterpart in ELLIPSE, content and task fulfilment, was added to cover response to the writing task, a criterion that is explicit in the CET writing specifications [19] but is not represented in the ELLIPSE analytic set. The resulting rubric was therefore a study-specific operational rubric rather than a direct reproduction of the ELLIPSE framework.
The machine rubric employed integer scores from 1 to 5 across the six dimensions. Integer scoring was adopted in this preliminary stage because the machine-agreed benchmark panels were defined by exact agreement between two independent models at each score level, and exact agreement is attainable at a useful rate only on a coarse scale. The choice has a measurable cost: an integer scale cannot represent the half-point distinctions the teachers used, and it therefore places a mechanical limit on agreement with a teacher reference expressed in quarter points. The machine-anchored stage is consequently reported as a preliminary comparison, and all substantive conclusions are drawn from the teacher-anchored stage, in which the models scored on the teachers’ own half-point scale.
Five machine-rubric dimensions had descriptor-level counterparts in the teacher rubric: overall quality, organisation and coherence, vocabulary and expression, grammar and syntax, and mechanics. The machine dimension of content and task fulfilment and the teacher dimensions of syntactic ability and phraseology had no one-to-one counterparts. The correspondences in Table 1 indicate overlap in descriptor focus rather than construct equivalence. The five mapped dimensions were used only for descriptive comparisons in the machine-anchored stage. In the teacher-anchored stage, the LLMs applied the seven-dimension teacher rubric directly, without cross-rubric mapping.

3.3. Teacher Reference Scores

Two College English instructors, each with more than eight years of teaching experience, independently rated all 414 essays on the seven-dimension analytic rubric after norming on the dimension descriptors and the distinction between adjacent score levels. Neither instructor had access to the other’s scores, and both original ratings were retained in the dataset.
Under the rating protocol, essays for which the two overall scores differed by more than one point were flagged as large-disagreement cases; this occurred for 10 of the 414 essays (2.4%). No third-rater adjudication was conducted, and neither original rating was revised or replaced. The disagreement flag was retained in the dataset as a quality indicator, and all affected essays remained in the analysis. For each essay–dimension pair, the arithmetic mean of the two independent ratings was used as the teacher reference score because the two raters had equivalent training and status. The resulting mean is treated as a practical local reference rather than an error-free gold standard.
Each individual teacher score lay on the half-point grid from 1 to 5, so the mean of two such scores could also take quarter-point values. This resolution was retained in the error analyses; for quadratic weighted kappa, which requires a common category set, reference scores were assigned to the nearest half-point category, and all reported kappa values for every condition were computed in the same way.
Table 2 reports agreement between the two instructors. Dimension-level quadratic weighted kappa (QWK) ranged from 0.376 for overall writing quality to 0.670 for vocabulary, with a mean of 0.561 across the seven dimensions. Exact agreement ranged from 0.345 to 0.437, while the proportion of ratings within half a point ranged from 0.819 to 0.860. Pooled QWK was 0.724 for the five mapped dimensions used in the machine-anchored comparison and 0.747 across all seven dimensions used in the teacher-anchored comparison. Approximately two thirds of the teacher reference scores fell between 2.5 and 3.5. Because chance-corrected agreement statistics depend partly on marginal score distributions and score variation, these values were interpreted together with exact agreement and half-point tolerance [31,32].
The teacher–teacher values are used throughout as an empirical agreement benchmark for this corpus and not as a strict statistical ceiling. A machine rater compared with the mean of two teacher ratings is compared with a more reliable quantity than either teacher rating alone, and can in principle agree with that mean more closely than the two teachers agree with each other. Section 4.5 therefore also reports agreement with each teacher’s original ratings separately.
Across the seven teacher-rubric dimensions, the instructors differed by more than half a point on 15.8% of essay–dimension pairs on average (15.7% across the five mapped dimensions). This proportion was used as the pair-level workload reference for the teacher-review rule defined in Section 3.5.2.

3.4. LLM Raters and Rating Design

LLM scoring was organised in two stages that differed in the source of the rubric and benchmark essays. The machine-anchored stage examined scoring based on the six-dimension machine rubric and machine-agreed benchmark essays. The teacher-anchored stage examined scoring based on the local seven-dimension teacher rubric and teacher-agreed benchmark essays. Both stages included a rubric-only condition so that the contribution of benchmark referencing could be examined separately.
Figure 1 summarises the overall procedure, which combined fold-specific benchmark selection, repeated scoring, score aggregation, statistical calibration, and teacher-review flagging. No model was fine-tuned, and no model parameters were accessed. All benchmark panels and calibration functions were constructed from the training partition of each cross-validation fold, and target essays in the held-out partition were scored without access to their teacher ratings.
The two stages differed in rubric structure, score resolution, benchmark source, and model set. Their results are therefore interpreted within stage rather than as a controlled estimate of the independent effect of benchmark source. The machine-anchored stage examined whether agreement among models provided an adequate basis for defining the rating standard; the teacher-anchored stage examined scoring based directly on the local rubric and benchmark essays accepted by the two College English teachers.

3.4.1. Models, Parameters, and Prompt Design

Three deployable LLM raters from different model families were evaluated: DeepSeek-V3, GLM-4-Air, and Qwen3-Max. The machine-consensus benchmark-selection panel consisted of DeepSeek-R1 and GLM-4-Plus. The models were accessed through their providers’ public OpenAI-compatible interfaces. Provider-side model identifiers, endpoints, access dates, recorded decoding settings, prompt templates, and retry procedures are reported in Appendix B.
The deployable raters were called with a temperature of 0.3, and the two benchmark-selection models with a temperature of 0.2. In the machine-anchored stage, DeepSeek-V3 and GLM-4-Air applied the six-dimension machine rubric and returned integer scores from 1 to 5. In the teacher-anchored stage, these two models and Qwen3-Max applied the seven-dimension teacher rubric and returned scores from 1 to 5 in half-point increments. Each call evaluated one dimension of one target essay.
The prompt contained a role instruction, the descriptors for the relevant dimension, the permitted score levels, and the target essay. In benchmark-referenced conditions, the prompt also presented the selected benchmark essays and their agreed scores in ascending order. The model was required to return a score in a structured format together with a one-sentence rationale.
The machine-consensus panel rated every essay once on the six-dimension machine rubric without access to teacher scores or to the other panel model’s output. Its ratings were used to select machine-agreed benchmark essays and to examine separately whether agreement between models corresponded to agreement with the teachers.

3.4.2. Benchmark Essay Panels

Benchmark panels were constructed separately within the training partition of each fold. No essay from the held-out partition could be selected as a benchmark for its own evaluation fold. Once selected, the same panel was used for all target essays rated under the corresponding fold, dimension, and experimental condition. Fixed consensus panels of this kind reproduce the benchmark-script logic of operational rater training, in which all raters are shown the same exemplars, and they differ deliberately from example-selection strategies that retrieve a different reference text for each target by lexical similarity [33]. A fixed panel can also be read in full by a teacher, so the standard given to the model is inspectable.
For the machine-anchored stage, two benchmark essays were selected for each of the five integer score levels on each of the six machine-rubric dimensions. Eligible essays were those for which DeepSeek-R1 and GLM-4-Plus assigned exactly the same score on the relevant dimension. When more than two eligible essays were available, the two essays closest to the corpus median length were selected to reduce systematic length differences across benchmark levels. This procedure produced a ten-essay panel for each machine-rubric dimension.
For the teacher-anchored stage, benchmark essays were selected from scripts on which the two teachers had assigned exactly the same score on the relevant teacher-rubric dimension. One essay was selected for each half-point level represented among the eligible scripts, and where several essays were eligible at the same level the essay closest to the corpus median length was selected. Because exact teacher agreement was not observed at every half-point level of every dimension, the resulting panels contained between four and seven essays rather than the nine that a fully populated scale would allow; the attested levels for each dimension are listed in Appendix B, Table A6.
The machine-agreed and teacher-agreed panels were not assumed to embody the same standard merely because their component essays had received internally consistent ratings. Machine agreement indicated consistency between the two panel models, whereas teacher agreement linked the benchmark essays to the standards applied in the local College English programme.

3.4.3. Repeated Rating and Score Aggregation

Each essay–dimension pair was rated on five occasions under the same prompt condition. The five occasions were independent in the following sense: each occasion was issued as a separate stateless request to the provider’s chat-completions endpoint, containing a single user message with the full prompt and no conversation history; no score, rationale, or identifier from any previous occasion was included in the request; and no provider-side session, thread, or caching option was enabled. The five requests for a given pair were therefore identical in content and differed only in the model’s sampling behaviour at a temperature of 0.3.
The median of the five occasion scores was used as the aggregated raw machine score, because the median retains the original score scale and limits the influence of a single atypical response. Two additional indicators described consistency across occasions: the occasion range, defined as the difference between the highest and lowest of the five scores, and the modal proportion, defined as the proportion of the five occasions producing the most frequently assigned score, which can therefore take the values 0.2, 0.4, 0.6, 0.8, and 1.0. A smaller range and a higher modal proportion indicate greater consistency across repeated ratings. These three quantities are defined formally in Appendix C, Equations (A1) and (A2).
These indicators describe within-model stability rather than agreement with teacher judgement. A model may repeatedly assign the same score while applying a rating standard that differs from that of the teachers. The one-sentence rationales were retained as part of the scoring record but were not used in score aggregation, calibration, or statistical analysis.

3.5. Score Calibration and Teacher Review

3.5.1. Calibration Methods

Three calibration methods were compared: location–scale correction, isotonic calibration, and equipercentile linking. Calibration was conducted separately for each model, rubric dimension, prompt condition, and cross-validation fold. All calibration parameters and mappings were estimated from the training partition and then applied without modification to the held-out partition. The three mappings are defined formally in Appendix C, Equations (A3)–(A5), and their treatment of interpolation, boundary values, and rounding is summarised in Appendix A, Table A3.
Location–scale correction provided the simplest comparison method. It adjusted the mean and standard deviation of the machine scores so that they matched those of the teacher reference scores in the training partition. The method therefore addressed general differences in rating severity and score spread, but it did not account for uneven relationships between individual score levels.
Isotonic calibration used paired machine and teacher reference scores in the training partition to estimate a nondecreasing score relationship that minimises squared error. The adjustment could vary across different parts of the scale, but higher raw machine scores could not be mapped below lower raw scores. Because it minimises squared error, the method also shrinks predictions towards the conditional mean of the reference scores, a property examined directly in Section 4.2.
Equipercentile linking related machine and teacher scores through their positions in the two training-fold score distributions. A machine score was assigned the value occupying approximately the same percentile position in the teacher reference distribution, with linear interpolation where percentile positions fell between observed score points. Linked values were restricted to the 1–5 reference range and rounded to the score grid used for reporting. Unlike the other two methods, equipercentile linking is estimated from the two marginal distributions rather than from paired observations, and the resulting mapping can be tabulated and inspected by a teacher.
The three methods represent different levels of score adjustment: average severity and spread, the relationship between paired scores, and the relative positions of the two distributions. Their comparison was intended to determine whether the observed improvements resulted specifically from equipercentile linking or could also be obtained through simpler or alternative forms of score correction.
A fixed strictly increasing score transformation does not change the rank ordering of the essays to which it is applied. In the present analysis, however, the mappings were estimated separately within each combination of fold and dimension, and rounding to a discrete score grid could create additional ties. A correlation computed over observations pooled across dimensions is therefore not the image of a single monotone transformation, and can change even when no within-dimension ordering is reversed. Spearman correlations were consequently examined at three levels: within each combination of fold and dimension, within each fold, and across the pooled out-of-fold predictions.
All three methods operate on score information rather than on the linguistic content of the essays. Reductions in severity, mean absolute error, or distributional mismatch are therefore evidence about score-scale alignment and not about the model’s judgement of content, argumentation, coherence, communicative effect, or overall writing quality.

3.5.2. Instability-Based Teacher-Review Rule

Repeated-rating instability was evaluated as a possible basis for directing selected scores to teacher review. An essay–dimension score was flagged when the range across the five rating occasions reached two scale points or when the modal proportion fell below a specified threshold; the rule is stated formally in Appendix C, Equation (A6). Rules of this general form, in which a computed indicator assigns items to differentiated downstream handling rather than treating all items alike, are used in other applied settings, for example in the priority-based classification of data units for protected transmission [34].
The modal-proportion threshold was determined from the training partition so that the proportion of flagged essay–dimension pairs remained as close as possible to, but did not exceed, the 15.8% pair-level workload reference derived from teacher disagreement. The resulting threshold was then applied to the held-out partition without further adjustment. Because the modal proportion takes only five distinct values, the attainable flagging rates form a coarse grid, and the rate obtained under the budget therefore differs across models. An essay was counted as requiring review if at least one of its seven dimension scores was flagged.
The review rule was evaluated by comparing teacher-referenced errors for flagged and unflagged scores. Quadratic weighted kappa, mean absolute error, and half-point tolerance accuracy were reported for both groups, together with essay-level paired bootstrap confidence intervals for the difference between them. The analysis also reported the accuracy of the scores that would have been released without teacher review, the proportion of essay–dimension pairs flagged, and the proportion of essays containing at least one flagged dimension. Repeated-score instability was treated as a screening indicator rather than as a complete measure of uncertainty, since a stable score is not necessarily accurate.

3.6. Evaluation Metrics and Statistical Analysis

Following psychometric guidelines for automated scoring evaluation [23,24], machine scores were compared with the teacher reference scores using five complementary indices: quadratic weighted kappa, Spearman rank correlation, mean absolute error, signed bias, and half-point tolerance accuracy. Complete definitions are given in Appendix A, Table A2.
Quadratic weighted kappa was used as the principal agreement statistic because the scores were ordinal, larger discrepancies required greater penalties, and chance agreement was non-negligible given the concentration of teacher scores within a restricted range. QWK is also widely used in automated essay-scoring evaluation. Because it is sensitive to marginal score distributions and score-range restriction [31,32], it was interpreted together with Spearman correlation, MAE, signed bias, and half-point tolerance accuracy.
Spearman correlation measured the similarity of the rank ordering produced by the machine and teacher reference scores, and was reported separately because a model may rank essays in a similar order while assigning scores at a different level of severity. Fold-level and within-dimension Spearman correlations were reported before and after calibration to distinguish changes caused by fold-specific mappings, by aggregation across dimensions, or by additional tied scores from genuine changes in rank ordering. Mean absolute error reported the average distance, in score points, between the machine score and the teacher reference score. Signed bias reported the average machine-minus-teacher difference, with negative values indicating greater severity than the teacher reference. Half-point tolerance accuracy reported the proportion of machine scores falling within half a point of the teacher reference score, the scoring interval used by the individual teachers.
Results were reported both across pooled essay–dimension observations and separately for each writing dimension. Dimension-level results provided the primary basis for educational interpretation because pooled statistics could conceal substantial differences among vocabulary, grammar, cohesion, conventions, and overall writing quality.
All analyses used stratified five-fold cross-validation at the essay level, stratified on the rounded teacher reference score for overall writing quality, with a fixed random seed. The five held-out partitions contained 83, 83, 83, 83, and 82 essays. Benchmark panels, calibration parameters, review thresholds, and score mappings were estimated only from the training partition of each fold. Each essay received one set of held-out predictions, and no teacher score from a held-out essay was used to construct the resources applied to that essay. The three calibration methods were compared using the same held-out raw machine scores.
Statistical uncertainty for the main comparisons was estimated through paired bootstrap resampling at the essay level with 1000 resamples. Resampling was conducted by essay rather than by individual dimension score so that the dependence among ratings of different dimensions from the same essay was retained. Ninety-five per cent confidence intervals were reported for the main differences between raw and calibrated conditions, between calibration methods, and between flagged and unflagged scores.
In addition to comparison with the mean teacher reference score, machine scores were compared separately with the original ratings of each teacher as a sensitivity analysis, and subgroup consistency was examined descriptively by gender and by essay-length terciles. The subgroup analyses were used to identify variation within the present corpus rather than to establish general fairness across student populations.

4. Results

4.1. Inter-Model Agreement and Teacher-Referenced Alignment of Raw LLM Scores

The preliminary machine-anchored comparison examined whether agreement between LLMs provided evidence that their ratings reflected the standards applied by the two College English teachers. DeepSeek-R1 and GLM-4-Plus showed high inter-model agreement across the six machine-rubric dimensions, with QWK values ranging from 0.695 to 0.778 and exact agreement of 77.7%. These values were higher than the corresponding teacher–teacher agreement values, indicating that the two models applied broadly similar rating standards.
Their agreement with the teacher reference scores was substantially lower. When the two panel ratings were combined for comparison with the teacher reference, mean teacher-referenced QWK was 0.152 across the five mapped dimensions. Dimension-level QWK was 0.097 for overall writing quality, 0.022 for organisation and cohesion, 0.122 for vocabulary, 0.427 for grammar, and 0.090 for mechanics and conventions. Grammar showed the closest correspondence with teacher ratings, whereas agreement was limited for overall writing quality, discourse organisation, vocabulary, and conventions.
The difference was also evident in rating severity. Across the five mapped dimensions, the panel models assigned scores 0.95 points below the teacher reference scores on average, with dimension-level signed bias ranging from −1.64 to −0.40. Only 26.1% of the machine-panel scores fell within half a point of the teacher reference score, compared with 84.3% of the paired teacher ratings on the same dimensions. As shown in Figure 2, the machine-panel score distribution was concentrated at lower score levels, while the degree of teacher-referenced agreement varied considerably across dimensions.
The raw rubric-only scores produced by the two deployable raters showed a similar directional pattern. DeepSeek-V3 obtained a pooled QWK of 0.086, a signed bias of −0.989, and a half-point tolerance accuracy of 0.223. GLM-4-Air showed closer, but still limited, teacher-referenced agreement, with a pooled QWK of 0.249, a signed bias of −0.656, and a half-point tolerance accuracy of 0.360. Both deployable raters therefore assigned systematically lower raw scores than the teachers, although the extent of severity and agreement differed between models.
These results show that inter-model agreement alone was insufficient to define the local rating standard. Because the preliminary stage used an integer machine rubric, its values are not directly comparable with those from the teacher-anchored analysis.

4.2. Effects of Benchmark Referencing and Score Calibration

Table 3 compares raw scores obtained with and without benchmark essays under the machine-anchored and teacher-anchored conditions. In the machine-anchored stage, machine-agreed benchmark panels produced small and inconsistent changes in teacher-referenced performance. For DeepSeek-V3, pooled QWK increased only from 0.086 in the rubric-only condition to 0.107 in the benchmark-referenced condition, while Spearman correlation increased from 0.164 to 0.181. MAE decreased from 1.180 to 1.133, and signed bias changed from −0.989 to −0.935. For GLM-4-Air, machine-agreed benchmarks reduced MAE from 0.926 to 0.802 and reduced signed severity from −0.656 to −0.432, and half-point tolerance accuracy increased from 0.360 to 0.456; however, pooled QWK decreased from 0.249 to 0.225 and Spearman correlation decreased from 0.352 to 0.251. The machine-agreed panels therefore affected score location and absolute error but did not consistently improve teacher-referenced agreement or rank ordering.
A clearer pattern was observed in the teacher-anchored stage. Teacher-agreed benchmark panels improved raw QWK for all three deployable raters. For DeepSeek-V3, QWK increased from 0.161 under direct scoring with the teacher rubric to 0.286 with teacher-agreed benchmarks. The corresponding increases were from 0.249 to 0.379 for GLM-4-Air and from 0.230 to 0.401 for Qwen3-Max. Spearman correlation also increased from 0.183 to 0.322, from 0.309 to 0.409, and from 0.301 to 0.468, respectively. The teacher-agreed panels reduced absolute error and systematic severity as well: MAE decreased from 0.808 to 0.747 for DeepSeek-V3, from 0.783 to 0.684 for GLM-4-Air, and from 0.835 to 0.686 for Qwen3-Max, and signed bias moved from −0.445 to −0.302, from −0.332 to −0.115, and from −0.418 to −0.194, respectively.
Teacher-agreed benchmark panels were consistently associated with closer raw alignment in the teacher-anchored stage. Because the two stages differed in several design features, the contrast is descriptive rather than causal.
Table 4 compares location–scale correction, isotonic calibration, and equipercentile linking using the same held-out raw scores under the teacher-benchmark condition. All three methods produced large and broadly similar improvements over the raw scores. Across the three models, pooled QWK after location–scale correction ranged from 0.614 to 0.637, compared with 0.668 to 0.697 after isotonic calibration and 0.629 to 0.644 after equipercentile linking, against raw values of 0.286 to 0.401. Corresponding MAE values were 0.475–0.484, 0.378–0.385, and 0.462–0.487, against raw values of 0.684–0.747. Signed bias, which ranged from −0.302 to −0.115 across the three models in the raw condition, was reduced to −0.027 to +0.028 after location–scale correction, −0.034 to +0.031 after isotonic calibration, and −0.150 to −0.072 after equipercentile linking. Half-point tolerance accuracy ranged from 0.688 to 0.699, 0.794 to 0.803, and 0.704 to 0.723 respectively, against raw values of 0.481–0.543. The same ordering of methods was obtained in the rubric-only condition, where pooled QWK reached 0.593–0.609 after location–scale correction, 0.661–0.687 after isotonic calibration, and 0.617–0.638 after equipercentile linking.
The improvement was therefore not specific to equipercentile linking. Paired bootstrap comparisons at the essay level showed that the differences among the three methods were an order of magnitude smaller than the difference between any of them and the raw scores. Relative to raw scores, equipercentile linking increased pooled QWK by +0.358 (95% CI [+0.337, +0.378]) for DeepSeek-V3, +0.265 ([+0.247, +0.284]) for GLM-4-Air, and +0.228 ([+0.212, +0.245]) for Qwen3-Max. Relative to location–scale correction, the corresponding differences were only +0.030 ([+0.018, +0.042]), +0.021 ([+0.009, +0.033]), and −0.008 ([−0.022, +0.007]). Relative to isotonic calibration, equipercentile linking was lower by −0.025 ([−0.043, −0.007]), −0.041 ([−0.057, −0.025]), and −0.068 ([−0.085, −0.050]), and also produced higher MAE by +0.077 to +0.103 points and lower half-point tolerance accuracy by 0.073 to 0.090. Because a correction based on two summary statistics recovers almost the whole of the gain, the improvement reflects the removal of systematic severity and scale-use differences rather than a property of any individual method.
Figure 3 shows the distributional effect of calibration. Before calibration, machine scores were concentrated below the teacher reference distribution; after calibration their location and spread more closely approximated the teacher scores.
The pooled advantage of isotonic calibration is accompanied by a compression of the score distribution, shown in Figure 4. Pooled across the seven dimensions, the standard deviation of the teacher reference scores was 0.719, and the standard deviations of the calibrated machine scores were 0.710–0.753 after location–scale correction and 0.694–0.717 after equipercentile linking, but only 0.553–0.594 after isotonic calibration. At the dimension level the compression was severe: for conventions, isotonic calibration returned the single value 3.0 for every essay in every model, giving a standard deviation of 0 and a dimension-level QWK of 0, while its MAE for that dimension was low because a constant close to the reference mean minimises absolute deviation in a narrow distribution. Its higher pooled QWK is obtained because differences in level between dimensions survive the pooling even when within-dimension discrimination has been removed. Equipercentile linking retained score dispersion more closely than isotonic calibration and was therefore used in the subsequent dimension-level, review-rule, and subgroup analyses. The contrast also shows that pooled agreement alone is insufficient for selecting a calibration method for classroom reporting.
Pooled Spearman correlations increased after calibration, including increases from 0.322 to 0.657 for DeepSeek-V3, from 0.409 to 0.647 for GLM-4-Air, and from 0.468 to 0.634 for Qwen3-Max in the teacher-benchmark condition after equipercentile linking. Figure 5 separates the levels at which this change occurs. Within each fold, the correlation computed over all seven dimensions rose in the same way as the pooled value, from 0.309–0.335 to 0.638–0.686 for DeepSeek-V3, from 0.369–0.447 to 0.577–0.701 for GLM-4-Air, and from 0.456–0.491 to 0.618–0.671 for Qwen3-Max. Within each individual combination of fold and dimension, where the calibration is a single monotone mapping, the correlation was essentially unchanged: the mean change across the 35 cells was +0.0004 for DeepSeek-V3, −0.0110 for GLM-4-Air, and −0.0011 for Qwen3-Max, and the largest change observed in any single cell was 0.113. Averaged over the seven dimensions, the within-dimension correlation did not improve, moving from 0.322 to 0.316 for DeepSeek-V3, from 0.343 to 0.328 for GLM-4-Air, and from 0.325 to 0.316 for Qwen3-Max.
Calibration was estimated separately for each dimension, so a correlation computed over observations pooled across dimensions is not the image of one monotone transformation. Before calibration, the seven dimensions were displaced from the teacher scale by different amounts, which misaligned the pooled ranking; correcting each dimension separately removes that misalignment and raises the pooled correlation without reordering any essay within any dimension. The small residual changes within individual cells follow from rounding to the discrete reporting grid, which increased the mean number of tied values per cell from 75.7–75.8 to 77.8–78.2. The increase in pooled rank correlation therefore indicates improved comparability of scores across dimensions rather than improved discrimination among essays.

4.3. Dimension-Level Teacher-Referenced Performance

Table 5 and Figure 6 present dimension-level results for the teacher-benchmark condition. Agreement differed substantially across the seven dimensions, and broadly similar patterns were observed for DeepSeek-V3, GLM-4-Air, and Qwen3-Max. The values quoted in this section are those obtained after equipercentile linking.
Vocabulary showed the highest teacher-referenced agreement, with QWK values ranging from 0.449 to 0.571 across the three models, compared with a teacher–teacher agreement benchmark of 0.670. Syntactic ability produced QWK values of 0.410–0.419, compared with a benchmark of 0.607. Grammatical accuracy ranged from 0.373 to 0.487, compared with a benchmark of 0.663. Phraseology showed more moderate agreement, with QWK values of 0.316–0.358, compared with a benchmark of 0.587.
Overall writing quality produced QWK values of 0.277–0.331. These values were lower in absolute terms than those for vocabulary, syntax, and grammar, but the corresponding teacher–teacher benchmark was also relatively low at 0.376, indicating that the reference score for this dimension is itself measured with more error than the others. The results do not establish that the calibrated models reproduced the teachers’ interpretation of overall writing quality, particularly for judgements involving content development, argumentation, and communicative effect.
The weakest teacher-referenced agreement was found for cohesion and coherence and for conventions. Cohesion and coherence produced QWK values of 0.057–0.130, compared with a benchmark of 0.580. Conventions produced values of 0.057–0.102, compared with a benchmark of 0.440. These low values were observed across all three model families and under every calibration method tested, including the raw scores, which produced QWK values of 0.008–0.080 for cohesion and 0.065–0.103 for conventions. The pattern is therefore not a by-product of the choice of calibration method.
For cohesion and conventions, calibrated MAE remained between 0.429 and 0.591 scale points, and signed bias remained within ±0.19 points. The relatively moderate absolute error, together with low QWK, indicates that distributional calibration placed many scores close to the teacher scale without reproducing the teachers’ distinctions among individual essays.
The comparatively stronger results for vocabulary, syntactic ability, and grammatical accuracy indicate that calibrated machine scores in these dimensions may provide more stable information for teacher-reviewed formative profiles, whereas the substantially lower agreement for cohesion, conventions, and aspects of overall writing quality indicates that these judgements should continue to depend primarily on teacher evaluation. The pooled values of 0.63–0.64 reported in Section 4.2 are consistent with dimension-level agreement ranging from near zero to approximately 0.57, so statements about classroom use must be made at the dimension level.

4.4. Instability-Based Review and Automatically Released Scores

The instability-based review analysis examined whether variation across the five rating occasions identified scores with larger teacher-referenced errors. The analysis was conducted under the teacher-benchmark condition after equipercentile linking, and the proportion of flagged essay–dimension pairs was constrained by the 15.8% pair-level workload reference derived from disagreements between the two teachers. Table 6 reports the results.
Because the modal proportion of five occasions takes only the values 0.2, 0.4, 0.6, 0.8, and 1.0, the attainable flagging rates are coarse, and the largest threshold satisfying the budget differed across models. The selected rule flagged 15.1% of essay–dimension pairs for DeepSeek-V3 (438 of 2898), 0.3% for GLM-4-Air (10 of 2898), and 10.2% for Qwen3-Max (296 of 2898). For GLM-4-Air the next available threshold would have flagged 16.9% of pairs, which exceeds the budget. At the essay level, at least one of the seven dimension scores was flagged for 66.7% of essays for DeepSeek-V3 (276 of 414), 2.4% for GLM-4-Air (10 of 414), and 53.6% for Qwen3-Max (222 of 414). The essay-level rates were much higher than the pair-level rates because a single essay contained seven independently evaluated dimensions; a pair-level flagging rate of about 15% distributed independently across seven dimensions implies that roughly two thirds of essays contain at least one flagged score.
For DeepSeek-V3, flagged scores showed somewhat lower teacher-referenced agreement than unflagged scores, with QWK values of 0.595 and 0.650, MAE of 0.489 and 0.457, and half-point tolerance accuracy of 0.674 and 0.730, respectively. The direction of this difference is consistent with the intended use of the rule, but its magnitude was small and the essay-level bootstrap interval included zero: the difference in MAE between flagged and unflagged scores was +0.032 (95% CI [−0.007, +0.070]).
The same pattern was not observed for the other two models. For GLM-4-Air, the ten flagged scores obtained a QWK of 0.503, an MAE of 0.475, and a tolerance accuracy of 0.700, compared with 0.644, 0.466, and 0.723 for the 2888 unflagged scores; the difference in MAE was +0.012 (95% CI [−0.148, +0.162]). For Qwen3-Max, flagged and unflagged scores produced QWK values of 0.637 and 0.628, MAE values of 0.490 and 0.486, and tolerance accuracy values of 0.699 and 0.704; the difference in MAE was +0.005 (95% CI [−0.039, +0.047]). At a common threshold applied to all three models, the flagged group was slightly more accurate than the unflagged group for two of the three: flagging every pair with a modal proportion of 0.6 or below gave flagged-versus-unflagged QWK values of 0.595 versus 0.650 for DeepSeek-V3 but 0.661 versus 0.641 for GLM-4-Air and 0.660 versus 0.627 for Qwen3-Max. Consistently with this, the rank correlation between the occasion range and the absolute teacher-referenced error was +0.038, −0.037, and +0.007 for the three models respectively.
The automatically released, unflagged scores obtained QWK values of 0.650 for DeepSeek-V3, 0.644 for GLM-4-Air, and 0.628 for Qwen3-Max, with MAE values of 0.457, 0.466, and 0.486 and half-point tolerance accuracy values of 0.730, 0.723, and 0.704. These figures describe the accuracy of the scores that would have been returned to students without teacher review under the specified rule. They are barely distinguishable from the accuracy of the full score set, which is the expected consequence of a screening rule that does not separate accurate from inaccurate scores.
The relationship between occasion-level instability and teacher-referenced error was therefore weak for DeepSeek-V3 and absent for GLM-4-Air and Qwen3-Max, and Figure 7 shows the flagged-minus-unflagged difference in mean absolute error, with its bootstrap confidence interval, for each model. The pair-level flag rates remained at or below the prespecified workload reference by construction, but this does not establish a reduction in teacher marking time: the practical review burden depends on the number of complete essays requiring review, and the essay-level rates of 66.7% and 53.6% for two of the three models indicate that a rule calibrated to a pair-level budget can nevertheless send the majority of essays to a teacher.

4.5. Comparison with Each Teacher Separately

The calibrated scores were compared with each teacher’s original ratings as well as with their mean, because a comparison with a two-rater mean is not equivalent to a comparison with an individual rater. Table 7 reports the results for the teacher-benchmark condition after equipercentile linking.
Agreement with either individual teacher was lower than agreement with their mean. Pooled QWK against Teacher 1 was 0.615, 0.619, and 0.597 for DeepSeek-V3, GLM-4-Air, and Qwen3-Max, and against Teacher 2 it was 0.603, 0.603, and 0.571, compared with 0.643, 0.644, and 0.629 against the mean. All of these values were below the teacher–teacher value of 0.747. The mean is a more reliable quantity than either constituent rating, so agreement with it is expected to be higher; no model reached the teacher–teacher level against either individual teacher.
The pattern was reversed for half-point tolerance accuracy, which was higher against each individual teacher (0.738–0.770) than against the mean (0.704–0.723), because an individual teacher rating falls on the half-point grid used by the models whereas the mean can take quarter-point values that no machine score can match exactly. The two indices therefore respond to different properties of the reference.
The two teachers did not differ greatly in the severity with which the models matched them. Signed bias against Teacher 1 was −0.036, −0.030, and −0.108 for the three models, and against Teacher 2 it was −0.119, −0.113, and −0.191, indicating that Teacher 2 rated slightly more leniently than Teacher 1 by approximately 0.08 scale points. Dimension-level agreement showed the same ordering of dimensions against each teacher separately as against the mean, with cohesion and conventions weakest in every comparison. The main findings therefore do not depend on the choice between the mean and either individual teacher as the reference.

4.6. Descriptive Subgroup and Essay Length Checks

Calibrated performance was examined descriptively across gender groups and essay-length terciles under the teacher-benchmark condition after equipercentile linking. For DeepSeek-V3, pooled QWK was 0.642 for the 300 essays written by male students and 0.648 for the 114 essays written by female students. The corresponding values were 0.641 and 0.654 for GLM-4-Air and 0.618 and 0.658 for Qwen3-Max. Signed bias remained within ±0.16 scale points in every group. Absolute error was slightly larger for essays written by male students in all three models, with MAE of 0.471, 0.475, and 0.496 against 0.437, 0.443, and 0.461 for essays written by female students. Signed bias did not follow a single direction: for male and female groups it was −0.092 and −0.040 for DeepSeek-V3 and −0.075 and −0.064 for GLM-4-Air, whereas for Qwen3-Max it was −0.147 and −0.157. All six values lie within a range of 0.12 scale points.
The differences between gender groups were modest, but they should be interpreted cautiously because the subgroup sizes were unequal, with 300 male and 114 female writers. These comparisons do not establish the absence of gender-related effects and do not provide evidence that would generalise beyond the study population.
Essay-length terciles were formed at 122 and 145 words, giving groups of 140, 140, and 134 essays. Across terciles, DeepSeek-V3 produced calibrated QWK values of 0.653, 0.654, and 0.607 for short, middle, and long essays; GLM-4-Air produced 0.655, 0.657, and 0.602; and Qwen3-Max produced 0.632, 0.644, and 0.598. Agreement was therefore slightly lower for the longest third of the corpus in all three models, but the differences were small and no monotonic relationship was observed. Signed bias remained within ±0.17 points in every length group for every model. Within this corpus, calibration did not appear to improve aggregate alignment by introducing a consistent advantage for longer or shorter essays.

5. Discussion

5.1. Local Teacher Standards and Statistical Calibration

The findings show that agreement among LLMs does not necessarily indicate alignment with the standards used in a local College English programme. Models may apply similar scoring criteria while differing from teachers in their interpretation of the rubric and their use of the score scale. For instructional assessment, the relevant reference is therefore the locally defined writing construct rather than consistency among machine raters.
Teacher-approved rubrics and benchmark essays provide a practical basis for anchoring LLM scoring to local expectations. The benchmark essays make score-level interpretations more concrete and allow teachers to inspect the standards presented to the model. Although teacher ratings are not error-free, they represent the operational standard of the programme in which the scores will be interpreted and used. Because the two experimental stages differed in several design features, their contrast remains descriptive rather than a controlled estimate of the effect of benchmark provenance.
Statistical calibration has a narrower function. It can reduce systematic differences in severity and make machine scores more interpretable on the local scale, but it does not change how the model evaluates the linguistic or rhetorical quality of an essay. That a two-parameter location–scale correction recovered almost the whole of the observed gain indicates how general this adjustment is. A calibrated score may be closer to the teacher score distribution without reproducing teacher judgements about content, argumentation, coherence, or communicative effect. Calibration should therefore be understood as score-scale adjustment rather than evidence that the model has acquired a teacher-like understanding of writing quality.

5.2. Dimension-Specific Formative Use and Teacher Governance

The educational value of LLM-assisted scoring varies across writing dimensions. Vocabulary, syntactic ability, and grammatical accuracy showed comparatively stronger alignment with teacher ratings. Under local validation, scores in these dimensions may provide preliminary information about lexical range, sentence construction, and grammatical control. Such information could support more frequent analytic profiles, help teachers identify recurring language problems, and provide a basis for targeted revision activities.
These scores should not be treated as complete diagnoses or as substitutes for written feedback. A dimension score does not explain why a particular expression is ineffective, how a sentence should be revised, or whether a language choice is appropriate for the task. Its instructional value depends on teacher interpretation, the student’s previous work, and the learning objectives of the course. The model-generated rationales were not evaluated in this study and should not be assumed to provide reliable instructional feedback.
Greater caution is required for overall writing quality, content development, cohesion and coherence, argumentation, and communicative effect. These judgements depend on the relevance and development of ideas, relations among claims, task fulfilment, discourse organisation, and the needs of the intended reader. They cannot be secured through score calibration alone. Teachers should therefore retain primary responsibility for these dimensions and for the final interpretation of student performance.
The appropriate division of labour is one in which the LLM provides preliminary, low-stakes information while teachers remain responsible for consequential and complex judgements. Machine-assisted profiles may help teachers direct attention towards disputed essays, weak arguments, discourse-level problems, and individual learning needs. They may also support more frequent formative assessment between assignments receiving full teacher marking. This potential benefit, however, requires classroom evaluation rather than being inferred from agreement statistics.
Teacher control should begin with the design of the assessment. Teachers should determine the rubric, approve the benchmark essays, and decide which dimensions are suitable for machine assistance. Benchmark panels should reflect current course objectives and the writing profile of the local student population. Changes in the model, prompt, rubric, writing task, or learner group should lead to renewed local validation.
Operational use also requires continuing oversight. Schools should document the model version, prompts, benchmark panels, calibration procedures, and conditions under which scores are reported. Teachers should periodically check a sample of machine-assisted scores, and dimensions with weak local validation should receive mandatory human review. Repeated-score stability should not be treated as sufficient evidence of accuracy.
Students should have access to a clear procedure for requesting teacher reconsideration of a machine-assisted score. Institutional policies should also address the storage and transmission of student writing, access permissions, retention periods, and possible reuse of texts or ratings. These arrangements ensure that the model remains an auxiliary assessment resource rather than an independent scoring authority.

5.3. Limitations and Future Research

The findings are limited to 414 essays from one institution, five writing topics, two teacher raters, and the models and prompts examined in this study. The teacher reference was based on the mean of two independent ratings without third-rater adjudication, and the benchmark panels did not cover every score level in every dimension. Teacher workload, cost, latency, student responses, and the instructional quality of model-generated feedback were not measured.
Future research should examine the procedure across institutions, writing genres, topics, learner groups, and larger panels of trained raters. Classroom studies are also needed to determine whether dimension-level profiles support revision, increase the frequency of useful feedback, and change how teachers allocate marking time. Until such evidence is available, LLM-assisted scoring is best positioned as a teacher-governed, locally calibrated, low-stakes formative resource rather than as a replacement for professional judgement.

6. Conclusions

This study examined whether LLM-assisted scoring could support more frequent formative assessment of College English writing while remaining accountable to the standards applied by local teachers. Using 414 classroom essays independently rated by two trained College English teachers, the study found that high agreement between LLMs did not indicate corresponding alignment with the teacher reference scores. The raw LLM ratings were generally more severe, and the degree of teacher-referenced agreement varied substantially across writing dimensions.
Use of the local teacher rubric and teacher-agreed benchmark essays was associated with closer raw alignment than the preliminary machine-anchored conditions. Statistical calibration further reduced differences in score severity and scale use. Under the teacher-anchored conditions, pooled QWK reached 0.614–0.637 after location–scale correction, 0.668–0.697 after isotonic calibration, and 0.629–0.644 after equipercentile linking, compared with a teacher–teacher agreement benchmark of 0.747 and raw values of 0.286–0.401. Because a two-parameter correction recovered almost the whole of the gain, these improvements are attributed to the correction of score severity and scale use rather than to any particular calibration method. The pooled advantage of isotonic calibration was obtained by compressing the score distribution and in one dimension by returning a constant score, which makes it unsuitable for classroom score reporting despite its aggregate accuracy. Calibration did not establish that the models understood writing quality in the same way as the teachers, nor did it remove the marked differences among writing dimensions.
Teacher-referenced agreement was relatively stronger for vocabulary, syntactic ability, and grammatical accuracy, suggesting that these dimensions may provide preliminary information for targeted formative work when interpreted by teachers. Agreement remained weak for cohesion and coherence and was limited for aspects of overall writing quality involving content development, argumentation, discourse meaning, and communicative effect. These judgements should remain under teacher control. Repeated-score instability did not reliably identify higher-error ratings for any of the three models and should not be used as a criterion for releasing scores without human review.
The findings therefore support a limited role for LLMs in College English writing assessment. When constrained by a locally developed rubric, teacher-approved benchmark essays, statistical calibration, repeated scoring, and continuing teacher review, LLM-assisted scoring may help provide dimension-level information more frequently and allow teachers to direct greater attention to disputed essays, complex writing problems, and individual guidance. It should be used for low-stakes formative purposes rather than as an independent scoring authority, and teaching programmes must retain responsibility for defining the assessment standard, determining which scores may be reported, reconsidering disputed results, and interpreting students’ writing development. Because all evidence reported here comes from one institution, five topics, and two raters, these conclusions are conditional on the setting studied and should be re-established locally before the procedure is adopted elsewhere.

Author Contributions

Conceptualization, Y.W. (Yongping Wang); methodology, Y.W. (Yongping Wang), X.C. (Xizhi Chu) and X.C. (Xuan Cheng); software, X.C. (Xuan Cheng); validation, X.C. (Xizhi Chu), T.W. and Y.W. (Yapeng Wang); formal analysis, Y.W. (Yongping Wang), X.C. (Xizhi Chu) and X.C. (Xuan Cheng); investigation, T.W.; resources, T.W. and X.C. (Xuan Cheng); data curation, Y.W. (Yongping Wang), N.L., T.W. and X.C. (Xuan Cheng); writing—original draft preparation, Y.W. (Yongping Wang); writing—review and editing, X.C. (Xizhi Chu) and Y.W. (Yapeng Wang); visualization, Y.W. (Yongping Wang); supervision, Y.W. (Yapeng Wang); project administration, N.L. and Y.W. (Yapeng Wang); funding acquisition, X.C. (Xizhi Chu). All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the 2025 Annual Project of the Shaanxi Provincial “14th Five-Year Plan” Education Science Planning, “Research on an AI-Empowered College English Writing Instruction Model Based on the Production-Oriented Approach”, grant number SGH25Y3282.

Institutional Review Board Statement

Ethical review and approval were waived for this study because the research used anonymized student writing collected as part of regular course activities, and no personally identifiable information was retained.

Informed Consent Statement

Informed consent was waived because the study involved the analysis of anonymized student writing collected as part of regular course activities, with no personally identifiable information retained or disclosed. Students were informed that their writing would be included in the corpus and used for research purposes, and they agreed to its use.

Data Availability Statement

The data and reproducibility materials supporting the findings of this study are publicly available in the GitHub repository at https://github.com/EV4-CX/China_English-Writing-Dataset (accessed on 18 August 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AESAutomated essay scoring
AWEAutomated writing evaluation
CETCollege English Test
CSEChina’s Standards of English Language Ability
EFLEnglish as a foreign language
ELLIPSEEnglish Language Learner Insight, Proficiency and Skills Evaluation
L2Second language
LLMLarge language model
MAEMean absolute error
MTLDMeasure of textual lexical diversity
QWKQuadratic weighted kappa

Appendix A. Definitions

Appendix A provides formal definitions of the assessment terms, evaluation indices, calibration procedures, and review rule used in the study.
Table A1. Glossary of the assessment terms used in this article.
Table A1. Glossary of the assessment terms used in this article.
TermDefinition as Used in This Article
Formative assessmentAssessment whose results are used to adapt teaching and to give feedback that guides subsequent work, as distinct from summative assessment, which certifies attainment. The distinction lies in the use made of the evidence, not in the instrument.
Rating rubricA documented scoring instrument specifying the dimensions on which a performance is judged, the ordered score levels available on each dimension, and the descriptors stating what performance at each level looks like.
Analytic rubricA rubric yielding one score per dimension, producing a profile, as opposed to a holistic rubric, which yields a single score for the whole text.
DimensionOne scored aspect of writing within an analytic rubric, for example vocabulary or cohesion and coherence.
Essay–dimension pairThe unit of analysis: one essay scored on one dimension. With 414 essays and seven dimensions there are 2898 such pairs.
Machine raterAn automated system considered in the role a human rater would otherwise occupy: it receives a script and a rubric and returns an ordinal score on that rubric.
Benchmark scriptA previously rated essay whose agreed score illustrates one level of the scale, used during rater training or, here, presented in the scoring prompt. Also called an exemplar or sample script.
Benchmark panelThe set of benchmark scripts covering the score levels of one dimension, held fixed for all essays scored in one fold under one condition.
NormingThe training procedure in which raters discuss and score a common set of scripts before operational marking, in order to establish a shared interpretation of the descriptors.
Double ratingIndependent rating of the same script by two raters, which makes disagreement visible.
AdjudicationThe procedure by which substantial disagreements between raters are reviewed and resolved, normally by a third rater. No adjudication was carried out in this study.
Teacher reference scoreThe arithmetic mean of the two teachers’ independent ratings for an essay–dimension pair, used here as the local comparison standard.
Rater severityA rater’s general tendency to award lower scores than other raters applying the same rubric to the same scripts; leniency is the opposite tendency.
Calibration (of scores)A transformation applied to scores after rating, estimated from reference data, that adjusts their location, spread, or distribution so that they can be read on the reference scale. It does not alter how the rater judged the text.
Equipercentile linkingA calibration in which a score is replaced by the score occupying the same percentile position in the reference distribution.
OccasionOne independent rating of one essay–dimension pair by one model; each pair was rated on five occasions.
Modal proportionThe proportion of the five occasions that produced the most frequently assigned score; it can equal 0.2, 0.4, 0.6, 0.8, or 1.0.
Occasion rangeThe difference between the highest and lowest of the five occasion scores.
Held-out partitionThe fifth of the corpus not used to construct the benchmark panels, calibration mappings, or review threshold applied to it.
Definitions follow standard usage in language-assessment and educational-measurement sources [13,16,21,23,24].
Table A2. Definitions of the five evaluation indices.
Table A2. Definitions of the five evaluation indices.
IndexWhat Is ComputedRangeReading
Quadratic weighted kappa (QWK)Both score sets are assigned to the nine half-point categories from 1.0 to 5.0. A square table of observed joint frequencies is formed, together with a second table of the frequencies expected from the two sets of marginal proportions. Each cell carries a weight equal to the squared distance between its row and column categories, divided by the squared distance between the extreme categories. QWK is one minus the ratio of the weight-summed observed table to the weight-summed expected table.≤1; 1 = identical, 0 = agreement no better than the marginals predict, negative = worsePrincipal ordinal agreement index. Depends on the marginal distributions and on score variation, so it is read together with the error indices.
Spearman rank correlationThe Pearson correlation of the mid-ranks of the machine scores and the mid-ranks of the reference scores over the observations in the stated set.−1 to 1Similarity of ordering only; unaffected by the level at which scores are placed.
Mean absolute error (MAE)The mean over pairs of the absolute difference between the machine score and the reference score.0 upwards, in scale pointsTypical size of the discrepancy for an individual score.
Signed biasThe mean over pairs of the machine score minus the reference score.Negative to positive, in scale pointsNegative indicates systematic severity, positive systematic leniency.
Half-point tolerance accuracy (Acc ± 0.5)The proportion of pairs for which the absolute difference between machine and reference score is at most 0.5.0 to 1The proportion of scores close enough to the reference to be usable for feedback, at the resolution the individual teachers used.
In every case the machine score is the calibrated or raw aggregated score for one essay–dimension pair, and the reference is the teacher reference score for the same pair unless another reference is named. Pooled values are computed over all 2898 out-of-fold essay–dimension pairs; dimension-level values over the 414 pairs of one dimension.
Table A3. Definitions of the three calibration methods and of the review rule.
Table A3. Definitions of the three calibration methods and of the review rule.
ProcedureEstimated fromDefinitionNotes
Location–scale correctionThe mean and standard deviation of the machine scores and of the teacher reference scores in the training partition.Each held-out machine score is shifted so that the machine mean coincides with the reference mean, and scaled so that the machine standard deviation coincides with the reference standard deviation. Where the machine standard deviation is numerically zero, the scaling factor is set to one and only the shift is applied.Two free parameters. Corrects average severity and spread only.
Isotonic calibrationPaired machine and reference scores in the training partition.A nondecreasing step function of the raw machine score is fitted to the reference scores by minimising the sum of squared deviations, using the pool-adjacent-violators algorithm. Raw values outside the range observed in training are mapped to the nearest fitted endpoint.Nonparametric and order-preserving. Because it minimises squared error, its fitted values shrink towards the conditional mean, which compresses the reported score range.
Equipercentile linkingThe marginal distributions of the machine scores and of the reference scores in the training partition; no pairing is used.For a given raw machine score, the proportion of training machine scores at or below that value is computed. The calibrated value is the point of the reference distribution below which the same proportion of training reference scores falls, obtained by linear interpolation where the required percentile position falls between two observed points.Order-preserving and tabulable: the whole mapping for one dimension and fold can be printed as a short lookup table.
Instability-based review ruleThe distribution of occasion range and modal proportion in the training partition, together with the pair-level workload reference of 15.8%.A held-out essay–dimension score is flagged for teacher review if the occasion range is at least two scale points, or if the modal proportion is below the threshold τ , chosen in the training partition as the largest value for which the proportion of flagged pairs does not exceed the workload reference. An essay is counted as requiring review if at least one of its seven dimension scores is flagged.The modal proportion takes only five values, so the attainable flagging rates are coarse and the selected τ differs across models.
All three calibrations were estimated separately within every combination of model, dimension, prompt condition, and fold, using only the training partition of that fold, and were then applied unchanged to the held-out partition. Calibrated values were restricted to the interval from 1.0 to 5.0 and rounded to the nearest half point before any statistic was computed. The corresponding formal expressions are given in Appendix C.

Appendix B. Reproducibility Information

Appendix B reports the model identifiers, API endpoints, decoding settings, access dates, request and retry procedures, cross-validation structure, and benchmark panels used in the study. The rating data, fold assignments, benchmark identifiers, calibration mappings, prompt templates, and analysis scripts referred to below are available as described in the Data Availability Statement.
Table A4. Models, provider-side identifiers, endpoints, decoding settings, and access dates.
Table A4. Models, provider-side identifiers, endpoints, decoding settings, and access dates.
RoleModelProvider-Side IdentifierEndpointTemp.OccasionsAccess Dates (2026)
Deployable raterDeepSeek-V3deepseek-chathttps://api.deepseek.com (accessed on 2 July 2026)0.352 July; 4 July
Deployable raterGLM-4-Airglm-4-air-250414https://open.bigmodel.cn/api/paas/v4 (accessed on 2 July 2026)0.352 July; 5 July
Deployable raterQwen3-Maxqwen3-maxhttps://dashscope.aliyuncs.com/compatible-mode/v1 (accessed on 4 July 2026)0.354 July
Benchmark panelDeepSeek-R1deepseek-reasonerhttps://api.deepseek.com (accessed on 2 July 2026)0.212 July
Benchmark panelGLM-4-Plusglm-4-plushttps://open.bigmodel.cn/api/paas/v4 (accessed on 2 July 2026)0.212 July
All models were accessed through the providers’ public OpenAI-compatible chat-completions interfaces from mainland China. API credentials were supplied through a local configuration file. Default provider values were used for all decoding parameters other than temperature; no top-p, penalty, seed, or maximum-token setting was specified, and no system message was sent. The addresses in the Endpoint column are API base URLs that were called programmatically with authenticated POST requests; they are not browsable web pages and return no content when opened in a browser. The date in parentheses after each endpoint is the first date on which it was called, while the final column lists every date on which scoring runs were executed. The corresponding publicly accessible documentation pages are https://api-docs.deepseek.com for DeepSeek, https://docs.bigmodel.cn for Zhipu AI (GLM), and https://help.aliyun.com/zh/model-studio/ for Alibaba Cloud Model Studio (Qwen), all accessed on 18 August 2026. Provider-side model catalogues change over time, so the identifiers and endpoints are reported as they were on the access dates given.
Table A5. Request protocol, independence of rating occasions, and handling of failures.
Table A5. Request protocol, independence of rating occasions, and handling of failures.
ItemSpecification
Unit of a requestOne dimension of one essay. Seven requests per essay in the teacher-anchored stage, six in the machine-anchored stage.
Message structureA single user message containing the completed prompt template. No system message, no assistant turn, no conversation history.
Independence of the five occasionsEach occasion is a separate stateless request with identical content. No score, rationale, or identifier from a previous occasion is included. No provider-side session, thread, conversation identifier, or prompt-caching option is enabled. The five responses differ only through sampling at temperature 0.3.
Prompt contentRole instruction; the descriptor set for the dimension being rated; the permitted score levels; in benchmark-referenced conditions, the benchmark essays with their agreed scores in ascending order; the target essay; and a request for a structured response containing the score and a one-sentence rationale.
Request timeout180 s.
Response parsingThe score is read from the score field of the returned JSON object. A value already on the permitted grid is accepted unchanged; a value inside the interval 1.0–5.0 but off the grid is assigned to the nearest permitted level; a value outside 1.0–5.0, a missing field, or unparsable output is treated as a failed attempt.
Retry policyUp to three attempts per occasion, with waits of 3 s, 6 s, and 9 s. The prompt is not modified between attempts.
Handling of incomplete setsAn essay–dimension pair was written to the score file only if all five occasions returned a valid score. No pair was excluded on this ground: all 414 essays × 7 dimensions × 3 models × 2 conditions produced complete five-occasion sets.
RationalesA one-sentence rationale was requested and stored with every score but was not used in aggregation, calibration, flagging, or any reported statistic.
Table A6. Cross-validation structure and composition of the teacher-agreed benchmark panels.
Table A6. Cross-validation structure and composition of the teacher-agreed benchmark panels.
DimensionEssays with Exact Teacher Agreement (of 414)Half-Point Levels Attested in the Full CorpusPanel Size per Fold
Overall writing quality1542.5, 3.0, 3.5, 4.0, 4.54–5
Cohesion and coherence1521.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.56–7
Syntactic ability1631.5, 2.0, 2.5, 3.0, 3.5, 4.06
Vocabulary1722.5, 3.0, 3.5, 4.0, 4.5, 5.06
Phraseology1701.0, 1.5, 2.0, 2.5, 3.0, 3.56
Grammatical accuracy1431.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.07
Conventions1811.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.06–7
Folds were produced by stratified five-fold splitting at the essay level, stratified on the rounded teacher reference score for overall writing quality, with a fixed random seed, giving held-out partitions of 83, 83, 83, 83, and 82 essays. Exact-agreement essays are those on which the two teachers assigned identical scores on that dimension. A panel contains at most one essay per attested half-point level, selected as the eligible essay closest to the corpus median length within the training partition of the fold; panel sizes therefore vary by fold where a level is attested in one training partition but not another.

Appendix C. Formal Specification of the Scoring, Calibration, and Review Procedures

In the expressions below, essays are indexed by e, rubric dimensions by d, the five independent rating occasions by r, and the five cross-validation folds by k. The set of permitted score levels, denoted S , is the half-point grid from 1.0 to 5.0 in the teacher-anchored stage and the integer grid from 1 to 5 in the machine-anchored stage. The score returned on one occasion is the occasion score, denoted s and carrying the essay, dimension, and occasion subscripts; the teacher reference score, denoted h, is the mean of the two teachers’ ratings for the same essay–dimension pair. Every estimate is obtained from the training partition of the fold and dimension concerned, denoted T with the dimension and fold subscripts, whose elements are the essay–dimension pairs of that partition and are indexed by i. The reporting operator, written as a capital pi, Π, clips a value to the interval from 1 to 5 and rounds it to the nearest element of the half-point grid.
The aggregated raw machine score, distinguished by a caret, is the median of the five occasion scores:
s ^ e d = median s e d 1 , , s e d 5 , s e d r S = { 1.0 , 1.5 , , 5.0 } ,
and the two consistency indicators computed from those same five scores are the occasion range R, the difference between the largest and the smallest of them, and the modal proportion m, the number of occasions on which the most frequent score level v was returned, divided by five:
R e d = max r s e d r min r s e d r , m e d = 1 5 max v S { r : s e d r = v } .
The vertical bars denote the number of elements of the set they enclose, so the modal proportion takes values in the set 0.2, 0.4, 0.6, 0.8, 1.0.
Each calibration method defines a mapping, denoted c with a superscript naming the method and with the dimension and fold subscripts, from a raw machine score s to a reported score; the mapping is estimated on the training partition of a fold and dimension and applied unchanged to the corresponding held-out partition. Location–scale correction matches the first two moments of the two training distributions:
c d , k LS ( s ) = Π μ d , k H + σ d , k H σ d , k M s μ d , k M .
Here μ and σ are the mean and the standard deviation over the training partition, the superscript M denotes the machine scores and the superscript H the teacher reference scores, and the scaling factor is set to one when the machine standard deviation is numerically zero. Isotonic calibration fits a nondecreasing function, denoted g with a caret, to the paired training observations by least squares, each pair consisting of a raw machine score and the teacher reference score for the same essay–dimension pair i:
c d , k ISO ( s ) = Π g ^ d , k ( s ) , g ^ d , k = arg min g nondecreasing i T d , k h i g ( s i ) 2 ,
where the minimisation is over nondecreasing functions and is solved by the pool-adjacent-violators algorithm; raw values outside the training range are mapped to the nearest fitted endpoint. Equipercentile linking composes the empirical distribution function of the machine scores with the quantile function of the teacher reference scores:
c d , k EQ ( s ) = Π F d , k H 1 F d , k M ( s ) .
Here F with the superscript M is the empirical distribution function of the machine scores in the training partition and F with the superscript H that of the teacher reference scores, so the inverse of the latter is the corresponding quantile function; that inverse is obtained by linear interpolation between adjacent observed score points. All three mappings are nondecreasing, so none reverses the ordering of two essays within a given dimension and fold.
The instability-based review rule flags an essay–dimension score, and refers an essay for teacher review if any of its dimension scores is flagged:
flag e d = 1 R e d 2 or m e d < τ , review e = 1 d flag e d 1 .
Each right-hand side equals one when the condition in brackets holds and zero otherwise, so a score is flagged when its occasion range is at least two scale points or its modal proportion falls below the threshold τ , and an essay is referred for review when at least one of its dimension scores is flagged. The threshold τ is selected on the training partition as the largest value in the attainable grid for which the expected proportion of flagged pairs does not exceed the pair-level workload reference of 15.8% obtained from teacher disagreement, and is then applied unchanged to the held-out partition.

References

  1. Xu, Y.; Harfitt, G.J. Is Assessment for Learning Feasible in Large Classes? Challenges and Coping Strategies from Three Case Studies. Asia-Pac. J. Teach. Educ. 2019, 47, 472–486. [Google Scholar] [CrossRef] [Scilit]
  2. Wu, W.; Huang, J.; Han, C.; Zhang, J. Evaluating Peer Feedback as a Reliable and Valid Complementary Aid to Teacher Feedback in EFL Writing Classrooms: A Feedback Giver Perspective. Stud. Educ. Eval. 2022, 73, 101140. [Google Scholar] [CrossRef] [Scilit]
  3. National Advisory Committee on Foreign Language Teaching in Higher Education. Guidelines on College English Teaching (2020 Edition); Higher Education Press: Beijing, China, 2020. (In Chinese) [Google Scholar]
  4. Black, P.; Wiliam, D. Assessment and Classroom Learning. Assess. Educ. Princ. Policy Pract. 1998, 5, 7–74. [Google Scholar] [CrossRef] [Scilit]
  5. Bai, L.; Hu, G. In the Face of Fallible AWE Feedback: How Do Students Respond? Educ. Psychol. 2017, 37, 67–81. [Google Scholar] [CrossRef] [Scilit]
  6. Zhang, Z.V.; Hyland, K. Student Engagement with Teacher and Automated Feedback on L2 Writing. Assess. Writ. 2018, 36, 90–102. [Google Scholar] [CrossRef] [Scilit]
  7. Mizumoto, A.; Eguchi, M. Exploring the Potential of Using an AI Language Model for Automated Essay Scoring. Res. Methods Appl. Linguist. 2023, 2, 100050. [Google Scholar] [CrossRef] [Scilit]
  8. Yancey, K.P.; LaFlair, G.T.; Verardi, A.R.; Burstein, J. Rating Short L2 Essays on the CEFR Scale with GPT-4. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), Toronto, ON, Canada, 13 July 2023; pp. 576–584. [Google Scholar] [CrossRef] [Scilit]
  9. Kim, Y. Automated Essay Scoring with GPT-4 for a Local Placement Test: Investigating Prompting Strategies, Intra-Rater Reliability, and Alignment with Human Scores. TESOL Q. 2025, 59, S318–S329. [Google Scholar] [CrossRef] [Scilit]
  10. Yavuz, F.; Çelik, Ö.; Yavaş Çelik, G. Utilizing Large Language Models for EFL Essay Grading: An Examination of Reliability and Validity in Rubric-Based Assessments. Br. J. Educ. Technol. 2025, 56, 150–166. [Google Scholar] [CrossRef] [Scilit]
  11. Gomes, A.; Mendes, A.J.; Marcelino, M.J.; Ke, W.; Im, S.K.; Siu, A. A Teacher’s View about Introductory Programming Teaching and Learning: Portuguese and Macanese Perspectives. In Proceedings of the 2017 IEEE Frontiers in Education Conference (FIE), Indianapolis, IN, USA, 18–21 October 2017; pp. 1–8. [Google Scholar] [CrossRef] [Scilit]
  12. Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Adv. Neural Inf. Process. Syst. 2023, 36, 46595–46623. [Google Scholar] [CrossRef] [Scilit]
  13. Weigle, S.C. Assessing Writing; Cambridge University Press: Cambridge, UK, 2002. [Google Scholar] [CrossRef] [Scilit]
  14. Eckes, T. Introduction to Many-Facet Rasch Measurement: Analyzing and Evaluating Rater-Mediated Assessments, 2nd ed.; Peter Lang: Frankfurt am Main, Germany, 2015. [Google Scholar] [CrossRef] [Scilit]
  15. Linacre, J.M. Many-Facet Rasch Measurement; MESA Press: Chicago, IL, USA, 1989. [Google Scholar]
  16. Kolen, M.J.; Brennan, R.L. Test Equating, Scaling, and Linking: Methods and Practices, 3rd ed.; Springer: New York, NY, USA, 2014. [Google Scholar] [CrossRef] [Scilit]
  17. Livingston, S.A. Equating Test Scores (Without IRT), 2nd ed.; Educational Testing Service: Princeton, NJ, USA, 2014. [Google Scholar]
  18. Ministry of Education of the People’s Republic of China. China’s Standards of English Language Ability (GF 0018–2018); Higher Education Press: Beijing, China, 2018. (In Chinese) [Google Scholar]
  19. National College English Testing Committee. Syllabus for the College English Test Band 4 and Band 6 (2016 Revision); Shanghai Jiao Tong University Press: Shanghai, China, 2016. (In Chinese) [Google Scholar]
  20. Hyland, K.; Hyland, F. Feedback on Second Language Students’ Writing. Lang. Teach. 2006, 39, 83–101. [Google Scholar] [CrossRef] [Scilit]
  21. Knoch, U. Rating Scales for Diagnostic Assessment of Writing: What Should They Look Like and Where Should the Criteria Come From? Assess. Writ. 2011, 16, 81–96. [Google Scholar] [CrossRef] [Scilit]
  22. Faseeh, M.; Jaleel, A.; Iqbal, N.; Ghani, A.; Abdusalomov, A.; Mehmood, A.; Cho, Y.I. Hybrid Approach to Automated Essay Scoring: Integrating Deep Learning Embeddings with Handcrafted Linguistic Features for Improved Accuracy. Mathematics 2024, 12, 3416. [Google Scholar] [CrossRef] [Scilit]
  23. Williamson, D.M.; Xi, X.; Breyer, F.J. A Framework for Evaluation and Use of Automated Scoring. Educ. Meas. Issues Pract. 2012, 31, 2–13. [Google Scholar] [CrossRef] [Scilit]
  24. Ramineni, C.; Williamson, D.M. Automated Essay Scoring: Psychometric Guidelines and Practices. Assess. Writ. 2013, 18, 25–39. [Google Scholar] [CrossRef] [Scilit]
  25. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; Volume 70, pp. 1321–1330. [Google Scholar]
  26. Sheng, H.; Liu, X.; He, H.; Zhao, J.; Kang, J. Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, 4–9 November 2025; pp. 11286–11328. [Google Scholar] [CrossRef] [Scilit]
  27. Crossley, S.; Tian, Y.; Baffour, P.; Franklin, A.; Kim, Y.; Morris, W.; Benner, M.; Picou, A.; Boser, U. The English Language Learner Insight, Proficiency and Skills Evaluation (ELLIPSE) Corpus. Int. J. Learn. Corpus Res. 2023, 9, 248–269. [Google Scholar] [CrossRef] [Scilit]
  28. Crossley, S.A.; Tian, Y.; Baffour, P.; Franklin, A.; Benner, M.; Boser, U. A Large-Scale Corpus for Assessing Written Argumentation: PERSUADE 2.0. Assess. Writ. 2024, 61, 100865. [Google Scholar] [CrossRef] [Scilit]
  29. Crossley, S.A.; Baffour, P.; Burleigh, L.; King, J. A Large-Scale Corpus for Assessing Source-Based Writing Quality: ASAP 2.0. Assess. Writ. 2025, 65, 100954. [Google Scholar] [CrossRef] [Scilit]
  30. McCarthy, P.M.; Jarvis, S. MTLD, vocd-D, and HD-D: A Validation Study of Sophisticated Approaches to Lexical Diversity Assessment. Behav. Res. Methods 2010, 42, 381–392. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Brennan, R.L.; Prediger, D.J. Coefficient Kappa: Some Uses, Misuses, and Alternatives. Educ. Psychol. Meas. 1981, 41, 687–699. [Google Scholar] [CrossRef] [Scilit]
  32. Feinstein, A.R.; Cicchetti, D.V. High Agreement but Low Kappa: I. The Problems of Two Paradoxes. J. Clin. Epidemiol. 1990, 43, 543–549. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Zebaze, A.R.; Sagot, B.; Bawden, R. In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, NM, USA, 29 April–4 May 2025; pp. 1222–1252. [Google Scholar]
  34. Im, S.K.; Pearmain, A. Error Resilient Video Coding with Priority Data Classification Using H.264 Flexible Macroblock Ordering. IET Image Process. 2007, 1, 197–204. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Cross-validated procedure for teacher-referenced LLM-assisted writing assessment. Within each fold, the training partition supplies the benchmark panels and the reference score distribution, and the held-out partition is scored without access to its teacher ratings. The system produces five independent benchmark-referenced ratings and an audit record, aggregates them by the median, calibrates the aggregated score to the programme reference scale by equipercentile linking, and routes ratings flagged as unstable to teacher review. The diagram shows the teacher-anchored stage, in which the analytic rubric contains seven CSE-informed dimensions; the preliminary machine-anchored stage uses a six-dimension rubric.
Figure 1. Cross-validated procedure for teacher-referenced LLM-assisted writing assessment. Within each fold, the training partition supplies the benchmark panels and the reference score distribution, and the held-out partition is scored without access to its teacher ratings. The system produces five independent benchmark-referenced ratings and an audit record, aggregates them by the median, calibrates the aggregated score to the programme reference scale by equipercentile linking, and routes ratings flagged as unstable to teacher review. The diagram shows the teacher-anchored stage, in which the analytic rubric contains seven CSE-informed dimensions; the preliminary machine-anchored stage uses a six-dimension rubric.
Mathematics 14 03033 g001
Figure 2. Inter-model agreement and teacher-referenced alignment in the preliminary machine-anchored comparison (n = 414 essays; five mapped dimensions). (a) Pooled score distributions for the combined machine-panel ratings and the teacher reference scores; the machine panel scores about one band lower. (b) Teacher–teacher and machine-panel–teacher QWK by mapped writing dimension. Inter-model QWK among the two panel models was 0.695–0.778.
Figure 2. Inter-model agreement and teacher-referenced alignment in the preliminary machine-anchored comparison (n = 414 essays; five mapped dimensions). (a) Pooled score distributions for the combined machine-panel ratings and the teacher reference scores; the machine panel scores about one band lower. (b) Teacher–teacher and machine-panel–teacher QWK by mapped writing dimension. Inter-model QWK among the two panel models was 0.695–0.778.
Mathematics 14 03033 g002
Figure 3. Effects of statistical calibration on teacher-referenced score distributions and pooled agreement (teacher-anchored stage, teacher-benchmark condition, five-fold cross-validation). (a) Raw and equipercentile-calibrated machine-score distributions for GLM-4-Air in relation to the teacher reference distribution. (b) Pooled QWK against the teacher reference for raw scores and for the three calibration methods. The teacher–teacher QWK of 0.747 is included as a corpus-specific agreement reference.
Figure 3. Effects of statistical calibration on teacher-referenced score distributions and pooled agreement (teacher-anchored stage, teacher-benchmark condition, five-fold cross-validation). (a) Raw and equipercentile-calibrated machine-score distributions for GLM-4-Air in relation to the teacher reference distribution. (b) Pooled QWK against the teacher reference for raw scores and for the three calibration methods. The teacher–teacher QWK of 0.747 is included as a corpus-specific agreement reference.
Mathematics 14 03033 g003
Figure 4. Dispersion of calibrated scores by dimension for DeepSeek-V3 under the teacher-benchmark condition, compared with the dispersion of the teacher reference scores. Isotonic calibration returns a single constant value for conventions, for which the standard deviation is therefore zero. The same pattern was obtained for GLM-4-Air and Qwen3-Max.
Figure 4. Dispersion of calibrated scores by dimension for DeepSeek-V3 under the teacher-benchmark condition, compared with the dispersion of the teacher reference scores. Isotonic calibration returns a single constant value for conventions, for which the standard deviation is therefore zero. The same pattern was obtained for GLM-4-Air and Qwen3-Max.
Mathematics 14 03033 g004
Figure 5. Rank correlation with the teacher reference score before and after calibration, at three levels of aggregation (teacher-benchmark condition; five-fold cross-validation). Each panel is one model. Small dots are the five fold-level correlations, filled squares the correlation pooled across the seven dimensions, and open diamonds the mean of the seven within-dimension correlations. Calibration raises the pooled and fold-level correlations but leaves the within-dimension correlation almost unchanged.
Figure 5. Rank correlation with the teacher reference score before and after calibration, at three levels of aggregation (teacher-benchmark condition; five-fold cross-validation). Each panel is one model. Small dots are the five fold-level correlations, filled squares the correlation pooled across the seven dimensions, and open diamonds the mean of the seven within-dimension correlations. Calibration raises the pooled and fold-level correlations but leaves the within-dimension correlation almost unchanged.
Mathematics 14 03033 g005
Figure 6. Dimension-level teacher-referenced performance under the teacher-benchmark condition after equipercentile linking (five-fold cross-validation; all seven teacher-rubric dimensions). (a) QWK by teacher-rubric dimension, with teacher–teacher agreement included as a corpus-specific agreement reference. (b) Calibrated MAE by dimension for DeepSeek-V3, GLM-4-Air, and Qwen3-Max.
Figure 6. Dimension-level teacher-referenced performance under the teacher-benchmark condition after equipercentile linking (five-fold cross-validation; all seven teacher-rubric dimensions). (a) QWK by teacher-rubric dimension, with teacher–teacher agreement included as a corpus-specific agreement reference. (b) Calibrated MAE by dimension for DeepSeek-V3, GLM-4-Air, and Qwen3-Max.
Mathematics 14 03033 g006
Figure 7. Difference in mean absolute error between flagged and unflagged scores for each model (teacher-benchmark condition after equipercentile linking), with 95% confidence intervals from 1000 essay-level paired bootstrap resamples. A positive value means flagged scores are less accurate. Every interval includes zero, so occasion-level instability does not reliably identify the less accurate scores.
Figure 7. Difference in mean absolute error between flagged and unflagged scores for each model (teacher-benchmark condition after equipercentile linking), with 95% confidence intervals from 1000 essay-level paired bootstrap resamples. A positive value means flagged scores are less accurate. Every interval includes zero, so occasion-level instability does not reliably identify the less accurate scores.
Mathematics 14 03033 g007
Table 1. Sources, adaptations, and descriptor-based correspondences of the preliminary six-dimension machine rubric.
Table 1. Sources, adaptations, and descriptor-based correspondences of the preliminary six-dimension machine rubric.
Adapted Machine DimensionELLIPSE Source Dimension(s)Local Teacher-Rubric CounterpartNature of Correspondence
Overall qualityOverall proficiencyOverall writing qualityBroad holistic correspondence
Organisation and coherenceCohesionCohesion and coherenceRelated but not identical; renamed to match College English terminology
Vocabulary and expressionVocabulary; PhraseologyVocabulary; PhraseologyTwo source traits combined into one machine dimension
Grammar and syntaxGrammar; SyntaxGrammatical accuracy; Syntactic abilityTwo source traits combined into one machine dimension
MechanicsConventionsConventionsClose descriptor correspondence; renamed
Content and task fulfilmentNot derived from ELLIPSE; added with reference to the CET writing specifications [19]No direct counterpartAdded for task-response evaluation; excluded from all mapped comparisons
The source framework is the ELLIPSE analytic rating scheme [27]. The correspondences indicate overlap in descriptor focus rather than construct equivalence. The machine rubric was used only in the preliminary machine-anchored stage.
Table 2. Inter-rater agreement of the two College English teachers by rubric dimension (n = 414 essays; half-point scale).
Table 2. Inter-rater agreement of the two College English teachers by rubric dimension (n = 414 essays; half-point scale).
DimensionQWKExact AgreementWithin 0.5Diff. > 0.5
Overall writing quality0.3760.3720.8410.159
Cohesion and coherence0.5800.3670.8410.159
Syntactic ability0.6070.3940.8410.159
Vocabulary0.6700.4150.8530.147
Phraseology0.5870.4110.8430.157
Grammatical accuracy0.6630.3450.8190.181
Conventions0.4400.4370.8600.140
Mean of seven dimensions0.5610.3920.8420.158
Pooled, five mapped dimensions0.7240.3870.8430.157
Pooled, all seven dimensions0.7470.3920.8420.158
Within 0.5 = proportion of essay–dimension pairs on which the two teachers differed by at most half a point; Diff. > 0.5 = the complementary proportion, which defines the pair-level workload reference used in Section 3.5.2. Pooled QWK is 0.724 over the five mapped dimensions and 0.747 over all seven dimensions; both are included as corpus-specific agreement references.
Table 3. Effects of benchmark referencing on raw teacher-referenced performance under the machine-anchored and teacher-anchored rating designs (five-fold cross-validation).
Table 3. Effects of benchmark referencing on raw teacher-referenced performance under the machine-anchored and teacher-anchored rating designs (five-fold cross-validation).
StageModelConditionQWKSpearmanMAEBiasAcc ± 0.5
Machine-anchoredDeepSeek-V3Rubric-only0.0860.1641.180−0.9890.223
B0.1070.1811.133−0.9350.241
GLM-4-AirRubric-only0.2490.3520.926−0.6560.360
B0.2250.2510.802−0.4320.456
Teacher-anchoredDeepSeek-V3Rubric-only0.1610.1830.808−0.4450.455
Teacher-B0.2860.3220.747−0.3020.481
GLM-4-AirRubric-only0.2490.3090.783−0.3320.449
Teacher-B0.3790.4090.684−0.1150.543
Qwen3-MaxRubric-only0.2300.3010.835−0.4180.424
Teacher-B0.4010.4680.686−0.1940.539
All values are for uncalibrated aggregated machine scores compared with the teacher reference score. B = machine-agreed benchmark panels; Teacher-B = teacher-agreed benchmark panels. Results should be interpreted within each stage because the two stages used different rubrics and score resolutions.
Table 4. Comparison of location–scale correction, isotonic calibration, and equipercentile linking on the same held-out raw scores (teacher-anchored stage, teacher-benchmark condition, five-fold cross-validation, all seven dimensions).
Table 4. Comparison of location–scale correction, isotonic calibration, and equipercentile linking on the same held-out raw scores (teacher-anchored stage, teacher-benchmark condition, five-fold cross-validation, all seven dimensions).
ModelCalibrationQWKSpearmanMAEBiasAcc ± 0.5
DeepSeek-V3None (raw)0.2860.3220.747−0.3020.481
Location–scale0.6140.6320.475+0.0120.699
Isotonic0.6680.7110.385+0.0310.794
Equipercentile0.6430.6570.462−0.0780.722
GLM-4-AirNone (raw)0.3790.4090.684−0.1150.543
Location–scale0.6240.6390.479+0.0280.699
Isotonic0.6850.7220.378+0.0180.803
Equipercentile0.6440.6470.466−0.0720.723
Qwen3-MaxNone (raw)0.4010.4680.686−0.1940.539
Location–scale0.6370.6530.484−0.0270.688
Isotonic0.6970.7080.383−0.0340.794
Equipercentile0.6290.6340.487−0.1500.704
Teacher–teacher pooled QWK on the same dimensions is 0.747, with 84.2% of teacher pairs within half a point; both are included as corpus-specific agreement references. All calibration parameters were estimated from the training partition of each fold, separately for each model and dimension, and applied unchanged to the held-out partition. Under the rubric-only condition the same ordering of methods was obtained: pooled QWK 0.593–0.609 after location–scale correction, 0.661–0.687 after isotonic calibration, and 0.617–0.638 after equipercentile linking.
Table 5. Dimension-level QWK against the teacher reference score by calibration method (teacher-benchmark condition, five-fold cross-validation).
Table 5. Dimension-level QWK against the teacher reference score by calibration method (teacher-benchmark condition, five-fold cross-validation).
DimensionT–TModelRawLocation–ScaleIsotonicEquipercentile
Overall writing quality0.376DeepSeek-V30.2600.3080.1400.331
GLM-4-Air0.3290.3060.1990.325
Qwen3-Max0.3050.3170.1530.277
Cohesion and coherence0.580DeepSeek-V30.0370.0840.0680.130
GLM-4-Air0.0800.1110.0970.097
Qwen3-Max0.0080.0180.0010.057
Syntactic ability0.607DeepSeek-V30.4320.3920.3470.410
GLM-4-Air0.4170.4140.3400.419
Qwen3-Max0.4080.4080.2630.412
Vocabulary0.670DeepSeek-V30.2840.6090.5850.571
GLM-4-Air0.3750.5370.6220.533
Qwen3-Max0.4040.5120.6010.449
Phraseology0.587DeepSeek-V30.2760.3050.2530.326
GLM-4-Air0.2780.3390.2700.358
Qwen3-Max0.2630.3170.2680.316
Grammatical accuracy0.663DeepSeek-V30.4140.4260.2930.373
GLM-4-Air0.4440.4600.3580.433
Qwen3-Max0.4890.5370.4820.487
Conventions0.440DeepSeek-V30.0650.0790.0000.057
GLM-4-Air0.0780.1130.0000.059
Qwen3-Max0.1030.1370.0000.102
T–T = teacher–teacher QWK on the same dimension, included as a corpus-specific agreement reference. Isotonic calibration returned a single constant value for conventions in all three models, so its QWK for that dimension is zero. Cohesion and coherence and conventions remain weak under every method and in the raw scores.
Table 6. Flagging rates and teacher-referenced performance of flagged, unflagged, and automatically released scores (teacher-benchmark condition after equipercentile linking).
Table 6. Flagging rates and teacher-referenced performance of flagged, unflagged, and automatically released scores (teacher-benchmark condition after equipercentile linking).
QuantityDeepSeek-V3GLM-4-AirQwen3-Max
Selected modal-proportion threshold τ 0.800.601.00
Flagged essay–dimension pairs15.1% (438/2898)0.3% (10/2898)10.2% (296/2898)
Essays with at least one flagged dimension66.7% (276/414)2.4% (10/414)53.6% (222/414)
QWK, flagged scores0.5950.5030.637
QWK, unflagged (released) scores0.6500.6440.628
MAE, flagged scores0.4890.4750.490
MAE, unflagged (released) scores0.4570.4660.486
Acc ± 0.5, flagged scores0.6740.7000.699
Acc ± 0.5, unflagged (released) scores0.7300.7230.704
ΔMAE, flagged minus unflagged [95% CI]+0.032 [−0.007, +0.070]+0.012 [−0.148, +0.162]+0.005 [−0.039, +0.047]
ρ (occasion range, absolute error)+0.038−0.037+0.007
Flagged pairs at the common threshold τ = 0.8015.1%16.9%4.2%
QWK flagged/unflagged at τ = 0.800.595/0.6500.661/0.6410.660/0.627
A score was flagged when the range across the five rating occasions reached two scale points or the modal proportion fell below the threshold τ , as defined in Appendix C, Equation (A6). τ was chosen in the training partition as the largest value for which the flagged proportion did not exceed the 15.8% pair-level workload reference. Because the modal proportion of five occasions takes only five values, the attainable flagging rates are coarse and the selected τ differs across models. ΔMAE is the flagged-minus-unflagged difference in mean absolute error with a 95% confidence interval from 1000 essay-level paired bootstrap resamples. The last two rows apply a common threshold (τ = 0.80, that is, flag whenever the modal proportion is 0.6 or below) to all three models.
Table 7. Agreement of calibrated scores with each teacher separately and with the teacher reference score (teacher-benchmark condition after equipercentile linking).
Table 7. Agreement of calibrated scores with each teacher separately and with the teacher reference score (teacher-benchmark condition after equipercentile linking).
ModelReference UsedQWKSpearmanMAEBiasAcc ± 0.5
DeepSeek-V3Teacher 10.6150.6130.488−0.0360.770
Teacher 20.6030.6090.492−0.1190.758
Reference (mean)0.6430.6570.462−0.0780.722
GLM-4-AirTeacher 10.6190.6050.493−0.0300.767
Teacher 20.6030.6000.495−0.1130.753
Reference (mean)0.6440.6470.466−0.0720.723
Qwen3-MaxTeacher 10.5970.5970.507−0.1080.758
Teacher 20.5710.5860.524−0.1910.738
Reference (mean)0.6290.6340.487−0.1500.704
Teacher 1 and Teacher 2 denote the two teachers’ original half-point ratings; Reference denotes their arithmetic mean, which can take quarter-point values. Teacher–teacher pooled QWK is 0.747 and teacher–teacher half-point agreement is 0.842. QWK is higher against the reference because the mean of two ratings is a more reliable quantity than either rating alone; half-point tolerance accuracy is higher against an individual teacher because that reference lies on the same half-point grid as the machine scores.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, Y.; Liu, N.; Chu, X.; Wang, T.; Cheng, X.; Wang, Y. LLM-Assisted Scoring for College English Writing Assessment: Statistical Calibration Against Teacher Standards. Mathematics 2026, 14, 3033. https://doi.org/10.3390/math14173033

AMA Style

Wang Y, Liu N, Chu X, Wang T, Cheng X, Wang Y. LLM-Assisted Scoring for College English Writing Assessment: Statistical Calibration Against Teacher Standards. Mathematics. 2026; 14(17):3033. https://doi.org/10.3390/math14173033

Chicago/Turabian Style

Wang, Yongping, Ning Liu, Xizhi Chu, Tuo Wang, Xuan Cheng, and Yapeng Wang. 2026. "LLM-Assisted Scoring for College English Writing Assessment: Statistical Calibration Against Teacher Standards" Mathematics 14, no. 17: 3033. https://doi.org/10.3390/math14173033

APA Style

Wang, Y., Liu, N., Chu, X., Wang, T., Cheng, X., & Wang, Y. (2026). LLM-Assisted Scoring for College English Writing Assessment: Statistical Calibration Against Teacher Standards. Mathematics, 14(17), 3033. https://doi.org/10.3390/math14173033

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop