1. Introduction
Artificial Intelligence (AI) has rapidly emerged as a transformative force in education, reshaping teaching, learning, and assessment practices. Advances in machine learning and natural language processing have enabled the development of automated systems capable of evaluating student performance, providing feedback, and supporting decision-making processes in educational contexts (
Bulut et al., 2024;
Kaur & Kumari, 2025;
Munaye et al., 2025;
Raza & Alam, 2025). AI-driven grading systems—such as Automated Essay Scoring (AES) and Large Language Model (LLM)-based evaluators—have attracted growing attention due to their potential to increase efficiency, consistency, and scalability in student assessment (
Gobrecht et al., 2024;
Yang et al., 2024).
Student assessment, however, remains a complex and critical component of the educational process. Traditional human grading is inherently influenced by subjectivity, cognitive biases, and inter-rater variability, which may lead to inconsistencies and perceived unfairness in evaluation outcomes (
Litman et al., 2021;
Ntumi et al., 2025). Empirical evidence suggests that different teachers often assign different scores to the same student work, even when using standardized rubrics, due to variations in interpretation, expectations, and contextual factors (
Najafi & Subramanian, 2025). Additionally, grading large volumes of student responses imposes a significant cognitive and time burden on educators, further increasing the likelihood of inconsistency and error (
Pohn et al., 2025).
In response to these limitations, AI-based grading systems have been proposed as tools capable of enhancing objectivity and consistency in assessment. By applying standardized evaluation criteria across all responses, AI systems can reduce random variation and mitigate certain forms of human bias (
Gobrecht et al., 2024). Empirical studies have demonstrated that automated grading systems can achieve moderate to high levels of agreement with human evaluators, often reaching reliability levels comparable to those of trained raters (
Ntumi et al., 2025). At the same time, these systems offer substantial efficiency gains, enabling rapid processing of large datasets and immediate feedback generation (
Bulut et al., 2024).
Despite these advantages, the use of AI in educational assessment raises significant concerns. A growing body of research highlights the potential for algorithmic bias, lack of transparency, and limitations in evaluating complex cognitive skills such as creativity and critical thinking (
Yang et al., 2024;
Litman et al., 2021). AI grading systems may exhibit systematic deviations from human judgments such as proportional bias—where automated scores differ depending on performance level, raising questions about fairness and validity in evaluation (
Ntumi et al., 2025). Furthermore, the “black box” nature of many AI models complicates the interpretation of grading decisions and may reduce trust among educators (
Sessler et al., 2025).
Although existing research provides valuable insights into the performance of AI grading systems, important gaps remain. First, much of the empirical evidence is derived from higher education or English-language contexts, with limited research conducted in school-level environments or non-English educational systems (
Yang et al., 2024;
Ntumi et al., 2025). Second, relatively few studies combine statistical comparison of grading outcomes with an analysis of teachers’ perceptions, despite the importance of user acceptance for the implementation of AI in educational practice. Third, while prior research has established general agreement between AI and human grading, less attention has been given to the systematic patterns of divergence between the two approaches, particularly in relation to grading consistency and subjectivity.
The field of docimology has long documented the inherent variability of human grading practices. Research has shown that scores assigned to the same student work may differ across raters and even across repeated evaluations by the same rater due to contextual, cognitive, and procedural influences (
Al jadaan et al., 2025). Consequently, human assessment cannot always be regarded as a perfectly objective or error-free benchmark. This perspective is particularly relevant when evaluating AI-assisted grading systems, as comparisons between human and AI evaluations should be interpreted in terms of agreement and consistency rather than absolute correctness.
Based on the literature review and the identified research gap, the present study addresses the following research questions:
RQ1. To what extent do AI-generated grades agree with grades assigned by human teachers in primary school mathematics assessments?
RQ2. Do statistically significant differences exist between human grading and AI-generated grading outcomes across different AI systems?
RQ3. Do AI grading systems exhibit systematic scoring patterns that differ from human evaluators?
RQ4. How do teachers perceive the use of AI-assisted grading in terms of reliability, fairness, efficiency, and pedagogical value?
RQ5. Does teachers’ familiarity with AI technologies influence their acceptance of AI-assisted grading?
The contribution of this research is twofold. Empirically, it provides evidence for the reliability and consistency of AI-assisted grading in a real educational setting. Theoretically, the study contributes to the discussion on grading consistency and systematic scoring differences by examining how AI-generated evaluations compare with human judgments under a common assessment framework. By integrating statistical findings with pedagogical perspectives, the study contributes to the development of more balanced, human-centered approaches to AI-supported assessment.
Overall, this research seeks to clarify whether artificial intelligence can function as a reliable support tool in educational evaluation and whether systematic differences emerge between AI-generated and human-assigned grades that warrant continued human over-sight.
3. Materials and Methods
3.1. Research Design
The present study adopts a quantitative comparative research design to examine the differences between human and artificial intelligence (AI) grading, and to explore the role of AI in the assessment of student-written responses. Specifically, the study compares grading outcomes produced by human teachers with those generated by AI systems.
A comparative design is appropriate for identifying systematic differences and levels of agreement between two evaluation methods applied to the same dataset. In this study, student responses were independently evaluated by human raters and three AI systems (GPT-5.0, Gemini 2.x variants, and DeepSeek V3/R1), allowing for direct comparison of scoring patterns. In addition, a survey component was incorporated to capture educators’ attitudes toward AI-assisted grading, providing a complementary perspective on the practical applicability of such systems.
To ensure clarity, transparency, and replicability, the research followed a structured experimental procedure consisting of three core components: (a) a written assessment task, (b) a dual grading process (human and AI), and (c) a teacher perception survey.
Initially, a structured written assessment task was designed and administered to 10–11-year-old students under controlled classroom conditions. The assessment was conducted during regular classroom hours in a primary school setting. Students completed the task within a fixed 45 min period under standardized conditions. The duration corresponded to a typical instructional period in Greek primary education and was identical for all participants. The task consisted of mathematical problems (true/false and multiple-choice exercises, simple calculations and open-ended problems) requiring both computational accuracy and short written explanations of reasoning. This format was intentionally selected to simulate authentic classroom assessment conditions, while enabling the evaluation of both objective and semi-interpretative performance criteria.
Following data collection, all student responses were independently evaluated through a dual grading process. Human grading was conducted by qualified primary school teachers using a common analytical rubric consisting of five criteria: (a) Content/Calculation Accuracy (40%), (b) Application of Methodology/Solution Process (30%), (c) Justification/Interpretation of Reasoning (10%), (d) Presentation/Organization of Work (10%), and (e) Use of Terminology and Mathematical Language (10%). Each teacher evaluated a subset of responses, with partial overlap across raters to enable inter-rater comparison.
In parallel, the same responses were submitted to three AI systems (GPT-5.0, Gemini 2.x variants, and DeepSeek V3/R1), which were prompted using an identical evaluation protocol and the same rubric criteria. This parallel grading structure ensured that the only independent variable in the evaluation process was the type of grader (human vs. AI).
To enhance methodological rigor, AI evaluations were standardized through a fixed prompting framework and a repeated scoring procedure, with the mean score used to re-duce stochastic variation in model outputs. All scores (human and AI) were recorded in a unified dataset for subsequent statistical analysis.
Finally, a structured questionnaire was administered to participating teachers to capture their perceptions regarding the use of AI in grading. The questionnaire included Likert-scale items examining perceived fairness, reliability, time efficiency, and the pedagogical implications of AI-assisted assessment.
The study did not employ methodological triangulation in the strict mixed-methods sense, as all data sources were quantitative in nature. However, two complementary sources of evidence were incorporated: (a) the comparison of AI-generated and human-assigned grades, and (b) teachers’ perceptions collected through the questionnaire. The combination of these datasets allowed the findings from the grading analyses to be interpreted alongside educators’ views regarding the use of AI in assessment.
Overall, this multi-stage design enables a comprehensive investigation of both the quantitative differences in grading outcomes and the subjective perceptions of educators, thereby allowing the technical findings to be interpreted alongside their pedagogical implications.
3.2. Sample
The student sample consisted of 139 written assessment responses collected from primary school students aged 10–11 years. All responses were anonymized prior to analysis.
Teacher participation occurred in two different phases of the study. First, twelve primary school teachers from seven public primary schools participated in the grading process. Those teachers from the participating schools independently evaluated student responses using the common analytical rubric. The participating schools were 1st Primary School of Lefkimmi (Corfu), 2nd Primary School of Lefkimmi (Corfu), 4th Primary School of Lefkimmi (Corfu), Primary School of Vragkaniotika (Corfu), 1st Primary School of Var-tholomio (Ilia County), 2nd Primary School of Vartholomio (Ilia County), and 1st Primary School of Gastouni (Ilia County).
Second, a broader sample of 95 teachers completed the questionnaire examining perceptions of AI-assisted grading. Questionnaire participation was anonymous; therefore, it was not possible to determine whether teachers involved in the grading process were also among the questionnaire respondents.
The distinction between the grading sample (
n = 12) and the questionnaire sample (
n = 95) is important for interpreting the results. Analyses concerning grading agreement between human and AI evaluators were based on the grading sample, whereas analyses of teacher perceptions and familiarity with AI tools were based on the questionnaire sample. The characteristics of the study samples are presented in
Table 1.
3.3. Research Instruments
3.3.1. Written Assessment Task
The primary dataset consisted of 139 student assessment scripts collected from fifth- and sixth-grade primary school students.
Two curriculum-aligned mathematics assessment forms were developed, one for Grade 5 and one for Grade 6. The assessment tasks reflected the expected learning outcomes of Grades 5 and 6. The items ranged from basic procedural exercises to more demanding open-ended problem-solving tasks requiring justification and mathematical reasoning. The Grade 5 assessment focused on operations with natural numbers, whereas the Grade 6 assessment focused on decimal numbers. The complete assessment instruments, including anonymized examples of student responses, are provided in
Appendix A (Grade 5) and
Appendix B (Grade 6).
The Grade 5 assessment consisted of four sections: (a) multiple-choice arithmetic items, (b) computational exercises, (c) conceptual true/false questions, and (d) a multi-step word problem requiring mathematical reasoning and written explanation.
The Grade 6 assessment followed a similar structure and included number comparison tasks, decimal calculations, conceptual true/false items, and two applied problem-solving activities requiring both numerical solutions and written justification.
Specifically, the Grade 5 assessment consisted of a total of 15 individual scoring items, while the Grade 6 assessment comprised a total of 19 individual scoring items. These items ranged from discrete procedural exercises to multi-step open-ended problem-solving tasks.
The maximum score was 100 points. Student responses were evaluated using a common analytical rubric that considered both mathematical accuracy and the quality of reasoning demonstrated in the written responses.
The assessments were administered under standardized classroom conditions during regular school hours. Students completed the assessment individually within a 45 min session, corresponding to a typical instructional period in Greek primary education.
AI systems evaluated the complete assessment scripts rather than only the open-ended components, allowing comparison between human and AI grading across the full range of assessment tasks.
3.3.2. Assessment Rubric
A standardized analytical rubric was developed and provided to all participating teachers. The rubric was specifically designed for the assessment of primary school mathematics tasks and included five evaluation dimensions: (a) Accuracy of Content and Calculations (40%), (b) Application of Methodology and Problem-Solving Procedures (30%), (c) Justification and Interpretation of Reasoning (10%), (d) Organization and Presentation of Work (10%), and (e) Use of Mathematical Language and Terminology (10%).
Each criterion was evaluated using a five-level performance scale ranging from 0 (insufficient performance) to 4 (excellent performance). Teachers received identical scoring forms and applied the same rubric when evaluating student responses. The final score for each criterion was weighted according to its predefined percentage contribution, and the total grade was calculated as the weighted sum of all rubric dimensions.
The same rubric and weighting scheme were incorporated into the prompting protocol used for the AI grading systems to ensure methodological equivalence between human and AI evaluations. The complete assessment rubric is presented in
Table 2, whereas the original assessment rubric used during the evaluation process is provided in
Appendix C.
Final Score Calculation:
Twelve qualified primary school teachers from seven participating schools served as human raters. Since mathematics is taught by general classroom teachers in Greek primary education, all raters were regularly involved in mathematics instruction and assessment. Each student script was independently evaluated by two teachers from the same school or from neighboring schools depending on local availability, using the common analytical rubric. The teachers worked independently and were blinded to the evaluations of other raters and the AI systems. Teachers did not grade all scripts included in the study; instead, each pair of raters evaluated a subset of responses, allowing the calculation of inter-rater agreement while maintaining the authenticity of the local assessment context.
For the purposes of comparison with AI-generated scores, the final score for each response was calculated as the arithmetic mean of the two teacher ratings. This procedure was adopted to reduce the influence of individual grading variability and to provide a more stable human reference score.
The total score was calculated as a weighted sum of all criteria, with each criterion contributing proportionally according to its predefined weight.
3.3.3. AI Grading Systems
Three large language models (LLM)-based AI systems were used for the automated evaluation of student responses:
OpenAI GPT-5.0 (web-based ChatGPT Plus interface, accessed between November and December 2025).
Google Gemini 2.x variants (web-based Gemini interface, accessed between November and December 2025).
DeepSeek V3/R1 Chat (web-based DeepSeek interface, accessed between December 2025 and January 2026).
All systems were used with default platform settings. Because evaluations were conducted through publicly available web interfaces rather than APIs, temperature and other generation parameters were not accessible to the researcher.
Data collection was conducted between November and December 2025. The versions available through the respective web interfaces during this period were used. All systems operated under their default platform settings, and no manual adjustments were made to generation parameters.
To ensure methodological consistency, all student responses were converted into digital text through an OCR procedure. The OCR output was manually reviewed and corrected by the researcher before submission to the AI systems, ensuring that all models receive identical input data.
A standardized evaluation protocol was applied across all AI systems. Each model received the same grading rubric, identical instructions, identical student responses, and the same scoring framework. The rubric consisted of five weighted criteria: (a) Accuracy of Content and Calculations (40%), (b) Application of Methodology and Problem-Solving Procedures (30%), (c) Justification of Reasoning (10%), (d) Organization and Presentation (10%), and (e) Use of Mathematical Language (10%).
Each response was evaluated twice by every AI system. The Mean of the two generated scores was used as the final AI score to reduce potential stochastic variation in model outputs.
No external answer key was provided to the AI systems. The models evaluated student responses using their internal mathematical reasoning capabilities together with the analytical rubric and grading instructions.
The complete evaluation prompt is provided in
Appendix D.
3.3.4. Teacher Questionnaire
A structured questionnaire was administered to teachers as mentioned above, to assess their perceptions of AI-assisted grading. The questionnaire consisted of six dimensions assessing teachers’ perceptions of AI-assisted grading: (a) Consistency and Reliability (4 items), (b) Fairness and Equity (4 items), (c) Pedagogical Judgment and Human Dimension (4 items), (d) Transparency, Risks and Limitations (3 items), (e) Human-in-the-Loop Use (4 items), and (f) Acceptance and Intention to Use (2 items).
All items were measured using a five-point Likert scale ranging from 1 (Strongly Disagree) to 5 (Strongly Agree). Several negatively worded items were reverse-coded prior to analysis. Composite scores were calculated as the arithmetic means of the items belonging to each dimension.
The questionnaire was developed by the researcher specifically for the purposes of the present study, following an extensive review of the literature on AI-assisted assessment. Item construction was theory-driven and informed by previous research on grading consistency, human grading variability, perceived fairness, algorithmic bias, transparency, trust, technology acceptance, and human–AI collaboration. Rather than adopting an existing instrument in its entirety, items were synthesized and adapted from multiple studies to address the specific objectives of the present research. The psychometric properties of the resulting instrument were subsequently examined through exploratory factor analysis and reliability analyses.
A mapping of questionnaire dimensions to the relevant theoretical constructs and supporting literature was performed during the instrument development process. The questionnaire structure is summarized in
Table 3.
Prior to the main analysis, the psychometric properties of the questionnaire were examined. Exploratory Factor Analysis (EFA) using Principal Axis Factoring with Promax rotation was conducted to investigate the underlying structure of the instrument.
The analysis supported a multidimensional structure consisting of six conceptual domains: Consistency and Reliability, Fairness and Equity, Pedagogical Judgment, Transparency and Ethics, Human-in-the-Loop Use, and Intention to Use.
Internal consistency was assessed using Cronbach’s alpha coefficients. Reliability estimates ranged from 0.706 to 0.869 across questionnaire dimensions, indicating acceptable to good internal consistency for most scales. The reliability coefficients for each questionnaire dimension are presented in
Table 4.
The Fairness and Equity dimension demonstrated lower internal consistency (α = 0.453) compared with the remaining scales. Inspection of the item statistics suggested conceptual heterogeneity among fairness-related perceptions, particularly regarding concerns about potential disadvantages for students from diverse linguistic or social backgrounds. Because the questionnaire was developed specifically for the purposes of the present study, and the exploratory factor analysis was intended to provide preliminary psychometric evidence rather than to redefine the theoretically derived questionnaire structure, this dimension was retained in the analyses despite its lower reliability. Consequently, findings related to the Fairness and Equity dimension should be regarded as exploratory rather than confirmatory and interpreted with caution.
3.4. Data Collection Procedure
The data collection process consisted of two major parts, which were further divided into four sequential stages.
The first part of the data collection, concerning student responses, included the following steps. First, student responses were collected through a standardized written assessment task. Second, the responses were independently graded by human teachers using an established rubric. Third, the same responses were evaluated by AI systems under equivalent conditions.
The second part of the data collection concerned the teachers. Specifically, in the fourth stage, participating teachers completed a perception questionnaire regarding AI-assisted grading.
All data were anonymized and stored in a structured dataset, including separate variables for human scores, AI scores, and questionnaire responses.
3.5. Statistical Analysis
Data analysis was conducted using IBM SPSS Statistics for Windows, Version 26.0 (IBM Corp., Armonk, NY, USA). Quantitative statistical methods included descriptive statistics, paired-samples t-tests, Intraclass Correlation Coefficient (ICC), and Bland–Altman analysis. Statistical significance was set at p < 0.05.
3.5.1. Descriptive Statistics
Descriptive statistics (mean, standard deviation, minimum, maximum) were calculated to summarize grading outcomes.
3.5.2. Comparative Analysis
Paired samples t-tests were used to examine differences between human and AI grading scores. Statistical significance was set at p < 0.05. Because three pairwise comparisons between human and AI grading systems were conducted, a Bonferroni-adjusted significance threshold of p < 0.017 was additionally considered to control for Type I error inflation. Effect sizes were calculated using Cohen’s d to assess the magnitude of differences.
3.5.3. Agreement and Reliability
The level of agreement between human and AI grading was assessed using the Intraclass Correlation Coefficient (ICC, two-way mixed effects model, absolute agreement). Values above 0.70 were interpreted as indicating acceptable reliability.
3.5.4. Bias Analysis
Bland–Altman analysis was employed to assess agreement between grading methods and to detect potential systematic bias. This method involves plotting the differences between two measurement methods against their mean values and calculating the limits of agreement (mean difference ± 1.96 SD). The presence of proportional bias is examined by assessing whether differences vary systematically across the range of scores. Narrow limits of agreement indicate high agreement between methods, whereas wider limits suggest greater variability and lower agreement at the individual level.
3.5.5. Questionnaire Analysis
Questionnaire data were analyzed using descriptive statistics and inferential tests (independent samples t-tests and one-way ANOVA), depending on the distribution of demographic variables.
3.6. Validity and Reliability of Research Instruments
The validity and reliability of the research instruments were examined separately for each tool used in the study.
Regarding the assessment rubric, content validity was ensured through its alignment with the Primary Education Mathematics Curriculum and the relevant literature on educational assessment and automated grading systems. The criteria were designed to reflect core dimensions of student performance, ensuring conceptual consistency and applicability across both human and AI grading.
In contrast, the teacher questionnaire was evaluated for internal consistency using Cronbach’s alpha coefficient. The analysis was conducted separately for each scale of the questionnaire, with values above 0.70 considered acceptable for reliability. This procedure ensured that the questionnaire items consistently measured the intended constructs, such as perceived fairness, reliability, and efficiency of AI-assisted grading.
3.7. Ethical Considerations
All procedures complied with ethical standards for educational research. Participation was voluntary, and informed consent was obtained from all participants. Student data were anonymized, and no personal identifiable information was included in the analysis.
4. Results
4.1. Descriptive Statistics
Descriptive statistical analysis was conducted to provide an overview of the grading outcomes produced by human evaluators and artificial intelligence systems.
As presented in
Table 5, the mean score assigned by teachers was M = 6.97 (SD = 2.48), while the corresponding mean scores for the AI systems were M = 6.11 (SD = 2.14) for GPT-5.0, M = 6.97 (SD = 2.30) for Gemini, and M = 6.57 (SD = 2.41) for DeepSeek.
Overall, the results indicate that AI-generated scores were broadly comparable to teacher evaluations. Gemini produced nearly identical mean scores to human raters, while GPT-5.0 and DeepSeek showed slightly lower average scores.
In terms of variability, AI systems generally demonstrated lower or comparable standard deviations relative to teachers, suggesting more consistent grading patterns.
4.2. Human Inter-Rater Reliability
Inter-rater reliability between the two human evaluators was assessed using Pearson correlation and the Intraclass Correlation Coefficient (ICC).
The analysis revealed substantial agreement between the two teacher ratings (r = 0.794,
p < 0.001; ICC = 0.750), indicating satisfactory consistency between independent human raters. These findings support the use of the averaged teacher score as the human reference measure for subsequent comparisons with AI-generated grades. The inter-rater reliability results are presented in
Table 6.
4.3. Comparison Between Human and AI Grading
To examine whether statistically significant differences existed between human and AI grading, paired samples t-tests were conducted.
As shown in
Table 7, significant differences were identified between teacher scores and two AI systems. Specifically, GPT-5.0 and DeepSeek tended to assign systematically different scores compared to human evaluators, whereas Gemini demonstrated near-identical scoring behavior.
The largest difference was observed for GPT-5.0 (Mean Difference = 0.853, 95% CI [0.619, 1.088]), followed by DeepSeek (Mean Difference = 0.395, 95% CI [0.155, 0.634]). In contrast, no statistically significant difference was found between human ratings and Gemini scores (Mean Difference = −0.007, 95% CI [−0.232, 0.218]).
Effect sizes (Cohen’s d) were moderate for GPT-5.0 (d = 0.61), small for DeepSeek (d = 0.28), and negligible for Gemini (d ≈ 0.00).
Overall, these findings suggest that AI systems approximated human grading outcomes at the aggregate level, although meaningful differences remained across models and individual student responses.
All statistically significant findings remained significant after Bonferroni correction (adjusted α = 0.017).
4.4. Agreement and Reliability Analysis
Beyond mean comparisons, the level of agreement between human and AI grading was examined using the Intraclass Correlation Coefficient (ICC).
The ICC for single measures was 0.777 (95% CI [0.716, 0.829], p < 0.001), indicating a good level of agreement at the level of individual raters. In contrast, the ICC for average measures was 0.946 (95% CI [0.927, 0.960], p < 0.001), demonstrating an excellent level of overall agreement across all evaluators.
These findings suggest that, although some variability exists at the level of individual grading, the overall consistency between human and AI grading is particularly high when aggregated scores are considered.
4.5. Bland–Altman Analysis
Bland–Altman analyses were conducted to further examine agreement between teacher evaluations and AI-generated scores. Unlike correlation coefficients and ICC values, Bland–Altman plots provide information about systematic differences between raters and allow the identification of potential proportional bias across the scoring range.
For GPT, the mean difference indicated a tendency toward stricter grading compared with human evaluators. Although most observations fell within the limits of agreement, the dispersion of differences increased at higher score levels. Subsequent regression analysis revealed a statistically significant proportional bias (β = 0.159, 95% CI [0.056, 0.263],
p = 0.003), indicating that discrepancies between GPT and teacher scores tended to increase as average student performance increased. The corresponding Bland–Altman plot is presented in
Figure 1.
For Gemini, the mean difference was close to zero and the distribution of differences appeared relatively balanced across the scoring range. Regression analysis did not reveal significant proportional bias (β = 0.080, 95% CI [−0.017, 0.178],
p = 0.106), suggesting stable agreement with teacher evaluations regardless of performance level. The corresponding Bland–Altman plot is presented in
Figure 2.
Similarly, DeepSeek demonstrated relatively small deviations from teacher scores across the full range of performance. No significant proportional bias was detected (β = 0.032, 95% CI [−0.072, 0.135],
p = 0.546), indicating that differences between DeepSeek and human evaluations remained largely constant. The corresponding Bland–Altman plot is presented in
Figure 3.
Overall, the Bland–Altman findings suggest that agreement patterns differed across AI systems. GPT exhibited evidence of performance-dependent divergence from teacher grading, whereas Gemini and DeepSeek maintained more stable agreement across the scoring spectrum.
4.6. Teacher Questionnaire Results
Descriptive statistical analysis of the questionnaire responses indicated generally positive but cautious attitudes toward AI-assisted grading.
Teachers reported relatively high perceptions of the consistency and reliability of AI systems (M = 3.77, SD = 0.59), while perceived fairness was moderately high (M = 3.54, SD = 0.74). In contrast, low scores were observed in the dimension of pedagogical adequacy (M = 1.85, SD = 0.61), indicating concerns regarding the ability of AI to evaluate higher-order cognitive skills such as critical thinking and creativity.
Concerns regarding risks and limitations of AI use were also notable (M = 3.77, SD = 0.82), particularly in relation to transparency and potential bias. Conversely, the highest mean values were observed in the dimensions of hybrid use (M = 4.07, SD = 0.67) and overall acceptance (M = 3.89, SD = 0.84), indicating a strong preference for a human-in-the-loop approach.
To further examine whether teacher perceptions differed across levels of familiarity with AI tools, a one-way ANOVA was conducted. The analysis revealed a statistically significant effect of familiarity on overall acceptance of AI (F (4, 90) = 3.186, p = 0.017). The assumption of homogeneity of variances was satisfied (Levene’s test, p = 0.142), and the robustness of the result was confirmed using the Welch test (p = 0.021). The effect size was moderate (η2 = 0.124).
Post hoc comparisons using Tukey HSD indicated that teachers with very high familiarity reported significantly higher acceptance compared to those with low familiarity (p = 0.044), while no significant differences were observed between the remaining groups.
As illustrated in
Figure 4, a clear increasing trend is observed, with higher levels of familiarity associated with higher levels of acceptance. The visual pattern of confidence intervals further supports the statistical findings.
Overall, these results suggest that familiarity with AI technologies plays a significant role in shaping teachers’ acceptance of AI-assisted grading, while broader pedagogical concerns remain.
4.7. Relationships Among Questionnaire Dimensions
Pearson correlation analyses were conducted to examine relationships among the questionnaire dimensions. Significant positive correlations were observed between perceived consistency and fairness (r = 0.658, p < 0.001), consistency and acceptance of AI-assisted grading (r = 0.551, p < 0.001), and fairness and acceptance (r = 0.626, p < 0.001). Human-in-the-loop use was also positively associated with consistency (r = 0.580, p < 0.001), fairness (r = 0.531, p < 0.001), and acceptance (r = 0.602, p < 0.001).
Conversely, the Pedagogical Judgment dimension was negatively associated with Human-in-the-Loop use (r = −0.474, p < 0.001) and Transparency/Risk concerns (r = −0.319, p = 0.002), suggesting that teachers who expressed stronger reservations regarding the ability of AI systems to capture pedagogical aspects of assessment were more likely to favor continued human involvement in grading processes.
Overall, the findings indicate that perceptions of consistency and fairness are closely linked to teachers’ willingness to accept AI-assisted grading, whereas concerns regarding pedagogical adequacy remain an important source of caution. The Pearson correlation coefficients among the questionnaire dimensions are presented in
Table 8.
5. Discussion
5.1. Interpretation of Findings
The present study examined the role of artificial intelligence in student grading by comparing AI-generated scores with human teacher evaluations and by exploring teachers’ perceptions regarding AI-assisted assessment.
The findings indicate that AI systems do not behave uniformly. While Gemini produced scores that were statistically indistinguishable from those of human evaluators, GPT-5.0 and DeepSeek demonstrated significant differences, suggesting model-specific grading tendencies. This result aligns with previous studies demonstrating strong agreement between automated grading systems and human raters (
Gobrecht et al., 2024;
Ntumi et al., 2025).
Beyond the observed differences between AI models, the descriptive analysis revealed subtle but consistent differences in scoring patterns. AI systems tended to exhibit lower variability compared to human evaluators, indicating more standardized grading behavior. This finding is consistent with the argument that automated systems may apply evaluation criteria more consistently, as they are not affected by cognitive fatigue, emotional factors, or contextual influences (
Bulut et al., 2024).
However, consistency should not be conflated with accuracy. Although AI systems produced stable grading outcomes, the observed deviations indicate that agreement with human evaluators varied according to both the assessment task and the AI model, suggesting systematic rather than purely random scoring tendencies. This interpretation is consistent with prior research highlighting limitations of AI systems when evaluating higher-order cognitive skills and complex reasoning processes (
Yang et al., 2024).
Among them, Gemini demonstrated the highest level of agreement with the human reference scores in the present study. However, the design of this research does not allow conclusions regarding the specific mechanisms responsible for this outcome. Since the study focused on comparing grading outcomes rather than investigating model architecture or training characteristics, any explanation of the observed differences between AI systems would remain speculative. Future research could explore the factors that contribute to variations in agreement across AI models under different assessment conditions.
5.2. Agreement, Reliability, and Bias
The agreement analysis demonstrated moderate to high levels of reliability between AI-generated scores and teacher evaluations. This finding suggests that AI systems may function as useful support tools in educational assessment, particularly when combined with human oversight.
Nevertheless, the Bland–Altman analysis revealed that agreement between AI and human grading was not uniform across all cases. While most differences fell within acceptable limits, larger discrepancies were observed in responses requiring nuanced interpretation. This pattern suggests the presence of systematic scoring differences rather than purely random variation.
Such differences may reflect characteristics of the grading approach adopted by AI systems, including differential sensitivity to specific response features or performance levels. Therefore, although AI systems can contribute to greater consistency in scoring, their evaluations may still diverge systematically from human judgments under certain conditions. This distinction is critical for understanding the role of AI in assessment, as it challenges the assumption that automated grading is inherently objective (
Litman et al., 2021;
Ntumi et al., 2025).
This finding aligns with previous research reporting proportional scoring tendencies in automated assessment systems, whereby grading behavior may vary across levels of student performance. Consequently, agreement statistics should be interpreted alongside analyses of systematic scoring patterns, as high overall agreement does not necessarily imply interchangeability at the individual response level, highlighting the importance of complementary analyses such as Bland–Altman plots and bias assessment procedures.
The present findings should not be interpreted as evidence of demographic or sociocultural algorithmic bias. Rather, they indicate the existence of systematic differences between AI-generated and human-assigned scores that warrant careful consideration when AI systems are used in educational assessment contexts. Accordingly, the present study evaluates agreement, consistency, and systematic scoring patterns between human and AI grading rather than algorithmic fairness in the broader sense used in the AI ethics literature.
Moreover, these findings should be interpreted within the broader docimological perspective. Human grading itself is characterized by a degree of variability, as documented in the assessment literature (
Al jadaan et al., 2025). Differences observed between AI-generated and human-assigned scores therefore do not necessarily indicate deficiencies in either system but rather reflect the complexity of educational assessment and the absence of a universally accepted grading standard.
Additional proportional bias analyses demonstrated that GPT was the only AI system exhibiting performance-dependent divergence from teacher evaluations. Specifically, discrepancies between GPT and human scores increased at higher levels of student performance (β = 0.159, 95% CI [0.056, 0.263], p = 0.003), whereas Gemini (β = 0.080, p = 0.106) and DeepSeek (β = 0.032, p = 0.546) maintained relatively stable agreement throughout the scoring range. This finding suggests that systematic grading tendencies may differ across AI models and should be considered when interpreting automated assessment outcomes.
Although the average measures for ICC indicated excellent overall agreement, the individual measures for ICC were notably lower. This distinction suggests that agreement is substantially stronger when scores are aggregated than when individual evaluations are considered separately. Therefore, the findings should be interpreted as evidence of strong overall consistency rather than perfect correspondence at the level of individual grading decisions.
Similarly, the Bland–Altman analyses demonstrated that, despite satisfactory agreement at the group level, meaningful discrepancies remained possible for individual student responses. These findings highlight the importance of distinguishing between population-level consistency and the practical implications of grading decisions applied to individual students.
5.3. Teacher Perceptions and Practical Implications
The findings from the teacher questionnaire provide important insights into the practical implementation of AI-assisted grading. Descriptive results showed that teachers perceive AI systems as relatively consistent (M = 3.77, SD = 0.59) and moderately fair (M = 3.54, SD = 0.74), indicating a cautiously positive stance toward their evaluative reliability.
At the same time, the pedagogical dimension received a notably low mean score (M = 1.85, SD = 0.61), suggesting that educators do not consider AI capable of adequately capturing deeper aspects of student performance, such as critical thinking, creativity, and conceptual understanding. This finding highlights a fundamental limitation of AI-assisted grading from a pedagogical perspective.
In addition, teachers reported relatively high levels of concern regarding risks and limitations (M = 3.77, SD = 0.82), particularly in relation to transparency, potential sources of bias, and over-reliance on automated systems. Despite these concerns, the highest mean score was observed in the dimension of hybrid use of AI (M = 4.07, SD = 0.67), followed by general acceptance (M = 3.89, SD = 0.84), indicating a strong preference for a human-in-the-loop approach.
Taken together, these findings suggest that teachers are willing to adopt AI technologies if they remain under human supervision and are used as supportive rather than autonomous evaluative tools. Hybrid assessment models, in which AI performs preliminary grading and teachers retain final evaluative authority, appear to offer a balanced approach that leverages the strengths of both human and automated evaluation (
Varghese et al., 2025).
Perhaps the most noteworthy finding of the questionnaire concerns the exceptionally low rating assigned to the pedagogical adequacy of AI-assisted grading. While statistical analyses indicated substantial agreement between AI systems and human evaluators, teachers remained unconvinced that AI could adequately capture deeper educational dimensions of student performance. This contrast highlights a fundamental distinction between measurement consistency and pedagogical validity and suggests that technical performance alone may be insufficient to secure educators’ trust in AI-supported assessment.
5.4. Theoretical Contribution
This study contributes to the existing literature in several ways.
First, it provides empirical evidence from a school-level context, addressing a gap in research that has predominantly focused on higher education environments.
Second, it integrates statistical analysis of grading outcomes with teacher perception data, offering a more comprehensive understanding of both the technical performance and practical acceptance of AI-assisted grading systems.
Third, the findings highlight that AI grading systems may combine high internal consistency with systematic patterns of deviation. This dual characteristic suggests that AI does not simply replicate human grading but introduces distinct evaluative dynamics that must be critically examined.
In particular, the study demonstrates that high consistency in AI grading does not necessarily translate into perceived pedagogical validity, highlighting a critical gap between statistical reliability and educational acceptance.
5.5. Limitations
Despite its contributions, the study has several limitations.
The findings are based on a specific dataset and educational context, which may limit generalizability to other subjects, age groups, or assessment formats.
The analysis focused on specific AI models, and results may vary depending on the architecture, training data, and configuration of alternative systems.
Although the statistical analyses examined agreement between human and AI grading, the study does not fully capture the qualitative aspects of student responses, such as creativity or originality, which remain challenging for automated evaluation systems.
Two distinct teacher samples were used in the study, each with its own limitations. The grading component relied on a relatively small group of human raters (n = 12), which may limit the generalizability of the inter-rater reliability findings. In contrast, the questionnaire component included a larger sample of teachers (n = 95), providing a broader perspective on educator perceptions; however, these findings remain context-specific and should be interpreted with caution when generalized to other educational settings or populations.
A further limitation concerns the use of human ratings as the reference framework. Although the study employed dual independent raters and averaged scores to improve reliability, previous docimological research has demonstrated that human grading is inherently variable and cannot be considered a perfectly objective standard (
Al jadaan et al., 2025).
Another methodological limitation concerns the potential dependency structure of the data. Student scripts were evaluated by teachers within partially overlapping grading groups rather than by a fully crossed rating design. Consequently, some degree of statistical dependency may have been present because student responses were nested within teacher grading groups, a feature that was not explicitly modeled in the statistical analyses. Future studies employing larger samples could benefit from multilevel modelling or other analytical approaches capable of explicitly accounting for the hierarchical structure of the data (e.g., student responses nested within teachers and schools). Such approaches would allow a more precise estimation of rater-level effects on human–AI grading comparisons.
An important methodological consideration concerns the dynamic nature of publicly available AI systems. The models used in this study were accessed through their web interfaces using the default platform configurations available during the data collection period. Because these systems are continuously updated by their developers, identical prompts submitted later may not necessarily produce identical grading outputs. Consequently, although the evaluation protocol can be replicated, the exact reproducibility of the AI-generated scores cannot be guaranteed across future model versions.
The findings should also be interpreted within the specific context of the study. The assessment tasks consisted of relatively structured mathematics activities administered to primary school students. Consequently, the results cannot be generalized directly to other subject areas, educational levels, or assessment formats that require more extensive written production, creativity, or complex reasoning.
Finally, the exclusive reliance on quantitative measures may overlook contextual factors influencing both human and AI grading decisions.
5.6. Implications for Practice and Future Research
The findings suggest that AI-assisted grading can enhance efficiency and consistency in educational assessment, particularly in large-scale or time-constrained contexts.
However, the presence of systematic differences and teacher concerns highlights the need for careful implementation. Educational institutions should adopt AI systems within structured frameworks that ensure transparency, fairness, and human oversight.
Future research should focus on expanding the analysis to diverse educational settings, exploring model-specific biases, and integrating qualitative approaches to better understand how AI systems evaluate complex student responses.
Additionally, further investigation is needed into the long-term pedagogical impact of AI-assisted grading on teaching practices and student learning outcomes.
6. Conclusions
This study examined the role of artificial intelligence in student assessment by comparing AI-generated grading outcomes with human teacher evaluations and by exploring teachers’ perceptions of AI-assisted grading.
The findings indicate that AI systems can produce grading results broadly comparable to those of human evaluators within the structured primary-school mathematics assessment context examined in this study, although differences were observed depending on the specific model used. While some systems demonstrated high alignment with human grading (e.g., Gemini), others exhibited statistically significant deviations (e.g., GPT-5.0 and DeepSeek), highlighting the importance of model selection in AI-assisted assessment.
In addition, AI systems showed more consistent grading behavior, as reflected in lower variability compared to human raters. However, this consistency does not necessarily imply pedagogical adequacy. The results suggest that AI may follow systematic evaluation patterns, especially in tasks requiring deeper conceptual understanding.
The analysis of teacher perceptions revealed a cautiously positive stance toward the use of AI in grading. Educators acknowledged its potential to enhance efficiency and reduce workload, but expressed concerns regarding fairness, transparency, and the ability of AI to capture complex aspects of student performance. The strong preference for hybrid approaches indicates that teachers view AI as a supportive tool rather than a replacement for human judgment.
Overall, the findings suggest that artificial intelligence can contribute to educational assessment when used within structured frameworks that ensure human oversight, transparency, and pedagogical validity.
Future research should further investigate the use of AI in diverse educational contexts, as well as its capacity to evaluate higher-order cognitive skills and its long-term impact on teaching and learning processes.