Next Article in Journal
Are There Inclusive Pedagogical Perspectives? Teachers’ Perceptions and UDL Training in Early Childhood Education (Ages 0–3)
Previous Article in Journal
Enacting Minority Language Rights Through High-Stakes Assessment: Task Design in the Slovene-Language First Written Paper of Italy’s School-Leaving Examination
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

From Concurrent Self-Assessment to Postdiction: Grade-Judgment Calibration Across Two Assessments in a First-Year Computer Science Course

1
Faculty of Informatics, University of Debrecen, 4028 Debrecen, Hungary
2
Faculty of Mathematics and Computer Science, Babeş-Bolyai University, 400084 Cluj-Napoca, Romania
*
Author to whom correspondence should be addressed.
Educ. Sci. 2026, 16(9), 1476; https://doi.org/10.3390/educsci16091476
Submission received: 31 May 2026 / Revised: 5 September 2026 / Accepted: 7 September 2026 / Published: 10 September 2026

Abstract

First-year computer-science students judged their grade on two assessments differing in judgment type: a concurrent self-assessment embedded in an early quiz (Assessment 1, N = 101 ) and a postdiction after a mid-course written-and-laboratory exam (Assessment 2, N = 117 ), with 94 students linked across both. The concurrent judgment was optimistic (mean signed error + 9.10 pp; 70.3 % overestimators; Spearman ρ = 0.533 ). The postdictions reversed the bias: students underestimated the written component by 10.08 pp and the laboratory by 4.21 pp, rank-order calibration being markedly better for the laboratory task ( ρ = 0.804 ). In the linked subsample, the reversal was large (paired d z = 1.00 ) but absolute error did not improve ( p = 0.582 ): the error changed direction, not magnitude. A tertile split shows broad-based quiz optimism and, on the postdictions, a gradient compatible with regression to the mean. An exploratory post hoc re-banding of the laboratory scores, excluding the administrative 0–1 segment, yields a Dunning–Kruger-compatible between-band difference of 1.80 raw points (95% CI [ 0.94 , 2.66 ] , p = 0.0001 ), robust to alternative bandings, though the low-band overestimation is not. The Assessment 1 overestimation replicates in two further cohorts. Calibration therefore differed markedly across two contexts that differ simultaneously in judgment type, timing, format, content and diagnostic cues, which these data cannot disentangle.

1. Introduction

The ability to evaluate one’s own knowledge and expected performance is a central component of self-regulated learning (Flavell, 1979; Schunk & Zimmerman, 1997). In higher education, students continuously judge how well they understand course content, how they have performed on an assessment, and what grade they are likely to obtain, and these judgments shape study behavior, effort allocation, help-seeking, and disciplinary self-concept (Dunlosky & Rawson, 2012). Inaccurate judgments may therefore have consequences beyond a single assessment, particularly in first-year courses that act as a gatekeeping experience for later academic progression (Falchikov & Boud, 1989; Karaca & Bowers, 2023).
Within metacognition research, the alignment between judgment and performance has been studied through the construct of calibration—the degree to which students’ confidence or performance judgments correspond to their actual outcomes (Bol & Hacker, 2012; Nietfeld et al., 2005). Better calibration has been associated with stronger academic performance, while lower-performing students are often found to be more miscalibrated and, in many settings, more overestimating (Feld et al., 2017; Kruger & Dunning, 1999; Mahmood, 2016). The phenomenon has most often been discussed through the Dunning–Kruger framework, which proposes that the skills required to perform well also enable accurate self-evaluation, creating a “dual burden” for low performers (Dunning, 2011; Dunning et al., 2003; Ehrlinger et al., 2008; Kruger & Dunning, 1999). The pattern has also been reported in computing-adjacent settings (Gibbs et al., 2017) and in Hungarian higher-education samples (Kun et al., 2023), the same broader context from which the present data are drawn. However, the framework has attracted substantial methodological criticism: several reanalyses (Gignac & Zajenkowski, 2020; Jansen et al., 2021; Lebuda et al., 2024; Magnus & Peresetsky, 2022; McIntosh et al., 2019) have shown that much of the canonical pattern can be reproduced by bounded scales, regression to the mean, and the mathematical structure of difference scores, while cross-cultural and clinical replications have yielded mixed evidence (Coutinho et al., 2024; Surdilović et al., 2022; Yang Hansen et al., 2024). For this reason, the present paper treats Dunning–Kruger as a secondary lens. The primary framing is calibration of grade judgments, direction, magnitude, and rank-order, examined in an authentic classroom context.
The empirical setting is a first-year computer science (CS1) course, a context that sharpens the metacognitive question for at least three reasons. First, although most students in this setting arrive with several years of secondary-school programming (here, M = 3.70 years; see Section 4), limited class time and uneven student motivation mean that the theoretical components of secondary CS curricula are rarely worked through in depth, so explicit reasoning over algorithm patterns (Fekete et al., 2020; Fóthi, 2012; Szlávi & Zsakó, 2016) is largely new at university (Csernoch et al., 2015; Lodi & Martini, 2021). Second, computing-education research has documented that early-undergraduate students monitor the quality of their own and their peers’ work inaccurately (Cloude et al., 2024) and that their reported confidence is only loosely tied to actual correctness (Strickroth, 2024). Third, miscalibration in an early CS1 assessment may interact with self-efficacy and persistence in the discipline (Chen et al., 2024; Han et al., 2021; Leppänen et al., 2025). Self-assessment in Hungarian and broader Central European CS contexts in particular has produced similar overconfidence findings (Kun, 2016; Máté & Darabos, 2017; Nagy et al., 2021).
The motivation for the present study is a specific classroom design in which the same students’ grade judgments were collected under two methodologically distinct conditions. The Assessment 1 estimate was the final item of a quiz administered early in the semester, so the student was making a judgment informed by the experience of having attempted the items but without any score or feedback—an immediate concurrent self-assessment (de Bruin et al., 2017). The Assessment 2 estimates were collected after the mid-course written and laboratory exam components, after engagement with the task was complete but before grades were returned, and were therefore postdictions in the sense of Hacker et al. (2000). Calibration research has repeatedly shown that timing and context shape both bias and precision (Bol & Hacker, 2012; Hartwig & Dunlosky, 2014; Nietfeld et al., 2005), but the literature rarely contrasts an embedded concurrent judgment with a subsequent postdiction in the same course. This is the gap the present paper addresses.
The manuscript therefore pursues two aims: (i) characterize the direction, magnitude, and rank-order accuracy of first-year CS students’ grade judgments separately for the Assessment 1 concurrent quiz judgment, the Assessment 2 written postdiction, and the Assessment 2 laboratory postdiction; and (ii) examine, in the 94-student linked subsample, whether the shift from a concurrent to a postdictive judgment is associated with a systematic change in calibration. A tertile-based “raw Dunning–Kruger” analysis is reported as a secondary, cautiously interpreted view.

2. Theoretical and Empirical Background

2.1. Metacognition and Calibration

Metacognition concerns learners’ awareness and regulation of their own cognitive processes (Flavell, 1979; Schunk & Zimmerman, 1997). Within this domain, calibration research treats self-assessment as an empirical, measurable relationship rather than as a subjective impression: two students may both feel confident, but calibration analysis asks whether that confidence is warranted (Bol & Hacker, 2012). Students who systematically overestimate or underestimate may make suboptimal decisions about preparation, revision, and help-seeking, with downstream effects on achievement (Callender et al., 2016; Dunlosky & Rawson, 2012; Osterhage, 2021).
Best practice in the calibration literature reports the direction of miscalibration (signed error), its magnitude (absolute error), and the rank-order correspondence between judgment and outcome (Pearson r or Spearman ρ ), because each captures a different facet of judgment quality (Hartwig & Dunlosky, 2014; McGuire, 2023; Nietfeld et al., 2005). The literature has also emphasized that calibration is heterogeneous across discipline, task complexity, experience level, and the conditions under which the judgment is made (Falchikov & Boud, 1989; Mahmood, 2016; Máté & Darabos, 2017; Prims & Moore, 2017), and meta-analytic work points to the same conclusion in non-adult populations (Xia et al., 2024). The contrast of the present study of two different judgment contexts within the same course is designed in this spirit.

2.2. Predictions, Concurrent Judgments, and Postdictions

A particularly relevant distinction is among predictions, concurrent judgments, and postdictions. Predictions are made before a task; postdictions are made after a task has been completed (Hacker et al., 2000). Postdictions are often expected to be more accurate because students can draw on experiential cues, perceived difficulty, retrieval fluency, and awareness of uncertainty gathered during the task (Bol & Hacker, 2012). Yet that expectation is not guaranteed: students may also rely on non-diagnostic cues, wishful thinking, or relief after finishing an exam, and persistent miscalibration has been documented even when practice-test feedback is provided (Osterhage, 2021; Serra & DeMarree, 2016).
Between pure predictions and standard postdictions lies a less-studied category of concurrent judgments (Máté et al., 2025), self-assessments made during, or immediately at the end of, the task and before any external feedback (de Bruin et al., 2017). When the final question of a quiz asks students to estimate how well they have performed, the student is making a judgment that is post-task in the sense that the items have been seen but pre-feedback in the sense that no score, peer comparison, or reflective period has intervened. Comparing such a concurrent judgment with a later postdiction is therefore not a clean pre/post manipulation: the two judgments differ on timing, judgment type, assessment content, and information available at the moment of judgment. The present study acknowledges this confound and treats the cross-assessment contrast as descriptive.

2.3. Grade Prediction, Desired Grades, and Bias

Grade estimates are not purely cognitive. Serra and DeMarree (2016) showed that students’ grade predictions are biased toward the grade they want to receive, a wish-based component that may be amplified after an exam in which the student has just invested effort. Overconfidence has further been decomposed into overestimation, overplacement, and overprecision (Moore & Healy, 2008); the present study addresses overestimation, the signed discrepancy between estimated and achieved score, which is the index most frequently reported in classroom calibration research (Bol & Hacker, 2012; Karaca & Bowers, 2023). Broader work on cognitive biases in educational decision-making (Friedman, 2023; Máté et al., 2025) further suggests that students arriving in a first-year course may anchor their estimates on preconceptions formed in secondary school rather than on task-specific evidence.

2.4. The Dunning–Kruger Debate

The Dunning–Kruger effect, the observation that lower-performing individuals overestimate while higher-performing individuals underestimate, is one of the most widely cited findings in psychology (Ehrlinger et al., 2008; Kruger & Dunning, 1999). It has, however, attracted sustained methodological criticism. Gignac and Zajenkowski (2020) and Magnus and Peresetsky (2022) demonstrate that much of the canonical pattern is reproducible from bounded scales and regression to the mean. Jansen et al. (2021) reinterpret it through a rational model with noisy evidence. Hartwig and Dunlosky (2014) and McIntosh et al. (2019) show that the effect’s apparent magnitude depends on the chosen judgment scale and on which metacognitive measure is used. Lebuda et al. (2024) and Coutinho et al. (2024) find more specific patterns when alternative analytic approaches are applied, and cross-cultural and applied replications likewise produce mixed evidence (Surdilović et al., 2022; Yang Hansen et al., 2024). The present paper does not interpret tertile patterns as direct evidence for the dual-burden hypothesis; it reports them transparently as the “raw” pattern, alongside the calibration metrics that are the main focus.

2.5. Relevance to First-Year Computer Science

The CS1 context sharpens the metacognitive question. Cloude et al. (2024) document that novice programmers inaccurately monitor the quality of their own and their peers’ code; Strickroth (2024) reports that self-confidence in programming solutions correlates with performance only loosely. Gorson and O’Rourke (2020) show that CS1 students’ negative self-assessments are shaped by peer comparison and the experience of struggle; Chen et al. (2024) describe the qualitative reasoning underlying those self-assessments. The “confidence gap” in university programming has been re-examined recently (Leppänen et al., 2025), and Sobral (2021) specifically reports persistent optimism bias in CS1 grade prediction. Studies of computational-thinking and algorithmic-skill assessments have noted that assessment format itself shapes how students judge their performance (Csernoch et al., 2015; Kang et al., 2023; Zhang et al., 2024). In the Hungarian setting from which the present data are drawn, self-assessment and self-efficacy have been studied in spreadsheet and Sprego contexts (Csapó et al., 2020, 2021; Máté et al., 2025; Nagy et al., 2021), with comparable overconfidence patterns. The present study sits in this strand of work while contributing a methodologically distinct judgment-type pairing within a single course.

3. Aim, Research Questions, and Hypotheses

3.1. Aim

The study examines how accurately first-year CS students judge their own performance across two course assessments that differ in judgment type, a concurrent self-assessment embedded in an early quiz (Assessment 1) and a postdiction made after a mid-course written-and-laboratory exam (Assessment 2). It evaluates direction, magnitude, and rank-order calibration separately for Assessment 1, for each Assessment 2 component, and for the 94-student linked subsample present in both assessments, and reports a secondary tertile-based “raw Dunning–Kruger” view.

3.2. Research Questions

RQ1. 
How accurately do first-year CS students estimate their performance on an early-course quiz when the estimate is an embedded concurrent self-assessment (Assessment 1)?
RQ2. 
How accurately do they estimate their performance on a mid-course written and laboratory exam when the estimate is a postdiction (Assessment 2, written and laboratory components)?
RQ3. 
Does the direction of miscalibration suggest systematic overconfidence, underconfidence, or a mixture at each judgment occasion?
RQ4. 
Within the 94-student linked subsample, does the shift from concurrent self-assessment (A1) to postdiction (A2) produce a systematic within-student change in signed error and in absolute error, each tested by a paired comparison? By contrast, how do the rank-order calibration coefficients of the same students compare across the two occasions, where the comparison is descriptive rather than a paired test of a within-student difference?
RQ5. 
Do lower-performing students show greater miscalibration—particularly in the direction of overestimation—than higher-performing students at either judgment occasion, when examined through a tertile split of actual performance?

3.3. Hypotheses

H1. 
The Assessment 1 mean signed error will be positive (overestimation), consistent with overconfidence findings in novices CS students (Cloude et al., 2024; Sobral, 2021; Strickroth, 2024).
H2. 
Absolute calibration error will be smaller on the Assessment 2 postdiction than on the Assessment 1 concurrent judgment, reflecting the additional task experience and, for the laboratory component, the availability of an immediate performance signal (Hacker et al., 2000; Karaca & Bowers, 2023; Osterhage, 2021).
H3. 
Tertile-based analysis will show a negative association between actual performance and signed error (lower performers overestimating more), consistent with the calibration literature (Karaca & Bowers, 2023; Serra & DeMarree, 2016); the pattern is interpreted as raw evidence rather than as confirmation of the dual-burden hypothesis (Gignac & Zajenkowski, 2020; Magnus & Peresetsky, 2022).
H4. 
Rank-order calibration will be stronger for the executable laboratory component than for the written component, because the laboratory task affords clearer self-assessment cues (Gorson & O’Rourke, 2020).

4. Materials and Methods

4.1. Design

The study is a classroom-based observational analysis of grade-judgment data collected as a routine part of two assessments in a first-year CS course. The two assessments differ in judgment type, not in cohort. Assessment 1 is an early-course quiz in which the self-estimate is the final item, made without feedback, an immediate concurrent self-assessment. Assessment 2 is a mid-course exam with a written and a laboratory component; after each component the student records a self-estimate before grades are returned, constituting a postdiction. Because the two assessments also differ in timing, format, and content, the cross-assessment contrast is treated as descriptive rather than as a controlled manipulation.

4.2. Participants and Samples

Participants were first-year students in a CS course at a Central European university covering introductory programming (C), algorithmic thinking, and basic problem solving. Attendance at the course was open to students on several first-year programs. The study population is the computer-science cohort, so the analysis includes only students enrolled in an informatics-bearing program (informatics, informatics teacher training, mathematics–informatics, and engineering informatics); students enrolled in the mathematics program, who sat the same quiz, fall outside the population and are excluded. Section 4.8 states this and the remaining inclusion criteria, with the resulting counts.
The primary cohort is from the 2023–2024 academic year. Two additional independent cohorts from 2024 to 2025 (fall) and from 2025 to 2026 (spring) supply replication data for Assessment 1 only.
Three analytic samples are distinguished. The Assessment 1 sample comprises N = 101 students with valid, normalized self-estimates and actual quiz scores; this is the basis for every Assessment 1 statistic reported below. The Assessment 2 sample comprises N = 117 students with valid self-estimates and actual scores on both exam components; this is the basis for every Assessment 2 statistic. The linked subsample comprises 94 students present in both assessments ( 93.1 % of A1, 80.3 % of A2) and is the only sample used for paired within-subject analyses. The replication cohorts ( N = 100 for 2024–2025; N = 56 for 2025–2026) are independent samples with no possible linkage to the primary cohort and are reported separately. Throughout the manuscript, every reported figure makes the underlying sample explicit, so that statistics for A1, for A2, and for the 94-student linked subsample are never combined.

4.3. Measures

The Assessment 1 self-estimate was elicited by a single Hungarian-language item at the end of the quiz asking the student to record, on a 0–100% scale, the percentage they believed they had solved correctly. Two further items recorded years of prior programming study and a 0–100% self-rated programming ability. Because raw responses were a mix of percentages, proportions (0–1), Hungarian 1–10 grades, and free-text entries, the analysis uses the normalized dataset prepared from the raw file by the rule stated in Section 4.8: proportions are scaled to percentages, unparseable entries are excluded, and the result is a 0–100 score per student. Actual Assessment 1 performance is the quiz percentage score (0–100).
For Assessment 2, students recorded two postdictive estimates on 0–100% scales, one for the written component and one for the laboratory component, after completing each component but before grades were returned. Actual Assessment 2 performance was recorded as raw points on different maxima (30 for the written component, 10 for the laboratory component). For bias analyses, actual scores were rescaled to 0–100 by dividing by the maximum and multiplying by 100. Rank-order analyses (Spearman ρ ) are scale-invariant and are reported on raw scores.
The dimensions used in the analyses reported below are therefore self-estimate (0–100%), actual score (0–100%), signed error (estimate − actual), absolute error, performance tertile (lower/middle/upper, defined within sample by actual score), and a participant indicator for the within-subject linkage. No information that could identify a student appears in the analytic files; identifiers are de-identified internal codes generated for course administration and were used here only to align records across the two assessments. Pseudonymization was a standard part of the assessment process: no student names were ever provided to the analysts. Because the design aligns each student’s records across the two assessments by internal identifier, re-linkage within the dataset is possible by construction, so the records are pseudonymized rather than anonymized.

4.4. Procedure

The early-course quiz was administered through the course learning management system. The self-estimate appeared at the end of the quiz; no score or item-level feedback was visible to the student at the moment of judgment, so the estimate is a concurrent self-assessment in the sense defined above. The mid-course exam took place at the midpoint of the semester and consisted of a written component followed by a laboratory component administered in person. After each component, students recorded a percentage self-estimate; grades were not communicated before the estimate was collected, so each estimate is a postdiction.

4.5. Mid-Course Exam: Structure, Grading, and Information Available to the Students

The mid-course exam was the same for every participant and its structure was published in advance in the course syllabus. Both components draw on a single, common pool of programming theorems and algorithmic patterns (Fekete et al., 2020; Fóthi, 2012; Szlávi & Zsakó, 2016) that had been worked through in class.
The written component was a theory-oriented paper with a maximum of 30 points; a 50 % score (≥15 points) was required to pass. The course grade contribution of the written component is points divided by 3, mapping the raw 0–30 scale onto the 0–10 scale used for the laboratory component. The items tested definitions, hand-tracing of given algorithms, short calculations whose answer depends on tracing an algorithm step by step, and pseudocode-style writing of taught algorithmic patterns on elementary input data. No automatic feedback of any kind was provided during or immediately after the written component.
The laboratory component was a practical programming paper with a maximum of 10 raw points, of which one point was awarded by default to every participant who attended the laboratory session. The 0–1 segment of the laboratory scale is therefore administrative and not a substantive low-performance band (see Section 5.5). Submissions were marked by an automated judge of the DOMjudge family (Eldering et al., 2022): the system compiled and ran each submission against a hidden battery of test cases prepared in advance, awarded full credit for a test that produced the expected output, partial credit for tests that succeeded but failed others, and no credit for tests on which the submission did not produce the expected output. Critically, the automated judge does not evaluate style, efficiency, or whether the algorithm chosen matches the one prescribed for the task; it only checks observable input–output behavior. The score returned by the judge at the moment of submission was therefore the principal in-task cue available to the student when the laboratory postdiction was recorded. The final laboratory grade was set only after a subsequent manual review by the instructor, in which points could be deducted for poor efficiency, weak style, or, for items that explicitly prescribed a specific algorithm or method, for using a different approach. Alternative correct solutions are otherwise accepted. The full rubric, including the criteria used in the manual review, was published in the course syllabus and was therefore common knowledge to every participant before they sat the exam. This combination of (i) an immediate but rubric-incomplete automated signal during the exam, (ii) a delayed manual adjustment against a publicly known rubric, and (iii) a one-point floor on the laboratory scale is what motivates the floor-aware re-banding analysis reported in Section 5.5 and is referenced throughout the results and discussion as “the laboratory cue structure”.

4.6. Ethical Considerations

The estimation activity was embedded in routine course assessment and did not introduce additional experimental manipulations beyond normal instructional activities. No student names or other personally identifiable information were collected for analysis; the course used internal identifiers that allowed assessments to be aligned for the same student but did not identify any individual. Pseudonymization was therefore a standard property of the assessment process rather than a step added for the study; because those identifiers make re-linkage possible by construction, the records are pseudonymized rather than anonymized. Students were informed at the start of the course that their course-assessment records might be used, in pseudonymized form, for research on the teaching of programming, and gave verbal consent to that use. Participation in the estimation items did not affect grades, and the analyses were conducted on the de-identified records only.

4.7. Use of Generative AI Tools

Generative AI assistants (GitHub Copilot Chat, versions available during 2025–2026) were used solely for two purposes. The first was English-language styling: prompts supplied a passage of the authors’ own draft and requested improvements to grammar, concision and register, with the instruction to preserve meaning and terminology. The second was cross-checking numerical statements against the underlying analysis outputs: prompts supplied a numerical statement from the manuscript together with the corresponding analysis output and asked whether the two agreed, so that transcription errors could be detected. No scientific content, results, figures, tables, citations, interpretations, or text generation was produced by AI, and no analysis was carried out by an AI tool. All suggestions were reviewed and edited by the authors, who take full responsibility for the manuscript’s content. No AI tool is listed as an author.

4.8. Data Analysis

Analyses were conducted in Python 3.13.3 using pandas 2.2.3, numpy 2.2.6 and scipy 1.16.3; the comparison of dependent correlations was additionally cross-checked in R 4.6.1 with the cocor package, version 1.1-4 (Diedenhofen & Musch, 2015).
  • Normalization of the Assessment 1 self-estimate.
The item asked students what percentage of the quiz they believed they had solved correctly, and offered a worked example in both of the notations in use locally: a percentage, or the equivalent mark on the 1–10 grading scale (85% being 8.5 on that scale). Students used both formats. Responses were normalized to a 0–100 scale by the following rule: where a response stated a range, its midpoint was taken; where a percentage sign was present, the value was treated as already a percentage; otherwise a bare value of 1 or less was treated as a proportion and multiplied by 100, and a bare value of 10 or less as a 1–10 grade and multiplied by 10; a value falling outside 0–100 after this step, and a response containing no numeral, were treated as missing. Applied to the raw export, the rule reproduces the independently hand-normalized value for every response that carries a value.
  • Inclusion criteria and exclusions.
Three criteria were applied in order. First, the student had to be enrolled in an informatics-bearing program, as set out in Section 4.2. Second, the response had to yield a self-estimate on the 0–100 scale under the rule above and to carry a recorded quiz score. Third, where a student sat the quiz more than once, the record with the higher quiz score among those meeting the second criterion was retained, so that each student contributes one observation. The raw Assessment 1 export contains 119 responses from 112 students, 7 of whom submitted twice. Six responses were excluded by the first criterion (five students enrolled in the mathematics program and one with no program recorded), seven by the second (three of these also recorded a zero quiz score), and five by the third. The analysis therefore uses 101 responses from 101 students. The same three criteria were applied to the replication cohorts.
  • Missing data.
Missing data were handled by complete-case analysis; no imputation was performed, and a blank self-estimate was treated as missing rather than as a score of zero.
  • Calibration measures.
For each student and each judgment, calibration was characterized by the signed error (estimate − actual, with positive values indicating overestimation), the absolute error, and the rank-order association between estimate and actual score using Pearson r and Spearman ρ . The proportion of overestimators, exact matches, and underestimators, together with the root mean squared error (RMSE), provided complementary summaries. Assessment 2 actual scores were rescaled to 0–100 against the confirmed maxima of 30 (written) and 10 (laboratory), and replication-cohort scores against the 23 scored quiz items; both are direct arithmetic transformations rather than modeling choices.
  • Within-subject comparisons.
For the within-subject comparison in the 94-student linked subsample, paired t-tests and Wilcoxon signed-rank tests (Wilcoxon, 1945) compared signed and absolute errors between A1 and each A2 component, with Cohen’s d z (the mean difference divided by the standard deviation of the differences) as the standardized effect size. Both tests are reported throughout because the normality assumption holds for one contrast and is mildly violated for the other: Shapiro–Wilk gives W = 0.979 , p = 0.131 for the A1 − A2 written difference and W = 0.971 , p = 0.034 for A1 − A2 laboratory. The two procedures agree on every contrast reported.
  • Comparison of dependent correlations.
The laboratory and written correlations are computed on four distinct variables measured on the same students, so they are dependent and non-overlapping, and Fisher’s z for independent samples does not apply. The analytic procedures for this design (Dunn & Clark, 1969; Pearson & Filon, 1898; Raghunathan et al., 1996) assume multivariate normality, which these variables do not satisfy: all four depart from normality (Shapiro–Wilk p between 0.0005 and 0.013 ) and the laboratory score is a 0–10 integer scale taking eleven distinct values. The comparison therefore uses a nonparametric bootstrap over students (50,000 resamples) (Efron & Tibshirani, 1993), reporting the percentile confidence interval for the difference between correlations and a two-sided bootstrap p-value; a permutation test exchanging the two component pairs within student is reported as a check, and the analytic values are given alongside. Bootstrap and jackknife standard errors agree with each other and exceed the normal-theory value, which is why the resampling result is the one reported.
  • Multiple comparisons.
RQ1–RQ4 constitute the confirmatory family and are reported without adjustment; every analysis from the symmetric-tertile split onward, including the floor-aware banding, its sensitivity contrasts and the continuous alternative, is exploratory and is reported as such, without adjustment and without confirmatory claims. Fifteen inferential tests are reported in total. A Holm–Bonferroni adjustment within the confirmatory family leaves every confirmatory conclusion unchanged, the largest confirmatory p value being 5.5 × 10 8 .
  • Tertile and banding analyses.
For the tertile analyses reported in Section 5.4, students were split into lower, middle, and upper thirds of actual performance within each assessment, and mean signed and absolute error were compared across tertiles. To avoid placing weight on arbitrary cut points (Gignac & Zajenkowski, 2020), the Spearman correlation between actual performance and signed error is reported alongside the tertile summary, so that the continuous pattern can be inspected directly. The floor-aware banding of the laboratory postdiction reported in Section 5.5 groups students by raw laboratory score into an administrative floor band (0–1 points), a low band (2–3), and a high band (7–10). The floor band is set aside because one laboratory point was awarded to every student who attended, so that segment of the scale cannot index low-performance metacognition. The cut points were chosen after the laboratory distribution had been inspected; the analysis is therefore exploratory throughout, and its sensitivity to the choice of cut points is reported alongside a continuous alternative.
  • Sensitivity checks
Sensitivity checks included exclusion of extreme outliers (signed errors exceeding ± 3 SD) and a descriptive comparison of linked and unlinked students. Because participation in each assessment depended on attendance, students simply did or did not attend the quiz and the exam; the linked/unlinked split reflects routine attendance variation rather than systematic selection, and is reported as context rather than as a bias correction.

5. Results

The results are based on verified computations on the source files. Assessment 2 actual scores are rescaled to 0–100 using the confirmed maxima of 30 (written) and 10 (laboratory). The three samples are reported in turn, followed by the within-subject linked-subsample comparison and the tertile-based view.

5.1. Assessment 1: Concurrent Self-Assessment ( N = 101 )

In the N = 101 sample with valid normalized Assessment 1 records, students estimated their quiz performance at a mean of 46.13 % ( S D = 20.01 ), while their actual percentage score averaged 37.04 % ( S D = 13.24 ). The mean signed error was therefore + 9.10 percentage points (Table 1), with a substantial spread ( S D = 16.63 , range 33.82 to + 49.17 ). Of the 101 students, 71 ( 70.3 % ) overestimated their performance and 30 ( 29.7 % ) underestimated; no exact matches occurred (Table 2). Rank-order calibration was moderate, with Pearson r = 0.565 and Spearman ρ = 0.533 (both p < 0.001 ), indicating that students had some awareness of their relative standing but mapped that awareness imprecisely to the percentage scale (see Figure 1a). The picture from the concurrent quiz judgment is thus one of systematic optimism: a clear majority of first-year students overestimated their early-course quiz performance by roughly nine percentage points, with the estimate more dispersed than the actual score.

5.2. Assessment 2: Postdiction ( N = 117 )

The N = 117 sample for Assessment 2 yielded a markedly different picture from the concurrent quiz judgment (Table 3). On the written component, the mean self-estimate was 48.39 % ( S D = 19.50 ) against a rescaled actual mean of 58.47 % ( S D = 20.13 ); the mean signed error was 10.08 pp, with 85 of 117 students ( 72.6 % ) classified as underestimators. On the laboratory component, the mean self-estimate ( 52.29 % , S D = 27.82 ) was closer to the rescaled actual mean ( 56.50 % , S D = 27.71 ), yielding a smaller bias of 4.21 pp. The laboratory component was also notable for 28 exact matches ( 23.9 % of the sample), most plausibly reflecting the laboratory cue structure described in Section 4.5: at submission time, the automated judge returned an immediate, observable test-case score that students could carry over almost unchanged into the postdiction. Rank-order calibration was moderate for the written component ( ρ = 0.633 ) and strong for the laboratory component ( ρ = 0.804 ); both correlations are highly significant ( p < 0.001 ).
The two correlations are computed on four distinct variables measured on the same students and are therefore dependent and non-overlapping, so the difference between them requires a procedure for dependent correlations rather than Fisher’s z for independent samples. A nonparametric bootstrap over students (50,000 resamples) gives a difference in Pearson correlations of 0.21 , 95% CI [ 0.040 , 0.368 ] , p = 0.014 , with a permutation test exchanging the two component pairs within students giving p = 0.012 ; on Spearman correlations the difference is 0.17 , 95% CI [ 0.026 , 0.313 ] , p = 0.022 . Analytic normal-theory procedures for dependent non-overlapping correlations give substantially smaller p values ( 0.0001 to 0.0004 ), but all four variables depart from normality and both bootstrap and jackknife standard errors exceed the normal-theory value, so the resampling result is the one reported. Rank-order calibration is therefore reliably stronger for the laboratory component than for the written component.
The reversal of bias direction from the concurrent quiz judgment to both postdiction components is the central descriptive result of the study (Figure 1b,c and Figure 2).

5.3. Within-Subject Comparison in the 94-Student Linked Subsample

The two assessments differ in judgment type, format, content, and timing, so any cross-assessment comparison is exploratory. The cleanest available view is the within-subject contrast in the 94 students who provided valid estimates on both occasions (Table 4). In this subsample, the mean signed error was + 8.99 pp on the Assessment 1 concurrent quiz judgment, 11.06 pp on the Assessment 2 written postdiction, and 4.70 pp on the Assessment 2 laboratory postdiction. The within-student difference between A1 and the written postdiction averaged + 20.05 pp, with paired t ( 93 ) = 9.73 , p < 0.001 , Wilcoxon p < 0.001 , and Cohen’s d z = 1.00 . The within-student difference between A1 and the laboratory postdiction averaged + 13.69 pp, with paired t ( 93 ) = 6.21 , p < 0.001 and Cohen’s d z = 0.64 . The same students therefore shifted from systematic overestimation on the concurrent quiz judgment to systematic underestimation on each postdictive component, and the shift is large in standardized terms. This is the most informative summary the data afford on whether judgment type matters, since within-subject comparison removes between-student variation and confirms that the bias reversal is not an artifact of who happened to appear in each sample (Figure 2).
The pattern for absolute error differs from that for bias. In the linked subsample, the mean absolute error was 14.95 percentage points on the Assessment 1 concurrent judgment, 15.93 on the written postdiction, and 11.72 on the laboratory postdiction. The written comparison shows no detectable change ( Δ = 0.98 , 95% CI [ 4.49 , 2.54 ] , t ( 93 ) = 0.55 , p = 0.582 , d z = 0.06 ); the laboratory comparison is at the margin ( Δ = + 3.23 , 95% CI [ 0.33 , 6.78 ] , t ( 93 ) = 1.80 , p = 0.075 ; Wilcoxon p = 0.017 ). Rank-order calibration within the same students was ρ = 0.532 on Assessment 1, 0.657 on the written postdiction, and 0.805 on the laboratory postdiction.
Taken together with the sign-error reversal, this indicates that the same students did not become better calibrated between the two occasions so much as differently miscalibrated: the direction of the error reversed while its magnitude, at least for the written component, did not change. Only the laboratory component, where an immediate automated signal was available at the moment of judgment, shows weak evidence of improved absolute accuracy alongside markedly better rank-order correspondence.

5.4. Raw Dunning–Kruger View: Tertile Split by Actual Performance

A tertile-based “raw” view of the lower-vs-upper pattern is presented in Table 5. On the Assessment 1 concurrent quiz judgment, all three tertiles overestimated their performance by a similar amount (lower tertile + 9.79 pp, middle + 11.21 pp, upper + 6.06 pp), and the continuous Spearman association between actual score and signed error was essentially zero ( ρ = 0.095 , p = 0.34 ). The canonical dual-burden pattern, in which lower performers overestimate while upper performers underestimate, is therefore not present in this concurrent quiz judgment; what is present is broad-based optimism. The Assessment 2 postdictions tell a different story. On the written component, the bias becomes progressively more negative across tertiles (lower 1.64 pp, middle 10.34 pp, upper 18.68 pp), with ρ ( actual , signed err ) = 0.465 , p < 0.001 . The laboratory component shows the same monotonic gradient (lower + 2.57 pp, middle 5.98 pp, upper 11.61 pp), with ρ = 0.325 , p < 0.001 . The tertile pattern visible on the postdictions thus involves not lower-tertile overestimation but progressively larger underestimation by higher performers. This is the pattern most commonly produced by bounded scales and regression to the mean (Gignac & Zajenkowski, 2020; Magnus & Peresetsky, 2022; McIntosh et al., 2019). Under the equally sized actual-score tertile split alone, the present results are therefore not a positive test of the Dunning–Kruger dual-burden hypothesis; Section 5.5 reports a complementary, floor-aware re-banding that yields a Dunning–Kruger-compatible pattern in the laboratory postdiction.

5.5. Floor-Aware Asymmetric Banding of the Postdictions

A refinement of the symmetric-tertile split is motivated by the structure of the laboratory grading scheme itself. Because one point was awarded by default to every participant of the laboratory exam, the 0–1 segment of the 0–10 raw laboratory scale is administrative rather than substantive: no student who attempted anything could fall below it, and it therefore cannot index meaningful low-performance metacognition. Treating the 0–1 interval as a floor band and defining the first substantive low-performance band as 2–3 on the raw laboratory scale (the “floor-aware asymmetric banding”), the laboratory data yield a Dunning–Kruger-compatible pattern.
The cut points were chosen after the laboratory distribution had been inspected, and this analysis is reported as exploratory throughout. Its motivation is a documented feature of the grading scheme rather than of the data, as set out above and defined in Section 4.8, but it was not specified in advance. The raw laboratory score distribution across the 117 students was 3, 7, 5, 12, 19, 12, 9, 17, 9, 13 and 11 students at scores 0 to 10, respectively.
Students whose actual laboratory score fell in the 2–3 band ( n = 17 ) overestimated their performance by + 0.78 raw points on average, with a 95% confidence interval of [ 0.07 , 1.49 ] and p = 0.033 against zero, (approximately + 7.8 pp on the 0–100 scale used in Table 5). Students in the 7–10 band ( n = 50 ) underestimated by 1.02 raw points (approximately 10.2 pp), 95% CI [ 1.539 , 0.497 ] , p < 0.001 ; the administrative floor band (0–1, n = 10 ) underestimated by 0.450 , 95% CI [ 0.878 , 0.022 ] . The between-band gap of about 1.80 raw points (∼18 pp on the 0–100 scale) has a 95% confidence interval of [ 0.942 , 2.658 ] , Welch t ( 36.6 ) = 4.253 , p = 0.00014 , Cohen’s d = 1.04 ; a Mann–Whitney test agrees, U = 698 , p = 6.7 × 10 5 .
Because the banding was chosen after the data were seen, its sensitivity to the choice of cut points is reported in full (Table 6), together with a continuous analysis that avoids cut points altogether: the Spearman correlation between actual laboratory score and signed error is 0.401 ( p = 1.9 × 10 5 ) with the administrative floor excluded ( n = 107 ) and 0.325 with it included ( n = 117 ); an ordinary least-squares fit gives a slope of 0.282 error-points per score-point ( p = 7.4 × 10 5 ).
This re-banding result is substantively interpretable because the laboratory grading rubric was common knowledge to the students: the automated judging system compared submissions against reference outputs, and the human grader could further deduct points for poor style, weak efficiency, or—where a specific algorithm or method was explicitly prescribed for an item—use of a different approach (alternative correct solutions are otherwise accepted). Because the grading criteria were published in the course syllabus, a simple lack of access to them is an implausible explanation for the overestimation observed in the low-but-non-floor band. This does not establish that students understood, internalized, or correctly applied those criteria when judging their own work: availability and comprehension are different things, and differential understanding or differential use of a published rubric remains a viable explanation that these data cannot rule out. It is consistent with a calibration or metacognitive failure of the kind originally described by Kruger and Dunning (Dunning, 2011; Dunning et al., 2003; Ehrlinger et al., 2008; Kruger & Dunning, 1999): partially working solutions may have been judged by their authors as more adequate than they actually were under the known professional requirements.
For the written postdiction, the evidence points in the same direction but is weaker. Using pedagogically motivated bands of 0–3, 3–5, 5–7, and 7–10 on the 0–10 scaled actual score, each interval closed below and open above and the last closed at both ends, the lowest band overestimated and the highest band underestimated, qualitatively mirroring the laboratory finding, although the written-component result is more sensitive to the exact band cut-offs than the laboratory result.
The corresponding mean signed errors are + 10.99 pp ( n = 11 ), 6.10 ( n = 27 ), 10.62 ( n = 43 ) and 18.86 ( n = 36 ). The written component has no administrative floor to set aside, and its lowest band contains only eleven students, so this result is reported as descriptive only.
The result carries two qualifications. First, the symmetric-tertile result in Section 5.4 stands: the postdiction tertile gradient is dominated by progressive underestimation among the upper tertile, the pattern produced by bounded scales and regression to the mean (Gignac & Zajenkowski, 2020; Jansen et al., 2021; Lebuda et al., 2024; Magnus & Peresetsky, 2022; McIntosh et al., 2019), so the floor-aware re-banding is a complementary descriptive view, not a replacement. Second, grouping participants by estimated rather than actual scores flips the sign of the calibration error, as regression-to-the-mean predicts (Magnus & Peresetsky, 2022; McIntosh et al., 2019).
The two components of this result differ in robustness. The between-band difference is robust: it remains significant under every alternative banding examined (Table 6), with p between 0.00004 and 0.00058 and differences between 1.18 and 1.94 raw points. The component that specifically distinguishes a Dunning–Kruger-compatible pattern from ordinary regression to the mean, namely the low band overestimating significantly above zero, is not equally robust: it survives for the narrow 2–3 and 2–4 bands, and its confidence interval includes zero as soon as the administrative floor is folded in, or the band is widened. The exclusion of the 0–1 segment is justified by the grading scheme, but the cut points were chosen after the data had been inspected. The finding is therefore reported as exploratory, and as compatible with rather than evidential for the dual-burden account.

5.6. Cross-Cohort Replication of the Assessment 1 Pattern

The Assessment 1 concurrent quiz judgment was administered, with the same instrument and procedure, in two further independent cohorts. In the 2024–2025 cohort ( N = 100 with valid normalized estimates), students overestimated by a mean of + 5.40 pp, with 62.0 % overestimators and Spearman ρ = 0.390 . In the 2025–2026 cohort ( N = 56 ), students overestimated by a mean + 7.08 pp, with 69.6 % overestimators and ρ = 0.434 . All three correlations are statistically significant. The replication cohorts were filtered by the same three inclusion criteria as the primary cohort (Section 4.8). Across three consecutive academic years, then, the concurrent quiz judgment is consistently optimistic by approximately five to nine percentage points (Figure 3 and Table 7). The replication across three cohorts provides reassurance that the Assessment 1 finding is a robust feature of the early-course concurrent self-assessment in this CS1 setting and is not an artifact of the 2023–24 cohort.

5.7. Additional Context

Assessment 1 also recorded years of prior programming study ( M = 3.70 , S D = 1.07 , n = 102 ) and a 0–100% self-rated programming ability ( M = 51.77 , S D = 24.06 ). Two records reported an impossible 1 years and are excluded from the first figure only; the variable is descriptive and enters no analysis. These variables may serve as covariates in follow-up analyses of whether prior experience moderates the relationship between self-estimate and actual performance; they are not the focus of the present manuscript.

5.8. Hypothesis Outcomes

The four hypotheses can now be resolved. H1 is supported: the Assessment 1 mean signed error is positive. H2 is rejected for the written component ( p = 0.582 ) and only weakly supported for the laboratory component ( p = 0.075 ; Wilcoxon p = 0.017 ), so absolute calibration did not improve between the two occasions in the way that additional task experience would predict. H3 is not supported on the concurrent judgment ( ρ = 0.10 , p = 0.34 ) and is supported on both postdiction components ( ρ = 0.47 and 0.33 , both p < 0.001 ), with the caveat that bounded scales and regression to the mean would produce that gradient in any case. H4 is supported: rank-order calibration is stronger for the laboratory component, the difference in Pearson correlations being 0.21 , 95% CI [ 0.04 , 0.37 ] , p = 0.014 .

6. Discussion

6.1. Principal Findings

Three findings stand out. First, the Assessment 1 concurrent quiz judgment was characterized by broad-based optimism: most students overestimated by roughly eight percentage points, with no strong evidence that lower performers overestimated more than higher performers. This overconfidence may reflect a mismatch between students’ prior experience in secondary education and the demands of higher education. Contemporary secondary-school curricula increasingly include programming, algorithms, and computational thinking as explicit learning goals (European Commission/EACEA/Eurydice, 2022; Government of Hungary, 2020; K–12 Computer Science Framework Steering Committee, 2016; Lee & Lee, 2021). However, curricular exposure to programming-oriented content does not necessarily imply calibrated mastery of formal algorithmic reasoning. Students may enter higher education with confidence shaped by school-level tasks and by a broader sense of digital familiarity, while university assessment requires more precise tracing, stricter control-flow reasoning, and reliable manipulation of variable states. This interpretation is also consistent with critiques of the assumption that familiarity with digital tools automatically transfers to genuine digital or computational problem solving (Csernoch, 2017, 2025; Kirschner & De Bruyckere, 2017; Prensky, 2001). Second, the Assessment 2 postdictions reversed the bias direction, with a substantial mean underestimation on the written component and a smaller one on the laboratory component. Rank-order calibration was strongest for the executable laboratory task. This reversal may indicate a recalibration effect: after initial exposure to university-level expectations, students may have become more aware of the precision required in formal algorithmic reasoning. The recalibration reading is available for the direction of the error and for rank-order correspondence, but not for absolute accuracy, which did not improve for the written component. In other words, they may have learned not only some of the required content, but also the limits of their own competence. Such awareness is pedagogically important, because recognizing what one does not yet know is a prerequisite for more accurate self-monitoring and further learning. This can be read as a shift from overconfidence based on prior school-level experience toward more realistic monitoring of schema availability and execution control (Csernoch, 2017, 2025; Kahneman, 2011; Sweller et al., 2011). Third, within the 94-student linked subsample, the bias shift between A1 and each A2 component is large in standardized terms ( d z = 1.00 for the written component) and is not explained by between-student differences. Absolute accuracy, by contrast, did not improve for the written component, so the shift is a change in the direction of bias rather than a gain in accuracy. The three-cohort replication makes the Assessment 1 optimism finding broadly relevant rather than cohort-specific, while the within-subject linked-subsample analysis documents the A1→A2 bias reversal directly within the same students, so the A1 and A2 conclusions rest on complementary forms of evidence.

6.2. Relation to Prior Literature

The Assessment 1 overestimation is consistent in direction and magnitude with calibration findings in other classroom contexts (Bol & Hacker, 2012; Karaca & Bowers, 2023; Osterhage, 2021; Serra & DeMarree, 2016) and with CS-specific evidence that early-undergraduate students, despite prior secondary-school programming, over-rate the correctness of their programming solutions (Cloude et al., 2024; Sobral, 2021; Strickroth, 2024). Within Central European CS education, the result also echoes earlier Hungarian work documenting overconfidence in spreadsheet and programming self-assessment (Csernoch et al., 2015; Kun, 2016; Máté & Darabos, 2017; Máté et al., 2025; Nagy et al., 2021).
The literature expects postdictions to be more accurate than predictions because experiential cues become available (Bol & Hacker, 2012; Hacker et al., 2000). The present results show a more specific pattern: rank-order calibration does improve from concurrent to postdiction, particularly for the laboratory component, but in the case of the written component, the price of this improvement is a substantial mean underestimation. One reading is that the experience of an unfamiliar written CS exam can suppress confidence beyond what is warranted, in line with anxiety and exam-difficulty interpretations discussed in earlier postdiction work (Hartwig & Dunlosky, 2014; McGuire, 2023; Nietfeld et al., 2005). Another reading is that students lacked diagnostic cues for the written component and fell back on conservative defaults. The stronger calibration of the laboratory postdiction is consistent with claims that more concrete, executable tasks afford clearer self-assessment cues (Chen et al., 2024; Gorson & O’Rourke, 2020), and is specifically supported here by the automated DOMjudge-style signal returned at submission time (see Section 4.5), which provided a concrete, if rubric-incomplete, in-task cue that the written component lacked.
The tertile pattern reported in Section 5.4 deserves cautious interpretation. The classic dual-burden pattern is absent on the concurrent quiz judgment, and the symmetric-tertile view of the postdictions is dominated by progressively larger underestimation among high performers, exactly what bounded scales and regression to the mean would produce, even in the absence of any metacognitive deficit (Gignac & Zajenkowski, 2020; Jansen et al., 2021; Lebuda et al., 2024; Magnus & Peresetsky, 2022; McIntosh et al., 2019). The floor-aware re-banding reported in Section 5.5, however, yields a Dunning–Kruger-compatible pattern on the laboratory postdiction. Once the administrative 0–1 segment of the laboratory scale is set aside, students in the 2–3 band overestimate by about + 0.78 raw points and students in the 7–10 band underestimate by about 1.02 raw points, a between-band gap of roughly 1.80 raw points (about 18 pp on the 0–100 scale) with a 95% confidence interval of [ 0.94 , 2.66 ] and p = 0.0001 .
One qualification applies to that result. The continuous negative gradient on which the floor-aware analysis rests is the same gradient attributed to regression to the mean in the symmetric-tertile section. What distinguishes a Dunning–Kruger-compatible reading is specifically the low band crossing zero into overestimation, and that is the part the sensitivity analysis shows to be fragile: it survives only for the narrow 2–3 and 2–4 bands. The pattern is therefore compatible with the dual-burden account rather than evidence for it.
The written and combined postdictions show weaker but directionally consistent gradients. Because the laboratory grading rubric (automated tests plus deductions for style, efficiency, and use of the taught algorithmic patterns) was common knowledge to the students (Csernoch et al., 2015; Gorson & O’Rourke, 2020), the low-band overestimation is unlikely to be explained by simple lack of access to the criteria, although availability does not establish that the criteria were understood or correctly applied; it is consistent with the metacognitive calibration deficit originally described by Kruger and Dunning (Dunning et al., 2003; Ehrlinger et al., 2008; Kruger & Dunning, 1999). The present data are therefore best described in calibration terms overall, with the laboratory postdiction additionally providing descriptive support for a Dunning–Kruger-compatible pattern under a floor-aware banding.

6.3. The Concurrent Self-Assessment as a Distinct Judgment Type

A methodological contribution of the study is the explicit separation of the concurrent self-assessment from the postdiction. The concurrent judgment differs from a prediction in that the student has already attempted the items and possesses experiential information about difficulty, fluency, and uncertainty, and it differs from a standard postdiction in that no time for reflection, no opportunity to consult materials or peers, and no shift out of the assessment mindset has intervened. The in-the-moment character of the concurrent judgment may help to explain why overestimation was so broadly distributed: students may have used the fluency of having just answered the questions as a proxy for correctness, a well-documented metacognitive heuristic (de Bruin et al., 2017; Hartwig & Dunlosky, 2014; Moore & Healy, 2008). The finding invites further work on how the timing and embedding of self-assessment items within assessments shape the resulting judgment.

6.4. Possible Explanations

Several non-exclusive explanations are compatible with the data. Fluency-based overestimation can plausibly account for the broad optimism of the concurrent quiz judgment (de Bruin et al., 2017). Anchoring on preconceptions from secondary school is consistent with the dispersion of A1 estimates relative to actual scores (Friedman, 2023; Máté et al., 2025), and wish-based judgment may have contributed in either direction (Serra & DeMarree, 2016). The stronger laboratory postdiction is most parsimoniously read in terms of the laboratory cue structure of Section 4.5: the immediate DOMjudge-style test-case score (Eldering et al., 2022) furnishes a concrete in-task cue that the written paper lacks, while remaining rubric-incomplete; this asymmetry, combined with the publicly available full rubric, is what allows the low-band overestimation reported in Section 5.5 to be read as consistent with a metacognitive calibration deficit rather than as ignorance of the criteria (Dunning, 2011; Dunning et al., 2003; Ehrlinger et al., 2008; Kruger & Dunning, 1999). Finally, by the time of the mid-course exam, students had several weeks of university-level exposure to the intended algorithmic patterns (Fekete et al., 2020; Fóthi, 2012; Szlávi & Zsakó, 2016) and the corresponding professional requirements, so their postdictive estimates were made against a better-defined target than their pre-task expectations had been at the start of the course (Csernoch, 2017, 2025).

6.5. Pedagogical Implications

The study is observational, but its descriptive results support several practical recommendations. Brief calibration prompts can be embedded within assessments so that self-monitoring is practiced as a regular activity (Angell et al., 2024; Callender et al., 2016; MacNeil et al., 2024). Returning estimate-versus-actual feedback after each assessment can help students update their self-models (Janssen & Lazonder, 2024). In CS1 specifically, the contrast between the better-calibrated laboratory postdiction and the more biased written postdiction (which lacks any in-task cue analogous to the laboratory automated judge) suggests that written algorithmic components benefit from explicit self-assessment support, such as published rubrics, annotated worked examples, or examples of typical errors (Chen et al., 2024; Han et al., 2021).

6.6. Limitations

The cross-assessment comparison is confounded by simultaneous differences in judgment type, timing, format, and content; the within-subject linked-subsample analysis mitigates but does not eliminate this. The primary cross-component analysis uses one cohort at one institution; the Assessment 1 finding replicates across two further cohorts, but the postdiction comparison does not. Not all students appear in both assessments because attendance at each assessment was free; some students simply did not attend the quiz, the exam, or both, so the 94-student linked subsample reflects who attended both occasions rather than systematic selection. The Assessment 1 estimate is a single percentage judgment, which is subject to anchoring and rounding, and the original responses required normalization that may have introduced minor measurement error.
Three Assessment 1 responses that had been recorded as a self-estimate of zero were in fact left blank in the raw export; they are treated as missing here, in line with the complete-case policy stated in Section 4.8. The floor-aware banding is exploratory, and its cut points were chosen after the data were inspected. Only bias and rank-order calibration, not absolute accuracy, differ in the direction that additional task experience would predict.
Finally, the study tests no intervention.

6.7. Future Work

Future work should collect concurrent and postdictive estimates for the same assessment in order to isolate the effect of judgment type from the effect of assessment differences, extend the design across further cohorts and institutions, and compare calibration before and after explicit metacognitive instruction or estimate-versus-actual feedback (Janssen & Lazonder, 2024; MacNeil et al., 2024). It would also be informative to follow students longitudinally to test whether early miscalibration predicts later academic outcomes or persistence (Schaefer et al., 2023), and to compare calibration across task formats (programming versus conceptual) within a single course (Csapó et al., 2020; Csernoch et al., 2015).

7. Conclusions

This study examined the calibration of grade judgments among first-year CS students across two assessment contexts that differ in judgment type—a concurrent self-assessment embedded in an early-course quiz and a postdiction made after a mid-course written-and-laboratory exam. On the concurrent quiz judgment, students were broadly optimistic (mean signed error + 9.10 pp; 70.3 % overestimators) with moderate rank-order calibration (Spearman ρ = 0.533 ). On the mid-course postdictions, the bias reversed: students underestimated their written exam performance by an average of 10.08 pp and their laboratory performance by 4.21 pp, with rank-order calibration markedly stronger for the executable laboratory task ( ρ = 0.804 ). Within the 94-student linked subsample present in both assessments, the bias reversal is large in standardized terms (Cohen’s d z = 1.00 for the written component) and is not explained by between-student differences. Absolute accuracy, however, did not improve for the written component ( p = 0.582 ). The tertile-based “raw Dunning–Kruger” view shows essentially no lower-performer overestimation gradient on the concurrent quiz judgment and a gradient on the postdictions that is dominated by growing underestimation among higher performers, a pattern more compatible with regression to the mean than with the dual-burden hypothesis. A complementary floor-aware re-banding of the laboratory postdiction, which sets aside the administrative 0–1 segment of the 0–10 raw scale (one point was awarded by default to every participant), yields a Dunning–Kruger-compatible pattern on the laboratory component: the 2–3 band overestimates by + 0.78 raw points and the 7–10 band underestimates by 1.02 raw points, a between-band gap of about 1.80 raw points (95% CI [ 0.94 , 2.66 ] , p = 0.0001 ) that is robust to alternative bandings, although the low band’s overestimation above zero is not, and with a descriptive counterpart on the written component. Because the laboratory rubric (automated DOMjudge-style test cases (Eldering et al., 2022) plus manual deductions) was published in the syllabus, low-band overestimation is unlikely to be explained by simple ignorance of the criteria and is consistent with the metacognitive calibration deficit described by Kruger and Dunning (Dunning, 2011; Dunning et al., 2003; Ehrlinger et al., 2008; Kruger & Dunning, 1999), although availability of a rubric does not establish that it was understood or correctly applied. The Assessment 1 overestimation pattern replicated in two further independent cohorts.
The strongest conclusion these data support is that the same students show substantially different bias patterns across two assessment contexts, including a reversal from mean overestimation on the concurrent quiz judgment to mean underestimation on both postdictive components, while absolute accuracy did not improve for the written component. The study cannot establish whether this pattern is attributable primarily to judgment timing, judgment type, assessment content, task transparency, accumulated course experience, or interactions among them, because the two assessments differ in all of these at once. Contrasting concurrent self-assessments with postdictions for the same assessment remains the design needed to isolate the effect of judgment type.

Author Contributions

Conceptualization, G.V. and M.C.; methodology, G.V. and M.C.; formal analysis, G.V.; investigation, G.V.; data curation, G.V.; writing—original draft preparation, G.V.; writing—review and editing, G.V. and M.C.; visualization, G.V.; supervision, M.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

This study analyzed only pseudonymized records generated by routine course assessment, with no additional experimental manipulation of the participants, and was conducted without a separate ethics review.

Informed Consent Statement

Verbal informed consent for the use of pseudonymized course-assessment records for research purposes was obtained from all participating students at the start of the course. Pseudonymization was a standard property of the assessment process; no student names were ever provided to the analysts, and records were linkable only by internal course identifier.

Data Availability Statement

The pseudonymized dataset and analysis code are available from the corresponding author upon reasonable request. They are not publicly available due to institutional privacy considerations.

Acknowledgments

The authors thank the participating students for their engagement with the course assessments.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Angell, J. B., Wilton, M., Dahlberg, C. L., & Crowe, A. J. (2024). Metacognitive exam preparation assignments in an introductory biology course improve exam scores for lower ACT students compared with assignments that focus on terms. CBE—Life Sciences Education, 23(1), ar6. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Bol, L., & Hacker, D. J. (2012). Calibration research: Where do we go from here? Frontiers in Psychology, 3, 229. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Callender, A. A., Franco-Watkins, A. M., & Roberts, A. S. (2016). Improving metacognition in the classroom through instruction, training, and feedback. Metacognition and Learning, 11(2), 215–235. [Google Scholar] [CrossRef] [Scilit]
  4. Chen, Y., Xie, B., & Li, Y. (2024). Understanding the reasoning behind students’ self-assessments of ability in introductory computer science courses. In Proceedings of the 2024 ACM conference on international computing education research. ACM. [Google Scholar] [CrossRef] [Scilit]
  5. Cloude, E. B., Wortha, F., Azevedo, R., & Winne, P. H. (2024). Novice programmers inaccurately monitor the quality of their work and their peers’ work in an introductory computer science course. In Proceedings of the 14th learning analytics and knowledge conference. ACM. [Google Scholar] [CrossRef] [Scilit]
  6. Coutinho, M. V. C., Thomas, J., Alsuwaidi, A. S., & Couchman, J. J. (2024). Unskilled and unaware: Second-order judgments increase with miscalibration for low performers. Frontiers in Psychology, 15, 1252520. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Csapó, G., Csernoch, M., & Abari, K. (2020). Sprego: Case study on the effectiveness of teaching spreadsheet management with schema construction. Education and Information Technologies, 25, 1585–1605. [Google Scholar] [CrossRef] [Scilit]
  8. Csapó, G., Sebestyén, K., & Csernoch, M. (2021). Case study: Developing long-term knowledge with Sprego. Education and Information Technologies, 26, 8033–8055. [Google Scholar] [CrossRef] [Scilit]
  9. Csernoch, M. (2017). Thinking fast and slow in computer problem solving. Journal of Software Engineering and Applications, 10(1), 11–40. [Google Scholar] [CrossRef]
  10. Csernoch, M. (2025). Lean digital education to resolve the paradox of the illusion of digital prosperity. Journal of Innovation and Knowledge, 10(2), 100676. [Google Scholar] [CrossRef] [Scilit]
  11. Csernoch, M., Biró, P., Máth, J., & Abari, K. (2015). Testing algorithmic skills in traditional and non-traditional programming environments. Informatics in Education, 14(2), 175–197. [Google Scholar] [CrossRef] [Scilit]
  12. de Bruin, A. B. H., Kok, E. M., Lobbestael, J., & de Grip, A. (2017). The impact of an online tool for monitoring and regulating learning at university: Overconfidence, learning strategy, and personality. Metacognition and Learning, 12, 21–43. [Google Scholar] [CrossRef] [Scilit]
  13. Diedenhofen, B., & Musch, J. (2015). cocor: A comprehensive solution for the statistical comparison of correlations. PLoS ONE, 10(4), e0121945. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Dunlosky, J., & Rawson, K. A. (2012). Overconfidence produces underachievement: Inaccurate self-evaluations undermine students’ learning and retention. Learning and Instruction, 22(4), 271–280. [Google Scholar] [CrossRef] [Scilit]
  15. Dunn, O. J., & Clark, V. (1969). Correlation coefficients measured on the same individuals. Journal of the American Statistical Association, 64(325), 366–377. [Google Scholar] [CrossRef]
  16. Dunning, D. (2011). The Dunning–Kruger effect: On being ignorant of one’s own ignorance. In J. M. Olson, & M. P. Zanna (Eds.), Advances in experimental social psychology (Vol. 44, pp. 247–296). Elsevier. [Google Scholar] [CrossRef] [Scilit]
  17. Dunning, D., Johnson, K., Ehrlinger, J., & Kruger, J. (2003). Why people fail to recognize their own incompetence. Current Directions in Psychological Science, 12(3), 83–87. [Google Scholar] [CrossRef] [Scilit]
  18. Efron, B., & Tibshirani, R. J. (1993). An introduction to the bootstrap (Vol. 57). Chapman and Hall. [Google Scholar] [CrossRef] [Scilit]
  19. Ehrlinger, J., Johnson, K., Banner, M., Dunning, D., & Kruger, J. (2008). Why the unskilled are unaware: Further explorations of (absent) self-insight among the incompetent. Organizational Behavior and Human Decision Processes, 105(1), 98–121. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Eldering, J., Kinkhorst, T., & van de Warken, P. (2022). DOMjudge: Programming contest jury system v8.1.3 [Computer software]. Available online: https://www.domjudge.org/ (accessed on 30 May 2026).
  21. European Commission/EACEA/Eurydice. (2022). Informatics education at school in Europe (Tech. Rep.). Publications Office of the European Union. Available online: https://eurydice.eacea.ec.europa.eu/publications/informatics-education-school-europe (accessed on 25 May 2026). [CrossRef]
  22. Falchikov, N., & Boud, D. (1989). Student self-assessment in higher education: A meta-analysis. Review of Educational Research, 59(4), 395–430. [Google Scholar] [CrossRef]
  23. Fekete, I., Gregorics, T., Kovácsné Pusztai, K., & Veszprémi, A. (2020). Programming theorems and their applications. Teaching Mathematics and Computer Science, 17(2), 213–241. [Google Scholar] [CrossRef] [Scilit]
  24. Feld, J., Sauermann, J., & de Grip, A. (2017). Estimating the relationship between skill and overconfidence. Journal of Behavioral and Experimental Economics, 68, 18–24. [Google Scholar] [CrossRef] [Scilit]
  25. Flavell, J. H. (1979). Metacognition and cognitive monitoring: A new area of cognitive-developmental inquiry. American Psychologist, 34(10), 906–911. [Google Scholar] [CrossRef]
  26. Fóthi, Á. (2012). Bevezetés a programozáshoz (3rd ed.). ELTE Eötvös Kiadó. (Original work published 1983). Available online: https://bzsr.web.elte.hu/progmod/konyv.pdf (accessed on 31 May 2026).
  27. Friedman, H. H. (2023). Cognitive biases and their influence on critical thinking and scientific reasoning: A practical guide. SSRN Electronic Journal, 2958800. [Google Scholar] [CrossRef] [Scilit]
  28. Gibbs, S., Moore, K., Steel, G., & McKinnon, A. (2017). The Dunning–Kruger effect in a workplace computing setting. Computers in Human Behavior, 72, 589–595. [Google Scholar] [CrossRef] [Scilit]
  29. Gignac, G. E., & Zajenkowski, M. (2020). The Dunning–Kruger effect is (mostly) a statistical artefact: Valid approaches to testing the hypothesis with individual differences data. Intelligence, 80, 101449. [Google Scholar] [CrossRef] [Scilit]
  30. Gorson, J., & O’Rourke, E. (2020). Why do CS1 students think they’re bad at programming? Investigating self-efficacy and self-assessments at three universities. In Proceedings of the 2020 ACM conference on international computing education research (pp. 170–181). ACM. [Google Scholar] [CrossRef] [Scilit]
  31. Government of Hungary. (2020). 110/2012. (VI. 4.) Korm. rendelet a Nemzeti alaptanterv kiadásáról, bevezetéséről és alkalmazásáról: Digitális kultúra. Government Decree/National Core Curriculum. Available online: https://www.oktatas.hu/pub_bin/dload/kozoktatas/kerettanterv/NAT.pdf (accessed on 25 May 2026).
  32. Hacker, D. J., Bol, L., Horgan, D. D., & Rakow, E. A. (2000). Test prediction and performance in a classroom context. Journal of Educational Psychology, 92(1), 160–170. [Google Scholar] [CrossRef]
  33. Han, J., Kelley, T., & Knowles, J. G. (2021). Factors influencing student STEM learning: Self-efficacy and outcome expectancy, 21st century skills, and career awareness. Journal for STEM Education Research, 4(2), 117–137. [Google Scholar] [CrossRef] [Scilit]
  34. Hartwig, M. K., & Dunlosky, J. (2014). The contribution of judgment scale to the unskilled-and-unaware phenomenon: How evaluating others can exaggerate over- (and under-) confidence. Memory & Cognition, 42(1), 164–173. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Jansen, R. A., Rafferty, A. N., & Griffiths, T. L. (2021). A rational model of the Dunning–Kruger effect supports insensitivity to evidence in low performers. Nature Human Behaviour, 5, 756–763. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Janssen, E. M., & Lazonder, A. W. (2024). Meta-analysis of interventions for monitoring accuracy in problem solving. Educational Psychology Review, 36, 96. [Google Scholar] [CrossRef] [Scilit]
  37. K–12 Computer Science Framework Steering Committee. (2016). K–12 computer science framework (Tech. Rep.). Available online: https://k12cs.org/wp-content/uploads/2016/09/K%E2%80%9312-Computer-Science-Framework.pdf (accessed on 25 May 2026).
  38. Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux. Available online: https://us.macmillan.com/books/9781429969352/thinkingfastandslow (accessed on 25 May 2026).
  39. Kang, Z., Xiong, Y., Weng, Y., Wang, M., & Jiang, K. (2023). Developing college students’ computational thinking multidimensional test based on Life Story situations. Education and Information Technologies, 28(3), 2661–2679. [Google Scholar] [CrossRef] [Scilit]
  40. Karaca, M., & Bowers, A. J. (2023). Low-performing students confidently overpredict their grade performance throughout the semester. Journal of Intelligence, 11(10), 188. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Kirschner, P. A., & De Bruyckere, P. (2017). The myths of the digital native and the multitasker. Teaching and Teacher Education, 67, 135–142. [Google Scholar] [CrossRef] [Scilit]
  42. Kruger, J., & Dunning, D. (1999). Unskilled and unaware of it: How difficulties in recognizing one’s own incompetence lead to inflated self-assessments. Journal of Personality and Social Psychology, 77(6), 1121–1134. [Google Scholar] [CrossRef] [PubMed]
  43. Kun, A. I. (2016). A comparison of self versus tutor assessment among Hungarian undergraduate business students. Assessment & Evaluation in Higher Education, 41(3), 350–367. [Google Scholar] [CrossRef] [Scilit]
  44. Kun, A. I., Farkas, J., & Juhász, C. (2023). The Dunning–Kruger effect in knowledge management: Examination of BSc level business students. International Journal of Engineering and Management Sciences, 8(1), 14–21. [Google Scholar] [CrossRef]
  45. Lebuda, I., Karwowski, M., & Forthmann, B. (2024). No strong support for a Dunning–Kruger effect in creativity: Analyses of self-assessment in absolute and relative terms. Personality and Individual Differences, 218, 112503. [Google Scholar] [CrossRef] [Scilit]
  46. Lee, M., & Lee, J. (2021). Enhancing computational thinking skills in informatics in secondary education: The case of South Korea. Educational Technology Research and Development, 69(5), 2869–2893. [Google Scholar] [CrossRef] [Scilit]
  47. Leppänen, L., Leinonen, J., & Hellas, A. (2025). Revisiting the confidence gap in university-level programming courses. In Proceedings of ITiCSE 2025. ACM. [Google Scholar] [CrossRef] [Scilit]
  48. Lodi, M., & Martini, S. (2021). Computational thinking, between Papert and Wing. Science & Education, 30, 883–908. [Google Scholar] [CrossRef] [Scilit]
  49. MacNeil, S. L., Wood, E., & Arslantas, F. (2024). Development of a metacognition co-curriculum for a university course in introductory organic chemistry. Frontiers in Education, 9, 1402599. [Google Scholar] [CrossRef] [Scilit]
  50. Magnus, J. R., & Peresetsky, A. A. (2022). A statistical explanation of the Dunning–Kruger effect. Frontiers in Psychology, 13, 840180. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  51. Mahmood, K. (2016). Do people overestimate their information literacy skills? A systematic review of empirical evidence. Communications in Information Literacy, 10(2), 199–213. [Google Scholar] [CrossRef] [Scilit]
  52. Máté, D., & Darabos, É. (2017). Measuring the accuracy of self-assessment among undergraduate students in higher education to enhance their competence. Journal of Competitiveness, 9(2), 78–92. [Google Scholar] [CrossRef] [Scilit]
  53. Máté, D., Kiss, J. T., & Csernoch, M. (2025). Cognitive biases in user experience and spreadsheet programming. Education and Information Technologies, 30, 14821–14851. [Google Scholar] [CrossRef] [Scilit]
  54. McGuire, S. A. (2023). Question format biases college students’ metacognitive judgments for exam performance. International Journal for the Scholarship of Teaching and Learning, 17(1), 15. [Google Scholar] [CrossRef] [Scilit]
  55. McIntosh, R. D., Fowler, E. A., Lyons, T., & Hamilton, C. (2019). Wise up: Clarifying the role of metacognition in the Dunning–Kruger effect. Journal of Experimental Psychology: General, 148(11), 1882–1897. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502–517. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Nagy, T., Csernoch, M., & Biró, P. (2021). The comparison of students’ self-assessment, gender, and programming-oriented spreadsheet skills. Education Sciences, 11(10), 590. [Google Scholar] [CrossRef] [Scilit]
  58. Nietfeld, J. L., Cao, L., & Osborne, J. W. (2005). Metacognitive monitoring accuracy and student performance in the postsecondary classroom. The Journal of Experimental Education, 74(1), 7–28. [Google Scholar]
  59. Osterhage, J. L. (2021). Persistent miscalibration for low and high achievers despite practice test feedback in an introductory biology course. Journal of Microbiology & Biology Education, 22(2), e00139-21. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Pearson, K., & Filon, L. N. G. (1898). Mathematical contributions to the theory of evolution. IV. On the probable errors of frequency constants and on the influence of random selection on variation and correlation. Philosophical Transactions of the Royal Society of London: Series A, 191, 229–311. [Google Scholar] [CrossRef] [Scilit]
  61. Prensky, M. (2001). Digital natives, digital immigrants part 1. On the Horizon, 9(5), 1–6. [Google Scholar] [CrossRef] [Scilit]
  62. Prims, J. P., & Moore, D. A. (2017). Overconfidence over the lifespan. Judgment and Decision Making, 12(1), 29–41. [Google Scholar] [CrossRef] [Scilit]
  63. Raghunathan, T. E., Rosenthal, R., & Rubin, D. B. (1996). Comparing correlated but nonoverlapping correlations. Psychological Methods, 1(2), 178–183. [Google Scholar] [CrossRef]
  64. Schaefer, S., Riediger, M., Li, S.-C., & Lindenberger, U. (2023). Too easy, too hard, or just right: Lifespan age differences and gender effects on task difficulty choices. International Journal of Behavioral Development, 47(3), 253–264. [Google Scholar] [CrossRef] [Scilit]
  65. Schunk, D. H., & Zimmerman, B. J. (1997). Social origins of self-regulatory competence. Educational Psychologist, 32(4), 195–208. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  66. Serra, M. J., & DeMarree, K. G. (2016). Unskilled and unaware in the classroom: College students’ desired grades predict their biased grade predictions. Memory & Cognition, 44, 1127–1137. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  67. Sobral, S. R. (2021). CS1 student grade prediction: Unconscious optimism vs insecurity? International Journal of Information and Education Technology, 11(8), 387–391. [Google Scholar] [CrossRef] [Scilit]
  68. Strickroth, S. (2024). Exploring students’ self-confidence in their programming solutions. In Proceedings of ITiCSE 2024. ACM. [Google Scholar] [CrossRef] [Scilit]
  69. Surdilović, D., Petrović, D., Pucar, A., Marković, D., Vasiljević, M., & Vukićević, B. (2022). Evaluation of the Dunning–Kruger effects among dental students at an academic training institution. Healthcare, 10(12), 2418. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  70. Sweller, J., Ayres, P., & Kalyuga, S. (2011). Cognitive load theory. Springer. [Google Scholar] [CrossRef] [Scilit]
  71. Szlávi, P., & Zsakó, L. (2016). A programozás gondolkodási eszköztára—Algoritmikus absztrakció, dekompozíció-szuperpozíció. In Infodidact 2016 (pp. 1–8). Webdidaktika Alapítvány. Available online: https://konferenciak.inf.elte.hu/infodidact/InfoDidact16/Manuscripts/SzPZsL.pdf (accessed on 25 May 2026).
  72. Wilcoxon, F. (1945). Individual comparisons by ranking methods. Biometrics Bulletin, 1(6), 80–83. [Google Scholar] [CrossRef] [Scilit]
  73. Xia, M., Poorthuis, A. M. G., & Thomaes, S. (2024). Children’s overestimation of performance across age, task, and historical time: A meta-analysis. Child Development, 95(3), 1001–1022. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. Yang Hansen, K., Thorsen, C., Radišić, J., Peixoto, F., Laine, A., & Liu, X. (2024). When competence and confidence are at odds: A cross-country examination of the Dunning–Kruger effect. European Journal of Psychology of Education, 39, 1537–1559. [Google Scholar] [CrossRef] [Scilit]
  75. Zhang, Y., Cheung, S. W., & Giovanelli, E. (2024). A systematic umbrella review on computational thinking assessment in higher education. European Journal of STEM Education, 9(1), 2. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Self-estimate versus actual score for (a) the Assessment 1 concurrent quiz judgment ( N = 101 ), (b) the Assessment 2 written postdiction ( N = 117 ), and (c) the Assessment 2 laboratory postdiction ( N = 117 ). The dashed line shows perfect calibration ( y = x ); points above the line indicate overestimation, points below indicate underestimation.
Figure 1. Self-estimate versus actual score for (a) the Assessment 1 concurrent quiz judgment ( N = 101 ), (b) the Assessment 2 written postdiction ( N = 117 ), and (c) the Assessment 2 laboratory postdiction ( N = 117 ). The dashed line shows perfect calibration ( y = x ); points above the line indicate overestimation, points below indicate underestimation.
Education 16 01476 g001
Figure 2. Distribution of signed errors (estimate − actual) for the three judgments: Assessment 1 concurrent quiz judgment, Assessment 2 written postdiction, and Assessment 2 laboratory postdiction. Bias direction reverses from positive (overestimation) on the concurrent quiz judgment to negative (underestimation) on both postdiction components. Within each violin the red line marks the mean and the black line the median, and the dashed horizontal line marks zero error.
Figure 2. Distribution of signed errors (estimate − actual) for the three judgments: Assessment 1 concurrent quiz judgment, Assessment 2 written postdiction, and Assessment 2 laboratory postdiction. Bias direction reverses from positive (overestimation) on the concurrent quiz judgment to negative (underestimation) on both postdiction components. Within each violin the red line marks the mean and the black line the median, and the dashed horizontal line marks zero error.
Education 16 01476 g002
Figure 3. Cross-cohort replication of the Assessment 1 concurrent quiz judgment in three consecutive academic years: mean signed error in percentage points, and the Spearman rank-order correlation between estimate and actual score, scaled by ten for display on a common axis. The direction and magnitude of the overestimation are consistent across cohorts.
Figure 3. Cross-cohort replication of the Assessment 1 concurrent quiz judgment in three consecutive academic years: mean signed error in percentage points, and the Spearman rank-order correlation between estimate and actual score, scaled by ten for display on a common axis. The direction and magnitude of the overestimation are consistent across cohorts.
Education 16 01476 g003
Table 1. Descriptive statistics for Assessment 1 (concurrent self-assessment), N = 101 .
Table 1. Descriptive statistics for Assessment 1 (concurrent self-assessment), N = 101 .
VariableMean SD Min Q 1 Median Q 3 Max
Self-estimate (%)46.1320.010.0030.0050.0063.1290.00
Actual score (%)37.0413.248.3328.1336.4645.4968.75
Signed error (E − A) + 9.10 16.63 33.82 + 8.33 + 49.17
Table 2. Calibration metrics for Assessment 1 ( N = 101 ).
Table 2. Calibration metrics for Assessment 1 ( N = 101 ).
MetricValue
Mean signed error (bias) + 9.10  pp (overestimation)
Mean absolute error15.01 pp
RMSE18.88
Pearson r0.565 ( p < 0.001 )
Spearman ρ 0.533 ( p < 0.001 )
Overestimators (n, %)71 (70.3%)
Underestimators (n, %)30 (29.7%)
Exact matches0 (0.0%)
Table 3. Calibration metrics for Assessment 2 components ( N = 117 ).
Table 3. Calibration metrics for Assessment 2 components ( N = 117 ).
MetricWritten ExamLaboratory
Self-estimate M ( S D )48.39 (19.50)52.29 (27.82)
Actual (rescaled) M ( S D )58.47 (20.13)56.50 (27.71)
Mean signed error (bias) 10.08  pp 4.21  pp
S D signed error17.8517.53
Mean absolute error15.55 pp11.39 pp
RMSE20.4317.95
Pearson r0.595 ***0.801 ***
Spearman ρ 0.633 ***0.804 ***
Overestimators (n, %)30 (25.6%)32 (27.4%)
Underestimators (n, %)85 (72.6%)57 (48.7%)
Exact matches (n, %)2 (1.7%)28 (23.9%)
*** p < 0.001 .
Table 4. Within-subject comparison in the 94-student linked subsample, on all three calibration dimensions.
Table 4. Within-subject comparison in the 94-student linked subsample, on all three calibration dimensions.
ContrastMean Diff.Paired t ( 93 ) pCohen’s d z
Signed error
A1 signed err + 8.99  pp
A2 written signed err 11.06  pp
A2 laboratory signed err 4.70  pp
A1 − A2 written (paired) + 20.05  pp9.73<0.0011.00
A1 − A2 laboratory (paired) + 13.69  pp6.21<0.0010.64
Absolute error
A1 absolute err14.95 pp
A2 written absolute err15.93 pp
A2 laboratory absolute err11.72 pp
A1 − A2 written (paired) 0.98  pp 0.55 0.582 0.06
A1 − A2 laboratory (paired) + 3.23  pp1.800.0750.19
Rank-order calibration, Spearman  ρ
A10.532
A2 written0.657
A2 laboratory0.805
Table 5. Tertile split by actual score: Mean signed error per tertile and continuous association between actual score and signed error.
Table 5. Tertile split by actual score: Mean signed error per tertile and continuous association between actual score and signed error.
SampleLower TertileMiddle TertileUpper Tertile ρ (act, err)
A1 quiz, concurrent ( N = 101 ) + 9.79 ( n = 34 ) + 11.21 ( n = 35 ) + 6.06 ( n = 32 ) 0.10 , p = 0.34
A2 written, postdiction ( N = 117 ) 1.64 ( n = 39 ) 10.34 ( n = 41 ) 18.68 ( n = 37 ) 0.47 , p < 0.001
A2 laboratory, postdiction ( N = 117 ) + 2.57 ( n = 46 ) 5.98 ( n = 38 ) 11.61 ( n = 33 ) 0.33 , p < 0.001
Table 6. Sensitivity of the floor-aware banding to the choice of cut points. The between-band difference is significant under every alternative examined; the low band’s overestimation above zero is not.
Table 6. Sensitivity of the floor-aware banding to the choice of cut points. The between-band difference is significant under every alternative examined; the low band’s overestimation above zero is not.
Low BandHigh Bandn Low/HighGapp
2–37–1017/501.8000.00014
2–47–1036/501.4710.00004
1–37–1024/501.3850.00058
0–37–1027/501.3440.00037
2–38–1017/331.9430.00015
2–36–1017/591.7860.00012
2–56–1048/591.3050.00012
0–56–1058/591.1760.00020
The low band’s mean signed error excludes zero only for the 2–3 and 2–4 bands.
Table 7. Cross-cohort replication of Assessment 1 calibration.
Table 7. Cross-cohort replication of Assessment 1 calibration.
CohortNEstimate MActual MSigned Error% Overest.Spearman ρ
2023–202410146.1337.04 + 9.10  pp70.3%0.533
2024–2025 (fall)10044.8239.41 + 5.40  pp62.0%0.390
2025–2026 (spring)5648.2041.12 + 7.08  pp69.6%0.434
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Vekov, G.; Csernoch, M. From Concurrent Self-Assessment to Postdiction: Grade-Judgment Calibration Across Two Assessments in a First-Year Computer Science Course. Educ. Sci. 2026, 16, 1476. https://doi.org/10.3390/educsci16091476

AMA Style

Vekov G, Csernoch M. From Concurrent Self-Assessment to Postdiction: Grade-Judgment Calibration Across Two Assessments in a First-Year Computer Science Course. Education Sciences. 2026; 16(9):1476. https://doi.org/10.3390/educsci16091476

Chicago/Turabian Style

Vekov, Géza, and Maria Csernoch. 2026. "From Concurrent Self-Assessment to Postdiction: Grade-Judgment Calibration Across Two Assessments in a First-Year Computer Science Course" Education Sciences 16, no. 9: 1476. https://doi.org/10.3390/educsci16091476

APA Style

Vekov, G., & Csernoch, M. (2026). From Concurrent Self-Assessment to Postdiction: Grade-Judgment Calibration Across Two Assessments in a First-Year Computer Science Course. Education Sciences, 16(9), 1476. https://doi.org/10.3390/educsci16091476

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop