Next Article in Journal
Artificial Intelligence and Academic Integrity in Virtual Higher Education: A Descriptive-Comparative Study of Student and Faculty Perceptions in Ecuador
Next Article in Special Issue
Designing an Immersive 360° Video Learning Environment for Virtual Chemistry Study Visits: Exploring Authentic Chemistry Research
Previous Article in Journal
Higher Education for Sustainability—Intergenerational Comparative Analysis of the Perceptions of Students at the University of the Basque Country Regarding Socioecological Transitions
Previous Article in Special Issue
Research at the Core: How Philippine Science Faculty in State Universities Enact the Research Function Within Trifocal Roles
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

From Teacher Workshop to Self-Study Online Module: Developing a Video-Based Resource for Diagnostic Accuracy in Primary Mathematics Teacher Education

1
Department of Inclusive Education, University of Potsdam, 14476 Potsdam, Germany
2
Department of Primary School Education, University of Potsdam, 14476 Postdam, Germany
*
Author to whom correspondence should be addressed.
Trends High. Educ. 2026, 5(3), 66; https://doi.org/10.3390/higheredu5030066
Submission received: 30 May 2026 / Revised: 9 July 2026 / Accepted: 14 July 2026 / Published: 16 July 2026

Abstract

Diagnostic competence is central to primary mathematics teacher education but existing approaches tend to depend on individual personnel and module structures, limiting their sustainability. This study combines a research-informed design approach with quasi-experimental intervention cycles: a fully asynchronous, video-based Moodle self-study module was developed to foster diagnostic competence in early arithmetic, consistently grounded in an empirically validated cognitive-developmental model. The module was implemented and iteratively refined across four phases—an in-service teacher workshop, a module pilot, and two quasi-experimental intervention cycles with pre-service primary teachers, including matched control groups. Across phases, the most consistent finding was a significant pre–post improvement in diagnostic accuracy for a below-average-performing pupil profile in the intervention group, while effects for average- and above-average profiles were less uniform. Log-data analyses showed that actual engagement with video vignettes, rather than structural completion, was the active ingredient of learning gains. Three design principles are proposed: model-based mathematical design, vertical connection of in-service and pre-service contexts, and guided asynchronous self-study with structured feedback. The study provides promising, context-bound evidence that a theoretically grounded, personnel-independent digital module can be integrated sustainably into primary teacher education and can foster diagnostic accuracy under the institutional and sample conditions investigated, while yielding initial design knowledge for comparable development efforts.

1. Introduction

Research-based teacher education is increasingly called upon to produce not only conceptual knowledge but durable, scalable resources that remain in use beyond individual projects and personnel [1,2]. In practice, however, pedagogical innovations in higher education tend to be tied to specific modules and lecturers, and the structural disconnect between pre-service and in-service phases further limits their reach and longevity. Diagnostic competence—the ability to accurately perceive, interpret, and respond to students’ learning—is particularly affected by this structural problem: it is widely recognised as central to instructional quality in heterogeneous classrooms, yet remains underdeveloped in initial teacher education [3,4].
The present study addresses this gap by developing and evaluating a fully asynchronous, video-based self-study module for primary mathematics teacher education. The module is based on an empirically validated cognitive-developmental model of early arithmetic [5]. It was developed and refined across four phases: an in-service teacher workshop, a module pilot, and two quasi-experimental intervention cycles with pre-service primary teachers. The module is designed as a personnel-independent resource that can be attached to existing courses. It aims to bridge pre-service and in-service contexts. Additionally, it is intended to remain sustainable within regular university structures.

2. Theoretical Framework

2.1. Teacher Education Structures and Diagnostic Competence

Despite growing investment in empirically grounded learning resources, pedagogical innovations in higher education are seldom examined for their long-term sustainability and often fail to become embedded in academic practice once project funding ends [1]. This is limiting their integration as durable components of regular programmes. A related challenge is the limited coherence between university-based preparation and subsequent phases of teacher education, so that competences developed in pre-service programmes do not consistently connect to school practice. This issue has long been problematised as contributing to the theory–practice gap in the development of teacher knowledge [6]. In two-phase systems, it is explicitly described as a misalignment between university curricula and induction programmes [7].
Diagnostic competence is particularly vulnerable to this fragmentation: its development requires sustained, practice-connected learning opportunities across pre-service and in-service phases that are currently inconsistently embedded in university-based initial teacher education and rarely continued systematically in subsequent phases [3,4]. Drawing on the tradition of pedagogical diagnostics, diagnostic competence refers to teachers’ capability to carry out the diagnostic tasks of their profession in a way that yields high-quality information about students’ learning prerequisites, processes, and outcomes for pedagogical decision-making [8]. Its development is anchored as a core objective in national and European competence frameworks for teachers, which emphasise assessment and diagnostic competencies and call for coherent, career-long professional learning [9,10].
Building on recent work, diagnostic competence is often modelled as the interplay between a professional knowledge base, diagnostic activities (e.g., problem identification, hypothesis generation, and evidence evaluation), and the accuracy of the resulting judgments [7]. Following the framework proposed by Shulman (1986) [11], this professional knowledge base comprises content knowledge (CK), pedagogical content knowledge (PCK), and general pedagogical knowledge (PK). CK refers to knowledge of the subject matter, PCK integrates knowledge of mathematics with knowledge of how students learn specific mathematical concepts, and PK encompasses general principles of teaching and classroom management that are applicable across subjects [11]. In this framework, applying conceptual knowledge in diagnostic activities is considered essential for developing diagnostic competences. Simulation-based approximations of practice are therefore highlighted as a promising approach. They allow repeated engagement with realistic cases under reduced time pressure while involving fewer ethical and organisational constraints than real classrooms [4]. A meta-analysis by Chernikova et al. (2020) confirmed robust positive effects of such interventions on diagnostic competence in higher education and showed that the effectiveness of instructional support within such simulations depends systematically on learners’ prior knowledge—with more structured scaffolding benefiting novice learners in particular [12]. Scaffolding is commonly defined as contingent and temporary instructional support that enables learners to complete tasks beyond their independent capability and is gradually withdrawn as competence increases [13]. In contemporary research on simulation- and technology-based learning environments, scaffolding is further conceptualised as structured and adaptive support that reduces task complexity and guides learners’ cognitive processing through prompts, cues, or representations [12].
Diagnostic judgement accuracy, in contrast, denotes the degree to which a teacher’s assessment of individual students corresponds to an external criterion such as standardised test scores or expert ratings, typically operationalised via correlations or discrepancy indices, and thus captures the correctness of the outcome of diagnostic activity in a given situation rather than the underlying competence itself [14,15]. Because it is a quantifiable, situation-specific performance indicator that can be reliably measured through pre–post comparisons, judgement accuracy remains the most widely used outcome measure in intervention research evaluating the effects of training on diagnostic performance [15,16]. Empirical findings consistently show that accuracy is not uniform across the performance spectrum: teachers judge pupils at the upper and lower ends of the distribution less accurately than those in the middle, a performance-level dependence that is especially consequential in heterogeneous primary classrooms [16,17,18]. These systematic biases underscore the need for targeted training environments that confront novice teachers with the full range of student performance levels and provide structured opportunities for reflective diagnostic practice, ideally in formats that can be integrated sustainably into regular teacher education programmes.

2.2. Video-Based and Digital Self-Study Formats in Teacher Education

Simulation-based training has emerged as a particularly promising approach to addressing such deficits, given accumulating evidence that it can foster both diagnostic activities and judgement accuracy [12,19,20]. This evidence is particularly relevant given the structural fragmentation described in Section 2.1, as such formats offer a scalable route to practice-connected diagnostic learning that does not depend on stable personnel or institutional arrangements. Building on this evidence base, video-based and digital self-study formats have become central design approaches for diagnostic learning environments in higher teacher education. Recent research documents a broad and growing body of work on video-based learning environments in teacher education [21], with a substantial strand focusing specifically on mathematics and STEM through the use of video and simulation-based approximations of practice [4,22]. These environments build directly on the knowledge base described by Heitzmann et al. (2019) [4], as the video-vignette format is specifically designed to engage PCK and diagnostic activities in a situated, low-stakes context. A central quality feature of these environments is the use of video vignettes that present authentic diagnostic tasks and function as approximations of classroom practice, in the sense of allowing novices to engage with core elements of teaching in reduced, lower-stakes form [23,24].
Simulation- and video-based formats have shown positive effects on teachers’ diagnostic activities and diagnostic accuracy across multiple studies. Meta-analytic evidence indicates that instructional support has a moderate positive effect in simulation-based learning [12]. Subject-specific studies in biology and mathematics confirm these findings. In biology teacher education, PCK scaffolding in a video-based simulation significantly improved pre-service teachers’ diagnostic competence [19]. Similarly, in primary mathematics, content-related scaffolding in a simulation improved diagnostic accuracy [20]. Taken together, these subject-specific replications confirm that the general meta-analytic pattern holds in the domain of diagnostic competence specifically, and that PCK-relevant scaffolding is a robust active ingredient regardless of subject context. The delivery of these formats as asynchronous digital self-study environments is well supported. A meta-analysis of 45 controlled studies found that online and blended learning conditions produced modest but consistent improvements in learning outcomes compared to face-to-face instruction alone [25]. Blended approaches showed a particularly clear advantage. Moreover, the effects increased with longer engagement time. Evidence further suggests that structured, timely feedback is a key moderator of learning gains [26]. It is particularly effective when it targets the decision-making process and the reasoning underlying specific decisions rather than overall performance alone [26]. Feedback that targets the decision-making process and the reasoning underlying specific judgments, rather than overall performance, has been identified as among the most powerful instructional supports in general learning research [26]. Its effectiveness as a moderator of learning outcomes has been confirmed specifically in simulation-based formats [12].
At the same time, subgroup analyses and profile studies indicate that these effects are not uniform but differ between learner groups in theoretically meaningful ways. Differences in learners’ professional knowledge constitute one theoretically relevant source of variation. Studies consistently show that learners with stronger pedagogical content knowledge (PCK) achieve greater diagnostic accuracy and engage in more differentiated diagnostic activities. For example, Kramer et al. (2021) found that PCK was the strongest knowledge-based predictor of diagnostic activities among pre-service biology teachers [3]. Pedagogical knowledge also contributed significantly. In addition, PCK significantly predicted diagnostic accuracy. Complementarily, Chernikova et al. (2020) showed in their meta-analysis that the type of scaffolding required for effective learning shifts systematically as a function of learners’ prior knowledge base [12]. Together, these findings indicate that PCK functions not only as a predictor of training outcomes but also as a moderator of what type of scaffolding is effective. Learners with weaker knowledge bases require more structured support to benefit. Regarding engagement and usage intensity, log-data analyses reveal that learners who engage actively with video-based simulation tasks achieve substantially different outcomes than those who engage superficially. Schons et al. (2023) [20] identified three engagement modes—passive, active, and constructive—based on pre-service teachers’ task selection behaviour and diagnostic interpretation patterns. Constructive engagement, distinguished by strategic selection of diagnostically relevant tasks and elaborated hypothesis generation, was the only mode that significantly predicted higher assessment accuracy [20]. Radkowitsch et al. (2023) used latent profile analysis to identify three qualitatively different diagnostic process profiles [27]. The profiles differed in the depth of describing, explaining, and decision-making in response to simulated student behaviour [27]. More elaborative profiles were associated with higher diagnostic accuracy, particularly among learners with stronger professional knowledge [27]. The convergence of these findings across two independent studies, using log data and latent profile analysis, suggests that engagement depth and prior knowledge interact, such that deeper engagement appears especially productive for learners who already possess the domain-specific knowledge needed to interpret what they observe. Taken together, these findings support the assumption that structured feedback, study-programme-related knowledge background, and the depth of individual engagement with digital learning materials are each important predictors of the effectiveness of video- and simulation-based online modules. PCK in particular functions as both a predictor of training outcomes and a boundary condition for meaningful diagnostic interpretation [3,4].

2.3. Cognitive-Developmental Model of Arithmetical Concepts

In primary mathematics, this domain-specific knowledge base must include a principled understanding of how numerical and arithmetical concepts develop in young children, as this provides the interpretive frame against which pupils’ responses can be assessed [4,28]. This is consistent with the finding reviewed in Section 2.2 that learners with insufficient PCK struggle to identify conceptually relevant patterns in pupils’ responses even under supportive instructional conditions. The cognitive-developmental model of arithmetic concepts [5] constitutes precisely such a content-specific referent for the domain of early arithmetic. In this model, the development of basic numerical and arithmetical concepts between roughly ages four and eight is described as a hierarchy of six qualitatively distinct levels: (1) count number, (2) mental number line, (3) cardinality and decomposability, (4) class inclusion and embeddedness, (5) relationality, and (6) units in numbers (bundling and unbundling) [5,29]. Each level in this model is associated with specific conceptual knowledge structures, typical solution strategies, and characteristic error patterns, so that pupils’ responses on diagnostic tasks can be interpreted as indicators of their current conceptual level rather than as isolated right-or-wrong answers [5,30]. In terms of the diagnostic competence framework outlined in Section 2.1, this means that the model directly supports the hypothesis generation and evidence evaluation activities that constitute the core of diagnostic practice.
The model has been validated in longitudinal and cross-sectional studies demonstrating that developmental trajectories follow the proposed level sequence and predict later mathematics achievement. It underpins the standardised instrument MARKO-D1+, designed for use in Grade 1 and the beginning of Grade 2 [29,30,31]. For teacher education, the model’s level descriptors can be used to design diagnostic tasks and interpret pupils’ strategies in theory-grounded terms [5,29]. The descriptors specify knowledge structures, typical solution strategies, and characteristic error patterns [5,29]. In this way, the model constitutes a topic-specific instantiation of PCK that links knowledge of children’s mathematical learning progressions to practical diagnostic work in primary mathematics [28], and provides a coherent conceptual backbone for a video-based diagnostic training module. Together, the cognitive-developmental model and the evidence on simulation-based training reviewed above identify a content-specific, theoretically grounded foundation for diagnostic training in early arithmetic. However, existing simulation- and video-based formats for diagnostic training have predominantly been examined as researcher-designed interventions embedded in specific course settings and dependent on dedicated facilitation [1,32]. This limits their transferability across programme tracks and their sustainability beyond individual projects [1]. There is, to date, comparatively little evidence on fully asynchronous, personnel-independent self-study formats that are tightly grounded in a validated domain-specific developmental model and empirically evaluated across multiple independent cohorts under controlled conditions.

2.4. Research-Informed Design for Digital Learning Resources

Translating this converging evidence base into an educationally sound and practically usable resource requires a principled design methodology capable of integrating theoretical knowledge with iterative, practice-oriented development. Design-based research has been highlighted as a particularly appropriate approach for the research-grounded development of curricula, courses, and learning resources, given its commitment to iterative refinement and the integration of theoretical and practical knowledge [2,33]. Quality in this tradition is commonly evaluated in terms of content validity, practicality, and effectiveness [34,35].
Research-informed design is a methodological approach situated within this broader field that incorporates influences from existing research and in which specific design choices are shaped by their contribution to practice in education [36]. In Stokes’s (1997) terms, it sits closer to the applied end of the quadrant model, prioritising the development of high-quality, practically useful resources over explicit theory building while remaining more strongly grounded in research than standard instructional systems design [37,38]. This position is particularly appropriate for iterative development studies where the primary contribution lies in systematic, evidence-informed design and evaluation rather than in dual-purpose theory generation [39]. The existing literature provides several concrete examples of research-informed design in mathematics and teacher education. The Cambridge Mathematics Framework applies research-informed design at the level of a broad, curriculum-independent reference framework, mapping conceptual progressions across the full domain of school mathematics to support curriculum design, resource development, and teacher education [36]. The FALEDIA platform operationalises a related logic at the level of a subject-specific, digital learning environment, using case-based modules and diagnostic tasks to foster pre-service primary teachers’ diagnostic skills in mathematics, with evidence of significant pre–post gains in diagnostic competence across a sample of 695 pre-service teachers from two universities [40]. Both projects exemplify how research-informed design can translate domain-specific research on mathematical learning into usable tools for curriculum and resource development. However, they differ in scope and granularity: the Cambridge Mathematics Framework provides a comprehensive, curriculum-independent map of mathematical ideas, whereas FALEDIA offers a targeted, platform-based environment for diagnostic case work in primary mathematics. Rillero and Camposeco (2018) illustrate the iterative development and use of an online problem-based learning module for pre-service and in-service teachers, demonstrating how repeated cycles of formative evaluation and revision can produce a practically usable, research-grounded online resource [41]. Together, these examples show how research-informed design can be applied at framework, platform, and module levels. The present study contributes a further instance by developing and evaluating a fully asynchronous, personnel-independent self-study module anchored in a validated developmental model and designed for flexible integration across programme tracks. Research-informed design therefore provides an appropriate methodological framework for the iterative development and evaluation of video-based digital learning resources in higher teacher education.

2.5. Research Questions

The module developed in this study pursues the dual goal of research-informed design: producing a high-quality, practically usable digital resource and generating cumulative evidence that can inform comparable development efforts in primary mathematics teacher education [2,36]. It integrates the theoretical strands reviewed above as follows: (1) It responds to structural constraints in teacher education (Section 2.1) by providing a personnel-independent, asynchronous format embeddable across programme tracks. (2) It operationalises diagnostic competence (Section 2.1) by targeting accuracy as the measurable outcome through tasks designed to activate the underlying knowledge and process dimensions. (3) It draws on design evidence for video-based, feedback-enriched simulation formats (Section 2.2) to support constructive engagement and differentiated learning across prior knowledge backgrounds. (4) It uses the cognitive-developmental model of arithmetical concepts (Section 2.3) as the operative content framework, providing the domain-specific PCK identified as a boundary condition for accurate diagnostic judgement. (5) It follows a research-informed design logic (Section 2.4) through iterative empirical evaluation across four phases. Taken together, these design choices position the module as a topic-specific instantiation of research-informed design for early arithmetic that combines a validated developmental model, video-based diagnostic simulations, and structured feedback within a personnel-independent, asynchronous format.
Quality is evaluated along four dimensions: content validity, practicality, effectiveness, and iterative refinement [34,35]. Content validity is treated as a design-level prerequisite, established through systematic grounding of all tasks and feedback in the Fritz et al. (2018) model [5]. Practicality is examined through participant evaluations of relevance, usability, and perceived learning across project phases. Effectiveness is examined through pre–post diagnostic accuracy under quasi-experimental conditions with a control group. It is further differentiated by pupil performance level, learner background, and engagement depth, consistent with evidence that the effects of video-based learning environments vary along these dimensions [3,12,20]. Iterative refinement is documented through transparent reporting of design changes made between phases and their observable consequences for outcome patterns. This evidence is interpreted in the discussion rather than addressed through a separate research question. Following the iterative design logic and the quality dimensions outlined above, the following research questions were addressed:
  • RQ1 (Formative Effectiveness): What exploratory evidence regarding pre–post gains in diagnostic accuracy can be derived from the initial workshop and pilot phases to inform the module’s iterative refinement?
  • RQ2 (Summative Effectiveness): To what extent does participation in the finalised module lead to pre–post improvements in diagnostic accuracy compared to a control group across the main intervention cycles?
  • RQ2a (Differential Background): How do these pre–post effects differ between pre-service teachers from inclusive education, mathematics specialisation, and other primary education programmes?
  • RQ2b (Behavioural Engagement): Within the intervention group, to what extent is behavioural engagement with the module’s tasks, as indicated by task completion rates and video-vignette skipping behaviour, associated with variations in pre–post diagnostic accuracy gains?
  • RQ3 (Practicality): How do participants across all project phases perceive the relevance, usability, and perceived learning effects of the diagnostic training concept?
The theoretical framework and research questions outlined above motivate a study design that combines iterative module development with empirical evaluation across multiple phases. Existing approaches to diagnostic training in primary mathematics teacher education depend heavily on specific personnel, time slots, and module assignments, limiting their reach and sustainability [1,7]. Two features of the content domain suggested a viable path forward: the cognitive-developmental model of arithmetical concepts provides structured building blocks for diagnostic training activities, and video-based formats allow authentic diagnostic practice without requiring live classroom access. Together, these conditions motivated the development of a video-based online self-study module organised around three core design principles: (1) consistent model-based mathematical design, where the cognitive-developmental model functions not only as theoretical background but as the operative framework for designing video tasks, structuring activities, and evaluating diagnostic judgments; (2) vertical connection of in-service and pre-service contexts, where the module was first developed and tested with practising primary teachers so that its content reflects genuine diagnostic challenges before being transferred to pre-service settings; and (3) structural independence through a self-study format, where implementation as a self-paced, asynchronous course makes the resource independent of specific staff and enables integration across different courses and programme tracks over time.

3. Materials and Methods

3.1. Research-Informed Design Process

The study follows a research-informed design logic ([36]; see Section 2.4), with the primary aim of producing a high-quality, practically usable, and theoretically grounded online self-study module. The approach involves the productive application of existing theory and evidence to design decisions, iterative refinement informed by formative evaluation, and transparent documentation of design rationale [36]. The cognitive-developmental model [5] serves as the shared referential framework for diagnostic tasks, design goals, and outcome measurement throughout all four phases.
Figure 1 summarises the four phases, their samples, and their contribution to both design knowledge and empirical evaluation. Across all four phases, theoretical and design knowledge accumulates iteratively, while the two intervention cycles additionally generate a combined sample for effectiveness and engagement analyses.
Phase 1—Exploratory in-service workshop: An in-service workshop with primary teachers (N = 9) introduced the cognitive-developmental model and engaged participants with authentic pupil video vignettes, alongside short pre- and post-assessments and a course evaluation. This phase served to examine the approach in a practice-oriented setting and to generate design-relevant insights before constructing the digital module.
Phase 2—Pilot of the module: Workshop findings informed the development of an eight-week Moodle-based self-study module as a supplementary, self-paced component for pre-service primary teachers, combining units on diagnostic competence and children’s arithmetical development with interactive video-based judgement tasks and activities on task selection and adaptation. The 14 student videos were rated by two expert raters with substantial agreement (Cohen’s κ = 0.72). A pilot with bachelor’s students (N = 4) examined feasibility, technical usability, content clarity, and workload; structured feedback informed documented revisions [36].
Phase 3—First intervention cycle: The revised module was implemented with a bachelor’s cohort (N = 88) enrolled in a compulsory mathematics education lecture, alongside a parallel control group (N = 73) following the same course without module access. Both groups completed identical diagnostic pre- and post-assessments; learning platform usage data were logged for the intervention group.
Phase 4—Second intervention cycle: A structurally refined module version was implemented with a new bachelor’s cohort (n = 37) and a matched control group (n = 40) in the same institutional setting. Identical pre- and post-assessments and log data were collected, enabling comparison across two intervention cycles under stable content conditions but with an altered engagement architecture.

3.2. Setting, Participants, and Study Design

The study was conducted across four iterative phases at a German university (see Table 1). The two formative phases are marked in grey. Phase 1): The exploratory workshop comprised three 3 h sessions with in-service primary teachers (N = 9) and included a pre–post assessment. Phase 2): The pilot introduced the first module version to a small group of pre-service teachers (N = 4), whose feedback informed technical and didactical revisions. Both intervention phases were embedded in the same compulsory mathematics education lecture; students self-selected into an intervention group with module access or a control group without (quasi-experimental design), and both groups completed identical diagnostic pre- and post-assessments. Phase 3): The first intervention cycle drew on three consecutive cohorts (N = 161). Phase 4): The second intervention cycle, one year later, used the refined module (N = 77). Across both intervention phases, the combined sample comprised N = 238 pre-service teachers (207 female, 30 male, 1 diverse; M age = 23.7 years, SD = 3.1), with study programme specialisations including inclusive education (n = 84), mathematics (n = 73), and other primary education tracks (n = 81).

3.3. Instruments

Diagnostic judgement accuracy was assessed using a video-based instrument developed and validated by Wagner and Ehlert (2019) [42]. The instrument presents three video vignettes, each showing a different child solving tasks from the standardised arithmetic test MARKO-D1+ (internal consistency r = 0.91; test–retest r = 0.89) and satisfactory construct validity, with concurrent correlations to established mathematics tests ranging from r = 0.56 to r = 0.77) [30]. One video vignette is showing a child performing at below-average, one at average, and one at above-average level. For each vignette, participants provide three types of judgement: (a) task-specific item-level judgements—for each of seven individual tasks, participants indicate whether the child answered correctly, incorrectly, or whether they are unsure (specific judgement in Karst et al.’s, 2018 [43] typology); (b) a performance-category assignment—the child’s overall mathematical competence level (below-average, average, above-average; global judgement); and (c) a prescriptive support judgement—whether the child requires support at a higher or lower performance level (global, forward-oriented judgement). Each participant response is coded as correct (1) or incorrect (0) by comparison with expert ratings and the child’s verified MARKO-D1+ test performance, yielding a child-specific accuracy score (range 0–3 across the three non-task-specific items, with mathematical competence scored across seven sub-items). Expert validation of the instrument confirmed substantial inter-rater reliability (κ = 0.83) and a strong correlation between the expert-rated vignette scores and the children’s actual MARKO-D1+ results (r = 0.94, p < 0.001), supporting both the reliability and criterion validity of the instrument [42]. Following Schrader and Helmke (1987) and Karst et al. (2018), the three judgement types map onto different accuracy components: the task-specific items primarily engage the level and differentiation components (how accurately the participant calibrates the child’s item-by-item performance), while the performance-category assignment addresses the rank component (correct ordering across performance groups), and the prescriptive support judgement captures absolute level accuracy (correct identification of the child’s developmental zone for intervention) [14,43]. Consistent with the documented performance-level dependence of judgement accuracy [16]—the three judgement targets are aggregated into a single composite accuracy score per child and per performance level. This yields three child-specific diagnostic judgment accuracy scores (for below-average, average, and above-average performers) as the primary outcome variables of the present study.
Evaluation questionnaires were administered via Moodle across phases. The in-service workshop used a 13-item scale with 4-point Likert responses addressing content, format, and satisfaction. The pilot phase included a detailed module evaluation covering structure, usability, and perceived knowledge gains. Intervention phases used brief items on perceived relevance, prior knowledge sufficiency, and module engagement (4- or 5-point agreement formats).
Semi-structured group interview: A group interview with the pilot cohort (N = 4), conducted two weeks after module completion and one week after the post-test, addressed usability and workload, clarity of vignettes, perceived learning, and suggestions for improvement. The interview was audio-recorded and used exploratorily to identify salient impressions informing design refinement, without formal coding or claims to qualitative generalisability.

3.4. Analysis

All instruments were administered online and analysed using IBM SPSS (Version 31.0.0.0). Diagnostic accuracy was analysed separately for each of the three pupil profiles (below-, average-, and above-average performing). For the main pre–post comparisons between intervention and control groups, mixed ANOVAs were conducted with time (pre-test vs. post-test) as the within-subjects factor and group (intervention vs. control) as the between-subjects factor; where significant Time × Group interactions emerged, follow-up paired-samples t-tests were computed within groups. Exploratory subgroup analyses within the intervention group used pre–post t-tests to compare diagnostic accuracy changes between higher- and lower-engagement subgroups (non-skipping vs. skipping video vignettes; completed vs. incomplete tasks) and repeated-measures ANOVAs to test Time × Group interactions across study programme specialisations (inclusive education, mathematics, other). For the main ANOVA models, formal multiple-testing corrections were applied within conceptually coherent families of effects. For RQ2, this concerned the three child profiles; for RQ2a, it concerned the three study programme specialisations. Bonferroni-adjusted p-value thresholds were used to identify effects that remain robust under conservative correction. In the Results Section, we report both the conventional p-values and indicate explicitly where Time × Group interactions or subgroup effects do not persist after Bonferroni adjustment. Exploratory engagement analyses are reported as uncorrected and are used primarily to inform iterative design decisions rather than to support strong effectiveness claims.

4. Results

The results are reported in two parts corresponding to the research-informed design logic and the chronological progression of the study. Part A (Section 4.1.1 and Section 4.1.2) presents formative evidence from the exploratory workshop and pilot phases. It provides descriptive data on the initial design promise and baseline trends, directly addressing the module’s iterative refinement (RQ1). Part B (Section 4.2.1, Section 4.2.2 and Section 4.2.3) reports the empirical evaluation findings from both intervention cycles. This part addresses overall module effectiveness compared to the control group (RQ2), elaborates on differential effects by study programme (RQ2a) and behavioural engagement (RQ2b), and culminates in the comprehensive longitudinal evaluation of participant practicality across all project phases (RQ3).

4.1. Formative Phases

4.1.1. In-Service Teacher Workshop

Nine primary teachers participated in three 3 h workshop sessions that combined an introduction to pupils’ arithmetical development with guided video-based diagnostic activities.
Addressing RQ1, which asked what exploratory pre–post evidence from the workshop and pilot phases could inform the module’s iterative refinement, mean diagnostic accuracy scores increased descriptively from pre- to post-test across all three child profiles (see Table 2) for workshop participants. For the below-average-performing child, this gain was statistically significant, with scores rising from M = 1.87 to M = 2.94, t(8) = −3.10, p = 0.015. For the above-average and average child profiles, the increases did not reach significance.
Addressing RQ3, module evaluation indicated high acceptance: all nine participants selected “strongly agree” or “rather agree” for perceived usefulness and clear structure toward objectives. Open-ended responses emphasised the value of authentic student material; two participants requested stronger links between diagnostic activities and textbook tasks. These findings informed three design implications carried forward: use of real-student videos as the central diagnostic format, retention of the same test instrument for cross-context comparability, and inclusion of activities explicitly connecting diagnostic judgments to task selection.

4.1.2. Pilot

The pilot study served as a formative check of the first complete module version with four pre-service teachers (all inclusive education specialisation).
Addressing RQ1, which asked what exploratory pre–post evidence from the formative phases could inform iterative refinement, pre–post changes in diagnostic accuracy were small and statistically non-significant across all three child profiles (Table 3) and are therefore interpreted descriptively. Mean scores showed marginal or negligible change across all profiles, with no consistent directional pattern. Mean scores for the below-average-performing child decreased slightly from M = 2.57 to 2.50, while they increased slightly for the average-performing child from M = 2.14 to 2.50. For the above-average-performing child, scores rose from M = 2.75 to M = 3.00 and the standard deviation decreased from SD = 0.50 to SD = 0.00, indicating that all participants reached the maximum possible score on this profile at post-test (ceiling effect), which constrained the sensitivity of this scale to further improvement for that profile in this small pilot group.
Addressing RQ3, questionnaire responses indicated increased confidence in assigning children to competence levels for three of four participants. Students strongly agreed that the module aligned with their prior mathematics experience, that early arithmetic content was relevant to future practice (all M = 4.00), that content knowledge had been deepened (M = 3.75), and that the teaching videos were relevant for practical phases (M = 3.50). In the group interview, participants valued the immediate feedback on diagnostic decisions, while recommending clearer vignettes for ambiguous pupil responses, more directly integrated feedback, additional textbook task examples aligned with competence levels, and a concise overview of the developmental model as an orientation aid. These findings informed three revisions carried forward to the intervention cycles: integration of direct feedback into video tasks, extension of the task-and-materials section, and addition of a developmental model script.

4.2. Intervention Phases

4.2.1. First Intervention Cycle

The first larger-scale implementation used a quasi-experimental design with an intervention group (N = 88) and a control group (N = 73). Baseline characteristics for all intervention cycles are presented in Table 4. Both groups were similar in semester stage (M = 4.48 in the bachelor’s programme for both groups) and showed comparable distributions of mathematics, inclusive education, and other specialisations, with slightly higher mean age and a somewhat higher proportion of inclusive education students in the intervention group. The following results provide initial phase-level evidence addressing RQ2; the combined sample analysis in Section 4.2.3 provides the primary evidence base for this question.
Addressing RQ2, which asked whether participation in the finalised module improves diagnostic accuracy compared to a control group across the intervention cycles, the mixed ANOVA results are presented in Table 5. Overall, evidence of differential pre–post change between intervention and control groups in the first cycle appeared only for diagnostic accuracy concerning the below-average child, and this effect is tentative after Bonferroni correction. For the above-average-performing child, the Time × Group interaction was non-significant; a significant main effect of group indicated overall higher diagnostic accuracy in the intervention group across both measurement points. For the below-average-performing child, the Time × Group reached the conventional significance threshold (F(1, 155) = 4.53, p = 0.035, ηp2 = 0.028); follow-up analyses showed a significant pre–post gain in the intervention group, t(86) = −2.50, p = 0.014, but not in the control group, t(69) = 0.62, p = 0.537. However, this interaction does not meet the Bonferroni-adjusted threshold of α = 0.017 across the three pupil profiles and is therefore treated as tentative phase-level evidence. For the average-performing child, both the interaction and main effect of time were non-significant, while the group effect again indicated overall higher accuracy in the intervention group without differential change over time.
Addressing RQ2b, exploratory post hoc analyses examined the role of engagement within the intervention group (Table 6). Engagement was operationalised along two dimensions: task completion and video-vignette engagement. The incomplete-tasks subgroup comprised participants who had five or more missing responses out of 12 tasks; the remaining participants formed the completed-tasks subgroup. For video-vignette engagement, participants who completed the video-vignette section in one hour or less, against an estimated workload of approximately three hours, were classified as the skipping subgroup, a classification corroborated by their logged video-watching times indicating that videos were not viewed in full; all remaining participants formed the non-skipping subgroup. These engagement analyses are exploratory, based on uncorrected p-values without Bonferroni or similar multiple-testing adjustments, and their results are used primarily to inform iterative design decisions rather than to support strong effectiveness claims.
Among task-completion subgroups, the completed-tasks subgroup showed a significant pre–post gain only for the below-average-performing child; the incomplete-tasks subgroup showed no significant changes (though the above-average profile approached significance, p = 0.053). For video-vignette engagement, only the non-skipping subgroup showed significant pre–post increases across all three child profiles; no significant changes emerged in the skipping subgroup.
In line with the iterative refinement logic of RQ1, the findings motivated a structural design revision: mandatory completion of all video-based tasks and clearer expectations for the temporal distribution of work across the module period.

4.2.2. Second Intervention Cycle

The second intervention cycle implemented the structurally revised module with a new cohort (intervention n = 37, control n = 40). Baseline characteristics are presented in Table 4. The two groups were comparable in semester stage in the bachelor’s programme (M = 5.31 vs. 5.13) and gender distribution, with slightly higher mean age in the intervention group and similar proportions of mathematics and other specialisations.
Addressing RQ2 (Summative Effectiveness), mixed ANOVA results for the second intervention cycle are presented in Table 7. Overall, differential pre–post change between intervention and control groups in this cycle was found only for diagnostic accuracy concerning the average-performing child, and this interaction remains significant after Bonferroni correction across the three pupil profiles. For the above-average-performing child, no significant interaction or group effect was found. For the below-average-performing child, the interaction and group effect were non-significant, but a significant main effect of time (F(1, 68) = 4.88, p = 0.030, ηp2 = 0.067) indicated a general pre–post improvement across both groups, irrespective of condition. This effect does not reach the Bonferroni-adjusted threshold of α = 0.017 across the three profiles and is therefore interpreted as tentative. For the average-performing child, the ANOVA revealed a significant Time × Group interaction F(1, 68) = 6.91, p = 0.011, ηp2 = 0.092, indicating that changes over time differed between the groups. There was also a significant main effect of Time, F(1, 68) = 6.40, p = 0.014, ηp2 = 0.086, suggesting that scores changed over time overall. Both effects remain significant under Bonferroni correction within this cycle.
Follow-up paired-samples t-tests indicated that diagnostic accuracy in the intervention group did not change significantly from pre-test to post-test, t(32) = −0.08, p = 0.938 (M pre = 2.16, M_post = 2.17), while the control group showed a significant decline, t(36) = 3.39, p = 0.002 (M_pre = 2.39, M_post = 1.93). The significant interaction for the average-performing child thus reflects a pattern in which control-group participants declined in diagnostic accuracy over the measurement period, whereas intervention-group participants remained stable.

4.2.3. Combined Sample Across Intervention Cycles

The combined sample (N = 238; intervention N = 125, control N = 113) pooled data from both intervention cycles to address the research questions fully. Baseline characteristics are combined from Table 4. The intervention and control groups were similar in semester stage (M = 4.72 vs. 4.71 in the bachelor’s programme) and in the proportions of mathematics, inclusive education, and other specialisations, with a slightly higher mean age and a somewhat higher proportion of inclusive education students in the intervention group.
Addressing RQ2 (Summative Effectiveness), mixed ANOVA results for the combined sample are presented in Table 8. Overall, differential pre–post change between intervention and control groups in the combined sample was most pronounced for the below-average- and average-performing children, with the interaction effect for the below-average child remaining tentative after Bonferroni correction. For the above-average-performing child, no significant Time × Group interaction or main effect of time was found, though a small significant group effect (F(1, 225) = 4.84, p = 0.029, ηp2 = 0.092) indicated overall higher diagnostic accuracy in the intervention group across measurement points. For the below-average-performing child, time (F(1, 225) = 4.63, p = 0.032), group (F(1, 225) = 5.63, p = 0.018), and their interaction (F(1, 225) = 4.06, p = 0.045) were all significant. Taken together, this pattern indicates that diagnostic accuracy for the below-average child improved over time and that this improvement occurred primarily in the intervention group, which also showed higher accuracy overall than the control group. However, the Time × Group interaction for the below-average child does not meet the Bonferroni-adjustment threshold of α = 0.017 across the three child profiles and is therefore interpreted as exploratory evidence of a possible selective gain for this profile. Follow-up paired-samples t-tests confirmed that diagnostic accuracy increased significantly in the intervention group, t(119) = −3.06, p = 0.003 (M_pre = 2.39, M_post = 2.63), while the control group showed no reliable change, t(106) = −0.10, p = 0.925 (M_pre = 2.34, M_post = 2.35). For the average-performing child, the Time × Group interaction and group effect were significant, and both effects remained significant under the Bonferroni-adjusted threshold of α = 0.017 across the three profiles. Follow-up paired-samples t-tests showed that the intervention group’s accuracy increased descriptively but non-significantly, t(119) = −1.56, p = 0.120 (M_pre = 2.28, M_post = 2.40), while the control group exhibited a small but significant decline, t(106) = 2.07, p = 0.041 (M_pre = 2.26, M_post = 2.06), reflecting a protective effect consistent with the Cycle 2 pattern.
Addressing RQ2a, subgroup analyses by study programme specialisation (Table 9) showed the strongest Time × Group interactions for inclusive education students, with effects at or below the conventional significance threshold across all three pupil profiles and the largest effect sizes, particularly for the below-average child (F(1, 82) = 8.22, p = 0.005, ηp2 = 0.095). This interaction for the below-average child remains significant under the Bonferroni-adjusted threshold of α = 0.017 across the three specialisations. For the above-average child in the inclusive education subgroup, the interaction (F(1, 82) = 4.45, p = 0.038, ηp2 = 0.054) reaches the conventional α = 0.05 level but does not meet the Bonferroni-adjusted threshold and is therefore interpreted as tentative. In the mathematics specialisation subgroup, a significant interaction emerged only for the average-performing child (F(1, 71) = 4.03, p = 0.049, ηp2 = 0.056), which likewise does not remain significant under a conservative Bonferroni adjustment across the three specialisations and should therefore be interpreted with caution. Effects for the other two profiles were near zero. In the other-specialisation subgroup, all interactions were non-significant with near-zero effect sizes. Overall, the pattern in Table 9 therefore shows that measurable pre–post differences between intervention and control groups are concentrated in the inclusive education subgroup and, to a lesser and more tentative extent, in mathematics students for the average-performing child, while no clear effects appear in the other programme tracks.
Addressing RQ2b, log data and self-reported engagement were analysed for 113 intervention-group participants. Self-reported engagement was high (M = 3.31–3.64 on a four-point scale; >90% selected “rather agree” or “strongly agree” on all four engagement items). The skipping indicator was not significantly correlated with any self-reported engagement measure (−0.13 ≤ r ≤ 0.09, all p > 0.25), whereas engagement items were moderately intercorrelated (e.g., r = 0.56 between investing time in the module overall and investing time in interactive videos, p < 0.001). Taken together with the subgroup pre–post results in Table 6, in line with RQ2b, these findings indicate that behavioural engagement in the form of completing tasks and not skipping video vignettes is associated with diagnostic accuracy gains within the intervention group, whereas self-reported engagement is uniformly high and only weakly related to these behavioural indicators.
Addressing RQ3, perceived learning evaluations were consistently positive: students agreed that prior knowledge was sufficient (M = 3.36, SD = 0.73), that the module deepened content knowledge (M = 3.73, SD = 0.55), that it connected well to prior mathematics experience (M = 3.59, SD = 0.51), and that both the early-arithmetic focus (M = 3.84, SD = 0.39) and work with competence levels (M = 3.87, SD = 0.39) were highly relevant for future practice. Subgroup analyses showed no significant differences across specialisations for most items, except perceived sufficiency of prior knowledge, where inclusive education students reported higher sufficiency than mathematics and other-specialisation students, F(2, 111) = 6.48, p = 0.002. The only significant association between behavioural engagement and evaluation items was a negative correlation between belonging to the skipping subgroup and perceived relevance of competence levels for future professional life, r = −0.22, p = 0.044, corroborated by an ANOVA, F(1, 81) = 4.18, p = 0.044.
Taken together, the RQ3 results across all four project phases show consistently high acceptance ratings: workshop participants unanimously selected “strongly agree” or “rather agree” for perceived usefulness and clear structure; pilot participants reported increased confidence in assigning children to competence levels; and intervention cycle participants reported high perceived relevance, deepened content knowledge, and strong alignment with future professional practice across all evaluation items. Subgroup analyses in the intervention cycles showed no significant differences across study programme specialisations for most items, with the exception of perceived sufficiency of prior knowledge. Open-ended and interview responses across the formative phases identified recurring recommendations regarding vignette clarity and stronger links to instructional materials, which were addressed through iterative design revisions.

5. Discussion

5.1. Effectiveness

RQ1 asked what exploratory evidence regarding pre–post gains in diagnostic accuracy could be derived from the initial workshop and pilot phases. In the in-service workshop, a significant pre–post gain emerged for the below-average-performing child, whereas changes for the average- and above-average-performing children were smaller and non-significant in this small sample. This pattern is consistent with previous work that showed that teachers’ diagnostic judgements are least accurate for pupils at the lower and upper ends of the performance distribution [16,17]. In the subsequent pilot with four pre-service teachers, pre–post changes were small and inconsistent across profiles and are therefore interpreted descriptively. Taken together, these formative findings suggest that authentic video-based diagnostic training holds initial promise for fostering judgment accuracy, particularly for challenging pupil profiles, in line with meta-analytic and subject-specific evidence that simulation-based training can improve diagnostic activities and accuracy in higher education [12,19,20].
For RQ2, the primary indicators are the Time × Group interactions in the mixed ANOVAs, which test whether diagnostic accuracy changed differently over time for intervention and control groups. Across development phases, the clearest and most robust interaction effects emerged for the average-performing child: in the second intervention cycle and in the combined sample, the Time × Group interaction remained significant under Bonferroni correction across the three pupil profiles, reflecting a pattern in which the intervention group’s accuracy stayed stable or improved slightly, whereas the control group’s accuracy declined. For the below-average-performing child, significant Time × Group interaction emerged in the first intervention cycle, and the combined sample showed significant effects of time, group, and their interaction, with growth confined to the intervention group. However, the interaction does not meet the Bonferroni-adjusted threshold of α = 0.017 across the three profiles and is therefore interpreted as tentative. The second cycle did not show a corresponding interaction. For the above-average-performing child, Time × Group interactions were non-significant in both cycles and in the combined sample. Taken together, these findings suggest that module participation was associated with differential pre–post change primarily at the average performance level, where Time × Group interactions remained significant after Bonferroni correction, with repeated but statistically less robust indications of a gain for the below-average child and little evidence of change for the above-average child. Although the Time × Group interaction for the below-average child does not meet the Bonferroni-adjusted threshold, the recurrence of conventionally significant effects across the first intervention cycle and the combined sample indicates a consistent pattern within the investigated setting, despite the structural revision and compositional differences between cohorts. Together, these align with and extend prior evidence suggesting that video- and simulation-based formats can foster diagnostic accuracy in higher teacher education [12,19,20] and are broadly consistent with the performance-level dependence of diagnostic judgement accuracy established in the theoretical framework [16,17].
Two measurement-related caveats qualify these conclusions. The composite score per child profile aggregates three judgement facets—task-level, categorical, and prescriptive—that are theoretically distinct but not psychometrically separated in this instrument, so the study cannot determine which facet drives the observed effects [43]. Additionally, diagnostic accuracy captures only the outcome of diagnostic activity; process-oriented operationalisations would be needed to examine the reasoning underlying these judgments in future work [44].

5.2. Differential Effectiveness

Addressing RQ2a, subgroup analyses revealed a theoretically meaningful gradient in module effectiveness across study programmes, albeit with limited robustness under Bonferroni adjustment. Pre-service teachers specialising in inclusive education showed Time × Group interactions for the below-average-performing child and additional effects at the above-average and average profiles that are weaker and partly tentative, with the largest effect sizes, particularly for the below-average child. The mathematics specialisation subgroup, by contrast, showed a significant interaction only for the average-performing child, which did not remain significant under Bonferroni adjustment and is interpreted cautiously. Students from other primary education tracks showed no significant interactions. This gradient is consistent with the role of practice-related knowledge orientation as a moderator of diagnostic learning gains [3,12]: students in programmes foregrounding developmental heterogeneity tended to show larger gains, indicating an association between programme orientation and uptake of the model-based feedback in this sample. A compositional difference between cycles is relevant here: the first cohort included all three programme tracks, while the second comprised only the mathematics specialisation and other tracks. Given the strong inclusive education effects, their absence in the second cycle likely contributes to the shift in the significant interaction from the below-average to the average-performing child across cycles. The combined sample therefore provides the most complete picture across performance profiles and study programmes.
Addressing RQ2b, log-data analyses indicated that actual engagement with the video-vignette component, rather than module completion as such, seems to be closely associated with learning gains. In the first intervention cycle, only non-skipping participants showed significant pre–post gains across all three child profiles, whereas skipping participants showed no significant changes. In the combined evaluation sample, self-reported engagement was uniformly high and uncorrelated with the skipping indicator, revealing a marked discrepancy between perceived and behavioural engagement. The only significant link between log data and self-report was that non-skippers rated the relevance of competence-level work higher than skippers, suggesting that genuine interaction with the video tasks also shapes perceived professional value. This discrepancy suggests that self-report alone was, in this sample, an insufficient proxy for learning activity in asynchronous formats and points to the need for embedded monitoring mechanisms in future module versions.
A further limitation applies to both subgroup analyses: subgroup samples are too small for multivariate modelling that could disentangle programme orientation from prior content knowledge, and the evaluation instrument was self-developed without published psychometric validation. These constraints limit the precision and generalisability of the observed subgroup differences and underline that the gradients reported for RQ2a and RQ2b should be interpreted as exploratory patterns rather than definitive causal effects.

5.3. Iterative Refinement

Addressing RQ1 (formative effectiveness), the formative phases provide exploratory evidence on pre–post gains in diagnostic accuracy that informed subsequent module refinement. The workshop produced a significant pre–post gain for the below-average-performing child, and pilot participants reported increased confidence in assigning children to competence levels. While these findings cannot be interpreted as effectiveness evidence given the absence of a control group and the small samples involved, they showed that the model-based diagnostic training concept was workable in practice and perceived as useful by participants, providing a sufficient basis for developing and testing a fully online module in controlled intervention cycles. In this sense, RQ1 is answered here not primarily in terms of formative effectiveness, as in Section 5.1, but in terms of its contribution to iterative refinement. The formative pre–post signals across the workshop and pilot phases indicate where preliminary gains emerged and how these signals were used to shape the design of the fully online module for the intervention cycles, rather than to provide generalisable effect estimates.
The cross-cycle comparison suggests both the value and the limits of structural refinement as a design strategy. The first cycle’s engagement data showed clearly that skipping video vignettes was associated with less learning gains, motivating a targeted structural revision: mandatory completion of all video-based tasks and clearer workload expectations for the second cycle. This revision was evidence-informed and precisely targeted. However, mandatory progression rules did not reliably replicate or strengthen the effectiveness pattern of the first cycle. This suggests that structural enforcement addresses a necessary but not sufficient condition for engagement quality: learners in this context can complete tasks without engaging with them in the cognitively active way that appears to drive accuracy gains. Future refinement cycles should therefore complement structural constraints with adaptive feedback mechanisms or embedded reflection prompts that make the diagnostic demands of the video tasks more cognitively visible.

5.4. Design Principles Revisited

The four-phase trajectory generates cumulative empirical evidence that can now be evaluated against each of the three design principles guiding the module’s development, yielding an initial, context-bound design knowledge contribution from this implementation. Participant evaluations across all phases (RQ3) consistently supported the module’s practical usability and perceived relevance: across the combined evaluation sample, students from all programme tracks rated the focus on competence levels and early arithmetic as highly relevant for future professional practice, and evaluation data from the workshop and pilot fed directly into successive design revisions. In what follows, the findings for RQ1, RQ2, RQ2a, RQ2b, and RQ3 are explicitly related to each design principle.
The model-based mathematical design receives the strongest empirical support within this study. Grounding all video tasks, activities, and feedback in the cognitive-developmental model [5] created a content-specific referent that rendered diagnostic judgments trainable across contexts and participant groups. The formative phases (RQ1) confirmed model-content alignment; participants specifically valued immediate feedback tied to model levels. In the intervention cycles (RQ2), gains were most reliable for the average and below-average child—precisely where the model’s level descriptors provide their clearest differentiating power—while effects were weaker and less consistent for the above-average profiles. For comparable module development projects, our findings tentatively support the principle that a domain-specific developmental model, rather than generic diagnostic process guidance, can serve as an effective operative framework for task design and feedback under conditions similar to those investigated here.
The vertical connection of in-service and pre-service contexts is supported by the trajectory of the design process, drawing primarily on RQ1 and RQ3. The in-service workshop (RQ1) established that model-anchored video diagnostics are perceived as relevant in school practice before any pre-service implementation, providing an authentic basis for the subsequent digital module. Across the combined evaluation sample in the intervention cycles (RQ3), students from all programme tracks rated the early arithmetic focus and competence-level work as highly relevant for future professional practice, suggesting that relevance established in the workshop transferred to the pre-service context. This cross-context validation strategy—testing the concept with practitioners before scaling to pre-service cohorts—emerges as a promising design decision for comparable digital resource development in teacher education.
The guided asynchronous self-study module with structured feedback and progression received encouraging, but qualified, support that was closely tied to RQ2, RQ2a, and RQ2b. The intervention cycles (RQ2) showed that the personnel-independent digital module can be embedded in regular primary mathematics teacher education and is associated with diagnostic accuracy gains for the average- and below-average-performing child in multiple cohorts. Subgroup analyses across programme tracks (RQ2a) suggest that these gains are not uniform but appear to be moderated by programme orientation, with the strongest effects in inclusive education students and more tentative effects in the mathematics specialisation subgroup. The engagement findings are consistent with the RQ2b result that behavioural engagement, rather than nominal completion, is associated with learning gains and suggest that the format’s effectiveness depends on learners actively engaging with the video-vignette tasks rather than completing the module nominally. Structural enforcement alone proved insufficient in the second cycle. Future design iterations should therefore embed adaptive support mechanisms that respond to observable engagement behaviour, complementing structural progression requirements with active monitoring of the quality of interaction with video content.

5.5. Limitations and Conclusions

Several limitations qualify the generalisability of the findings. The quasi-experimental design with non-randomised group allocation means that pre-existing differences between intervention and control groups cannot be fully excluded, as observed in some of the ANOVAs. All cycles were conducted within a single compulsory lecture at one German university, restricting transferability to other institutional contexts, subject domains, and national teacher education systems. No follow-up measurements were conducted, so it remains unknown whether diagnostic accuracy gains persist over time or translate into changed classroom practice. In addition, diagnostic accuracy was measured with a video-based instrument that is still undergoing further validation. The ceiling effects for high-performing pupil profiles in small pre-service samples indicate restricted sensitivity under specific conditions and point to the need for continued psychometric work. These ceiling effects may signal a potential measurement validity issue, in the sense that the instrument’s item range for diagnostic judgments may not sufficiently differentiate among already high-performing raters in these contexts, so that the apparent absence of further pre–post gains on the above-average profile could partly reflect limited scale sensitivity rather than a genuine lack of learning effects. The research-informed design trajectory also did not involve systematic collaborative co-design with practitioners across all phases, which limits the ecological validity of the resource within diverse school and university contexts. Together, these design and measurement constraints mean that the answers to RQ1, RQ2, RQ2a, and RQ2b should be regarded as context-bound and preliminary.
Despite these limitations, the findings converge on a consistent picture. The study provides promising, context-bound evidence that a theoretically grounded, personnel-independent digital module can be embedded in regular primary mathematics teacher education and can foster diagnostic accuracy under the specific institutional and sample conditions investigated. The iterative four-phase trajectory illustrates that, in this setting, evidence-informed refinement in line with research-informed design principles can lead to both a practically usable digital resource and transferable design knowledge. This addresses the dual-output goal that motivated the study and offers an empirically grounded basis for comparable development efforts in primary mathematics and adjacent domains of teacher education.

Author Contributions

Conceptualization, T.R.; methodology, T.R.; formal analysis, T.R.; investigation, T.R., N.R. and A.E.; resources, T.R.; data curation, T.R.; writing—original draft preparation, T.R.; writing—review and editing, T.R., N.R. and A.E.; visualization, T.R.; supervision, A.E.; project administration, A.E.; funding acquisition, A.E. All authors have read and agreed to the published version of the manuscript.

Funding

This project is part of the “Qualitätsoffensive Lehrerbildung”, a joint initiative of the Federal Government and the Länder which aims to improve the quality of teacher training. The programme is funded by the Federal Ministry of Education and Research, grant number 01JA1816. The authors are responsible for the content of this publication.

Institutional Review Board Statement

Ethical review and approval were waived for this study due to the research involving voluntary participation by adult pre-service and in-service teachers in an educational setting, without the collection of sensitive personal data or any interventions associated with psychological or physical risk. All participants provided informed consent, and the data were analysed anonymously.

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Guerra, C.; Costa, N. Can Pedagogical Innovations Be Sustainable? One Evaluation Outlook for Research Developed in Portuguese Higher Education. Educ. Sci. 2021, 11, 725. [Google Scholar] [CrossRef] [Scilit]
  2. McKenney, S.; Reeves, T.C. Conducting Educational Design Research; Routledge: London, UK, 2012; ISBN 978-0-415-61803-8. [Google Scholar]
  3. Kramer, M.; Förtsch, C.; Boone, W.J.; Seidel, T.; Neuhaus, B.J. Investigating Pre-Service Biology Teachers’ Diagnostic Competences: Relationships between Professional Knowledge, Diagnostic Activities, and Diagnostic Accuracy. Educ. Sci. 2021, 11, 89. [Google Scholar] [CrossRef] [Scilit]
  4. Heitzmann, N.; Seidel, T.; Opitz, A.; Hetmanek, A.; Wecker, C.; Fischer, M.R.; Ufer, S.; Schmidmaier, R.; Neuhaus, B.; Siebeck, M.; et al. Facilitating Diagnostic Competences in Simulations in Higher Education: A Framework and a Research Agenda. Frontline Learn. Res. 2019, 7, 1–24. [Google Scholar] [CrossRef] [Scilit]
  5. Fritz, A.; Ehlert, A.; Leutner, D. Arithmetische Konzepte aus kognitiv-entwicklungspsychologischer Sicht. J. Für Math.-Didakt. 2018, 39, 7–41. [Google Scholar] [CrossRef] [Scilit]
  6. Korthagen, F.A.J.; Kessels, J.P.A.M. Linking Theory and Practice: Changing the Pedagogy of Teacher Education. Educ. Res. 1999, 28, 4–17. [Google Scholar] [CrossRef]
  7. Voss, T.; Wittwer, J.; Nückles, M. Kohärenz Zwischen Theorie Und Praxis Durch Fokussierung Auf Core Practices—Ein Instruktionspsychologischer Ansatz Zur Abstimmung Der Phasen Der Lehrerbildung. In Proceedings of the Profilbildung Im Lehramtsstudium; Bundesministerium für Bildung und Forschung: Berlin, Germany, 2020; pp. 123–131. [Google Scholar]
  8. Schrader, F.-W. Diagnostische Kompetenz von Lehrpersonen. Beiträge Zur Lehrerbildung 2013, 31, 154–165. [Google Scholar] [CrossRef] [Scilit]
  9. KMK. Lehrerbildung für eine Schule der Vielfalt Gemeinsame Empfehlung von Hochschulrektorenkonferenz und Kultusministerkonferenz; KMK: Bonn, Germany, 2015. [Google Scholar]
  10. European Commission. Equity in School Education in Europe: Structures, Policies and Student Performance; Eurydice Report; Publications Office of the European Union: Luxembourg, 2020. [Google Scholar]
  11. Shulman, L.S. Those Who Understand: Knowledge Growth in Teaching. Educ. Res. 1986, 15, 4. [Google Scholar] [CrossRef] [Scilit]
  12. Chernikova, O.; Heitzmann, N.; Stadler, M.; Holzberger, D.; Seidel, T.; Fischer, F. Simulation-Based Learning in Higher Education: A Meta-Analysis. Rev. Educ. Res. 2020, 90, 499–541. [Google Scholar] [CrossRef] [Scilit]
  13. Wood, D.; Bruner, J.S.; Ross, G. The Role of Tutoring in Problem Solving. J. Child Psychol. Psychiatry 1976, 17, 89–100. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Schrader, F.; Helmke, A. Diagnostische Kompetenz von Lehrern: Komponenten Und Wirkungen. Empirische Pädagogik 1987, 1, 27–52. [Google Scholar]
  15. Südkamp, A.; Kaiser, J.; Möller, J. Accuracy of Teachers’ Judgments of Students’ Academic Achievement: A Meta-Analysis. J. Educ. Psychol. 2012, 104, 743–762. [Google Scholar] [CrossRef] [Scilit]
  16. Urhahne, D.; Wijnia, L. A Review on the Accuracy of Teacher Judgments. Educ. Res. Rev. 2021, 32, 100374. [Google Scholar] [CrossRef] [Scilit]
  17. Machts, N.; Kaiser, J.; Schmidt, F.T.C.; Möller, J. Accuracy of Teachers’ Judgments of Students’ Cognitive Abilities: A Meta-Analysis. Educ. Res. Rev. 2016, 19, 85–103. [Google Scholar] [CrossRef] [Scilit]
  18. Spinath, B. Akkuratheit der Einschätzung von Schülermerkmalen durch Lehrer und das Konstrukt der diagnostischen Kompetenz: Accuracy of Teacher Judgments on Student Characteristics and the Construct of Diagnostic Competence. Z. Für Pädagogische Psychol. 2005, 19, 85–95. [Google Scholar] [CrossRef] [Scilit]
  19. Irmer, M.; Traub, D.; Kramer, M.; Förtsch, C.; Neuhaus, B.J. Scaffolding Pre-Service Biology Teachers’ Diagnostic Competences in a Video-Based Learning Environment: Measuring the Effect of Different Types of Scaffolds. Int. J. Sci. Educ. 2022, 44, 1506–1526. [Google Scholar] [CrossRef] [Scilit]
  20. Schons, C.; Obersteiner, A.; Reinhold, F.; Fischer, F.; Reiss, K. Developing a Simulation to Foster Prospective Mathematics Teachers’ Diagnostic Competencies: The Effects of Scaffolding. J. Math. Didakt. 2023, 44, 59–82. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Gaudin, C.; Chaliès, S. Video Viewing in Teacher Education and Professional Development: A Literature Review. Educ. Res. Rev. 2015, 16, 41–67. [Google Scholar] [CrossRef] [Scilit]
  22. van Es, E.; Tekkumru-Kisa, M. Leveraging the Power of Video for Teacher Learning: A Design Framework for Mathematics Teacher Educators; Brill: Leiden, The Netherlands, 2019; pp. 23–54. ISBN 978-90-04-41895-0. [Google Scholar]
  23. Giek, T.; Eichentopf, P.; Seifried, J.; Burda-Zoyke, A. The Potential of Video Vignettes to Promote Diagnostic Skills of Prospective Teachers at Vocational Schools. Vocat. Learn. 2025, 18, 22. [Google Scholar] [CrossRef] [Scilit]
  24. Grossman, P.; Compton, C.; Igra, D.; Ronfeldt, M.; Shahan, E.; Williamson, P.W. Teaching Practice: A Cross-Professional Perspective. Teach. Coll. Rec. Voice Scholarsh. Educ. 2009, 111, 2055–2100. [Google Scholar] [CrossRef] [Scilit]
  25. Means, B.; Toyama, Y.; Murphy, R.; Baki, M. The Effectiveness of Online and Blended Learning: A Meta-Analysis of the Empirical Literature. Teach. Coll. Rec. Voice Scholarsh. Educ. 2013, 115, 1–47. [Google Scholar] [CrossRef] [Scilit]
  26. Hattie, J.; Timperley, H. The Power of Feedback. Rev. Educ. Res. 2007, 77, 81–112. [Google Scholar] [CrossRef] [Scilit]
  27. Radkowitsch, A.; Sommerhoff, D.; Nickl, M.; Codreanu, E.; Ufer, S.; Seidel, T. Exploring the Diagnostic Process of Pre-Service Teachers Using a Simulation—A Latent Profile Approach. Teach. Teach. Educ. 2023, 130, 104172. [Google Scholar] [CrossRef] [Scilit]
  28. Baumert, J.; Kunter, M. Stichwort: Professionelle Kompetenz von Lehrkräften. Z. Für Erzieh. 2006, 9, 469–520. [Google Scholar] [CrossRef] [Scilit]
  29. Fritz, A.; Ehlert, A.; Balzer, L. Development of Mathematical Concepts as Basis for an Elaborated Mathematical Understanding. S. Afr. J. Child. Educ. 2013, 3, 38–67. [Google Scholar] [CrossRef] [Scilit]
  30. Fritz, A.; Ehlert, A.; Ricken, G.; Balzer, L. MARKO-D1+: Mathematik- und Rechenkonzepte bei Kindern der Ersten Klassenstufe—Diagnose; Hogrefe: Wilhelmsplatz, Göttingen, 2017. [Google Scholar]
  31. Balt, M.; Fritz, A.; Ehlert, A. Insights Into First Grade Students’ Development of Conceptual Numerical Understanding as Drawn From Progression-Based Assessments. Front. Educ. 2020, 5, 80. [Google Scholar] [CrossRef] [Scilit]
  32. Cohen, J.; Jones, N.; Wong, V.C.; Quinn, A.; Erickson, S.; Parten, H.; Suhr, M.P. Practice-Based, Online Modules for Expediting Teacher Skill Development; Annenberg Institute at Brown University: Providence, RI, USA, 2026. [Google Scholar] [CrossRef]
  33. Design Based Research Collective Design-Based Research: An Emerging Paradigm for Educational Inquiry. Educ. Res. 2003, 32, 5–8. [CrossRef] [Scilit]
  34. Nieveen, N.M. Prototyping to reach product quality. In Design Approaches and Tools in Education and Training; Kluwer: Singapore, 1999; pp. 125–135. [Google Scholar]
  35. Nieveen, N.M.; Folmer, E. Formative Evaluation in Educational Design Research. In Educational Design Research—Part A: An Introduction; Plomp, T., Nieveen, N.M., Eds.; SLO: Enschede, The Netherlands, 2013; pp. 152–169. [Google Scholar]
  36. Jameson, E. Methodology: Research-Informed Design; Cambridge Mathematics: Cambridge, UK, 2019. [Google Scholar]
  37. Oh, E.; Reeves, T.C. The Implications of the Differences between Design Research and Instructional Systems Design for Educational Technology Researchers and Practitioners. Educ. Media Int. 2010, 47, 263–275. [Google Scholar] [CrossRef] [Scilit]
  38. Stokes, D.E. Pasteur’s Quadrant: Basic Science and Technological Innovation; Brookings Institution Press: Washington, DC, USA, 1997; ISBN 978-0-8157-8178-3. [Google Scholar]
  39. Shavelson, R.J.; Phillips, D.C.; Towne, L.; Feuer, M.J. On the Science of Education Design Studies. Educ. Res. 2003, 32, 25–28. [Google Scholar] [CrossRef] [Scilit]
  40. Walter, D.; Bergmann, A.; Maibach, M.; Huethorst, L.; Reinartz, L.; Grünewald, N.; Selter, C.; Harrer, A. How Pre-Service Teachers Can Be Supported to Increase Their Diagnostic Skills in Mathematics—Design and Evaluation of a Learning Platform for University Teacher Training. Front. Educ. 2025, 10, 1510828. [Google Scholar] [CrossRef] [Scilit]
  41. Rillero, P.; Camposeco, L. The Iterative Development and Use of an Online Problem-Based Learning Module for Preservice and Inservice Teachers. Interdiscip. J. Probl.-Based Learn. 2018, 12, 7. [Google Scholar] [CrossRef] [Scilit]
  42. Wagner, L.; Ehlert, A. Diagnostic Competence of Math Teacher Students—An Important Skill in Inclusive Settings. In Inclusive Mathematics Education; Kollosche, D., Marcone, R., Knigge, M., Penteado, M.G., Skovsmose, O., Eds.; Springer: Berlin/Heidelberg, Germany, 2019; p. 18. [Google Scholar]
  43. Karst, K.; Dotzel, S.; Dickhäuser, O. Comparing Global Judgments and Specific Judgments of Teachers About Students’ Knowledge: Is the Whole the Sum of Its Parts? Teach. Teach. Educ. 2018, 76, 194–203. [Google Scholar] [CrossRef] [Scilit]
  44. Herppich, S.; Praetorius, A.-K.; Förster, N.; Glogger-Frey, I.; Karst, K.; Leutner, D.; Behrmann, L.; Böhmer, M.; Ufer, S.; Klug, J.; et al. Teachers’ Assessment Competence: Integrating Knowledge-, Process-, and Product-Oriented Approaches into a Competence-Oriented Conceptual Model. Teach. Teach. Educ. 2018, 76, 181–193. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Overview of the four phases of the process in this study.
Figure 1. Overview of the four phases of the process in this study.
Higheredu 05 00066 g001
Table 1. Samples in different phases.
Table 1. Samples in different phases.
PhaseContextSampleStudy ProgrammesGroup & DesignEvaluation
(1) In-service workshopUniversity-based for primary teachers
2018
N = 9 (8 f, 1 m)In-service primary teachersSingle group; three 3 h workshops; pre–post assessmentDA; EQ1
(2) PilotPilot self-study module
2021
N = 4 fBachelor’s pre-service, primary, inclusive education specialisationSingle group; 8-week online self-study; pre–postDA; EQ2; group interview
(3) First interventionCompulsory mathematics course + self-study module
2021–2022
N = 161 (141 f, 19 m, 1 d)Mathematics n = 39; inclusive education n = 84; other primary tracks n = 38Intervention (module) n = 88 vs. control n = 73; 8-week online; quasi-experimental pre–postDA; EQ3
(4) Second interventionSame mathematics course + refined module
2023
N = 77 (68 f, 9 m)Mathematics n = 34; other primary tracks n = 43Intervention n = 37 vs. control n = 40; 8-week online; quasi-experimental pre–postDA; EQ3
Note: DA: Diagnostic accuracy; EQ1–3: Evaluation Questionnaire 1–3. The two formative phases are marked in grey.
Table 2. Pre- and post-test scores on diagnostic competence—in-service teachers (N = 9).
Table 2. Pre- and post-test scores on diagnostic competence—in-service teachers (N = 9).
Performance LevelPre-Test MPre-Test SDPost-Test MPost-Test SDt (df)p
Above-average2.300.932.760.47−1.79 (8)0.111
Below-average1.871.042.940.10−3.10 (8)0.015 *
Average2.141.022.650.70−1.24 (8)0.251
Note. Scores range from 0 to 3. * p < 0.05 (uncorrected, exploratory).
Table 3. Pre- and post-test scores on diagnostic accuracy—pilot group (N = 4).
Table 3. Pre- and post-test scores on diagnostic accuracy—pilot group (N = 4).
Performance LevelPre-Test MPre-Test SDPost-Test MPost-Test SDt (df)p
Above-average2.750.503.000.00−1.00 (3)0.391
Below-average 2.570.392.501.000.23 (3)0.833
Average2.141.052.501.00−0.58 (3)0.604
Note. Scores range from 0 to 3 (uncorrected, exploratory).
Table 4. Baseline characteristics of intervention and control group across intervention cycles.
Table 4. Baseline characteristics of intervention and control group across intervention cycles.
CharacteristicFirst Intervention Cycle
(N = 161)
Second Intervention Cycle
(N = 77)
GroupIntervention
(n = 88)
Control
(n = 73)
Intervention
(n = 37)
Control
(n = 40)
Age,
M (SD)
M = 23.93 (SD = 4.69)M = 21.89 (SD = 2.78)M = 24.11 (SD = 6.55)M = 22.18 (SD = 5.43)
Gender,
n (%)
female 81 (92%), male 7 (8%)female 60 (82%), male 12 (17%),
diverse 1 (1%)
31 female (84%),
6 male (16%)
35 female (88%),
5 male (12%)
Study programme specialisation,
n (%)
21 Mathematics (24%),
53 Inclusive
Education (60%),
14 other (16%)
18 Mathematics (25%),
31 Inclusive
Education (42%),
24 other (33%)
17 Mathematics (46%),
20 other (54%)
17 Mathematics (43%),
23 other (57%)
Semester,
M (SD)
Bachelor’s
M = 4.48 (SD = 0.85)
Bachelor’s
M = 4.48 (SD = 0.91)
Bachelor’s
M = 5.31 (SD = 1.47)
Bachelor’s
M = 5.13 (SD = 0.94)
Table 5. Mixed ANOVAs for diagnostic accuracy—first intervention cycle (N = 161).
Table 5. Mixed ANOVAs for diagnostic accuracy—first intervention cycle (N = 161).
Performance LevelEffectdfFpηp2
Above-averageTime × Group1, 1552.650.1060.017
Time1, 1551.410.2370.009
Group1, 15510.390.002 **0.063
Below-averageTime × Group1, 1554.530.035 *0.028
Time1, 1551.450.2310.009
Group1, 1553.270.0730.021
AverageTime × Group1, 1552.060.1530.013
Time1, 1550.380.5360.002
Group1, 1559.150.003 **0.056
Note. Three mixed ANOVAs (one per pupil profile) were conducted in this cycle; Bonferroni-adjusted α = 0.017 within this three-profile family. * p < 0.05. ** p < 0.005.
Table 6. Exploratory pre–post t-tests by task completion and video engagement—first intervention cycle.
Table 6. Exploratory pre–post t-tests by task completion and video engagement—first intervention cycle.
Performance LevelTask-CompletionNt (df)pVideo-Vignette
Engagement
Nt (df)p
Above-averageCompleted tasks59−1.71 (58)0.093Non-skipping68−2.72 (67)0.008 *
Uncompleted tasks28−2.03 (27)0.053Skipping19−0.65 (18)0.524
Below-averageCompleted tasks59−2.69 (58)0.009 *Non-skipping68−2.53 (67)0.014 *
Uncompleted tasks28−0.98 (27)0.337Skipping19−0.63 (18)0.539
AverageCompleted tasks59−1.44 (58)0.155Non-skipping68−2.25 (67)0.028 *
Uncompleted tasks28−0.84 (27)0.407Skipping190.52 (18)0.607
Note. Negative t-values indicate pre–post improvement. * p < 0.05 (uncorrected, exploratory).
Table 7. Mixed ANOVAs for diagnostic accuracy—second intervention cycle (N = 77).
Table 7. Mixed ANOVAs for diagnostic accuracy—second intervention cycle (N = 77).
Performance LevelEffectdfFpηp2
Above-averageTime × Group1, 680.770.3830.011
Time1, 680.650.4220.009
Group1, 680.040.8370.001
Below-averageTime × Group1, 680.140.7130.002
Time1, 684.880.030 *0.067
Group1, 682.600.1110.037
AverageTime × Group1, 686.910.011 *0.092
Time1, 686.400.014 *0.086
Group1, 680.010.941≈0.000
Note. Three mixed ANOVAs (one per pupil profile) were conducted in this cycle; Bonferroni-adjusted α = 0.017 within this three-profile family. * p < 0.05.
Table 8. Mixed ANOVAs for diagnostic accuracy—combined sample (N = 238, intervention n = 125).
Table 8. Mixed ANOVAs for diagnostic accuracy—combined sample (N = 238, intervention n = 125).
Performance LevelEffectdfFpηp2
Above-averageTime × Group1, 2251.460.2290.006
Time1, 2252.300.1310.010
Group1, 2254.840.029 *0.021
Below-averageTime × Group1, 2254.060.045 *0.018
Time1, 2254.630.032 *0.020
Group1, 2255.630.018 *0.024
AverageTime × Group1, 2256.880.009 *0.030
Time1, 2250.450.5030.002
Group1, 2256.530.011 *0.028
Note. Three mixed ANOVAs (one per pupil profile) were conducted in this cycle; Bonferroni-adjusted α = 0.017 within this three-profile family. * p < 0.05.
Table 9. Subgroup ANOVAs and descriptive statistics by study specialisation—combined sample.
Table 9. Subgroup ANOVAs and descriptive statistics by study specialisation—combined sample.
Performance LevelSpecialisationF(1, df)pηp2GroupPre-Test MPost-Test
M
Above-averageInclusive education (N = 84)4.450.038 *0.054Intervention2.352.52
Control2.252.05
Mathematics (N = 73)0.040.8380.001Intervention2.422.52
Control2.342.40
Other specialisation (N = 81)0.010.9450.000Intervention2.312.44
Control2.172.29
Below-averageInclusive education (N = 84)8.220.005 *0.095Intervention2.252.54
Control2.532.14
Mathematics (N = 73)0.000.9610.000Intervention2.502.70
Control2.272.46
Other specialisation (N = 81)0.200.6570.003Intervention2.512.70
Control2.272.39
AverageInclusive education (N = 84)2.840.0960.035Intervention2.342.57
Control2.342.20
Mathematics (N = 73)4.030.049 *0.056Intervention2.332.47
Control2.412.14
Other specialisation (N = 81)0.180.6700.002Intervention2.152.05
Control2.091.90
Note. Three mixed ANOVAs (one per specialisation) were conducted in this cycle; Bonferroni-adjusted α = 0.017 within this three-specialisation family. * p < 0.05.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Radke, T.; Reinsdorf, N.; Ehlert, A. From Teacher Workshop to Self-Study Online Module: Developing a Video-Based Resource for Diagnostic Accuracy in Primary Mathematics Teacher Education. Trends High. Educ. 2026, 5, 66. https://doi.org/10.3390/higheredu5030066

AMA Style

Radke T, Reinsdorf N, Ehlert A. From Teacher Workshop to Self-Study Online Module: Developing a Video-Based Resource for Diagnostic Accuracy in Primary Mathematics Teacher Education. Trends in Higher Education. 2026; 5(3):66. https://doi.org/10.3390/higheredu5030066

Chicago/Turabian Style

Radke, Thea, Nicole Reinsdorf, and Antje Ehlert. 2026. "From Teacher Workshop to Self-Study Online Module: Developing a Video-Based Resource for Diagnostic Accuracy in Primary Mathematics Teacher Education" Trends in Higher Education 5, no. 3: 66. https://doi.org/10.3390/higheredu5030066

APA Style

Radke, T., Reinsdorf, N., & Ehlert, A. (2026). From Teacher Workshop to Self-Study Online Module: Developing a Video-Based Resource for Diagnostic Accuracy in Primary Mathematics Teacher Education. Trends in Higher Education, 5(3), 66. https://doi.org/10.3390/higheredu5030066

Article Metrics

Back to TopTop