1. Introduction
Episiotomy is a controlled surgical incision made in the perineal tissue during childbirth, representing an invasive intervention that disrupts maternal tissue integrity. Inadequate repair of perineal trauma following birth may result in serious complications, including infection, hematoma, severe perineal pain, dyspareunia, urinary and fecal incontinence, aesthetic concerns, and decreased quality of life [
1]. Therefore, proper repair of obstetric perineal trauma using correct techniques is critical to protecting maternal health and preventing such adverse outcomes. Episiotomy and perineal laceration repair are among the fundamental clinical competencies expected of midwives and other members of the birth team, and their appropriate performance is essential to the quality of postpartum care. Although the World Health Organization (WHO) does not recommend routine episiotomy, it emphasizes that when indicated, episiotomy and its repair should be performed by healthcare professionals with adequate training and clinical competence [
2,
3]. Accordingly, midwifery students are expected to acquire the ability to appropriately assess the need for episiotomy and to competently perform episiotomy and perineal repair before graduation.
Despite its clinical importance, episiotomy repair is an invasive procedure that poses significant educational challenges for students. The process of acquiring this skill may lead to emotional and psychological difficulties, particularly anxiety and fluctuations in self-efficacy perceptions [
4]. High levels of anxiety among students may result in reduced confidence, fear of making mistakes in clinical settings, and consequently impaired learning processes. Previous studies have shown that students experiencing intense stress and anxiety during clinical practice demonstrate lower learning performance, reduced self-confidence and self-efficacy, and diminished capacity for rapid and effective decision-making [
5]. These findings highlight the necessity of employing appropriate educational strategies and providing supportive learning environments in skills-based education.
From an educational perspective, enabling students to practice in a safe environment without excessive cognitive load is essential for effective learning. Traditional clinical education, which involves performing procedures on real patients under instructor supervision, may generate stress due to the feeling of being constantly evaluated. In contrast, simulation-based environments allow students to practice without fear of harming patients, thereby reducing anxiety and extraneous cognitive load while promoting more efficient learning [
6]. Such environments encourage learners to engage more openly with the learning process, fostering motivation and a willingness to learn from errors.
Simulation-based education is widely used in clinical skills training because it provides a safe and controlled environment in which theoretical knowledge can be translated into practical competence. Previous studies have shown that simulation-based training can enhance students’ self-efficacy, confidence, and technical skills while reducing anxiety [
7,
8,
9]. For example, Demirel et al. (2020) reported significant improvements in self-efficacy and reductions in anxiety among midwifery students following episiotomy repair simulation training, while Srisukho et al. (2025) demonstrated improved surgical skills and confidence after obstetric suturing training using animal tissue models [
8,
9].
The degree of realism of simulation materials plays a crucial role in enriching the learning experience and facilitating skill acquisition. Among low-cost simulation models, sponge models are commonly used as an introductory tool for learning incision and basic suturing techniques in episiotomy repair [
10]. Although sponge models are accessible and inexpensive, they do not fully replicate anatomical tissue characteristics, which limits their realism. Consequently, animal tissue models have been incorporated into training to provide more realistic tactile and visual experiences. Several studies have documented the use of chicken breast and calf tongue in episiotomy and perineal laceration repair training, highlighting their structural similarity to human perineal tissue and their effectiveness in simulating incision and suturing procedures. Repetitive practice using such realistic materials has been shown to enhance skill learning and retention [
8,
11].
Beyond traditional physical models, simulation technologies have increasingly expanded into digital environments. In recent years, virtual reality (VR) has gained prominence in healthcare education by offering three-dimensional, immersive learning experiences. VR applications enable students to observe or experience episiotomy procedures within a simulated birth environment. In a qualitative study, Demir-Kaymak et al. (2024) reported that midwifery students perceived existing episiotomy training as limited in realism and opportunities for repetition, and expressed expectations that VR-based training could provide experiences more closely resembling real clinical settings [
12]. Tosun and Özkan (2025) demonstrated that interactive online episiotomy simulation significantly improved manual skills and clinical self-confidence among midwifery students who were unable to access clinical practice during the COVID-19 pandemic [
13]. These findings suggest that VR-based education, when appropriately designed, may compensate for the lack of hands-on practice.
Nevertheless, the primary limitation of VR is the absence of tactile and haptic feedback, which restricts its effectiveness in teaching psychomotor skills such as episiotomy repair when used as a standalone educational method. Therefore, VR is generally considered more suitable for supporting cognitive learning, spatial orientation, and procedural sequencing rather than hands-on skill acquisition. In contrast, high-fidelity simulation approaches that incorporate physical task trainers allow students to practice manual skills in a realistic environment. Consequently, sequential simulation-based approaches that combine physical and digital learning modalities may provide complementary educational experiences. Recent evidence suggests that such blended approaches can improve learners’ self-efficacy, engagement, and satisfaction while supporting clinical skill development.
Ensuring that midwifery students acquire episiotomy and perineal repair skills safely and effectively is essential for the quality of postpartum care after graduation. However, students often face limited opportunities to practice these skills in clinical environments due to restricted case diversity and the stress associated with performing invasive procedures on real patients. Simulation-based education in the preclinical phase therefore provides a critical opportunity for skill development. Sequential simulation-based approaches that integrate hands-on simulation with digital learning tools may support different components of clinical skills education. High-fidelity simulations provide realistic clinical experiences and opportunities for repetitive practice, while digital tools such as virtual reality support cognitive learning, procedural understanding, and learner engagement. Evidence suggests that combining these approaches can enhance students’ confidence, satisfaction, and clinical skill development more effectively than using a single educational modality alone.
The rationale of the present study was to compare the cognitive and affective outcomes associated with two different educational extensions administered after standardized sponge-based episiotomy repair training. Because all participants had already completed foundational hands-on instruction and supervised practice using the sponge model, the digital comparator was intentionally selected as a standardized cognitive–visual and observational learning condition rather than as a second psychomotor practice environment. The non-interactive immersive video was intended to reinforce visualization of anatomical landmarks, procedural sequencing, suturing steps, and post-repair assessment while providing identical content and viewing conditions for all participants. A fully interactive or haptically enabled VR simulation was outside the scope of the present design, which focused on comparing an additional active tactile practice package with an additional immersive observational learning package. The educational premise underlying this comparison was that the two extensions might support different components of learning after a common foundational practice phase. Additional chicken-tissue practice was expected to provide sensorimotor rehearsal, tactile familiarity, mastery experience, and corrective instructor feedback, whereas immersive video observation was expected to reinforce anatomical visualization, procedural sequencing, and cognitive organization without requiring further physical practice. Accordingly, the comparison was intended to explore whether these contrasting learning experiences were associated with different short-term patterns of self-efficacy, anxiety, and cognitive load, rather than to test the isolated effect of physical versus virtual modality. Accordingly, the digital condition did not include virtual instrument manipulation, active procedural practice, individualized feedback, or haptic input and should not be interpreted as a substitute for interactive VR-based psychomotor training. Specifically, the study compared chicken tissue–based hands-on practice with headset-delivered non-interactive immersive video instruction in terms of students’ self-efficacy, anxiety, and cognitive load.
Hypotheses
H1a. Students’ episiotomy skills self-efficacy scores will differ significantly between the baseline and post-sponge assessments.
H1b. Students’ anxiety scores will differ significantly between the baseline and post-sponge assessments.
H1c. Students’ cognitive load scores will differ significantly between the baseline and post-sponge assessments.
H2a. There will be a statistically significant difference between the chicken tissue–based training group and the VR-based video training group in episiotomy skills self-efficacy scores (ESSES) at the final assessment (T2).
H2b. There will be a statistically significant difference between the chicken tissue–based training group and the VR-based video training group in anxiety levels (STAI) at the final assessment (T2).
H2c. There will be a statistically significant difference between the chicken tissue–based training group and the VR-based video training group in cognitive load scores (DTCLS) at the final assessment (T2).
2. Materials and Methods
2.1. Study Design and Participants
This randomized comparative educational study examined short-term self-reported outcomes following two educational packages administered after standardized sponge-based training. The randomized conditions differed not only in instructional modality but also in duration, learner activity, tactile exposure, instructor contact, and feedback; therefore, the study was not designed as a time- or attention-matched comparison of modality-specific effects. The study population comprised all third-year students enrolled in the Department of Midwifery, Faculty of Health Sciences, Istanbul Atlas University, during the 2025–2026 academic year. Students were eligible if they were enrolled in the third year of the midwifery program, had completed the theoretical coursework on episiotomy and perineal repair, and provided written informed consent. Because the accessible population was limited to a single academic cohort, all 52 eligible students were invited to participate. Of these, 51 agreed to participate and completed the study. No a priori sample size calculation was performed because the sample size was determined by the number of eligible students in the available academic cohort. Thus, the study sample was not selected to provide a prespecified level of statistical power for detecting a particular between-group effect size. The randomized between-group analyses were therefore regarded as exploratory rather than confirmatory. Effect estimates and 95% confidence intervals were emphasized alongside p values, and non-significant findings were not interpreted as evidence of equivalence, non-inferiority, or absence of an intervention effect.
2.2. Randomization and Allocation Concealment
Following completion of the standardized sponge-based training and the post-sponge assessment (T1), participants were randomly allocated in a 1:1 ratio to either the chicken tissue practice group or the VR video group. Eligibility assessment, informed consent, and participant enrollment were conducted by the researchers responsible for data collection before randomization. The computer-generated simple randomization sequence was created using Randomizer.org by a member of the research team who was not involved in delivering the educational interventions. The same researcher placed the allocation assignments in sequentially numbered, identical, opaque, sealed envelopes and stored them securely until allocation. After each participant had completed the T1 assessment, the next envelope in numerical order was opened by this researcher, who assigned the participant to the indicated group and communicated the allocation to the intervention instructors. Participants, the researchers responsible for enrollment and T1 assessment, and the intervention instructors had no access to the allocation sequence before the envelope was opened. Thus, allocation concealment was maintained until completion of all pre-randomization procedures. Because of the visible nature of the educational interventions, blinding of participants and instructors after allocation was not feasible.
2.3. Inclusion and Exclusion Criteria
All participants were enrolled in the same academic year and had completed the same theoretical coursework before participation. None had previous independent experience performing episiotomy repair. Students who did not meet the eligibility criteria or declined to participate were not enrolled in the study.
2.4. Educational Interventions
All educational activities were conducted in the same skills laboratory environment and delivered by the same instructors using standardized instructional protocols. The educational interventions were delivered by two members of the research team from the Department of Midwifery. The third researcher, who generated the random allocation sequence and managed the sealed envelopes, was not involved in intervention delivery. The training process consisted of three sequential phases.
2.4.1. Standardized Sponge-Based Training (All Participants)
In the first phase, all participants received standardized theoretical and practical training on episiotomy repair using sponge models. This training was designed to provide students with foundational familiarity with episiotomy repair procedures, including identification of the incision line, selection of appropriate suture materials, layer-by-layer suturing techniques, and post-repair assessment. The training content, duration, and instructional sequence were identical for all participants. Upon completion of this phase, intermediate measurements (T1) were conducted. The standardized sponge-based training session lasted approximately 60 min and included both theoretical instruction and supervised practical application.
2.4.2. Chicken Tissue-Based Practical Training Group
Participants randomized to this group received hands-on episiotomy repair training using chicken tissue, selected for its structural and elastic similarity to human perineal tissue. Students individually performed incision and suturing procedures under instructor supervision, receiving immediate feedback throughout the session. Immediate verbal feedback was provided during practice, and students were guided to correct observed procedural errors. All procedures were conducted under standardized conditions and within predefined time limits. Following completion of this intervention, final measurements (T2) were obtained. The use of animal tissue complied with institutional biosafety regulations and was approved by the ethics committee. The chicken tissue–based practical training session lasted approximately 60 min. During this period, students individually performed incision and suturing procedures under instructor supervision and received immediate feedback.
2.4.3. Headset-Delivered Non-Interactive Immersive Video Instruction Group
Participants assigned to this group received headset-delivered non-interactive immersive video instruction through a virtual reality headset, hereafter referred to as the VR video condition. Participants viewed a single 10 min immersive instructional video. The condition was observational and non-interactive: participants did not manipulate virtual objects or instruments, perform simulated incision or suturing, receive individualized procedural feedback, or experience haptic feedback. The video demonstrated relevant anatomical landmarks, the sequence of episiotomy repair, suturing steps, and post-repair assessment. It was intended to support visual learning and procedural understanding rather than active psychomotor skill acquisition. The video content was selected following a structured quality assessment to ensure its educational appropriateness and quality. All participants viewed the same video once under standardized conditions. Final measurements (T2) were obtained immediately after completion of the VR video instruction.
2.4.4. Standardization of Interventions
To promote procedural consistency, all instructors followed a predefined instructional protocol. The training setting, core learning objectives, instructional sequence within each condition, and timing of outcome assessments were standardized. However, the randomized conditions were not matched for duration, learner activity, instructor contact, tactile exposure, or feedback. The chicken tissue condition involved approximately 60 min of supervised hands-on incision and suturing practice with immediate instructor feedback, whereas the VR condition consisted of a single 10 min non-interactive observational video without procedural practice or individualized feedback. Therefore, the randomized comparison should be interpreted as a comparison of two composite educational packages rather than as an isolated test of physical versus digital instructional modality. Any observed between-group differences could reflect the combined influence of training duration, active practice, tactile experience, instructor interaction, and feedback, in addition to the instructional medium itself. The intervention durations were not artificially equalized because each condition was implemented according to its planned educational format. The chicken tissue condition required sufficient time for individual incision, layer-by-layer suturing, knot placement, instructor observation, and immediate corrective feedback, whereas the VR video condition consisted of a fixed-duration instructional recording that was viewed once without interactive practice. Thus, the difference in duration reflected the structure and delivery requirements of the two educational packages rather than an attempt to provide equivalent instructional exposure. Nevertheless, this design does not permit the independent effect of instructional modality to be separated from the effects of educational dose, active participation, instructor contact, feedback, and tactile experience.
2.5. Ethical Considerations
Ethical approval was obtained from the Non-Interventional Scientific Research Ethics Committee of Istanbul Atlas University (Decision No: 31, Meeting No: 11, Date: 22 December 2025). The study was conducted in accordance with the Declaration of Helsinki. Written informed consent was obtained from all participants, and confidentiality and data security were strictly maintained. Data were used exclusively for scientific purposes.
2.6. Data Collection Tools
The outcomes evaluated in this study were self-reported psychological and educational outcomes, namely anxiety, perceived episiotomy skills self-efficacy, and cognitive load. No objective performance-based assessment, such as an Objective Structured Clinical Examination (OSCE), procedural checklist, expert-rated suturing score, or assessment of repair accuracy and quality, was administered. Therefore, the study was designed to evaluate learners’ perceived and affective responses to the educational interventions rather than their objectively demonstrated psychomotor competence.
2.6.1. Descriptive Information Form
The Descriptive Information Form was developed by the researchers based on the relevant literature to collect sociodemographic and educational characteristics of the participants. The form includes items related to age, perceived income level, place of residence, prior observation of episiotomy procedures and prior experience with episiotomy application. The form was administered to describe the baseline characteristics of the participants and to examine potential group comparability before and after the interventions.
2.6.2. State–Trait Anxiety Inventory (STAI)
The State–Trait Anxiety Inventory (STAI) was developed by Spielberger et al. to assess individuals’ state and trait anxiety levels [
14]. The Turkish adaptation and standardization of the inventory were conducted by Öner and Le Compte (1983) [
15]. The inventory consists of 40 items and includes two subscales: State Anxiety (STAI-S) and Trait Anxiety (STAI-T), each comprising 20 items. Items are rated on a 4-point Likert scale, and both direct and reverse-scored items are included [
16]. Total scores for each subscale range from 20 to 80, with higher scores indicating higher levels of anxiety. Previous studies have reported Cronbach’s alpha coefficients ranging from 0.83 to 0.87 for the state anxiety subscale and from 0.86 to 0.92 for the trait anxiety subscale in the Turkish version. In the present study, the internal consistency of the STAI was assessed separately for each subscale. Cronbach’s alpha coefficients for the State Anxiety subscale (STAI-S) were 0.93 at baseline, 0.84 following the sponge-based training, and 0.93 after completion of the randomized interventions (T2). For the Trait Anxiety subscale (STAI-T), Cronbach’s alpha coefficients were 0.87 at baseline, 0.85 following the sponge-based training, and 0.89 at T2.
2.6.3. Episiotomy Skills Self-Efficacy Scale (ESSES)
The Episiotomy Skills Self-Efficacy Scale (ESSES) was developed by Hadımlı et al. to evaluate midwifery students’ self-efficacy perceptions related to episiotomy application and repair [
4]. The scale was developed and psychometrically tested among third- and fourth-year midwifery students. The ESSES consists of 19 items structured into two subdimensions: Preparation and Application of Episiotomy (11 items) and Episiotomy Repair and Control (8 items). Items are rated on a 4-point Likert scale ranging from 1 (“Strongly disagree”) to 4 (“Strongly agree”). The total score ranges from 19 to 76, with higher scores indicating higher levels of self-efficacy related to episiotomy skills. No reverse-scored items are included. The original study reported a high internal consistency for the total scale (Cronbach’s alpha = 0.97). The ESSES is specifically designed for use in midwifery education contexts and is appropriate for evaluating skill-related self-efficacy following episiotomy training. In this study, the Episiotomy Skills Self-Efficacy Scale demonstrated excellent internal consistency, with Cronbach’s alpha coefficients of 0.98 at baseline, 0.96 following the sponge-based training, and 0.96 after completion of the randomized interventions (T2).
2.6.4. Different Types of Cognitive Load Scale
The Different Types of Cognitive Load Scale was developed by Leppink et al. to assess cognitive load within the framework of cognitive load theory [
17]. The Turkish validity and reliability study was conducted by Türel and Alpsülün (2025) [
18]. The scale consists of 10 items rated from 1 (“very low”) to 10 (“very high”) and includes three dimensions: intrinsic cognitive load, extraneous cognitive load, and germane cognitive load. Higher scores indicate higher perceived cognitive load. Although the scale permits separate assessment of the three cognitive load dimensions, the overall scale score was used in the present study in accordance with the Turkish adaptation study, which reported evidence of validity and high internal consistency for both the overall scale and its subdimensions. The Turkish adaptation reported a Cronbach’s alpha coefficient of 0.93 for the overall scale. In the present study, Cronbach’s alpha coefficients for the overall cognitive load score were 0.87 at baseline, 0.72 following sponge-based training, and 0.74 following the randomized interventions.
2.7. Data Collection
Figure 1 illustrates the data collection process and study flow of the research. Data collection was conducted during the 2025–2026 fall semester, between 23 December 2025 and 10 January 2026. Baseline assessments (T0) were performed prior to the standardized episiotomy repair training using a sponge model. Following the intermediate assessment (T1), conducted immediately after completion of the standardized training on the same day, participants were randomized into two groups: chicken tissue–based practical training and VR-based video training. Approximately one week after randomization, participants received the assigned intervention. Final assessments (T2) were carried out immediately after completion of the respective randomized interventions.
2.8. Study Registration
The study was retrospectively registered at ClinicalTrials.gov on 12 May 2026 (NCT07582302), after completion of participant recruitment and data collection. Retrospective registration occurred because the investigators initially classified the study as an educational intervention rather than as a trial requiring prospective registration. The measured outcomes—episiotomy skills self-efficacy, state and trait anxiety, and cognitive load—and the three assessment points were determined before data collection and were not selected on the basis of the observed results. However, the detailed statistical analysis strategy, including the final ANCOVA specifications, covariate adjustment, interaction testing, conditional comparisons, and multiplicity corrections, was not prospectively registered in a publicly accessible protocol or statistical analysis plan. This distinction should be considered when interpreting the analyses, particularly the exploratory interaction and conditional findings.
2.9. Statistical Analysis
Statistical analyses were performed using IBM SPSS Statistics version 28.0 (IBM Corp., Armonk, NY, USA). Continuous variables were summarized using means and standard deviations or medians and ranges, as appropriate, and categorical variables were presented as frequencies and percentages. Absolute standardized mean differences (SMDs) were calculated to describe the balance of participant characteristics and outcome scores between the randomized groups before the group-specific interventions. For continuous variables, the difference between group means was divided by the pooled standard deviation. For binary variables, the difference between group proportions was standardized using the pooled Bernoulli variance. SMDs were used descriptively rather than as significance tests, with values closer to zero indicating greater balance. Because randomization occurred after the post-sponge assessment, the T1 scores represented the immediate pre-intervention values for the randomized comparison. Distributional characteristics were evaluated using histograms, Q–Q plots, and the Shapiro–Wilk test. Changes from baseline to the post-sponge assessment were examined using paired-samples t-tests. Cohen’s d values were calculated as effect-size estimates. Holm adjustment was applied across the four baseline-to-post-sponge outcome comparisons defined for the present analysis. Final outcome scores were analyzed using analysis of covariance models. In each model, the final outcome score was entered as the dependent variable, training group and prior episiotomy observation experience were entered as fixed factors, and the corresponding post-sponge score was entered as a covariate. The homogeneity of regression slopes assumption was examined using group-by-covariate and prior-observation-by-covariate interaction terms. When the group-by-covariate interaction was not statistically significant, a standard ANCOVA model containing the main effects was used. When this interaction was statistically significant, the interaction term was retained, and conditional group comparisons were performed at the mean and at one standard deviation below and above the mean of the corresponding post-sponge score. Adjusted means adjusted mean differences, 95% confidence intervals, F statistics, and partial eta-squared values were reported. Levene’s test, standardized residuals, studentized residuals, Cook’s distances, and leverage values were examined as model diagnostics. Although some residual distributions departed from normality, no observations were excluded because diagnostic statistics did not indicate excessive leverage or a Cook’s distance greater than 1. Holm adjustment was applied across the four principal outcome-level tests included in the present analysis. Conditional comparisons arising from significant group-by-covariate interactions were additionally adjusted across the three examined covariate levels within each outcome and were treated as exploratory. All tests were two-tailed, and statistical significance was set at p < 0.05.
3. Results
A total of 51 students completed all study assessments. No protocol deviations occurred after randomization. All participants received the intervention to which they were allocated, completed the post-intervention assessment, and were included in the analysis according to their randomized group. The mean age of the participants was 21.43 ± 1.59 years, with an age range of 18–27 years. Twenty-eight participants (54.9%) had previously observed an episiotomy procedure, whereas 23 (45.1%) had not. Participant characteristics and outcome scores before the randomized interventions are presented in
Table 1. The mean age was similar in the chicken tissue and VR video groups. Previous episiotomy observation experience was reported by 56.0% of participants in the chicken tissue group and 53.8% of those in the VR video group. Descriptive values for the baseline and post-sponge outcome scores are also presented in
Table 1. Absolute SMDs were 0.01 for age and 0.04 for previous episiotomy observation, indicating close balance in these characteristics. At T0, the absolute SMDs ranged from 0.19 to 0.47, with the largest difference observed for state anxiety. At the immediate pre-intervention assessment (T1), absolute SMDs ranged from 0.20 to 0.26 across the four outcomes. These values indicate that some chance differences in outcome scores remained between the groups despite randomization. Accordingly, the primary models adjusted for the corresponding T1 outcome score, as well as previous episiotomy observation experience.
Within-participant changes from baseline to the post-sponge assessment are presented in
Table 2. Episiotomy skills self-efficacy increased from 50.41 ± 17.80 at baseline to 62.22 ± 12.97 after sponge-based training, t(50) = −4.044, unadjusted
p < 0.001, Holm-adjusted
p < 0.001,
d = 0.57. Total cognitive load scores also increased from 36.98 ± 17.69 to 46.22 ± 9.60, t(50) = −3.397, unadjusted
p = 0.001, Holm-adjusted
p = 0.003,
d = 0.48. No statistically significant changes were observed in state or trait anxiety after Holm adjustment.
The adjusted analyses of final outcome scores are presented in
Table 3. The homogeneity of regression slopes assumption was met for episiotomy self-efficacy and trait anxiety. After adjustment for post-sponge self-efficacy and prior episiotomy observation experience, the adjusted mean self-efficacy score was 65.15 (95% CI [60.64, 69.66]) in the chicken tissue group and 62.22 (95% CI [57.82, 66.62]) in the VR video group. The adjusted mean difference was 2.93 points (95% CI [−3.39, 9.24]), F(1, 47) = 0.870, unadjusted
p = 0.356, Holm-adjusted
p = 0.356, partial η
2 = 0.018. For trait anxiety, the adjusted mean was 41.23 (95% CI [37.13, 45.33]) in the chicken tissue group and 47.61 (95% CI [43.60, 51.62]) in the VR video group. The adjusted mean difference was −6.38 points (95% CI [−12.12, −0.64]), F(1, 47) = 5.003, unadjusted
p = 0.030, partial η
2 = 0.096. However, this difference did not remain statistically significant after Holm adjustment across the four outcome-level tests, adjusted
p = 0.112. As shown in
Table 3, the group-by-post-sponge cognitive load interaction was statistically significant before correction, F(1, 46) = 4.447, unadjusted
p = 0.040, partial η
2 = 0.088, but not after Holm adjustment, adjusted
p = 0.112. Similarly, the group-by-post-sponge state anxiety interaction was significant before correction, F(1, 46) = 5.120, unadjusted
p = 0.028, partial η
2 = 0.100, but not after Holm adjustment, adjusted
p = 0.112.
The conditional group comparisons were treated as exploratory and are presented in
Table 4. At a low post-sponge cognitive load score, the adjusted final cognitive load score was nominally higher in the chicken tissue group than in the VR video group. However, this comparison did not remain statistically significant after adjustment across the three cognitive load comparisons. No statistically significant group differences were observed at the mean or high post-sponge cognitive load levels. At a high post-sponge state anxiety level, the chicken tissue group had a lower adjusted final state anxiety score than the VR video group. This conditional comparison remained statistically significant after adjustment across the three state anxiety comparisons, within-outcome Holm-adjusted
p = 0.036. Nevertheless, the finding was interpreted as exploratory because the corresponding outcome-level interaction did not remain statistically significant after Holm adjustment across the four outcome-level tests.
4. Discussion
Episiotomy repair is an invasive and technically demanding procedure that requires manual dexterity, procedural knowledge, and accurate performance. Accordingly, simulation-based education may provide students with opportunities to become familiar with the procedural steps of episiotomy repair before performing the skill in clinical settings. The present study examined changes in midwifery students’ perceived self-efficacy, anxiety, and cognitive load during a sequential simulation-based training process. Following standardized sponge-based training, students reported higher self-efficacy of episiotomy skills and higher total cognitive load scores, whereas state and trait anxiety did not change significantly. Because all participants received the sponge-based instruction and no concurrent control or alternative-training group was included during this phase, the baseline-to-post-sponge comparisons represent uncontrolled within-participant changes. The study therefore cannot establish that the observed changes were caused specifically by the sponge model. Other factors, including theoretical instruction, instructor attention, supervised practice, repeated completion of the measures, expectancy or demand effects, and the passage of time between assessments, may also have contributed. Accordingly, these findings should be interpreted descriptively as changes observed across the foundational training period rather than as estimates of the causal effectiveness of sponge-based simulation.
The higher self-efficacy scores observed at the post-sponge assessment are consistent with previous evidence indicating that simulation-based education may support students’ confidence and perceived preparedness by providing structured, controlled, and repeatable learning experiences. Previous studies have reported that simulation supports confidence, self-efficacy, and readiness for clinical practice by allowing learners to become familiar with invasive procedures in a safe educational environment [
8,
19,
20,
21]. Similarly, studies involving episiotomy and perineal repair training have found improvements in students’ perceived confidence and preparedness following simulation-based instruction [
22,
23]. These findings are compatible with the potential value of low-cost simulation materials as an introductory component of episiotomy repair education, but they do not isolate the effect of the sponge model from the other elements of the instructional session. However, the ESSES measures students’ perceived capability rather than their objectively demonstrated technical performance. Self-efficacy is educationally relevant because it may influence motivation, willingness to practice, and readiness to engage in clinical learning; nevertheless, perceived confidence does not necessarily correspond to procedural competence. Accordingly, even a moderate change in self-efficacy may be meaningful for learner engagement, but it should not be interpreted as a clinically meaningful improvement in episiotomy repair performance. The observed increase in self-efficacy therefore cannot be interpreted as evidence that students produced anatomically accurate repairs, selected and placed sutures correctly, maintained appropriate tissue approximation, or completed the procedure safely. Objective measures such as expert-rated procedural checklists, suturing-quality scores, error counts, completion time, global rating scales, or an Objective Structured Clinical Examination would have been required to determine whether the training improved psychomotor performance.
Total cognitive load scores were also higher at the post-sponge assessment. This finding suggests that students experienced greater mental demand across the foundational instructional period; however, in the absence of a control group, the change cannot be attributed specifically to the sponge material itself. Such an increase may be expected when learners encounter a complex and unfamiliar clinical skill that requires simultaneous attention to anatomical landmarks, suture selection, procedural sequencing, and layer-by-layer repair. Nevertheless, the present study analyzed the total cognitive load score rather than the intrinsic, extraneous, and germane cognitive load dimensions separately. Consequently, the increase cannot be attributed specifically to productive cognitive engagement, schema construction, or ineffective instructional design. It should therefore be interpreted more cautiously as an increase in the overall mental demands experienced during the training process.
State and trait anxiety did not change significantly following sponge-based training. Thus, hypothesis H1b was not supported. The small effect sizes observed for both anxiety outcomes indicate that the foundational training phase was not associated with a substantial increase or decrease in anxiety. However, the absence of a statistically significant change should not be taken as evidence that the learning environment was psychologically safe or that anxiety was maintained within an optimal range. Instead, it indicates that anxiety levels remained broadly similar across the two assessment points.
The post-randomization analyses did not provide robust evidence that chicken tissue–based hands-on practice or non-interactive VR-based video instruction was superior across the evaluated outcomes. Nominal group or group-by-covariate effects were observed for trait anxiety, cognitive load, and state anxiety; however, none of the four outcome-level tests remained statistically significant after Holm correction. The adjusted difference in episiotomy self-efficacy was also not statistically significant. Therefore, the findings do not establish the superiority of either advanced instructional approach. They should also not be interpreted as evidence of equivalence or non-inferiority, because the study was not designed or powered for those purposes.
The magnitude and practical meaning of the observed effects should also be considered alongside statistical significance. During the uncontrolled foundational phase, the increase in episiotomy skills self-efficacy showed a moderate effect size (d = 0.57), while the increase in total cognitive load showed a small-to-moderate effect size (d = 0.48). These findings may be educationally relevant because they suggest a noticeable change in students’ perceived preparedness and mental effort across the training period. In contrast, the effect sizes for state and trait anxiety were small (d = 0.21 and d = 0.12, respectively), indicating little short-term change in anxiety. However, because the foundational phase lacked a control group, these effect sizes cannot be attributed specifically to the sponge model.
For the randomized comparisons, the adjusted group effect for episiotomy self-efficacy was small (partial η2 = 0.018), suggesting limited educational separation between the two conditions on perceived competence within this sample. The nominal effects for trait anxiety and the group-by-covariate interactions for cognitive load and state anxiety were larger (partial η2 ranging from 0.088 to 0.100), but they did not remain statistically significant after correction for multiple testing and may be unstable because of the small sample. In addition, no established minimally important differences are available for these measures in the context of episiotomy repair education. Therefore, the educational importance of these estimates remains uncertain. Their clinical relevance cannot be established because the study did not assess suturing quality, procedural accuracy, patient outcomes, or objectively demonstrated competence.
The small fixed cohort and the absence of an a priori power calculation further limit the interpretation of the randomized comparisons. Because the sample size was determined by cohort availability rather than by a prespecified minimum detectable effect, the study may have had insufficient power to identify small or moderate differences between the educational conditions. Accordingly, statistically non-significant findings may reflect limited precision rather than the true absence of an intervention effect. The confidence intervals around several estimates also remained compatible with effects in either direction. Conversely, nominally significant findings obtained in a small sample may be unstable and sensitive to sampling variability, particularly in the context of multiple outcome testing. The between-group findings should therefore be regarded as hypothesis-generating and require confirmation in an adequately powered multicenter trial.
The retrospective registration of the study represents an additional methodological consideration. Although the outcome domains and assessment points were determined before data collection, the detailed statistical analysis strategy was not prospectively registered in a publicly accessible protocol or statistical analysis plan. Consequently, readers cannot independently verify that all model specifications, covariate adjustments, interaction tests, conditional comparisons, and multiplicity procedures were established before examination of the data. This limitation is particularly relevant to the interaction and conditional analyses, which should be regarded as exploratory and hypothesis-generating. Retrospective registration does not invalidate the findings, but it reduces protection against selective analysis and selective outcome reporting and therefore warrants cautious interpretation. Future educational intervention studies should be prospectively registered before participant enrollment and should include a dated statistical analysis plan specifying outcomes, assessment times, covariates, model structures, multiplicity procedures, and planned subgroup or interaction analyses.
An important consideration in interpreting the randomized comparison is that the two conditions represented substantially different educational exposures. The chicken tissue condition combined a longer training duration with active psychomotor rehearsal, tactile input, instructor supervision, and immediate feedback. In contrast, the VR condition provided a shorter, non-interactive observational experience without hands-on practice or individualized feedback. The longer duration of the chicken tissue condition was closely linked to the time required for active procedural performance and instructor feedback, whereas the duration of the VR condition was determined by the fixed length of the observational video. Therefore, duration, learner activity, and instructor involvement were structurally intertwined in the present design and cannot be interpreted as independent intervention components. Consequently, the study cannot determine whether any observed differences were attributable to the instructional modality itself or to differences in educational dose, active participation, feedback, or tactile experience. Similarly, the absence of robust between-group differences should not be interpreted as evidence that the two approaches are educationally equivalent, because the interventions were neither time-matched nor designed as an equivalence or non-inferiority comparison. Future studies should compare interventions with equivalent training duration, comparable levels of instructor involvement and individualized feedback, and similar opportunities for repetition and active engagement. Such standardization would help separate the effect of instructional modality from the effects of educational dose, instructor contact, and practice intensity. Interactive or haptically enabled VR systems should also be evaluated under similarly matched conditions.
The nominally lower trait anxiety score in the chicken tissue group should be interpreted particularly cautiously. Trait anxiety is generally conceptualized as a relatively stable individual characteristic and would not ordinarily be expected to change substantially following a brief educational intervention. Moreover, the observed group difference did not remain statistically significant after adjustment for multiple outcomes. It may therefore reflect sampling variability, residual individual differences, or short-term contextual responses rather than a genuine intervention-related change. Although active hands-on practice with instructor feedback may provide greater procedural familiarity than observational instruction alone, the present findings are insufficient to establish this explanation.
Exploratory conditional analyses suggested that the association between post-sponge scores and final outcomes may have differed between groups for cognitive load and state anxiety. In particular, among students with relatively high post-sponge state anxiety, the chicken tissue group demonstrated lower adjusted final state anxiety than the VR video group. One possible explanation is that supervised hands-on practice may reduce uncertainty by allowing learners to actively perform the procedure and receive immediate feedback. However, this interpretation remains tentative because the corresponding outcome-level interaction did not remain statistically significant after adjustment across the four outcomes. The conditional findings should therefore be considered hypothesis-generating and require confirmation in adequately powered studies.
The results are broadly consistent with literature demonstrating that simulation-based education can support students’ perceived confidence, engagement, and preparedness for clinical practice [
19,
20,
21,
24]. Studies focusing specifically on episiotomy and perineal repair have likewise emphasized the educational value of controlled and repeatable simulation experiences [
8,
22,
23]. Additional exposure to a more realistic physical model or an immersive observational video did not produce clearly distinguishable short-term differences in the evaluated self-reported outcomes within this sample.
Taken together, the findings suggest that episiotomy repair education should be viewed as a staged learning process in which different instructional modalities may support different aspects of learning. Sponge-based models may provide an accessible introduction to procedural steps, whereas tissue-based practice and digital resources may offer additional tactile, visual, or observational experiences.
This distinction is particularly important because episiotomy repair is fundamentally a psychomotor and technically sequenced clinical skill. Competent performance requires more than knowledge of the procedural steps or confidence in performing them; it also requires accurate tissue handling, appropriate needle positioning, correct suture selection and placement, layer-by-layer anatomical approximation, knot security, and recognition of an inadequate repair. The outcomes used in the present study cannot determine whether students achieved these competencies. Moreover, because the chicken tissue condition included active suturing practice whereas the VR condition involved observational learning, objective performance assessment would have been especially important for determining whether the interventions had different effects on technical skill. The absence of such an assessment limits the educational and clinical interpretation of the findings.
Accordingly, the educational effectiveness of these approaches cannot be determined solely through self-reported self-efficacy, anxiety, and cognitive load. Future studies should combine these learner-reported outcomes with blinded expert assessment using validated procedural checklists or global rating scales, OSCE-based performance scores, measures of suturing and tissue-approximation quality, procedural errors, completion time, skill retention, and transfer to supervised clinical practice.
Limitations
Several limitations should be considered.
First, the study included a small single-center cohort whose size was determined by the number of eligible students rather than by an a priori power analysis. Therefore, the study may have had insufficient statistical power to detect small or moderate between-group effects, and the estimated effects may be imprecise. Statistically non-significant findings should not be interpreted as demonstrating the absence of a difference, equivalence, or non-inferiority between the interventions. Conversely, nominally significant findings may be unstable because of the small sample and multiple outcome testing. The limited sample and single-institution setting also restrict the generalizability of the findings to other student populations and educational contexts. The absence of an a priori power calculation further limits the interpretation of the randomized comparisons. Because the sample size was determined by cohort availability rather than by a prespecified minimum detectable effect, the study may have had insufficient power to identify small or moderate differences between the educational conditions. Accordingly, statistically non-significant findings may reflect limited precision rather than the true absence of an intervention effect. The confidence intervals around several estimates also remained compatible with effects in either direction. Conversely, nominally significant findings obtained in a small sample may be unstable and sensitive to sampling variability, particularly in the context of multiple outcome testing. The between-group findings should therefore be regarded as hypothesis-generating and require confirmation in an adequately powered multicenter trial.
Second, the randomized conditions differed not only in instructional modality but also in duration, active participation, tactile exposure, instructor contact, opportunities for repetition, and immediate feedback. The chicken tissue condition involved approximately 60 min of supervised hands-on practice, whereas the VR condition consisted of a single 10 min non-interactive observational video. These differences constitute important design-related confounders. Consequently, any between-group differences would represent the combined effects of the complete educational packages and could not be attributed specifically to the physical or digital modality. Conversely, the absence of statistically robust group differences cannot be interpreted as evidence of equivalence between the interventions. Future trials should compare interventions with equivalent training duration and comparable levels of instructor involvement, feedback, repetition, and learner engagement. This would permit a more valid assessment of whether observed differences are attributable to the instructional modality rather than to unequal educational exposure. Although a predefined instructional protocol was used, intervention fidelity was not independently assessed using a formal adherence checklist. In addition, the precise student-to-instructor exposure and number of practice attempts were not prospectively recorded, which may limit the reproducibility of the instructional procedures.
Third, the baseline-to-post-sponge phase did not include a concurrent control or alternative-training group. Therefore, the observed within-participant changes cannot be attributed specifically to the sponge model or interpreted as causal effects of sponge-based simulation. The instructional session included several components, including theoretical explanation, supervised practice, instructor attention, and repeated assessment, any of which may have contributed to the observed changes. Testing effects, expectancy effects, and other time-related influences also cannot be excluded. A controlled factorial or multi-arm design would be required to isolate the specific contribution of the sponge material from the broader instructional experience.
Fourth, all primary outcomes were based on self-report instruments, and no objective assessment of psychomotor performance was conducted. Although self-efficacy, anxiety, and cognitive load provide relevant information about students’ perceptions and learning experiences, they do not establish technical competence. The study therefore cannot determine whether either intervention improved suturing quality, procedural accuracy, tissue handling, anatomical layer approximation, knot security, procedural completion, or error rates. The absence of an OSCE, expert-rated procedural checklist, global competence rating, or blinded assessment of the completed repair is a major limitation, particularly because episiotomy repair is primarily a psychomotor skill. Consequently, the findings should not be generalized to students’ actual clinical performance, competence, or readiness to perform episiotomy repair independently.
Fifth, total cognitive load was analyzed rather than the individual intrinsic, extraneous, and germane cognitive load dimensions. Sixth, some model residuals departed from normality, although no observations were excluded because leverage and Cook’s distance diagnostics did not indicate an excessively influential case. Seventh, the study was registered after completion of recruitment and data collection. Although the outcome domains and assessment points had been determined before data collection, the detailed statistical analysis strategy was not prospectively documented in a publicly accessible protocol or statistical analysis plan. Therefore, the temporal pre-specification of the final model structures, covariate adjustments, interaction tests, conditional comparisons, and multiplicity procedures cannot be independently verified. This limitation increases the possibility of selective analytical decisions and is particularly relevant to the exploratory interaction and conditional findings. Finally, despite the use of multiplicity adjustments, the number of outcomes and exploratory conditional comparisons warrants cautious interpretation.