4.3.3. Qualitative Analysis
In this phase, two researchers independently read the full dataset multiple times to become familiar with the content and to develop an initial sense of recurring evaluative orientations. During this phase, both researchers produced short analytic memos to record early impressions, points of ambiguity and emergent patterns, without finalizing any coding decisions. Memos were kept alongside the dataset to support traceability of interpretive decisions.
For Phase 2 (segmentation into units of meaning), the unit of analysis was a “meaning unit,” defined as the smallest excerpt that expressed a complete evaluative claim about the framework. To ensure consistency in the qualitative analysis, clear segmentation rules were applied when identifying meaning units. A meaning unit was defined as the smallest text segment expressing a single evaluative claim, judgement, or suggestion about the XR2Learn framework. Responses containing multiple distinct ideas were segmented into separate units, even when these appeared within the same sentence. Contextual or descriptive statements were kept within the same unit when they directly supported an evaluative claim. Bullet-point responses were segmented so that each item constituted a separate meaning unit, while repetitions or rephrasing of the same point were merged unless a new emphasis was introduced.
To illustrate the application of these rules, consider the following examples. A statement such as
“The framework is conceptually strong and aligns well with established instructional design models”
was treated as a single meaning unit, as it expresses one coherent evaluative judgement. In contrast, a sentence like
“The framework is clear overall, but it lacks concrete guidance for real classroom implementation”
was segmented into two meaning units, as it contains two distinct evaluations related to clarity and usability. When experts addressed different dimensions within a single sentence, for example,
“The phases are logically structured, although the terminology related to XR could be simplified”
each evaluative component was segmented separately. Bullet-point responses were segmented so that each bullet constituted an individual meaning unit when it conveyed an independent idea. Conversely, statements that reiterated the same judgement across consecutive sentences, such as comments on flexibility or adaptability, were merged into a single meaning unit unless additional nuance was introduced. These examples illustrate how segmentation decisions were guided by analytical consistency while preserving the nuance and intent of expert feedback.
Across the full dataset, 565 meaning units were identified (sub-question range: (44–175)).
For Phase 3 (open coding), meaning units were first coded inductively. The two researchers coded independently, using descriptive labels that remained close to the participants’ wording where possible (111 codes). Following initial coding, in Phase 4 (code refinement and reduction), a structured comparison process was conducted. First, the two coders met for three consensus sessions to review overlaps, resolve definitional drift and merge synonymous labels. Codes were then reduced by:
(i) Merging semantically overlapping labels. Specifically, the initial codes “Support for engaging and meaningful learning experiences” and “Strong support for performance-based XR training” both reflected positive judgements about the framework’s experiential and engagement-oriented contribution. These were therefore consolidated into the final code “Engagement and experiential learning value of the framework.” This merged code captures expert recognition of XR2Learn’s capacity to support immersive, action-oriented and meaningful learning experiences.
(ii) Eliminating idiosyncratic codes that could not be supported beyond isolated instances unless they reflected a critical design concern raised with clear rationale. For example, “Content repetition leading to potential cognitive overload” was not maintained as a separate usability issue but was incorporated into the final code “Cognitive load and redundancy management”. This allowed individual observations about repetition to be interpreted within a broader instructional design concern rather than treated in isolation.
After refinement, the analysis retained fifty-six (56) codes.
For Phase 5 (theme development), refined codes were clustered into candidate themes within each sub-question, guided by conceptual coherence (i.e., whether codes addressed the same underlying evaluative issue) and analytical usefulness (i.e., whether a theme could inform a clear Delphi statement). Theme boundaries were iteratively reviewed against the coded extracts to ensure that each theme captured a distinct pattern and remained grounded in the data. Themes were labelled to reflect both the focus of the expert evaluation (e.g., procedural clarity, XR-specificity, usability constraints) and the direction of feedback (strength, weakness or area for refinement). Across all sub-questions, thirty-seven (37) themes were identified (sub-question range: 3–7). An overview of all identified themes, organized by sub-question, is provided in
Table 6.
In the final phase, Phase 6 (question-specific synthesis and item construction), for each sub-question, the identified themes were translated into candidate Delphi statements intended for Round 2. Item construction followed explicit criteria, which are (i) fidelity to the theme (the statement had to reflect the core meaning shared across the coded extracts), (ii) measurability (the statement had to be assessable with a Likert-type response without requiring additional interpretation), (iii) specificity (statements were phrased to avoid double-barreled claims) and (iv) non-redundancy (overlapping candidate statements were merged where they captured the same evaluative proposition).
To illustrate how the qualitative themes were operationalized into Delphi items, selected examples of the mapping between
Table 6 and
Table 7 are provided. For instance, Theme 1 (Pedagogical coherence and instructional structure) was translated into Statement Q7, which focuses explicitly on the logical sequencing and step-by-step structure of the framework, thereby preserving fidelity to the shared meaning of the coded extracts while remaining directly measurable through a Likert-type response. Similarly, Theme 5 (Practical enactment and educator usability) informed Statement Q11, which isolates the issue of actionable guidance through steps, examples or templates, avoiding double-barreled claims and ensuring specificity. Theme 3 (Integration of XR affordances with pedagogy and content) was reflected in Statement Q19, which assesses the extent to which immersive and interactive affordances are pedagogically translated into learning tasks, without overlapping with items addressing general engagement or structural clarity. Concerns captured in Theme 16 (Communication, representation and cognitive accessibility) were mapped to Statement Q10, focusing solely on language and presentation accessibility for non-expert users. Finally, Theme 4 (XR epistemological depth and theoretical specificity) was operationalized in Statement Q3, which directly addresses the perceived lack of explicit XR-specific learning theories. Across all cases, candidate statements were formulated to ensure fidelity to the underlying theme, measurability through Likert scale ratings, conceptual specificity and non-redundancy across the final item set.
This process produced nineteen (19) candidate statements. These were then reviewed by the research team in a structured synthesis meeting to ensure balanced coverage across the study dimensions (validity, clarity, usability and suitability) and across the eight sub-questions, resulting in nineteen final statements for Round 2.
4.3.4. Round 2—Quantitative Analysis and Results
In the second Delphi round, eighteen (N = 18) experts completed the quantitative rating of the nineteen statements. Minor attrition from the initial panel is common in Delphi studies due to expert availability and did not affect the expertise composition of the panel. To determine when the panel had reached agreement on an item, we adopted stringent yet literature-backed criteria. An interquartile range (IQR) ≤ 1, meaning that over half of all expert ratings fell within a single point on the five-point Likert scale, was used as a primary indicator of consensus, consistent with Delphi methodology guidelines [
25]. In fact, many Delphi studies treat an IQR ≤ 1 as evidence of high consensus on 5–7-point scales [
26]. We further required a standard deviation (SD) ≤ 1.5 as a complementary dispersion criterion, as some authors suggest using an SD threshold (≈1.5) to verify low variability in expert responses [
27]. This dual cutoff (IQR ≤ 1 and SD ≤ 1.5) is in line with recent Delphi designs of similar scope, which have defined consensus with these same thresholds for five-point expert ratings [
28]. The rationale is that a narrow IQR captures a tight clustering of opinions (indicating concentrated agreement), while the SD constraint guards against any large outlying disagreements, together ensuring robust convergence of the panel. Our chosen thresholds reflect a compromise found in Delphi practice, maximizing confidence in the included framework components while remaining attainable within three rounds [
29]. Each item meeting both dispersion criteria [
27] was considered to have achieved consensus and was retained for the final framework.
For each of the nineteen statements, the following indicators were computed, namely mean, median, SD, IQR and percentage of agreement (the proportion of responses rated 4 or 5). Most items achieved relatively tight clustering around the upper end of the scale. Seventeen out of nineteen statements met the consensus thresholds, while the remaining two displayed a little more variability, mainly in those linked to XR-specific design principles and implementation feasibility. Median values hovered mostly around 4, which shows a generally strong level of endorsement, though not blind agreement.
To make the overall pattern a bit clearer, the key descriptive and consensus statistics from Round 2 are summarized below (see
Table 7) and analyzed after.
The results of Q1 demonstrate strong agreement that the framework effectively adapts established instructional design models to XR contexts. Experts recognized its ability to translate the structure and logic of models such as ASSURE and ADDIE into immersive learning scenarios, while accommodating nonlinear flows and XR-specific constraints. High central tendency (M = 4.06, Md = 4.00) and low dispersion (SD = 0.64, IQR = 0.00), together with an agreement rate of 83.33%, confirm this as a clear strength.
Similarly, Q2 shows strong consensus that the framework is grounded in clear pedagogical principles and established learning theories. The median score of 4.00, agreement rate of 88.89% and low dispersion (SD = 0.80, IQR = 0.00) demonstrate a shared view that the framework’s instructional logic is theoretically sound, even without explicitly prescribing XR-specific theories.
In contrast, Q3 did not reach consensus regarding insufficient integration of XR-specific learning theories. Ratings clustered around neutrality (M = 2.89, Md = 3.00) with higher variability (SD = 1.08, IQR = 2.00) and a low agreement rate of 44.44%. This shows that the absence of explicit XR-specific theories is not widely perceived as a clear weakness, although divergent views point to the potential value of making such theories more visible and systematically articulated in future refinements.
The results of Q4 show strong agreement that the framework clearly aligns learning objectives, XR-based activities and assessment. This indicates solid internal coherence and support for constructive alignment. The median score of 4.00, agreement rate of 83.33% and low dispersion (SD = 0.87, IQR = 0.00) confirm this as a key strength.
Similarly, Q5 reveals strong consensus that the framework defines appropriate evaluation criteria and measurable indicators for assessing learning outcomes and instructional effectiveness in XR contexts. A median of 4.00, agreement rate of 83.33% and very low dispersion (SD = 0.68, IQR = 0.00) demonstrate high convergence among experts.
Q6 shows general agreement that the framework supports evidence-informed iterative revision. Although convergence is slightly lower than in other areas (agreement rate 72.22%, SD = 0.88, IQR = 0.75), the median value of 4.00 demonstrates a positive overall evaluation. This finding highlights iterative refinement as an acknowledged strength, while also showing that this mechanism could be made more explicit in future versions.
The results of Q7 demonstrate strong agreement that the framework is logically structured and easy to follow. High central tendency (M = 4.11, Md = 4.00), a strong agreement rate of 88.89% and low dispersion (SD = 0.76, IQR = 0.75) demonstrate broad consensus that users can clearly trace the instructional flow, confirming clarity and logical sequencing as a key strength.
Q8 shows strong agreement that the terminology aligns well with XR learning contexts. The high mean score (M = 4.22), median of 4.00 and agreement rate of 77.78%, with acceptable dispersion (SD = 0.81, IQR = 1.00), support the view that the framework uses conceptually appropriate and scientifically grounded XR terminology.
Q9 demonstrates general agreement that the framework provides sufficient examples, templates and case studies to support practical understanding. A median of 4.00, agreement rate of 77.78% and low dispersion (SD = 0.99, IQR = 0.00) show that practical guidance is perceived as a strength, while also leaving room for further enrichment in future iterations.
Q10 reveals strong consensus that the language and overall presentation are accessible to educators, EdTech experts and instructional designers beyond the XR domain. The mean score of 4.11, median of 4.00 and agreement rate of 77.78%, with moderate dispersion (SD = 0.90, IQR = 1.00), confirm accessibility and clarity across diverse professional audiences as another key strength of the framework.
The results of Q11 demonstrate moderate consensus that the framework currently lacks sufficiently concrete and actionable guidance for direct implementation in XR contexts. Low central tendency (M = 2.44, Md = 2.00) and an agreement rate of 61.11%, with moderate dispersion (SD = 0.92, IQR = 1.00), point to this aspect as a perceived weakness. Experts largely agree that additional operational detail, such as clearer steps, worked examples or explicit templates, would better support practical use.
For Q12, responses reflect a mixed but slightly leaning view regarding the need for additional training or professional development. The mean score of 2.72 and median of 2.00 show general disagreement that substantial extra training is required, yet the agreement rate of 55.56% and higher dispersion (SD = 1.27, IQR = 1.00) reveal notable variability. This demonstrates that while many experts find the framework accessible, others perceive time constraints and professional development needs as potential barriers, depending on prior XR experience and institutional context.
In contrast, Q13 did not reach consensus on resource-related constraints. Neutral central tendency (M = 3.06, Md = 3.00), high dispersion (SD = 1.21, IQR = 2.00) and a low agreement rate of 33.33% show that cost and infrastructure are not widely viewed as clear limitations of the framework. Instead, feasibility appears to depend on local conditions and institutional capacity rather than representing a systematic weakness. This finding is also in line with the evidence from earlier research findings in the field of innovation adoption and the economics of education, as has been discussed in
Section 2.1 above. However, the lack of consensus also revealed another point of potential misunderstanding in the way the question has been placed. The role of an ID framework should be to extend to all necessary dimensions in order to equip designers with the knowledge of all factors and the methods they need to apply, as well as all the involved risks in this process, rather than affecting these factors and risks per se. Thus, the role of the proposed framework is not intended to somehow result in optimization of resource consumption or institutional readiness, but rather to raise the awareness of the teams involved in instructional design on the need to account for these factors in the design process.
The results of Q14 show strong agreement that the framework aligns well with existing instructional design practices. High central tendency (M = 4.06, Md = 4.00), a strong agreement rate of 83.33% and very low dispersion (SD = 0.64, IQR = 0.00) demonstrate clear consensus that XR2Learn fits established workflows without requiring major shifts in professional practice, confirming this as a key strength.
Similarly, Q15 demonstrates strong agreement regarding the framework’s adaptability across educational contexts and XR technologies. A mean of 4.06, median of 4.00 and agreement of 77.78%, with acceptable dispersion (SD = 0.87, IQR = 1.00), show that the framework can be applied flexibly across different settings, learner groups and XR modalities.
The results of Q16 further demonstrate strong consensus that the framework supports core XR learning characteristics such as spatial interaction, multimodal engagement and adaptive interaction. The median score of 4.00, agreement rate of 77.78% and low dispersion (SD = 0.71, IQR = 0.00) reinforce the framework’s alignment with key affordances of immersive learning environments.
In contrast, Q17 reveals strong agreement that the framework would benefit from explicit extensions addressing AI-based adaptivity and personalization. With a mean of 4.06, median of 4.00 and agreement rate of 77.78%, this finding highlights a perceived gap rather than a current strength, indicating the need for more clearly defined AI-driven mechanisms to support personalized XR learning design.
The results of Q18 show very strong agreement that the framework offers sufficient guidance for designing engaging and meaningful XR learning experiences. High central tendency (M = 3.94, Md = 4.00), an exceptionally high agreement rate of 94.44% and very low dispersion (SD = 0.54, IQR = 0.00) demonstrate near consensus among experts. This confirms pedagogical clarity and experiential, engagement-focused guidance as one of the most strongly endorsed strengths of the framework.
Finally, Q19 demonstrates general agreement that the framework effectively leverages XR affordances in the design of learning activities. The mean score of 3.78, median of 4.00 and agreement rate of 66.67%, with moderate dispersion (SD = 0.81, IQR = 1.00), reflect a positive overall perception. Although initially framed as a potential weakness, experts largely viewed the pedagogical use of XR affordances as a strength, while also showing that this aspect could be further clarified and strengthened in future iterations.