3.1. Research Design
This study adopted a five-wave survey design, longitudinal at the level of each school’s teacher and Year 9 populations and repeated cross-sectional at the level of individual respondents. The same instrument was administered annually over five consecutive years (2021–2025) to Year 9 student cohorts and all teaching staff at four Australian secondary schools (
Table 1), with institutional ethics clearance by the University of Adelaide Human Research Ethics Committee (H-2021-159; consent procedures are described in the Informed Consent Statement).
Surveys were administered in Term 4 of each year at all four schools within a 3-week window, so that the waves are comparable in position within the school year. Neither strand is a panel. Surveys were anonymous, and responses were not linked across waves. The student strand is a repeated cross-section of successive Year 9 cohorts; the teacher strand is a series of repeated samples drawn from substantially overlapping, but not identical, school staff populations (
Section 3.2). The estimand throughout is, therefore, change in group-level means across waves, not within-person change.
Three features made this design appropriate. The phenomenon under investigation (teachers’ and students’ perceptions of classroom technologies) is dynamic, so repeated annual measurement was intrinsic to the research questions. This involved administering the same UTAUT-based instrument to both groups in the same schools over the same period to allow their trajectories to be compared directly, and the comparative design across four schools and two device-policy regimes required measurement consistent across contexts. Two contextual facts bear on the interpretation of the later waves. The Australian Framework for Generative Artificial Intelligence in Schools (
Australian Government Department of Education, 2023), which applied to all four schools, was released between the Y3 and Y4 waves and took effect from Term 1 of 2024, so Y4 was the first wave collected under a national policy on Gen-AI (
Section 2.5). Device provision at all four schools was unchanged across the five waves. Other hypothetical school-level events that could bear on the findings were not recorded.
3.3. Survey Instrument
The survey was organised into six sections: Section A gathered demographic information; Section B asked respondents to list in free text the technologies they considered most important in their classrooms (RQ1); Section C asked about choice and voluntariness in technology use; Section D presented the four UTAUT batteries (rated 1–7: PE, EE, SI and FC); Section E covered technology and learning impacts; and Section F invited final free-text thoughts. Parallel student and teacher versions of the instrument were identical in structure and differed only in role-appropriate wording (
Table S2). This paper analyses Sections B, C and D (
Table S2). Section A was used to characterise the samples (
Section 3.2). Section E is reported in a companion study. Section F responses are not analysed here. The instrument is available from the author, subject to the conditions of the ethics approval.
Section D items were adapted from
Venkatesh et al.’s (
2003) original UTAUT scales, with wording modifications to fit the secondary-school classroom context (
Table S2). Items were rated on a seven-point Likert scale anchored on “strongly disagree” (1) and “strongly agree” (7), following
Venkatesh et al.’s (
2003) original response format. The one reverse-worded item (facilitating conditions; Venkatesh et al.’s PBC5) was reverse-scored prior to subscale computation. Subscale scores were computed as the mean of the four construct items for respondents with at least three valid items; respondents with fewer than three valid items on a construct were excluded from that construct’s analyses only. No imputation was performed. The extent and distribution of item-level missingness, and a complete-case sensitivity analysis, are reported at the start of
Section 4. Internal consistency of each UTAUT subscale was assessed at each wave using Cronbach’s alpha, with the a priori expectation that α ≥ 0.70 would indicate acceptable reliability for each subscale at each wave.
3.4. Data Analysis
The first stage of analysis was descriptive, reflecting the study’s comparative and longitudinal aims. For each UTAUT subscale, means and standard deviations were computed per school, per wave, and for teacher and student groups separately. Group-level means reported in the text and in
Table 3 are the unweighted average of the four school means, so that each school context contributes equally to the aggregate irrespective of its cohort size. The mixed-effects models (below) are estimated on respondent-level data and, therefore, weight respondents equally. The two bases differ by at most 0.02 points at any wave. Change across the five waves was examined through plots of subscale means over time, disaggregated by school and by device-policy regime. Differences between schools and between device-policy regimes were examined descriptively through effect sizes (Cohen’s
d for pairwise comparisons, η
2 for multi-school comparisons). Device-policy comparisons were computed separately for teachers and students, since the BYOD/mandated distinction governs student device choice directly and teacher device provision only indirectly. Between-school comparisons are reported descriptively; individual- and group-level comparisons are tested inferentially in the second stage below. Cronbach’s alpha was computed for each UTAUT subscale in each school–wave–group cell as the instrument-reliability check described in
Section 3.3. The Section C voluntariness item was analysed descriptively as a validity check on the device-policy contrast used to address RQ4.
Free-text responses to the Section B item were analysed through structured content analysis. Responses were segmented into individual technology nominations, normalised to consolidate spelling variants and brand-name synonyms, and coded into technology categories using a frame developed inductively from the Y1 responses and extended, with earlier categories retained, as new technologies appeared. Coding was performed by the author using a written codebook in Y1 and thereafter extended only by adding categories. Categories were developed inductively from the Y1 responses by grouping normalised nominalisations by primary classroom function. Categories were only added in later waves when nominations that fitted no existing category exceeded 10% of responses in that wave. A random 20% sample of responses stratified by wave was re-coded by the author after an interval of 10 weeks, with intra-coder agreement of κ = 0.80. Where the two passes disagreed, the disagreement was resolved by reference to the codebook definition. Because coding was not independently replicated, the category prevalences should be read as the product of a single coder’s judgement; this is recorded as a limitation in
Section 5.7. Coding was not blind to respondent group or wave. Multi-technology responses generated one nomination per technology. Ambiguous nominations were resolved by assigning the category to the corresponding function the respondent described, and nominations too vague to classify were coded ‘unspecified’ and therefore excluded. Category prevalence was computed per wave and group as the percentage of respondents nominating at least one technology in the category; these figures underpin the three-phase periodisation in
Section 4.1.
The second stage tested the study’s longitudinal contrast between teacher and student trajectories. This was achieved formally through linear mixed-effects models, estimated separately for each UTAUT construct. This approach respects the nesting of respondents within schools and provides a formal test of the wave–group interaction, a stronger warrant than a series of pairwise effect sizes. Each model specified the respondent-level subscale mean as the outcome; wave (Y1–Y5), group (teacher-vs.-student) and their interaction as fixed effects; and a random intercept for school. Seeing as respondents were not linked across waves, each wave was treated as an independent sample within school, and the wave–group interaction estimates change in the difference between the teacher and student wave means, rather than change within individuals. The two series also differ in composition. The student series compares successive Year 9 cohorts, whose members differ in prior schooling, digital experience and exposure to Gen-AI, whereas the teacher series compares heavily overlapping samples of the same staff. A rising student series is, therefore, consistent with cohort replacement, as well as with attitude change, and the interaction should be read as a divergence between two population-level series rather than as a comparison of two within-population trajectories. The wave–group interaction (the model term corresponding to the teacher–student divergence) was evaluated through likelihood-ratio tests comparing maximum-likelihood fits with and without the interaction term. Unconditional (intercept-only) models were estimated first to partition variance within and between schools via the intraclass correlation coefficient (ICC).
Two constraints follow from the four-school design: school-level characteristics were not entered as predictors, and, because random-effect variance estimates are unstable with few clusters (
McNeish & Stapleton, 2016), all models were re-estimated with school as fixed effects, with substantive conclusions unchanged. Multilevel modelling conventions follow
Raudenbush and Bryk (
2002). A final caveat concerns measurement invariance: formal invariance testing across groups and waves was not conducted, so comparisons of construct levels between teachers and students assume the adapted items functioned equivalently in both populations (
Putnick & Bornstein, 2016). This assumption is least secure for PE, whose items differ in wording between student and teacher versions, and most secure for EE, SI and FC, whose items are identical.
All analyses used lme4 (
Bates et al., 2015) and lmerTest (
Kuznetsova et al., 2017). Mixed-effects models were specified as
yij = β
0 + Σβ_w Wave_w + β_g Group + Σβ_wg (Wave_w × Group) +
uj +
eij, with Wave entered as four dummy variables (reference Y1), Group as one dummy (reference Student),
uj ~ N(0, τ
2) the school random intercept, and
eij ~ N(0, σ
2) the residual. Models were estimated by maximum likelihood so that nested models could be compared by likelihood-ratio test. ICCs were computed as τ
2/(τ
2 + σ
2) from the unconditional models. Respondents with a valid subscale score (
Section 3.3) contributed to that construct’s model. Residual diagnostics (Q–Q plots and residual-versus-fitted plots) showed no material departure from normality. Cohen’s
d was computed for each pairwise comparison as the difference in cell means divided by the pooled standard deviation of the two cells being compared. Where a mean
d across waves is reported, it is the arithmetic mean of the five wave-specific values. η
2 for the four-school comparison was computed from a one-way ANOVA on respondent-level subscale scores within each wave–group cell.