Next Article in Journal
The Influence of Mentoring on Educational Attitudes, Beliefs, and Behaviors (EABBs): A Scoping Review
Previous Article in Journal
From Online Video-Based Professional Development to Differentiated Teaching: A Case Study of Mathematics Teacher
Previous Article in Special Issue
Use of Digital Technologies to Support Socioemotional Teacher Training: A Systematic Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Rethinking Student Engagement Datasets for AI in Virtual Learning: A Narrative Review

1
College of Engineering and Technology, American University of the Middle East, Egaila 54200, Kuwait
2
KITE Research Institute, Toronto Rehabilitation Institute, University Health Network, 550 University Ave., Toronto, ON M5G 2A2, Canada
3
Lawrence Bloomberg Faculty of Nursing, University of Toronto, Toronto, ON M5T 1P8, Canada
4
Rehabilitation Sciences Institute, University of Toronto, Toronto, ON M5G 1V7, Canada
*
Author to whom correspondence should be addressed.
Educ. Sci. 2026, 16(4), 548; https://doi.org/10.3390/educsci16040548
Submission received: 20 January 2026 / Revised: 3 March 2026 / Accepted: 13 March 2026 / Published: 1 April 2026
(This article belongs to the Special Issue Beyond Classroom Walls: Exploring Virtual Learning Environments)

Abstract

Student engagement (SE) is central to learning success in virtual environments, and artificial intelligence (AI) models increasingly rely on annotated datasets to estimate it automatically. However, existing datasets vary substantially in how engagement is defined and operationalized. This narrative review examined 31 SE datasets collected in computer-based virtual learning settings, analyzing them across seven dimensions of engagement annotation: sources, modality, timing, temporal resolution, level of abstraction, combination, and quantification. Substantial heterogeneity was observed, with only 3 of the 31 datasets explicitly grounding their annotations in psychometrically validated or theoretically established engagement constructs. These inconsistencies have important implications for AI model development. For example, mismatches between the modality used for annotation and that used for model training may reduce robustness, and divergent quantification schemes (binary, ordinal, or continuous) restrict cross-dataset comparability and transferability. By explicitly linking dataset design decisions to risks affecting model validity, generalization, and interpretability, this review highlights the need for stronger construct grounding, clearer reporting of annotation protocols, and greater standardization to support more reliable and transferable engagement-aware AI systems.

1. Introduction

With widespread internet access, virtual learning has become a mainstream mode of education, supported by interactive technologies and smart education platforms (Mukhtar et al., 2020). Compared to traditional instruction, virtual learning offers greater accessibility, lower costs, and personalized, adaptive learning (Dung, 2020). However, virtual environments pose challenges in measuring student engagement (SE). In face-to-face classrooms, instructors can directly observe behavioral cues, whereas online settings make engagement assessment more difficult, particularly in large-scale courses (Sümer et al., 2021). This presents a substantial limitation, given the strong association between SE and academic achievement and learning outcomes (Gray & DiLoreto, 2016). To mitigate this challenge, smart education systems are increasingly incorporating mechanisms for real-time engagement monitoring to facilitate timely feedback and interventions that support and enhance student participation.
Researchers in educational psychology have identified SE as comprising three primary components (Fredricks et al., 2004; Trowler, 2010). These include behavioral engagement, referring to actions such as attendance, participation, and staying on task; affective engagement, involving emotional responses such as enthusiasm, interest, and enjoyment; and cognitive engagement, reflecting a student’s investment in learning and willingness to tackle challenges. The concept of engagement varies based on the analytical perspective and level of detail, often referred to as “grain size” (Sinatra et al., 2015). At the microlevel, it concerns an individual’s involvement in a specific task, while at the macrolevel, it refers to group engagement, such as that of a class or institution.
Recent advances in artificial intelligence (AI) have enabled the development of algorithms that automatically and objectively measure SE in virtual learning environments (Abedi & Khan, 2023; Buono et al., 2023; Dewan et al., 2019; S. Gupta et al., 2023a, 2023b; S. K. Gupta et al., 2019; Karimah & Hasegawa, 2022; Liu et al., 2018; Mandia et al., 2023). Most approaches rely on machine learning (Van Engelen & Hoos, 2020), which requires annotated ground-truth data for training and evaluation (Dewan et al., 2019; Karimah & Hasegawa, 2022; Mandia et al., 2023). These models are typically trained on annotated subsets and evaluated on unseen samples to assess generalizability. Video and audio are the predominant data modalities (Dewan et al., 2019; Karimah & Hasegawa, 2022), from which features such as eye gaze, head pose, valence, arousal, and tonal variations are extracted (Abedi & Khan, 2023; Dewan et al., 2019; Karimah & Hasegawa, 2022). These features, paired with engagement annotations, are used to train machine learning and deep learning models (Abedi & Khan, 2021, 2024; Dewan et al., 2019; Karimah & Hasegawa, 2022). A major challenge is the absence of a universally accepted SE definition, resulting in inconsistent annotations across datasets. This inconsistency limits the development, validation, and generalizability of AI-driven SE measurement tools. High-quality, cohesive annotations aligned with educational psychology frameworks (Fredricks et al., 2004; Trowler, 2010) are essential for developing accurate and widely applicable models. Despite the growing reliance on such datasets, there has been limited structured examination of how engagement is defined and operationalized across existing datasets and how these variations influence AI model validity and generalizability. To address this gap, the present study provides a focused narrative synthesis of dataset design and annotation practices using seven dimensions of engagement annotation.
The objective of this narrative review was to identify and examine inconsistencies in the definitions and annotation protocols used in existing SE datasets. Rather than aiming for exhaustive coverage as in a systematic review, the goal was to provide an in-depth, focused examination of methodological practices in dataset construction and annotation. The guiding research question was as follows: How inconsistent are the definitions and annotation protocols of SE in existing datasets when analyzed through the lens of the seven established dimensions of engagement annotation? These dimensions include (1) sources, referring to the observers performing the annotation; (2) data modality, which indicates the type of information observed for annotation; (3) timing, denoting when the annotation occurs; (4) temporal resolution, referring to the timesteps at which annotations are made; (5) level of abstraction, indicating whether engagement is defined and annotated as a single or multi-component construct; (6) combination, which describes how different engagement components are integrated into a single value; and (7) quantification, which refers to how engagement is represented numerically. In the present narrative review, when the included articles were unclear, dimensions were assigned based on reported annotation procedures and predefined criteria, with ambiguous cases resolved through re-examination to ensure consistent coding.
This paper is structured as follows. Section 2 examines previous reviews on SE measurement and outlines the unique contribution of this work. Section 3 describes the methodology employed in the narrative review. Section 4 presents and analyzes the key findings. Finally, Section 5 provides a discussion of the implications and outlines directions for future research. In addition, Appendix A introduces alternative SE definitions and annotation protocols suitable for virtual learning.

2. Related Reviews and Contribution

A few reviews have examined automatic SE measurement, primarily focusing on AI techniques and algorithms. Dewan et al. (2019) categorized SE measurement in online learning by participation level: automatic, semi-automatic, and manual. They emphasized computer vision methods for video-based SE due to their non-invasive nature and promising results. The review also highlighted the challenges in SE techniques, touching upon the existing datasets and the evaluation metrics used for these techniques. Karimah and Hasegawa (2022) conducted a systematic review of SE definitions, datasets, and methods in smart education. They examined dataset features and steps in engagement measurement using supervised learning, including pre-processing, model development, and evaluation. Challenges included assessing cognitive engagement, personalizing measurement, and limitations of machine and deep learning techniques. Mandia et al. (2023) reviewed automatic SE measurement in traditional and online settings, focusing on machine learning methods, data collection, annotation procedures, and evaluation metrics. Booth et al. (2023) selectively reviewed techniques for engagement measurement in educational applications, discussing the strengths and limitations of traditional and affective computing approaches. They also explored proactive and reactive strategies for enhancing engagement to improve learning outcomes. From the educational perspective, several reviews have explored the broader theory of SE, both generally (Wong & Liem, 2021) and in the context of technology-mediated learning (Aslan et al., 2014; Henrie et al., 2015; Hu & Li, 2017; Schindler et al., 2017).
Earlier reviews in the AI domain mainly focused on machine-learning and deep-learning models for SE measurement (Dewan et al., 2019; Karimah & Hasegawa, 2021, 2022; Mandia et al., 2023; Salam et al., 2022). Dewan et al. (2019) targeted virtual learning, while others (Karimah & Hasegawa, 2022; Mandia et al., 2023; Salam et al., 2022) included both virtual and traditional settings. Although prior reviews considered SE datasets, they largely overlooked inconsistencies in engagement definitions and annotation protocols. To address this gap, the present review systematically analyzes these issues through seven dimensions of engagement annotation and examines their implications for AI model validity, comparability, and cross-context applicability.

3. Methods

This study was designed as a narrative review with a transparent but non-exhaustive search strategy, rather than as a systematic review. A narrative approach was selected because the objective was to critically examine and interpret methodological diversity in engagement definitions and annotation practices, rather than to quantify effect sizes or synthesize empirical outcomes. IEEE Xplore, ACM Digital Library, SpringerLink, ScienceDirect, and Google Scholar were searched for English journals and conference publications published between 2010 and 2025. This time window was selected because AI-based engagement measurement in virtual learning expanded substantially in the early 2010s, driven by advances in computer vision and deep learning. The aim of the search was to identify relevant datasets for in-depth methodological examination, not to achieve comprehensive coverage. Different combinations of the following keywords were adopted to search the databases: engagement measurement/detection/prediction/recognition/classification/regression, machine learning, deep learning, artificial intelligence, data, and dataset.
Studies were included if they introduced an original single- or multi-modal SE dataset intended for the development of AI models for SE measurement. Eligible datasets must have been collected from students participating in online or offline computer-based virtual learning sessions. The exclusion criteria were as follows: studies that developed AI models for SE measurement without introducing an original SE dataset; studies that introduced SE datasets involving students in in-person classroom settings or group contexts; and studies that presented student datasets focused on general affect detection, such as basic affect state recognition or valence and arousal recognition, rather than engagement measurement. In this review, virtual learning contexts were defined as computer-mediated environments in which engagement was measured during interaction with digital content, including fully online and computer-based hybrid settings, while datasets that focused exclusively on in-person classroom observation without a digital interface were excluded. These criteria were applied to ensure conceptual relevance to the methodological questions of this review, rather than to enforce formal systematic screening.
Consistent with the narrative nature of this review, we did not apply PRISMA-style procedures, flow diagrams, or formal risk-of-bias appraisal. The focus of this review was not on analyzing the use of AI techniques for SE measurement. Other reviews that delve into AI, machine learning, and deep learning (Booth et al., 2023; Dewan et al., 2019; Karimah & Hasegawa, 2022; Mandia et al., 2023; Salam et al., 2022) are available for interested readers; refer to Section 2.
The included datasets were analyzed with respect to the definition and annotation of SE, using the seven dimensions of engagement annotation outlined in Figure 1 and described in Section 3.1. Additionally, eight baseline characteristics of these datasets were reviewed: the nature of the activity (whether it was computer-based lecture viewing, reading, writing, or usage of educational software); the type of activity (interactive vs. non-interactive); the setting of data collection (in-the-wild or in-the-lab); the number of students; their gender; age composition; the country where the dataset was collected and annotated; and the distribution of samples across various engagement levels within the dataset.

3.1. Dimensions of Engagement Annotation

D’Mello (2015) identified five dimensions of affect annotation for developing affect detection systems: sources, data modality, timing, temporal resolution (timescale), and level of abstraction. There are two key differences between affect detection and engagement measurement. Initially, affect represents merely a singular component of engagement, which encompasses behavioral and cognitive elements as well. In a subsequent point of differentiation, while affect pertains to “detection”, engagement “measurement” seeks to not only differentiate between engagement and disengagement but also quantify the “level” of engagement (Liao et al., 2024). Considering the differences between affect and engagement, the aforementioned five dimensions were refined and supplemented with two novel dimensions: combination and quantification (scale), aiming to address the initial and subsequent disparities highlighted above. Employing these seven dimensions, a comprehensive examination of SE definition and annotation in existing datasets was undertaken.

3.1.1. Sources

The first dimension of engagement annotation, sources, refers to the types and number of individuals performing the annotation (D’Mello, 2015). Most SE datasets rely on observer-based annotation by either expert (trained) or non-expert (untrained) observers. Experts, such as educational psychology professionals, are costly, while non-experts, often students, annotate based on personal perception. Non-expert annotations are commonly collected through crowdsourcing (Brabham, 2013), which enables fast data collection but may introduce noise due to limited expertise (A. Gupta et al., 2016). Although observer-based methods are time-consuming and prone to bias (Minner et al., 2010; Whitehill et al., 2014), they do not interrupt learning and can capture SE as it occurs (Henrie et al., 2015), allowing for detailed annotation of engagement episodes.
SE annotations can also be obtained via self-report. While concurrent self-reports may interrupt tasks (O’Brien & Cairns, 2016), retrospective reports may be biased and vary by individual interpretation. However, because SE includes cognitive and emotional aspects, self-report is argued to be a valid approach to capture students’ perceived engagement (Henrie et al., 2015).

3.1.2. Data Modality

The dimension of data modality (D’Mello, 2015) refers to the type of information observed during annotation, such as video, audio, screen recordings, or mouse trajectory data. These modalities may differ from those used to deliver educational content. For example, observers may watch videos of students in virtual sessions to annotate engagement without access to their screen recordings.

3.1.3. Timing

Annotation timing (D’Mello, 2015) refers to when the annotation is performed. It can occur in real-time, such as when students self-report engagement via pop-up prompts during virtual sessions. In observer-based settings, annotation may be done offline by reviewing recorded data (Aslan et al., 2017), or online through live observation during the session (Ocumpaugh, 2015).

3.1.4. Temporal Resolution

Temporal resolution (timescale) (D’Mello, 2015) refers to the granularity of annotation: frame-level (e.g., still images from video), segment-level (e.g., self-reports every 10 min), or session-level (e.g., retrospective self-reports at the session’s end).

3.1.5. Level of Abstraction

Regarding the level of abstraction, engagement may be annotated at a high level without distinguishing its components, for example, as simply engaged or disengaged. Alternatively, affective, behavioral, and cognitive components may be annotated separately and then combined into a numerical engagement value. In affect annotation, Porayska-Pomsta et al. (2013) refer to these approaches as discrete and dimensional responses, respectively. Each component can also be annotated at varying levels; for instance, behavioral engagement may be labeled as Off-Task or On-Task, with On-Task further divided into categories such as On-Task Conversation or On-Task Giving Answers (Ocumpaugh, 2015).

3.1.6. Combination

To prepare a dataset for AI model development, each sample requires a numerical value or a class label. The combination dimension refers to how affective, behavioral, and cognitive components of engagement are combined to produce this output. For instance, Aslan et al. (2017) defined a student as engaged if they were behaviorally On-Task and affectively Highly Motivated. Naibert and Barbera (2022) proposed architectures combining these components from self-report data. However, combining components can obscure distinctions among them, potentially losing valuable information (McNeish, 2022). Combination strategies should also account for correlations between components (D’Mello & Graesser, 2012). As noted in the level of abstraction, if engagement is directly labeled without component analysis, no combination is needed.

3.1.7. Quantification

The quantification (scale) dimension addresses how engagement is numerically represented, reflecting the type of variable used for its definition and annotation. As a construct grounded in psychology, the quantification of engagement necessitates a robust foundation that upholds objectivity, precision, and rigorous standards (Tafreshi et al., 2016). SE has been annotated as a binary variable (engaged/disengaged), a categorical variable in some datasets, and more commonly as an ordinal or interval variable representing discrete or continuous levels.

4. Results

The database search retrieved approximately 118 records. As this review was conducted narratively rather than systematically, full-text articles were reviewed for conceptual relevance to dataset collection and annotation practices. Several studies were excluded because they focused on proposing new methodological approaches for AI-driven engagement measurement without introducing an original dataset, or because the datasets involved in-person classroom settings rather than virtual learning environments. To minimize potential selection bias, predefined inclusion and exclusion criteria were applied consistently during the screening process. This process resulted in 31 datasets included for analysis. Thirty-one studies introducing original datasets for SE measurement in virtual learning were included in the review. The seven dimensions of engagement annotation across these datasets are summarized in Table 1, and their baseline characteristics are presented in Table 2, which together form the basis for the subsequent in-depth analysis of methodological inconsistencies in engagement annotation and dataset characteristics. The results, therefore, go beyond description, highlighting patterns of variability and their implications for dataset comparability and AI model development.

4.1. Inconsistencies in Sources

The reviewed datasets employed diverse annotation sources. Some relied on expert annotators, while others used non-experts or crowdsourced contributors. Several datasets used self-reported questionnaires, and a few combined self-reports with observer-based annotations. In most studies, non-expert or crowdsourced observers were untrained students or freelancers. Booth et al. (2017) noted that they often received no guidance on interpreting engagement, leading them to rely on personal, potentially inaccurate definitions. This risks flawed annotations and, in turn, AI models that learn a distorted understanding of engagement. A. Gupta et al. (2016) reported noise issues in crowdsourced annotations, and some studies applied filtering to address this (Kaur et al., 2018). Using expert annotators with clear protocols can reduce such noise. Although costly and time-consuming, precise annotation is essential for developing broadly applicable AI models. Regarding the datasets with multiple annotators, different metrics were employed to determine inter-rater reliability or correlation (Hallgren, 2012). These metrics included Cronbach’s alpha, Cohen’s kappa, Fleiss’ kappa, and Krippendorff’s alpha, all designed to assess the annotation quality across different annotators. As shown in Table 1, varying numbers of expert and non-expert annotators were involved in the reviewed datasets. Different strategies were used to handle disagreements between annotators, including majority voting, removing data samples with high disagreement, and appointing a tie-breaker annotator to make the final decision.

4.2. Inconsistencies in Data Modality

A variety of information types were used by observers for engagement annotation in the reviewed datasets, including video, image, audio, screen capture, and mouse cursor tracking. While the use of data modalities is not inherently problematic, inconsistencies between the modalities used for annotation and those used to train AI models can create challenges. For instance, some studies used retrospective self-reports for annotation but developed models using video data. Since self-reports are collected after engagement occurs, they may not reliably reflect in situ engagement, unlike observer-based annotations with appropriate temporal resolution. Such mismatches between annotation modality and training modality may reduce model robustness, as models are optimized to reproduce labels derived from information they do not directly observe. This can limit transferability across datasets and contexts where modality alignment differs. For example, in (Y. Chen et al., 2015), engagement was annotated through retrospective self-reports, while the model relied on video-based features for prediction. Similarly, in (Bosch et al., 2016), engagement labels were derived from self-reports combined with observer judgments, whereas the AI model primarily utilized video input. Such modality mismatches illustrate how annotation sources may not align with the information available to the model during training.

4.3. Inconsistencies in Timing

There were inconsistencies in the timing of self-reports across existing engagement datasets. Most relied on retrospective self-reports, while only a few used concurrent self-reports, and one study combined both approaches. Among observer-based annotations, nearly all were conducted retrospectively using recorded data, with only one dataset employing in situ annotation during the learning session.

4.4. Inconsistencies in Temporal Resolution

D’Mello and Graesser (2012) distinguished mood states, which span entire sessions, from faster-changing affect states that fluctuate within seconds. Their experiments (Baker et al., 2007; D’Mello & Graesser, 2012; D’Mello et al., 2007) observed affect transitions every 10–60 s. Engagement annotation, with affect as one component, should therefore align with these temporal dynamics. Existing datasets use inconsistent temporal resolutions, ranging from frame-level to video segment-level (1–30 min) and session-level annotation. Some employed 10–60-s segments, which better match known affective state dynamics.
None of the reviewed high-resolution protocols specified how to label engagement during state transitions. In lower-resolution datasets (e.g., five-minute intervals), multiple engagement states may occur within a single segment. When segments are treated as multi-sets (bags of words) (Abedi & Khan, 2022), existing protocols do not clarify how many engagement or disengagement states are needed to classify a segment. An adaptive temporal resolution (Aslan et al., 2017; Ocumpaugh, 2015) (see Appendix A) presents a plausible solution. Temporal resolution can be coded as fixed when annotations were applied at predetermined intervals (e.g., single-frame, 10 s, 5 min, full session) and as adaptive when annotations were triggered by perceived engagement state changes. In the reviewed datasets presented in Table 1, adaptive temporal resolution was implemented as observer-judged, meaning that annotations were made when annotators perceived a change in the student’s engagement state.
This inconsistency in temporal resolution makes it difficult to develop and evaluate AI models across datasets. For example, a sequential machine-learning model with a specific architecture (Abedi & Khan, 2022, 2023; Liao et al., 2021) must handle segment lengths of 10 s in (A. Gupta et al., 2016), 100 s in (Thomas et al., 2022), and 5 min in (Kaur et al., 2018).

4.5. Inconsistencies in the Level of Abstraction

Even though engagement is inherently a multi-component state, most of the reviewed datasets defined it as a single-component construct without clarifying its underlying dimensions. In some cases, engagement was annotated based solely on one component, such as affective, behavioral, or cognitive engagement. Only a few works defined and annotated engagement as a multi-component state, incorporating more than one of these dimensions.

4.6. Inconsistencies in Combination

In most datasets where engagement was defined and annotated as a multi-component construct, the components were combined to produce a single numerical engagement value, as in (Alyuz et al., 2021; Mohamad Nezami et al., 2020), where various combinations of affective and behavioral components resulted in a dichotomous engagement state. In other studies, one component was treated as a prerequisite for others. For example, in (Alkabbany et al., 2019), the presence of behavioral engagement alone (without affective engagement) indicated lower engagement levels, while the presence of both affective and behavioral components corresponded to higher engagement levels. In some works, such as (Bosch et al., 2016), the affective and behavioral components were not combined. Collapsing multiple engagement components into a single label may obscure meaningful distinctions and constrain model interpretability, limiting the ability of AI systems to generalize across contexts where engagement components manifest differently.

4.7. Inconsistencies in Quantification

A major issue in the reviewed datasets is the inconsistency in the scales used to measure SE. Various quantification methods were applied, ranging from two-point to six-point scales, as well as continuous values. Accordingly, engagement was represented as a dichotomous (binary), ordinal, interval, or categorical variable. For example, some datasets used three categories such as Engaged, Not-engaged, and Unknown. This variation complicates comparisons and model generalization across datasets.
Moreover, the limited use of psychometrically validated engagement scales raises concerns about the construct fidelity of training labels. When annotation schemes lack validation against established engagement theory, AI models may internalize inconsistent or theoretically weak proxies of engagement, reducing both interpretability and cross-context generalizability.
Corresponding to the different engagement scales in the existing datasets, different types of machine-learning models were trained to solve binary or multi-class classification or regression problems (Dewan et al., 2019; Karimah & Hasegawa, 2022). Due to this scale inconsistency, it is infeasible to use a machine-learning model trained on one dataset to make engagement inferences on another dataset annotated with a different engagement scale. Moreover, the performance of machine-learning models trained on different datasets with different engagement scales (e.g., a binary classification model with a regression model) cannot be compared. Divergent quantification schemes also restrict cross-dataset benchmarking, as models trained under binary, ordinal, or continuous formulations are not directly comparable and may encode fundamentally different assumptions about engagement.

4.8. Baseline Characteristics

The SE datasets reviewed in Table 2 show notable variability in baseline characteristics. They span activities such as software use, lectures, reading, and writing, with different levels of interactivity and real-world applicability (i.e., in-the-wild settings). Student numbers vary widely across studies, and demographic reporting on factors such as gender and age is inconsistent, with many datasets not reporting this important information, as reflected by the NA entries in Table 2. About 25%, 23%, and 10% of the datasets were collected in the United States, India, and Turkey, respectively, with 42% from other countries. As of 2025, all publicly available datasets specifically designed for AI-based student engagement measurement in virtual learning environments (A. Gupta et al., 2016; Kaur et al., 2018; Singh et al., 2023) originate from India. Although culture and ethnicity are crucial in SE measurement (Jimerson & Chen, 2022; Verkuyten et al., 2019), most datasets involve students from only one country. Few include ethnically diverse participants (Alyuz et al., 2021), which may limit AI models’ generalizability across cultural backgrounds (Xu et al., 2020). This concentration of datasets within a limited number of countries raises concerns about cultural bias in engagement annotation. Cultural norms influence facial expressivity, gaze behavior, and classroom interaction patterns, which may affect how engagement is interpreted by annotators. AI models trained on culturally homogeneous datasets may therefore encode culture-specific patterns that do not generalize across educational contexts.
In addition to cultural concentration, other contextual factors such as demographic imbalance, camera placement, lighting conditions, and recording quality may introduce systematic biases into both annotation and model training. These factors can produce uneven error patterns, for example, reduced accuracy for underrepresented groups or sensitivity to environmental variations. Addressing these risks requires more transparent reporting of dataset characteristics and the use of mitigation strategies such as reweighting, domain adaptation, and calibration when evaluating AI-based engagement models.
Engagement-level class distributions also differ, with some datasets reporting detailed breakdowns, while others give minimal or no information. This diversity in baseline characteristics, especially inconsistent demographic and engagement-distribution reporting, complicates cross-study comparisons and underscores the need for standardized reporting in future research.
The characteristics of the virtual learning environment (see Table 2) where the engagement annotation is applied are an important but often missing element in existing datasets. Most do not justify the choice of SE definitions based on the learning context. For instance, it is important to distinguish between interactive environments, where students actively interact with the computer (e.g., mouse movements), and non-interactive ones, such as passive video viewing. The settings in the reviewed studies varied and included (i) live online courses, (ii) recorded lectures viewed offline, (iii) writing tasks completed on a computer, and (iv) writing tasks completed on paper. Since affective, behavioral, and cognitive engagement differ across these contexts, annotation protocols should be tailored accordingly.
The distribution of samples in different levels of engagement in the existing datasets is presented in Table 2. It can be observed that the number of samples in disengagement or low levels of engagement is typically much lower than the number of samples in high levels of engagement in almost all the existing datasets. The highly imbalanced data distribution in these datasets must be considered when developing AI models.
Only a few of the reviewed datasets are publicly available, including DAiSEE (A. Gupta et al., 2016), EmotiW-EW (Kaur et al., 2018), and EngageNet (Singh et al., 2023). Few studies have examined annotation issues in these datasets. Abedi and Khan (2021) highlighted annotation problems in DAiSEE that mislead temporal and non-temporal deep learning models. Liao et al. (2021) and Mehta et al. (2022) also identified inconsistencies, citing examples of a single student annotated at different engagement levels. Liao et al. (2021) further criticized the use of discrete engagement labels and advocated for continuous annotation.

5. Discussion

This narrative review examined existing datasets for measuring SE in virtual learning, focusing on inconsistencies in engagement definitions and annotation protocols. Using seven dimensions of engagement annotation, the review assessed their alignment with definitions from educational psychology and analyzed how annotation and dataset design choices influence dataset validity and downstream AI model behavior. It also outlined how such inconsistencies constrain model generalizability and explored adaptable SE definitions and protocols from non-virtual learning contexts.
Based on the synthesis structured around the seven dimensions of engagement annotation, guidelines are proposed in alignment with established definitions of SE (Fredricks et al., 2004; Trowler, 2010). These guidelines are intended to support the implementation of more transparent and standardized annotation protocols when datasets are constructed for AI-based automatic engagement estimation.
The most reliable sources for engagement annotation are multiple observers with expertise in educational psychology and a strong understanding of SE Aslan et al. (2017). If non-experts are involved, structured training and evaluation are essential (Aslan et al., 2017; Ocumpaugh, 2015). The HELP protocol (Aslan et al., 2017), for instance, includes observer training, a pre-labeling phase, performance evaluation, and selection of final annotators based on their accuracy. Regarding the data modality for annotation, an important consideration is the alignment between the modalities used for annotation and those used for AI model development. For instance, if observers annotate engagement based on both student videos and screen recordings, but the AI model is trained solely on video data, it may struggle to accurately estimate engagement. A student may appear focused on the screen while actually engaging with unrelated content, leading to misleading training signals.
Regarding the timing, a retrospective approach by observing recorded data of students (e.g., videos) is optimal instead of real-time observation. In this manner, annotators can replay and examine certain moments more than once to annotate engagement more reliably. An adaptive temporal resolution, where annotations occur at points of engagement state change (Aslan et al., 2017), is recommended to ensure each data sample reflects a consistent engagement state. Fixed-length windows (A. Gupta et al., 2016; Kaur et al., 2018; Singh et al., 2023; Thomas et al., 2022) may contain multiple states, potentially causing confusion for both annotators and AI models trained on such data.
For the level of abstraction, annotating engagement components individually, rather than as a single entity, can improve annotation accuracy and simplify AI model training. As in HELP (Aslan et al., 2017) and BROMP (Ocumpaugh, 2015), affective, behavioral, and cognitive engagement can be separately labeled: facial expressions for affective, head and body pose for behavioral, and conversations or responses for cognitive engagement. Separate or multi-label AI models can then be developed for each component. Although some studies have proposed combining engagement components into a single value, such as HELP (Aslan et al., 2017) and BROMP (Ocumpaugh, 2015), which consider students disengaged if the behavioral component is Off-Task, further research is needed to determine how these components should be combined. The quantification dimension also requires additional investigation. Protocols such as HELP and BROMP use binary labels (engaged versus disengaged), which may simplify model training but risk excluding intermediate levels of engagement that fall between full engagement and complete disengagement.
Beyond methodological inconsistency, these findings raise concerns about construct validity. Engagement is a theoretically grounded, multi-component construct in educational psychology, encompassing affective, behavioral, and cognitive dimensions (Fredricks et al., 2004). When datasets operationalize engagement without explicitly mapping annotation schemes to these established constructs, AI models may learn proxies that do not faithfully represent the intended theoretical construct. This weakens both interpretability and theoretical coherence across studies.
Based on the findings of this review, the following actionable guidelines can be derived. Ideally, engagement datasets should (1) employ multiple trained annotators with clear operational protocols; (2) align annotation modality with the modality used for AI model training; (3) adopt adaptive or theoretically justified temporal resolutions that reflect engagement state changes; (4) annotate affective, behavioral, and cognitive components separately before any combination; and (5) justify quantification schemes in relation to established engagement theory. In practice, large-scale data collection may require methodological compromises. When expert annotation is not feasible, structured training and reliability checks for non-expert annotators become essential. If fixed temporal windows are used for scalability, their limitations should be explicitly acknowledged. Similarly, simplified engagement scales may be acceptable for deployment purposes, provided their theoretical implications are clearly stated.
From a validity perspective, different annotation choices may threaten multiple forms of validity. Limited or incomplete engagement components may weaken content validity; unclear alignment between labels and theoretical definitions may undermine construct validity; and inconsistent quantification or modality mismatches may reduce criterion validity by weakening the relationship between annotated engagement and observable learning outcomes. Explicitly considering these validity dimensions when designing SE datasets would strengthen both theoretical coherence and AI model reliability.
Engagement datasets often include video, audio, and behavioral traces of students, raising important considerations regarding data protection, privacy, and informed consent. Even when datasets are publicly available, issues such as re-identification risk and long-term storage of sensitive data require careful governance. In addition, biased annotation practices or demographically unbalanced datasets may propagate or amplify inequities when AI models are deployed across diverse populations. Transparent reporting of data collection procedures, demographic composition, and annotation protocols is therefore essential to mitigate ethical risks and support responsible engagement-aware AI development.
Although this review aimed to provide an in-depth and focused examination of the literature, the approach taken was narrative rather than systematic. As a result, there is a potential for missing relevant resources and introducing selection bias. Nonetheless, the risk was substantially mitigated through multiple rounds of database and gray literature searches. Another limitation is the lack of empirical validation of the proposed recommendations, as they are based on synthesis and interpretation of existing studies rather than direct testing in real-world settings. Future research should examine these recommendations through empirical, experimental, or implementation studies to assess their feasibility, effectiveness, and generalizability.
The primary objective of this review was not to analyze AI techniques for SE measurement. Readers seeking insights into AI, machine learning, and deep learning methods for SE measurement may refer to other reviews available in the literature (Booth et al., 2023; Dewan et al., 2019; Karimah & Hasegawa, 2022; Mandia et al., 2023; Salam et al., 2022).
While the efforts of researchers in collecting SE datasets are commendable, concerns remain regarding annotation comparability and AI model generalizability across datasets. This underscores the need for closer collaboration among AI, affective computing, and educational psychology researchers to develop standardized and theoretically grounded annotation protocols for SE in virtual learning. The synthesis presented in this review, structured around seven dimensions of engagement annotation and key dataset characteristics, is intended to guide the development of more transparent and consistent practices in AI-based engagement measurement. Notably, SE in virtual learning shares overlapping characteristics with engagement in other domains, such as video games (X. Chen et al., 2019; McMahan, 2013), virtual reality (Huang et al., 2021), gamification (da Rocha Seixas et al., 2016), and tele-health applications (Guhan et al., 2020). Drawing connections across these fields may offer new perspectives and support the refinement of annotation standards tailored to virtual learning contexts. Ultimately, advancing well-annotated, culturally diverse, and theoretically aligned datasets is essential for building reliable, generalizable, and impactful engagement-aware AI systems that meaningfully enhance student learning outcomes.

Author Contributions

S.S.K. prepared the initial draft of the manuscript. A.A. revised and expanded the draft, added additional records to the scoping review, designed the seven dimensions of engagement annotation framework, and managed major and minor revisions. T.J.F.C. and S.S.K. supervised and oversaw the study. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the University Health Network and University of Toronto Lillian Love Chair in Women’s Health, held by Tracey J. F. Colella.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

Not applicable.

Conflicts of Interest

The authors declare no conflict of interest.

Appendix A. Student Engagement Annotation Protocols in Other Settings

This appendix explores SE definitions and annotation protocols from contexts outside virtual learning, highlighting their potential applicability for measuring SE in virtual learning environments.
The Baker Rodrigo Ocumpaugh Monitoring Protocol (BROMP) (Ocumpaugh, 2015) is an observation protocol for in vivo annotation of students’ affective and behavioral states. In the BROMP platform, observers (referred to as sources) undergo training on the BROMP annotation protocol. After training, they are tested to ensure their observations achieve a sufficient inter-rater agreement, with Cohen’s Kappa > 0.6. Once they meet this standard, they are granted a BROMP certification and can then participate in the annotation process. Students in an in-person classroom working with educational software on computers are observed by the observer in-person by a side glance to make a holistic judgment of a student’s state based on facial expressions, speech, body posture, gestures, and the student’s interaction with the educational software (data modality).
Observation is performed in a round-robin manner, observing and annotating one student and moving to the next. The frequency of observations per student varied between class periods depending on the number of students in the class (timing). Each student is observed for 20 s or until a visible state is detected (temporal resolution). The annotation is inserted in a mobile application. In the BROMP protocol, the affective and behavioral states of students are annotated separately (combination). Various affect states are included in the BROMP protocol; some commonly used are Boredom, Confusion, Delight, Engaged Concentration, Frustration, and Surprise. The main behavioral states are On-Task and Off-Task (level of abstraction and quantification).
Aslan et al. (2017) stated unaddressed challenges in BROMP (Ocumpaugh, 2015) as follows: (1) limited chance for revision as annotation is performed in vivo, (2) difficulty of making a decision about a student’s state in real-time, (3) fragmented annotation and disregarding state change in students due to the round-robin technique, (4) limited labels for model training, and (5) inevitable observer effect due to the presence of the observer in the learning settings.
To address the challenges mentioned above, Aslan et al. (2017) developed the Human Expert Labeling Process (HELP) for SE annotation, which consists of three stages: pre-labeling, labeling, and post-labeling. The pre-labeling stage involves a systematic process of training and evaluation for observers (source). Following this, the second stage focuses on the main annotation performed by observers. Finally, in the third stage, the annotations are evaluated, and the annotated dataset is created.
HELP uses an annotation software containing the recorded video, audio, screen capture of students’ computers and learning material contextual data, and demographic information of students (data modality). The timing is post-facto as observers watch the recorded data of students retrospectively. Temporal resolution is similar to BROMP, after observing the first state change of the student. The annotation of the affect and behavioral states are separate. The discrete affective states are Satisfied, Bored, and Confused, and the behavioral states are On-Task and Off-Task (level of abstraction and scale). Inspired by Woolf et al. (2009), different combinations of affect and behavioral states result in the dichotomous state of Engaged versus Not-engaged (combination).
Some studies have used BROMP (Bosch et al., 2016) and HELP (Alyuz et al., 2021, 2017; Okur et al., 2017) for engagement annotation. Aslan et al. (2017) did not clearly illustrate the effectiveness of addressing the fifth challenge in BROMP concerning the inevitable observer effect. Further research is needed to determine the more effective approach, whether it is continuously observing one student as in HELP or employing the round-robin technique in BROMP.

References

  1. Abedi, A., & Khan, S. S. (2021). Improving state-of-the-art in detecting student engagement with resnet and TCN hybrid network. In 2021 18th conference on robots and vision (CRV) (pp. 151–157). IEEE. [Google Scholar]
  2. Abedi, A., & Khan, S. S. (2022). Detecting disengagement in virtual learning as an anomaly. arXiv, arXiv:2211.06870. [Google Scholar]
  3. Abedi, A., & Khan, S. S. (2023). Affect-driven ordinal engagement measurement from video. Multimedia Tools and Applications, 83, 24899–24918. [Google Scholar] [CrossRef] [Scilit]
  4. Abedi, A., & Khan, S. S. (2024). Engagement measurement based on facial landmarks and spatial-temporal graph convolutional networks. In International conference on pattern recognition (pp. 321–338). IEEE. [Google Scholar]
  5. Alkabbany, I., Ali, A., Farag, A., Bennett, I., Ghanoum, M., & Farag, A. (2019). Measuring student engagement level using facial information. In 2019 IEEE international conference on image processing (ICIP) (pp. 3337–3341). IEEE. [Google Scholar]
  6. Altuwairqi, K., Jarraya, S. K., Allinjawi, A., & Hammami, M. (2021). A new emotion–based affective model to detect student’s engagement. Journal of King Saud University-Computer and Information Sciences, 33(1), 99–109. [Google Scholar] [CrossRef] [Scilit]
  7. Alyuz, N., Aslan, S., D’Mello, S. K., Nachman, L., & Esme, A. A. (2021). Annotating student engagement across grades 1–12: Associations with demographics and expressivity. In International conference on artificial intelligence in education (pp. 42–51). Springer. [Google Scholar]
  8. Alyuz, N., Okur, E., Genc, U., Aslan, S., Tanriover, C., & Esme, A. A. (2017). An unobtrusive and multimodal approach for behavioral engagement detection of students. In Proceedings of the 1st ACM sigchi international workshop on multimodal interaction for education (pp. 26–32). Association for Computing Machinery. [Google Scholar]
  9. Aslan, S., Cataltepe, Z., Diner, I., Dundar, O., Esme, A. A., Ferens, R., Kamhi, G., Oktay, E., Soysal, C., & Yener, M. (2014). Learner engagement measurement and classification in 1:1 learning. In 2014 13th international conference on machine learning and applications (pp. 545–552). IEEE. [Google Scholar]
  10. Aslan, S., Mete, S. E., Okur, E., Oktay, E., Alyuz, N., Genc, U. E., Stanhill, D., & Esme, A. A. (2017). Human expert labeling process (HELP): Towards a reliable higher-order user state labeling process and tool to assess student engagement. Educational Technology, 57(1), 53–59. [Google Scholar]
  11. Baker, R. S. d., Rodrigo, M. M. T., & Xolocotzin, U. E. (2007). The dynamics of affective transitions in simulation problem-solving environments. In International conference on affective computing and intelligent interaction (pp. 666–677). Spinger. [Google Scholar]
  12. Bhardwaj, P., Gupta, P., Panwar, H., Siddiqui, M. K., Morales-Menendez, R., & Bhaik, A. (2021). Application of deep learning on student engagement in e-learning environments. Computers & Electrical Engineering, 93, 107277. [Google Scholar]
  13. Booth, B. M., Ali, A. M., Narayanan, S. S., Bennett, I., & Farag, A. A. (2017). Toward active and unobtrusive engagement assessment of distance learners. In 2017 seventh international conference on affective computing and intelligent interaction (ACII) (pp. 470–476). IEEE. [Google Scholar]
  14. Booth, B. M., Bosch, N., & D’Mello, S. K. (2023). Engagement detection and its applications in learning: A tutorial and selective review. Proceedings of the IEEE, 111(10), 1398–1422. [Google Scholar] [CrossRef] [Scilit]
  15. Bosch, N. (2016). Detecting student engagement: Human versus machine. In Proceedings of the 2016 conference on user modeling adaptation and personalization (pp. 317–320). Association for Computing Machinery. [Google Scholar]
  16. Bosch, N., D’mello, S. K., Ocumpaugh, J., Baker, R. S., & Shute, V. (2016). Using video to automatically detect learner affect in computer-enabled classrooms. ACM Transactions on Interactive Intelligent Systems (TiiS), 6(2), 1–26. [Google Scholar] [CrossRef] [Scilit]
  17. Brabham, D. C. (2013). Crowdsourcing. MIT Press. [Google Scholar]
  18. Buono, P., De Carolis, B., D’Errico, F., Macchiarulo, N., & Palestra, G. (2023). Assessing student engagement from facial behavior in on-line learning. Multimedia Tools and Applications, 82(9), 12859–12877. [Google Scholar] [CrossRef] [Scilit]
  19. Chen, J., Luo, N., Liu, Y., Liu, L., Zhang, K., & Kolodziej, J. (2016). A hybrid intelligence-aided approach to affect-sensitive e-learning. Computing, 98(1), 215–233. [Google Scholar] [CrossRef] [Scilit]
  20. Chen, X., Niu, L., Veeraraghavan, A., & Sabharwal, A. (2019). FaceEngage: Robust estimation of gameplay engagement from user-contributed (YouTube) videos. IEEE Transactions on Affective Computing, 13(2), 651–665. [Google Scholar] [CrossRef] [Scilit]
  21. Chen, Y., Bosch, N., & D’Mello, S. (2015, June 26–29). Video-based affect detection in noninteractive learning environments. Proceedings of the 8th International Conference on Educational Data Mining, Madrid, Spain. [Google Scholar]
  22. da Rocha Seixas, L., Gomes, A. S., & de Melo Filho, I. J. (2016). Effectiveness of gamification in the engagement of students. Computers in Human Behavior, 58, 48–63. [Google Scholar] [CrossRef] [Scilit]
  23. De Carolis, B., D’Errico, F., Macchiarulo, N., & Palestra, G. (2019). “Engaged faces”: Measuring and monitoring student engagement from face and gaze behavior. In IEEE/WIC/ACM international conference on web intelligence-companion volume (pp. 80–85). Association for Computing Machinery. [Google Scholar]
  24. Delgado, K., Origgi, J. M., Hasanpoor, T., Yu, H., Allessio, D., Arroyo, I., Lee, W., Betke, M., Woolf, B., & Bargal, S. A. (2021). Student engagement dataset. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 3628–3636). IEEE. [Google Scholar]
  25. Dewan, M., Murshed, M., & Lin, F. (2019). Engagement detection in online learning: A review. Smart Learning Environments, 6(1), 1. [Google Scholar] [CrossRef] [Scilit]
  26. D’Mello, S. K. (2015). On the influence of an iterative affect annotation approach on inter-observer and self-observer reliability. IEEE Transactions on Affective Computing, 7(2), 136–149. [Google Scholar] [CrossRef] [Scilit]
  27. D’Mello, S. K., & Graesser, A. (2012). Dynamics of affective states during complex learning. Learning and Instruction, 22(2), 145–157. [Google Scholar] [CrossRef] [Scilit]
  28. D’Mello, S. K., Taylor, R., & Graesser, A. (2007). Monitoring affective trajectories during complex learning. In Proceedings of the annual meeting of the cognitive science society (Vol. 29). Springer. [Google Scholar]
  29. Dung, D. T. H. (2020). The advantages and disadvantages of virtual learning. IOSR Journal of Research & Method in Education, 10(3), 45–48. [Google Scholar]
  30. Fredricks, J. A., Blumenfeld, P. C., & Paris, A. H. (2004). School engagement: Potential of the concept, state of the evidence. Review of educational research, 74(1), 59–109. [Google Scholar] [CrossRef] [Scilit]
  31. Gray, J. A., & DiLoreto, M. (2016). The effects of student engagement, student satisfaction, and perceived learning in online learning environments. International Journal of Educational Leadership Preparation, 11(1), n1. [Google Scholar]
  32. Guhan, P., Awasthi, N., Bussell, K., Manocha, D., Reeves, G., & Bera, A. (2020). Developing an effective and automated patient engagement estimator for telehealth: A machine learning approach. arXiv, arXiv:2011.08690. [Google Scholar]
  33. Gupta, A., D’Cunha, A., Awasthi, K., & Balasubramanian, V. (2016). Daisee: Towards user engagement recognition in the wild. arXiv, arXiv:1609.01885. [Google Scholar]
  34. Gupta, S., Kumar, P., & Tekchandani, R. K. (2023a). A multimodal facial cues based engagement detection system in e-learning context using deep learning approach. Multimedia Tools and Applications, 82, 28589–28615. [Google Scholar]
  35. Gupta, S., Kumar, P., & Tekchandani, R. K. (2023b). Facial emotion recognition based real-time learner engagement detection system in online learning context using deep learning models. Multimedia Tools and Applications, 82(8), 11365–11394. [Google Scholar]
  36. Gupta, S. K., Ashwin, T., & Guddeti, R. M. R. (2019). Students’ affective content analysis in smart classroom environment using deep learning techniques. Multimedia Tools and Applications, 78, 25321–25348. [Google Scholar] [CrossRef] [Scilit]
  37. Hallgren, K. A. (2012). Computing inter-rater reliability for observational data: An overview and tutorial. Tutorials in Quantitative Methods for Psychology, 8(1), 23. [Google Scholar] [CrossRef] [Scilit]
  38. Henrie, C. R., Halverson, L. R., & Graham, C. R. (2015). Measuring student engagement in technology-mediated learning: A review. Computers & Education, 90, 36–53. [Google Scholar] [CrossRef] [Scilit]
  39. Hu, M., & Li, H. (2017). Student engagement in online learning: A review. In 2017 international symposium on educational technology (ISET) (pp. 39–43). IEEE. [Google Scholar]
  40. Huang, W., Roscoe, R. D., Johnson-Glenberg, M. C., & Craig, S. D. (2021). Motivation, engagement, and performance across multiple virtual reality sessions and levels of immersion. Journal of Computer Assisted Learning, 37(3), 745–758. [Google Scholar]
  41. Hutt, S., Grafsgaard, J. F., & D’Mello, S. K. (2019). Time to scale: Generalizable affect detection for tens of thousands of students across an entire school year. In Proceedings of the 2019 CHI conference on human factors in computing systems (pp. 1–14). Association for Computing Machinery. [Google Scholar]
  42. Jeong, Y.-S., & Cho, N.-W. (2023). Evaluation of e-learners’ concentration using recurrent neural networks. The Journal of Supercomputing, 79, 4146–4163. [Google Scholar]
  43. Jimerson, S. R., & Chen, C. (2022). Multicultural and cross-cultural considerations in understanding student engagement in schools: Promoting the development of diverse students around the world. In Handbook of research on student engagement (pp. 629–646). Springer. [Google Scholar]
  44. Kamath, A., Biswas, A., & Balasubramanian, V. (2016). A crowdsourced approach to student engagement recognition in e-learning environments. In 2016 IEEE winter conference on applications of computer vision (WACV) (pp. 1–9). IEEE. [Google Scholar]
  45. Karimah, S. N., & Hasegawa, S. (2021). Automatic engagement recognition for distance learning systems: A literature study of engagement datasets and methods. In International conference on human-computer interaction (pp. 264–276). Springer. [Google Scholar]
  46. Karimah, S. N., & Hasegawa, S. (2022). Automatic engagement estimation in smart education/learning settings: A systematic review of engagement definitions, datasets, and methods. Smart Learning Environments, 9(1), 31. [Google Scholar] [CrossRef] [Scilit]
  47. Kaur, A., Mustafa, A., Mehta, L., & Dhall, A. (2018). Prediction and localization of student engagement in the wild. In 2018 digital image computing: Techniques and applications (DICTA) (pp. 1–8). IEEE. [Google Scholar]
  48. Liao, J., Hao, Y., Zhou, Z., Pan, J., & Liang, Y. (2024). Sequence-level affective level estimation based on pyramidal facial expression features. Pattern Recognition, 145, 109958. [Google Scholar]
  49. Liao, J., Liang, Y., & Pan, J. (2021). Deep facial spatiotemporal network for engagement prediction in online learning. Applied Intelligence, 51(10), 6609–6621. [Google Scholar] [CrossRef] [Scilit]
  50. Liu, Y., Chen, J., Zhang, M., & Rao, C. (2018). Student engagement study based on multi-cue detection and recognition in an intelligent learning environment. Multimedia Tools and Applications, 77, 28749–28775. [Google Scholar] [CrossRef] [Scilit]
  51. Ma, X., Xu, M., Dong, Y., & Sun, Z. (2021). Automatic student engagement in online learning environment based on neural turing machine. International Journal of Information and Education Technology, 11(3), 107–111. [Google Scholar] [CrossRef] [Scilit]
  52. Mandia, S., Mitharwal, R., & Singh, K. (2023). Automatic student engagement measurement using machine learning techniques: A literature study of data and methods. Multimedia Tools and Applications, 83, 49641–49672. [Google Scholar] [CrossRef] [Scilit]
  53. McMahan, A. (2013). Immersion, engagement, and presence: A method for analyzing 3-D video games. In The video game theory reader (pp. 67–86). Routledge. [Google Scholar]
  54. McNeish, D. (2022). Limitations of the sum-and-alpha approach to measurement in behavioral research. Policy Insights from the Behavioral and Brain Sciences, 9(2), 196–203. [Google Scholar] [CrossRef] [Scilit]
  55. Mehta, N. K., Prasad, S. S., Saurav, S., Saini, R., & Singh, S. (2022). Three-dimensional DenseNet self-attention neural network for automatic detection of student’s engagement. Applied Intelligence, 52, 13803–13823. [Google Scholar]
  56. Minner, D. D., Levy, A. J., & Century, J. (2010). Inquiry-based science instruction—what is it and does it matter? Results from a research synthesis years 1984 to 2002. Journal of Research in Science Teaching: The Official Journal of the National Association for Research in Science Teaching, 47(4), 474–496. [Google Scholar]
  57. Mohamad Nezami, O., Dras, M., Hamey, L., Richards, D., Wan, S., & Paris, C. (2020). Automatic recognition of student engagement using deep learning and facial expression. In Joint european conference on machine learning and knowledge discovery in databases (pp. 273–289). Springer. [Google Scholar]
  58. Monkaresi, H., Bosch, N., Calvo, R. A., & D’Mello, S. K. (2016). Automated detection of engagement using video-based estimation of facial expressions and heart rate. IEEE Transactions on Affective Computing, 8(1), 15–28. [Google Scholar] [CrossRef] [Scilit]
  59. Mukhtar, K., Javed, K., Arooj, M., & Sethi, A. (2020). Advantages, limitations and recommendations for online learning during COVID-19 pandemic era. Pakistan Journal of Medical Sciences, 36(COVID19-S4), S27. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  60. Naibert, N., & Barbera, J. (2022). Development and evaluation of a survey to measure student engagement at the activity level in general chemistry. Journal of Chemical Education, 99(3), 1410–1419. [Google Scholar] [CrossRef] [Scilit]
  61. O’Brien, H., & Cairns, P. (2016). Why engagement matters: Cross-disciplinary perspectives of user engagement in digital media. Springer. [Google Scholar]
  62. Ocumpaugh, J. (2015). Baker rodrigo ocumpaugh monitoring protocol (BROMP) 2.0 technical and training manual (p. 60). Teachers College, Columbia University and Ateneo Laboratory for the Learning Sciences. [Google Scholar]
  63. Okur, E., Alyuz, N., Aslan, S., Genc, U., Tanriover, C., & Arslan Esme, A. (2017). Behavioral engagement detection of students in the wild. In International conference on artificial intelligence in education (pp. 250–261). Springer. [Google Scholar]
  64. Porayska-Pomsta, K., Mavrikis, M., D’Mello, S., Conati, C., & Baker, R. S. (2013). Knowledge elicitation methods for affect modelling in education. International Journal of Artificial Intelligence in Education, 22(3), 107–140. [Google Scholar] [CrossRef] [Scilit]
  65. Salam, H., Celiktutan, O., Gunes, H., & Chetouani, M. (2022). Automatic context-driven inference of engagement in HMI: A survey. arXiv, arXiv:2209.15370. [Google Scholar] [CrossRef] [Scilit]
  66. Schindler, L. A., Burkholder, G. J., Morad, O. A., & Marsh, C. (2017). Computer-based technology and student engagement: A critical review of the literature. International Journal of Educational Technology in Higher Education, 14(1), 25. [Google Scholar] [CrossRef] [Scilit]
  67. Sinatra, G. M., Heddy, B. C., & Lombardi, D. (2015). The challenges of defining and measuring student engagement in science (Vol. 50, No. 1). Taylor & Francis. [Google Scholar]
  68. Singh, M., Hoque, X., Zeng, D., Wang, Y., Ikeda, K., & Dhall, A. (2023). Do I have your attention: A large scale engagement prediction dataset and baselines. In Proceedings of the 25th international conference on multimodal interaction (pp. 174–182). Association for Computing Machinery. [Google Scholar]
  69. Sümer, Ö., Goldberg, P., D’Mello, S., Gerjets, P., Trautwein, U., & Kasneci, E. (2021). Multimodal engagement analysis from facial videos in the classroom. IEEE Transactions on Affective Computing, 14(2), 1012–1027. [Google Scholar] [CrossRef] [Scilit]
  70. Tafreshi, D., Slaney, K. L., & Neufeld, S. D. (2016). Quantification in psychology: Critical analysis of an unreflective practice. Journal of Theoretical and Philosophical Psychology, 36(4), 233. [Google Scholar] [CrossRef] [Scilit]
  71. Thomas, C., Sarma, K. P., Gajula, S. S., & Jayagopi, D. B. (2022). Automatic prediction of presentation style and student engagement from videos. Computers and Education: Artificial Intelligence, 3, 100079. [Google Scholar] [CrossRef] [Scilit]
  72. Trowler, V. (2010). Student engagement literature review (Vol. 11, No. 1, pp. 1–15). The Higher Education Academy. [Google Scholar]
  73. Van Engelen, J. E., & Hoos, H. H. (2020). A survey on semi-supervised learning. Machine Learning, 109(2), 373–440. [Google Scholar] [CrossRef] [Scilit]
  74. Vanneste, P., Oramas, J., Verelst, T., Tuytelaars, T., Raes, A., Depaepe, F., & Van den Noortgate, W. (2021). Computer vision and human behaviour, emotion and cognition detection: A use case on student engagement. Mathematics, 9(3), 287. [Google Scholar] [CrossRef] [Scilit]
  75. Verkuyten, M., Thijs, J., & Gharaei, N. (2019). Discrimination and academic (dis) engagement of ethnic-racial minority students: A social identity threat perspective. Social Psychology of Education, 22, 267–290. [Google Scholar] [CrossRef] [Scilit]
  76. Verma, M., Nakashima, Y., Takemura, N., & Nagahara, H. (2022). Multi-label disengagement and behavior prediction in online learning. In International conference on artificial intelligence in education (pp. 633–639). Springer. [Google Scholar]
  77. Whitehill, J., Serpell, Z., Lin, Y.-C., Foster, A., & Movellan, J. R. (2014). The faces of engagement: Automatic recognition of student engagementfrom facial expressions. IEEE Transactions on Affective Computing, 5(1), 86–98. [Google Scholar] [CrossRef] [Scilit]
  78. Wong, Z. Y., & Liem, G. A. D. (2021). Student engagement: Current state of the construct, conceptual refinement, and future research directions. Educational Psychology Review, 34(1), 1–32. [Google Scholar] [CrossRef] [Scilit]
  79. Woolf, B., Burleson, W., Arroyo, I., Dragon, T., Cooper, D., & Picard, R. (2009). Affect-aware tutors: Recognising and responding to student affect. International Journal of Learning Technology, 4(3/4), 129–164. [Google Scholar] [CrossRef] [Scilit]
  80. Xu, T., White, J., Kalkan, S., & Gunes, H. (2020). Investigating bias and fairness in facial expression recognition. In Computer vision—ECCV 2020 workshops: Glasgow, UK, August 23–28, Proceedings, Part VI 16 (pp. 506–523). Springer. [Google Scholar]
  81. Zaletelj, J., & Košir, A. (2017). Predicting students’ attention in the classroom from Kinect facial and body features. EURASIP Journal on Image and Video Processing, 2017(1), 80. [Google Scholar] [CrossRef] [Scilit]
  82. Zheng, X., Hasegawa, S., Tran, M.-T., Ota, K., & Unoki, T. (2021). Estimation of learners’ engagement using face and body features by transfer learning. In International conference on human-computer interaction (pp. 541–552). Springer. [Google Scholar]
Figure 1. Seven dimensions of engagement annotation.
Figure 1. Seven dimensions of engagement annotation.
Education 16 00548 g001
Table 1. The seven dimensions of engagement annotation in the reviewed student engagement datasets.
Table 1. The seven dimensions of engagement annotation in the reviewed student engagement datasets.
ReferenceSourcesData ModalityTimingTemporal ResolutionLevel of AbstractionCombinationQuantification
 Whitehill et al. (2014)Observers: 9 trained studentsVideoRetrospective10 and 60 sAffective, behavioral, and cognitive componentsRule-basedOrdinal 1–4: not at all engaged–very engaged
Aslan et al. (2014)Observers: 3 expertsVideo and Computer screenRetrospective20 minEngagementNACategorical: engaged, not engaged, and unknown
Y. Chen et al. (2015)self-reportsNARetrospective2 minAffective componentNAOrdinal 1–6: very little engagement–very much engagement
Bosch (2016)Observers: trained expertsVideoConcurrentAdaptive or 20 sAffective and behavioral componentsNo combinationCategorical: boredom, confusion, delight, frustration, engaged concentration and off-task, on-task conversation, and on-task
J. Chen et al. (2016)-Video and skin conductanceRetrospective-EngagementNADichotomous: attention ON and attention OFF
Monkaresi et al. (2016)self-reportsNAConcurrent and retrospective2 minEngagementNADichotomous: engaged and not-engaged
A. Gupta et al. (2016)Observers: 10 untrained crowdsourcersVideoRetrospective10 sAffective componentNAOrdinal 1–4: very low–very high
Kamath et al. (2016)Observers: 25 untrained crowdsourcersVideoRetrospectivesingle-frameEngagementNAOrdinal 1–3: not-engaged–very engaged
Bosch et al. (2016)self-reports + observers (untrained annotators)VideoConcurrent + retrospective12 sCognitive componentNADichotomous: mind-wandering (not-engaged) and no mind-wandering (engaged)
Booth et al. (2017)Observers: 9 untrained studentsVideoRetrospective20 minEngagementNAInterval 0–1
Okur et al. (2017)Observers: 3 expertsVideo, audio, computer screen, mouse cursor, and URL logsRetrospectiveAdaptiveBehavioral componentNACategorical: on-task, off-task, not applicable, cannot decide
Alyuz et al. (2017)Observers: 3 expertsVideo, audio, computer screen, and mouse cursorRetrospectiveAdaptiveBehavioral componentNACategorical: on-task, off-task, not applicable, cannot decide
Zaletelj and Košir (2017)Observers: 5 expertsVideoRetrospective1 sEngagementNAOrdinal 1–5: 5 levels of engagement
Kaur et al. (2018)Observers: 5 expertsVideoRetrospective5 minBehavioral componentNAOrdinal 0–3: completely disengaged–highly engaged
De Carolis et al. (2019)self-reportsNARetrospective9 minCognitive Engagement--
 Hutt et al. (2019)self-reportsNARetrospectiveRandomAffective componentNAOrdinal 1–5: not at all engaged to very engaged
Alkabbany et al. (2019)Observers: 4 annotatorsVideoRetrospectiveSingle-frameAffective and behavioral componentsNo combinationOrdinal 0–3: no detected face to emotionally engaged
Mohamad Nezami et al. (2020)Observers: 6 trained studentsVideoRetrospectiveSingle-frameAffective and behavioral componentsRule-basedDichotomous: engaged and not-engaged
Vanneste et al. (2021)self-reportsNAConcurrent5–12 minEngagementNAInterval 0–2: totally disengaged to totally engaged
Alyuz et al. (2021)Observers: 3 expertsAudio, video, and screen captureRetrospectiveAdaptiveAffective and behavioral componentsNo combinationCategorical: satisfied, bored, confused, on-task, off-task, not available, and cannot decide
Bhardwaj et al. (2021)Observers: 10 annotatorsVideoRetrospectiveEngagementNAOrdinal 0–5: six levels of engagement
Delgado et al. (2021)Observers: 3 untrained crowdsourcersVideoRetrospectiveSingle-frameBehavioral componentNo combinationCategorical: looking at their screen, looking at their paper, and wandering
Zheng et al. (2021)self-reports + observersVideoRetrospective12 minEngagementNAOrdinal 1–3: three levels of engagement
Altuwairqi et al. (2021)self-reports + observers (untrained students)VideoRetrospectiveSingle-frameAffective componentNAOrdinal 1–5: low to strong engagement
Ma et al. (2021)Observers: 3 untrained annotatorsVideoRetrospective30 minEngagementNAInterval 0–1
S. Gupta et al. (2023b)ImageRetrospectiveSingle-frameAffective engagementNAInterval 0–1: basic facial expressions converted to an interval variable
Buono et al. (2023)self-reportsNARetrospective9 minEngagementNAInterval 0–1
Jeong and Cho (2023)Observers: 3 expertsVideoRetrospective5 sEngagementNADichotomous: engaged and not-engaged
Thomas et al. (2022)self-reportsNARetrospective100 sEngagementNAOrdinal 1–5
Verma et al. (2022)Observers: 3 trained annotatorsVideoRetrospectiveAdaptiveBehavioral engagementNACategorical: disengagement, strange eye movements, presence of some kind of facial expression, yawning, face occlusion, body movements, and pressing keyboard
Singh et al. (2023)Observers: 3 untrained annotatorsVideoRetrospective10 sEngagementNAOrdinal 1–4: very low to very high
NA: For self-reports, students annotated their own engagement without utilizing any specific data modalities for annotation. In cases where the level of abstraction was “Engagement,” the observers did not annotate individual components of engagement, making no combination necessary. In several datasets, the individual components of engagement were not combined.
Table 2. The baseline characteristics of the reviewed student engagement datasets.
Table 2. The baseline characteristics of the reviewed student engagement datasets.
ReferenceActivityInteractiveIn-the-Wild# of Students# of FemalesAge (Years)CountryClass Distribution of Samples
Whitehill et al. (2014)SoftwareYesNo3425NAUnited States6% not engaged at all, 10% nominally engaged, 46% engaged in task, and 38% very engaged
Aslan et al. (2014)LectureNoNo9NANATurkeyNA
Y. Chen et al. (2015)ReadingNoYes88NANAUnited StatesNA
Bosch (2016)SoftwareYesYes1375713–15United States4% boredom, 2% confusion, 2% delight, 14% frustration, and 78% engaged concentration; 5% off-task, 21% on-task conversation, and 74% on-task
J. Chen et al. (2016)SoftwareYesNo3017NAChinaNA
Monkaresi et al. (2016)WritingNoNo23920–60Australia80% engaged and 20% not-engaged
A. Gupta et al. (2016)LectureNoYes11232NAIndia1% very low, 5% low, 41% high, and 45% very high level of engagement
Kamath et al. (2016)LectureNoYes23NA18–24India9% not-engaged, 51% nominally engaged, and 40% very engaged
Bosch et al. (2016)ReadingNoYes98NANAUnited StatesNA
Booth et al. (2017)LectureNoNo12NA25United StatesNA
Okur et al. (2017)LectureYesYes28NA14–15Turkey71% on-task and 29% off-task
Alyuz et al. (2017)LectureYesYes17NA14–15Turkey68% on-task and 32% off-task
Zaletelj and Košir (2017)LectureNoNo222NASloveniaNA
Kaur et al. (2018)LectureNoYes782519–27India5% completely disengaged, 27% barely engaged, 42% engaged, and 26% highly engaged
De Carolis et al. (2019)LectureYesNo19721ItalyNA
Hutt et al. (2019)LectureYesYes69,174NANAUnited StatesNA
Alkabbany et al. (2019)LectureYesYes14NANAUnited States0% no face detected, 12% behaviorally not engaged, 56% behaviorally engaged, emotionally not engaged, and 32% emotionally engaged
Mohamad Nezami et al. (2020)SoftwareYesNo201114–16Australia50% engaged and 50% not-engaged
Vanneste et al. (2021)LectureNoYes14418BelgiumNA
Alyuz et al. (2021)LectureYesNo6030NACanadaNA
Bhardwaj et al. (2021)LectureNoYes1000NANAIndiaNA
Delgado et al. (2021)SoftwareYesNo19NANAUnited States25% looking at their screen, 72% looking at their paper, and 3% wandering
Zheng et al. (2021)SoftwareYesNo19NANAJapan33% (1), 40% (2), and 27% (3)
Altuwairqi et al. (2021)LectureNoYes110NANASaudi ArabiaNA
Ma et al. (2021)LectureNoYes59NA20–32ChinaNA
S. Gupta et al. (2023b)NANANANANANAIndiaNA
Buono et al. (2023)LectureNoYes311419ItalyNA
Jeong and Cho (2023)LectureNoYes92NA20–31South KoreaNA
Thomas et al. (2022)LectureNoYes6NANAIndiaNA
Verma et al. (2022)ReadingYesNo26NANAJapanNA
Singh et al. (2023)LectureNoYes1274418–37IndiaNA
NA: The corresponding baseline characteristic of the dataset was not reported.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Khan, S.S.; Abedi, A.; Colella, T.J.F. Rethinking Student Engagement Datasets for AI in Virtual Learning: A Narrative Review. Educ. Sci. 2026, 16, 548. https://doi.org/10.3390/educsci16040548

AMA Style

Khan SS, Abedi A, Colella TJF. Rethinking Student Engagement Datasets for AI in Virtual Learning: A Narrative Review. Education Sciences. 2026; 16(4):548. https://doi.org/10.3390/educsci16040548

Chicago/Turabian Style

Khan, Shehroz S., Ali Abedi, and Tracey J. F. Colella. 2026. "Rethinking Student Engagement Datasets for AI in Virtual Learning: A Narrative Review" Education Sciences 16, no. 4: 548. https://doi.org/10.3390/educsci16040548

APA Style

Khan, S. S., Abedi, A., & Colella, T. J. F. (2026). Rethinking Student Engagement Datasets for AI in Virtual Learning: A Narrative Review. Education Sciences, 16(4), 548. https://doi.org/10.3390/educsci16040548

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop