Next Article in Journal
Conditional Counter-Inspection with Curriculum-Biased Experts for Lightweight 5G Intrusion Detection
Next Article in Special Issue
A Generative AI Architecture Integrating Retrieval-Augmented Generation and Low-Rank Adaptation for Knowledge-Intensive Medical Reasoning
Previous Article in Journal
Optimizing Collaborative Filtering for Accurate Rating Predictions in Very Sparse Datasets
Previous Article in Special Issue
Novel Video Understanding Approach for Embodied Learning of Robotics Technology
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment

by
Georgios Kostopoulos
1,2,
Sotiris Kotsiantis
1,*,
Theodor Panagiotakopoulos
3,4 and
Achilles Kameas
5
1
School of Social Sciences, Hellenic Open University, 26331 Patras, Greece
2
Department of Mathematics, University of Patras, 26504 Patras, Greece
3
Department of Management Science and Technology, University of Patras, 26334 Patras, Greece
4
School of Business, University of Nicosia, 1700 Nicosia, Cyprus
5
School of Technology and Science, Hellenic Open University, 26335 Patras, Greece
*
Author to whom correspondence should be addressed.
Future Internet 2026, 18(3), 115; https://doi.org/10.3390/fi18030115
Submission received: 1 February 2026 / Revised: 16 February 2026 / Accepted: 23 February 2026 / Published: 24 February 2026

Abstract

Multimodal Learning Analytics (MMLA) is an extension of Learning Analytics that combines multiple data streams such as audio, video, physiological signals, logs, and spatial trails to analyze learning processes that cannot be easily captured through any single modality. This review synthesizes research on sensing and instrumentation, feature extraction, multimodal fusion, modeling approaches, and end-to-end systems that provide feedback and support reflection. We also discuss how generative AI and Large Language Models (LLMs) increasingly improve MMLA pipelines by enabling scalable semantic and pragmatic analysis of learner discourse and interaction. In addition, we review robustness issues that arise when working with real-world data (e.g., noise, missing data, and scalability) and responsible deployment issues such as privacy and student-focused views of fairness, accountability, transparency, and ethics (FATE).

Graphical Abstract

1. Introduction

Multimodal Learning Analytics (MMLA) has emerged as a critical response to the inherent limitations of traditional Learning Analytics (LA). While conventional LA often relies on interaction logs and clickstream data, these unimodal records fail to capture the complexity of central learning processes, such as collaborative dynamics, emotional engagement, attention management, and physiological regulation [1,2]. As education increasingly shifts towards complex physical, social, and blended environments, the exclusive reliance on single-modality digital traces poses serious challenges to the ecological validity and educational relevance of analytical inferences.
By integrating heterogeneous data streams, including eye movements, body posture, gestures, speech, spatial location, and physiological signals, MMLA seeks to achieve a comprehensive and dynamically nuanced understanding of learning as it unfolds in real-time [3]. This approach allows researchers to model latent phenomena that are otherwise unobservable. Crucially, recent scholarship indicates a paradigm shift in the field: moving beyond descriptive modeling toward the development of closed-loop systems that support actionable interventions, such as automated feedback, reflection, and formative assessment in authentic settings [4,5].
The rapid evolution of MLA is currently driven by three converging forces. First, advances in sensing technologies and computer vision have substantially improved the feasibility of capturing rich behavioral data in real-world environments. Innovations such as 360-degree cameras are now being employed to mitigate occlusion and enhance face and pose estimation in crowded collaborative scenarios [6]. Second, progress in Machine Learning (ML) has revolutionized feature extraction. Notably, the integration of Large Language Models (LLMs) and generative Artificial Intelligence (AI) has expanded the analytic scope of MMLA, enabling sophisticated semantic and pragmatic analyses of student discourse and identifying teamwork patterns with unprecedented accuracy [7,8]. Third, the deployment of MMLA in authentic contexts, ranging from K-12 classrooms to clinical simulations, has foregrounded the necessity of responsible AI. Issues of privacy, consent, and fairness are no longer secondary concerns but central design constraints, motivating a surge in research focused on student-centered FATE (Fairness, Accountability, Transparency, and Ethics) principles [9].
Despite its promise, MMLA remains a technically and conceptually complex domain. The field is characterized by significant variation in modality selection, sensing configurations, and fusion techniques [10,11]. Furthermore, critical challenges persist: fusion protocols are often under-specified; robustness to noise and incomplete data remains unevenly addressed [12]; and system assessment is frequently biased towards predictive accuracy at the expense of interpretability and pedagogical utility.
Unlike previous reviews that have adopted primarily bibliometric or modality-focused approaches [13,14], this survey adopts a systems-focused, pipeline-level perspective. Rather than examining data sources in isolation, we synthesize how sensing architectures, data fusion strategies, modeling techniques, and feedback loops interact to support or hinder the deployment of MMLA in real-world educational settings.
The current survey provides a structured synthesis of MMLA research with five major aims. First, we classify common modalities and sensor configurations. Second, we dissect end-to-end MMLA systems, tracing the pipeline from data collection to high-level reasoning. Third, we critically examine modeling and assessment methods, with a focus on robustness and validity. Fourth, we review systems designed to formally close the loop, enabling real-time feedback and self-regulated learning support. Finally, we emphasize the growing body of research dedicated to responsible deployment, privacy preservation, and ethical governance. Throughout, we categorize research across diverse learning contexts (K-12, higher education, and professional training) to identify emerging trends and delineate future research directions.

2. Research Questions and Review Method

This article is a survey that targets the systems-level deployment of Multimodal Learning Analytics (MMLA), with an emphasis on end-to-end pipelines, classroom-facing tools that close the feedback loop, and the socio-technical requirements for trustworthy adoption. To make the contribution auditable and clearly differentiated from a high-level position piece, we articulate explicit research questions (RQs) and define traceable synthesis criteria that govern study inclusion and analysis.

2.1. Research Questions

The review is guided by three research questions:
  • RQ1 (Pipeline and architecture): What are the recurring architectural components and design patterns of MMLA systems deployed beyond controlled laboratory settings (e.g., sensing configurations, feature extraction, fusion strategies, and reliability mechanisms)?
  • RQ2 (Loop-closing and pedagogical integration): How do MMLA systems translate multimodal inferences into actionable feedback for learners and educators (e.g., dashboards, debriefing tools, formative assessment workflows, and human-in-the-loop processes)?
  • RQ3 (Deployment constraints and gaps): What technical, organizational, and ethical constraints most frequently impede systems-level deployment, and which research gaps remain under-addressed by existing reviews (e.g., fusion reporting standards, robustness in authentic settings, validity/fairness evidence, and governance)?

2.2. Traceable Synthesis Criteria

To ensure that claims in later sections can be traced to identifiable evidence, we apply the following synthesis criteria when selecting and interpreting the literature:
  • Scope: Studies involve multimodal data (at least two modalities) and connect measurements to learning or collaboration constructs in an educational or training context.
  • Systems-level focus: We prioritize work that specifies an end-to-end pipeline (from sensing through inference to interpretation/feedback), reports synchronization and fusion decisions, or evaluates a tool used in practice (e.g., dashboard, debriefing workflow, automated feedback).
  • Deployment evidence: We distinguish between (i) controlled studies and (ii) authentic or semi-authentic deployments (e.g., classrooms, simulations, longitudinal use). Claims about feasibility or impact are grounded primarily in the latter category when available.
  • Extraction template: For each included work, we code (a) context and participants, (b) modalities and sensors, (c) target constructs and operationalizations, (d) feature extraction approach (including LLM-based coding when applicable), (e) fusion strategy (early/late/hybrid/temporal), (f) validation/evaluation design, and (g) reported risks/mitigations (privacy, fairness, transparency, governance).
  • Synthesis logic: Findings are synthesized by mapping coded evidence onto the reference architecture and the loop-closing taxonomy, and by summarizing recurring limitations as open gaps.

2.3. Study Identification and Screening (PRISMA-Inspired Workflow)

To mitigate selection bias and improve reproducibility, we followed a PRISMA-inspired workflow adapted to a structured narrative survey. Records were retrieved from Scopus using the query families described in Section 3. After de-duplication, titles and abstracts were screened against the inclusion/exclusion criteria. Full-text screening was then conducted for borderline cases and for studies that appeared to report end-to-end systems, multimodal fusion decisions, or classroom-facing tools. Figure 1 summarizes identification, screening, eligibility assessment, and inclusion. Counts are reported for transparency and updated to match the final corpus.

2.4. Quality/Rigor Appraisal

Because MMLA studies vary widely in design (e.g., tool deployments, controlled experiments, case studies), we did not apply a single numeric risk-of-bias score. Instead, during extraction, we recorded whether each study reported: (i) participant/context details, (ii) modality/sensor specifications, (iii) synchronization and fusion decisions, (iv) validation strategy (including construct validity when applicable), and (v) deployment evidence (controlled vs. authentic/semi-authentic settings). When reporting was incomplete on these dimensions, the corresponding findings are treated as tentative and discussed as gaps rather than generalized conclusions.

2.5. Positioning Relative to Prior Reviews and the Gap This Survey Fills

Prior MMLA reviews have advanced (a) modality/sensor taxonomies and (b) methodological overviews of feature extraction and modeling. However, three gaps motivate the present survey’s contribution:
  • From components to deployable systems: We synthesize evidence at the end-to-end pipeline level (sensing, preprocessing, alignment, feature extraction, fusion, and deployment), highlighting how temporal synchronization and robustness decisions constrain downstream pedagogical use.
  • Closing-the-loop as a first-class object of analysis: Beyond prediction and classification, we analyze how studies implement actionability (feedback, reflection, and orchestration) and what forms of human-in-the-loop integration and evaluation are reported.
  • Trustworthy deployment as socio-technical orchestration: We consolidate deployment constraints spanning technical reliability, validity/fairness, privacy/consent, and governance, and we surface under-reported aspects such as fusion reporting, uncertainty communication, and longitudinal evidence of impact.

3. Scope and Methodology

This survey adopts a structured narrative review approach to synthesize research on MMLA in educational contexts. To support transparency and reproducibility, we specify the search sources, time window, screening process, and selection criteria used to construct the final corpus.
Search strategy and sources: the literature search was conducted in Scopus, Web of Science, and Google Scholar, covering publications from 2015 to early 2025. The core query combined terms for multimodality with terms for learning analytics and education (e.g., “multimodal learning analytics”, “multimodal data” AND education, sensor-based” AND learning analytics, “embodied ” AND learning analytics). We additionally performed backward searching by inspecting the reference lists of key surveys and highly cited empirical studies to identify further relevant work.
Screening and eligibility: records were screened at the title/abstract level, followed by full-text assessment for eligibility. When multiple versions of the same study were identified (e.g., technical report and extended journal article), we retained the most complete version. Figure 1 summarizes the identification and screening process; the final counts were updated to match the included corpus.
Inclusion criteria (studies were included) if they:
  • Focused on educational or training contexts (formal, informal, or professional learning);
  • Employed two or more data modalities (e.g., video, audio, physiological, interaction logs);
  • Reported empirical findings, system designs, or methodological frameworks that connect multimodal evidence to learning constructs, instructional design, or learning outcomes.
Exclusion criteria (studies were excluded) if they:
  • Addressed multimodal sensing without a learning or instructional focus;
  • Focused exclusively on affect detection or surveillance without pedagogical grounding;
  • Were purely technical computer vision or signal-processing papers without an educational setting, learning construct/outcome, or learning-oriented interpretation of the extracted features.
Synthesis approach: rather than organizing studies by modality alone, we synthesize findings across the multimodal data pipeline (sensing, preprocessing, feature extraction, fusion, modeling, and interpretation), application contexts, and ethical considerations. This organization highlights cross-cutting design patterns, recurrent methodological bottlenecks, and evidence gaps relevant to the deployment of MMLA systems in authentic educational settings.

4. Background and Conceptual Foundations

4.1. Defining Multimodality in Learning Analytics

In the context of LA, multimodality refers to the systematic integration of heterogeneous channels of evidence to model learning processes that cannot be adequately captured through a single data source. These channels typically span digital interaction traces (e.g., logs, clickstreams), observable physical behaviors (e.g., posture, gesture, proximity, eye gaze), verbal and paralinguistic communications (e.g., speech, discourse structure), and physiological signals (e.g., heart rate, and electrodermal activity) [1,13,15].
Crucially, within the MMLA domain, multimodality is defined not merely by the simultaneous presence of multiple sensors, but by the analytic intention to triangulate these data streams to infer meaningful latent constructs. Different modalities often capture complementary aspects of the learning experience. For instance, gaze and posture may serve as proxies for attention and physical engagement, while speech and discourse provide insight into epistemic and social coordination [3,16]. The primary advantage of MMLA lies in capitalizing on these complementarities to resolve ambiguities inherent in unimodal data, rather than treating modalities as redundant predictors.
However, the literature reveals considerable variation in the operationalization of multimodality. Implementations range from loose coupling, where data streams are analyzed in parallel, to tight integration, involving fusion at the feature, model, or decision levels. This heterogeneity poses challenges for comparative analysis and underscores the need for a rigorous methodological basis for defining and validating multimodal constructs in educational settings [10].

4.2. From Descriptive Analytics to Evidence-Based Interventions

A long-standing ambition of MMLA research is to transcend descriptive diagnostics and achieve closing the loop: linking data collection and modeling to timely, interpretable, and pedagogically actionable interventions [4,5]. This transition from sensing to supporting manifests primarily in two forms. First, human-in-the-loop systems provide reflective dashboards and debriefing tools that assist learners and instructors in making sense of complex evidence. Recent longitudinal studies suggest that such visualization tools can significantly enhance collaborative reflection and formative assessment when aligned with instructional goals [4,17]. Second, a nascent but rapidly growing body of work leverages generative AI and LLMs to automate the feedback loop. Unlike static dashboards, these systems can generate contextualized, natural-language feedback in real time, addressing the scalability limitations of human-led debriefing [7,18].
Despite these advances, the effectiveness of MMLA interventions is not solely determined by algorithmic accuracy. Emerging research emphasizes that user perception, specifically regarding usefulness, trust, and cognitive load, is a critical success factor. This has led to a focus on the principles of FATE, positing that sophisticated models must be explainable and transparent to be trusted by stakeholders [9,19]. Consequently, the field is moving towards a user-centered evaluation paradigm that assesses not just predictive performance, but also the educational impact and ethical reception of MMLA systems.

4.3. Design and Specification Frameworks

The technical and socio-pedagogical complexity inherent in MMLA systems has necessitated the development of specialized frameworks to explicate design assumptions and facilitate cross-disciplinary collaboration. Unlike traditional, language-agnostic data mining pipelines, MMLA systems involve extended value chains ranging from physical sensing infrastructure to signal processing, fusion, modeling, and interpretation.
The Multimodal Data Value Chain (M-DVC) was introduced as a conceptual tool to structure this complexity [20]. By articulating each stage of the pipeline, the M-DVC helps stakeholders surface hidden dependencies, identify potential sources of noise, and manage data quality issues early in the design process [12].
Building on this foundation, recent scholarship has adopted a broader, design-oriented perspective. The Multimodal Learning Analytics Design Framework (MDF), grounded in empirical studies of expert practitioners, proposes a set of recurring design considerations that guide system development across diverse contexts [11,21]. These frameworks collectively argue that MMLA solutions must be conceptualized as socio-technical systems, where success depends as much on learning design, institutional context, and ethical governance as it does on sensor accuracy and algorithmic sophistication.

5. Modalities, Sensors, and Data Sources

5.1. Visual and Spatial Data: Embodied Interaction and Proximity

Visual and spatial modalities are among the most prevalent data sources in MMLA, particularly for investigating embodied, collaborative, and physically mediated learning. Advances in video sensing and motion capture enable researchers to quantify postures, gestures, movement trajectories, and interpersonal proximity, providing empirical proxies for non-verbal engagement and social coordination. Early foundational work focused on developing specialized tools for posture recognition and visualization to decode embodied behaviors in situated learning [22,23].
Beyond individual behavioral tracking, spatial analytics facilitate the examination of how physical environments and learning designs shape interaction. For instance, motion capture data has demonstrated that furniture configurations and spatial affordances directly influence participation dynamics and group problem-solving efficiency [24]. This underscores the dual utility of visual data: modeling the learner while simultaneously evaluating the learning environment’s design.
Recent research has prioritized enhancing the robustness of visual feature extraction in in-the-wild group settings. Comparative analyses indicate that alternative capture configurations, such as 360-degree cameras, significantly mitigate occlusion and field-of-view constraints compared to traditional setups, yielding higher-fidelity pose and facial estimates [6]. Despite these technical leaps, vision-based modalities remain sensitive to environmental lighting, pose significant privacy concerns, and entail high computational costs, necessitating a strategic approach to sensor selection.

5.2. Audio and Speech: Discourse Structure and Participation Dynamics

Auditory data provides a critical window into the verbal and paralinguistic dimensions of learning. Speech features are extensively used to map participation balance and turn-taking patterns, particularly in collaborative contexts. By treating speech patterns as social relations data, researchers utilize social network analysis (SNA) to detect emergent roles, leadership dominance, and overall collaboration quality [25].
The integration of Automated Speech Recognition (ASR) and Natural Language Processing (NLP) has shifted the focus from simple participation metrics to the semantic and epistemic depth of discourse. Modern pipelines now support the automated transcription and coding of teamwork communication, even in fast-paced or physically distributed environments [8]. However, these automated approaches introduce specific technical challenges, such as speaker diarization errors in noisy classrooms and the difficulty of parsing domain-specific jargon, both of which can impact the validity of downstream analysis.

5.3. Physiological and Affective Signals: Unveiling Internal States

Physiological modalities are employed to infer latent cognitive and affective states that are otherwise inaccessible through observation. Common signals include heart rate (HR), electrodermal activity (EDA), electroencephalography (EEG), and eye-tracking measures. Triangulating physiological signals with behavioral data has proven more effective in detecting learner wakefulness, cognitive load, and attention than unimodal approaches [26].
Moreover, physiological sensing is increasingly utilized to support adaptive and self-regulatory interventions. For example, EEG-based neurofeedback loops integrated into MMLA pipelines can help learners monitor and regulate their cognitive states during demanding online tasks [27]. Systematic reviews in the field highlight a growing trend toward using biosensors for personalized learning, while simultaneously cautioning against significant hurdles such as signal noise, individual baseline variability, and the intrusiveness-utility trade-off [1].

5.4. Digital Traces: The Contextual Anchor

Despite the rise of sophisticated physical sensors, traditional digital traces, such as log files, clickstreams, and learning artifacts, remain the backbone of MMLA. These traces provide the necessary contextual grounding and temporal anchors for interpreting high-frequency multimodal signals. In programming education, for instance, process traces from Integrated Development Environments (IDEs) are fused with behavioral data to analyze the impact of instructor scaffolding on students’ problem-solving strategies [28,29,30].
In game-based learning, digital logs are often synchronized with gaze and facial expressions to model the relationship between performance and emotional valence [15]. Ultimately, digital traces play a foundational role in ensuring the interpretability and validity of MMLA systems, confirming that multimodal data extends, rather than replaces, traditional analytics.
The synthesis of Table 1 and Table 2 demonstrates that the value and feasibility of multimodal data sources are inherently context-dependent. While visual, auditory, and physiological modalities provide access to latent learning constructs that are highly complementary, their adoption remains heterogeneous across educational sectors. This uneven distribution is primarily driven by divergent logistical constraints, ethical mandates, and pedagogical objectives.
While digital traces offer a low-friction entry point for analytics, more sophisticated multimodal configurations are currently concentrated within tertiary education and research laboratories. In these environments, infrastructure readiness, robust consent protocols, and specialized educational goals are better aligned to support high-fidelity data collection. These observations underscore the critical necessity of strategic modality selection, a design consideration that balances data richness with ethical and practical viability. The specific mechanisms and frameworks for integrating these diverse data streams will be analyzed in the following section on multimodal fusion and modeling.
Finally, Figure 2 presents a comprehensive taxonomy of multimodal learning data utilized in MMLA systems, categorizing sources into visual, auditory, interactional, physiological, and textual modalities.

6. The Multimodal Learning Analytics Pipeline

6.1. Reference Architecture: A Modular Perspective

Despite the diversity in sensing technologies and learning domains, a common six-stage reference architecture has emerged: (1) data acquisition and temporal alignment, (2) preprocessing, (3) modality-specific feature extraction, (4) multimodal fusion, (5) modeling and inference, and (6) pedagogical interpretation. This architecture emphasizes that multimodality is not merely an attribute of the modeling step but a property that must be maintained throughout the entire analytical process.
Recent scholarship increasingly prioritizes scalability and modularity to ensure these systems function reliably outside controlled laboratory environments. For instance, smart classroom architectures now emphasize reconfigurability and the dynamic deployment of sensing elements to accommodate the fluid nature of authentic educational spaces [31]. Such considerations are vital for supporting longitudinal studies and fostering sustained instructional use.

6.2. Feature Extraction: Bridging Raw Signals and Learning Constructs

Feature extraction serves as the translation layer between raw sensor data and theoretically grounded learning constructs. Extracted features typically range from low-level signal descriptors to high-level behavioral representations:
  • Visual Pipelines: Derive gaze distributions for attention mapping [18], pose keypoints for coordination analysis [6], and discretized posture classes to visualize embodied engagement [22].
  • Audio Pipelines: Focus on speech activity and conversational turn-taking, often operationalized through participation graphs or social network measures [25].
  • Semantic and Pragmatic Features: Most notably, the integration of LLMs has revolutionized this layer, enabling the automated coding of communication quality and interactional competence at a scale previously impossible with manual annotation [7,18].
Despite these advancements, pipeline noise, including diarization errors and overlapping speech, remains a persistent challenge that can propagate through the system, potentially compromising model validity [8,18].

6.3. From Multimodal Indicators to Learning Constructs: Validity and Theoretical Grounding

A central challenge in MMLA is that multimodal measurements are not learning constructs. Raw signals (e.g., video, audio, physiology, logs) are transformed into features (e.g., gaze dispersion, turn-taking rates, EDA peaks), which may then be interpreted as indicators(e.g., attention, participation, arousal). Learning-science constructs such as cognitive load, engagement, self-regulation, or collaboration quality are higher-level theoretical entities that require explicit operationalization and validation. Treating indicators as direct proxies for constructs risks over-interpretation and weakens cumulative comparison across studies.

6.3.1. Why the Mapping Is Non-Trivial

The same observable pattern can reflect different underlying processes depending on task demands, social norms, and instructional design (context dependence). Multiple different behaviors can lead to similar outcomes (equifinality), and learners may change their behavior when sensing is present (reactivity). As a result, indicator-to-construct mappings should be treated as hypotheses that require evidence, not as fixed correspondences.

6.3.2. Problematizing Common Constructs

For example, physiological arousal may correlate with cognitive load, stress, or excitement; speech quantity may reflect collaboration quality or simply role assignment; and gaze direction can indicate attention, confusion, or off-task monitoring depending on the activity. Even when predictive performance is high, it does not guarantee construct validity if models rely on context-specific proxies.

6.3.3. Minimum Validity and Reporting Expectations

To strengthen theoretical support and reduce construct drift, MMLA studies should report:
  • Construct definition: the learning-theory definition adopted (and why it fits the context).
  • Operationalization: how the construct is mapped to indicators (features, windows, thresholds).
  • Triangulation: alignment with independent evidence (e.g., expert ratings, validated instruments, performance measures, qualitative observations).
  • Boundary conditions: when the indicator is expected to fail (task types, populations, environments).
  • Interpretation limits: what the model output does not justify (e.g., causal claims, stable traits).
This perspective frames multimodal modeling as a socio-technical measurement problem: progress depends not only on better sensors and fusion, but also on transparent construct definitions, theory-consistent operationalizations, and validation designs that connect multimodal indicators to learning processes in context.

6.4. Multimodal Fusion Strategies

Fusion lies at the conceptual heart of MMLA, yet it remains one of the least standardized components in empirical research. Recent scoping reviews indicate that while data collection is often documented in detail, the specific mechanisms of fusion and their implications for model transparency are frequently under-reported [10]. We classify fusion techniques into three primary categories:
  • Early Fusion (Feature Level): Combines features into a single high-dimensional vector before modeling. While it captures cross-modal interactions early, it is highly sensitive to noise and missing data.
  • Late Fusion (Decision Level): Aggregates the outputs of independent unimodal models (e.g., via ensemble learning). This approach offers superior modularity and robustness to sensor failure but may overlook fine-grained inter-modality dependencies.
  • Hybrid and Temporal Fusion: Align modalities along a temporal axis to study the dynamic coordination of multimodal events. This provides the most nuanced view of learning but significantly increases computational and modeling complexity.
Beyond this high-level taxonomy, fusion choices have direct implications for validity, interpretability, robustness, and fairness in deployed systems. Fusion is not merely a modeling preference; it is a commitment about (i) how modalities are assumed to relate, (ii) how missing or noisy channels are handled, and (iii) what evidence is ultimately presented to learners and teachers. Below we highlight several under-discussed implications.

6.4.1. Construct Validity and Cross-Modal Confounding

Early fusion can encourage models to exploit modality-specific shortcuts (e.g., microphone placement, camera angle, or classroom layout) that correlate with outcomes but are not theoretically meaningful learning constructs. Because correlations become entangled across modalities, it becomes harder to show that the fused representation operationalizes the intended construct rather than proxy signals. Late fusion preserves unimodal pathways and can make it easier to audit which modalities drive predictions, but it may still mask cross-modal confounds if the decision rule and modality weights are not reported.

6.4.2. Interpretability and the Evidence Trail

For loop-closing applications, stakeholders need to understand what evidence supports a claim about engagement, collaboration, or affect. Late fusion typically supports clearer evidence trails (separate discourse, posture, and participation indicators), whereas early fusion can produce accurate predictions with limited explanatory value. Hybrid approaches can balance these needs only if authors report which cross-modal interactions are learned and how they are communicated at the interface level.

6.4.3. Robustness to Missingness and Sensor Failure

Early fusion often assumes synchronized feature vectors and can degrade sharply under missingness, requiring imputation choices that meaningfully affect conclusions. Late fusion is structurally more robust to sensor dropout (unimodal models can still operate), but it can fail silently if the aggregation stage does not explicitly represent uncertainty. For deployment, studies should report (i) missingness rates by modality, (ii) fallback behaviors, and (iii) how uncertainty is propagated to dashboards or feedback.

6.4.4. Temporal Alignment as a Modeling Assumption

Temporal and hybrid fusion require decisions about time windows, lag, and alignment. These choices implicitly encode a theory of learning dynamics; different window sizes can change which cross-modal relationships appear stable or “significant”. Under-reporting alignment decisions therefore risks overgeneralization and limits comparability across studies.

6.4.5. Fairness, Privacy, and Differential Modality Burden

Modalities differ in intrusiveness and error profiles across contexts and populations (e.g., vision under occlusion/lighting; physiological baselines across individuals). Early fusion can amplify these disparities if errors from one modality dominate the fused representation. A more critical reporting standard should therefore include subgroup performance where feasible and a justification of modality selection relative to privacy burden.
Taken together, these implications suggest that under-reporting of fusion mechanisms is not a minor documentation issue but a barrier to cumulative knowledge. At minimum, studies should specify (a) the fusion stage(s), (b) synchronization/alignment choices, (c) missing-data handling, (d) modality contribution or ablation evidence, and (e) how fused outputs are translated into actionable and uncertainty-aware feedback.

6.5. Robustness and Reliability in Authentic Settings

In real-world classrooms, MMLA pipelines must be resilient to sensor noise, occlusion, and missingness. Empirical evidence suggests that while attribute noise can degrade performance, certain algorithmic families, particularly tree-based ensemble methods, demonstrate higher robustness in noisy environments [12].
Furthermore, robustness is increasingly viewed through a socio-technical lens. Technical resilience directly impacts student and teacher trust; misrepresentations caused by data gaps can lead to perceived unfairness. Ensuring transparency regarding data uncertainty is now recognized as a prerequisite for the ethical deployment of MMLA systems [4,19].

6.6. End-to-End Interpretability and Explainability

Interpretability is not a post-hoc feature but an end-to-end design requirement. In MMLA, multiple levels of abstraction separate raw signals from derived constructs, making explainability particularly challenging.
  • At the Sensing Level: Interpretability requires transparency regarding what data is captured and under what conditions.
  • At the Modeling Level: A known tension exists between the high accuracy of black-box deep learning models (often used in early fusion) and the interpretability of late fusion models.
  • At the Interface Level: Dashboards and feedback tools must communicate not only the result but also the underlying uncertainty and data limitations.
Research confirms that stakeholders in collaborative settings demand to understand the why behind analytical outputs to engage in meaningful reflection [4,19]. Consequently, the field is shifting toward Explainable Multimodal Analytics, where system transparency is treated as a core pedagogical affordance.
Figure 3 summarizes the canonical MMLA workflow, highlighting the role of multimodal fusion as a key intermediary between data preprocessing and LA models.

7. Modeling Approaches and Evaluation Practices

The selection of a modeling strategy in MMLA is inherently tied to the pedagogical objectives of the study. This section reviews the transition from predictive benchmarks to process-oriented and human-centered evaluation frameworks.

7.1. Supervised Learning: Predicting Outcomes and Detecting States

Supervised learning remains the dominant paradigm within the MMLA community, particularly for modeling discrete learning outcomes such as academic performance, engagement levels, affective states, and team success. Researchers employ a spectrum of algorithms, ranging from classical classifiers, including Decision Trees, Random Forests, and Support Vector Machines, to sophisticated Deep Learning architectures [15,26,32].
Empirical evidence suggests that model selection is highly context-dependent. In collaborative and project-based learning, multimodal features derived from computer vision and interaction with physical artifacts can predict team success with high accuracy [32]. While deep models excel in scenarios with high-resolution temporal data and large-scale datasets, traditional machine learning models are often preferred for their robustness to noise and inherent interpretability, especially when data are limited. Consequently, the selection of a model in MMLA must balance predictive power with the need for pedagogical transparency.

7.2. Unsupervised Learning and Pattern Discovery

Unsupervised methods are instrumental in MMLA for discovering latent behavioral patterns without predefined labels. Clustering techniques have been successfully applied to identify distinct learner profiles, such as active, passive, or transitional states, within complex environments like programming classrooms [33]. Comparative analyses further demonstrate how multimodal process data can reveal systematic differences in behavioral trajectories across different instructional modalities, such as block-based versus text-based coding environments [28].
Earlier work underscores the value of unsupervised approaches for uncovering how physical actions and verbalizations co-evolve in embodied learning contexts [34]. While these methods are powerful for hypothesis generation and exploratory research, they are highly sensitive to feature selection and time granularity. Therefore, unsupervised findings require rigorous triangulation with theoretical frameworks or qualitative evidence to ensure educational validity.

7.3. Temporal and Network-Based Analytics for Collaboration

Modeling collaboration necessitates analytical strategies that account for both interactional structure and temporal evolution. Temporal analyses provide a window into the micro-genetics of learning, capturing phase transitions, social regulation, and patterns of coordination that emerge over time. Studies in embodied collaboration demonstrate that integrating verbal, spatial, and physiological streams allows for a granular distinction between successful and unsuccessful collaboration that aggregate measures often overlook [16].
Complementing these temporal views, network-based models, such as Social and Epistemic Network Analysis (SNA/ENA), offer interpretable representations of collaborative dynamics. By operationalizing discourse and participation as relational structures, these models enable the examination of influence, role distribution, and knowledge construction processes [8,25]. These representations serve as a critical bridge, translating raw multimodal data into constructs that are actionable for educators, such as teamwork competence and collaborative quality.

7.4. Evaluation Paradigms: Beyond Predictive Accuracy

Evaluation in MMLA is undergoing a significant shift, moving away from an exclusive focus on technical metrics like accuracy, precision, and F1-scores. While these remain relevant for model comparison, they provide little insight into a system’s pedagogical utility or ethical standing. Recent research advocates for a multi-dimensional evaluation framework that incorporates perceived utility, interpretability, data transparency, and stakeholder trust [4].
Student-centered evaluations have highlighted the urgency of integrating FATE principles into MMLA deployment. Qualitative evidence suggests that learners are acutely sensitive to data representation, levels of access, and the ongoing nature of consent [19]. This shift implies that a successful MMLA system is one that is not only statistically resilient but also socio-technically reliable and ethically sound.
Table 3 reinforces that evaluation in mmlacannot be decoupled from modeling intent. A predictive model optimized for outcome prediction requires validation evidence that is not equivalent to that needed by an exploratory or process-oriented method. Thus, any evaluation framework that is narrowly focused on predictive performance may end up misunderstanding the utility of unsupervised network analyses or temporal analyses, whose primary utility is in interpretation and theoretical and educational value. To move mmlafrom a state of technical viability toward meaningful and responsible use, it is critical that objectives and methods be aligned.

7.5. Strength of Evidence for Pedagogical Impact

A recurring limitation in the MMLA literature is that evidence of pedagogical impact often lags behind advances in sensing and modeling. Many studies convincingly demonstrate detection, classification, or descriptive analytics (e.g., engagement state inference, collaboration-role detection), yet fewer evaluate whether these inferences lead to sustained improvements in learning outcomes, instructional decision-making, or equity-sensitive benefits over time. To avoid overgeneralization, we distinguish three evidence tiers:
  • Tier 1: Technical proof-of-concept. Studies that establish the feasibility of multimodal capture and prediction (accuracy/F1, robustness under noise) but do not test downstream instructional use or learning effects.
  • Tier 2: Classroom-facing tools with usability/acceptability evidence. Studies that integrate MMLA into dashboards, debriefing workflows, or feedback tools and report user perceptions (usefulness, trust, interpretability), but provide limited causal or longitudinal evidence of learning gains.
  • Tier 3: Validated instructional interventions. Studies that evaluate MMLA-enabled interventions with stronger designs (e.g., longitudinal deployments, quasi-experimental or experimental comparisons, pre/post learning measures, or instructor decision outcomes) enable more credible claims about pedagogical impact.
Across the reviewed literature, Tier 1 and Tier 2 studies are substantially more prevalent than Tier 3. Therefore, claims about “transformative” impact should be interpreted primarily as potential rather than established effect, with the strongest evidence currently concentrated in specific contexts (e.g., simulation-based training and structured debriefing) where feedback loops and assessment practices are already well-defined.

8. Applications by Learning Context

8.1. Professional Simulations: The Frontier of High-Stakes Training

Simulation-based professional training represents the most mature application domain for MMLA. In high-stakes fields such as healthcare, maritime, and aviation, the stringent requirements for effective communication, coordination, and teamwork necessitate comprehensive sensing environments. These settings provide the controlled conditions required for complex data acquisition and offer high instructional returns through structured, data-driven debriefing.
Recent evidence from clinical simulations demonstrates that end-to-end MMLA pipelines, integrating visual attention, speech processing, and discourse semantics, can now generate automated feedback on interaction quality with high fidelity [18]. Complementary design ethnographies emphasize that success in these environments depends on embedding sensors within recurring cycles of observation and intervention without undermining the “fidelity” or realism of the simulation [35]. Thus, professional simulations serve as a primary proving ground for scalable, action-oriented MMLA deployments.

8.2. STEM and Programming: Decoding Process over Product

STEM and computer programming education constitute the second major pillar of MMLA research, focusing on the cognitive and behavioral dimensions of technical skill acquisition. By enriching traditional development logs with multimodal data, researchers can gain insight into the “hidden” aspects of the coding process [21].
Current research utilizes MMLA to disentangle the complex dynamics of pair programming, identifying interaction patterns that directly correlate with learning outcomes [29]. Furthermore, MMLA has been instrumental in evaluating instructor scaffolding, revealing how different types of intervention exert immediate or delayed effects on group problem-solving strategies [30]. Comparative studies have also utilized multimodal traces to contrast block-based versus text-based modalities, uncovering differences in debugging behavior and time allocation that remain invisible to performance-based metrics alone [28].

8.3. Online Learning: Enhancing Observability and Self-Regulation

In online environments, MMLA addresses the observability gap regarding a learner’s internal cognitive and affective states. Interaction data is frequently supplemented with physiological measures to monitor attention, fatigue, and self-regulatory processes. For instance, combining heart rate, pressure sensors, and facial analysis has significantly improved the detection of learner wakefulness and drowsiness compared to unimodal interaction logs [26].
Beyond mere detection, physiological signals are increasingly explored as a foundation for adaptive neurofeedback mechanisms designed to support self-regulated learning [27]. However, as noted in recent systematic reviews, online applications often favor lightweight sensing configurations to mitigate challenges related to signal reliability, individual variability, and the inherent intrusiveness of wearable hardware [1].

8.4. K–12 Education: Personalization Under Ethical Constraints

Compared to higher education and professional training, MMLA applications in K-12 are more varied and exploratory. While the potential for personalization is significant, allowing teachers to tap into learning facets that traditional assessments miss, the field faces substantial hurdles [14].
A scoping study of MMLA in K-8 environments identified persistent gaps in data fusion transparency and a lack of focus on ethical considerations [10]. Consequently, current research in this sector emphasizes the development of developmentally sound and ethically grounded design principles. The goal is to ensure that personalization through MMLA does not compromise student privacy or instructional equity [14,19].

8.5. Synthesis of Deployment Maturity

The comparative analysis in Table 4 underscores that the maturity and deployment readiness of MMLA are markedly heterogeneous across learning environments. This variation reflects fundamental differences in pedagogical stakes, infrastructural support, and the specific governance frameworks of each sector.
Currently, professional and simulation-based training emerges as the most advanced application area. By leveraging controlled settings, standardized scenarios, and high incentives for performance, these environments successfully utilize MMLA to provide actionable, high-fidelity feedback on teamwork and communication. In stark contrast, K–12 and informal learning environments remain in an exploratory phase. In these contexts, the complex interplay of classroom logistics, stringent ethical mandates, and the demand for algorithmic transparency significantly constrains the extent of sensorization and system integration.
Ultimately, the trajectory of MMLA is driven not only by advancements in sensing and modeling but also by the socio-technical readiness of the learning context itself. This underscores the necessity of designing MMLA systems that are context-sensitive and strictly aligned with the affordances and constraints of the target environment.

9. Systems That Close the Loop: Feedback, Reflection, and Formative Assessment

A fundamental objective of MMLA is to transcend descriptive diagnostics by engineering systems that actively facilitate teaching and learning. Closing the loop in this context refers to the intentional transformation of complex multimodal data into representations that are not only theoretically sound but also interpretable, trustworthy, and actionable for both learners and educators [4,20]. By prioritizing feedback, reflection, and formative assessment, MMLA shifts from being a mere analytical framework to becoming a robust pedagogical infrastructure capable of supporting real-time and post-hoc instructional interventions.

9.1. Dashboards and Reflective Debriefing Tools

The mediation between intricate analytic pipelines and pedagogical decision-making is frequently facilitated through visualization and structured reflection. Dashboards that synthesize non-verbal collaborative behaviors, such as speech distribution, physical posture, and interactional balance, exemplify how latent multimodal signals can be converted into actionable insights without overwhelming users with raw sensor data [3,22]. Such tools are most effective when integrated into facilitated debriefing sessions, where analytics serve as data-informed prompts for dialogue rather than definitive evaluative judgments [35].
Longitudinal evaluations of evidence-based MMLA systems indicate a predominantly positive reception, with stakeholders reporting enhanced clarity in learning objectives and strengthened support for reflective practices [4]. Nevertheless, these deployments surface persistent tensions regarding system complexity, perceived data accuracy, and the “black-box” nature of some models. These challenges underscore that closing the loop is a socio-technical endeavor that requires a delicate calibration between analytical depth and user-facing simplicity to foster sustained trust [19,36].

9.2. Formative Assessment in Collaborative Learning

MMLA provides a unique vantage point for formative assessment in collaborative environments, where traditional, unimodal metrics often fail to capture the nuances of group dynamics. By triangulating discourse patterns, facial expressions, and affective states, MMLA can identify behavioral markers that correlate with long-term learning outcomes and professional competency development [5,15].
In teacher training and professional simulations, for instance, MMLA acts as an automated observer that provides a continuous, process-driven stream of evidence [5,18]. However, the transition toward automated formative assessment necessitates rigorous scrutiny of validity and fairness. Scholars argue for a design-for-formative-processes approach, ensuring that automation supports rather than replaces human judgment, thereby addressing concerns related to algorithmic bias and the oversimplification of complex social interactions [10,19].

9.3. Generative AI and MMLA-Enabled Learning Aids

The convergence of generative AI and MMLA defines the next frontier of loop-closing systems. Integrating LLMs with multimodal pipelines allows for the creation of sophisticated learning aids that offer personalized, natural-language guidance [7,8]. This integration suggests a dual role for GenAI in the MMLA ecosystem:
  • Adaptive Intervention: GenAI agents can function as real-time tutors that adapt their feedback based on a multimodal understanding of the learner’s cognitive load, engagement, and emotional state [1,7].
  • Enriched Data Acquisition: Interactions with GenAI tools generate rich, high-fidelity data that can be re-analyzed to refine models of self-directed and collaborative learning, essentially making the AI tool both a catalyst for learning and a sophisticated sensor for analysis [7,18].
As illustrated in Table 5, closing the MMLA loop involves navigating distinct trade-offs between interpretability, scalability, and ethical risk. While visualization-heavy systems (dashboards) prioritize human-in-the-loop transparency, AI-driven agents prioritize personalized responsiveness. A successful MMLA infrastructure is likely to be one that effectively orchestrates these diverse feedback channels, ensuring that the analytical insights generated are ethically grounded and pedagogically transformative [9,19].

10. Ethics, Privacy, and Trustworthy Deployment

Ethical, legal, and social implications (ELSI) are not peripheral concerns but reside at the very core of Multimodal Learning Analytics (MMLA). Unlike traditional analytics, MMLA necessitates high-resolution, often continuous sensing of a learner’s physical body and behavior, frequently blurring the lines between digital and physical learning environments. Consequently, responsible MMLA deployment requires a coordinated orchestration of privacy, fairness, transparency, and governance throughout the data lifecycle.
As illustrated in Figure 4, responsible MMLA deployment requires coordinated consideration of privacy, fairness, transparency, security, and institutional governance throughout the data lifecycle.

10.1. Privacy Within “In-Between” Learning Spaces

MMLA often operates in “in-between” spaces where the boundaries between public and private learning, research and instruction, and voluntary participation and institutional oversight are inherently fluid [9]. The deployment of cameras, microphones, and wearables introduces surveillance dynamics that may exceed a learner’s expectations, particularly when data is repurposed for secondary analysis.
These concerns are amplified in large-scale implementations and K–12 settings, where power asymmetries between the institution and the learner (especially minors) are most pronounced. While data collection may meet legal standards, its ethical acceptability hinges on the principles of proportionality, autonomy, and transparency. Privacy in MMLA must therefore be conceptualized not merely as a matter of data protection, but as a contextual negotiation of boundaries and trust.

10.2. Student-Centered FATE Principles

Current scholarship emphasizes a student-centered approach to Fairness, Accountability, Transparency, and Ethics (FATE). Empirical evidence indicates that student trust is predicated on several interrelated factors: the non-deceptive representation of behavior, role-based differentiated access, and consent processes that are continuous and revisable rather than static, one-time events [19].
Furthermore, transparency in MMLA extends beyond explaining algorithmic functionality. Trust is significantly undermined when systems obscure data uncertainty, signal loss, or the inherent limits of automated inference. Studies in authentic settings suggest that acknowledging system fallibility, rather than presenting analytics as absolute authority, fosters more effective reflective use and stakeholder buy-in [4]. In this light, uncertainty communication emerges as an ethical mandate rather than a mere technical enhancement.

10.3. Practical Recommendations for Trustworthy Deployment

Based on the surveyed literature, we distill the following recommendations for advancing the maturity of MMLA deployments:
  • Explicit Fusion Reporting: Researchers should document not only which modalities are collected, but precisely how and why they are fused. Obscure fusion practices impede both interpretability and scientific reproducibility [10].
  • Socio-Technical Alignment: Systems must be treated as socio-technical artifacts, ensuring that sensing choices and analytical models are strictly aligned with pedagogical goals and the contextual constraints of the learning environment [11,35].
  • Multidimensional Evaluation: Assessment must transcend predictive metrics (Accuracy, F1) to incorporate measures of perceived utility, stakeholder acceptance, and ethical resonance [4,19].
  • Design for Imperfection: Authentic classroom data is inherently noisy. Systems should be architected for robustness against signal missingness and bias, utilizing idealized datasets only for initial benchmarking [12].

10.4. Synthesis of Ethical Accumulation

As mapped in Table 6, ethical risks in MMLA are not localized in data collection but accumulate and transform across the analytic pipeline. Decisions made during sensing determine which behaviors are made visible, while subsequent choices in fusion and modeling dictate how these behaviors are interpreted and reinforced. This systemic view necessitates an Ethics-by-Design approach, ensuring that transparency, proportionality, and learner agency are integrated into the pipeline’s architecture rather than being addressed post-hoc.

11. Open Challenges and Research Directions

Despite rapid methodological advancements, the transition of Multimodal Learning Analytics (MMLA) from an experimental discipline to a mature field of educational practice hinges on addressing several interrelated technical, methodological, and socio-ethical hurdles. We identify five high-leverage research directions that are poised to shape the next generation of MMLA systems.
  • Standardized Benchmarks and Reproducible Data Ecosystems: the development of MMLA is currently constrained by the scarcity of open-access, well-documented multimodal datasets captured in authentic collaborative settings. The heterogeneity of sensing hardware, coupled with stringent privacy mandates, limits cross-study comparability. Establishing shared benchmarks, standardized evaluation protocols, and privacy-preserving data-sharing mechanisms (e.g., synthetic data or federated learning) is essential for cumulative knowledge building and the objective comparison of modeling architectures.
  • From Correlation to Causality: a significant portion of current research prioritizes predictive accuracy over pedagogical explanation. Advancing the field requires a shift toward models that explicitly incorporate temporal dynamics, causal dependencies, and theoretically motivated constructs [16]. Integrating learning sciences theory with sequence modeling and causal inference will enable researchers to move beyond identifying “what” happened toward explaining how and why specific multimodal patterns lead to learning gains.
  • Operationalizing FATE as System Requirements: while transparency and consent are frequently advocated, they are seldom implemented as measurable or testable system properties. Future research must operationalize FATE within the system’s architecture. This includes developing empirical metrics for explainability, effective uncertainty communication, and mechanisms for continuous, dynamic consent, ensuring these principles are treated as core functional requirements rather than post-hoc normative additions [19].
  • Resilient and Reconfigurable Infrastructure for Authentic Deployment: most current MMLA systems are tailored for bespoke, laboratory-style experimental setups. Scaling these technologies to diverse classrooms and professional training centers requires modular architectures that support dynamic sensor reconfiguration and graceful degradation, the ability of a system to maintain partial functionality under conditions of high noise or data missingness [31]. Infrastructure resilience is a prerequisite for moving beyond pilot studies toward sustainable, longitudinal deployment.
  • Human-in-the-Loop and Participatory Analytics: fully automated interpretations in MMLA carry the risk of misrepresentation and the subsequent erosion of stakeholder trust, particularly in high-stakes environments. Human-in-the-loop pipelines, where instructors and learners can inspect, contextualize, or override automated inferences, offer a robust path toward reconciling computational scalability with professional pedagogical judgment. A key design challenge lies in creating interfaces that support this participatory agency without inducing cognitive overload for the end-user.
In conclusion, the evolution of MMLA necessitates a paradigm shift from isolated technical advancements toward an integrated ecosystem of theoretically grounded and socially responsible analytics. Addressing these challenges is critical for transforming MMLA from a specialized research tool into a legitimate and transformative component of modern educational infrastructure.

12. Conclusions

Multimodal Learning Analytics represents a transformative expansion of the Learning Analytics field, providing the technical and conceptual tools to analyze learning as a multidimensional process—spanning embodied, linguistic, artifact-mediated, and social dimensions. By integrating heterogeneous data streams, MMLA renders visible critical facets of learning that remain inaccessible to traditional analytics, such as the nuances of non-verbal collaboration, the temporal regulation of affect, and the micro-dynamics of physical-digital interactions.
The current state of the art, as synthesized in this review, demonstrates significant methodological maturity. This is evident in the deployment of sophisticated sensing suites, the automation of high-level feature extraction via LLMs, and the development of fusion architectures capable of handling the inherent complexity of collaborative environments. Empirical evidence across tertiary, vocational, and K–12 settings further underscores the potential of MMLA to revolutionize formative assessment, providing real-time feedback and scaffolding reflection in ways that were previously unfeasible.
However, the transition from experimental success to systemic educational impact is contingent upon overcoming several critical barriers. The lack of standardized reporting for multimodal fusion, the sensitivity of models to environmental noise, and the absence of shared benchmarks remain significant technical bottlenecks. Perhaps more importantly, the ethical implications of high-resolution sensing necessitate a rigorous “Ethics-by-Design” approach to ensure that student agency and privacy are not compromised in the pursuit of analytical depth.
Moving forward, the maturation of MMLA will be driven by a shift toward socio-technical orchestration. This involves the development of scalable infrastructures, the adoption of participatory design frameworks, and an evaluative paradigm that prioritizes interpretability and pedagogical utility alongside predictive performance. By aligning technical capacity with ethical governance and instructional theory, MMLA is poised to evolve from a specialized research methodology into a foundational pillar of evidence-based teaching and learning in complex, real-world environments.

Author Contributions

Conceptualization, G.K. and S.K.; methodology, T.P.; validation, A.K.; writing—original draft preparation, G.K. and T.P.; writing—review and editing, S.K. and A.K. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the European Union through the Competitiveness Programme (ESPA 2021–2027).

Data Availability Statement

No new data were created or analyzed in this study.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Becerra, A.; Cobos, R.; Lang, C. Enhancing online learning by integrating biosensors and multimodal learning analytics for detecting and predicting student behaviour: A review. Behav. Inf. Technol. 2025, in press. [Google Scholar] [CrossRef] [Scilit]
  2. Patarakin, E.; Kutuzov, A.I.; Dvoretskaya, I. Multimodal learning analytics: A bibliometric and ontological analysis. Educ. Sci. J. 2025, 27, 33–71. [Google Scholar] [CrossRef] [Scilit]
  3. Noel, R.A.; Miranda, D.; Cechinel, C.; Riquelme, F.; Primo, T.T.; Munoz-Soto, R. Visualizing Collaboration in Teamwork: A Multimodal Learning Analytics Platform for Non-Verbal Communication. Appl. Sci. 2022, 12, 7499. [Google Scholar] [CrossRef] [Scilit]
  4. Yan, L.; Echeverria, V.; Jin, Y.; Fernandez-Nieto, G.M.; Zhao, L.; Li, X.; Alfredo, R.D.; Swiecki, Z.L.; Gašević, D.; Maldonado, R.M. Evidence-based multimodal learning analytics for feedback and reflection in collaborative learning. Br. J. Educ. Technol. 2024, 55, 1900–1925. [Google Scholar] [CrossRef] [Scilit]
  5. Moon, J.; Yeo, S.; Banihashem, S.K.; Noroozi, O. Using multimodal learning analytics as a formative assessment tool: Exploring collaborative dynamics in mathematics teacher education. J. Comput. Assist. Learn. 2024, 40, 2753–2771. [Google Scholar] [CrossRef] [Scilit]
  6. Rajarathinam, R.J.; Kang, J.; Palaguachi, C. 360-Degree Cameras vs Traditional Cameras in Multimodal Learning Analytics: Comparative Study of Facial Recognition and Pose Estimation. J. Educ. Data Min. 2025, 17, 157–182. [Google Scholar] [CrossRef]
  7. Lin, C.J.; Wang, W.; Lee, H.Y.; Li, P.; Huang, Y.M.; Wu, T.T. Advancing self-directed learning in STEM education: Integrating GPT-based learning aid with multimodal learning analytics. J. Res. Technol. Educ. 2025, in press. [Google Scholar] [CrossRef] [Scilit]
  8. Zhao, L.; Gašević, D.; Swiecki, Z.L.; Li, Y.; Lin, J.; Sha, L.; Yan, L.; Alfredo, R.D.; Li, X.; Maldonado, R.M. Towards automated transcribing and coding of embodied teamwork communication through multimodal learning analytics. Br. J. Educ. Technol. 2024, 55, 1673–1702. [Google Scholar] [CrossRef] [Scilit]
  9. Prinsloo, P.; Slade, S.; Khalil, M. Multimodal learning analytics—In-between student privacy and encroachment: A systematic review. Br. J. Educ. Technol. 2023, 54, 1566–1586. [Google Scholar] [CrossRef] [Scilit]
  10. Caskurlu, S.; Ocak, C.; Dai, C. The Scope of Multimodal Learning Analytics in K–8: A Systematic Review. J. Learn. Anal. 2025, 12, 224–236. [Google Scholar] [CrossRef] [Scilit]
  11. Ouhaichi, H.; Bahtijar, V.; Spikol, D. Exploring design considerations for multimodal learning analytics systems: An interview study. Front. Educ. 2024, 9, 1356537. [Google Scholar] [CrossRef] [Scilit]
  12. Chejara, P.; Prieto, L.P.; Dimitriadis, Y.A.; Rodríguez-Triana, M.J.; Ruiz-Calleja, A.; Kasepalu, R.; Shankar, S.K. The Impact of Attribute Noise on the Automated Estimation of Collaboration Quality Using Multimodal Learning Analytics in Authentic Classrooms. J. Learn. Anal. 2024, 11, 73–90. [Google Scholar] [CrossRef] [Scilit]
  13. Pei, B.; Xing, W.; Wang, M. Academic development of multimodal learning analytics: A bibliometric analysis. Interact. Learn. Environ. 2023, 31, 3543–3561. [Google Scholar] [CrossRef] [Scilit]
  14. Khor, E.T.; Tan, L.; Chan, S.H.L. Systematic Review on the Application of Multimodal Learning Analytics to Personalize Students’ Learning. Asten J. Teach. Educ. 2024, 1–14. [Google Scholar] [CrossRef] [Scilit]
  15. Emerson, A.J.; Cloude, E.B.; Azevedo, R.; Lester, J.C. Multimodal learning analytics for game-based learning. Br. J. Educ. Technol. 2020, 51, 1505–1526. [Google Scholar] [CrossRef] [Scilit]
  16. Yan, L.; Maldonado, R.M.; Swiecki, Z.L.; Zhao, L.; Li, X.; Gašević, D. Dissecting the Temporal Dynamics of Embodied Collaborative Learning Using Multimodal Learning Analytics. J. Educ. Psychol. 2024, 117, 106–133. [Google Scholar] [CrossRef] [Scilit]
  17. Huang, L.; Doleck, T.; Chen, B.; Huang, X.; Tan, C.; Lajoie, S.P.; Wang, M.H. Multimodal learning analytics for assessing teachers’ self-regulated learning in planning technology-integrated lessons in a computer-based environment. Educ. Inf. Technol. 2023, 28, 15823–15843. [Google Scholar] [CrossRef] [Scilit]
  18. Popov, V.; Nguyen, S.; Ochoa, X. Applying multimodal learning analytics to naturalistic recordings of clinical simulations: Towards an accurate and scalable pipeline for automated feedback generation. Learn. Instr. 2026, 102, 102267. [Google Scholar] [CrossRef] [Scilit]
  19. Jin, Y.; Echeverria, V.; Yan, L.; Zhao, L.; Alfredo, R.D.; Tsai, Y.S.; Gašević, D.; Maldonado, R.M. FATE in MMLA: A Student-Centred Exploration of Fairness, Accountability, Transparency, and Ethics in Multimodal Learning Analytics. J. Learn. Anal. 2024, 11, 6–23. [Google Scholar] [CrossRef] [Scilit]
  20. Shankar, S.K.; Rodríguez-Triana, M.J.; Ruiz-Calleja, A.; Prieto, L.P.; Chejara, P.; Martínez-Monés, A. Multimodal Data Value Chain (M-DVC): A Conceptual Tool to Support the Development of Multimodal Learning Analytics Solutions. Rev. Iberoam. Tecnol. Aprendiz. 2020, 15, 113–122. [Google Scholar] [CrossRef] [Scilit]
  21. Mangaroska, K.; Sharma, K.; Gašević, D.; Giannakos, M.N. Multimodal learning analytics to inform learning design: Lessons learned from computing education. J. Learn. Anal. 2020, 7, 79–97. [Google Scholar] [CrossRef] [Scilit]
  22. Munoz-Soto, R.; Barcelos, T.S.; Villarroel, R.H.; Guiñez, R.; Merino, E. Body Posture Visualizer to Support Multimodal Learning Analytics. IEEE Lat. Am. Trans. 2018, 16, 2706–2715. [Google Scholar] [CrossRef]
  23. Munoz-Soto, R.; Villarroel, R.H.; Barcelos, T.S.; de Souza, A.A.; Merino, E.; Guiñez, R.; Silva, L.A. Development of a software that supports multimodal learning analytics: A case study on oral presentations. J. Univers. Comput. Sci. 2018, 24, 149–170. [Google Scholar]
  24. Vujović, M.; Hernández-Leo, D.; Tassani, S.; Spikol, D. Round or rectangular tables for collaborative problem solving? A multimodal learning analytics study. Br. J. Educ. Technol. 2020, 51, 1597–1614. [Google Scholar] [CrossRef] [Scilit]
  25. Riquelme, F.; Munoz-Soto, R.; Lean, R.M.; Villarroel, R.H.; Barcelos, T.S.; de Albuquerque, V.H.C. Using multimodal learning analytics to study collaboration on discussion groups: A social network approach. Univers. Access Inf. Soc. 2019, 18, 633–643. [Google Scholar] [CrossRef] [Scilit]
  26. Kawamura, R.; Shirai, S.; Takemura, N.; Alizadeh, M.; Cukurova, M.; Takemura, H.; Nagahara, H. Detecting Drowsy Learners at the Wheel of e-Learning Platforms with Multimodal Learning Analytics. IEEE Access 2021, 9, 115165–115174. [Google Scholar] [CrossRef] [Scilit]
  27. Han, I.; Obeid, I.; Greco, D. Multimodal Learning Analytics and Neurofeedback for Optimizing Online Learners’ Self-Regulation. Technol. Knowl. Learn. 2023, 28, 1937–1943. [Google Scholar] [CrossRef] [Scilit]
  28. Sun, D.; Gutiérrez-Castillo, J.J.; Li, Y.; Zhu, C.; Zhou, Y. Using multimodal learning analytics to understand effects of block-based and text-based modalities on computer programming. J. Comput. Assist. Learn. 2024, 40, 1123–1136. [Google Scholar] [CrossRef] [Scilit]
  29. Xu, W.; Wu, Y.; Gutiérrez-Castillo, J.J. Multimodal learning analytics of collaborative patterns during pair programming in higher education. Int. J. Educ. Technol. High. Educ. 2023, 20, 8. [Google Scholar] [CrossRef] [Scilit]
  30. Gutiérrez-Castillo, J.J.; Dai, X.; Chen, S. Applying multimodal learning analytics to examine the immediate and delayed effects of instructor scaffoldings on small groups’ collaborative programming. Int. J. STEM Educ. 2022, 9, 45. [Google Scholar] [CrossRef] [Scilit]
  31. Huertas Celdrán, A.; Ruipérez-Valiente, J.A.; García Clemente, F.J.; Rodríguez-Triana, M.J.; Shankar, S.K.; Martínez Pérez, G.M. A scalable architecture for the dynamic deployment of multimodal learning analytics applications in smart classrooms. Sensors 2020, 20, 2923. [Google Scholar] [CrossRef] [Scilit]
  32. Spikol, D.; Ruffaldi, E.; Dabisias, G.; Cukurova, M. Supervised machine learning in multimodal learning analytics for estimating success in project-based learning. J. Comput. Assist. Learn. 2018, 34, 366–377. [Google Scholar] [CrossRef] [Scilit]
  33. Yusuf, A.; Md Noor, N.; Bello, S. Using multimodal learning analytics to model students’ learning behavior in animated programming classroom. Educ. Inf. Technol. 2024, 29, 6947–6990. [Google Scholar] [CrossRef] [Scilit]
  34. Smith, C.P.; King, B.; Gonzalez, D. Using multimodal learning analytics to identify patterns of interactions in a body-based mathematics activity. J. Interact. Learn. Res. 2016, 27, 355–379. [Google Scholar]
  35. Sellberg, C.; Sharma, A. Toward multimodal learning analytics in simulation-based collaborative learning: A design ethnography of maritime training. Int. J. Comput. Support. Collab. Learn. 2025, 20, 201–221. [Google Scholar] [CrossRef] [Scilit]
  36. Cornide-Reyes, H.C.C.; Noel, R.A.; Riquelme, F.; Gajardo, M.; Cechinel, C.; Lean, R.M.; Becerra, C.; Villarroel, R.H.; Munoz-Soto, R. Introducing low-cost sensors into the classroom settings: Improving the assessment in agile practices with multimodal learning analytics. Sensors 2019, 19, 3291. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. PRISMA-inspired flow diagram for study identification, screening, eligibility assessment, and inclusion.
Figure 1. PRISMA-inspired flow diagram for study identification, screening, eligibility assessment, and inclusion.
Futureinternet 18 00115 g001
Figure 2. Taxonomy of multimodal data sources in MMLA, including visual, auditory, interactional, physiological, and textual modalities.
Figure 2. Taxonomy of multimodal data sources in MMLA, including visual, auditory, interactional, physiological, and textual modalities.
Futureinternet 18 00115 g002
Figure 3. End-to-end pipeline of Multimodal Learning Analytics (MMLA), from heterogeneous data acquisition and preprocessing to multimodal fusion, analytics modeling and educational interventions.
Figure 3. End-to-end pipeline of Multimodal Learning Analytics (MMLA), from heterogeneous data acquisition and preprocessing to multimodal fusion, analytics modeling and educational interventions.
Futureinternet 18 00115 g003
Figure 4. Conceptual framework for responsible Multimodal Learning Analytics (MMLA), illustrating the relationships among data collection, privacy and consent, fairness, explainability, security, and institutional governance.
Figure 4. Conceptual framework for responsible Multimodal Learning Analytics (MMLA), illustrating the relationships among data collection, privacy and consent, fairness, explainability, security, and institutional governance.
Futureinternet 18 00115 g004
Table 1. Cross-modal comparison of common data sources in MMLA.
Table 1. Cross-modal comparison of common data sources in MMLA.
ModalityLearning Constructs Commonly ModeledKey Challenges and Limitations
Visual and spatial dataEmbodied engagement, participation, coordination, collaboration quality, spatial organization, physical interaction patternsOcclusion and field-of-view issues; sensitivity to lighting and classroom layout; computational cost; privacy concerns; ambiguity in mapping low-level features to learning constructs
Audio and speech dataParticipation balance, turn-taking, social roles, discourse structure, epistemic engagement, collaboration dynamicsSpeech overlap and diarization errors; domain-specific vocabulary; transcription inaccuracies; limited access to non-verbal meaning; privacy and consent challenges
Physiological and affective signalsAttention, cognitive load, affect, arousal, drowsiness, self-regulationSignal noise and individual variability; calibration requirements; intrusiveness; ethical and privacy concerns; weak construct validity without contextual grounding
Digital traces and learning productsTask progress, strategy use, performance outcomes, persistence, self-regulated behaviorsLimited observability of embodied and affective processes; coarse temporal resolution; construct under-specification when used in isolation
Multimodal combinationsHolistic learning processes combining cognitive, affective, social, and embodied dimensionsData synchronization and fusion complexity; missing or noisy modalities; interpretability of fused models; increased system complexity and deployment cost
Table 2. Typical deployment of multimodal data sources across learning contexts.
Table 2. Typical deployment of multimodal data sources across learning contexts.
ModalityK–12 EducationHigher EducationProfessional and Simulation-Based Training
Visual and spatial dataUsed selectively for posture, attention, and group interaction; often constrained by privacy regulations and classroom logisticsCommon in labs and collaborative classrooms to study embodied and group learning behaviorsWidely used in simulations (e.g., healthcare, maritime, aviation) to capture embodied teamwork and professional practices
Audio and speech dataApplied cautiously to measure participation and classroom talk; consent and data protection are central concernsFrequently used to analyze collaborative discourse, discussion quality, and team communicationCore modality for teamwork assessment, communication skills, and debriefing in high-fidelity simulations
Physiological and affective signalsRare; primarily used in controlled or short-term studies due to intrusiveness and ethical constraintsEmerging in experimental and online settings to study attention, cognitive load, and self-regulationMore feasible in simulations and training contexts where wearables are normalized and stakes justify richer sensing
Digital traces and learning productsDominant modality due to low intrusiveness and ease of deployment (e.g., LMS logs, assignments)Foundational data source, often combined with other modalities in mmlastudiesUsed alongside sensor data to anchor performance outcomes, task completion, and assessment
Multimodal combinationsLimited adoption; typically small-scale or exploratory deploymentsIncreasingly common in research-oriented courses and lab-based studiesMost mature and integrated deployments, supporting feedback, reflection, and formative assessment
Table 3. Mapping modeling approaches in MMLAs to learning goals and evaluation criteria.
Table 3. Mapping modeling approaches in MMLAs to learning goals and evaluation criteria.
Modeling ApproachPrimary Learning GoalsTypical Evaluation Criteria
Supervised prediction and classificationOutcome prediction (performance, success, engagement); state detection (attention, affect, drowsiness); formative assessment supportPredictive accuracy (e.g., F1, AUC); robustness to noise; generalization across contexts; interpretability of features; alignment with pedagogical constructs
Unsupervised learning and pattern miningDiscovery of behavioral profiles; exploration of learning strategies; hypothesis generation; learning design insightsCluster coherence and stability; interpretability; theoretical plausibility; triangulation with qualitative or outcome data
Temporal and sequential modelingModeling learning dynamics; phase transitions; regulation and coordination over time; process-oriented understandingTemporal validity; sensitivity to timing and granularity; explanatory value of sequences; correspondence with learning phases
Network-based modelingAnalysis of collaboration structure; participation balance; discourse and influence patterns; teamwork qualityStructural validity; alignment with social and epistemic theory; interpretability for educators; robustness to missing or noisy interactions
Hybrid and multimodal ensemble modelsHolistic modeling of cognitive, affective, social and embodied processes; actionable feedback generationBalance between performance and transparency; modality contribution analysis; user trust and perceived fairness; deployment feasibility
Table 4. Maturity and deployment readiness of multimodal learning analytics across learning contexts.
Table 4. Maturity and deployment readiness of multimodal learning analytics across learning contexts.
Learning ContextResearch MaturityTypical Deployment CharacteristicsKey Constraints and Risks
Professional and simulation-based trainingHighEnd-to-end pipelines; rich multimodal sensing; structured debriefing and feedback; repeated and longitudinal useSystem complexity; sensor calibration; scalability beyond simulation centers; maintaining realism and trust
Higher education (collaborative and lab-based learning)Medium–highMultimodal studies in labs and selected courses; increasing focus on feedback and reflection toolsFragmented architectures; limited longitudinal evidence; instructor workload and integration challenges
Online and blended learningMediumLightweight multimodal augmentation of logs (e.g., gaze, physiology); focus on detection and personalizationIntrusiveness; signal noise; learner consent; uneven access to sensing hardware
K–12 educationLow–mediumSmall-scale and exploratory deployments; emphasis on personalization and engagementStrong ethical and privacy constraints; limited transparency and fusion reporting; classroom logistics
Informal and ubiquitous learningLowMostly conceptual or prototype-driven work; opportunistic sensingData sparsity; lack of instructional alignment; governance and consent challenges
Table 5. Comparison of feedback modalities in multimodal learning analytics, summarizing their primary pedagogical functions, key strengths and main limitations or risks.
Table 5. Comparison of feedback modalities in multimodal learning analytics, summarizing their primary pedagogical functions, key strengths and main limitations or risks.
Feedback ModalityPrimary FunctionStrengths and AffordancesKey Limitations and Risks
Analytic dashboardsSummarize multimodal indicators for teachers and learners; support post-hoc reflection and instructional decision-makingHigh transparency; supports human interpretation and discussion; aligns well with formative and reflective practices; relatively low automation riskCognitive overload; requires analytic literacy; limited adaptivity; effectiveness depends on facilitation and context
Structured debriefing toolsScaffold guided reflection using multimodal evidence in facilitated settings (e.g., simulations)Strong pedagogical alignment; integrates analytics into existing instructional practices; supports sensemaking and shared interpretationResource-intensive; limited scalability; dependent on instructor expertise and time availability
AI tutors and generative learning aidsProvide personalized, adaptive feedback and guidance during or after learning activitiesHigh responsiveness and personalization; scalable; can support self-directed learning; continuous interaction generates rich analytic dataOpacity of model reasoning; risk of over-automation; trust, bias and accountability concerns; requires strong governance and transparency
Hybrid systemsCombine dashboards, debriefing and AI-driven feedback within a single pipelineBalances interpretability and adaptivity; supports multiple stakeholders and time scales; flexible deploymentIncreased system complexity; integration challenges; higher design and maintenance costs
Table 6. Mapping ethical risks and mitigation considerations across the multimodal learning analytics pipeline.
Table 6. Mapping ethical risks and mitigation considerations across the multimodal learning analytics pipeline.
Pipeline StagePrimary Ethical RisksKey Mitigation and Design Considerations
Sensing and data collectionIntrusive surveillance; blurred public/private boundaries; unequal power relations; inadequate or coerced consentProportional sensing; clear communication of purpose; opt-in and revisitable consent; context-sensitive deployment, especially with minors
Preprocessing and segmentationLoss of contextual meaning; biased filtering or exclusion of behaviors; amplification of sensor errorsTransparent preprocessing choices; documentation of signal loss and uncertainty; validation across diverse learners and settings
Feature extractionWeak or unjustified construct validity; misrepresentation of learner states; hidden assumptions in automated codingTheoretically grounded feature definitions; reporting extraction accuracy and limitations; triangulation with qualitative or human judgment
Multimodal fusionOpacity in how modalities are combined; unequal weighting of data sources; reduced interpretabilityExplicit reporting of fusion strategies; analysis of modality contributions; alignment of fusion choices with pedagogical intent
Modeling and inferenceBias and unfair outcomes; overgeneralization; automation bias in decision-makingRobustness testing; bias and fairness audits; preference for interpretable models in high-stakes uses; human-in-the-loop designs
Visualization and feedback deliveryOverconfidence in analytics; misinterpretation; stigmatization of learners; cognitive overloadUncertainty visualization; explanatory narratives; formative framing; role-appropriate access controls
Deployment and reuseFunction creep; secondary use without consent; erosion of trust over timeGovernance frameworks; data minimization and retention policies; continuous consent and stakeholder engagement
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Kostopoulos, G.; Kotsiantis, S.; Panagiotakopoulos, T.; Kameas, A. A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment. Future Internet 2026, 18, 115. https://doi.org/10.3390/fi18030115

AMA Style

Kostopoulos G, Kotsiantis S, Panagiotakopoulos T, Kameas A. A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment. Future Internet. 2026; 18(3):115. https://doi.org/10.3390/fi18030115

Chicago/Turabian Style

Kostopoulos, Georgios, Sotiris Kotsiantis, Theodor Panagiotakopoulos, and Achilles Kameas. 2026. "A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment" Future Internet 18, no. 3: 115. https://doi.org/10.3390/fi18030115

APA Style

Kostopoulos, G., Kotsiantis, S., Panagiotakopoulos, T., & Kameas, A. (2026). A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment. Future Internet, 18(3), 115. https://doi.org/10.3390/fi18030115

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop