A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment
Abstract
1. Introduction
2. Research Questions and Review Method
2.1. Research Questions
- RQ1 (Pipeline and architecture): What are the recurring architectural components and design patterns of MMLA systems deployed beyond controlled laboratory settings (e.g., sensing configurations, feature extraction, fusion strategies, and reliability mechanisms)?
- RQ2 (Loop-closing and pedagogical integration): How do MMLA systems translate multimodal inferences into actionable feedback for learners and educators (e.g., dashboards, debriefing tools, formative assessment workflows, and human-in-the-loop processes)?
- RQ3 (Deployment constraints and gaps): What technical, organizational, and ethical constraints most frequently impede systems-level deployment, and which research gaps remain under-addressed by existing reviews (e.g., fusion reporting standards, robustness in authentic settings, validity/fairness evidence, and governance)?
2.2. Traceable Synthesis Criteria
- Scope: Studies involve multimodal data (at least two modalities) and connect measurements to learning or collaboration constructs in an educational or training context.
- Systems-level focus: We prioritize work that specifies an end-to-end pipeline (from sensing through inference to interpretation/feedback), reports synchronization and fusion decisions, or evaluates a tool used in practice (e.g., dashboard, debriefing workflow, automated feedback).
- Deployment evidence: We distinguish between (i) controlled studies and (ii) authentic or semi-authentic deployments (e.g., classrooms, simulations, longitudinal use). Claims about feasibility or impact are grounded primarily in the latter category when available.
- Extraction template: For each included work, we code (a) context and participants, (b) modalities and sensors, (c) target constructs and operationalizations, (d) feature extraction approach (including LLM-based coding when applicable), (e) fusion strategy (early/late/hybrid/temporal), (f) validation/evaluation design, and (g) reported risks/mitigations (privacy, fairness, transparency, governance).
- Synthesis logic: Findings are synthesized by mapping coded evidence onto the reference architecture and the loop-closing taxonomy, and by summarizing recurring limitations as open gaps.
2.3. Study Identification and Screening (PRISMA-Inspired Workflow)
2.4. Quality/Rigor Appraisal
2.5. Positioning Relative to Prior Reviews and the Gap This Survey Fills
- From components to deployable systems: We synthesize evidence at the end-to-end pipeline level (sensing, preprocessing, alignment, feature extraction, fusion, and deployment), highlighting how temporal synchronization and robustness decisions constrain downstream pedagogical use.
- Closing-the-loop as a first-class object of analysis: Beyond prediction and classification, we analyze how studies implement actionability (feedback, reflection, and orchestration) and what forms of human-in-the-loop integration and evaluation are reported.
- Trustworthy deployment as socio-technical orchestration: We consolidate deployment constraints spanning technical reliability, validity/fairness, privacy/consent, and governance, and we surface under-reported aspects such as fusion reporting, uncertainty communication, and longitudinal evidence of impact.
3. Scope and Methodology
- Focused on educational or training contexts (formal, informal, or professional learning);
- Employed two or more data modalities (e.g., video, audio, physiological, interaction logs);
- Reported empirical findings, system designs, or methodological frameworks that connect multimodal evidence to learning constructs, instructional design, or learning outcomes.
- Addressed multimodal sensing without a learning or instructional focus;
- Focused exclusively on affect detection or surveillance without pedagogical grounding;
- Were purely technical computer vision or signal-processing papers without an educational setting, learning construct/outcome, or learning-oriented interpretation of the extracted features.
4. Background and Conceptual Foundations
4.1. Defining Multimodality in Learning Analytics
4.2. From Descriptive Analytics to Evidence-Based Interventions
4.3. Design and Specification Frameworks
5. Modalities, Sensors, and Data Sources
5.1. Visual and Spatial Data: Embodied Interaction and Proximity
5.2. Audio and Speech: Discourse Structure and Participation Dynamics
5.3. Physiological and Affective Signals: Unveiling Internal States
5.4. Digital Traces: The Contextual Anchor
6. The Multimodal Learning Analytics Pipeline
6.1. Reference Architecture: A Modular Perspective
6.2. Feature Extraction: Bridging Raw Signals and Learning Constructs
- Audio Pipelines: Focus on speech activity and conversational turn-taking, often operationalized through participation graphs or social network measures [25].
6.3. From Multimodal Indicators to Learning Constructs: Validity and Theoretical Grounding
6.3.1. Why the Mapping Is Non-Trivial
6.3.2. Problematizing Common Constructs
6.3.3. Minimum Validity and Reporting Expectations
- Construct definition: the learning-theory definition adopted (and why it fits the context).
- Operationalization: how the construct is mapped to indicators (features, windows, thresholds).
- Triangulation: alignment with independent evidence (e.g., expert ratings, validated instruments, performance measures, qualitative observations).
- Boundary conditions: when the indicator is expected to fail (task types, populations, environments).
- Interpretation limits: what the model output does not justify (e.g., causal claims, stable traits).
6.4. Multimodal Fusion Strategies
- Early Fusion (Feature Level): Combines features into a single high-dimensional vector before modeling. While it captures cross-modal interactions early, it is highly sensitive to noise and missing data.
- Late Fusion (Decision Level): Aggregates the outputs of independent unimodal models (e.g., via ensemble learning). This approach offers superior modularity and robustness to sensor failure but may overlook fine-grained inter-modality dependencies.
- Hybrid and Temporal Fusion: Align modalities along a temporal axis to study the dynamic coordination of multimodal events. This provides the most nuanced view of learning but significantly increases computational and modeling complexity.
6.4.1. Construct Validity and Cross-Modal Confounding
6.4.2. Interpretability and the Evidence Trail
6.4.3. Robustness to Missingness and Sensor Failure
6.4.4. Temporal Alignment as a Modeling Assumption
6.4.5. Fairness, Privacy, and Differential Modality Burden
6.5. Robustness and Reliability in Authentic Settings
6.6. End-to-End Interpretability and Explainability
- At the Sensing Level: Interpretability requires transparency regarding what data is captured and under what conditions.
- At the Modeling Level: A known tension exists between the high accuracy of black-box deep learning models (often used in early fusion) and the interpretability of late fusion models.
- At the Interface Level: Dashboards and feedback tools must communicate not only the result but also the underlying uncertainty and data limitations.
7. Modeling Approaches and Evaluation Practices
7.1. Supervised Learning: Predicting Outcomes and Detecting States
7.2. Unsupervised Learning and Pattern Discovery
7.3. Temporal and Network-Based Analytics for Collaboration
7.4. Evaluation Paradigms: Beyond Predictive Accuracy
7.5. Strength of Evidence for Pedagogical Impact
- Tier 1: Technical proof-of-concept. Studies that establish the feasibility of multimodal capture and prediction (accuracy/F1, robustness under noise) but do not test downstream instructional use or learning effects.
- Tier 2: Classroom-facing tools with usability/acceptability evidence. Studies that integrate MMLA into dashboards, debriefing workflows, or feedback tools and report user perceptions (usefulness, trust, interpretability), but provide limited causal or longitudinal evidence of learning gains.
- Tier 3: Validated instructional interventions. Studies that evaluate MMLA-enabled interventions with stronger designs (e.g., longitudinal deployments, quasi-experimental or experimental comparisons, pre/post learning measures, or instructor decision outcomes) enable more credible claims about pedagogical impact.
8. Applications by Learning Context
8.1. Professional Simulations: The Frontier of High-Stakes Training
8.2. STEM and Programming: Decoding Process over Product
8.3. Online Learning: Enhancing Observability and Self-Regulation
8.4. K–12 Education: Personalization Under Ethical Constraints
8.5. Synthesis of Deployment Maturity
9. Systems That Close the Loop: Feedback, Reflection, and Formative Assessment
9.1. Dashboards and Reflective Debriefing Tools
9.2. Formative Assessment in Collaborative Learning
9.3. Generative AI and MMLA-Enabled Learning Aids
10. Ethics, Privacy, and Trustworthy Deployment
10.1. Privacy Within “In-Between” Learning Spaces
10.2. Student-Centered FATE Principles
10.3. Practical Recommendations for Trustworthy Deployment
- Explicit Fusion Reporting: Researchers should document not only which modalities are collected, but precisely how and why they are fused. Obscure fusion practices impede both interpretability and scientific reproducibility [10].
- Design for Imperfection: Authentic classroom data is inherently noisy. Systems should be architected for robustness against signal missingness and bias, utilizing idealized datasets only for initial benchmarking [12].
10.4. Synthesis of Ethical Accumulation
11. Open Challenges and Research Directions
- Standardized Benchmarks and Reproducible Data Ecosystems: the development of MMLA is currently constrained by the scarcity of open-access, well-documented multimodal datasets captured in authentic collaborative settings. The heterogeneity of sensing hardware, coupled with stringent privacy mandates, limits cross-study comparability. Establishing shared benchmarks, standardized evaluation protocols, and privacy-preserving data-sharing mechanisms (e.g., synthetic data or federated learning) is essential for cumulative knowledge building and the objective comparison of modeling architectures.
- From Correlation to Causality: a significant portion of current research prioritizes predictive accuracy over pedagogical explanation. Advancing the field requires a shift toward models that explicitly incorporate temporal dynamics, causal dependencies, and theoretically motivated constructs [16]. Integrating learning sciences theory with sequence modeling and causal inference will enable researchers to move beyond identifying “what” happened toward explaining how and why specific multimodal patterns lead to learning gains.
- Operationalizing FATE as System Requirements: while transparency and consent are frequently advocated, they are seldom implemented as measurable or testable system properties. Future research must operationalize FATE within the system’s architecture. This includes developing empirical metrics for explainability, effective uncertainty communication, and mechanisms for continuous, dynamic consent, ensuring these principles are treated as core functional requirements rather than post-hoc normative additions [19].
- Resilient and Reconfigurable Infrastructure for Authentic Deployment: most current MMLA systems are tailored for bespoke, laboratory-style experimental setups. Scaling these technologies to diverse classrooms and professional training centers requires modular architectures that support dynamic sensor reconfiguration and graceful degradation, the ability of a system to maintain partial functionality under conditions of high noise or data missingness [31]. Infrastructure resilience is a prerequisite for moving beyond pilot studies toward sustainable, longitudinal deployment.
- Human-in-the-Loop and Participatory Analytics: fully automated interpretations in MMLA carry the risk of misrepresentation and the subsequent erosion of stakeholder trust, particularly in high-stakes environments. Human-in-the-loop pipelines, where instructors and learners can inspect, contextualize, or override automated inferences, offer a robust path toward reconciling computational scalability with professional pedagogical judgment. A key design challenge lies in creating interfaces that support this participatory agency without inducing cognitive overload for the end-user.
12. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Becerra, A.; Cobos, R.; Lang, C. Enhancing online learning by integrating biosensors and multimodal learning analytics for detecting and predicting student behaviour: A review. Behav. Inf. Technol. 2025, in press. [Google Scholar] [CrossRef] [Scilit]
- Patarakin, E.; Kutuzov, A.I.; Dvoretskaya, I. Multimodal learning analytics: A bibliometric and ontological analysis. Educ. Sci. J. 2025, 27, 33–71. [Google Scholar] [CrossRef] [Scilit]
- Noel, R.A.; Miranda, D.; Cechinel, C.; Riquelme, F.; Primo, T.T.; Munoz-Soto, R. Visualizing Collaboration in Teamwork: A Multimodal Learning Analytics Platform for Non-Verbal Communication. Appl. Sci. 2022, 12, 7499. [Google Scholar] [CrossRef] [Scilit]
- Yan, L.; Echeverria, V.; Jin, Y.; Fernandez-Nieto, G.M.; Zhao, L.; Li, X.; Alfredo, R.D.; Swiecki, Z.L.; Gašević, D.; Maldonado, R.M. Evidence-based multimodal learning analytics for feedback and reflection in collaborative learning. Br. J. Educ. Technol. 2024, 55, 1900–1925. [Google Scholar] [CrossRef] [Scilit]
- Moon, J.; Yeo, S.; Banihashem, S.K.; Noroozi, O. Using multimodal learning analytics as a formative assessment tool: Exploring collaborative dynamics in mathematics teacher education. J. Comput. Assist. Learn. 2024, 40, 2753–2771. [Google Scholar] [CrossRef] [Scilit]
- Rajarathinam, R.J.; Kang, J.; Palaguachi, C. 360-Degree Cameras vs Traditional Cameras in Multimodal Learning Analytics: Comparative Study of Facial Recognition and Pose Estimation. J. Educ. Data Min. 2025, 17, 157–182. [Google Scholar] [CrossRef]
- Lin, C.J.; Wang, W.; Lee, H.Y.; Li, P.; Huang, Y.M.; Wu, T.T. Advancing self-directed learning in STEM education: Integrating GPT-based learning aid with multimodal learning analytics. J. Res. Technol. Educ. 2025, in press. [Google Scholar] [CrossRef] [Scilit]
- Zhao, L.; Gašević, D.; Swiecki, Z.L.; Li, Y.; Lin, J.; Sha, L.; Yan, L.; Alfredo, R.D.; Li, X.; Maldonado, R.M. Towards automated transcribing and coding of embodied teamwork communication through multimodal learning analytics. Br. J. Educ. Technol. 2024, 55, 1673–1702. [Google Scholar] [CrossRef] [Scilit]
- Prinsloo, P.; Slade, S.; Khalil, M. Multimodal learning analytics—In-between student privacy and encroachment: A systematic review. Br. J. Educ. Technol. 2023, 54, 1566–1586. [Google Scholar] [CrossRef] [Scilit]
- Caskurlu, S.; Ocak, C.; Dai, C. The Scope of Multimodal Learning Analytics in K–8: A Systematic Review. J. Learn. Anal. 2025, 12, 224–236. [Google Scholar] [CrossRef] [Scilit]
- Ouhaichi, H.; Bahtijar, V.; Spikol, D. Exploring design considerations for multimodal learning analytics systems: An interview study. Front. Educ. 2024, 9, 1356537. [Google Scholar] [CrossRef] [Scilit]
- Chejara, P.; Prieto, L.P.; Dimitriadis, Y.A.; Rodríguez-Triana, M.J.; Ruiz-Calleja, A.; Kasepalu, R.; Shankar, S.K. The Impact of Attribute Noise on the Automated Estimation of Collaboration Quality Using Multimodal Learning Analytics in Authentic Classrooms. J. Learn. Anal. 2024, 11, 73–90. [Google Scholar] [CrossRef] [Scilit]
- Pei, B.; Xing, W.; Wang, M. Academic development of multimodal learning analytics: A bibliometric analysis. Interact. Learn. Environ. 2023, 31, 3543–3561. [Google Scholar] [CrossRef] [Scilit]
- Khor, E.T.; Tan, L.; Chan, S.H.L. Systematic Review on the Application of Multimodal Learning Analytics to Personalize Students’ Learning. Asten J. Teach. Educ. 2024, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Emerson, A.J.; Cloude, E.B.; Azevedo, R.; Lester, J.C. Multimodal learning analytics for game-based learning. Br. J. Educ. Technol. 2020, 51, 1505–1526. [Google Scholar] [CrossRef] [Scilit]
- Yan, L.; Maldonado, R.M.; Swiecki, Z.L.; Zhao, L.; Li, X.; Gašević, D. Dissecting the Temporal Dynamics of Embodied Collaborative Learning Using Multimodal Learning Analytics. J. Educ. Psychol. 2024, 117, 106–133. [Google Scholar] [CrossRef] [Scilit]
- Huang, L.; Doleck, T.; Chen, B.; Huang, X.; Tan, C.; Lajoie, S.P.; Wang, M.H. Multimodal learning analytics for assessing teachers’ self-regulated learning in planning technology-integrated lessons in a computer-based environment. Educ. Inf. Technol. 2023, 28, 15823–15843. [Google Scholar] [CrossRef] [Scilit]
- Popov, V.; Nguyen, S.; Ochoa, X. Applying multimodal learning analytics to naturalistic recordings of clinical simulations: Towards an accurate and scalable pipeline for automated feedback generation. Learn. Instr. 2026, 102, 102267. [Google Scholar] [CrossRef] [Scilit]
- Jin, Y.; Echeverria, V.; Yan, L.; Zhao, L.; Alfredo, R.D.; Tsai, Y.S.; Gašević, D.; Maldonado, R.M. FATE in MMLA: A Student-Centred Exploration of Fairness, Accountability, Transparency, and Ethics in Multimodal Learning Analytics. J. Learn. Anal. 2024, 11, 6–23. [Google Scholar] [CrossRef] [Scilit]
- Shankar, S.K.; Rodríguez-Triana, M.J.; Ruiz-Calleja, A.; Prieto, L.P.; Chejara, P.; Martínez-Monés, A. Multimodal Data Value Chain (M-DVC): A Conceptual Tool to Support the Development of Multimodal Learning Analytics Solutions. Rev. Iberoam. Tecnol. Aprendiz. 2020, 15, 113–122. [Google Scholar] [CrossRef] [Scilit]
- Mangaroska, K.; Sharma, K.; Gašević, D.; Giannakos, M.N. Multimodal learning analytics to inform learning design: Lessons learned from computing education. J. Learn. Anal. 2020, 7, 79–97. [Google Scholar] [CrossRef] [Scilit]
- Munoz-Soto, R.; Barcelos, T.S.; Villarroel, R.H.; Guiñez, R.; Merino, E. Body Posture Visualizer to Support Multimodal Learning Analytics. IEEE Lat. Am. Trans. 2018, 16, 2706–2715. [Google Scholar] [CrossRef]
- Munoz-Soto, R.; Villarroel, R.H.; Barcelos, T.S.; de Souza, A.A.; Merino, E.; Guiñez, R.; Silva, L.A. Development of a software that supports multimodal learning analytics: A case study on oral presentations. J. Univers. Comput. Sci. 2018, 24, 149–170. [Google Scholar]
- Vujović, M.; Hernández-Leo, D.; Tassani, S.; Spikol, D. Round or rectangular tables for collaborative problem solving? A multimodal learning analytics study. Br. J. Educ. Technol. 2020, 51, 1597–1614. [Google Scholar] [CrossRef] [Scilit]
- Riquelme, F.; Munoz-Soto, R.; Lean, R.M.; Villarroel, R.H.; Barcelos, T.S.; de Albuquerque, V.H.C. Using multimodal learning analytics to study collaboration on discussion groups: A social network approach. Univers. Access Inf. Soc. 2019, 18, 633–643. [Google Scholar] [CrossRef] [Scilit]
- Kawamura, R.; Shirai, S.; Takemura, N.; Alizadeh, M.; Cukurova, M.; Takemura, H.; Nagahara, H. Detecting Drowsy Learners at the Wheel of e-Learning Platforms with Multimodal Learning Analytics. IEEE Access 2021, 9, 115165–115174. [Google Scholar] [CrossRef] [Scilit]
- Han, I.; Obeid, I.; Greco, D. Multimodal Learning Analytics and Neurofeedback for Optimizing Online Learners’ Self-Regulation. Technol. Knowl. Learn. 2023, 28, 1937–1943. [Google Scholar] [CrossRef] [Scilit]
- Sun, D.; Gutiérrez-Castillo, J.J.; Li, Y.; Zhu, C.; Zhou, Y. Using multimodal learning analytics to understand effects of block-based and text-based modalities on computer programming. J. Comput. Assist. Learn. 2024, 40, 1123–1136. [Google Scholar] [CrossRef] [Scilit]
- Xu, W.; Wu, Y.; Gutiérrez-Castillo, J.J. Multimodal learning analytics of collaborative patterns during pair programming in higher education. Int. J. Educ. Technol. High. Educ. 2023, 20, 8. [Google Scholar] [CrossRef] [Scilit]
- Gutiérrez-Castillo, J.J.; Dai, X.; Chen, S. Applying multimodal learning analytics to examine the immediate and delayed effects of instructor scaffoldings on small groups’ collaborative programming. Int. J. STEM Educ. 2022, 9, 45. [Google Scholar] [CrossRef] [Scilit]
- Huertas Celdrán, A.; Ruipérez-Valiente, J.A.; García Clemente, F.J.; Rodríguez-Triana, M.J.; Shankar, S.K.; Martínez Pérez, G.M. A scalable architecture for the dynamic deployment of multimodal learning analytics applications in smart classrooms. Sensors 2020, 20, 2923. [Google Scholar] [CrossRef] [Scilit]
- Spikol, D.; Ruffaldi, E.; Dabisias, G.; Cukurova, M. Supervised machine learning in multimodal learning analytics for estimating success in project-based learning. J. Comput. Assist. Learn. 2018, 34, 366–377. [Google Scholar] [CrossRef] [Scilit]
- Yusuf, A.; Md Noor, N.; Bello, S. Using multimodal learning analytics to model students’ learning behavior in animated programming classroom. Educ. Inf. Technol. 2024, 29, 6947–6990. [Google Scholar] [CrossRef] [Scilit]
- Smith, C.P.; King, B.; Gonzalez, D. Using multimodal learning analytics to identify patterns of interactions in a body-based mathematics activity. J. Interact. Learn. Res. 2016, 27, 355–379. [Google Scholar]
- Sellberg, C.; Sharma, A. Toward multimodal learning analytics in simulation-based collaborative learning: A design ethnography of maritime training. Int. J. Comput. Support. Collab. Learn. 2025, 20, 201–221. [Google Scholar] [CrossRef] [Scilit]
- Cornide-Reyes, H.C.C.; Noel, R.A.; Riquelme, F.; Gajardo, M.; Cechinel, C.; Lean, R.M.; Becerra, C.; Villarroel, R.H.; Munoz-Soto, R. Introducing low-cost sensors into the classroom settings: Improving the assessment in agile practices with multimodal learning analytics. Sensors 2019, 19, 3291. [Google Scholar] [CrossRef] [Scilit] [PubMed]




| Modality | Learning Constructs Commonly Modeled | Key Challenges and Limitations |
|---|---|---|
| Visual and spatial data | Embodied engagement, participation, coordination, collaboration quality, spatial organization, physical interaction patterns | Occlusion and field-of-view issues; sensitivity to lighting and classroom layout; computational cost; privacy concerns; ambiguity in mapping low-level features to learning constructs |
| Audio and speech data | Participation balance, turn-taking, social roles, discourse structure, epistemic engagement, collaboration dynamics | Speech overlap and diarization errors; domain-specific vocabulary; transcription inaccuracies; limited access to non-verbal meaning; privacy and consent challenges |
| Physiological and affective signals | Attention, cognitive load, affect, arousal, drowsiness, self-regulation | Signal noise and individual variability; calibration requirements; intrusiveness; ethical and privacy concerns; weak construct validity without contextual grounding |
| Digital traces and learning products | Task progress, strategy use, performance outcomes, persistence, self-regulated behaviors | Limited observability of embodied and affective processes; coarse temporal resolution; construct under-specification when used in isolation |
| Multimodal combinations | Holistic learning processes combining cognitive, affective, social, and embodied dimensions | Data synchronization and fusion complexity; missing or noisy modalities; interpretability of fused models; increased system complexity and deployment cost |
| Modality | K–12 Education | Higher Education | Professional and Simulation-Based Training |
|---|---|---|---|
| Visual and spatial data | Used selectively for posture, attention, and group interaction; often constrained by privacy regulations and classroom logistics | Common in labs and collaborative classrooms to study embodied and group learning behaviors | Widely used in simulations (e.g., healthcare, maritime, aviation) to capture embodied teamwork and professional practices |
| Audio and speech data | Applied cautiously to measure participation and classroom talk; consent and data protection are central concerns | Frequently used to analyze collaborative discourse, discussion quality, and team communication | Core modality for teamwork assessment, communication skills, and debriefing in high-fidelity simulations |
| Physiological and affective signals | Rare; primarily used in controlled or short-term studies due to intrusiveness and ethical constraints | Emerging in experimental and online settings to study attention, cognitive load, and self-regulation | More feasible in simulations and training contexts where wearables are normalized and stakes justify richer sensing |
| Digital traces and learning products | Dominant modality due to low intrusiveness and ease of deployment (e.g., LMS logs, assignments) | Foundational data source, often combined with other modalities in mmlastudies | Used alongside sensor data to anchor performance outcomes, task completion, and assessment |
| Multimodal combinations | Limited adoption; typically small-scale or exploratory deployments | Increasingly common in research-oriented courses and lab-based studies | Most mature and integrated deployments, supporting feedback, reflection, and formative assessment |
| Modeling Approach | Primary Learning Goals | Typical Evaluation Criteria |
|---|---|---|
| Supervised prediction and classification | Outcome prediction (performance, success, engagement); state detection (attention, affect, drowsiness); formative assessment support | Predictive accuracy (e.g., F1, AUC); robustness to noise; generalization across contexts; interpretability of features; alignment with pedagogical constructs |
| Unsupervised learning and pattern mining | Discovery of behavioral profiles; exploration of learning strategies; hypothesis generation; learning design insights | Cluster coherence and stability; interpretability; theoretical plausibility; triangulation with qualitative or outcome data |
| Temporal and sequential modeling | Modeling learning dynamics; phase transitions; regulation and coordination over time; process-oriented understanding | Temporal validity; sensitivity to timing and granularity; explanatory value of sequences; correspondence with learning phases |
| Network-based modeling | Analysis of collaboration structure; participation balance; discourse and influence patterns; teamwork quality | Structural validity; alignment with social and epistemic theory; interpretability for educators; robustness to missing or noisy interactions |
| Hybrid and multimodal ensemble models | Holistic modeling of cognitive, affective, social and embodied processes; actionable feedback generation | Balance between performance and transparency; modality contribution analysis; user trust and perceived fairness; deployment feasibility |
| Learning Context | Research Maturity | Typical Deployment Characteristics | Key Constraints and Risks |
|---|---|---|---|
| Professional and simulation-based training | High | End-to-end pipelines; rich multimodal sensing; structured debriefing and feedback; repeated and longitudinal use | System complexity; sensor calibration; scalability beyond simulation centers; maintaining realism and trust |
| Higher education (collaborative and lab-based learning) | Medium–high | Multimodal studies in labs and selected courses; increasing focus on feedback and reflection tools | Fragmented architectures; limited longitudinal evidence; instructor workload and integration challenges |
| Online and blended learning | Medium | Lightweight multimodal augmentation of logs (e.g., gaze, physiology); focus on detection and personalization | Intrusiveness; signal noise; learner consent; uneven access to sensing hardware |
| K–12 education | Low–medium | Small-scale and exploratory deployments; emphasis on personalization and engagement | Strong ethical and privacy constraints; limited transparency and fusion reporting; classroom logistics |
| Informal and ubiquitous learning | Low | Mostly conceptual or prototype-driven work; opportunistic sensing | Data sparsity; lack of instructional alignment; governance and consent challenges |
| Feedback Modality | Primary Function | Strengths and Affordances | Key Limitations and Risks |
|---|---|---|---|
| Analytic dashboards | Summarize multimodal indicators for teachers and learners; support post-hoc reflection and instructional decision-making | High transparency; supports human interpretation and discussion; aligns well with formative and reflective practices; relatively low automation risk | Cognitive overload; requires analytic literacy; limited adaptivity; effectiveness depends on facilitation and context |
| Structured debriefing tools | Scaffold guided reflection using multimodal evidence in facilitated settings (e.g., simulations) | Strong pedagogical alignment; integrates analytics into existing instructional practices; supports sensemaking and shared interpretation | Resource-intensive; limited scalability; dependent on instructor expertise and time availability |
| AI tutors and generative learning aids | Provide personalized, adaptive feedback and guidance during or after learning activities | High responsiveness and personalization; scalable; can support self-directed learning; continuous interaction generates rich analytic data | Opacity of model reasoning; risk of over-automation; trust, bias and accountability concerns; requires strong governance and transparency |
| Hybrid systems | Combine dashboards, debriefing and AI-driven feedback within a single pipeline | Balances interpretability and adaptivity; supports multiple stakeholders and time scales; flexible deployment | Increased system complexity; integration challenges; higher design and maintenance costs |
| Pipeline Stage | Primary Ethical Risks | Key Mitigation and Design Considerations |
|---|---|---|
| Sensing and data collection | Intrusive surveillance; blurred public/private boundaries; unequal power relations; inadequate or coerced consent | Proportional sensing; clear communication of purpose; opt-in and revisitable consent; context-sensitive deployment, especially with minors |
| Preprocessing and segmentation | Loss of contextual meaning; biased filtering or exclusion of behaviors; amplification of sensor errors | Transparent preprocessing choices; documentation of signal loss and uncertainty; validation across diverse learners and settings |
| Feature extraction | Weak or unjustified construct validity; misrepresentation of learner states; hidden assumptions in automated coding | Theoretically grounded feature definitions; reporting extraction accuracy and limitations; triangulation with qualitative or human judgment |
| Multimodal fusion | Opacity in how modalities are combined; unequal weighting of data sources; reduced interpretability | Explicit reporting of fusion strategies; analysis of modality contributions; alignment of fusion choices with pedagogical intent |
| Modeling and inference | Bias and unfair outcomes; overgeneralization; automation bias in decision-making | Robustness testing; bias and fairness audits; preference for interpretable models in high-stakes uses; human-in-the-loop designs |
| Visualization and feedback delivery | Overconfidence in analytics; misinterpretation; stigmatization of learners; cognitive overload | Uncertainty visualization; explanatory narratives; formative framing; role-appropriate access controls |
| Deployment and reuse | Function creep; secondary use without consent; erosion of trust over time | Governance frameworks; data minimization and retention policies; continuous consent and stakeholder engagement |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kostopoulos, G.; Kotsiantis, S.; Panagiotakopoulos, T.; Kameas, A. A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment. Future Internet 2026, 18, 115. https://doi.org/10.3390/fi18030115
Kostopoulos G, Kotsiantis S, Panagiotakopoulos T, Kameas A. A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment. Future Internet. 2026; 18(3):115. https://doi.org/10.3390/fi18030115
Chicago/Turabian StyleKostopoulos, Georgios, Sotiris Kotsiantis, Theodor Panagiotakopoulos, and Achilles Kameas. 2026. "A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment" Future Internet 18, no. 3: 115. https://doi.org/10.3390/fi18030115
APA StyleKostopoulos, G., Kotsiantis, S., Panagiotakopoulos, T., & Kameas, A. (2026). A Survey of Multimodal Learning Analytics: Data, Methods, Systems, and Responsible Deployment. Future Internet, 18(3), 115. https://doi.org/10.3390/fi18030115

