Next Issue
Volume 10, February
Previous Issue
Volume 9, December
 
 

Multimodal Technol. Interact., Volume 10, Issue 1 (January 2026) – 11 articles

Cover Story (view full-size image): Extended Reality (XR) is transforming training in high-stakes domains like medicine and emergency response. However, AI-driven personalization can be brittle without human oversight. This article presents a framework for “Adaptive Realities”, integrating Human-in-the-Loop (HITL) supervision with real-time AI adaptation. By coupling multimodal sensing with adjustable autonomy and explainable AI, the system ensures trainers remain central to the decision-making process. Across six studies, the authors demonstrate how pedagogical guardrails and transparent dashboards co-produce training effectiveness and calibrated trust. The results chart a path toward more accountable, human-centered XR systems for safety-critical practice. View this paper
  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Select all
Export citation of selected articles as:
24 pages, 6152 KB  
Article
Adaptive Realities: Human-in-the-Loop AI for Trustworthy XR Training in Safety-Critical Domains
by Daniele Pretolesi, Georg Regal, Helmut Schrom-Feiertag and Manfred Tscheligi
Multimodal Technol. Interact. 2026, 10(1), 11; https://doi.org/10.3390/mti10010011 - 22 Jan 2026
Viewed by 2224
Abstract
Extended Reality (XR) technologies have matured into powerful tools for training in high-stakes domains, from emergency response to search and rescue. Yet current systems often struggle to balance real-time AI-driven personalisation with the need for human oversight and calibrated trust. This article synthesizes [...] Read more.
Extended Reality (XR) technologies have matured into powerful tools for training in high-stakes domains, from emergency response to search and rescue. Yet current systems often struggle to balance real-time AI-driven personalisation with the need for human oversight and calibrated trust. This article synthesizes the programmatic contributions of a multi-study doctoral project to advance a design-and-evaluation framework for trustworthy adaptive XR training. Across six studies, we explored (i) recommender-driven scenario adaptation based on multimodal performance and physiological signals, (ii) persuasive dashboards for trainers, (iii) architectures for AI-supported XR training in medical mass-casualty contexts, (iv) theoretical and practical integration of Human-in-the-Loop (HITL) supervision, (v) user trust and over-reliance in the face of misleading AI suggestions, and (vi) the role of interaction modality in shaping workload, explainability, and trust in human–robot collaboration. Together, these investigations show how adaptive policies, transparent explanation, and adjustable autonomy can be orchestrated into a single adaptation loop that maintains trainee engagement, improves learning outcomes, and preserves trainer agency. We conclude with design guidelines and a research agenda for extending trustworthy XR training into safety-critical environments. Full article
Show Figures

Figure 1

24 pages, 476 KB  
Article
APAR: A Structural Design and Guidance Framework for Gamification in Education Based on Motivation Theories
by J. Carlos López-Ardao, Miguel Rodríguez-Pérez, Sergio Herrería-Alonso, M. Estrella Sousa-Vieira, Alfonso Lago Ferreiro, Andrés Suárez-González and Raúl F. Rodríguez-Rubio
Multimodal Technol. Interact. 2026, 10(1), 10; https://doi.org/10.3390/mti10010010 - 10 Jan 2026
Cited by 1 | Viewed by 2550
Abstract
Gamification is widely used to enhance student motivation, yet many educational design proposals remain conceptual and provide limited operational guidance for digital learning environments. This paper introduces APAR (Activities, Points, Achievements and Rewards), a content-independent structural framework for designing and implementing educational gamification [...] Read more.
Gamification is widely used to enhance student motivation, yet many educational design proposals remain conceptual and provide limited operational guidance for digital learning environments. This paper introduces APAR (Activities, Points, Achievements and Rewards), a content-independent structural framework for designing and implementing educational gamification in learning platforms. Grounded in motivation theories (including Self-Determination Theory and Relatedness–Autonomy–Mastery–Purpose) and reward taxonomies (Status, Access, Power and Stuff), APAR distinguishes high-level design constructs from concrete game elements (e.g., points, badges and leaderboards) and provides a systematic design loop linking learning activities, feedback, intermediate goals and reinforcement. The contribution includes (i) a mapping table relating each APAR construct to motivation models, supported dynamics and typical learning-platform implementations; (ii) an actionable design guide; and (iii) an empirical illustration implemented in Moodle in a higher-education Computer Networks course. In this setting, the proportion of enrolled students taking the final exam increased from 58% to 72% in the first year, and the proportion of enrolled students passing increased from 17% to 38%; in 2022–2023 these values were 70% and 39%, respectively (56% of exam takers passed). While the use case relies on quantitative course-level indicators and is observational, the findings support the potential of structural gamification as an integrated methodological tool and motivate further mixed-method validations. Full article
Show Figures

Figure 1

26 pages, 1616 KB  
Systematic Review
AI-Powered Procedural Haptics for Narrative VR: A Systematic Literature Review
by Vimala Perumal and Zeeshan Jawed Shah
Multimodal Technol. Interact. 2026, 10(1), 9; https://doi.org/10.3390/mti10010009 - 9 Jan 2026
Viewed by 2334
Abstract
Haptic feedback is important for narrative virtual reality (VR), yet authoring remains costly and difficult to scale due to device-specific tuning, placement constraints, and the need for semantically congruent timing. We systematically reviewed user studies on haptics in narrative VR to establish an [...] Read more.
Haptic feedback is important for narrative virtual reality (VR), yet authoring remains costly and difficult to scale due to device-specific tuning, placement constraints, and the need for semantically congruent timing. We systematically reviewed user studies on haptics in narrative VR to establish an empirical baseline and identify gaps for AI-powered procedural haptics. Following PRISMA 2020, we searched IEEE Xplore, ACM Digital Library, Scopus, Web of Science, PubMed, and PsycINFO (English; human participants; haptics synchronized to narrative events) and performed backward/forward citation chasing (final search: 31 July 2025). We also conducted a parallel scoping scan of grey literature (arXiv and CHI/SIGGRAPH workshops/demos), finalized on 7 September 2025; these records are summarized separately and were not included in the evidence synthesis. Of 493 records screened, 26 full texts were assessed, and 10 studies were included. Quantitatively, presence improved in 6/8 studies that measured it and immersion improved in 3/3; sample sizes ranged 8–108. Across varied modalities and placements, haptics improved presence and immersion and often enhanced affect; validated measures of narrative comprehension were rare. None of the included studies evaluated AI-generated procedural haptics in user studies. We conclude by proposing a structured, three-phase research roadmap designed to bridge this critical gap, moving the field from theoretical promise to the empirical validation of intelligent systems capable of making rich, adaptive, and scalable haptic narratives a reality. Full article
Show Figures

Graphical abstract

28 pages, 515 KB  
Review
From Cues to Engagement: A Comprehensive Survey and Holistic Architecture for Computer Vision-Based Audience Analysis in Live Events
by Marco Lemos, Pedro J. S. Cardoso and João M. F. Rodrigues
Multimodal Technol. Interact. 2026, 10(1), 8; https://doi.org/10.3390/mti10010008 - 8 Jan 2026
Cited by 2 | Viewed by 2199
Abstract
The accurate measurement of audience engagement in real-world live events remains a significant challenge, with the majority of existing research confined to controlled environments like classrooms. This paper presents a comprehensive survey of Computer Vision AI-driven methods for real-time audience engagement monitoring and [...] Read more.
The accurate measurement of audience engagement in real-world live events remains a significant challenge, with the majority of existing research confined to controlled environments like classrooms. This paper presents a comprehensive survey of Computer Vision AI-driven methods for real-time audience engagement monitoring and proposes a novel, holistic architecture to address this gap, with this architecture being the main contribution of the paper. The paper identifies and defines five core constructs essential for a robust analysis: Attention, Emotion and Sentiment, Body Language, Scene Dynamics, and Behaviours. Through a selective review of state-of-the-art techniques for each construct, the necessity of a multimodal approach that surpasses the limitations of isolated indicators is highlighted. The work synthesises a fragmented field into a unified taxonomy and introduces a modular architecture that integrates these constructs with practical, business-oriented metrics such as Commitment, Conversion, and Retention. Finally, by integrating cognitive, affective, and behavioural signals, this work provides a roadmap for developing operational systems that can transform live event experience and management through data-driven, real-time analytics. Full article
Show Figures

Figure 1

18 pages, 4285 KB  
Article
Eye-Tracking and Emotion-Based Evaluation of Wardrobe Front Colors and Textures in Bedroom Interiors
by Yushu Chen, Wangyu Xu and Xinyu Ma
Multimodal Technol. Interact. 2026, 10(1), 7; https://doi.org/10.3390/mti10010007 - 6 Jan 2026
Cited by 1 | Viewed by 949
Abstract
Wardrobe fronts form a major visual element in bedroom interiors, yet material selection for their colors and textures often relies on intuition rather than evidence. This study develops a data-driven framework that links gaze behavior and affective responses to occupants’ preferences for wardrobe [...] Read more.
Wardrobe fronts form a major visual element in bedroom interiors, yet material selection for their colors and textures often relies on intuition rather than evidence. This study develops a data-driven framework that links gaze behavior and affective responses to occupants’ preferences for wardrobe front materials. Forty adults evaluated color and texture swatches and rendered bedroom scenes while eye-tracking data capturing attraction, retention, and exploration were collected. Pairwise choices were modeled using a Bradley–Terry approach, and visual-attention features were integrated with emotion ratings to construct an interpretable attention index for predicting preferences. Results show that neutral light colors and structured wood-like textures consistently rank highest, with scene context reducing preference differences but not altering the order. Shorter time to first fixation and longer fixation duration were the strongest predictors of desirability, demonstrating the combined influence of rapid visual capture and sustained attention. Within the tested stimulus set and viewing conditions, the proposed pipeline yields consistent preference rankings and an interpretable attention-based score that supports evidence-informed shortlisting of wardrobe-front materials. The reported relationships between gaze, affect, and choice are associative and are intended to guide design decisions within the scope of the present experimental settings. Full article
Show Figures

Figure 1

17 pages, 1388 KB  
Article
HISF: Hierarchical Interactive Semantic Fusion for Multimodal Prompt Learning
by Haohan Feng and Chen Li
Multimodal Technol. Interact. 2026, 10(1), 6; https://doi.org/10.3390/mti10010006 - 6 Jan 2026
Viewed by 930
Abstract
Recent vision-language pre-training models, like CLIP, have been shown to generalize well across a variety of multitask modalities. Nonetheless, their generalization for downstream tasks is limited. As a lightweight adaptation approach, prompt learning could allow task transfer by optimizing only several learnable vectors [...] Read more.
Recent vision-language pre-training models, like CLIP, have been shown to generalize well across a variety of multitask modalities. Nonetheless, their generalization for downstream tasks is limited. As a lightweight adaptation approach, prompt learning could allow task transfer by optimizing only several learnable vectors and thus is more flexible for pre-trained models. However, current methods mainly concentrate on the design of unimodal prompts and ignore effective means for multimodal semantic fusion and label alignment, which limits their representation power. To tackle these problems, this paper designs a Hierarchical Interactive Semantic Fusion (HISF) framework for multimodal prompt learning. On top of frozen CLIP backbones, HISF injects visual and textual signals simultaneously in intermediate layers of a Transformer through a cross-attention mechanism as well as fitting category embeddings. This architecture realizes the hierarchical semantic fusion at the modality level with structural consistency kept at each layer. In addition, a Label Embedding Constraint and a Semantic Alignment Loss are proposed to promote category consistency while alleviating semantic drift in training. Extensive experiments across 11 few-shot image classification benchmarks show that HISF improves the average accuracy by around 0.7% compared to state-of-the-art methods and has remarkable robustness in cross-domain transfer tasks. Ablation studies also verify the effectiveness of each proposed part and their combination: hierarchical structure, cross-modal attention, and semantic alignment collaborate to enrich representational capacity. In conclusion, the proposed HISF is a new hierarchical view for multimodal prompt learning and provides a more lightweight and generalizable paradigm for adapting vision-language pre-trained models. Full article
Show Figures

Figure 1

61 pages, 4117 KB  
Systematic Review
Neuroplasticity-Informed Learning Under Cognitive Load: A Systematic Review of Functional Imaging, Brain Stimulation, and Educational Technology Applications
by Evgenia Gkintoni, Andrew Sortwell, Stephanos P. Vassilopoulos and Georgios Nikolaou
Multimodal Technol. Interact. 2026, 10(1), 5; https://doi.org/10.3390/mti10010005 - 31 Dec 2025
Cited by 11 | Viewed by 14089
Abstract
Background/Objectives: This systematic review examines neuroplasticity-informed approaches to learning under cognitive load, synthesizing evidence from functional imaging, brain stimulation, and educational technology research. As digital learning environments increasingly challenge learners with complex cognitive demands, understanding how neuroplasticity principles can inform adaptive educational design [...] Read more.
Background/Objectives: This systematic review examines neuroplasticity-informed approaches to learning under cognitive load, synthesizing evidence from functional imaging, brain stimulation, and educational technology research. As digital learning environments increasingly challenge learners with complex cognitive demands, understanding how neuroplasticity principles can inform adaptive educational design becomes critical. This review examines how neural mechanisms underlying learning under cognitive load can inform the development of evidence-based educational technologies that optimize neuroplastic potential while mitigating cognitive overload. Methods: Following PRISMA guidelines, we synthesized 94 empirical studies published between 2005 and 2025 across PubMed, Scopus, Web of Science, and PsycINFO. Studies were selected based on rigorous inclusion criteria that emphasized functional neuroimaging (fMRI, EEG), non-invasive brain stimulation (tDCS, TMS), and educational technology applications, which examined learning outcomes under varying cognitive load conditions. Priority was given to research with translational implications for adaptive learning systems and personalized educational interventions. Results: Functional imaging studies reveal an inverted-U relationship between cognitive load and neuroplasticity, with a moderate challenge in optimizing prefrontal-parietal network activation and learning-related neural adaptations. Brain stimulation research demonstrates that tDCS and TMS can enhance neuroplastic responses under cognitive load, particularly benefiting learners with lower baseline abilities. Educational technology applications demonstrate that neuroplasticity-informed adaptive systems, which incorporate real-time cognitive load monitoring and dynamic difficulty adjustment, significantly enhance learning outcomes compared to traditional approaches. Individual differences in cognitive capacity, neurodiversity, and baseline brain states substantially moderate these effects, necessitating the development of personalized intervention strategies. Conclusions: Neuroplasticity-informed learning approaches offer a robust framework for educational technology design that respects cognitive load limitations while maximizing adaptive neural changes. Integration of functional imaging insights, brain stimulation protocols, and adaptive algorithms enables the development of inclusive educational technologies that support diverse learners under cognitive stress. Future research should focus on scalable implementations of real-time neuroplasticity monitoring in authentic educational settings, as well as on developing ethical frameworks for deploying neurotechnology-enhanced learning systems across diverse populations. Full article
Show Figures

Figure 1

17 pages, 1120 KB  
Article
Neuroception of Psychological Safety and Attitude Towards General AI in uHealth Context
by Anca-Livia Panfil, Simona C. Tamasan, Claudia C. Vasilian, Raluca Horhat and Diana Lungeanu
Multimodal Technol. Interact. 2026, 10(1), 4; https://doi.org/10.3390/mti10010004 - 30 Dec 2025
Viewed by 2020
Abstract
Interest in general AI is widespread, and much is expected from its large-scale adoption in the healthcare sector. However, the success of uHealth implementations relies on genuine trust, beyond technical performance. Neuroception of psychological safety (NPS), grounded in polyvagal theory, encompasses the human [...] Read more.
Interest in general AI is widespread, and much is expected from its large-scale adoption in the healthcare sector. However, the success of uHealth implementations relies on genuine trust, beyond technical performance. Neuroception of psychological safety (NPS), grounded in polyvagal theory, encompasses the human subconscious and automatic processes of safety and risk detection. We conducted a cross-sectional survey to explore a hypothetical connection between NPS and the perception of general AI in the uHealth context, by an anonymous online questionnaire comprising the following: Neuroception of Psychological Safety Scale (NPSS), four-item AI Attitude Scale (AIAS-4), and questions on AI threat, age, gender, and level of education. Multivariate analysis was performed using covariance-based structural equation modeling. We received 201 responses: 73 (36.3%) males vs. 128 (63.7%) females, all adults with varying levels of education (from 0 = basic formal education to 4 = master’s degree). Respondents belonged to four demographic cohorts: from Baby boomers to Generation Z. SEM results indicated that attitudes towards AI-driven health interventions are significantly impacted by social engagement and compassion (NPSS factors). Gender, education, and demographic cohort were confirmed as significant covariates. NPS-related attitudes towards AI should be considered and analyzed by healthcare providers, application developers, and policy or regulatory authorities. Full article
Show Figures

Figure 1

20 pages, 4100 KB  
Article
Mixed Reality Game Design for the Effectiveness and Application Research of Integrating Sustainable Concepts into Blended Learning
by Zhengqing Wang, Chenxi Xiao and Pengwei Hsiao
Multimodal Technol. Interact. 2026, 10(1), 3; https://doi.org/10.3390/mti10010003 - 30 Dec 2025
Viewed by 1556
Abstract
This study explores how mixed reality (MR) game environments, enabled by sensor-based motion tracking and interactive visualization technologies, can be effectively integrated into blended learning to promote sustainability education. Using eight Macau bakeries as empirical cases, field investigations collected and categorized surplus bread [...] Read more.
This study explores how mixed reality (MR) game environments, enabled by sensor-based motion tracking and interactive visualization technologies, can be effectively integrated into blended learning to promote sustainability education. Using eight Macau bakeries as empirical cases, field investigations collected and categorized surplus bread samples, while carbon emission frameworks informed pedagogical design. Employing a multidimensional research methodology combining questionnaires and semi-structured interviews, the study delved into the intrinsic link between bread waste and carbon emissions. Through perceptual interaction design and task-oriented challenge modes within the MR environment, users were immersed in experiencing the pathway of sustainable behavioral impact. Post-instructional engagement with the MR game revealed that >90% of participants expressed strong affinity for the system design, and >85% perceived it as intuitively operable. Analysis of user feedback and performance data demonstrates the system’s potential to deliver solutions for reducing bread waste and carbon emissions. By establishing a replicable MR game framework and technical mechanisms, this research offers novel perspectives for future sustainability education studies in the field of behavioral mixed reality design. Full article
Show Figures

Graphical abstract

11 pages, 555 KB  
Article
Human–AI Feedback Loop for Pronunciation Training: A Mobile Application with Phoneme-Level Error Highlighting
by Aleksei Demin, Georgii Vorontsov and Dmitrii Chaikovskii
Multimodal Technol. Interact. 2026, 10(1), 2; https://doi.org/10.3390/mti10010002 - 26 Dec 2025
Cited by 4 | Viewed by 3024
Abstract
This paper presents an AI-augmented pronunciation training approach for Russian language learners through a mobile application that supports an interactive learner–system feedback loop. The system combines a pre-trained Wav2Vec2Phoneme neural network with Needleman–Wunsch global sequence alignment to convert reference and learner speech into [...] Read more.
This paper presents an AI-augmented pronunciation training approach for Russian language learners through a mobile application that supports an interactive learner–system feedback loop. The system combines a pre-trained Wav2Vec2Phoneme neural network with Needleman–Wunsch global sequence alignment to convert reference and learner speech into aligned phoneme sequences. Rather than producing an overall pronunciation score, the application provides localized, interpretable feedback by highlighting phoneme-level matches and mismatches in a red/green transcription, enabling learners to see where sounds were substituted, omitted, or added. Implemented as a WeChat Mini Program with a WebSocket-based backend, the design illustrates how speech-to-phoneme models and alignment procedures can be integrated into a lightweight mobile interface for autonomous pronunciation practice. We further provide a feature-level comparison with widely used commercial applications (Duolingo, HelloChinese, Babbel), emphasizing differences in feedback granularity and interpretability rather than unvalidated accuracy claims. Overall, the work demonstrates the feasibility of alignment-based phoneme-level feedback for mobile pronunciation training and motivates future evaluation of recognition reliability, latency, and learning outcomes on representative learner data. Full article
Show Figures

Figure 1

22 pages, 1413 KB  
Systematic Review
Motion Capture as an Immersive Learning Technology: A Systematic Review of Its Applications in Computer Animation Training
by Xinyi Jiang, Zainuddin Ibrahim, Jing Jiang and Gang Liu
Multimodal Technol. Interact. 2026, 10(1), 1; https://doi.org/10.3390/mti10010001 - 23 Dec 2025
Cited by 5 | Viewed by 3123
Abstract
Motion capture (MoCap) is increasingly recognized as a powerful multimodal immersive learning technology, providing embodied interaction and real-time motion visualization that enrich educational experiences. Although MoCap is gaining prominence within educational research, its pedagogical value and integration into computer animation training environments have [...] Read more.
Motion capture (MoCap) is increasingly recognized as a powerful multimodal immersive learning technology, providing embodied interaction and real-time motion visualization that enrich educational experiences. Although MoCap is gaining prominence within educational research, its pedagogical value and integration into computer animation training environments have received relatively limited systematic investigation. This review synthesizes findings from 17 studies to analyze how MoCap supports instructional design, creative development, and workflow efficiency in animation education. Results show that MoCap enables a multimodal learning process by combining visual, kinesthetic, and performative modalities, strengthening learners’ sense of presence, agency, and perceptual–motor understanding. Furthermore, we identified five key technical affordances of MoCap, including precision and fidelity, multi-actor and creative control, interactivity and immersion, perceptual–motor learning, and emotional expressiveness, which together shape both cognitive and creative learning outcomes. Emerging trends highlight MoCap’s growing convergence with VR/AR, XR, real-time rendering engines, and AI-augmented motion analysis, expanding its role in the design of immersive and interactive educational systems. This review offers insights into the use of MoCap in animation education research and provides a springboard for future work on more immersive and industry-relevant training. Full article
(This article belongs to the Special Issue Educational Virtual/Augmented Reality)
Show Figures

Figure 1

Previous Issue
Next Issue
Back to TopTop