Next Issue
Volume 10, October
Previous Issue
Volume 10, August
 
 

Multimodal Technol. Interact., Volume 10, Issue 9 (September 2026) – 12 articles

Cover Story (view full-size image): Game-based learning environments offer engaging and interactive ways to support learning. However, accommodating players with different skill levels remains a challenge. Static difficulty settings do not sufficiently address individual differences in performance and progress. Dynamic difficulty adjustment (DDA) adapts challenges to the player to support a balanced learning experience. The study explores DDA in a cooperative multiplayer environment that teaches programming through block-based puzzles. A rating-based system adapts task difficulty to player performance and was compared with a static version using a fixed medium difficulty. The largest observed differences concerned perceived competence and gameplay progression. The findings highlight the potential of adaptive game mechanics for supporting diverse skill levels. View this paper
  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Select all
Export citation of selected articles as:
31 pages, 12537 KB  
Article
Designing and Evaluating: A Multimodal AI Story Co-Creation System for Supporting Children’s Post-Conflict Reflection
by Zihui Jiang, Yi Li, Yanfei Xu, Yu Gao and Xueyi Li
Multimodal Technol. Interact. 2026, 10(9), 98; https://doi.org/10.3390/mti10090098 - 21 Sep 2026
Viewed by 272
Abstract
Peer conflict is a common and developmentally meaningful social experience in children’s school life. However, existing interactive technologies have mainly focused on immediate mediation or conflict skills training, whereas established approaches such as attributional intervention, restorative practices, reflective learning, perspective taking, conflict debriefing, [...] Read more.
Peer conflict is a common and developmentally meaningful social experience in children’s school life. However, existing interactive technologies have mainly focused on immediate mediation or conflict skills training, whereas established approaches such as attributional intervention, restorative practices, reflective learning, perspective taking, conflict debriefing, and social information processing have rarely been integrated within a single AI-supported system for children’s post-conflict reflection. Based on social information processing theory, this paper presents Kindom, a multimodal AI-supported story co-creation system for children aged 6–8 years. Through neutral prompts, role-based perspective switching, revision-oriented narrative branches, and Qwen-generated co-created narrative endings and researcher-mediated visual summaries, Kindom supports children in event review, emotion recognition, intention inference, and strategy generation and evaluation. This study adopted a two-stage Research through Design approach. First, formative interviews were conducted with 10 primary school teachers and 4 educational experts to identify design needs and develop system design goals. A one-month mixed-methods quasi-experimental study was then conducted to compare the performance of 40 children under the Kindom condition and a matched picture-book condition. The results showed that children in the Kindom group had significantly lower hostile attribution than those in the control group (p < 0.001, d = 1.175), and significantly higher positive response generation (p < 0.001, d = 2.109) and positive response evaluation (p < 0.001, d = 1.577). Qualitative analysis further showed that children shifted from hostile interpretations to contextualized interpretations, from a single-character perspective to an understanding of both parties’ emotions, and from direct responses to constructive strategies such as negotiation, repair, and help-seeking. These scenario-based task results suggest that multimodal AI-supported story co-creation may provide structured, low-pressure, process-oriented scaffolding for children’s social information processing and post-conflict reflection. This study proposes a four-stage support model for children’s post-peer-conflict reflection, designs and implements a child-centered multimodal AI-supported story co-creation system, and provides preliminary mixed-methods evidence based on scenario tasks. Its main contribution lies in integrating AI-supported multimodal story co-creation, sequential SIP-based reflection, structured revision of children’s initial interpretations, and traceable interaction logs. Full article
►▼ Show Figures

Figure 1

18 pages, 11693 KB  
Article
Effects of Robot Speed, Number, and Orientation on Operator Stress and Mental Workload in Human–Robot Collaboration: A Multimodal Assessment Using Physiological and Subjective Measures
by Qian Zhang, Nisa Mareldiya Soltani, Jana Jovcheva, Jenna Snead and Mia Yaqin Wang
Multimodal Technol. Interact. 2026, 10(9), 97; https://doi.org/10.3390/mti10090097 - 20 Sep 2026
Viewed by 215
Abstract
This study examined the effects of robot speed, number, and spatial orientation on operator stress, mental workload, attention, and excitement during human–robot collaboration. As collaborative robots (cobots) become increasingly prevalent in modern manufacturing, operator safety concerns persist. Despite Industry 5.0’s focus on human-centered [...] Read more.
This study examined the effects of robot speed, number, and spatial orientation on operator stress, mental workload, attention, and excitement during human–robot collaboration. As collaborative robots (cobots) become increasingly prevalent in modern manufacturing, operator safety concerns persist. Despite Industry 5.0’s focus on human-centered production, the impact of robot configuration on operators’ cognitive and emotional states remains underexplored. Thirty participants completed eight sessions of a Lego block stacking task using a mixed-factor (3 × 2 × 2) within-subjects design that varied robot speed (slow, fast, or mixed), robot number (one or two), and left–right spatial orientation. Physiological data were collected continuously via an Empatica E4 wristband and analyzed using linear mixed models; subjective data were collected via a post-session questionnaire and analyzed using non-parametric tests. Slower robot speeds reduced stress, workload, attention, and excitement relative to higher speeds. Two robots increased stress, mental workload, attention, and excitement relative to one robot, with consistent elevations in electrodermal activity, shorter inter-beat intervals, and lower skin temperature confirming sympathetic arousal. The fast two-robot condition produced the highest subjective demand across all measures. Notably, one robot at fast speed and two robots at slow speed produced statistically comparable subjective workload, suggesting both configurations offer productive alternatives with manageable operator strain. Robot spatial orientation did not significantly affect any outcome measure. These findings suggest that, under the conditions of this laboratory paradigm, a configuration of one robot at fast speed may offer a favorable balance of workload and engagement for safe and effective collaboration, with two robots at slow speed as a viable secondary alternative; direct objective productivity data were not collected, and these recommendations should be validated against task throughput in future work. Design implications for collaborative workstations are discussed. Full article
►▼ Show Figures

Figure 1

37 pages, 5630 KB  
Article
Vellum of Lies: Designing and Evaluating a Multimodal Tangible Interaction System for Embodied Serious Historical Narratives
by Yuli Hou, Shuting Wang, Zhoutong Su, Leying Bi, Xiaopei Ye and Yifan Zhang
Multimodal Technol. Interact. 2026, 10(9), 96; https://doi.org/10.3390/mti10090096 - 20 Sep 2026
Viewed by 242
Abstract
Many existing historical narrative experiences still rely on screen-based viewing and limited multimodal interaction, which tends to keep users in a passive receiving role and limits their cognitive understanding, reflective thinking, user engagement, and emotional response. To address this, we designed Vellum of [...] Read more.
Many existing historical narrative experiences still rely on screen-based viewing and limited multimodal interaction, which tends to keep users in a passive receiving role and limits their cognitive understanding, reflective thinking, user engagement, and emotional response. To address this, we designed Vellum of Lies, an immersive tangible user interface for anti-feudal historical narratives. The work integrates physical props, sensors, motion-based input, and projection mapping, organizing the story into continuous interaction nodes that allow users to participate in narrative progression through embodied interaction, including spatial movement, object manipulation, and final choice-making. The final choice mechanism operationalizes Tragic Agency by allowing users to make meaningful choices while constraining their ability to alter the final tragic outcome, with the aim of evoking tension between perceived choice and narrative inevitability. A mixed-methods user experience evaluation (N = 32) compared Vellum of Lies with a video-watching narrative experience. In this small-scale exploratory study, the experimental condition showed significant advantages in the HUQ Overall score, Reflection, the UES-SF Overall score, Aesthetic Appeal, Interest/Enjoyment, Pleasure/Valence, and Arousal, while also showing lower NASA-TLX Overall scores, Mental Demand, and Effort. Other measured subdimensions did not reach statistical significance. These preliminary findings suggest that multimodal tangible interaction may function not only as a presentation technique, but also as a narrative mechanism that coordinates physical input, bodily movement, projected feedback, and constrained choice-making to support selected aspects of understanding, engagement, and emotional reflection in serious historical narratives. Full article
►▼ Show Figures

Figure 1

33 pages, 3038 KB  
Review
Evidence-Grounded, Role-Adaptive Conversational AI for Occupational Therapists and Caregivers in Neurodevelopmental Care: A Comprehensive Narrative Review and Sociotechnical Reference Architecture
by Pantelis Pergantis, Konstantinos Georgiou, Nikolaos Bardis, Charalabos Skianis and Athanasios Drigas
Multimodal Technol. Interact. 2026, 10(9), 95; https://doi.org/10.3390/mti10090095 - 19 Sep 2026
Viewed by 336
Abstract
Generative conversational AI can make health information easier to access, but its responses may still be unsupported, outdated, or poorly matched to the user. These risks are important in neurodevelopmental care, where occupational therapists and parents or caregivers require different levels of detail, [...] Read more.
Generative conversational AI can make health information easier to access, but its responses may still be unsupported, outdated, or poorly matched to the user. These risks are important in neurodevelopmental care, where occupational therapists and parents or caregivers require different levels of detail, language, and guidance. This narrative review used a structured multidisciplinary search to examine knowledge-base governance, retrieval-augmented generation, source verification, role adaptation, uncertainty communication, safety routing, privacy, human oversight, and lifecycle control. The evidence was synthesized into a sociotechnical reference architecture for adult-facing conversational AI serving occupational therapists and parents or caregivers. The architecture comprises five connected layers: intended use and request classification; knowledge governance and retrieval; claim-level verification; role-specific response design and safety review; and response delivery and lifecycle control. Occupational relevance, documentation, accountability, and continuous evaluation operate across the layers. Each layer is linked to operational requirements, responsible actors, and evaluation indicators. The framework is intended to guide future development and evaluation rather than function as a validated clinical tool. It excludes diagnosis, autonomous treatment planning, direct child-facing interaction, and replacement of occupational therapy assessment. The review provides a practical basis for developing conversational AI that is evidence-grounded, traceable, role-appropriate, and compatible with professional oversight. Full article
►▼ Show Figures

Figure 1

23 pages, 8876 KB  
Article
Designing an Equity-Oriented Dietary Self-Management System for Long-Term Care Residents with Type 2 Diabetes: Participatory Co-Design and Usability Testing
by Xueyi Li, Yonghong Liu and Yangcheng Wang
Multimodal Technol. Interact. 2026, 10(9), 94; https://doi.org/10.3390/mti10090094 - 17 Sep 2026
Viewed by 177
Abstract
Older adults with type 2 diabetes (T2D) in long-term care face dietary self-management challenges arising from standardized routines, sensory and cognitive variability, and reliance on caregivers. This study used participatory co-design to develop and preliminarily evaluate an equity-oriented digital dietary management prototype. From [...] Read more.
Older adults with type 2 diabetes (T2D) in long-term care face dietary self-management challenges arising from standardized routines, sensory and cognitive variability, and reliance on caregivers. This study used participatory co-design to develop and preliminarily evaluate an equity-oriented digital dietary management prototype. From June 2024 to December 2025, a four-phase study was conducted at a long-term care facility in Beijing, China. Semi-structured interviews with 20 stakeholders (12 residents, 5 family caregivers, and 3 nurses) and a one-day shadowing observation of one of these residents informed a 90 min co-creation workshop with 20 stakeholders (12 residents, 4 family caregivers, and 4 nurses) and iterative prototyping. All 12 residents in the workshop had participated in the interviews, including the resident involved in the shadowing observation. The resulting high-fidelity prototype integrated a consolidated reminder dashboard, an adaptive portion-control slider, and glucose–meal feedback visualization. Task-based usability testing with an independent sample of 12 participants (6 residents, 3 family caregivers, and 3 nurses) who had not participated in the preceding phases yielded a mean task-success rate of 94.44% (SD 0.43), a mean task-completion time of 1.9 min (SD 0.40), a mean error count of 0.33 per session (SD 0.47), a mean System Usability Scale score of 81.0 (SD 5.23), and satisfaction of 4.5/5 (SD 0.50). Qualitative feedback indicated that participants perceived the prototype as easy to navigate and potentially supportive of resident autonomy and coordinated record-keeping. These findings provide preliminary support for its usability and acceptability in the evaluated tasks. Field studies are needed to assess implementation, workflow effects, and behavioral or clinical outcomes. Full article
►▼ Show Figures

Figure 1

21 pages, 5629 KB  
Article
Nephron Flow Lab: Vibe Coding a Visual Nephrology Learning App
by Isaac Calderwood and Tyler Bland
Multimodal Technol. Interact. 2026, 10(9), 93; https://doi.org/10.3390/mti10090093 - 12 Sep 2026
Viewed by 306
Abstract
Nephrology requires students to integrate nephron transport, hormonal regulation, medication effects, and serum and urine electrolyte patterns, which can be difficult to represent with static instructional materials. This study describes the development and early evaluation of Nephron Flow Lab, a web-based interactive nephrology [...] Read more.
Nephrology requires students to integrate nephron transport, hormonal regulation, medication effects, and serum and urine electrolyte patterns, which can be difficult to represent with static instructional materials. This study describes the development and early evaluation of Nephron Flow Lab, a web-based interactive nephrology learning application created through a generative artificial intelligence-assisted “vibe coding” workflow. Initial visual prototyping was performed using Gemini 3.1 Pro, with Codex using GPT 5.5 used for complex physiologic logic, animations, and implementation. Idaho WWAMI students received optional access to the application before Exam 5 and the course final examination, whereas students at other WWAMI sites served as controls. Aggregate item-level cross-sectional difference-in-differences contrasts compared focal-versus-baseline performance at the Experimental site with corresponding contrasts at individual and pooled control sites. Pooled-control estimates showed no measurable advantage for Idaho on Exam 5 (−0.88 percentage points; 95% CI, −5.96 to 4.87), the final examination (0.66 percentage points; 95% CI, −4.57 to 5.48), or both assessments combined (0.47 percentage points; 95% CI, −3.36 to 4.52). Survey respondents rated the application favorably using a modified uMARS survey, with an overall app-quality score of 4.54/5. Nephron Flow Lab demonstrates the feasibility of faculty-built AI-assisted educational software and was favorably perceived by survey respondents; however, access to the application was not associated with a measurable aggregate achievement advantage in this exploratory evaluation. Full article
►▼ Show Figures

Figure 1

20 pages, 614 KB  
Article
Synchronized Listening-While-Reading in EFL Vocabulary Learning: Immediate Gains and Exploratory Heart-Rate Patterns
by Heeseong Ahn, Nahkyoung Han, Jong-su Park, Young Seok Oh, Hubert H. Pak and Chungwan Lim
Multimodal Technol. Interact. 2026, 10(9), 92; https://doi.org/10.3390/mti10090092 - 7 Sep 2026
Viewed by 298
Abstract
Multimodal input is often assumed to facilitate second-language learning by reducing cognitive demands, but this account may not fully explain digital listening-while-reading. This study examined vocabulary learning and task-related physiological activation in 40 Korean university EFL learners randomly assigned to text-only or synchronized [...] Read more.
Multimodal input is often assumed to facilitate second-language learning by reducing cognitive demands, but this account may not fully explain digital listening-while-reading. This study examined vocabulary learning and task-related physiological activation in 40 Korean university EFL learners randomly assigned to text-only or synchronized text-with-audio conditions. Vocabulary knowledge was assessed using a 20-item test before and immediately after the task, while heart rate was recorded during task performance and recovery. The text-with-audio group showed substantial vocabulary improvement (M = 67.00 to 80.25, p < 0.001), whereas the text-only group showed no significant improvement (p = 0.096). More importantly, vocabulary gain differed significantly between conditions, with greater immediate gains in the text-with-audio condition. Mean task heart rate (d = 0.59, p = 0.069) and elevation from baseline (d = 0.63, p = 0.053) were numerically higher in the text-with-audio condition, but neither difference reached statistical significance. Thus, the physiological findings are suggestive rather than confirmatory and should not be interpreted as direct evidence of engagement or cognitive load. The findings support a provisional balanced-efficiency framework in which improved multimodal learning may occur without reduced physiological activation, requiring validation with direct cognitive, behavioral, and complementary physiological measures. Full article
►▼ Show Figures

Figure 1

22 pages, 5282 KB  
Article
Beyond the Archive: Designing a Digital Showroom for Exploring the Living Heritage of Nakhon Si Thammarat Brocade
by Supaporn Chai-Arayalert, Ketsaraporn Suttapong, Supattra Puttinaovarat, Nattaporn Thongsri and Jariya Seksan
Multimodal Technol. Interact. 2026, 10(9), 91; https://doi.org/10.3390/mti10090091 - 30 Aug 2026
Viewed by 519
Abstract
Handwoven textiles constitute a form of intangible craft heritage whose cultural value resides not only in finished artifacts but also in the embodied skills, motif knowledge, practitioner narratives, and community-based meanings through which they are produced and transmitted. However, many digital heritage platforms [...] Read more.
Handwoven textiles constitute a form of intangible craft heritage whose cultural value resides not only in finished artifacts but also in the embodied skills, motif knowledge, practitioner narratives, and community-based meanings through which they are produced and transmitted. However, many digital heritage platforms remain oriented toward visual presentation, product cataloguing, and archival access, offering limited support for communicating the interpretive, practice-based, and relational dimensions of craft knowledge. Addressing this gap, this study designed, developed, and evaluated a four-room digital showroom for Nakhon Si Thammarat Brocade, a traditional brocade textile of southern Thailand. Guided by Design Science Research methodology and heritage interpretation theory, the digital showroom was conceptualized as a cultural communication interface rather than as an online product display. Knowledge elicitation with master weavers and a user-needs survey informed the development of four connected spaces: the Historical Room, Aesthetic Room, Livelihood and Craft Value Room, and Learning-by-Doing Room. Together, these rooms organize user engagement through a progressive interpretive pathway, moving from cultural orientation and motif understanding to artisan-value recognition and guided creative participation. An end-user evaluation with 100 participants showed positive perceptions across functionality and usability (M = 4.13, SD = 0.72), content structure and organization (M = 4.06, SD = 0.79), cultural content quality (M = 4.34, SD = 0.73), and interactive learning-by-doing experience (M = 4.17, SD = 0.81). Users particularly valued the culturally grounded content, practitioner-centered narratives, and opportunities to apply cultural knowledge through virtual design activities. The findings suggest that a practitioner-informed and theoretically grounded digital showroom can extend digital heritage practice beyond passive product representation by supporting contextualized, interpretive, and participatory engagement with intangible craft heritage. Full article
►▼ Show Figures

Figure 1

24 pages, 38505 KB  
Article
Dynamic Difficulty Adjustment in a Multiplayer Learning Environment: An Exploratory Evaluation of Player Experience
by Michael Holly, Alexander Kassil and Johanna Pirker
Multimodal Technol. Interact. 2026, 10(9), 90; https://doi.org/10.3390/mti10090090 - 29 Aug 2026
Viewed by 483
Abstract
Balancing the learning experience for players with diverse skill levels, particularly in multi-user learning environments, remains a challenge. Many game-based learning systems rely on static difficulty settings that do not adapt to individual abilities, leading to frustration or disengagement. Dynamic difficulty adjustment (DDA) [...] Read more.
Balancing the learning experience for players with diverse skill levels, particularly in multi-user learning environments, remains a challenge. Many game-based learning systems rely on static difficulty settings that do not adapt to individual abilities, leading to frustration or disengagement. Dynamic difficulty adjustment (DDA) aims to create a more personalized experience by dynamically adjusting the task’s difficulty to match the player’s evolving skill level, while also accommodating players with lower performance capabilities. This paper explores the potential of a DDA system in a multiplayer learning environment that includes block-based puzzles to teach programming concepts. We used a rating algorithm to evaluate the player’s performance and dynamically adjust in-game objectives. To explore the player experience, we conducted an exploratory quasi-experimental between-group study comparing the adaptive version with a fixed-difficulty implementation. The largest observed between-group difference concerned perceived competence, with higher scores in the DDA group (DDA: AVG = 2.31, SD = 0.76; Non-DDA: AVG = 1.62, SD = 1.07). However, this difference did not remain statistically significant after adjustment. No significant differences were found in perceived levels of challenge, workload, flow, or team involvement. Full article
►▼ Show Figures

Figure 1

26 pages, 3699 KB  
Article
Immersive Virtual Reality for Occupational Track Safety Training: A Quasi-Experimental Field Study Comparing Six Training Formats
by Oliver Christ, Aaron Ettlin, Pascal Valentin Meier, Pascal Duchêne, Paul Hügli, Stefan Adam and Stephan Gut
Multimodal Technol. Interact. 2026, 10(9), 89; https://doi.org/10.3390/mti10090089 - 28 Aug 2026
Viewed by 490
Abstract
Conventional occupational safety training is often criticized for weak transfer of declarative knowledge into workplace behavior. Immersive virtual reality (iVR) has been proposed as a promising complement, yet comparisons of different iVR formats remain scarce in real occupational settings. This quasi-experimental field study [...] Read more.
Conventional occupational safety training is often criticized for weak transfer of declarative knowledge into workplace behavior. Immersive virtual reality (iVR) has been proposed as a promising complement, yet comparisons of different iVR formats remain scarce in real occupational settings. This quasi-experimental field study (N = 127) compared six training formats—training-as-usual (control), desktop single-player, iVR single-player, iVR multi-player, iVR multi-player with retrospective perspective change (RPC), and iVR classroom—each embedded as an add-on to the Swiss Federal Railways (SBB) track safety course (SstA). The primary outcome was achieving the full score on a safety-critical checklist exam item; secondary outcomes were technology acceptance (TAM: perceived usefulness, ease of use, behavioral intention) and five practice-oriented items. Binomial logistic regression showed a significant overall model (χ2(5) = 15.80, p = 0.007): participants in all five digital conditions had significantly higher odds of achieving the full score than participants in training-as-usual (OR = 5.13–14.62). Kruskal–Wallis tests revealed significant group differences on all three TAM scales and three practice-oriented items, with the most favorable descriptive pattern observed for iVR classroom. The observed pattern suggests that instructional configuration may play an important role alongside technological immersion in occupational iVR training. However, given the quasi-experimental design and the presence of residual confounding, these findings should be interpreted as associative rather than causal. Full article
(This article belongs to the Special Issue Educational Virtual/Augmented Reality)
►▼ Show Figures

Figure 1

26 pages, 3414 KB  
Article
Semantic-Enhanced Underwater Videos Multi-Label Classification Network Based on Structural Graph Convolution
by Yun Li, Jun Yang, Hui Guo, Junfeng Wei, Kunsheng Wu and Peiguang Jing
Multimodal Technol. Interact. 2026, 10(9), 88; https://doi.org/10.3390/mti10090088 - 27 Aug 2026
Viewed by 297
Abstract
Underwater visual degradation makes it difficult for the image modality to represent video semantics, while label sparsity in underwater scenes leads to weak inter-category correlations, thereby degrading the performance of multi-label classification. To address these issues, this paper proposes a Semantic-Enhanced Underwater Videos [...] Read more.
Underwater visual degradation makes it difficult for the image modality to represent video semantics, while label sparsity in underwater scenes leads to weak inter-category correlations, thereby degrading the performance of multi-label classification. To address these issues, this paper proposes a Semantic-Enhanced Underwater Videos Multi-Label Classification Network Based on Structural Graph Convolution (SEMGCN). Specifically, the proposed method first disentangles the image and text modalities into shared and private representations, and enhances feature representation capability through orthogonal constraints and feature reconstruction. Moreover, a Cross-Modal Category-Aware Module (CCAM) is constructed to model interactions between image and text features and perform bidirectional cross-attention with category-label text embeddings, thereby generating category-aware initial node representations. Furthermore, to alleviate the limitation of semantic propagation caused by sparse label co-occurrence, a Structural Graph Convolutional Network (SGCN) is proposed. By integrating explicit co-occurrence relationships with implicit structural similarity relationships, the proposed model collaboratively captures both explicit and latent semantic associations, thereby improving multi-label classification performance under label-sparse conditions. Experiments were conducted on the self-constructed Underwater Video Multi-label Classification Dataset (UVMC) and the public MLSV2018 dataset. The experimental results show that SEMGCN achieves Average Precision scores of 0.8645 and 0.8388 on UVMC and MLSV2018, respectively, demonstrating the effectiveness of the proposed method. Full article
►▼ Show Figures

Figure 1

23 pages, 523 KB  
Article
A Persistent Multi-User Virtual Reality Garden: Architecture, Traceability, and Technical Validation
by Giovanni Giuliodori, Erica Santaguida, Chiara Evangelista and Massimo Bergamasco
Multimodal Technol. Interact. 2026, 10(9), 87; https://doi.org/10.3390/mti10090087 - 23 Aug 2026
Viewed by 731
Abstract
Virtual reality (VR) applications are often designed as episodic experiences, with limited support for persistence, longitudinal revisitation, and structured integration of interaction data across sessions. This paper presents a persistent multi-user VR garden architecture that combines snapshot-based state restoration, structured event, movement, and [...] Read more.
Virtual reality (VR) applications are often designed as episodic experiences, with limited support for persistence, longitudinal revisitation, and structured integration of interaction data across sessions. This paper presents a persistent multi-user VR garden architecture that combines snapshot-based state restoration, structured event, movement, and transcript records, an asymmetric owner–visitor workflow, cloud-mediated speech transcription, and deferred AI-supported synthesis. The architecture separates the current spatial configuration of the environment from the interaction traces through which it evolves, supporting repeated access, state restoration, historical consultation, and post-hoc processing. A controlled technical validation using synthetic or researcher-generated inputs was conducted through Unity Editor/backend tests and on Meta Quest 3 hardware. Persistence was evaluated at 20, 100, and 250 objects in the Editor and at 1, 50, and 100 objects on Quest, with two Quest replicas per load. Application-level visitor restrictions were verified across nine prohibited write operations. A frozen production speech-to-text corpus completed 40/40 requests with a micro-averaged word error rate of 9.50% and a median end-to-end latency of 1.619 s. The deferred AI pipeline was additionally verified as a functioning technical integration, while limitations in semantic-detail preservation were observed. The results support the implementation-level feasibility of the proposed persistent and traceable VR architecture. The present evaluation does not establish backend-level authorization guarantees, general AI or NPC-grounding performance, user outcomes, or clinical effectiveness. Full article
(This article belongs to the Topic AI-Based Interactive and Immersive Systems)
►▼ Show Figures

Figure 1

Previous Issue
Next Issue
Back to TopTop