Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

Search Results (141)

Search Parameters:
Keywords = simulation of multimodal trains

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
34 pages, 6523 KB  
Article
A Blockchain-Enabled Federated Neuro-Symbolic Framework for Secure Wearable Biosensor-Based Health Monitoring
by Khulud Salem Alshudukhi and Noshina Tariq
Biosensors 2026, 16(8), 442; https://doi.org/10.3390/bios16080442 - 16 Aug 2026
Viewed by 238
Abstract
Wearable biosensors generate continuous physiological data in smart Internet of Disease (IoD) environments. These data can support early disease detection and remote patient monitoring. However, wearable data are often noisy, sensitive, and distributed across different devices. This paper proposes a multimodal neuro-symbolic model [...] Read more.
Wearable biosensors generate continuous physiological data in smart Internet of Disease (IoD) environments. These data can support early disease detection and remote patient monitoring. However, wearable data are often noisy, sensitive, and distributed across different devices. This paper proposes a multimodal neuro-symbolic model to overcome these limitations and incorporates it into a secure Edge–Fog–Cloud framework for anomaly detection in smart healthcare applications. The proposed system integrates the semantic analysis of clinical text using Bio-ClinicalBERT with temporal numerical data using an LSTM-based model, creating a unified neuro-symbolic artificial intelligence (AI) pipeline. Initial data processing is performed at the Edge, whereas inference is carried out at distributed Fog nodes for low-latency anomaly detection. Model training is handled in the Cloud, and privacy-preserving federated learning (FL) is supported through Homomorphic Encryption (HomEnc) to facilitate collaborative model training without sharing raw patient data. A sharded Tangle ledger is also used, with transactions broadcast by the Fog nodes and validated in the Cloud to create tamper-evident transaction logs. Furthermore, Honey Encryption (HoneyEnc) is integrated into the Fog layer to enhance security against brute-force attacks. Experimental results show that the proposed framework achieved 99.22% accuracy and a 99.31% F1-score on the held-out test set, with bootstrap 95% confidence intervals of 98.96–99.47% for accuracy and 99.08–99.53% for the F1-score. It also reduced detection latency from 185 ms in the baseline setting to approximately 50 ms in the Fog-inference setting. The blockchain layer achieved approximately 500 Transactions Per Second (TPS), while higher throughput was observed under increased transaction load and shard parallelism. Because the evaluation is based on synthetic multimodal EHR-like data and controlled simulations, the reported findings should be interpreted as proof-of-concept internal validation rather than evidence of deployment-ready clinical generalizability; external validation using real wearable biosensor data, hospital IoMT streams, or public clinical datasets such as MIMIC-III/MIMIC-IV is required before clinical deployment. These results highlight the potential of the proposed system for secure data processing and trustworthy anomaly detection in smart healthcare environments. Full article
(This article belongs to the Special Issue Wearable Biosensors and Health Monitoring)
Show Figures

Figure 1

25 pages, 8734 KB  
Article
CORRECT-Net: A Multimodal Vibration–Current Fusion Network for Coal–Rock Cutting State Recognition in Shearers
by Lijuan Zhao, Zhanpeng Zhang, Yadong Wang, Tiangu Wu and Jie Hao
Sensors 2026, 26(16), 5181; https://doi.org/10.3390/s26165181 - 16 Aug 2026
Viewed by 281
Abstract
Coal–rock cutting state recognition is essential for adaptive cutting and intelligent speed regulation in shearers. To address the limited representational capability of individual signals, the confusion between adjacent gangue-bearing cutting conditions, and the domain discrepancy between simulation and experimental data, a vibration–current multimodal [...] Read more.
Coal–rock cutting state recognition is essential for adaptive cutting and intelligent speed regulation in shearers. To address the limited representational capability of individual signals, the confusion between adjacent gangue-bearing cutting conditions, and the domain discrepancy between simulation and experimental data, a vibration–current multimodal fusion method based on CORRECT-Net is proposed. First, an EDEM–RecurDyn–MATLAB/Simulink co-simulation system was developed to generate cutting records for four coal–rock states. After screening for physical equivalence and label conflicts, 158 valid records were retained and grouped into 150 physical-condition groups, which were partitioned at the group level into training, validation, and test sets. Subsequently, SincNet was employed to extract frequency-band-constrained features, a Transformer was used to model long-range temporal dependencies, and a residual importance-guided GATv2 module was introduced to perform cross-modal fusion of vibration-impact and current-load features. On 2500 test windows, CORRECT-Net achieved an accuracy of 96.20% ± 0.11%, a macro-F1 score of 95.14% ± 0.21%, and a hazardous-condition miss rate of 0.58% ± 0.13%. Compared with the multimodal 1D-CNN, TCN, and Bi-LSTM models, CORRECT-Net improved the accuracy by 8.80, 4.00, and 2.08 percentage points, respectively. In the progressive ablation study, the accuracy increased from 87.40% ± 0.26% to 96.20% ± 0.11%, while the macro-F1 score increased from 84.57% ± 0.34% to 95.14% ± 0.21%. Under Gaussian noise with a standard deviation of 0.05, the model retained an accuracy of 92.76% ± 0.24%. When the vibration and current modalities were separately unavailable, the corresponding accuracies were 86.56% ± 0.37% and 92.44% ± 0.25%, respectively. A five-fold simulation-to-experiment transfer evaluation was further conducted at the independent-run level using five experimental records per class. Without adaptation using experimental samples, the model achieved an accuracy of 91.33% ± 5.19%. When 20% and 50% of the experimental windows were used for adaptation, the accuracy increased to 96.33% ± 0.75% and 98.67% ± 1.39%, respectively. These results demonstrate that CORRECT-Net effectively integrates mechanical vibration responses and motor-load information and, under the present simulation and experimental conditions, achieves high recognition accuracy, a low hazardous-condition miss rate, and effective adaptability to the experimental domain. Full article
(This article belongs to the Section Industrial Sensors)
Show Figures

Figure 1

13 pages, 535 KB  
Review
Artificial Intelligence in Cardiac Surgery and Surgical Training: Opportunities, Risks, and Safeguards for Preserving Expertise
by Lazar Velicki, Aleksandra Milovancev, Andrej Preveden, Jelena Vuckovic, Miodrag Belopavlovic, Milan Rodic, Nenad Filipovic and Djordje Jakovljevic
J. Clin. Med. 2026, 15(16), 6313; https://doi.org/10.3390/jcm15166313 - 15 Aug 2026
Viewed by 198
Abstract
Artificial intelligence (AI) is entering cardiac surgery through predictive modelling, multimodal imaging, perioperative monitoring, workflow automation, and emerging computer-vision applications. The most mature evidence concerns risk prediction before and after surgery. Even in this domain, however, systematic reviews show that improvements over conventional [...] Read more.
Artificial intelligence (AI) is entering cardiac surgery through predictive modelling, multimodal imaging, perioperative monitoring, workflow automation, and emerging computer-vision applications. The most mature evidence concerns risk prediction before and after surgery. Even in this domain, however, systematic reviews show that improvements over conventional statistical models are often modest and that routine clinical implementation remains limited. In surgical education, simulation, automated video analysis, and objective performance metrics may expand opportunities for deliberate practice and provide feedback that is less dependent on individual observers. Most of this evidence comes from general, laparoscopic, urological, and robotic surgery rather than cardiac-specific training, and its transferability should not be assumed. The same technologies also create risks. Automation bias, cognitive off-loading, reduced exposure to failure management, and displacement of mentor–trainee interaction may weaken the independent judgement on which safe cardiac surgery depends. Opaque models, dataset shift, inequitable performance, and uncertain accountability add further clinical and ethical concerns. This narrative review examines the current and emerging roles of AI across the cardiac surgical pathway and in cardiothoracic training, while distinguishing demonstrated applications from plausible but unproven uses. We propose a human-in-command framework based on external validation, local performance testing, transparent intended use, preserved manual and crisis-management competencies, simulation of technology failure, faculty oversight, competency-based credentialing, and continuous audit. AI should be judged not by technical novelty alone but by whether it improves care while preserving the ability of surgeons and teams to operate safely when the technology is unavailable or wrong. Full article
(This article belongs to the Special Issue Current Advances and Future Perspectives in Cardiothoracic Surgery)
Show Figures

Graphical abstract

27 pages, 23489 KB  
Article
Toward Self-Evolving Lunar Robotic Autonomy Through Contract-Governed Skill Registration
by Bingqi Huang, Bingchuan Wei, Yingkai Cai and Zhaokui Wang
Astronautics 2026, 1(3), 15; https://doi.org/10.3390/astronautics1030015 - 11 Aug 2026
Viewed by 183
Abstract
Permanent lunar habitation will require robotic systems that can maintain infrastructure, recover from local failures, and acquire new operational capabilities under limited Earth supervision. Existing planetary robots are largely fixed-function specialists, while end-to-end foundation-model policies remain difficult to validate and extend for safety-critical [...] Read more.
Permanent lunar habitation will require robotic systems that can maintain infrastructure, recover from local failures, and acquire new operational capabilities under limited Earth supervision. Existing planetary robots are largely fixed-function specialists, while end-to-end foundation-model policies remain difficult to validate and extend for safety-critical surface operations. We present SELENE (Self-Evolving Lunar Embodied ageNt Ecosystem), an architectural proposal for contract-governed lunar robotic autonomy centered on a shared Atomic Action Library A. The key abstraction is the Atomic Action Contract: a typed skill interface that specifies parameters, preconditions, goal predicates, execution bindings, safety envelopes, runtime reports, and validation metadata. Through this contract, a VLM-driven Cognitive Agent plans over executable skills, a multi-modal Execution Agent realizes them through optimization-based controllers, Vision–Language–Action (VLA) policies, Vision–Language–Navigation (VLN) policies, or reinforcement-learned policies, and an offline Evolutionary Agentic Framework synthesizes and registers new candidate contracts without modifying the planner or the execution interface. This paper presents an architecture-level validation of that contract mechanism. We instantiate SELENE across two heterogeneous pathways on LunarBot and its simulation counterpart, with optimization-based control supported as a third execution modality. A pre-trained VLA policy adapted from 100 teleoperated demonstrations achieves 29/30 task success (96.7 percent) in in-domain trials on the physical LunarBot. A curriculum–RL policy instantiates the traversal pathway in simulated lunar-gravity terrain. Together, these results show that the Atomic Action Contract can serve as a common registration and dispatch interface across heterogeneous control modalities. The same contract layer also defines the path toward runtime gap-triggered self-evolution, mission-grade admission, and lunar-environment validation in subsequent system-level studies. Full article
Show Figures

Figure 1

27 pages, 21309 KB  
Article
Integrating Real Tool Interaction and Multimodal Operator Monitoring in Immersive Simulation for Human-Centred Assessment
by Davide Fabiocchi, Marco Carnevale and Hermes Giberti
Electronics 2026, 15(16), 3525; https://doi.org/10.3390/electronics15163525 - 8 Aug 2026
Viewed by 218
Abstract
The transition from Industry 4.0 to Industry 5.0 is increasing the need for design approaches that place human centrality, safety, and ergonomics at the core of system development. In this context, immersive simulation is evolving from a training-oriented technology into a controlled environment [...] Read more.
The transition from Industry 4.0 to Industry 5.0 is increasing the need for design approaches that place human centrality, safety, and ergonomics at the core of system development. In this context, immersive simulation is evolving from a training-oriented technology into a controlled environment for the observation of human behaviour and task execution. However, many Virtual Reality applications still rely on generic interaction devices that are not sufficiently representative of real tool-mediated operations, limiting the reliability of ergonomic and behavioural assessment. This paper proposes an anthropocentric framework for human-centred assessment in immersive simulation. The framework integrates four main components: a task-oriented Scenario Digital Twin, a Physical-Tool-in-the-Loop module based on a real instrument synchronised with its virtual counterpart, an Interaction Engine for state-dependent action management, and an Operator-in-the-Loop module coupled with a Human-Centred Assessment Layer. The framework is instantiated through an immersive hazelnut pruning simulator, selected as a representative case study because pruning involves non-neutral postures, irreversible actions, and strong dependence on tool handling. The reference implementation combines reality-based reconstruction of hazelnut trees, interactive branch-cutting logic, an instrumented electric pruning shear, full upper-body embodiment, and multimodal data acquisition. In particular, the proposed architecture supports the extraction of body motion through markerless multi-camera pose estimation and the acquisition of eye-related variables through the head-mounted display. A preliminary experimental session demonstrates the technical feasibility of the proposed architecture by showing that multimodal data, including reconstructed 3D body kinematics and eye-tracking signals, can be successfully acquired during immersive task execution, providing a basis for subsequent ergonomic analysis. The results show that a real instrumented tool, a task-oriented digital twin, and a continuous monitoring pipeline can be integrated within a single immersive platform, as well as that the assessment pipeline is sensitive to differing postural configurations during simulated task execution. They do not establish equivalence between behaviour in the simulator and behaviour during real-world pruning; dedicated cross-modal validation is therefore required. The framework is intended to support the evolution of immersive simulation toward assessment-oriented applications, contributing to a broader Safety-by-Design perspective in which human behaviour and interaction quality are considered as early-stage design inputs. Full article
Show Figures

Figure 1

24 pages, 10878 KB  
Review
Artificial Intelligence for Hydraulic-Fracturing Decision Support: A Workflow-Oriented Critical Review
by Xiaobing Bian, Jiaxing Zhou, Liang Fu, Aoran Jin and Wei Zhang
Processes 2026, 14(16), 2537; https://doi.org/10.3390/pr14162537 - 7 Aug 2026
Viewed by 538
Abstract
Hydraulic fracturing is a critical technology for unconventional oil and gas development, but its performance is strongly affected by geological heterogeneity, complex fracture propagation, operational uncertainty, and nonlinear interactions among engineering parameters. Artificial intelligence (AI) provides tools for extracting relationships from geological, geophysical, [...] Read more.
Hydraulic fracturing is a critical technology for unconventional oil and gas development, but its performance is strongly affected by geological heterogeneity, complex fracture propagation, operational uncertainty, and nonlinear interactions among engineering parameters. Artificial intelligence (AI) provides tools for extracting relationships from geological, geophysical, operational, and production data. This structured narrative review synthesizes AI applications across four sequential stages of the hydraulic-fracturing workflow: sweet-spot identification, fracturing-parameter optimization, operational diagnosis and risk warning, and post-fracturing flowback prediction and control. Representative studies reported sweet-spot classification accuracy of 97.5% and R2 = 0.97 for production-performance prediction; a simulator-coupled optimization study reported a 13% economic improvement, and field-data models used cohorts of up to 295 wells. Operational studies reported point-event recognition above 97%, pressure forecasting 30 s ahead, and risk forecasts over three consecutive 60 s intervals. A post-fracturing model trained on 286 wells predicted responses over 30-, 90-, 180-, and 360-day horizons. These values are study-specific and are not directly comparable because the datasets, targets, partitions, and metrics differ. Collectively, the evidence indicates measurable but uneven progress; field readiness remains limited by data quality, multimodal alignment, physical consistency, uncertainty quantification, external validation, and weak coupling between model outputs and operational decisions. The review contributes a reproducible workflow-oriented coding framework and defines validation and deployment priorities for reliable, interpretable, and executable AI-assisted fracturing decision support. Full article
(This article belongs to the Special Issue Application of Artificial Intelligence in Oil and Gas Engineering)
Show Figures

Figure 1

30 pages, 1256 KB  
Article
Multimodal History-Window Gated-Attention Soft Actor-Critic for Urban Low-Altitude UAV Navigation
by Xi You and Wenjun Yi
Drones 2026, 10(8), 605; https://doi.org/10.3390/drones10080605 - 5 Aug 2026
Viewed by 540
Abstract
Urban low-altitude unmanned aerial vehicle (UAV) navigation combines partial observability, building occlusion, wind disturbance, and continuous control. This study develops and evaluates HW-GA-SAC, a multimodal history-window Soft Actor-Critic (SAC) policy for procedurally generated three-dimensional MuJoCo cities. A Gated Transformer-XL (GTrXL)-inspired gated-attention encoder processes [...] Read more.
Urban low-altitude unmanned aerial vehicle (UAV) navigation combines partial observability, building occlusion, wind disturbance, and continuous control. This study develops and evaluates HW-GA-SAC, a multimodal history-window Soft Actor-Critic (SAC) policy for procedurally generated three-dimensional MuJoCo cities. A Gated Transformer-XL (GTrXL)-inspired gated-attention encoder processes a fixed eight-step navigation history, while a current-frame safety branch supplies vertical clearance, sparse Light Detection and Ranging (LiDAR)-like range sectors, and handcrafted safety cues directly to the actor and critic. The policy uses obstacle-related observations and reward shaping to support collision avoidance; it does not include constrained policy optimization or a separate runtime safety filter. In a seven-method comparison using five training seeds and five evaluation layouts, HW-GA-SAC achieved a 96% ± 3% success rate, 207 ± 16 average return, and 3% ± 4% timeout rate. Feedforward SAC achieved 92% ± 11% success and a 7% ± 10% timeout rate, but its successful paths were more direct. Five-seed learning curves, city-split evaluation, wind sensitivity, sensing perturbations, inference profiling, and ablation studies further characterize the method. Within this simulation protocol, HW-GA-SAC provides the strongest completion-oriented performance, with a measurable trade-off between task completion and path directness. Full article
(This article belongs to the Section Innovative Urban Mobility)
Show Figures

Figure 1

17 pages, 229 KB  
Article
From Alignment to Evocation: On the Capability Boundaries and Collaborative Paths of AI Art Creation—A Framework Based on the Neuroaesthetic “Ring Scale” and Prompt Engineering
by Xianqun Yi and Hongsheng Li
Arts 2026, 15(8), 179; https://doi.org/10.3390/arts15080179 - 3 Aug 2026
Viewed by 305
Abstract
Recent generative art outputs across music, literature, painting and moving-image media have attracted extensive scholarly and public interest, yet evaluations of their creative capacities are mostly limited to informal observational accounts. Drawing on neuroaesthetic reasoning, this paper puts forward a dual-layer analytical framework [...] Read more.
Recent generative art outputs across music, literature, painting and moving-image media have attracted extensive scholarly and public interest, yet evaluations of their creative capacities are mostly limited to informal observational accounts. Drawing on neuroaesthetic reasoning, this paper puts forward a dual-layer analytical framework that differentiates two distinct modes of aesthetic reception: Alignment, defined as statistical template matching, and evocation, referring to the novel association of scattered embodied memory fragments. Building on this binary categorization, the study introduces the tentative Ring Scale taxonomy—a figurative target-shooting metaphor rather than quantitative metric—as a purely descriptive tool for stratifying relative aesthetic evocation intensity. This framework further unpacks the neurocognitive underpinnings of auditory, visual and textual aesthetic pathways, alongside their combined multimodal interactions within film and television works. It tentatively accounts for why generative systems tend to deliver more cohesive aesthetic outcomes within the auditory domain, and hypothesises a present functional limitation of current large models: these systems perform comparatively well within Alignment-driven aesthetic effects, while layered high-order evocation remains constrained by inherent structural limitations of statistical training architectures. From this diagnostic observation, three directional paradigm shifts for human–AI collaborative creation are outlined: shifting from human substitution to human–machine complementarity, shifting from exhaustive template imagery generation to targeted latent fragment elicitation, and shifting from optimising figurative Ring-tier descriptive labels to pursuing transformative aesthetic fission effects. The study frames imaginative cognition as the central driving force behind fruitful human–AI co-creation, and positions prompt engineering as the actionable operational bridge connecting human imaginative thought to machine-executable generative parameters. Three tentative prompt design tactics are then elaborated: physiological arousal framing, multisensory scenario simulation prompts, and intentional strategic blank-leaving. Additionally, this work discusses the plausible constructive functions of model hallucination phenomena when viewed through the lens of high-tier aesthetic evocation, rather than merely framing such outputs as technical errors. All judgments and tier comparisons raised throughout the paper are framed as unvalidated observational hypotheses open to empirical testing. To facilitate follow-up empirical scrutiny, the paper collates a full set of testable hypotheses derived from its theoretical reasoning and outlines feasible experimental validation pipelines, with an open call for controlled empirical research to corroborate or refine the proposed qualitative framework. Full article
20 pages, 10023 KB  
Article
Multimodal Adversarial Transfer Learning for Bearing Fault Diagnosis of Unmanned Mining Trucks in Realistic Noisy Environments
by Haifeng Han, Rui Yang, Jianjian Yang and Chenyu Liu
Sensors 2026, 26(15), 4890; https://doi.org/10.3390/s26154890 - 3 Aug 2026
Viewed by 296
Abstract
Unmanned mining trucks operate in harsh environments such as those in open-pit mines, where online fault diagnosis of critical drivetrain bearings faces severe challenges including slow response, high precision requirements, and strong interference from realistic on-site noise. To address the insufficient generalization capability [...] Read more.
Unmanned mining trucks operate in harsh environments such as those in open-pit mines, where online fault diagnosis of critical drivetrain bearings faces severe challenges including slow response, high precision requirements, and strong interference from realistic on-site noise. To address the insufficient generalization capability of existing diagnostic methods in real-world noisy scenarios, this paper proposes a multimodal adversarial transfer learning framework for bearing fault diagnosis in unmanned mining trucks. First, to bridge the domain shift gap between laboratory data and on-site truck data, an augmented multimodal dataset is constructed based on real-vehicle noise grafting. This approach fuses authentic background noise collected from the field with clean laboratory fault signals, thereby simulating graded on-site interference. Second, a deep feature extraction network integrating CNN, ViT, and CBAM attention mechanisms is designed. Building upon this backbone, an adversarial training scheme combined with a hierarchical adaptive fine-tuning strategy is introduced to formulate a domain-adversarial transfer learning model. This model is capable of extracting robust features that are both fault-discriminative and domain-invariant from multimodal signals (vibration and current). Experimental results on the constructed noise-augmented dataset demonstrate that the proposed method maintains high diagnostic accuracy in cross-domain scenarios with strong noise and limited samples, significantly outperforming conventional approaches. This study provides an effective technical pathway for real-time and highly reliable “edge-terminal” fault diagnosis of unmanned mining trucks operating in realistic noisy environments. Full article
(This article belongs to the Section Fault Diagnosis & Sensors)
Show Figures

Figure 1

31 pages, 4419 KB  
Article
An Approach to Improving Lane-Changing Competence of Novice Drivers Based on Observational Learning
by Jing Liu, Yi Feng and Kang Jiang
Systems 2026, 14(8), 910; https://doi.org/10.3390/systems14080910 - 1 Aug 2026
Viewed by 258
Abstract
Novice drivers often exhibit poor lane-changing competence due to ineffective risk identification and control instability. Existing driver training primarily focuses on mechanical skills, lacking systematic psychological interventions. To address this gap, this study constructs a four-stage intervention program (attention, retention, reproduction, and motivation) [...] Read more.
Novice drivers often exhibit poor lane-changing competence due to ineffective risk identification and control instability. Existing driver training primarily focuses on mechanical skills, lacking systematic psychological interventions. To address this gap, this study constructs a four-stage intervention program (attention, retention, reproduction, and motivation) utilizing Bandura’s observational learning theory. A within-subject pre–post design driving simulator experiment was conducted with 46 drivers (16 experienced and 30 novices). Multimodal data—including eye movements, physiological responses, and vehicle kinematics—were collected and analyzed. The results demonstrate that the proposed intervention effectively bridged the initial significant capability gap between the two groups. Post-intervention, novice drivers exhibited significantly optimized visual search strategies, reduced mental workload, and enhanced vehicle control stability during lane-changing maneuvers. Ultimately, no statistical differences remained between novices and experienced drivers across all core indicators, fully validating the research hypotheses. These findings offer robust theoretical support and a practical paradigm for reforming novice driver training, contributing to reduced lane-changing accidents and enhanced active road safety. Full article
Show Figures

Figure 1

26 pages, 6642 KB  
Article
MiniUAV-VLA: A Compact Vision–Language–Action Model for Cooperative Multi-UAV Search and Elimination via MARL Expert Distillation
by Hongwei Han, Guanghong Gong and Ni Li
Drones 2026, 10(8), 572; https://doi.org/10.3390/drones10080572 - 27 Jul 2026
Viewed by 456
Abstract
Coordinating multiple unmanned aerial vehicles (UAVs) for cooperative missions requires agents that perceive their environment, reason about objectives, and generate joint actions. Vision–language–action (VLA) models unify these capabilities but lack a principled source of multi-agent training data and suffer from a training–inference discrepancy [...] Read more.
Coordinating multiple unmanned aerial vehicles (UAVs) for cooperative missions requires agents that perceive their environment, reason about objectives, and generate joint actions. Vision–language–action (VLA) models unify these capabilities but lack a principled source of multi-agent training data and suffer from a training–inference discrepancy in closed-loop control. We propose MiniUAV-VLA, a compact centralized VLA controller for simulated multi-UAV search-and-elimination based on multi-agent reinforcement learning (MARL) expert distillation. A QMIX expert policy achieving 100% mission success generates multimodal demonstrations pairing rendered tactical map images with structured textual state prompts. A 158 M-parameter VLA model with approximately 65 M trainable parameters in the MiniMind-3V backbone and vision projection is fine-tuned with a multi-agent discrete action head that jointly predicts actions for all UAVs in a single forward pass. We identify a training–inference feature mismatch in behavior cloning and address it via prompt-end action pooling, which extracts action-relevant hidden states at the user–prompt boundary rather than after the generated response. In closed-loop evaluation with four drones and six mobile targets averaged over five evaluation seeds, MiniUAV-VLA reaches 74.4 ± 4.6% mission success against 9.4 ± 2.1% for a random policy and 16.2 ± 3.2% for an observation-limited greedy baseline. Across five independent training runs, prompt-end action pooling improves mean closed-loop success from 40.6% to 76.2% over the last-token alternative. These results support MARL expert distillation as a data-efficient route to compact multi-agent VLA control in this simulated setting. Full article
Show Figures

Figure 1

22 pages, 4181 KB  
Article
Latency-Aware Hybrid Transformer–Capsule Network for Audio-Visual Emotion Recognition in Edge–Fog–Cloud Environments
by Abhinav Shukla, Deepika Pahuja, Ayush Kumar Agrawal, R Kanesaraj Ramasamy and Parul Dubey
Algorithms 2026, 19(8), 626; https://doi.org/10.3390/a19080626 - 27 Jul 2026
Viewed by 293
Abstract
Audio-visual emotion recognition (AVER) is central to affective computing systems that require reliable, real-time interpretation of human emotions. However, many existing multimodal models treat feature learning and deployment efficiency separately, limiting their ability to preserve hierarchical facial relationships, capture long-range speech dynamics, and [...] Read more.
Audio-visual emotion recognition (AVER) is central to affective computing systems that require reliable, real-time interpretation of human emotions. However, many existing multimodal models treat feature learning and deployment efficiency separately, limiting their ability to preserve hierarchical facial relationships, capture long-range speech dynamics, and operate with low latency in distributed settings. This study proposes a latency-aware hybrid Transformer–capsule network for audio-visual emotion recognition in a simulated edge–fog–cloud environment. The visual stream employs a CNN–Capsule branch to retain spatial hierarchies in facial expressions, while the audio stream uses a CNN–Transformer branch to learn local spectral patterns and long-range temporal dependencies from speech. A cross-modal Transformer fusion module integrates complementary emotional cues, and a latency-aware task-allocation mechanism allocates preprocessing, inference, and training-related operations across edge, fog, and cloud layers according to workload, node capacity, and communication delay. Unlike approaches that optimize multimodal representation learning and distributed deployment as separate problems, the proposed framework adopts a deployment-aware co-design in which spatial visual representation, temporal acoustic modeling, multimodal interaction, and deterministic latency-aware task allocation are coordinated within a unified processing pipeline. The framework is evaluated on RAVDESS, CREMA-D, and SAVEE using a subject-independent protocol. Experimental results show an average accuracy of 91.5%, an F1-score of 90.7%, an MCC of 0.894, and an AUC of 0.950. The framework further incorporates a deterministic latency-aware task-allocation mechanism for coordinating operations across edge, fog, and cloud resources. Physical-device deployment and comprehensive resource profiling remain subjects for future validation. Full article
Show Figures

Figure 1

28 pages, 952 KB  
Article
An IoT-Ready Context-Aware Patient State Framework with LLM-Driven Recommendation for Shoulder Rehabilitation
by Jonghyeok Mun, Nackhwan Kim and Jongsun Choi
Electronics 2026, 15(15), 3307; https://doi.org/10.3390/electronics15153307 - 27 Jul 2026
Viewed by 344
Abstract
In Internet-of-Things (IoT)-based rehabilitation, patient data from wearable sensors (IMU, EMG, heart rate), clinical assessments, and surveys differ in format and granularity, complicating unified patient-state construction. Existing AI-based systems often omit fatigue level or rehabilitation stage, or they place large language models (LLMs) [...] Read more.
In Internet-of-Things (IoT)-based rehabilitation, patient data from wearable sensors (IMU, EMG, heart rate), clinical assessments, and surveys differ in format and granularity, complicating unified patient-state construction. Existing AI-based systems often omit fatigue level or rehabilitation stage, or they place large language models (LLMs) in the clinical decision-making role, exposing patients to hallucination risk. We propose an IoT-ready framework comprising three components: an input-source-independent context abstraction pipeline, a digital twin that simulates clinical-score trajectories, and an LLM Agent that interprets the outputs of a deterministic algorithm and a digital twin as (subject, predicate, object) Triplets to produce a retrieval-augmented clinical report. The algorithm—not the LLM—selects the 13-exercise sequence from Shoulder Pain and Disability Index (SPADI) item-level responses. Explicit per-layer schemas let IoT-sensor branches be added without changing the downstream interface. Here we evaluate the framework on the clinical-score path (three patient-reported outcome measures and six range-of-motion measures); the multi-modal IoT branches maintain deployment-target functionality. On 48 IRB-approved shoulder rehabilitation patients (144 longitudinal records; augmented to 7200 only to train the trajectory generator), real-only leave-one-subject-out evaluation gave VAS RMSE 0.658 (0–10) and SPADI RMSE 6.499 (0–100). Ablation across four surface forms revealed a fidelity–accuracy trade-off—Narrative highest on template fidelity (BERTScore F1 0.250), raw JSON highest on judge-rated accuracy—and the framework adopts the balanced-midpoint Triplet form. A claim-level audit of the Triplet-form reports left 22–31% of atomic claims unsupported (Claude Sonnet and GPT-4o judges) regardless of guideline retrieval. A raw-LLM control never reproduced the algorithm-defined sequence exactly (0 of 48), supporting deterministic, auditable sequence selection and the need for clinician review before clinical use. Full article
Show Figures

Figure 1

17 pages, 1460 KB  
Article
A Curriculum-Embedded Two-Session AI Chatbot-Based History-Taking Practicum in Korean Medicine Diagnostics
by In-Young Choi, Jundong Kim, Ji-Hwan Kim, Hye-Yoon Lee, Won-Hwan Park, Chang-Eop Kim and Dong-Woo Lim
Appl. Sci. 2026, 16(14), 7223; https://doi.org/10.3390/app16147223 - 19 Jul 2026
Viewed by 320
Abstract
Background: History-taking is a core clinical competency in Korean medicine diagnostics, but conventional training methods such as peer role-play and standardized patient-based education have limitations in providing repeated, individualized, and scalable practice opportunities. This study aimed to evaluate the feasibility and educational value [...] Read more.
Background: History-taking is a core clinical competency in Korean medicine diagnostics, but conventional training methods such as peer role-play and standardized patient-based education have limitations in providing repeated, individualized, and scalable practice opportunities. This study aimed to evaluate the feasibility and educational value of a two-session AI chatbot-based history-taking practicum with automated feedback in Korean medicine diagnostics. Methods: This prospective single-arm repeated-measures educational study was conducted with fourth-year students at the College of Korean Medicine, Dongguk University, in May and June 2026. A total of 76 students participated in two chatbot-assisted history-taking sessions using dizziness and shoulder pain scenarios. Students completed surveys on baseline AI familiarity, chatbot experience, usability, and self-efficacy. Self-efficacy was assessed at three time points: before the first session, after the first session, and after the second session. Chatbot-generated feedback scores were compared between session 1 and 2 for each scenario using paired complete-case analyses. Open-ended responses were descriptively categorized. Results: Students rated the chatbot-based practicum positively in terms of active participation, perceived usefulness, accessibility, and convenience. Item-level self-efficacy analysis showed significant time effects in two domains: planning the conversation, and closing the conversation appropriately. The overall mean self-efficacy score gradually increased from 3.739 ± 0.546 before the first session to 3.887 ± 0.591 after the second session; however, the overall time effect did not reach statistical significance. Chatbot-generated feedback scores showed scenario-dependent patterns. Scores for the shoulder pain scenario increased from session 1 to session 2 before adjustment, but this change did not remain significant after Holm correction; scores for the dizziness scenario showed a non-significant decreasing trend. Open-ended responses indicated that students valued repeated practice and immediate feedback, while also noting limitations related to feedback accuracy, realism of patient responses, and the lack of physical examination or multimodal diagnostic information. Conclusions: The chatbot-assisted practicum was feasible and favorably perceived by students, with selected item-level changes and a modest non-significant upward trend in overall self-efficacy in this exploratory educational study. These findings support the potential role of AI chatbot-based simulation as a supplementary, scalable tool for repeated history-taking practice and formative feedback, rather than as a replacement for performance-based clinical skills training. Given the single-arm design and reliance on learner-reported outcomes, these findings should be interpreted as exploratory. Future controlled studies should incorporate objective performance outcomes, expert-validated automated scoring, non-AI comparison groups, and more realistic multimodal clinical scenarios to determine the educational effectiveness of chatbot-assisted history-taking training. Full article
(This article belongs to the Special Issue New Insights in Artificial Intelligence and E-Learning)
Show Figures

Figure 1

29 pages, 2871 KB  
Article
Federated Energy-Aware Deep Reinforcement Learning for GNSS-Independent Swarm UAV Autonomy
by Nikolaos Almalis, George Tsihrintzis, George Baris and Nikolaos Armenakis
Electronics 2026, 15(14), 3064; https://doi.org/10.3390/electronics15143064 - 13 Jul 2026
Viewed by 610
Abstract
Achieving scalable swarm autonomy in Global Navigation Satellite System (GNSS)-denied and communication-constrained environments remains an open challenge at the intersection of robotics, distributed optimization, and reinforcement learning. Existing unmanned aerial vehicle (UAV) autonomy frameworks typically decouple navigation, perception, and distributed learning, while assuming [...] Read more.
Achieving scalable swarm autonomy in Global Navigation Satellite System (GNSS)-denied and communication-constrained environments remains an open challenge at the intersection of robotics, distributed optimization, and reinforcement learning. Existing unmanned aerial vehicle (UAV) autonomy frameworks typically decouple navigation, perception, and distributed learning, while assuming centralized coordination or reliable global positioning. This paper introduces a unified federated deep reinforcement learning architecture that enables GNSS-independent multi-UAV autonomy through the principled integration of multi-modal perception, decentralized policy optimization, energy-aware control, and edge-compliant inference. The proposed framework formulates joint navigation and dynamic target tracking as a partially observable Markov decision process optimized via Proximal Policy Optimization (PPO) over structured motion primitives. A communication-efficient federated learning mechanism enables distributed policy convergence under non-independent and identically distributed (non-IID) agent experiences without sharing raw data, establishing a scalable alternative to centralized training. To address sim-to-real discrepancies, the architecture incorporates domain randomization, structured sensor noise modeling, and curriculum-based training to promote robust zero-shot deployment. Multi-agent simulation experiments evaluate the swarm-level and federated-learning behavior of the proposed framework, while single-UAV field deployment evidence using a DJI Matrice 100 platform supports the feasibility of the onboard sensing, perception, and edge-inference pipeline under realistic outdoor conditions. The evaluation demonstrates stable decentralized convergence, improved energy efficiency relative to centralized baselines, robust target-tracking performance under GNSS-denied conditions, and real-time edge-compliant inference. The results establish that federated reinforcement learning can serve as a viable systems-level foundation for resilient, energy-aware, and scalable aerial swarm intelligence, advancing the state of the art in distributed autonomous robotics. Full article
Show Figures

Figure 1

Back to TopTop