Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (209)

Search Parameters:
Keywords = real-time action recognition

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
26 pages, 6192 KB  
Article
Evaluating the Effectiveness of AI-Generated Data for Video-Based Action Recognition
by Kamil Gomulka, Piotr Wozniak and Tomasz Krzeszowski
Electronics 2026, 15(16), 3712; https://doi.org/10.3390/electronics15163712 - 19 Aug 2026
Viewed by 144
Abstract
Human action recognition relies heavily on large-scale annotated video datasets, which are costly and time-consuming to curate, while AI-generated videos offer a promising alternative data source, their effectiveness for training action recognition models remains insufficiently explored. This study evaluates AI-generated videos for action [...] Read more.
Human action recognition relies heavily on large-scale annotated video datasets, which are costly and time-consuming to curate, while AI-generated videos offer a promising alternative data source, their effectiveness for training action recognition models remains insufficiently explored. This study evaluates AI-generated videos for action recognition across convolutional and transformer-based architectures using real, synthetic, and hybrid datasets. To ensure generative diversity and consistency, a structured prompt engineering pipeline combining action descriptions, environmental contexts, and camera viewpoints was developed. Synthetic datasets were generated using the Grok Imagine and Meta AI Vibes video generation models and paired with a 15-class subset of the Human Motion Database 51 (HMDB51) to construct the Generative Synthetic Human Action Recognition Dataset (GenSynth-HARD). To mitigate domain shift arising from discrepancies between real and AI-generated videos, a Conditional Domain Adversarial Network (CDAN) with a dynamically scaled Gradient Reversal Layer (GRL) was integrated for domain feature alignment. The best-performing hybrid model achieved a Top-1 accuracy of 79.92% on the HMDB51 subset, demonstrating that incorporating synthetic videos effectively supports model performance while significantly reducing annotation overhead. Full article
(This article belongs to the Special Issue Convolutional Neural Networks and Vision Applications, 4th Edition)
Show Figures

Figure 1

24 pages, 17590 KB  
Article
Automation of Monitoring Compliance with Technological Regulations Using the Example of the Process of Filling Petroleum Products
by Anatoly Sidorov, Alexey Zaripov, Ivan Tikshaev and Vladislav Pozdyshev
J. Imaging 2026, 12(8), 349; https://doi.org/10.3390/jimaging12080349 - 3 Aug 2026
Viewed by 217
Abstract
In this paper, an approach to automating the monitoring of compliance with regulated technological operations using computer vision methods is proposed and investigated, with the loading of petroleum products considered as a case study. A distinctive feature of this work is the integration [...] Read more.
In this paper, an approach to automating the monitoring of compliance with regulated technological operations using computer vision methods is proposed and investigated, with the loading of petroleum products considered as a case study. A distinctive feature of this work is the integration of a formalized description of the production process in BPMN notation with an object detection system, enabling not only object recognition but also interpretation of the sequence of technological actions performed by personnel. Based on the collected and annotated dataset containing more than 6000 images, a YOLOv11 neural network model was trained to monitor key stages of a technological operation. The experimental results show that the trained model provides high accuracy in detecting objects during the daytime (mAP50 is approximately 0.98), while maintaining the ability to work in real time. The results obtained confirm their applicability in industrial conditions. The work revealed the dependence of the quality of computer vision system functioning on the illumination conditions of the production area. It has been established that at night there is a significant decrease in recognition accuracy due to the presence of glare from lighting sources directed at the camera area. The results obtained make it possible to substantiate the need to take into account lighting factors when designing video monitoring systems for technological processes. To move from the level of object detection to monitoring compliance with regulations, an algorithm for interpreting detected objects has been developed, which ensures the fixation and analysis of the sequence of operations performed. Experimental tests conducted at the existing production site have confirmed the possibility of automated detection of violations of technological regulations and an increase in the level of industrial safety. The directions for further development of the proposed approach have also been identified, including the expansion of the training sample, taking into account a variety of production scenarios, and the development of methods to increase the stability of the system in difficult light conditions. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

38 pages, 3059 KB  
Review
Review: Techniques in Egocentric Multi-View Image Analysis: Advances, Challenges, and Future Directions
by Duc Tri Phan and Hong Duc Nguyen
J. Imaging 2026, 12(7), 324; https://doi.org/10.3390/jimaging12070324 - 17 Jul 2026
Viewed by 529
Abstract
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion [...] Read more.
Egocentric multi-view image analysis refers to the processing of utilizing synchronized video streams captured from multiple wearable cameras worn on the head or body, providing complementary first-person perspectives of dynamic, real-world interactions. Unlike single-view egocentric vision, which may suffer from severe occlusions, motion blur, and limited field-of-view or traditional fixed-camera multi-view setups (assuming static geometry and controlled environments), egocentric multi-view systems leverage body-worn rigs to enable a more robust and flexible 3D understanding in open-world, mobile scenarios. In this work, we present a systematic survey of advancements in cross-view feature fusion, geometric consistency enforcement, open-world detection, human–object interaction (HOI) modeling, action segmentation, 3D reconstruction, and novel-view synthesis specifically tailored to wearable multi-camera platforms. Key datasets released between 2024 and 2026—including HOT3D (833 min of synchronized multi-view hand/object interactions from Project Aria and Quest 3), MultiEgo (first multi-egocentric dataset for 4D social scene reconstruction), and Ego-1K (large-scale 12-camera rig for dynamic 3D video synthesis) are thoroughly examined alongside an analysis of integrations with large language models (LLMs) and vision–language models that drive performance gains, typically in the 15–30% range over single-view baselines in hand tracking, HOI recognition, and reconstruction fidelity, although we show through a consolidated meta-analysis that this gain is task-dependent: larger for geometry-bottlenecked tasks such as in-hand object lifting, and smaller, method-dependent, or occasionally negative for semantic-recognition tasks such as keystep recognition under naive view fusion. These methods cover work in multi-view stereo, cross-view learning, and novel-view synthesis while addressing several real-time wearable constraints. Practical applications such as immersive Augmented Reality/Virtual Reality (AR/VR), assistive robotics, and healthcare monitoring are also discussed together with the challenges in motion calibration, benchmark diversity, and edge deployment ability. Thus, in this review, we attempt to fill a critical gap by focusing exclusively on wearable multi-view systems in an open-world setting, synthesizing the latest literature to chart future directions toward more embodied and continual learning agents. Full article
(This article belongs to the Special Issue Techniques in Multi-View Image Analysis)
Show Figures

Figure 1

37 pages, 35970 KB  
Review
A Survey on Action Recognition: Multimodal Approaches, Ethical Considerations, and Feedback Mechanisms
by Bilal AbdulRahman, Zhigang Zhu and Alison Conway
Electronics 2026, 15(14), 3139; https://doi.org/10.3390/electronics15143139 - 16 Jul 2026
Viewed by 377
Abstract
Action recognition has emerged as a critical area of research within the realm of computer vision, driven by the increasing demand for intelligent human–machine systems capable of understanding and interpreting human behaviors in the real world. The ability to decipher intricate details of [...] Read more.
Action recognition has emerged as a critical area of research within the realm of computer vision, driven by the increasing demand for intelligent human–machine systems capable of understanding and interpreting human behaviors in the real world. The ability to decipher intricate details of human actions holds immense potential to improve system design, predictive modeling, data-informed decision-making, and real-time operational improvements across a wide variety of domains. Some examples of applications range from surveillance and real-time management of public spaces and infrastructure systems, to development of predictive modeling and robotic systems for individualized healthcare interventions, to implementing effective human–computer interaction in both professional and recreational settings. This paper provides a comprehensive survey of the current state of action recognition, focusing specifically on three open-world challenges: the integration of multimodalities, the ethical and social implications of these technologies, and the utilization of feedback mechanisms to enhance model performance. We delve into the evolution of action recognition, from early feature-based approaches to the deep learning revolution, emphasizing how the incorporation of multiple sensory modalities—such as visual, audio, and depth data as well as other cues—has advanced the field. Furthermore, we examine the ethical challenges associated with deploying these technologies in the public domain, particularly regarding privacy, bias, and societal impact, and discuss the need for responsible development and regulation. The third focus of the paper is the use of top-down and bottom-up feedback mechanisms within deep learning architectures, exploring how these strategies can mimic human cognitive processes to improve accuracy and reliability in action recognition systems. By identifying current gaps and proposing future research directions, this paper aims to inspire continued innovation in this dynamic and impactful field for intelligent systems. Full article
Show Figures

Figure 1

31 pages, 6966 KB  
Review
Deep Learning for Sensor-Based Sport Performance and Health Monitoring: A Review of Wearable, Vision-Based, and Multimodal Sensing Approaches
by Liu Liu, Xinyu Hu, Hong Wei, Ziqian Yang and Tao Sun
Sensors 2026, 26(14), 4384; https://doi.org/10.3390/s26144384 - 10 Jul 2026
Viewed by 988
Abstract
Recent advances in wearable, vision-based, trajectory, physiological, and multimodal sensing technologies, together with deep learning, have enabled continuous, objective, and individualized assessment of sport performance and athlete health. Unlike prior reviews that primarily focus on a single sensing modality, sport, or algorithmic series, [...] Read more.
Recent advances in wearable, vision-based, trajectory, physiological, and multimodal sensing technologies, together with deep learning, have enabled continuous, objective, and individualized assessment of sport performance and athlete health. Unlike prior reviews that primarily focus on a single sensing modality, sport, or algorithmic series, this review integrates wearable, vision-based, trajectory, physiological, and multimodal sensing streams with deep learning models across both performance analysis and athlete health monitoring, thereby clarifying modality-task-model relationships and translational limitations. This review synthesizes recent progress in sensor-based sports intelligence, focusing on how heterogeneous data streams are transformed into performance- and health-related decision support. The reviewed applications include athlete and ball perception, multi-object tracking, pose estimation, action recognition, trajectory and tactical analysis, training-load and fatigue monitoring, injury-risk prediction, rehabilitation monitoring, and return-to-play support. Deep learning architectures, including CNNs, LSTMs, GRUs, TCNs, Transformers, attention mechanisms, graph neural networks, and multimodal fusion models, are discussed in relation to their suitability for visual, temporal, spatial, physiological, and multisource data. This review further identifies key challenges, including data heterogeneity, annotation scarcity, limited cross-sport and cross-device generalization, real-time deployment constraints, model interpretability, privacy protection, and ethical governance. Moving forward, research efforts should focus on the development of standardized datasets, reliable multimodal data fusion strategies, self-supervised and transfer learning approaches, and deployment on edge or cloud computing platforms. Additionally, enhancing interpretability through explainable AI and implementing closed-loop, individualized monitoring systems are critical. By synthesizing advances in sensing technologies, deep learning methodologies, and real-world applications, this review aims to provide a practical reference for optimizing athletic performance, preventing injuries, guiding rehabilitation, and supporting long-term health management of athletes. Full article
(This article belongs to the Section Wearables)
Show Figures

Figure 1

25 pages, 2842 KB  
Article
Artificial Intelligence-Based Insider-Threat Detection: A Hybrid Explainable Framework with Automated Response and Privilege Containment
by Abdel Rahman Alkharabsheh, Ghaya Binsalma, Mahra Alharmi, Ruqia Alshateri, Shahad Altaee and Mousa Sweidan
Computers 2026, 15(7), 426; https://doi.org/10.3390/computers15070426 - 2 Jul 2026
Viewed by 1231
Abstract
Insider threats continue to be the most persistent and most destructive threat to cybersecurity; malicious or negligent users work only in the real-time restricted area of the organization and are gradually breaking the boundaries of company norms. Conventional rule-based and statistical detection methods [...] Read more.
Insider threats continue to be the most persistent and most destructive threat to cybersecurity; malicious or negligent users work only in the real-time restricted area of the organization and are gradually breaking the boundaries of company norms. Conventional rule-based and statistical detection methods have difficulty detecting inconspicuous, context-dependent, and ever-changing behavior, leading to detection delays and high false-positive rates. Our paper introduces an explainable AI-based Insider-Threat Detection (AIB-ITD) model that integrates enterprise telemetry—including email, web, logon/VPN, and file events—into a unified behavioral framework. The effectiveness of combining heterogeneous behavioral indicators observed in AIB-ITD is consistent with recent behavioral analytics implementations that have demonstrated the value of multimodal user-behavior profiling for insider-threat identification in enterprise environments. The proposed AIB-ITD framework is based on anomaly-driven processing, unsupervised models (Isolation Forest, PCA reconstruction, and Autoencoder) are combined with sequential modeling (with an LSTM Autoencoder) to model both static and temporal deviations in behavior. An ensemble strategy is applied to combine the outputs of these models to yield a probabilistic insider risk score. To improve transparent analysis and to help the analyst gain trust, SHapley Additive Explanations (SHAP) is used to keep every detection outcome transparent and interpretable using the features. It also integrates feature correlation analysis, static vs sequential-model comparisons, and SHAP stability assessment to validate methodological robustness and reproducibility. An experimental review of the hybrid ensemble using the SEI/CMU CERT Insider Threat Dataset reveals that it performs better than single models for anomaly detection and stability, especially with the inclusion of temporal patterns. The assessment prioritizes anomaly score consistency and reliable risk ranking, rather than classification accuracy, to better reflect real deployment scenarios. In addition, an Automated Response and Privilege Containment (ARPC) feature automatically converts risk scores to multilevel mitigation actions that serve to protect the privacy of the user as the least privileged policies are enforced promptly. The proposed model showed superior robustness, stability, and operational effectiveness to classical methods, especially in the presence of scarce labeled data. Through hybrid anomaly recognition, explainable AI and automated response, AIB-ITD is a practical and scalable solution for next-generation insider-threat detection in enterprise systems. Full article
Show Figures

Figure 1

37 pages, 6867 KB  
Article
ITS-Vision: Autonomous Vehicles as Mobile Surveillance Nodes in Intelligent Transportation Systems—A Conceptual Framework and Proof-of-Concept Prototype
by Mirabela-Melinda Medvei, Denis Georgian Gurău and Mihai Coca
Future Internet 2026, 18(7), 349; https://doi.org/10.3390/fi18070349 - 1 Jul 2026
Viewed by 650
Abstract
Crime surveillance in urban environments faces increasing challenges due to dynamic conditions and the demand for real-time monitoring. This paper investigates the use of video data from autonomous vehicles to enhance situational awareness in public spaces through deep learning models optimized for edge [...] Read more.
Crime surveillance in urban environments faces increasing challenges due to dynamic conditions and the demand for real-time monitoring. This paper investigates the use of video data from autonomous vehicles to enhance situational awareness in public spaces through deep learning models optimized for edge processing. High-resolution vehicle-mounted cameras serve as mobile surveillance units capable of real-time object detection, human action recognition, and anomaly detection, bridging the gap between autonomous mobility and urban monitoring. Building on this vision, we introduce ITS-Vision, a generic framework that operationalizes these use cases, enabling autonomous vehicles to function as mobile, context-aware sensing platforms. To validate this approach, we develop prototypes for key ITS-Vision components: a fight detection module using a fine-tuned X3D model, suspect identification via MediaPipe for detection combined with FaceNet for embedding extraction, and a dangerous items detection module using a fine-tuned YOLOv11n model. Due to the limited availability of real-world autonomous vehicle video datasets, experiments were conducted in controlled laboratory environments, demonstrating the feasibility of the proposed architecture and algorithms under simulated conditions. Future work will focus on collecting dedicated datasets and advancing the models toward deployment in real urban scenarios. Full article
(This article belongs to the Section Smart System Infrastructure and Applications)
Show Figures

Figure 1

26 pages, 15054 KB  
Article
Beef Cattle Behavior Recognition Based on Nighttime Farm Videos via Spatio-Temporal Enhancement and Dynamic Fusion
by Yamin Han, Zhenyu Zhang, Wenchao Zhang, Shichao Cao, Yang Sun, Zixin Jia, Danyang Wu, Lyuwen Huang and Hongming Zhang
Animals 2026, 16(12), 1881; https://doi.org/10.3390/ani16121881 - 17 Jun 2026
Viewed by 357
Abstract
Beef cattle behavior provides valuable information regarding their health status. Recently, deep convolutional network-based methods have achieved considerable results in beef cattle behavior recognition. However, their robustness under low-light or dark conditions remains limited, which restricts their application in real farm environments. To [...] Read more.
Beef cattle behavior provides valuable information regarding their health status. Recently, deep convolutional network-based methods have achieved considerable results in beef cattle behavior recognition. However, their robustness under low-light or dark conditions remains limited, which restricts their application in real farm environments. To address this issue, this study constructed a realistic beef cattle behavior dataset in the dark, named Dark Beef Cattle Actions, which was collected under real nighttime farm conditions. The constructed dataset contains 1097 video clips collected from 30 beef cattle and covers 6 behavioral classes, including running, feeding, drinking, grooming, mounting, and fighting. Based on this dataset, we proposed a novel neural network architecture based on spatio-temporal dark enhancement and dynamic fusion for beef cattle behavior recognition in the dark. First, a spatio-temporal dark enhancement module was designed to improve dark video quality while preserving motion features. Second, a dynamic fusion module was introduced to adaptively fuse features from different branches and obtain more discriminative representations. In addition, a joint loss was adopted to optimize both dark enhancement and action recognition. Experimental results on the constructed dataset show that the proposed method achieved a weighted-averaged precision score of 88.47%, a weighted-averaged recall score of 80.18%, an accuracy score of 83.80%, and a weighted-averaged F1-score of 84.12%. Compared with other state-of-the-art methods, the proposed method achieved competitive performance in the recognition of night-time beef cattle behavior. These findings would provide support for intelligent livestock behavior recognition and monitoring in precision farming. Full article
(This article belongs to the Special Issue Artificial Intelligence as a Useful Tool in Behavioural Studies)
Show Figures

Figure 1

24 pages, 8539 KB  
Article
Temporally Consistent Student Behavior Recognition in Smart Classrooms via Attention-Guided Perception and State Estimation
by Shuzhao Zong, Chenyang He, Peng Sun and Chenliang Ma
Electronics 2026, 15(12), 2644; https://doi.org/10.3390/electronics15122644 - 15 Jun 2026
Viewed by 340
Abstract
Recognizing student behaviors in classroom videos remains challenging due to complex backgrounds, frequent occlusions, subtle inter-class motion differences, and temporal jitter in frame-wise predictions. To address these issues, this paper proposes a hybrid student behavior recognition framework that integrates a Multi-branch Spatiotemporal Attention [...] Read more.
Recognizing student behaviors in classroom videos remains challenging due to complex backgrounds, frequent occlusions, subtle inter-class motion differences, and temporal jitter in frame-wise predictions. To address these issues, this paper proposes a hybrid student behavior recognition framework that integrates a Multi-branch Spatiotemporal Attention Network (MSTA-Net) with a Behavior State Kalman Filter (BSKF). At the perceptual level, MSTA-Net employs decoupled channel, spatial, and short-term temporal attention branches to enhance discriminative behavioral features while suppressing irrelevant background information. At the cognitive level, BSKF reformulates behavior recognition as a continuous state estimation problem in a high-dimensional probability space, where behavioral inertia is exploited to smooth noisy observations and improve temporal consistency. Experimental results on the SCB-Dataset and real-world classroom video sequences demonstrate that the proposed method achieves an accuracy of 94.7% and a real-time inference speed of 33 FPS. Compared with purely deep learning-based models, the proposed framework reduces the Action Category Switching (ACS) rate by 50%, indicating substantially improved robustness in long-term behavior recognition. These results suggest that coupling attention-based perception with Kalman-based state estimation provides an effective and efficient solution for reliable student behavior analysis in intelligent classroom environments. Full article
Show Figures

Figure 1

29 pages, 28758 KB  
Article
Spatio-Temporal Feature Enhancement for Recognizing Strongly Correlated Sequential Actions in Aircraft Assembly
by Jiaming Shi, Xiang Huang, Guoyi Hou, Chengda Guo, Qingxue Wang and Yumin Chen
Sensors 2026, 26(12), 3781; https://doi.org/10.3390/s26123781 - 13 Jun 2026
Viewed by 536
Abstract
The positioning and clamping process in aircraft assembly exhibits pronounced long-term temporal correlations and intense human–machine interactions. Consequently, assembly quality depends heavily on operator compliance and consistency. Capturing long-term, strongly correlated features in complex industrial environments remains a significant challenge. To overcome this, [...] Read more.
The positioning and clamping process in aircraft assembly exhibits pronounced long-term temporal correlations and intense human–machine interactions. Consequently, assembly quality depends heavily on operator compliance and consistency. Capturing long-term, strongly correlated features in complex industrial environments remains a significant challenge. To overcome this, this study proposes a Long-Term Strongly Associated Action Recognition Network (LTSA-Net) tailored for aircraft assembly positioning and clamping tasks. Based on the C3D backbone, the model first incorporates the SimAM attention mechanism and BN modules to significantly enhance focus on critical spatiotemporal features. To address the challenge of capturing long-term temporal dependencies, LTSFEM is designed to extract global temporal information accurately. Furthermore, to balance structural lightweight design with real-time inference requirements, the CWSTB module is integrated to achieve substantial parameter compression. In addition, a dedicated aircraft assembly positioning and clamping dataset was constructed, and a robust training framework was established using the AdamW optimizer and Mixup data augmentation. Experimental results demonstrate that LTSA-Net achieves a recognition accuracy of 98.82% on the LTSA-Dataset, with a per-frame inference time of 42 ms, successfully meeting the dual requirements of high precision and real-time performance in industrial scenarios, and providing a practical technical solution for intelligent monitoring of aircraft assembly processes. Full article
(This article belongs to the Section Industrial Sensors)
Show Figures

Figure 1

23 pages, 2117 KB  
Article
A Traffic Police Gesture Recognition Method Based on BiLSTM-Transformer Architecture
by Xiaoyu Zhang, Baohua Guo, Sen Wang, Anthony Sigama and David Bassir
Electronics 2026, 15(12), 2578; https://doi.org/10.3390/electronics15122578 - 11 Jun 2026
Viewed by 376
Abstract
To address the issues of insufficient real-time performance and inadequate modeling of temporal features in traffic police gesture recognition, this paper proposes a method based on skeleton keypoints and hybrid temporal modeling. First, YOLOv11m-Pose is employed to detect human skeleton keypoints in video [...] Read more.
To address the issues of insufficient real-time performance and inadequate modeling of temporal features in traffic police gesture recognition, this paper proposes a method based on skeleton keypoints and hybrid temporal modeling. First, YOLOv11m-Pose is employed to detect human skeleton keypoints in video sequences, extracting reliable two-dimensional skeleton features. Second, this study designs a temporal modeling network that integrates a bidirectional long short-term memory (BiLSTM) with a Transformer. The BiLSTM models local temporal continuity and action transition features between adjacent frames, capturing short-term dynamic changes. The Transformer, through its self-attention mechanism, models global temporal dependencies and weights critical time steps to extract long-range discriminative information. Experimental results demonstrate that the proposed method achieved 98.91% for both Accuracy and F1-Score. In terms of Accuracy, it outperformed the BiLSTM and Transformer models by 2.43% and 7.67%, respectively. It outperforms most methods based on recurrent neural networks and feature fusion. Meanwhile, the model achieves an average inference time of just 1.3299 s per gesture sequence. Consequently, this approach strikes a favorable balance between recognition accuracy and real-time performance, demonstrating significant practical value. Full article
(This article belongs to the Special Issue AI Innovations in Smart Transportation)
Show Figures

Figure 1

28 pages, 3423 KB  
Review
Hydrogel-Based Optical Sensors for Chemical and Biosensing: Materials, Selectivity, and Applications
by Hossein Omidian and Sumana Dey Chowdhury
Appl. Sci. 2026, 16(12), 5867; https://doi.org/10.3390/app16125867 - 10 Jun 2026
Viewed by 378
Abstract
Hydrogel-based optical sensors have emerged as a versatile class of analytical materials that combine soft-matter processability, tunable network chemistry, and compatibility with luminescent, colorimetric, photonic, and hybrid transduction strategies. Progress in the field is driven not by a single sensing mechanism, but by [...] Read more.
Hydrogel-based optical sensors have emerged as a versatile class of analytical materials that combine soft-matter processability, tunable network chemistry, and compatibility with luminescent, colorimetric, photonic, and hybrid transduction strategies. Progress in the field is driven not by a single sensing mechanism, but by the convergence of key advances in material functionalization, embedded selectivity, operation across diverse sample matrices, mechanical and analytical robustness, and usability beyond the laboratory. Current systems include framework-integrated, nanoparticle-doped, probe-functionalized, photonic-crystal, enzyme-immobilized, and device-coupled hydrogels, reflecting growing architectural diversity and application-oriented engineering. Selectivity has likewise advanced from basic interferent screening to recognition-specific, imprinted, and pattern-discriminative formats suited to complex environmental, food, biological, and wearable settings. Evidence of stability, reusability, and deformation tolerance further suggests that many platforms are moving beyond proof-of-concept demonstrations toward credible real-world operation. At the same time, translational priorities such as portability, smartphone readout, implantable and epidermal formats, and multifunctionality spanning antimicrobial action, adsorption, anti-counterfeiting, and device integration are becoming increasingly prominent. Together, these trends show that hydrogel-based optical sensing is maturing into a materially rich, application-responsive domain. The key challenge ahead is to unify materials design, selectivity control, durability, and deployability in standardized, reproducible, and clinically or environmentally credible sensing platforms. Full article
Show Figures

Figure 1

20 pages, 8187 KB  
Article
From IMU Streams to Real-Time Decisions: Past-Only Next-Window Badminton Action Prediction
by Qinglin Zhu, Jiao Wang and Bin Guo
Sensors 2026, 26(12), 3651; https://doi.org/10.3390/s26123651 - 8 Jun 2026
Viewed by 558
Abstract
We study real-time next-window badminton action prediction from wearable IMU streams where the system must predict the action label of the upcoming 100 ms window using past-only (causal) information. To handle severe class imbalance in continuous streams, we employ window-level downsampling of the [...] Read more.
We study real-time next-window badminton action prediction from wearable IMU streams where the system must predict the action label of the upcoming 100 ms window using past-only (causal) information. To handle severe class imbalance in continuous streams, we employ window-level downsampling of the dominant background class and compress multi-sensor time/frequency features using PCA before temporal modeling. We evaluate the full pipeline under a hop-based streaming protocol and show that our BiLSTM + MHSA model achieves high recognition performance (test accuracy 96.36%, Macro-F1 95.82%) while remaining deployable in real time, reaching 58.20 windows/s end to end (including preprocessing), i.e., 5.82× the real-time requirement (10 windows/s under a 100 ms output interval), on a Windows PC with an NVIDIA RTX 3080 GPU. These results support low-latency applications such as live coaching feedback and tactical analytics. Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

16 pages, 1810 KB  
Article
Gaze Tracking- and Facial Movement-Driven Human–Computer Interaction System
by Yue Liu, Yuxiang Li, Lu Leng and Cheonshik Kim
Appl. Sci. 2026, 16(11), 5653; https://doi.org/10.3390/app16115653 - 4 Jun 2026
Viewed by 413
Abstract
With the development of human–computer interaction technology, non-contact interaction based on gaze tracking and facial movements has become a research hotspot. Traditional mouse-and-keyboard methods pose challenges for people with disabilities or limited hand movements, while existing gaze-tracking systems often rely on expensive hardware [...] Read more.
With the development of human–computer interaction technology, non-contact interaction based on gaze tracking and facial movements has become a research hotspot. Traditional mouse-and-keyboard methods pose challenges for people with disabilities or limited hand movements, while existing gaze-tracking systems often rely on expensive hardware or lack sufficient accuracy. This paper designs and implements a real-time system using ordinary cameras, achieving natural, efficient interaction via multimodal input combination. The system uses an improved MobileNetV2 backbone to construct GazeTrackNet for gaze estimation. It adopts MediaPipe Face Mesh to detect facial landmarks. Meanwhile, it applies geometric feature analysis, including eye aspect ratio and mouth aspect ratio, to identify actions such as blinking and mouth opening. It adopts a hybrid control strategy that combines gaze jumping and head fine-tuning, using mouth state as the main control switch. Key contributions include a lightweight gaze-tracking algorithm that enables stable and efficient gaze detection on consumer-grade hardware, a multimodal interaction strategy based on facial movement that improves system stability and ease of use, and a complete prototype system that achieves real-time performance on standard laptops. Experimental results show an average gaze average angle error of 3.0°, 97% eye state recognition accuracy, and end-to-end latency below 70 ms. The system can satisfy the requirements of daily desktop interaction under normal indoor lighting, and shows potential for future barrier-free interaction applications after further validation with target users. Existing gaze-tracking methods either suffer from low precision on lightweight devices or bring heavy computational overhead. Common facial recognition approaches also face frequent false trigger interference. Compared with them, our scheme achieves balanced accuracy and real-time performance via an attention-enhanced structure, and the designed dual anti-shake mechanism effectively suppresses misjudgment, delivering a more stable hands-free interaction experience. Full article
(This article belongs to the Special Issue Image Processing: Technologies, Methods, Apparatus)
Show Figures

Figure 1

23 pages, 9952 KB  
Article
A Bio-Inspired Lightweight Human Action Recognition Method Based on Human Keypoint Detection
by Weihao Huang, Mianting Wu, Weixiong Chen and Qiang Zhou
Biomimetics 2026, 11(5), 355; https://doi.org/10.3390/biomimetics11050355 - 20 May 2026
Viewed by 424
Abstract
Recognizing human actions from static images in complex industrial environments remains challenging due to insufficient feature representation and high computational complexity. This issue is particularly critical in power-grid safety monitoring, where improper worker postures (e.g., bending, climbing, falling) can lead to severe accidents [...] Read more.
Recognizing human actions from static images in complex industrial environments remains challenging due to insufficient feature representation and high computational complexity. This issue is particularly critical in power-grid safety monitoring, where improper worker postures (e.g., bending, climbing, falling) can lead to severe accidents and personal injuries, necessitating automated monitoring systems that operate reliably on resource-constrained edge devices. This study proposes a bio-inspired lightweight recognition framework that integrates an improved YOLO-Pose model with a gated recurrent unit (GRU) network. The scientific motivation is grounded in the observation that the human musculoskeletal system achieves highly efficient motion perception through three key mechanisms: hierarchical muscle coordination providing intrinsic rotation invariance, proprioceptive feedback enabling real-time error correction, and selective neural gating reducing redundant information transmission. These biological principles directly inspire our technical contributions: polar-coordinate encoding provides rotation invariance, three-stage filtering mimics proprioceptive feedback, and GRU gating mirrors selective information propagation. Unlike prior approaches that treat pose-based action recognition as a generic computer vision problem, this work explicitly incorporates anatomical structural constraints into the computational pipeline. The framework addresses three research gaps: (1) existing methods lack biomechanically derived invariance properties; (2) GCN-based approaches use fixed topologies that fail to adapt to occlusion patterns; (3) the trade-off between model complexity and accuracy remains unsatisfactory for edge deployment. Experiments on the self-constructed SKPose dataset demonstrate that the proposed method achieves 95.04% accuracy, outperforming ST-GCN by 3.67 percentage points and 2s-AGCN by 1.94 percentage points, with an inference speed of 48 FPS on 8.7 M parameters in underground power-grid environments and provides practical support for biomimetic perception systems and industrial safety monitoring. Full article
(This article belongs to the Special Issue Bionic Intelligent Robots)
Show Figures

Figure 1

Back to TopTop