electronics-logo

Journal Browser

Journal Browser

Deep Learning Applications on Human Activity Recognition

A special issue of Electronics (ISSN 2079-9292). This special issue belongs to the section "Artificial Intelligence".

Deadline for manuscript submissions: closed (15 July 2026) | Viewed by 4880

Editors


E-Mail Website
Guest Editor
Department of Engineering, University of Messina, C.da di Dio - 98166 Sant'Agata, Messina, Italy
Interests: point cloud analysis and registration; differential entropy analysis; machine vision; human pose estimation; deep learning application

E-Mail Website
Guest Editor
Department of Industrial and Information Engineering and Economics, University of L’Aquila, 67040 L'Aquila, Italy
Interests: reverse engineering; digital twin; deep learning methods for computer vision; human pose estimation; digitalized risk assessment
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

Human activity recognition (HAR) has become a key technology with transformative applications in healthcare, smart environments, security, sports analytics, and human–computer interaction. The adoption of deep learning has significantly enhanced HAR capabilities, enabling the more accurate and scalable recognition of human activities from sensor data, video streams, and multimodal sources.

This Special Issue aims to showcase cutting-edge applications of deep learning in HAR, emphasizing real-world implementations and their impact on various domains. We invite high-quality original research articles and comprehensive reviews that explore how deep learning is being leveraged to improve activity recognition in practical settings. Contributions that address challenges related to data collection, deployment in real-world scenarios, and integration with emerging technologies such as the IoT, wearable devices, and smart cities are particularly encouraged.

Topics of interest include, but are not limited to, the following:

  • Industrial and workplace safety applications.
  • Healthcare applications of HAR, including rehabilitation monitoring.
  • Activity recognition in sports and fitness tracking.
  • HAR in human–computer interaction and augmented reality.
  • Real-time HAR applications in smart environments.
  • HAR in autonomous systems and robotics.
  • HAR for security and surveillance.
  • Smart home and smart city applications using HAR.
  • Sensor-based activity recognition using deep learning.
  • Computer-vision-based HAR.
  • Multimodal data fusion for HAR.

We welcome contributions that present novel applications, case studies, and implementations demonstrating the impact of deep learning on HAR across various domains.

Dr. Emmanuele Barberi
Dr. Emanuele Guardiani
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Electronics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2400 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • human activity recognition
  • deep learning applications
  • computer vision
  • healthcare and smart environments
  • sports analytics
  • wearable sensor technology
  • real-time HAR
  • security and surveillance
  • IoT and smart cities

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (4 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

17 pages, 2734 KB  
Article
Hand Gesture Recognition Based on Multi-Scale Attention Graph Convolutional Network
by Xiaowei Han, Tingshan Yan, Yunjing Lu, Ruize Liang, Honghui Zhang and Wei Chen
Electronics 2026, 15(16), 3649; https://doi.org/10.3390/electronics15163649 - 15 Aug 2026
Abstract
Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, and shallow spatio-temporal [...] Read more.
Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, and shallow spatio-temporal fusion. We propose a Multi-scale Attention Graph Convolutional Network (MA-GCN) that combines three components within one skeleton framework: a hybrid topology that augments physiological connections with semantic priors; a Gaussian Multi-Scale Channel Attention (GMCA) module for coordinate denoising and adaptive channel weighting; and a Local-Global Fusion Module (LGFM) that combines local convolutional features with channel-wise global attention. Ablation studies quantify the independent and joint contributions of these components. MA-GCN obtains Top-1 accuracies of 97.50%/95.95% on SHREC’17 Track and 94.29%/92.86% on DHG14/28 for the 14-/28-class settings. In a SHREC’17 Track-to-FPHA pre-train-then-fine-tune evaluation, it reaches 94.09% Top-1 accuracy, providing preliminary evidence that the proposed framework maintains effectiveness under cross-dataset transfer. Full article
(This article belongs to the Special Issue Deep Learning Applications on Human Activity Recognition)
Show Figures

Figure 1

19 pages, 5896 KB  
Article
Entropy-Gated Prediction Agreement for Two-View Video Action Recognition
by Young-Jin Park and Hui-Sup Cho
Electronics 2026, 15(13), 2844; https://doi.org/10.3390/electronics15132844 - 30 Jun 2026
Viewed by 294
Abstract
Human action recognition (HAR) often struggles to capture important temporal cues distributed across an entire video when relying solely on a single sampled clip. To overcome this limitation, this study proposes a framework that constructs two temporal views from the same video and [...] Read more.
Human action recognition (HAR) often struggles to capture important temporal cues distributed across an entire video when relying solely on a single sampled clip. To overcome this limitation, this study proposes a framework that constructs two temporal views from the same video and explicitly learns the prediction consistency between them. Specifically, the prediction-level agreement (AG) loss was introduced to align the class probability distributions of the two views. In addition, conditional gating was applied to adaptively control the contribution of AG loss according to the sample-wise prediction confidence, thereby reducing unstable alignment in temporally ambiguous or information-insufficient segments. The proposed framework was evaluated using both convolutional neural network (CNN)- and Transformer-based backbones on three representative action-recognition benchmark datasets, and it generally improved the performance over the single-view baseline across backbone–dataset combinations. Further empirical analyses, including training behavior, motion magnitude, temporal prediction stability, and qualitative case studies, were conducted to examine the effectiveness and behavior of the proposed two-view framework from multiple perspectives. Full article
(This article belongs to the Special Issue Deep Learning Applications on Human Activity Recognition)
Show Figures

Figure 1

22 pages, 1845 KB  
Article
Subset-Aware Dual-Teacher Knowledge Distillation with Hybrid Scoring for Human Activity Recognition
by Young-Jin Park and Hui-Sup Cho
Electronics 2025, 14(20), 4130; https://doi.org/10.3390/electronics14204130 - 21 Oct 2025
Cited by 4 | Viewed by 1268
Abstract
Human Activity Recognition (HAR) is a key technology with applications in healthcare, security, smart environments, and sports analytics. Despite advances in deep learning, challenges remain in building models that are both efficient and generalizable due to the large scale and variability of video [...] Read more.
Human Activity Recognition (HAR) is a key technology with applications in healthcare, security, smart environments, and sports analytics. Despite advances in deep learning, challenges remain in building models that are both efficient and generalizable due to the large scale and variability of video data. To address these issues, we propose a novel Dual-Teacher Knowledge Distillation (DTKD) framework tailored for HAR. The framework introduces three main contributions. First, we define static and dynamic activity classes in an objective and reproducible manner using optical-flow-based indicators, establishing a quantitative classification scheme based on motion characteristics. Second, we develop subset-specialized teacher models and design a hybrid scoring mechanism that combines teacher confidence with cross-entropy loss. This enables dynamic weighting of teacher contributions, allowing the student to adaptively balance knowledge transfer across heterogeneous activities. Third, we provide a comprehensive evaluation on the UCF101 and HMDB51 benchmarks. Experimental results show that DTKD consistently outperforms baseline models and achieves balanced improvements across both static and dynamic subsets. These findings validate the effectiveness of combining subset-aware teacher specialization with hybrid scoring. The proposed approach improves recognition accuracy and robustness, offering practical value for real-world HAR applications such as driver monitoring, healthcare, and surveillance. Full article
(This article belongs to the Special Issue Deep Learning Applications on Human Activity Recognition)
Show Figures

Figure 1

22 pages, 1404 KB  
Article
Deep-Learning-Based Human Activity Recognition: Eye-Tracking and Video Data for Mental Fatigue Assessment
by Batol Hamoud, Walaa Othman, Nikolay Shilov and Alexey Kashevnik
Electronics 2025, 14(19), 3789; https://doi.org/10.3390/electronics14193789 - 24 Sep 2025
Cited by 5 | Viewed by 2449
Abstract
This study addresses mental fatigue as a critical state arising from prolonged human activity and positions its detection as a valuable task within the broader scope of human activity recognition using deep learning. This work compares two models for mental fatigue detection: a [...] Read more.
This study addresses mental fatigue as a critical state arising from prolonged human activity and positions its detection as a valuable task within the broader scope of human activity recognition using deep learning. This work compares two models for mental fatigue detection: a model that uses eye-tracking data for fatigue predictions and a vision-based model that relies on vital signs and human activity indicators from facial video using deep learning and computer vision techniques. The eye-tracking model (based on TabNet architecture) achieved 82% accuracy, while the vision-based model (features were estimated using deep learning and computer vision) based on Random Forest architecture reached 78% accuracy. A correlation analysis revealed strong alignment between both models’ predictions, with 21 out of 27 sessions showing significant positive correlations on the collected dataset. Further comparison with an earlier-developed vision-based model trained on another dataset supported the generalizability of the vision-based model using physiological indicators extracted from a facial video for fatigue estimation. These findings highlight the potential of the vision-based model as a practical alternative to sensor and special-devices-based systems, especially in settings where non-intrusiveness and scalability are critical. Full article
(This article belongs to the Special Issue Deep Learning Applications on Human Activity Recognition)
Show Figures

Figure 1

Back to TopTop