Next Article in Journal
Graph Neural Networks for Software Vulnerability Mining: A Review
Next Article in Special Issue
Explainable Early Activity Recognition via Wearable Inertial Sensors for Human–Robot Collaboration in Agriculture
Previous Article in Journal
Attention Bidirectional Gated Fusion Based Multimodal Intent Recognition Under Uncertain Missing Modalities
Previous Article in Special Issue
Validating the Performance of VR Headset Eye-Tracking Using Gold Standard Eye-Tracker and MoCap System
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition

by
Kabul Khudaybergenov
1,* and
Avazjon Marakhimov
2
1
Department of Applied Informatics, Kimyo International University in Tashkent, Tashkent 100121, Uzbekistan
2
Department of Information Processing and Management Systems, Tashkent State Technical University, Tashkent 100174, Uzbekistan
*
Author to whom correspondence should be addressed.
Information 2026, 17(8), 729; https://doi.org/10.3390/info17080729
Submission received: 25 June 2026 / Revised: 12 July 2026 / Accepted: 21 July 2026 / Published: 28 July 2026

Abstract

Skeleton-based human action recognition is an important problem in applied vision systems, yet many existing approaches depend on a single skeleton descriptor or a single feature-learning mechanism. This restriction can weaken the representation of local body kinematics, long-range joint relations, and temporal dependencies within an action sequence. To address these limitations, this paper proposes HILF-GT (Hybrid Invariant Latent Feature Graph Transformer), a hybrid Graph Convolutional Network (GCN)-Transformer framework based on multiple spatio-temporal invariant latent features. The representation module constructs complementary structured tensors from skeleton graphs, inter-joint distances, adjacent-frame joint displacements, and inter-limb angles. Instead of transforming these descriptors into image-like maps for separate Convolutional Neural Network (CNN)-based classification, HILF-GT keeps their graph and temporal organization during learning. A local GCN branch models skeleton-aware kinematic patterns, whereas a graph-aware Transformer branch uses biased self-attention and cross-attention to capture dependencies among distant joints, frames, and latent-feature streams. A Perceiver-style latent bottleneck is further introduced to reduce the memory cost of global attention over frame-joint tokens. Experiments were conducted on four standard benchmark datasets, including NTU-RGB+D 60, NTU-RGB+D 120, NW-UCLA, and UTD-MHAD. The proposed method achieved 93.1% and 97.20% accuracy on the NTU-RGB+D 60 Cross-Subject and Cross-View protocols, 88.15% and 90.20% on the NTU-RGB+D 120 Cross-Subject and Cross-Setup protocols, 98.50% on NW-UCLA, and 97.50% on UTD-MHAD.
Keywords: human action; skeleton-based action recognition; invariant latent feature; spatial–temporal representation; deep learning; graph transformer; GCN-transformer; graph-aware attention human action; skeleton-based action recognition; invariant latent feature; spatial–temporal representation; deep learning; graph transformer; GCN-transformer; graph-aware attention

Share and Cite

MDPI and ACS Style

Khudaybergenov, K.; Marakhimov, A. Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition. Information 2026, 17, 729. https://doi.org/10.3390/info17080729

AMA Style

Khudaybergenov K, Marakhimov A. Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition. Information. 2026; 17(8):729. https://doi.org/10.3390/info17080729

Chicago/Turabian Style

Khudaybergenov, Kabul, and Avazjon Marakhimov. 2026. "Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition" Information 17, no. 8: 729. https://doi.org/10.3390/info17080729

APA Style

Khudaybergenov, K., & Marakhimov, A. (2026). Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition. Information, 17(8), 729. https://doi.org/10.3390/info17080729

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop