Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition
Abstract
1. Introduction
2. Related Work
2.1. Image-Based Skeleton Representations
2.2. GCN-Based Skeleton Action Recognition
2.3. Attention-Based Skeleton Action Recognition
2.4. Positioning of the Proposed Method with Respect to Prior Hybrid Architectures
3. Methodology
3.1. Action Representation
3.1.1. Skeleton Graph Matrix
3.1.2. Joints Distance Matrix
3.1.3. Adjacent Distance Matrix
3.1.4. Limbs Angle Matrix
3.1.5. Latent-Feature Tensors
3.2. GCN-Transformer Hybrid Action Classification
3.2.1. Local Graph Branch
3.2.2. Graph-Aware Attention Branch
3.2.3. Latent Cross-Attention Bottleneck
3.2.4. Classifier and Training Objective
| Algorithm 1 HILF-GT: latent-feature construction and hybrid classification |
| Require: Skeleton sequence , skeleton graph , trained parameters |
| Ensure: Predicted action class |
| Stage 1: Descriptor selection (offline, once per dataset, training split only) |
|
| Stage 2: Latent-feature construction (per sequence) |
| Stage 3: Hybrid forward pass |
4. Experiments
4.1. Datasets
4.2. Implementation Details
4.3. Comparison with the State-of-the-Art
4.4. Ablation Study
4.4.1. Effect of Each Component on the Prediction
4.4.2. Latent-Feature Tensors Versus Image-Based Transformer Classification
4.4.3. View Invariance
4.4.4. Velocity and Sequence Size Invariance
4.4.5. Hyperparameter Sensitivity
4.4.6. Robustness to Skeleton Degradations
4.5. Computation Complexity
4.6. Trade-Off Between Efficiency and Accuracy
4.7. Qualitative Evaluation
4.8. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Li, Z.; Li, F.; Hua, G. Dynamic graph attention network for skeleton-based action recognition. Appl. Sci. 2025, 15, 4929. [Google Scholar] [CrossRef]
- Zheng, N.; Du, Y.; Xia, H.; Liang, Z. Signal-SGN: A spiking graph convolutional network for skeleton action recognition via learning temporal-frequency dynamics. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025. [Google Scholar] [CrossRef]
- Chen, D.; Chen, M.; Wu, P.; Wu, M.; Zhang, T.; Li, C. Two-stream spatio-temporal GCN-transformer networks for skeleton-based action recognition. Sci. Rep. 2025, 15, 9482. [Google Scholar] [CrossRef] [PubMed]
- Cui, X.; Zhang, J.; He, Y.; Wang, Z.; Zhao, W. GCN-Former: A method for action recognition using graph convolutional networks and Transformer. Appl. Sci. 2025, 15, 4511. [Google Scholar] [CrossRef]
- Wang, J.; Sun, Y.; Tian, S. Deep learning for student behavior detection in smart classroom environments. Information 2025, 16, 949. [Google Scholar] [CrossRef]
- Wang, H.; Weng, W.; Wang, J.; Zhao, F.; Xie, G.-S.; Geng, X.; Wang, L. Foundation model for skeleton-based human action understanding. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 47–61. [Google Scholar] [CrossRef] [PubMed]
- Weng, W.; Wang, H.; Wang, J.; He, L.; Xie, G.-S. USDRL: Unified skeleton-based dense representation learning with multi-grained feature decorrelation. Proc. AAAI Conf. Artif. Intell. 2025, 39, 8332–8340. [Google Scholar] [CrossRef]
- Wang, H.; Ma, X.; Kuang, J.; Gui, J. Heterogeneous skeleton-based action representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025. [Google Scholar] [CrossRef]
- Li, X.; Lin, J.; Li, X.; Ye, Q. Dual-geometry prior frequency nonlinear graph convolutional network for human action recognition. In Proceedings of the ICASSP 2026 IEEE International Conference on Acoustics, Speech and Signal Processing, Barcelona, Spain, 10–15 May 2026. [Google Scholar] [CrossRef]
- Ye, Q.; Zhou, Y.; He, L.; Zhang, J.; Guo, X.; Zhang, J.; Tan, M.; Xie, W. SUGAR: Learning skeleton representation with visual-motion knowledge for action recognition. Proc. AAAI Conf. Artif. Intell. 2026, 40, 17930–17938. [Google Scholar] [CrossRef]
- Kabulov, A.; Babadzhanov, A.; Baizhumanov, A.; Saymanov, I.; Babadjanov, A. Algorithms for solving systems of Boolean equations based on the transformation of logical expressions. Mathematics 2026, 14, 594. [Google Scholar] [CrossRef]
- Abdusalomov, A.; Mukhiddinov, M.; Abdurashidova, K.; Kutlimuratov, A.; Marakhimov, A.; Seytnazarov, K.; Cho, Y.-I. SYMPHONIA–Enhanced multimodal emotion recognition with dual-branch dynamic attention and hierarchical adaptive fusion. Comput. Mater. Contin. 2026, 88, 077057. [Google Scholar] [CrossRef]
- Shao, D.; Shi, M.; Liu, L. FineTec: Fine-grained action recognition under temporal corruption via skeleton decomposition and sequence completion. Proc. AAAI Conf. Artif. Intell. 2026, 40, 8842–8850. [Google Scholar] [CrossRef]
- Patro, S.G.K.; Garg, S.; Riyazuddin, M.; Rachapudi, V.; Makharov, K.; Smerat, A.; Karimi, R. NeuroExplain-net for transparent multi-class lung cancer screening using computed tomography. Intell.-Based Med. 2026, 15, 100409. [Google Scholar] [CrossRef]
- Murugan, J.S.; Ramkumar, M.S.; Imambi, S.S.; Sivaramkrishnan, M.; Thangavelsamy, N.; Abass, K.S.; Rakhimova, M.; Khishe, M. Gated adaptive graph causal attention decision capsule transformers network with Meerkat optimization algorithm for EEG- and EMG-guided myoelectric control in upper limb rehabilitation. Intell.-Based Med. 2026, 14, 100382. [Google Scholar] [CrossRef]
- Liu, Y.; Yang, J.; Perera, M.; Ji, P.; Kim, D.; Xu, M.; Wang, T.; Anwar, S.; Gedeon, T.; Qin, Z. Representation-centric survey of supervised skeletal action recognition and the new benchmark. Pattern Recognit. 2026, 180, 114140. [Google Scholar] [CrossRef]
- Zhang, J.; Lin, L.; Yang, S.; Liu, J. Self-supervised skeleton-based action representation learning: A benchmark and beyond. Int. J. Comput. Vis. 2026, 134, 38. [Google Scholar] [CrossRef]
- Shuai, T.; Beng, S.; Khalid, F.B.; Rahmat, R.W.B.O.K. Advances in facial micro-expression detection and recognition: A comprehensive review. Information 2025, 16, 876. [Google Scholar] [CrossRef]
- Liu, Y.; Shi, T.; Zhai, M.; Liu, J. Frequency decoupled masked auto-encoder for self-supervised skeleton-based action recognition. IEEE Signal Process. Lett. 2025, 32, 546–550. [Google Scholar] [CrossRef]
- Wei, J.; Qin, L.; Yu, B.; Zou, T.; Yan, C.; Xiao, D.; Yu, Y.; Yang, L. VA-AR: Learning velocity-aware action representations with mixture of window attention. Proc. AAAI Conf. Artif. Intell. 2025, 39, 8286–8294. [Google Scholar] [CrossRef]
- Caetano, C.; Brémond, F.; Schwartz, W.R. Skeleton image representation for 3D action recognition based on tree structure and reference joints. In Proceedings of the 2019 32nd SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), Rio de Janeiro, Brazil, 28–31 October 2019; pp. 16–23. [Google Scholar]
- Zhou, Y.; Xu, T.; Wu, C.; Wu, X.; Kittler, J. Adaptive hyper-graph convolution network for skeleton-based human action recognition with virtual connections. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, HI, USA, 19–23 October 2025. [Google Scholar] [CrossRef]
- Chen, Z.; Li, S.; Yang, B.; Li, Q.; Liu, H. Multi-scale spatial temporal graph convolutional network for skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 2–9 February 2021; pp. 1113–1122. [Google Scholar]
- Chen, Y.; Zhang, Z.; Yuan, C.; Li, B.; Deng, Y.; Hu, W. Channel-wise topology refinement graph convolution for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 10–17 October 2021; pp. 13359–13368. [Google Scholar]
- Liu, Y.; Liu, R.; Hu, Y.; Wu, M.; Xin, W.; Miao, Q.; Wu, S.; Li, L. A systematic review of skeleton-based action recognition: Methods, challenges, and future directions. IEEE Trans. Neural Netw. Learn. Syst. 2026, 37, 2046–2065. [Google Scholar] [CrossRef] [PubMed]
- Zhu, A.; Zhu, J.; Bailey, J.; Gong, M.; Ke, Q. Semantic-guided cross-modal prompt learning for skeleton-based zero-shot action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025. [Google Scholar] [CrossRef]
- Bavil, A.F.; Damirchi, H.; Taghirad, H.D. Action capsules: Human skeleton action recognition. Comput. Vis. Image Underst. 2023, 233, 103722. [Google Scholar] [CrossRef]
- Pang, Y.; Ke, Q.; Rahmani, H.; Bailey, J.; Liu, J. IGFormer: Interaction graph transformer for skeleton-based human interaction recognition. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 605–622. [Google Scholar]
- Ma, N.; Sun, B.; Han, Y.; Xu, G. Kinematic enhanced hypergraph convolutional network for skeleton-based human action recognition with LLM training guides. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025. [Google Scholar] [CrossRef]
- Ma, N.; Xu, G.; Han, Y.; Sun, B. THTFormer: Topology-adaptive hypergraph Transformer network for skeleton-based action recognition. Pattern Recognit. 2026, 171, 112125. [Google Scholar] [CrossRef]
- Korban, M.; Youngs, P.; Acton, S.T. RL-GTN: A reinforced divergence-optimized graph Transformer network for skeleton-based action recognition. Pattern Recognit. 2026, 172, 112681. [Google Scholar] [CrossRef]
- Plizzari, C.; Cannici, M.; Matteucci, M. Skeleton-based action recognition via spatial and temporal transformer networks. Comput. Vis. Image Underst. 2021, 208, 103219. [Google Scholar] [CrossRef]
- Filali, H.; Boulealam, C.; El Fazazy, K.; Mahraz, A.M.; Tairi, H.; Riffi, J. Meaningful multimodal emotion recognition based on capsule graph transformer architecture. Information 2025, 16, 40. [Google Scholar] [CrossRef]
- Yang, Y.; Zhao, J.; Kuang, Z.; Ta, N.; Hong, J. Kinematic priors benefit skeleton-based action recognition. In Proceedings of the ICASSP 2026 IEEE International Conference on Acoustics, Speech and Signal Processing, Barcelona, Spain, 10–15 May 2026. [Google Scholar] [CrossRef]
- Li, X.; Geng, Q.; Huang, Q.; Li, X.; Tang, J.; Ye, Q. Spatial-temporal self-compensating graph convolutional network for skeleton-based action recognition under data constraints. IEEE Trans. Image Process. 2026, 35, 5818–5833. [Google Scholar] [CrossRef] [PubMed]
- Wang, P.; Li, Z.; Hou, Y.; Li, W. Action recognition based on joint trajectory maps using convolutional neural networks. In Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands, 15–19 October 2016; pp. 102–106. [Google Scholar]
- Yang, Y.; Chen, H.; Liu, Z.; Hu, S.; Jiao, Y. Dual space representation learning for skeleton-based action recognition. IEEE Signal Process. Lett. 2025, 32, 2104–2108. [Google Scholar] [CrossRef]
- Liu, R.; Chen, Y.; Gai, F.; Liu, Y.; Miao, Q.; Wu, S. Local and global spatial-temporal Transformer for skeleton-based action recognition. Neurocomputing 2025, 636, 129820. [Google Scholar] [CrossRef]
- Alavigharahbagh, A.; Hajihashemi, V.; Machado, J.J.M.; Tavares, J.M.R.S. Deep learning approach for human action recognition using a time saliency map based on motion features considering camera movement and shot in video image sequences. Information 2023, 14, 616. [Google Scholar] [CrossRef]
- Yan, S.; Xiong, Y.; Lin, D. Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018. [Google Scholar]
- Shi, L.; Zhang, Y.; Cheng, J.; Lu, H. Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 12026–12035. [Google Scholar]
- Kilis, N.; Papaioannidis, C.; Mademlis, I.; Pitas, I. An efficient framework for human action recognition based on graph convolutional networks. In Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP 2022), Bordeaux, France, 16–19 October 2022; pp. 1441–1445. [Google Scholar]
- Zhou, H.; Liu, Q.; Wang, Y. Learning discriminative representations for skeleton based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 10608–10617. [Google Scholar]
- Si, C.; Jing, Y.; Wang, W.; Wang, L.; Tan, T. Skeleton-based action recognition with spatial reasoning and temporal stack learning. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 103–118. [Google Scholar]
- Si, C.; Chen, W.; Wang, W.; Wang, L.; Tan, T. An attention enhanced graph convolutional LSTM network for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 1227–1236. [Google Scholar]
- Nikpour, B.; Armanfard, N. Spatial hard attention modeling via deep reinforcement learning for skeleton-based human activity recognition. IEEE Trans. Syst. Man Cybern. Syst. 2023, 53, 4291–4301. [Google Scholar] [CrossRef]
- Shahroudy, A.; Liu, J.; Ng, T.T.; Wang, G. NTU RGB+D: A large scale dataset for 3D human activity analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 1010–1019. [Google Scholar]
- Liu, J.; Shahroudy, A.; Perez, M.; Wang, G.; Duan, L.Y.; Kot, A.C. NTU RGB+D 120: A large-scale benchmark for 3D human activity understanding. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 42, 2684–2701. [Google Scholar] [PubMed]
- Wang, J.; Nie, X.; Xia, Y.; Wu, Y.; Zhu, S.C. Cross-view action modeling, learning and recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 2649–2656. [Google Scholar]
- Chen, C.; Jafari, R.; Kehtarnavaz, N. UTD-MHAD: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor. In Proceedings of the 2015 IEEE International Conference on Image Processing (ICIP), Quebec City, QC, Canada, 27–30 September 2015; pp. 168–172. [Google Scholar]
- Li, C.; Hou, Y.; Wang, P.; Li, W. Joint distance maps based action recognition with convolutional neural networks. IEEE Signal Process. Lett. 2017, 24, 624–628. [Google Scholar] [CrossRef]
- Song, Y.F.; Zhang, Z.; Shan, C.; Wang, L. Constructing stronger and faster baselines for skeleton-based action recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 1474–1488. [Google Scholar] [CrossRef]
- Li, M.; Chen, S.; Chen, X.; Zhang, Y.; Wang, Y.; Tian, Q. Actional-structural graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 3595–3603. [Google Scholar]
- Song, Y.F.; Zhang, Z.; Shan, C.; Wang, L. Richly activated graph convolutional network for robust skeleton-based action recognition. IEEE Trans. Circuits Syst. Video Technol. 2020, 31, 1915–1925. [Google Scholar] [CrossRef]
- Cheng, K.; Zhang, Y.; He, X.; Chen, W.; Cheng, J.; Lu, H. Skeleton-based action recognition with shift graph convolutional network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 183–192. [Google Scholar]
- Ye, F.; Pu, S.; Zhong, Q.; Li, C.; Xie, D.; Tang, H. Dynamic GCN: Context-enriched topology learning for skeleton-based action recognition. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020; pp. 55–63. [Google Scholar]
- Xu, K.; Ye, F.; Zhong, Q.; Xie, D. Topology-aware convolutional neural network for efficient skeleton-based action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 22 February–1 March 2022; pp. 2866–2874. [Google Scholar]
- Liu, Y.; Zhang, H.; Li, Y.; He, K.; Xu, D. Skeleton-based human action recognition via large-kernel attention graph convolutional network. IEEE Trans. Vis. Comput. Graph. 2023, 29, 2575–2585. [Google Scholar] [CrossRef] [PubMed]
- Liu, F.; Wang, C.; Tian, Z.; Du, S.; Zeng, W. Advancing skeleton-based human behavior recognition: Multi-stream fusion spatiotemporal graph convolutional networks. Complex Intell. Syst. 2025, 11, 94. [Google Scholar]
- Guo, T.; Liu, M.; Liu, H.; Wang, G.; Li, W. Improving self-supervised action recognition from extremely augmented skeleton sequences. Pattern Recognit. 2024, 150, 110333. [Google Scholar] [CrossRef]
- Li, M.; Chen, K.; Bai, Y.; Pei, J. Skeleton action recognition via graph convolutional network with self-attention module. Electron. Res. Arch. 2024, 32, 2848–2864. [Google Scholar] [CrossRef]
- Ran, R.; Yang, W. FD-GCN: Feedback directed graph convolutional network for skeleton-based action recognition. Graph. Models 2025, 142, 101306. [Google Scholar] [CrossRef]
- Tu, Z.; Zhang, Z.; Gong, J.; Yuan, J.; Du, B. Informative Sample Selection Model for Skeleton-Based Action Recognition With Limited Training Samples. IEEE Trans. Image Process. 2025, 34, 7335–7346. [Google Scholar] [CrossRef] [PubMed]
- Lasri, K.; El Fazazy, K.; Mohamed Mahraz, A.; Tairi, H.; Riffi, J. DPCA-GCN: Dual-Path Cross-Attention Graph Convolutional Networks for Skeleton-Based Action Recognition. Computation 2025, 13, 293. [Google Scholar] [CrossRef]
- Chi, H.G.; Ha, M.H.; Chi, S.; Lee, S.W.; Huang, Q.; Ramani, K. InfoGCN: Representation learning for human skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 20186–20196. [Google Scholar]
- Zhang, P.; Lan, C.; Xing, J.; Zeng, W.; Xue, J.; Zheng, N. View adaptive neural networks for high performance skeleton-based human action recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 1963–1978. [Google Scholar] [CrossRef] [PubMed]
- Liu, D.; Li, X.; Cai, Z.; Chen, P. TSGCNeXt: Dynamic-Static Multi-graph Convolution for efficient skeleton-based action recognition. Expert Syst. Appl. 2025, 276, 127081. [Google Scholar] [CrossRef]
- Han, X.-W.; Chen, X.-Y.; Cui, Y.; Guo, Q.-Y.; Hu, W. Adaptive Channel-Enhanced Graph Convolution for Skeleton-Based Human Action Recognition. Appl. Sci. 2024, 14, 8185. [Google Scholar] [CrossRef]
- Hou, Y.; Li, Z.; Wang, P.; Li, W. Skeleton optical spectra-based action recognition using convolutional neural networks. IEEE Trans. Circuits Syst. Video Technol. 2016, 28, 807–811. [Google Scholar] [CrossRef]
- Tu, Z.; Talebi, H.; Zhang, H.; Yang, F.; Milanfar, P.; Bovik, A.; Li, Y. MaxViT: Multi-axis vision Transformer. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 459–479. [Google Scholar]
- Kumar, R.; Corvisieri, G.; Fici, T.F.; Hussain, S.I.; Tegolo, D.; Valenti, C. Transfer learning for facial expression recognition. Information 2025, 16, 320. [Google Scholar] [CrossRef]
- Shi, L.; Zhang, Y.; Cheng, J.; Lu, H. Decoupled spatial-temporal attention network for skeleton-based action-gesture recognition. In Proceedings of the Asian Conference on Computer Vision, Kyoto, Japan, 30 November–4 December 2020. [Google Scholar]










| Method | Multi-Stream with Stream-Specific Graphs | Feature-Derived Attention Bias | Parallel GCN- Transformer Branches | Latent Bottleneck Over Cross-Stream Tokens |
|---|---|---|---|---|
| ST-TR [32] | – | – | – | – |
| IGFormer [28] | – | partial (topology) | – | – |
| Kinematic-HGCN [29] | – | partial (hyperedges) | – | – |
| THTFormer [30] | – | partial (hyperedges) | – | – |
| RL-GTN [31] | – | partial (topology) | – | – |
| 2s-AGCN/MST-GCN [23,41] | partial (shared graph) | – | – | – |
| Kinematic-prior models [34,35] | – | partial (fixed priors) | – | partial (single stream) |
| HILF-GT (proposed) | √ | √ | √ | √ |
| Module | Setting | Value |
|---|---|---|
| Stream GCN branches (×4) | GCN blocks L | 3 |
| Channel widths | 64/128/256 | |
| Activation /normalization | ReLU/BatchNorm | |
| Relation types | 3 (GLFs); 2 (JDLFs, ADLFs, LALFs) | |
| Graph-aware Transformer branch | Encoder layers | 4 |
| Attention heads | 8 | |
| Token dimension d | 256 | |
| Feed-forward dimension | 1024 | |
| Latent bottleneck | Latent tokens R | 64 |
| Bottleneck layers | 2 | |
| Attention heads | 8 | |
| Training | Optimizer | Adam |
| Initial learning rate | (halved after epoch 5) | |
| Batch size | 32 | |
| Epochs | 65 | |
| Dropout/weight decay | 0.1/ | |
| Sequence length T | 300 (NTU); 52 (NW-UCLA, UTD-MHAD) |
| Method | C-Sub | C-View |
|---|---|---|
| ST-GCN [40] | 81.51 | 88.31 |
| SR-TSL [44] | 84.82 | 92.42 |
| AS-GCN [53] | 86.84 | 94.23 |
| 2s-AGCN [41] | 88.58 | 95.18 |
| AGC-LSTM [45] | 89.23 | 95.10 |
| RA-GCN [54] | 87.33 | 93.67 |
| 4s-Shift-GCN [55] | 90.77 | 96.15 |
| Dynamic-GCN [56] | 91.52 | 96.10 |
| MST-GCN [23] | 91.75 | 96.63 |
| ST-TR [32] | 90.23 | 96.35 |
| Ta-CNN [57] | 90.37 | 95.24 |
| EfficientGCN [52] | 92.17 | 96.11 |
| ST-SLKA [58] | 90.17 | 96.32 |
| Action Capsules [27] | 90.10 | 96.34 |
| SHARL [46] | 91.14 | 96.24 |
| SCA-GCN [59] | 89.8 | 96.0 |
| AimCLR++ [60] | 80.9 | 85.4 |
| SA-GCN [61] | 91.5 | 84.7 |
| FD-GCN [62] | 90.6 | 96.30 |
| ISSM [63] | 81.1 | 86.7 |
| DPCA-GCN [64] | 88.72 | 94.31 |
| HILF-GT (Proposed) | 93.1 | 97.20 |
| Method | C-Sub | C-Set |
|---|---|---|
| ST-GCN [40] | 70.17 | 73.2 |
| AS-GCN [53] | 77.19 | 78.5 |
| 2s-AGCN [41] | 82.25 | 84.2 |
| RA-GCN [54] | 81.21 | 82.7 |
| 4s-Shift-GCN [55] | 85.39 | 87.6 |
| Dynamic-GCN [56] | 87.23 | 88.6 |
| MST-GCN [23] | 87.75 | 88.50 |
| ST-TR [32] | 85.71 | 87.10 |
| Ta-CNN [57] | 85.47 | 87.31 |
| IGFormer [28] | 85.44 | 86.52 |
| EfficientGCN [52] | 88.27 | 88.19 |
| ST-SLKA [58] | 86.13 | 87.28 |
| AimCLR++ [60] | 70.1 | 71.2 |
| SA-GCN [61] | 79.2 | 78.5 |
| FD-GCN [62] | 85.5 | 87.90 |
| DPCA-GCN [64] | 82.85 | 83.65 |
| HILF-GT (Proposed) | 88.15 | 90.20 |
| Method | Accuracy |
|---|---|
| AGC-LSTM [45] | 93.32 |
| VA-CNN [66] | 90.27 |
| 4s-shift-GCN [55] | 94.26 |
| Ta-CNN [57] | 96.21 |
| CTR-GCN [24] | 96.45 |
| InfoGCN [65] | 96.66 |
| FRHead [43] | 96.78 |
| Action Capsules [27] | 97.13 |
| FD-GCN [62] | 95.10 |
| TSGCNeXt [67] | 96.55 |
| ACE-GCN [68] | 96.6 |
| ISSM [63] | 87.9 |
| HILF-GT (Proposed) | 98.50 |
| Method | Accuracy |
|---|---|
| Kinect [50] | 66.11 |
| Inertial [50] | 67.22 |
| Kinect&Inertial [50] | 79.41 |
| JTM [36] | 85.58 |
| Optical Spectra [69] | 86.59 |
| JDM [51] | 88.41 |
| HILF-GT (Proposed) | 97.50 |
| Benchmark | Methods Exceeded | vs. Median | vs. Best |
|---|---|---|---|
| NTU-RGB+D 60 (C-Sub) | 21/21 | ||
| NTU-RGB+D 60 (C-View) | 21/21 | ||
| NTU-RGB+D 120 (C-Sub) | 15/16 | ||
| NTU-RGB+D 120 (C-Set) | 16/16 | ||
| NW-UCLA | 12/12 | ||
| UTD-MHAD | 6/6 |
| Model | NTU-RGB+D 60 (C-Sub) | NTU-RGB+D 60 (C-View) | NTU-RGB+D 120 (C-Sub) | NTU-RGB+D 120 (C-Set) | NW-UCLA |
|---|---|---|---|---|---|
| GLFs | 80.23 | 86.42 | 72.30 | 74.38 | 83.28 |
| JDLFs | 77.28 | 86.24 | 68.20 | 70.38 | 91.60 |
| ADLFs | 73.16 | 80.51 | 62.72 | 65.62 | 88.71 |
| LALFs | 73.41 | 77.19 | 61.24 | 63.27 | 87.27 |
| FDLFs | 81.51 | 85.77 | 72.51 | 74.64 | 92.81 |
| tensor-fusion | 87.22 | 91.22 | 76.57 | 90.32 | 95.50 |
| full hybrid | 93.10 | 97.20 | 88.15 | 90.20 | 98.50 |
| GLFs | JDLFs | ADLFs | LALFs | FDLFs | Hybrid | Accuracy |
|---|---|---|---|---|---|---|
| √ | √ | √ | √ | 88.21 | ||
| √ | √ | √ | √ | 89.64 | ||
| √ | √ | √ | √ | 88.16 | ||
| √ | √ | √ | √ | √ | 90.46 | |
| √ | √ | √ | √ | √ | 91.17 | |
| √ | √ | √ | √ | √ | 91.22 | |
| √ | √ | √ | √ | √ | 91.27 | |
| √ | √ | √ | √ | √ | 91.28 | |
| √ | √ | √ | √ | √ | 90.14 | |
| √ | √ | √ | √ | √ | √ | 93.1 |
| Latent Feature | Ours NTU60 | MaxVIT NTU60 | Ours NW-UCLA | MaxVIT NW-UCLA |
|---|---|---|---|---|
| GLFs | 80.23 | 74.34 | 83.28 | 73.61 |
| JDLFs | 77.28 | 60.31 | 91.60 | 87.23 |
| ADLFs | 73.16 | 68.62 | 88.71 | 78.73 |
| LALFs | 73.41 | 64.16 | 87.27 | 82.51 |
| FDLFs | 81.51 | 71.65 | 92.81 | 88.35 |
| Hyperparameter | Value | Accuracy |
|---|---|---|
| Latent tokens R | 32 | 97.20 |
| 64 * | 98.50 | |
| 128 | 98.05 | |
| Bottleneck layers | 1 | 98.25 |
| 2 * | 98.50 | |
| 3 | 98.31 | |
| Learning rate | 98.02 | |
| * | 98.50 | |
| 98.38 | ||
| Selected angles U | 10 | 97.85 |
| 14 * | 98.50 | |
| 18 | 98.42 |
| Perturbation | Severity | Accuracy | vs. Clean |
|---|---|---|---|
| Missing joints | 10% | 91.85 | |
| 20% | 89.94 | ||
| 30% | 87.41 | ||
| Gaussian noise | 92.71 | ||
| 91.66 | |||
| 90.38 | |||
| Part occlusion [54] | one arm | 90.52 | |
| both legs | 89.87 | ||
| trunk | 86.73 | ||
| Temporal occlusion | 10% frames | 92.34 | |
| 20% frames | 91.08 | ||
| 30% frames | 89.25 | ||
| Viewpoint rotation | 92.94 | ||
| 92.61 | |||
| 92.18 |
| Model | Parameters (M) | GFLOPs | Inference Time |
|---|---|---|---|
| ST-GCN [40] | 3.1 | 18 | 47 |
| RA-GCN [54] | 5.4 | 34 | 20 |
| 2s-AGCN [41] | 5.7 | 39 | 21 |
| 4s-ShiftGCN [55] | 3.0 | 16 | – |
| CTR-GCN [24] | 0.8 | 8 | – |
| DSTA-Net [72] | 4.2 | 68 | – |
| ST-TR [32] | 15.7 | 252 | – |
| HILF-GT (latent-feature stream) | 0.9 | 9 | 101 |
| HILF-GT (tensor fusion) | 2.8 | 28 | 33 |
| HILF-GT (full hybrid) | 8.2 | 82 | 19 |
| Component | Params (M) | GFLOPs | Peak Memory (GB) | Latency (ms) | Throughput (seq/s) |
|---|---|---|---|---|---|
| GCN branch: GLFs | 0.9 | 9 | 0.42 | 1.6 | 625 |
| GCN branch: JDLFs | 0.7 | 6 | 0.31 | 1.2 | 833 |
| GCN branch: ADLFs | 0.6 | 5 | 0.28 | 1.1 | 909 |
| GCN branch: LALFs | 0.7 | 6 | 0.30 | 1.2 | 833 |
| Graph-aware Transformer branch | 3.8 | 38 | 1.65 | 3.9 | 256 |
| Latent cross-attention bottleneck | 1.5 | 18 | 0.71 | 1.4 | 714 |
| Full hybrid (network only) | 8.2 | 82 | 3.67 | 10.4 | 96 |
| Full hybrid (incl. latent-feature construction) | 8.2 | 82 | 3.67 | 53.0 | 19 |
| Configuration | Latent-Feature Gen. | Backbone | Hybrid Head | Total | Accuracy |
|---|---|---|---|---|---|
| full hybrid | 93.10 | ||||
| tensor fusion | – | 87.22 | |||
| two latent features (FDLFs + GLFs) | 88.50 | ||||
| single latent-feature stream (FDLFs) | – | 82.22 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Khudaybergenov, K.; Marakhimov, A. Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition. Information 2026, 17, 729. https://doi.org/10.3390/info17080729
Khudaybergenov K, Marakhimov A. Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition. Information. 2026; 17(8):729. https://doi.org/10.3390/info17080729
Chicago/Turabian StyleKhudaybergenov, Kabul, and Avazjon Marakhimov. 2026. "Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition" Information 17, no. 8: 729. https://doi.org/10.3390/info17080729
APA StyleKhudaybergenov, K., & Marakhimov, A. (2026). Hybrid Invariant Latent Feature Graph Transformer for Skeleton-Based Human Action Recognition. Information, 17(8), 729. https://doi.org/10.3390/info17080729

