Gaussian Topology Refinement and Multi-Scale Shift Graph Convolution for Efficient Real-Time Sports Action Recognition
Abstract
1. Introduction
- We propose EMS-GCN, a highly efficient skeleton-based action recognition network tailored for edge sensor applications. This model reconstructs the underlying spatiotemporal interaction mechanisms. As illustrated in Figure 1, EMS-GCN achieves high-precision real-time inference with only 0.56 M parameters.
- We design a Gaussian topology refinement module. Distance-based symmetry constraints are utilized. Topological structure uncertainty induced by acquisition jitter and drift in sensor data is effectively resolved. Geometric stability of the skeleton graph is restored.
- We propose a temporal modeling scheme integrating multi-scale shift and linear attention. By employing a feature concatenation strategy, this approach eliminates the temporal aliasing risks inherent in traditional shift operations. Furthermore, it achieves precise capture of global temporal dependencies with linear computational complexity, striking an optimal balance between long-range context modeling and real-time response requirements.
2. Related Work
2.1. Skeleton-Based Action Recognition with GCNs
2.2. Lightweight Graph Convolutional Networks
3. Proposed Method
3.1. Overall Architecture
- Spatial Dimension: We introduce the Gaussian Topology Refinement Module (GTRM). By leveraging inter-node statistical distribution properties to constrain dynamic topology learning, this module addresses topological ambiguity caused by positional drift in sensor data.
- Temporal Dimension. We propose the Multi-scale Shift Linear Attention Module (MS-LTA). This component captures local temporal features via parameter-free shift operations and models global dependencies through a linear-complexity attention mechanism, effectively replacing computationally intensive standard temporal convolutions.
| Algorithm 1: Forward Procedure of the Gaussian-Refined CTR-GCN with Multi-Scale Shift and Linear Temporal Attention |
| Input: Skeleton tensor , predefined graph partitions |
| Output: Class logits |
| 1 Reorder to merge person and joint dimensions, then apply batch normalization over ; |
| 2 Reshape normalized features into ; |
| 3 for to do |
| 4 | // Spatial graph reasoning with multi-subset CTRGC |
| 5 | ; |
| 6 | for to do |
| 7 | | Project by convolutions to obtain relation features and node features; |
| 8 | | Compute dynamic topology ; |
| 9 | | Compute Gaussian prior from temporally averaged features: ; |
| 10 | | Refine topology: ; |
| 11 | | Aggregate subset response: ; |
| 12 | end |
| 13 | Fuse subset responses with residual graph mapping and activation: ; |
| 14 | // Temporal modeling block |
| 15 | ; |
| 16 | ; |
| 17 | ; |
| 18 | Apply dropout, channel-spatial attention, temporal downsampling (if stride ), and residual fusion: ; |
| 19 end |
| 20 Apply global average pooling over temporal and joint axes; |
| 21 Average features across the person dimension , then apply dropout; |
| 22 Return logits ; |
3.2. Gaussian-Refined Channel-Wise Topology Modeling
3.3. Multi-Scale Shift Linear Temporal Attention

3.4. Loss Function
4. Experiments
4.1. Experimental Setup
4.2. Performance Comparison
4.3. Ablation Study
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Xin, W.; Liu, R.; Liu, Y.; Chen, Y.; Yu, W.; Miao, Q. Transformer for skeleton-based action recognition: A review of recent advances. Neurocomputing 2023, 537, 164–186. [Google Scholar] [CrossRef]
- Yue, R.; Tian, Z.; Du, S. Action recognition based on RGB and skeleton data sets: A survey. Neurocomputing 2022, 512, 287–306. [Google Scholar] [CrossRef]
- Kong, Y.; Fu, Y. Human action recognition and prediction: A survey. Int. J. Comput. Vis. 2022, 130, 1366–1401. [Google Scholar] [CrossRef]
- Zhang, J.; Lin, L.; Yang, S.; Liu, J. Self-Supervised Skeleton-Based Action Representation Learning: A Benchmark and Beyond. arXiv 2024, arXiv:2406.02978. [Google Scholar] [CrossRef]
- Do, J.; Kim, M. Skateformer: Skeletal-temporal transformer for human action recognition. In Proceedings of the European Conference on Computer Vision (ECCV); Springer Nature: Cham, Switzerland, 2024; pp. 401–420. [Google Scholar]
- Wu, W.; Zheng, C.; Yang, Z.; Chen, C.; Das, S.; Lu, A. Frequency guidance matters: Skeletal action recognition by frequency-aware mixed transformer. In Proceedings of the 32nd ACM International Conference on Multimedia (ACM MM), Melbourne, Australia, 1–28 November 2024; pp. 4660–4669. [Google Scholar]
- Cheng, K.; Zhang, Y.; He, X.; Chen, W.; Cheng, J.; Lu, H. Skeleton-based action recognition with shift graph convolutional network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 183–192. [Google Scholar]
- Zhou, Y.; Yan, X.; Cheng, Z.Q.; Yan, Y.; Dai, Q.; Hua, X.S. BlockGCN: Redefine topology awareness for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 2049–2058. [Google Scholar]
- Zhou, Y.; Cheng, Z.Q.; He, J.Y.; Luo, B.; Geng, Y.; Xie, X. Overcoming topology agnosticism: Enhancing skeleton-based action recognition through redefined skeletal topology awareness. arXiv 2023, arXiv:2305.11468. [Google Scholar]
- Myung, W.S.; Su, N.; Xue, J.H.; Wang, G. DeGCN: Deformable graph convolutional networks for skeleton-based action recognition. IEEE Trans. Image Process. 2024, 33, 2477–2490. [Google Scholar] [CrossRef] [PubMed]
- Jiang, Y.; Deng, H. Lighter and faster: A multi-scale adaptive graph convolutional network for skeleton-based action recognition. Eng. Appl. Artif. Intell. 2024, 132, 107957. [Google Scholar] [CrossRef]
- Noor, N.; Jametoni, F.; Kim, J.; Hong, H.; Park, I.K. Efficient skeleton-based action recognition for real-time embedded systems. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 5889–5897. [Google Scholar]
- Kang, M.S.; Kang, D.; Kim, H.S. Efficient skeleton-based action recognition via joint-mapping strategies. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2–7 January 2023; pp. 3403–3412. [Google Scholar]
- Xia, Y.; Gao, Q.; Wu, W.; Cao, Y. Skeleton-based action recognition based on multidimensional adaptive dynamic temporal graph convolutional network. Eng. Appl. Artif. Intell. 2024, 127, 107210. [Google Scholar] [CrossRef]
- Xia, Y.; Gao, Q.; Wu, W.; Cao, Y. Chase: Learning convex hull adaptive shift for skeleton-based multi-entity action recognition. Adv. Neural Inf. Process. Syst. 2024, 37, 9388–9420. [Google Scholar]
- Li, X.; Kang, J.; Yang, Y.; Zhao, F. A lightweight attentional shift graph convolutional network for skeleton-based action recognition. Int. J. Comput. Commun. Control 2023, 18. [Google Scholar] [CrossRef]
- Liu, Z.; Xia, H.; Guo, T.; Sun, L.; Shao, M.; Xia, S.Y. Cross-Block Fine-Grained Semantic Cascade for Skeleton-Based Sports Action Recognition. In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition (FG), Istanbul, Turkey, 27–31 May 2024; pp. 1–10. [Google Scholar]
- Yan, S.; Xiong, Y.; Lin, D. Spatial temporal graph convolutional networks for skeleton-based action recognition. Proc. AAAI Conf. Artif. Intell. 2018, 32, 7444–7452. [Google Scholar] [CrossRef]
- Shi, L.; Zhang, Y.; Cheng, J.; Lu, H. Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 12026–12035. [Google Scholar]
- Shi, L.; Zhang, Y.; Cheng, J.; Lu, H. Skeleton-based action recognition with multi-stream adaptive graph convolutional networks. IEEE Trans. Image Process. 2020, 29, 9532–9545. [Google Scholar] [CrossRef] [PubMed]
- Li, M.; Chen, S.; Chen, X.; Zhang, Y.; Wang, Y.; Tian, Q. Actional-structural graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 3595–3603. [Google Scholar]
- Shi, L.; Zhang, Y.; Cheng, J.; Lu, H. Skeleton-based action recognition with directed graph neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 7912–7921. [Google Scholar]
- Korban, M.; Li, X. DDGCN: A dynamic directed graph convolutional network for action recognition. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2020; pp. 761–776. [Google Scholar]
- Li, C.; Huang, Q.; Mao, Y. DD-GCN: Directed diffusion graph convolutional network for skeleton-based human action recognition. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), Brisbane, Australia, 10–14 June 2023; pp. 786–791. [Google Scholar]
- Duan, H.; Wang, J.; Chen, K.; Lin, D. DG-STGCN: Dynamic spatial-temporal modeling for skeleton-based action recognition. arXiv 2022, arXiv:2210.05895. [Google Scholar]
- Zhang, X.; Xu, C.; Tao, D. Context aware graph convolution for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 14333–14342. [Google Scholar]
- Liu, Z.; Zhang, H.; Chen, Z.; Wang, Z.; Ouyang, W. Disentangling and unifying graph convolutions for skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 143–152. [Google Scholar]
- Chen, Y.; Zhang, Z.; Yuan, C.; Li, B.; Deng, Y.; Hu, W. Channel-wise topology refinement graph convolution for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 13359–13368. [Google Scholar]
- Peng, W.; Hong, X.; Chen, H.; Zhao, G. Learning graph convolutional network for skeleton-based human action recognition by neural searching. Proc. AAAI Conf. Artif. Intell. 2020, 34, 2669–2676. [Google Scholar] [CrossRef]
- Chi, H.; Ha, M.H.; Chi, S.; Lee, S.W.; Huang, Q.; Ramani, K. InfoGCN: Representation learning for human skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 20186–20196. [Google Scholar]
- Lee, J.; Lee, M.; Lee, D.; Lee, S. Hierarchically decomposed graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Vancouver, BC, Canada, 17–24 June 2023; pp. 10444–10453. [Google Scholar]
- Kose, H.T.; Nunez-Yanez, J.; Piechocki, R.; Pope, J. A survey of computationally efficient graph neural networks for reconfigurable systems. Information 2024, 15, 377. [Google Scholar] [CrossRef]
- Liu, J.; Chen, S.; Shen, L. A comprehensive survey on graph neural network accelerators. Front. Comput. Sci. 2025, 19, 192104. [Google Scholar] [CrossRef]
- Corradini, F.; Gerosa, F.; Gori, M.; Lucheroni, C.; Piangerelli, M.; Zannotti, M. A systematic literature review of spatio-temporal graph neural network models for time series forecasting and classification. Neural Netw. 2025, 195, 108269. [Google Scholar] [CrossRef]
- Roy, A.; Tiwari, A.; Saurav, S.; Singh, S. Enhancing skeleton-based action recognition using a knowledge-driven shift graph convolutional network. Comput. Electr. Eng. 2024, 120, 109633. [Google Scholar] [CrossRef]
- Lu, C.; Chen, H.; Li, M.; Jing, L. Attention-guided and topology-enhanced shift graph convolutional network for skeleton-based action recognition. Electronics 2024, 13, 3737. [Google Scholar] [CrossRef]
- Chaudhuri, S.; Bhattacharya, S. Simba: Mamba augmented U-ShiftGCN for skeletal action recognition in videos. arXiv 2024, arXiv:2404.07645. [Google Scholar]
- Wu, B.; Xue, M.; Jia, Y.; Zhang, N.; Zhao, G.; Wang, X.; Zhang, C. Lightweight and efficient skeleton-based sports activity recognition with ASTM-Net. PLoS ONE 2025, 20, e0324605. [Google Scholar] [CrossRef] [PubMed]
- Wang, L.; Zhang, X.; Zhang, C. Graph convolutional network with multi-view topology for lightweight skeleton-based action recognition. Symmetry 2025, 17, 1235. [Google Scholar] [CrossRef]
- Zhou, A.; Yang, J.; Qi, Y.; Qiao, T.; Shi, Y.; Duan, C.; Zhao, W.; Hu, C. HGNAS: Hardware-aware graph neural architecture search for edge devices. IEEE Trans. Comput. 2024, 73, 2693–2707. [Google Scholar] [CrossRef]
- Chen, Y.; Shi, Y.; Li, G.; Zhang, L.; Li, J.; Gao, J.; Chu, W. KGS-GCN: Enhancing Sparse Skeleton Sensing via Kinematics-Driven Gaussian Splatting and Probabilistic Topology for Action Recognition. arXiv 2026, arXiv:2603.16943. [Google Scholar]
- Wang, J.; Li, Z.; Liu, B.; Cai, H.; Saada, M.; Meng, Q. High-performance inference graph convolutional networks for skeleton-based action recognition. Neurocomputing 2025, 653, 131078. [Google Scholar] [CrossRef]
- Zhang, Y.Q.; Pang, C.; Geng, P.; Lu, X.Q.; Lyu, L. Multi-Scale Adaptive Large Kernel Graph Convolutional Network for Skeleton-Based Action Recognition. J. Comput. Sci. Technol. 2025, 40, 1285–1300. [Google Scholar] [CrossRef]
- Zhang, W.; Zhu, M.; Derpanis, K.G. From actemes to action: A strongly-supervised representation for action understanding. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Sydney, Australia, 1–8 December 2013; pp. 2656–2663. [Google Scholar]
- Shahroudy, A.; Liu, J.; Ng, T.-T.; Wang, G. NTU RGB+D: A large scale dataset for 3D human activity analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 1010–1019. [Google Scholar]
- Wang, J.; Nie, X.; Xia, Y.; Wu, Y.; Zhu, S.-C. Cross-view action recognition via view knowledge transfer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA, 23–28 June 2014; pp. 1891–1898. [Google Scholar]
- Song, Y.F.; Zhang, Z.; Shan, C.; Wang, L. Constructing stronger and faster baselines for skeleton-based action recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 1474–1488. [Google Scholar] [CrossRef]
- Zhou, H.; Liu, Q.; Wang, Y. Learning discriminative representations for skeleton based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 10608–10617. [Google Scholar]
- Plizzari, C.; Cannici, M.; Matteucci, M. Skeleton-based action recognition via spatial and temporal transformer networks. Comput. Vis. Image Underst. 2021, 208, 103219. [Google Scholar] [CrossRef]
- Liu, H.; Liu, Y.; Chen, Y.; Yuan, C.; Li, B.; Hu, W. TransSkeleton: Hierarchical spatial–temporal transformer for skeleton-based action recognition. IEEE Trans. Circuits Syst. Video Technol. 2023, 33, 4137–4148. [Google Scholar] [CrossRef]
- Zhou, Y.; Cheng, Z.-Q.; Li, C.; Fang, Y.; Geng, Y.; Xie, X.; Keuper, M. Hypergraph transformer for skeleton-based action recognition. arXiv 2022, arXiv:2211.09590. [Google Scholar]
- Xin, W.; Miao, Q.; Liu, Y.; Liu, R.; Pun, C.-M.; Shi, C. Skeleton mixformer: Multivariate topology representation for skeleton-based action recognition. In Proceedings of the 31st ACM International Conference on Multimedia (ACM MM), Ottawa, ON, Canada, 29 October–3 November 2023; pp. 2211–2220. [Google Scholar]



| Methods | NTU-60 (%) | NTU-120 (%) | Penn Action (%) | NW-UCLA (%) | Params (M) | Flops (G) | ||
|---|---|---|---|---|---|---|---|---|
| x-Sub | x-View | x-Sub | x-Set | |||||
| MS-G3D [27] | 91.5 | 96.2 | 86.9 | 88.4 | 96.1 | - | 2.8 (80%) | 5.2 |
| CTR-GCN [28] | 92.4 | 96.4 | 88.9 | 90.4 | 96.9 | 96.5 | 1.5 (62%) | 2.0 |
| EfficientGCN [47] | 91.7 | 95.7 | 88.3 | 89.1 | 96.7 | - | 2.0 (72%) | 15.2 |
| InfoGCN [30] | 92.8 | 96.7 | 89.2 | 90.7 | 96.5 | 96.6 | 1.6 (65%) | 1.8 |
| FRHead [48] | 93.1 | 96.8 | 89.5 | 90.9 | 97.0 | 96.8 | 2.0 (72%) | - |
| BlockGCN [8] | 92.4 | 97.0 | 90.3 | 91.5 | 96.8 | 96.9 | 1.3 (57%) | 1.6 |
| DeGCN [10] | 93.3 | 97.4 | 91.0 | 92.1 | 97.6 | 97.2 | 5.6 (90%) | - |
| ST-TR [49] | 90.8 | 96.3 | 85.1 | 87.1 | 96.3 | - | 12.1 (96%) | 259.4 |
| TranSkeleton [50] | 92.8 | 97.0 | 89.4 | 90.5 | 96.7 | - | 2.2 (75%) | 9.2 |
| Hyperformer [51] | 92.9 | 96.5 | 89.9 | 91.3 | 97.1 | 96.7 | 2.7 (80%) | 9.6 |
| SkeMixFormer [52] | 93.0 | 97.1 | 90.1 | 91.3 | 99.2 | 97.4 | 2.1 (73%) | 4.8 |
| SkateFormer [5] | 93.5 | 97.4 | 89.8 | 91.4 | 98.4 | 98.3 | 2.0 (72%) | 3.6 |
| FreqMixFormer [6] | 93.6 | 97.4 | 90.5 | 91.9 | 99.7 | 97.4 | 2.0 (72%) | 64.4 |
| EMS-GCN | 92.6 | 97.3 | 88.7 | 91.8 | 99.4 | 97.2 | 0.56 | 1.3 |
| Method | +GTRM | +MS-LTA | Action Penn (%) | NW-UCLA (%) |
|---|---|---|---|---|
| Baseline | × | × | 96.9 | 96.5 |
| Baseline + GTRM | √ | × | 98.2 | 96.7 |
| Baseline + MS-LTA | × | √ | 99.1 | 96.9 |
| EMS-GCN | √ | √ | 99.4 | 97.2 |
| Strategy | Complexity | Flops (G) | Penn Action (%) |
|---|---|---|---|
| TCN | 1.2 | 96.9 | |
| Self-Attention | 1.8 | 97.7 | |
| MS-LTA | 0.6 | 99.4 |
| Dilations Configuration | Receptive Field | Penn Action (%) |
|---|---|---|
| Single Scale | Small | 98.4 |
| Dual Scale | Medium | 99.4 |
| Triple Scale | Large | 99.2 |
| Branch Configuration | Focus Area | Penn Action (%) |
|---|---|---|
| w/o Attention | - | 98.4 |
| Global Only | Channel | 99.2 |
| Local Only | Temporal | 99.1 |
| Ours | Channel + Temporal | 99.4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, L.; Liu, H.; Jin, X. Gaussian Topology Refinement and Multi-Scale Shift Graph Convolution for Efficient Real-Time Sports Action Recognition. Symmetry 2026, 18, 639. https://doi.org/10.3390/sym18040639
Wang L, Liu H, Jin X. Gaussian Topology Refinement and Multi-Scale Shift Graph Convolution for Efficient Real-Time Sports Action Recognition. Symmetry. 2026; 18(4):639. https://doi.org/10.3390/sym18040639
Chicago/Turabian StyleWang, Longying, Hongyang Liu, and Xinyi Jin. 2026. "Gaussian Topology Refinement and Multi-Scale Shift Graph Convolution for Efficient Real-Time Sports Action Recognition" Symmetry 18, no. 4: 639. https://doi.org/10.3390/sym18040639
APA StyleWang, L., Liu, H., & Jin, X. (2026). Gaussian Topology Refinement and Multi-Scale Shift Graph Convolution for Efficient Real-Time Sports Action Recognition. Symmetry, 18(4), 639. https://doi.org/10.3390/sym18040639
