RGB Gait Recognition Using Large Vision Models for Industrial Access Control
Abstract
1. Introduction
- We establish an RGB evaluation on our previously introduced industrial gait dataset, using 1500 sequences from 150 workers to examine recognition under bidirectional views, gate-specific actions, and occlusion.
- We compare DINOv2-S, DINOv3-S, and DINOv3-S+ using a common token grid and the same downstream configuration, reporting recognition accuracy together with measured inference costs.
- We analyze human-prior transfer and compare raw RGB, background-suppressed RGB, and background-suppressed RGB with one-patch boundary expansion for same-view and cross-view recognition.
2. Related Work
2.1. Gait Representations and Robust Recognition
2.2. RGB-Based Gait Recognition
2.3. Gait Datasets and Industrial Access-Control Scenarios
3. Materials and Methods
3.1. Industrial Access-Control Dataset and Input Construction
3.1.1. Acquisition Setting and Protocol
3.1.2. Human-Prior Branch
3.1.3. Backbone-Independent Foreground Processing
3.2. Industrial RGB Gait Recognition with BiggerGait*
3.2.1. Grouped Layer-Wise Gait Representation
3.2.2. DINO Backbones and Token Alignment
3.2.3. Evaluation Protocol
3.2.4. Training and Inference Configuration
4. Results
4.1. Transfer of Fixed Human-Prior Weights
4.2. Backbone and Input Construction
4.3. Computational Efficiency
4.4. Background Suppression and Deployment Implications
5. Discussion
Application Scope and Limitations
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. DINOv2: Learning Robust Visual Features Without Supervision. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Ye, D.; Fan, C.; Ma, J.; Liu, X.; Yu, S. BigGait: Learning Gait Representation You Want by Large Vision Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2024; pp. 200–210. [Google Scholar]
- Ye, D.; Fan, C.; Huang, Z.; Luo, C.; Li, J.; Yu, S.; Liu, X. BiggerGait: Unlocking Gait Recognition with Layer-Wise Representations from Large Vision Models. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2025; Volume 38. [Google Scholar]
- Bai, J.; Wu, H.; Li, X.; Zhang, X. Occlusion-robust gait recognition in bidirectional industrial access-control scenarios. Pattern Recognit. Lett. 2026, 209, 8–14. [Google Scholar] [CrossRef] [Scilit]
- Siméoni, O.; Vo, H.V.; Seitzer, M.; Baldassarre, F.; Oquab, M.; Jose, C.; Khalidov, V.; Szafraniec, M.; Yi, S.; Ramamonjisoa, M.; et al. DINOv3. arXiv 2025, arXiv:2508.10104. [Google Scholar] [CrossRef] [Scilit]
- Han, J.; Bhanu, B. Individual Recognition Using Gait Energy Image. IEEE Trans. Pattern Anal. Mach. Intell. 2006, 28, 316–322. [Google Scholar] [CrossRef] [Scilit]
- Chao, H.; Wang, K.; He, Y.; Zhang, J.; Feng, J. GaitSet: Cross-View Gait Recognition Through Utilizing Gait as a Deep Set. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 3467–3478. [Google Scholar] [CrossRef] [Scilit]
- Fan, C.; Peng, Y.; Cao, C.; Liu, X.; Hou, S.; Chi, J.; Huang, Y.; Li, Q.; He, Z. GaitPart: Temporal Part-Based Model for Gait Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2020; pp. 14225–14233. [Google Scholar]
- Lin, B.; Zhang, S.; Yu, X. Gait Recognition via Effective Global-Local Feature Representation and Local Temporal Aggregation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2021; pp. 14648–14656. [Google Scholar]
- Fan, C.; Hou, S.; Huang, Y.; Yu, S. Exploring Deep Models for Practical Gait Recognition. arXiv 2023, arXiv:2303.03301. [Google Scholar]
- Fan, C.; Liang, J.; Shen, C.; Hou, S.; Huang, Y.; Yu, S. OpenGait: Revisiting Gait Recognition Toward Better Practicality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 9707–9716. [Google Scholar]
- Wang, Z.; Hou, S.; Zhang, M.; Liu, X.; Cao, C.; Huang, Y.; Li, P.; Xu, S. QAGait: Revisit Gait Recognition from a Quality Perspective. Proc. AAAI Conf. Artif. Intell. 2024, 38, 5785–5793. [Google Scholar] [CrossRef] [Scilit]
- Peng, G.; Wang, Y.; Zhang, S.; Li, R.; Zhao, Y.; Li, A. RSANet: Relative-Sequence Quality Assessment Network for Gait Recognition in the Wild. Pattern Recognit. 2025, 161, 111219. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Zheng, J.; Zhu, S.; Yan, C. TrackletGait: A Robust Framework for Gait Recognition in the Wild. IEEE Trans. Multimed. 2025, 27, 8875–8887. [Google Scholar] [CrossRef] [Scilit]
- Liao, R.; Yu, S.; An, W.; Huang, Y. A Model-Based Gait Recognition Method with Body Pose and Human Prior Knowledge. Pattern Recognit. 2020, 98, 107069. [Google Scholar] [CrossRef] [Scilit]
- Teepe, T.; Khan, A.; Gilg, J.; Herzog, F.; Hörmann, S.; Rigoll, G. GaitGraph: Graph Convolutional Network for Skeleton-Based Gait Recognition. In Proceedings of the 2021 IEEE International Conference on Image Processing (ICIP); IEEE: Piscataway, NJ, USA, 2021; pp. 2314–2318. [Google Scholar] [CrossRef] [Scilit]
- Teepe, T.; Gilg, J.; Herzog, F.; Hörmann, S.; Rigoll, G. Towards a Deeper Understanding of Skeleton-Based Gait Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops; IEEE: Piscataway, NJ, USA, 2022; pp. 1568–1576. [Google Scholar] [CrossRef] [Scilit]
- Fan, C.; Ma, J.; Jin, D.; Shen, C.; Yu, S. SkeletonGait: Gait Recognition Using Skeleton Maps. Proc. AAAI Conf. Artif. Intell. 2024, 38, 1662–1669. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Makihara, Y.; Xu, C.; Yagi, Y.; Yu, S.; Ren, M. End-to-End Model-Based Gait Recognition. In Proceedings of the Computer Vision–ACCV 2020; Springer: Cham, Switzerland, 2020; pp. 3–20. [Google Scholar] [CrossRef] [Scilit]
- Zheng, J.; Liu, X.; Gu, X.; Sun, Y.; Gan, C.; Yan, C.; Mei, T. Parsing is All You Need for Accurate Gait Recognition in the Wild. In Proceedings of the ACM International Conference on Multimedia; Association for Computing Machinery: New York, NY, USA, 2023; pp. 116–124. [Google Scholar]
- Peng, Y.; Ma, K.; Zhang, Y.; He, Z. Learning Rich Features for Gait Recognition by Integrating Skeletons and Silhouettes. Multimed. Tools Appl. 2024, 83, 7273–7294. [Google Scholar] [CrossRef] [Scilit]
- Jin, D.; Fan, C.; Chen, W.; Yu, S. Exploring More from Multiple Gait Modalities for Human Identification. Proc. AAAI Conf. Artif. Intell. 2025, 39, 4120–4128. [Google Scholar] [CrossRef] [Scilit]
- Deelaka, P.N.; De Silva, D.Y.; Wickramanayake, S.; Meedeniya, D.; Rasnayaka, S. TEZARNet: TEmporal Zero-Shot Activity Recognition Network. In Proceedings of the Neural Information Processing; Luo, B., Cheng, L., Wu, Z.G., Li, H., Li, C., Eds.; Springer: Singapore, 2024; pp. 444–455. [Google Scholar] [CrossRef] [Scilit]
- De Silva, D.Y.; Wickramanayake, S.; Meedeniya, D.; Rasnayaka, S. SEZ-HARN: Self-Explainable Zero-shot Human Activity Recognition Network. arXiv 2025, arXiv:2507.00050. [Google Scholar] [CrossRef] [Scilit]
- Liang, J.; Fan, C.; Hou, S.; Shen, C.; Huang, Y.; Yu, S. GaitEdge: Beyond Plain End-to-End Gait Recognition for Better Practicality. In Proceedings of the Computer Vision–ECCV 2022; Springer: Cham, Switzerland, 2022; pp. 375–390. [Google Scholar] [CrossRef] [Scilit]
- Castro, F.M.; Delgado-Escaño, R.; Hernández-García, R.; Marín-Jiménez, M.J.; Guil, N. AttenGait: Gait Recognition with Attention and Rich Modalities. Pattern Recognit. 2024, 148, 110171. [Google Scholar] [CrossRef] [Scilit]
- Jin, D.; Fan, C.; Ma, J.; Zhou, J.; Chen, W.; Yu, S. On Denoising Walking Videos for Gait Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2025; pp. 12347–12357. [Google Scholar]
- Huang, Z.; Ye, D.; Liu, X.; Kong, Y. Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2026; pp. 28379–28390. [Google Scholar]
- Yu, S.; Tan, D.; Tan, T. A Framework for Evaluating the Effect of View Angle, Clothing and Carrying Condition on Gait Recognition. In Proceedings of the 18th International Conference on Pattern Recognition; IEEE: Piscataway, NJ, USA, 2006; Volume 4, pp. 441–444. [Google Scholar]
- Takemura, N.; Makihara, Y.; Muramatsu, D.; Echigo, T.; Yagi, Y. Multi-View Large Population Gait Dataset and Its Performance Evaluation for Cross-View Gait Recognition. IPSJ Trans. Comput. Vis. Appl. 2018, 10, 4. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Z.; Guo, X.; Yang, T.; Huang, J.; Deng, J.; Huang, G.; Du, D.; Lu, J.; Zhou, J. Gait Recognition in the Wild: A Benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2021; pp. 14789–14799. [Google Scholar]
- Zheng, J.; Liu, X.; Liu, W.; He, L.; Yan, C.; Mei, T. Gait Recognition in the Wild with Dense 3D Representations and a Benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2022; pp. 20228–20237. [Google Scholar]
- Li, W.; Hou, S.; Zhang, C.; Cao, C.; Liu, X.; Huang, Y.; Zhao, Y. An In-Depth Exploration of Person Re-Identification and Gait Recognition in Cloth-Changing Conditions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 13824–13833. [Google Scholar]
- Shen, C.; Fan, C.; Wu, W.; Wang, R.; Huang, G.Q.; Yu, S. LidarGait: Benchmarking 3D Gait Recognition with Point Clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2023; pp. 1054–1063. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLOv8. Computer Software. 2023. Available online: https://github.com/ultralytics/ultralytics (accessed on 13 September 2026).






| Setting | Backbone | Params. (M) | FFN | Input | Patch | Token Grid |
|---|---|---|---|---|---|---|
| DINOv2-S | ViT-S/14 | 22.057 | MLP | |||
| DINOv3-S | ViT-S/16 | 21.597 | MLP | |||
| DINOv3-S+ | ViT-S+/16 | 28.693 | SwiGLU |
| Backbone | Prior Weights | Same-View R1 | Cross-View R1 | Mean R5 | |||||
|---|---|---|---|---|---|---|---|---|---|
| Avg. | Avg. | Same | Cross | ||||||
| DINOv2-S | Released weights | 95.0 | 91.7 | 93.3 | 87.5 | 85.0 | 86.3 | 99.6 | 97.9 |
| DINOv3-S | Direct transfer | 90.8 | 92.5 | 91.7 | 74.2 | 82.5 | 78.3 | 98.8 | 98.3 |
| DINOv3-S | Adapted weights | 95.8 | 93.3 | 94.6 | 87.5 | 88.3 | 87.9 | 99.2 | 98.8 |
| Backbone | Input | Same-View R1 | Cross-View R1 | Mean R5 | |||||
|---|---|---|---|---|---|---|---|---|---|
| Avg. | Avg. | Same | Cross | ||||||
| DINOv2-S | Raw RGB | 90.8 | 88.3 | 89.6 | 80.8 | 86.7 | 83.8 | 98.3 | 97.5 |
| Suppressed | 95.8 | 93.3 | 94.6 | 81.7 | 85.0 | 83.3 | 100.0 | 96.3 | |
| Expanded | 94.2 | 93.3 | 93.8 | 80.0 | 89.2 | 84.6 | 99.6 | 97.9 | |
| DINOv3-S | Raw RGB | 89.2 | 87.5 | 88.3 | 81.7 | 85.0 | 83.3 | 98.8 | 96.7 |
| Suppressed | 95.8 | 92.5 | 94.2 | 83.3 | 85.8 | 84.6 | 99.6 | 97.1 | |
| Expanded | 95.0 | 92.5 | 93.8 | 85.8 | 86.7 | 86.3 | 99.6 | 98.8 | |
| DINOv3-S+ | Raw RGB | 90.0 | 87.5 | 88.8 | 85.0 | 88.3 | 86.7 | 99.2 | 97.9 |
| Suppressed | 97.5 | 95.8 | 96.7 | 85.8 | 86.7 | 86.3 | 100.0 | 99.2 | |
| Expanded | 96.7 | 95.8 | 96.3 | 85.0 | 91.7 | 88.3 | 99.6 | 98.3 | |
| Backbone | Total Params. (M) | GFLOPs/Seq. | Latency (ms/Seq.) | Memory (MiB) | Throughput (Seq./s) |
|---|---|---|---|---|---|
| DINOv2-S | 42.668 | 2674.228 | 1556.879 | 26.604 | |
| DINOv3-S | 42.208 | 2688.474 | 2047.062 | 21.109 | |
| DINOv3-S+ | 49.304 | 3017.808 | 2073.632 | 19.154 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Bai, J.; Wu, H.; Liu, Y.; Qin, J.; Li, X. RGB Gait Recognition Using Large Vision Models for Industrial Access Control. J. Imaging 2026, 12, 458. https://doi.org/10.3390/jimaging12090458
Bai J, Wu H, Liu Y, Qin J, Li X. RGB Gait Recognition Using Large Vision Models for Industrial Access Control. Journal of Imaging. 2026; 12(9):458. https://doi.org/10.3390/jimaging12090458
Chicago/Turabian StyleBai, Jiaqi, Huijuan Wu, Yan Liu, Jinjiang Qin, and Xiaoying Li. 2026. "RGB Gait Recognition Using Large Vision Models for Industrial Access Control" Journal of Imaging 12, no. 9: 458. https://doi.org/10.3390/jimaging12090458
APA StyleBai, J., Wu, H., Liu, Y., Qin, J., & Li, X. (2026). RGB Gait Recognition Using Large Vision Models for Industrial Access Control. Journal of Imaging, 12(9), 458. https://doi.org/10.3390/jimaging12090458

