People Counting Using YOLO-Based Detection and Clustering for a Mobile Robot
Abstract
1. Introduction
- Proposing a novel algorithm for visual counting people based on clustering for a mobile robot;
- Incorporating optional modules, including an additional backbone for enhanced features and an encoder to reduce dimensionality;
- Extending the MDDRobots dataset [17] with manually labeled metadata for visual detection, re-identification, and people counting;
- Verifying the proposed method on the dataset containing indoor recordings from real mobile robots;
- Evaluating and comparing the method’s performance across platforms with varying computational capabilities.
2. Related Work
2.1. Non-Image-Based Approaches
2.2. YOLO-Based Detection
2.3. YOLO-Based People Counting
2.4. People Re-Identification and Clustering
2.5. Robotics and Real-World Applications
3. Proposed Method
| Algorithm 1 Visual People Counting Using YOLO and Feature Clustering |
|
3.1. YOLO-Based People Detection
3.2. Feature Extraction
3.3. Encoder
3.4. Clustering
4. Dataset
5. Experiments
5.1. Experiment Setup
5.2. Evaluation Metrics
5.3. People Detection
5.4. People Counting
5.5. Encoder Application
5.6. Additional Model Applications
6. Discussion
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Coşkun, A.; Kara, A.; Parlaktuna, M.; Ozkan, M.; Parlaktuna, O. People counting system by using kinect sensor. In Proceedings of the International Symposium on Innovations in Intelligent Systems and Applications (INISTA); IEEE: Piscataway, NJ, USA, 2015; pp. 1–7. [Google Scholar] [CrossRef]
- Choi, J.W.; Yim, D.H.; Cho, S.H. People Counting Based on an IR-UWB Radar Sensor. IEEE Sens. J. 2017, 17, 5717–5727. [Google Scholar] [CrossRef]
- Nguyen, H.H.; Ta, T.N.; Nguyen, N.C.; Bui, V.T.; Pham, H.M.; Nguyen, D.M. YOLO Based Real-Time Human Detection for Smart Video Surveillance at the Edge. In Proceedings of the 2020 IEEE 18th International Conference on Communications and Electronics (ICCE); IEEE: Piscataway, NJ, USA, 2021; pp. 439–444. [Google Scholar] [CrossRef]
- Kajabad, E.N.; Ivanov, S.V. People Detection and Finding Attractive Areas by the use of Movement Detection Analysis and Deep Learning Approach. Procedia Comput. Sci. 2019, 156, 327–337. [Google Scholar] [CrossRef]
- Ren, P.; Wang, L.; Fang, W.; Song, S.; Djahel, S. A novel squeeze YOLO-based real-time people counting approach. Int. J. Bio-Inspir. Comput. 2020, 16, 94–101. [Google Scholar] [CrossRef]
- Zheng, L.; Yang, Y.; Hauptmann, A. Person Re-identification: Past, Present and Future. arXiv 2016, arXiv:1610.02984. [Google Scholar] [CrossRef]
- Zhou, Y. Deep Learning Based People Detection, Tracking and Re-identification in Intelligent Video Surveillance System. In Proceedings of the International Conference on Computing and Data Science (CDS); IEEE: Piscataway, NJ, USA, 2020; pp. 443–447. [Google Scholar] [CrossRef]
- Ding, L.; Wang, S.; Li, R.; Chen, L.; Dong, J. PC-PINet: Partial Re-identification Network for People Counting with Overlapping Cameras. In Proceedings of the 6th International Conference on Image, Vision and Computing (ICIVC), Qingdao, China, 23–25 July 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 66–71. [Google Scholar] [CrossRef]
- Ndayishimiye, F.; Yoon, G.J.; Lee, J.; Yoon, S.M. Person re-identification transformer with patch attention and pruning. J. Vis. Commun. Image Represent. 2025, 106, 104348. [Google Scholar] [CrossRef]
- Fan, H.; Zheng, L.; Yan, C.; Yang, Y. Unsupervised Person Re-identification: Clustering and Fine-tuning. ACM Trans. Multimed. Comput. Commun. Appl. 2018, 14, 83. [Google Scholar] [CrossRef]
- Narayan, N.; Sankaran, N.; Arpit, D.; Dantu, K.; Setlur, S.; Govindaraju, V. Person Re-identification for Improved Multi-person Multi-camera Tracking by Continuous Entity Association. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Honolulu, HI, USA, 21–26 July 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 566–572. [Google Scholar] [CrossRef]
- Zhang, Z.Z.; Kayacan, E.; Thompson, B.G.; Chowdhary, G.V. High precision control and deep learning-based corn stand counting algorithms for agricultural robot. Auton. Robot. 2020, 44, 1289–1302. [Google Scholar] [CrossRef]
- Kejriwal, N.; Garg, S.; Kumar, S. Product counting using images with application to robot-based retail stock assessment. In Proceedings of the International Conference on Technologies for Practical Robot Applications (TePRA), Woburn, MA, USA, 11–12 May 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 1–6. [Google Scholar] [CrossRef]
- Bhangale, U.; Patil, S.; Vishwanath, V.; Thakker, P.; Bansode, A.; Navandhar, D. Near Real-time Crowd Counting using Deep Learning Approach. Procedia Comput. Sci. 2020, 171, 770–779. [Google Scholar] [CrossRef]
- Konrad, J.; Cokbas, M.; Ishwar, P.; Little, T.D.; Gevelber, M. High-accuracy people counting in large spaces using overhead fisheye cameras. Energy Build. 2024, 307, 113936. [Google Scholar] [CrossRef]
- Moussaïd, M.; Schinazi, V.R.; Kapadia, M.; Thrash, T. Virtual Sensing and Virtual Reality: How New Technologies Can Boost Research on Crowd Dynamics. Front. Robot. AI 2018, 5, 82. [Google Scholar] [CrossRef] [PubMed]
- Wozniak, P.; Krzeszowski, T.; Kwolek, B. Multi-Domain Indoor Dataset for Visual Place Recognition and Anomaly Detection by Mobile Robots. Sci. Data 2025, 12, 817. [Google Scholar] [CrossRef] [PubMed]
- Yuan, Y.; Qiu, C.; Xi, W.; Zhao, J. Crowd Density Estimation Using Wireless Sensor Networks. In Proceedings of the Seventh International Conference on Mobile Ad-Hoc and Sensor Networks, Beijing, China, 16–18 December 2011; IEEE: Piscataway, NJ, USA, 2011; pp. 138–145. [Google Scholar] [CrossRef]
- Hashimoto, M.; Tsuji, A.; Nishio, A.; Takahashi, K. Laser-based tracking of groups of people with sudden changes in motion. In Proceedings of the IEEE International Conference on Industrial Technology (ICIT), Seville, Spain, 17–19 March 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 315–320. [Google Scholar] [CrossRef]
- Raykov, Y.P.; Ozer, E.; Dasika, G.; Boukouvalas, A.; Little, M.A. Predicting room occupancy with a single passive infrared (PIR) sensor through behavior extraction. In Proceedings of the International Joint Conference on Pervasive and Ubiquitous Computing, UbiComp’16, New York, NY, USA, 12–16 September 2016; ACM: New York, NY, USA, 2016; pp. 1016–1027. [Google Scholar] [CrossRef]
- Kuchár, P.; Pirník, R.; Janota, A.; Malobický, B.; Kubík, J.; Šišmišová, D. Passenger Occupancy Estimation in Vehicles: A Review of Current Methods and Research Challenges. Sustainability 2023, 15, 1332. [Google Scholar] [CrossRef]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the Computer Vision—ECCV 2014; Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T., Eds.; Springer International Publishing: Cham, Switzerland, 2014; pp. 740–755. [Google Scholar] [CrossRef]
- Shao, S.; Li, Z.; Zhang, T.; Peng, C.; Yu, G.; Zhang, X.; Li, J.; Sun, J. Objects365: A Large-Scale, High-Quality Dataset for Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 8429–8438. [Google Scholar] [CrossRef]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Computer Vision—ECCV 2016; Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar] [CrossRef]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv 2017, arXiv:1704.04861. [Google Scholar] [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [PubMed]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-End Object Detection with Transformers. In Proceedings of the Computer Vision—ECCV; Vedaldi, A., Bischof, H., Brox, T., Frahm, J.M., Eds.; Springer International Publishing: Cham, Switzerland, 2020; pp. 213–229. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; IEEE: Piscataway, NJ, USA, 2016; pp. 779–788. [Google Scholar] [CrossRef]
- Vu, T.T.H.; Pham, D.L.; Chang, T.W. A YOLO-based Real-time Packaging Defect Detection System. Procedia Comput. Sci. 2023, 217, 886–894. [Google Scholar] [CrossRef]
- Ge, Y.; Lin, S.; Zhang, Y.; Li, Z.; Cheng, H.; Dong, J.; Shao, S.; Zhang, J.; Qi, X.; Wu, Z. Tracking and Counting of Tomato at Different Growth Period Using an Improving YOLO-Deepsort Network for Inspection Robot. Machines 2022, 10, 489. [Google Scholar] [CrossRef]
- Yang, X.; Gao, Y.; Yin, M.; Li, H. Automatic Apple Detection and Counting with AD-YOLO and MR-SORT. Sensors 2024, 24, 7012. [Google Scholar] [CrossRef] [PubMed]
- Krishna, N.M.; Reddy, R.Y.; Reddy, M.S.C.; Madhav, K.P.; Sudham, G. Object Detection and Tracking Using Yolo. In Proceedings of the 3rd International Conference on Inventive Research in Computing Applications (ICIRCA), Coimbatore, India, 2–4 September 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 1–7. [Google Scholar] [CrossRef]
- McCarthy, C.; Ghaderi, H.; Martí, F.; Jayaraman, P.; Dia, H. Video-based automatic people counting for public transport: On-bus versus off-bus deployment. Comput. Ind. 2025, 164, 104195. [Google Scholar] [CrossRef]
- Ye, S.; Bohush, R.; Chen, H.; Zakharava, I.; Ablameyko, S. Person Tracking and Reidentification for Multicamera Indoor Video Surveillance Systems. Pattern Recognit. Image Anal. 2020, 30, 827–837. [Google Scholar] [CrossRef]
- Fayyaz, M.; Yasmin, M.; Sharif, M.; Shah, J.H.; Raza, M.; Iqbal, T. Person re-identification with features-based clustering and deep features. Neural Comput. Appl. 2020, 32, 10519–10540. [Google Scholar] [CrossRef]
- Lin, Y.; Dong, X.; Zheng, L.; Yan, Y.; Yang, Y. A Bottom-Up Clustering Approach to Unsupervised Person Re-Identification. Proc. AAAI Conf. Artif. Intell. 2019, 33, 8738–8745. [Google Scholar] [CrossRef]
- Zhao, Y.; Shu, Q.; Shi, X.; Zhan, J. Unsupervised person re-identification by dynamic hybrid contrastive learning. Image Vis. Comput. 2023, 137, 104786. [Google Scholar] [CrossRef]
- Zhai, Y.; Lu, S.; Ye, Q.; Shan, X.; Chen, J.; Ji, R.; Tian, Y. AD-Cluster: Augmented Discriminative Clustering for Domain Adaptive Person Re-Identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–18 June 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 9018–9027. [Google Scholar] [CrossRef]
- Gao, Q.; Jia, M.; Chen, J.; Zhang, J. SAFA: Lifelong Person Re-Identification learning by statistics-aware feature alignment. J. Vis. Commun. Image Represent. 2025, 107, 104378. [Google Scholar] [CrossRef]
- Yıldırım, Ş.; Ulu, B. Deep Learning Based Apples Counting for Yield Forecast Using Proposed Flying Robotic System. Sensors 2023, 23, 6171. [Google Scholar] [CrossRef] [PubMed]
- Ye, H.; Zhao, J.; Zhan, Y.; Chen, W.; He, L.; Zhang, H. Person Re-Identification for Robot Person Following With Online Continual Learning. IEEE Robot. Autom. Lett. 2024, 9, 9151–9158. [Google Scholar] [CrossRef]
- Banerjee, S.; Kumar, A.; Shekhar, A. Indoor Surveillance Robot with Person Following and Re-identification. In Proceedings of the International Joint Conferenceon Neural Networks (IJCNN), Yokohama, Japan, 30 June–5 July 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1–8. [Google Scholar] [CrossRef]
- Li, S.; Hishiyama, R. An Indoor People Counting and Tracking System using mmWave sensor and sub-sensors. IFAC-PapersOnLine 2023, 56, 7096–7101. [Google Scholar] [CrossRef]
- Yang; Gonzalez-Banos; Guibas. Counting people in crowds with a real-time network of simple image sensors. In Proceedings of the Ninth IEEE International Conference on Computer Vision, Nice, France, 13–16 October 2003; IEEE: Piscataway, NJ, USA, 2003; Volume 1, pp. 122–129. [Google Scholar] [CrossRef]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR’14, Columbus, OH, USA, 23–28 June 2014; IEEE: Piscataway, NJ, USA, 2014; pp. 580–587. [Google Scholar] [CrossRef]
- Girshick, R. Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). IEEE Computer Society, Santiago, Chile, 7–13 December 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 1440–1448. [Google Scholar] [CrossRef]
- Redmon, J.; Farhadi, A. YOLOv3: An Incremental Improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef]
- Bochkovskiy, A.; Wang, C.Y.; Liao, H.Y.M. YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef]
- Weng, K.; Chu, X.; Xu, X.; Huang, J.; Wei, X. EfficientRep: An Efficient Repvgg-style ConvNets with Hardware-aware Neural Network Design. arXiv 2023, arXiv:2302.00386. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. In Proceedings of the Advances in Neural Information Processing Systems; Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2024; Volume 37, pp. 107984–108011. [Google Scholar] [CrossRef]
- Hartigan, J.A.; Wong, M.A. Algorithm AS 136: A K-Means Clustering Algorithm. J. R. Stat. Soc. Ser. C (Appl. Stat.) 1979, 28, 100–108. [Google Scholar] [CrossRef]
- Zhong, S. Efficient online spherical k-means clustering. In Proceedings of the IEEE International Conference on Neural Networks, Montreal, QC, Canada, 31 July–4 August 2005; IEEE: Piscataway, NJ, USA, 2005; Volume 5, pp. 3180–3185. [Google Scholar] [CrossRef]
- Comaniciu, D.; Meer, P. Mean shift: A robust approach toward feature space analysis. IEEE Trans. Pattern Anal. Mach. Intell. 2002, 24, 603–619. [Google Scholar] [CrossRef]
- Ester, M.; Kriegel, H.P.; Sander, J.; Xu, X. A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the Second International Conference on Knowledge Discovery and Data Mining, KDD’96, Portland, OR, USA, 2–4 August 1996; AAAI Press: Washington, DC, USA, 1996; pp. 226–231. [Google Scholar]
- Gomulka, K.; Wozniak, P.; Krzeszowski, T. People Counting Using YOLO-Based Detection and Clustering. 2026. Available online: https://github.com/GomulkaK/PeopleCounting (accessed on 14 July 2026).
- Everingham, M.; Van Gool, L.; Williams, C.K.I.; Winn, J.; Zisserman, A. The Pascal Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef]
- Powers, D. Evaluation: From Precision, Recall and F-Measure to ROC, Informedness, Markedness & Correlation. J. Mach. Learn. Technol. 2011, 2, 37–63. [Google Scholar]
- Ren, P.; Fang, W.; Djahel, S. A novel YOLO-Based real-time people counting approach. In Proceedings of the International Smart Cities Conference (ISC2), Wuxi, China, 14–17 September 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 1–2. [Google Scholar] [CrossRef]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 318–327. [Google Scholar] [CrossRef] [PubMed]
- Ge, Z.; Liu, S.; Wang, F.; Li, Z.; Sun, J. YOLOX: Exceeding YOLO Series in 2021. arXiv 2021, arXiv:2107.08430. [Google Scholar] [CrossRef]
- Khanam, R.; Hussain, M. What is YOLOv5: A deep look into the internal features of the popular object detector. arXiv 2024, arXiv:2407.20892. [Google Scholar] [CrossRef]
- Varghese, R.; Sambath, M. YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness. In Proceedings of the International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), Chennai, India, 18–19 April 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-time Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 16965–16974. [Google Scholar] [CrossRef]
- Tian, Z.; Shen, C.; Chen, H.; He, T. FCOS: A Simple and Strong Anchor-Free Object Detector. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 1922–1933. [Google Scholar] [CrossRef] [PubMed]
- Wojke, N.; Bewley, A.; Paulus, D. Simple online and realtime tracking with a deep association metric. In Proceedings of the IEEE International Conference on Image Processing (ICIP), Beijing, China, 17–20 September 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 3645–3649. [Google Scholar] [CrossRef]
- Aharon, N.; Orfaig, R.; Bobrovsky, B.Z. BoT-SORT: Robust Associations Multi-Pedestrian Tracking. arXiv 2022, arXiv:2206.14651. [Google Scholar] [CrossRef]
- Liang, D.; Chen, X.; Xu, W.; Zhou, Y.; Bai, X. Transcrowd: Weakly-supervised crowd counting with transformers. Sci. China Inf. Sci. 2022, 65, 160104. [Google Scholar] [CrossRef]
- Liang, D.; Xu, W.; Zhu, Y.; Zhou, Y. Focal Inverse Distance Transform Maps for Crowd Localization. Trans. Multi. 2023, 25, 6040–6052. [Google Scholar] [CrossRef]









| Subset | Camera | Resolution [px] | Aspect Ratio | mFPS | No. Images | No. BBox |
|---|---|---|---|---|---|---|
| iPhone | Human | 960 × 540 | 1.7 | 8.8 | 4500 | 8405 |
| GoPro | Human | 960 × 540 | 1.7 | 9.3 | 4500 | 6480 |
| P40Pro | Human | 640 × 360 | 1.7 | 4.3 | 3150 | 6849 |
| Pi Camera | Robot | 640 × 480 | 1.3 | 6.0 | 5400 | 8817 |
| Xtion | Robot | 640 × 480 | 1.3 | 1.3 | 1800 | 1843 |
| Camera | Corridor1 | Corridor2 | Corridor3 | D3A | D7 | F102 | F104 | F105 | F107 |
|---|---|---|---|---|---|---|---|---|---|
| GoPro | 1.88 | 1.71 | 1.15 | 0.93 | 0.79 | 2.10 | 2.80 | 2.48 | 2.98 |
| iPhone | 1.73 | 1.28 | 1.07 | 0.26 | 0.74 | 1.32 | 2.16 | 1.63 | 2.77 |
| P40Pro | 1.70 | 2.42 | 2.74 | 1.97 | 0.81 | 2.44 | 2.72 | 1.40 | 3.34 |
| Pi Camera | 1.70 | 2.89 | 2.72 | 1.52 | 0.56 | 1.05 | 1.67 | 0.90 | 1.70 |
| Xtion | 0.64 | 1.63 | 2.20 | 1.11 | 0.53 | 0.74 | 1.03 | 0.52 | 0.84 |
| Avg. | 1.73 | 1.99 | 1.98 | 1.36 | 0.69 | 1.53 | 2.08 | 1.39 | 2.33 |
| Feature | Platform | ||
|---|---|---|---|
| Mobile Workstation | Workstation | Jetson Nano | |
| Operating System | Windows | Ubuntu 22.04 LTS | Ubuntu 20.04 |
| CPU Model | Intel Core i7-11800H | AMD Ryzen Threadripper 3970X | ARM Cortex-A57 |
| CPU Clock Speed/Turbo | 2.3/4.6 GHz | 3.7/4.5 GHz | 1.43 GHz |
| CPU Cores/Threads | 8/16 | 32/64 | 4 |
| RAM | 16 GB | 64 GB | 4 GB |
| GPU Model | NVIDIA GeForce RTX 3060 M | NVIDIA GeForce RTX 2080 Ti | NVIDIA Maxwell GPU |
| GPU Memory | 6 GB | 11 GB | 0.5 GB |
| GPU Cores | 3840 | 4352 | 128 |
| GPU Clock Speed | 1.6 GHz | 2.2 GHz | 0.8 GHz |
| GPU Architecture | Ampere | Turing | Maxwell |
| AI Performance | 10.9 TFLOPS | 13.5 TFLOPS | 0.5 TFLOPS |
| Python Version | 3.8 | 3.10 | 3.8 |
| PyTorch Version | 2.2 | 2.2 | 1.13 |
| Compute Capability | 8.6 | 7.5 | 5.3 |
| CUDA Version | 11.8 | 12.4 | 10.8 |
| Model | Backbone | Params | Input Size [px] | 50 [%] | 50:95 [%] | [%] | Time [ms] * |
|---|---|---|---|---|---|---|---|
| DETR [27] | ResNet-50 | 41.3 M | 640 × 640 | 71.29 | 51.99 | 80.42 | 35.90 |
| 320 × 320 | 61.47 | 36.67 | 72.97 | 31.43 | |||
| FasterRCNN [26] | ResNet-50FPN | 43.7 M | 640 × 640 | 73.87 | 54.10 | 81.53 | 66.41 |
| 320 × 320 | 70.03 | 50.36 | 79.24 | 48.99 | |||
| FCOS [64] | ResNet-50FPN | 32.3 M | 640 × 640 | 63.81 | 44.21 | 75.24 | 42.00 |
| 320 × 320 | 42.55 | 29.08 | 57.49 | 38.88 | |||
| RT-DETR-L [63] | CSPResNet-50 | 33.0 M | 640 × 640 | 80.08 | 65.41 | 85.87 | 42.52 |
| 320 × 320 | 74.58 | 55.77 | 82.49 | 30.26 | |||
| RetinaNet [59] | ResNet-50FPN | 38.2 M | 640 × 640 | 60.29 | 42.23 | 71.83 | 47.59 |
| 320 × 320 | 51.15 | 34.39 | 65.07 | 27.89 | |||
| YOLOXm [60] | CSPDarknet53 | 25.3 M | 640 × 640 | 65.59 | 47.85 | 76.46 | 22.26 |
| 320 × 320 | 62.68 | 44.16 | 74.09 | 21.16 | |||
| YOLOv5m [61] | CSPDarknet53 | 21.2 M | 640 × 640 | 74.27 | 56.30 | 82.73 | 20.97 |
| 320 × 320 | 74.66 | 56.14 | 82.84 | 14.63 | |||
| YOLOv8m [62] | CSPDarknet53 | 25.9 M | 640 × 640 | 79.84 | 63.41 | 85.33 | 29.15 |
| 320 × 320 | 77.82 | 61.73 | 84.29 | 17.78 | |||
| YOLOv10m [50] | CSPDarknet53 | 15.4 M | 640 × 640 | 78.06 | 62.50 | 84.67 | 23.51 |
| 320 × 320 | 75.00 | 59.35 | 82.56 | 14.72 |
| Model Configuration | |||||||
|---|---|---|---|---|---|---|---|
| Model | Block | Min. ROI [px%] | Score Thr. | K-Means | Mean-Shift | DBSCAN | SK-Means |
| YOLOv10s | 15 | 1 | 0.7 | 1.47 | 2.00 | 2.58 | 1.31 |
| 20 | 1 | 0.7 | 1.36 | 1.96 | 2.00 | 1.29 | |
| 21 | 1 | 0.7 | 1.27 | 1.84 | 2.29 | 1.29 | |
| 21 | 2 | 0.7 | 1.36 | 1.93 | 2.27 | 1.62 | |
| 21 | 4 | 0.7 | 1.42 | 1.93 | 2.22 | 1.20 | |
| 21 | 1 | 0.5 | 1.31 | 1.87 | 1.96 | 1.51 | |
| 21 | 1 | 0.6 | 1.27 | 1.78 | 2.11 | 1.44 | |
| 21 | 4 | 0.8 | 1.38 | 2.07 | 2.53 | 1.51 | |
| Model | Method | GoPro | iPhone | P40Pro | Pi Camera | Xtion | Avg. |
| YOLOv10n | K-means | 1.44 | 1.11 | 0.67 | 1.55 | 1.11 | 1.18 |
| YOLOv10s | 1.15 | 1.11 | 0.88 | 1.89 | 0.89 | 1.26 | |
| YOLOv10m | 1.11 | 1.22 | 1.44 | 2.11 | 1.22 | 1.42 | |
| YOLOv10n | SK-means | 1.22 | 0.89 | 0.33 | 2.33 | 0.89 | 1.13 |
| YOLOv10s | 1.22 | 1.77 | 1.00 | 2.33 | 0.89 | 1.44 | |
| YOLOv10m | 1.00 | 1.22 | 1.00 | 2.11 | 1.33 | 1.33 | |
| YOLOv10n | DeepSORT [65] | 1.78 | 3.00 | 5.22 | 2.89 | 1.56 | 2.89 |
| YOLOv10n | BoT-SORT [66] | 3.78 | 4.22 | 3.33 | 4.89 | 2.78 | 3.80 |
| * | ||||||
|---|---|---|---|---|---|---|
| Model | Clustering Method | iPhone | P40Pro | Pi Camera | Xtion | Avg. |
| YOLOv10n | K-means | 1.11 | 0.67 | 1.55 | 1.11 | 1.11 |
| YOLOv10n + Enc. 32 | 1.00 | 0.67 | 2.22 | 0.78 | 1.16 | |
| YOLOv10n + Enc. 64 | 1.11 | 0.77 | 2.22 | 0.78 | 1.22 | |
| YOLOv10n | SK-means | 0.89 | 0.33 | 2.33 | 0.89 | 1.11 |
| YOLOv10n + Enc. 32 | 1.00 | 0.56 | 2.22 | 0.89 | 1.16 | |
| YOLOv10n + Enc. 64 | 1.11 | 0.66 | 2.22 | 1.00 | 1.25 | |
| Hardware | Model | Per-Frame [s] | Clustering [s] | Total * [s] | |||||
|---|---|---|---|---|---|---|---|---|---|
| Detection | Extraction | Encoder | Det. + Ex. + En. | 100 | 300 | 500 | |||
| Mobile workstation | YOLOv10n | 2.63 × | 0.70 × | — | 2.70 × | 1.24 | 1.36 | 2.23 | 2.26 |
| YOLOv10n + Enc. 32 | 2.51 × | 0.41 × | 1.52 × | 2.70 × | 1.14 | 1.23 | 1.38 | 1.41 | |
| YOLOv10n + Enc. 64 | 2.61 × | 0.45 × | 1.45 × | 2.80 × | 1.17 | 1.31 | 1.57 | 1.60 | |
| Workstation | YOLOv10n | 2.12 × | 0.22 × | — | 2.15 × | 0.28 | 0.67 | 0.96 | 0.98 |
| YOLOv10n + Enc. 32 | 1.86 × | 0.22 × | 1.03 × | 1.99 × | 0.22 | 0.29 | 0.48 | 0.50 | |
| YOLOv10n + Enc. 64 | 2.08 × | 0.21 × | 1.05 × | 2.20 × | 0.23 | 0.45 | 0.57 | 0.59 | |
| Jetson Nano | YOLOv10n | 1.36 × | 1.53 × | — | 1.37 × | 7.25 | 2.30 × 10 | 2.61 × 10 | 2.62 × 10 |
| YOLOv10n + Enc. 32 | 1.48 × | 1.72 × | 5.13 × | 1.55 × | 1.75 | 3.56 | 9.94 | 1.01 × 10 | |
| YOLOv10n + Enc. 64 | 1.88 × | 4.34 × | 5.42 × | 1.98 × | 2.97 | 7.73 | 1.04 × 10 | 1.06 × 10 | |
| Model | Additional Model | Feature Size * | ** | ||||
|---|---|---|---|---|---|---|---|
| iPhone | P40Pro | Pi Camera | Xtion | Avg. | |||
| YOLOv10n | – | 384 | 1.00 | 0.67 | 2.22 | 0.78 | 1.16 |
| YOLOv10n | Mobilenetv3 | 384 + 576 | 1.00 | 0.67 | 2.22 | 0.78 | 1.17 |
| YOLOv10n | ResNet-18 | 384 + 512 | 1.00 | 0.56 | 2.22 | 0.78 | 1.14 |
| YOLOv10n | SqueezeNet | 384 + 512 | 1.00 | 0.56 | 2.33 | 0.78 | 1.17 |
| YOLOv10n | ViT | 384 + 768 | 1.11 | 0.78 | 2.22 | 1.11 | 1.31 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gomulka, K.; Wozniak, P.; Krzeszowski, T. People Counting Using YOLO-Based Detection and Clustering for a Mobile Robot. Sensors 2026, 26, 4664. https://doi.org/10.3390/s26154664
Gomulka K, Wozniak P, Krzeszowski T. People Counting Using YOLO-Based Detection and Clustering for a Mobile Robot. Sensors. 2026; 26(15):4664. https://doi.org/10.3390/s26154664
Chicago/Turabian StyleGomulka, Kamil, Piotr Wozniak, and Tomasz Krzeszowski. 2026. "People Counting Using YOLO-Based Detection and Clustering for a Mobile Robot" Sensors 26, no. 15: 4664. https://doi.org/10.3390/s26154664
APA StyleGomulka, K., Wozniak, P., & Krzeszowski, T. (2026). People Counting Using YOLO-Based Detection and Clustering for a Mobile Robot. Sensors, 26(15), 4664. https://doi.org/10.3390/s26154664

