YOLO-CBD: Classroom Behavior Detection Method Based on Behavior Feature Extraction and Aggregation
Abstract
1. Introduction
- (1)
- This paper proposes a classroom behavior detection model, YOLO-CBD (YOLOv10s Classroom Behavior Detection), targeting crowd occlusion and small-target detection in dense classrooms. By constructing a novel FMBNet backbone, the model enhances the capture of behavior features in densely crowds.
- (2)
- To improve the detection capability for small classroom behaviors in distant instances, the AKConv deformable convolution is incorporated for feature aggregation structure to create the VACSP (VoVGSCSP + AKConv) module, replacing all C2f modules in the neck network. This modification allows the network to more precisely identify and locate occluded small targets, enhancing its handling capacity in complex scenarios.
- (3)
- To address the multi-scale variations of student targets in classroom environments, this study integrates the GSConv module into the Bi-FPN feature pyramid network to resolve the loss of feature diversity after feature fusion. This optimizes the network’s multi-scale features fusion, enhancing its ability to recognize and locate targets of different scales.
- (4)
- This paper modifies CIOU to Wise-IoU, enabling the network to adaptively evaluate the difficulty level of samples by dynamically adjusting the weight function, enhancing the detection performance and generalization ability.
2. Related Works
2.1. YOLOv10
2.2. Student Classroom Behavior Detection
2.3. Attention Mechanism
3. Methods
3.1. YOLO-CBD Architecture
3.2. FMBNet Backbone Network
3.3. VACSP Module
3.4. BGC-FPN Structure
3.5. Wise-IoU Loss Function
4. Experiment Setup
4.1. Dataset and Experimental Environment
4.2. Model Training
4.3. Evaluation Metrics
5. Experimental Results and Analysis
5.1. Ablation Study
5.2. Visualization Analysis
5.3. Performance Comparison with Mainstream Models on the Classroom Behavior Dataset
5.4. Performance Comparison with Mainstream Models on the Public Dataset
5.5. Experimental Results Analysis and Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Messeri, L.; Crockett, M.J. Artificial intelligence and illusions of understanding in scientific research. Nature 2024, 627, 49–58. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Saini, M.K.; Goel, N. How smart are smart classrooms? A review of smart classroom technologies. ACM Comput. Surv. (CSUR) 2019, 52, 130. [Google Scholar] [CrossRef] [Scilit]
- Chu, L.; Liu, Y.; Zhai, Y.; Wang, D.; Wu, Y. The use of deep learning integrating image recognition in language analysis technology in secondary school education. Sci. Rep. 2024, 14, 2888. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, R.; Chen, S.; Tian, G.; Wang, P.; Ying, S. Post-secondary classroom teaching quality evaluation using small object detection model. Sci. Rep. 2024, 14, 5816. [Google Scholar] [CrossRef] [Scilit]
- Dang, M.; Liu, G.; Li, H.; Xu, Q.; Wang, X.; Pan, R. Multi-object behaviour recognition based on object detection cascaded image classification in classroom scenes. Appl. Intell. 2024, 54, 4935–4951. [Google Scholar] [CrossRef] [Scilit]
- Zhou, H.; Jiang, F.; Si, J.; Xiong, L.; Lu, H. Stuart: Individualized classroom observation of students with automatic behavior recognition and tracking. In Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Cao, Q.; Qian, C.; Chen, D. YOLO-AMM: A Real-Time Classroom Behavior Detection Algorithm Based on Multi-Dimensional Feature Optimization. Sensors 2025, 25, 1142. [Google Scholar] [CrossRef] [Scilit]
- Dohare, S.; Hernandez-Garcia, J.F.; Lan, Q.; Rahman, P.; Mahmood, A.R.; Sutton, R.S. Loss of plasticity in deep continual learning. Nature 2024, 632, 768–774. [Google Scholar] [CrossRef] [Scilit]
- Wu, S.; Wang, B. DRSI-Net: Dual-residual spatial interaction network for multi-person pose estimation. Knowl.-Based Syst. 2024, 295, 111836. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Nie, J.; Wei, S.; Zhu, G.; Dai, W.; Yang, C. A Study of Classroom Behavior Recognition Incorporating Super-Resolution and Target Detection. Sensors 2024, 24, 5640. [Google Scholar] [CrossRef] [Scilit]
- Kim, B.; Kim, J.; Chang, H.J.; Oh, T.H. A unified framework for unsupervised action learning via global-to-local motion transformer. Pattern Recognit. 2025, 159, 111118. [Google Scholar] [CrossRef] [Scilit]
- Ren, S.Q.; He, K.M.; Girshick, R.; Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Redmon, J.; Farhadi, A. Yolov3: An incremental improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef] [Scilit]
- Khanam, R.; Hussain, M. What is yolov5: A deep look into the internal features of the popular object detector. arXiv 2024, arXiv:2407.20892. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv 2022, arXiv:2209.02976. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Bochkovskiy, A.; Liao, H. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv 2023, arXiv:2207.02696. [Google Scholar] [CrossRef] [Scilit]
- Liu, Q.; Jiang, R.; Xu, Q.; Wang, D.; Sang, Z.; Jiang, X. Yolov8n_bt: Research on classroom learning behavior recognition algorithm based on improved yolov8n. IEEE Access 2024, 12, 36391–36403. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Yeh, I.; Liao, H.M. Yolov9: Learning what you want to learn using programmable gradient information. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September—4 October 2024; pp. 1–21. [Google Scholar] [CrossRef] [Scilit]
- Cheng, T.; Song, L.; Ge, Y.; Liu, W.; Wang, X.; Shan, Y. Yolo-world: Real-time open-vocabulary object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024. [Google Scholar] [CrossRef] [Scilit]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. Ssd: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision 2016, Amsterdam, The Netherlands, 8–16 October 2016; pp. 21–37. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Qi, X.; Saudagar, A.K.J.; Badshah, A.M.; Muhammad, K.; Liu, S. Student behavior recognition for interaction detection in the classroom environment. Image Vis. Comput. 2023, 136, 104726. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Li, L.; Zeng, C.; Dong, S.; Sun, J. SLBDetection-Net: Towards closed-set and open-set student learning behavior detection in smart classroom of K-12 education. Expert Syst. Appl. 2025, 260, 125392. [Google Scholar] [CrossRef] [Scilit]
- Zhao, J.; Zhu, H.; Niu, L. BiTNet: A lightweight object detection network for real-time classroom behavior recognition with transformer and bi-directional pyramid network. J. King Saud Univ. Comput. Inf. Sci. 2023, 35, 101670. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Huang, Z.; Song, Q.; Bai, K. PV-YOLO: A lightweight pedestrian and vehicle detection model based on improved YOLOv8. Digit. Signal Process. 2025, 156, 104857. [Google Scholar] [CrossRef] [Scilit]
- Jiao, B.; Wang, Y.; Wang, P.; Wang, H.; Yue, H. RS-YOLO: An efficient object detection algorithm for road scenes. Digit. Signal Process. 2025, 157, 104889. [Google Scholar] [CrossRef] [Scilit]
- Zhao, J.; Zhu, H. Cbph-net: A small object detector for behavior recognition in classroom scenarios. IEEE Trans. Instrum. Meas. 2023, 72, 1–12. [Google Scholar] [CrossRef] [Scilit]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. arXiv 2024, arXiv:2405.14458. [Google Scholar] [CrossRef] [Scilit]
- Tan, M.; Le, Q. EfficientNetV2: Smaller Models and Faster Training. arXiv 2021, arXiv:2104.00298. [Google Scholar] [CrossRef] [Scilit]
- Zhu, L.; Wang, X.; Ke, Z.; Zhang, W.; Lau, R.W. Biformer: Vision transformer with bi-level routing attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2023, Vancouver, BC, Canada, 17–24 June 2023. [Google Scholar] [CrossRef] [Scilit]
- Tan, M.; Pang, R.; Le, Q.V. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2020, Seattle, WA, USA, 13–19 June 2020. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Li, J.; Wei, H.; Liu, Z.; Zhan, Z.; Ren, Q. Slim-neck by GSConv: A better design paradigm of detector architectures for autonomous vehicles. arXiv 2022, arXiv:2206.02424. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Song, Y.; Song, T.; Yang, D.; Ye, Y.; Zhou, J.; Zhang, L. AKConv: Convolutional kernel with arbitrary sampled shapes and arbitrary number of parameters. arXiv 2023, arXiv:2311.11587. [Google Scholar] [CrossRef] [Scilit]
- Tong, Z.; Chen, Y.; Xu, Z.; Yu, R. Wise-IoU: Bounding box regression loss with dynamic focusing mechanism. arXiv 2023, arXiv:2301.10051. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Zhu, Y.; Liu, J.; Yu, W.; Jiang, C. An Interpretability optimization method for deep learning networks based on grad-CAM. IEEE Internet Things J. 2025, 12, 3961–3970. [Google Scholar] [CrossRef] [Scilit]















| Models | Modules | P (%) | R (%) | mAP50 (%) | Params (M) | FPS | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| A | B | C | D | E | ||||||
| Baseline model | × | × | × | × | × | 87.8 | 86.2 | 89.9 | 8.92 | 435 |
| +A | √ | × | × | × | × | 89.4 | 82.2 | 91.4 | 9.65 | 455 |
| +A+B | √ | √ | × | × | × | 91.4 | 87.2 | 92.0 | 9.66 | 400 |
| +A+B+C | √ | √ | √ | × | × | 92.3 | 88.1 | 92.5 | 9.47 | 435 |
| +A+B+C+D | √ | √ | √ | √ | × | 93.1 | 89.7 | 93.1 | 9.43 | 417 |
| YOLO-CBD (ours) | √ | √ | √ | √ | √ | 93.5 | 89.9 | 93.4 | 9.43 | 435 |
| Models | P (%) | R (%) | mAP50 (%) | Params (M) | FPS |
|---|---|---|---|---|---|
| YOLOv10s [28] | 87.8 | 86.2 | 89.9 | 8.92 | 435 |
| YOLOv3-spp [13] | 83.0 | 77.6 | 84.6 | 62.55 | 135 |
| YOLOv5s [14] | 83.0 | 81.7 | 87.0 | 7.02 | 333 |
| YOLOv6s [15] | 85.0 | 82.3 | 87.7 | 16.30 | 909 |
| YOLOv7 [16] | 85.7 | 86.2 | 90.7 | 36.49 | 156 |
| YOLOv8s [17] | 85.3 | 85.1 | 89.4 | 11.13 | 909 |
| YOLOv9-c [18] | 82.5 | 79.8 | 85.9 | 3.60 | 132 |
| YOLOv8-Worldv2 [19] | 88.1 | 84.8 | 89.8 | 12.75 | 1250 |
| YOLOv8-DETR [20] | 84.4 | 86.2 | 89.9 | 28.79 | 556 |
| Faster R-CNN [12] | 75.3 | 83.2 | 84.7 | 136.71 | 48 |
| SSD [21] | 82.8 | 40.7 | 64.9 | 3.68 | 146 |
| YOLO-CBD (Ours) | 93.5 | 89.9 | 93.4 | 9.43 | 435 |
| Models | P (%) | R (%) | mAP50 (%) | Params (M) | FPS |
|---|---|---|---|---|---|
| YOLOv10s [28] | 74.4 | 58.8 | 64.5 | 8.94 | 455 |
| YOLOv3-spp [13] | 83.9 | 63.5 | 71.5 | 62.65 | 164 |
| YOLOv5s [14] | 76.1 | 59.6 | 66.9 | 7.06 | 169 |
| YOLOv6s [15] | 60.3 | 47.7 | 50.6 | 16.30 | 476 |
| YOLOv7 [16] | 67.3 | 45.0 | 49.3 | 36.58 | 128 |
| YOLOv8s [17] | 63.9 | 47.7 | 52.5 | 11.13 | 435 |
| YOLOv9-c [18] | 52.7 | 44.2 | 45.3 | 3.61 | 51 |
| YOLOv8-Worldv2 [19] | 74.3 | 60.4 | 67.1 | 12.75 | 333 |
| YOLOv8-DETR [20] | 58.8 | 46.5 | 48.8 | 28.80 | 208 |
| Faster R-CNN [12] | 57.6 | 72.8 | 70.6 | 137.08 | 50 |
| SSD [21] | 85.0 | 54.3 | 67.1 | 6.07 | 135 |
| YOLO-CBD (Ours) | 85.2 | 64.1 | 72.3 | 9.44 | 385 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Peng, S.; Zhang, X.; Zhou, L.; Wang, P. YOLO-CBD: Classroom Behavior Detection Method Based on Behavior Feature Extraction and Aggregation. Sensors 2025, 25, 3073. https://doi.org/10.3390/s25103073
Peng S, Zhang X, Zhou L, Wang P. YOLO-CBD: Classroom Behavior Detection Method Based on Behavior Feature Extraction and Aggregation. Sensors. 2025; 25(10):3073. https://doi.org/10.3390/s25103073
Chicago/Turabian StylePeng, Shuyun, Xiaopei Zhang, Luoyu Zhou, and Peng Wang. 2025. "YOLO-CBD: Classroom Behavior Detection Method Based on Behavior Feature Extraction and Aggregation" Sensors 25, no. 10: 3073. https://doi.org/10.3390/s25103073
APA StylePeng, S., Zhang, X., Zhou, L., & Wang, P. (2025). YOLO-CBD: Classroom Behavior Detection Method Based on Behavior Feature Extraction and Aggregation. Sensors, 25(10), 3073. https://doi.org/10.3390/s25103073

