RCAF-Net: Wildlife Target Detection in Complex Forest Scenarios
Simple Summary
Abstract
1. Introduction
- (1)
- A wildlife detection dataset was established for typical complex forest regions in Northeast China. This dataset encompasses various natural environments, including forests, snowy landscapes, and bushes, and features eight representative wildlife species. It effectively reflects the challenges encountered in real-world monitoring, such as strong background interference, significant scale variations, frequent occlusions, and high similarity between targets and their environments. This dataset provides a foundational resource for model training, performance evaluation, and subsequent method improvements.
- (2)
- An improved wildlife target detection framework based on YOLO11n was developed for complex forest monitoring scenarios. The proposed framework incorporates PMGHA, RFAConv, CSFCN, and ELGH to enhance feature representation, multi-scale feature fusion, and lightweight deployment capability under complex natural environments. By jointly optimizing detection performance and deployment efficiency, the proposed method improves the robustness and adaptability of wildlife target detection in complex forest scenes.
- (3)
- Extensive experiments demonstrate that RCAF-Net achieves improved detection performance while maintaining good deployment efficiency in complex forest scenarios. In addition, visualization analysis and deployment validation further confirm the effectiveness and practical applicability of the proposed method for wildlife monitoring tasks.
2. Materials and Methods
2.1. Data and Preprocessing
2.2. Construction Process of RCAF-Net
2.2.1. C3K2-RFAConv Backbone Feature Extraction Module Based on Receptive-Field Attention
2.2.2. PMGHA Shallow Feature Enhancement Module Based on Parallel Mixed Attention
2.2.3. Context Calibration and Spatial Alignment-Based CSFCN Feature Fusion Module
2.2.4. Group-Convolution-Based ELGH Lightweight Detection Head
3. Experiments and Results Analysis
3.1. Experimental Environment and Parameter Settings
3.2. Performance Evaluation Metrics
3.3. Ablation Study of the Improved Module
3.4. Random Seed Stability Analysis
3.5. Performance Comparison with Mainstream Detection Models
3.6. Class Recognition Analysis Based on Confusion Matrix
3.7. Visualization Analysis of Attention Regions Based on Grad-CAM
3.8. Comparative Analysis of Detection Results Visualization
3.9. Cross-Dataset Generalization Performance Analysis
3.10. Embedded Edge Device Deployment Validation
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- IPBES. Thematic Assessment Report on the Sustainable Use of Wild Species; Fromentin, J.M., Emery, M.R., Eds.; IPBES Secretariat: Bonn, Germany, 2022. [Google Scholar] [CrossRef]
- Tuia, D.; Kellenberger, B.; Beery, S.; Costelloe, B.R.; Zuffi, S.; Risse, B.; Mathis, A.; Mathis, M.W.; van Langevelde, F.; Burghardt, T.; et al. Perspectives in machine learning for wildlife conservation. Nat. Commun. 2022, 13, 792. [Google Scholar] [CrossRef]
- Ahumada, J.A.; Fegraus, E.; Birch, T.; Flores, N.; Kays, R.; O’Brien, T.G.; Palmer, J.; Schuttler, S.; Zhao, J.; Jetz, W.; et al. Wildlife Insights: A Platform to Maximize the Potential of Camera Trap and Other Passive Sensor Wildlife Data for the Planet. Environ. Conserv. 2020, 47, 1–6. [Google Scholar] [CrossRef]
- Lahoz-Monfort, J.J.; Magrath, M.J.L. A Comprehensive Overview of Technologies for Species and Habitat Monitoring and Conservation. BioScience 2021, 71, 1038–1062. [Google Scholar] [CrossRef]
- Zemanova, M.A. Towards More Compassionate Wildlife Research through the 3Rs Principles: Moving from Invasive to Non-Invasive Methods. Wildl. Biol. 2020, 2020, wlb.00607. [Google Scholar] [CrossRef]
- Tan, M.; Chao, W.; Cheng, J.-K.; Zhou, M.; Ma, Y.; Jiang, X.; Ge, J.; Yu, L.; Feng, L. Animal Detection and Classification from Camera Trap Images Using Different Mainstream Object Detection Architectures. Animals 2022, 12, 1976. [Google Scholar] [CrossRef] [PubMed]
- Caravaggi, A.; Banks, P.B.; Burton, A.C.; Finlay, C.M.V.; Haswell, P.M.; Hayward, M.W.; Rowcliffe, J.M.; Wood, M.D. A Review of Camera Trapping for Conservation Behaviour Research. Remote Sens. Ecol. Conserv. 2017, 3, 109–122. [Google Scholar] [CrossRef]
- Norouzzadeh, M.S.; Nguyen, A.; Kosmala, M.; Swanson, A.; Palmer, M.S.; Packer, C.; Clune, J. Automatically Identifying, Counting, and Describing Wild Animals in Camera-Trap Images with Deep Learning. Proc. Natl. Acad. Sci. USA 2018, 115, E5716–E5725. [Google Scholar] [CrossRef]
- Liu, L.; Ouyang, W.; Wang, X.; Fieguth, P.; Chen, J.; Liu, B.; Pietikäinen, M. Deep learning for generic object detection: A survey. Int. J. Comput. Vis. 2020, 128, 261–318. [Google Scholar] [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef]
- Ma, J.; Liu, Z.; Yao, W.; Xie, Z.; Gao, J.; Chen, W. Detection of Large Herbivores in UAV Images: A New Method for Small Target Recognition in Large-Scale Images. Diversity 2022, 14, 624. [Google Scholar] [CrossRef]
- Lyu, H.; Qiu, F.; An, L.; Stow, D.; Lewison, R.; Bohnett, E. Deer Survey from Drone Thermal Imagery Using Enhanced Faster R-CNN Based on ResNets and FPN. Ecol. Inform. 2024, 79, 102383. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar] [CrossRef]
- Lei, J.; Gao, S.; Rasool, M.A.; Fan, R.; Jia, Y.; Lei, G. Optimized Small Waterbird Detection Method Using Surveillance Videos Based on YOLOv7. Animals 2023, 13, 1929. [Google Scholar] [CrossRef]
- Ye, Q.; Ma, M.; Zhao, X.; Duan, B.; Wang, L.; Ma, D. ADD-YOLO: An Algorithm for Detecting Animals in Outdoor Environments Based on Unmanned Aerial Imagery. Measurement 2025, 242, 116019. [Google Scholar] [CrossRef]
- Yang, W.; Liu, Y.; Wang, J.; Yan, Z.; Ma, Y.; Feng, L. A Forest Wildlife Detection Algorithm Based on Improved YOLOv5s. Animals 2023, 13, 3134. [Google Scholar] [CrossRef]
- Zhu, Y.; Zhao, Y.; He, Y.; Wu, B.; Su, X. YOLO-WildASM: An Object Detection Algorithm for Protected Wildlife. Animals 2025, 15, 2699. [Google Scholar] [CrossRef] [PubMed]
- He, A.; Li, X.; Wu, X.; Su, C.; Chen, J.; Xu, S.; Guo, X. ALSS-YOLO: An Adaptive Lightweight Channel Split and Shuffling Network for TIR Wildlife Detection in UAV Imagery. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 17308–17326. [Google Scholar] [CrossRef]
- Tabak, M.A.; Norouzzadeh, M.S.; Wolfson, D.W.; Sweeney, S.J.; VerCauteren, K.C.; Snow, N.P.; Halseth, J.M.; Di Salvo, P.A.; Lewis, J.S.; White, M.D.; et al. Machine learning to classify animal species in camera trap images: Applications in ecology. Methods Ecol. Evol. 2019, 10, 585–590. [Google Scholar] [CrossRef]
- Song, Q.; Guan, Y.; Guo, X.; Guo, X.; Chen, Y.; Wang, H.; Ge, J.; Wang, T.; Bao, L. Benchmarking Wild Bird Detection in Complex Forest Scenes. Ecol. Inform. 2024, 80, 102466. [Google Scholar] [CrossRef]
- Ma, Z.; Dong, Y.; Xia, Y.; Xu, D.; Xu, F.; Chen, F. Wildlife Real-Time Detection in Complex Forest Scenes Based on YOLOv5s Deep Learning Network. Remote Sens. 2024, 16, 1350. [Google Scholar] [CrossRef]
- Beery, S.; Morris, D.; Yang, S. Efficient Pipeline for Camera Trap Image Review. arXiv 2019, arXiv:1907.06772. [Google Scholar] [CrossRef]
- Wang, T.; Feng, L.; Mou, P.; Wu, J.; Smith, J.L.D.; Xiao, W.; Yang, H.; Dou, H.; Zhao, X.; Cheng, Y.; et al. Opportunities for Amur Tiger Recovery in China. Biol. Conserv. 2018, 217, 269–279. [Google Scholar] [CrossRef]
- Wang, T.; Feng, L.; Mou, P.; Wu, J.; Smith, J.L.D.; Xiao, W.; Yang, H.; Dou, H.; Zhao, X.; Cheng, Y.; et al. Amur tigers and leopards returning to China: Direct evidence and a landscape conservation plan. Landsc. Ecol. 2016, 31, 491–503. [Google Scholar] [CrossRef]
- NCTLNP Dataset Contributors. Northeast China Tiger and Leopard National Park Wildlife Monitoring Dataset; GitHub Repository. 2023. Available online: https://github.com/myyyyw/NTLNP (accessed on 20 March 2026).
- Jocher, G.; Qiu, J. Ultralytics YOLO11; GitHub Repository: 2024. Available online: https://github.com/ultralytics/ultralytics (accessed on 20 March 2026).
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar] [CrossRef]
- Chollet, F. Xception: Deep Learning with Depthwise Separable Convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1251–1258. [Google Scholar] [CrossRef]
- Zhang, X.; Liu, C.; Song, T.; Yang, D.; Ye, Y.; Li, K.; Song, Y. RFAConv: Innovating Spatial Attention and Standard Convolutional Operation. arXiv 2023, arXiv:2304.03198. [Google Scholar] [CrossRef]
- Geng, Q.; Wan, M.; Cao, X.; Zhou, Z. Context and Spatial Feature Calibration for Real-Time Semantic Segmentation. IEEE Trans. Image Process. 2023, 32, 5465–5477. [Google Scholar] [CrossRef]
- Ouyang, D.; He, S.; Zhang, G.; Luo, Z.; Guo, H.; Zhan, J.; Huang, Z. Efficient Multi-Scale Attention Module with Cross-Spatial Learning. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; pp. 1–5. [Google Scholar] [CrossRef]
- Wan, D.; Lu, R.; Shen, S.; Xu, T.; Lang, X.; Ren, Z. MLCA: A Mixed Local-Channel Attention Module for Improving Object Detection in Complex Scenes. Eng. Appl. Artif. Intell. 2023, 123, 106442. [Google Scholar] [CrossRef]
- Ioannou, Y.; Robertson, D.; Cipolla, R.; Criminisi, A. Deep Roots: Improving CNN Efficiency with Hierarchical Filter Groups. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 5977–5986. [Google Scholar] [CrossRef]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-Time Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 16965–16974. [Google Scholar] [CrossRef]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar] [CrossRef]
- Roboflow. Wildlife Dataset. Roboflow Universe, 2023. Available online: https://universe.roboflow.com/project-vbv5j/wildlife-yj7t1 (accessed on 21 April 2026).


















| Class | Train | Valid | Test |
|---|---|---|---|
| Amur Tiger | 439 | 125 | 64 |
| Amur Leopard | 313 | 89 | 46 |
| Sika Deer | 398 | 113 | 58 |
| Wild Boar | 444 | 127 | 64 |
| Red Fox | 352 | 100 | 51 |
| Roe Deer | 443 | 126 | 64 |
| Leopard Cat | 215 | 61 | 32 |
| Badger | 261 | 74 | 39 |
| Total | 2865 | 815 | 418 |
| Parameter | Setting |
|---|---|
| Epochs | 200 |
| Patience | 50 |
| Batch size | 8 |
| Images size | 640 |
| Workers | 8 |
| Optimizer | SGD |
| Close mosaic | 10 |
| Warmup epochs | 3 |
| Initial Learning Rate | 0.01 |
| Final Learning Rate | 0.01 |
| Momentum | 0.937 |
| Weight decay | 0.0005 |
| Model | P/% | R/% | mAP@0.5/% | mAP@0.5:0.95/% | FLOPs/G | Params/M |
|---|---|---|---|---|---|---|
| YOLO11n | 85.2 | 75.8 | 83.4 | 63.9 | 6.3 | 2.58 |
| PMGHA | 86.7 | 76.4 | 85.3 | 65.2 | 6.5 | 2.59 |
| RFAConv | 86.3 | 76.2 | 85.2 | 65.1 | 6.9 | 2.69 |
| CSFCN | 87.4 | 75.9 | 85.9 | 65.8 | 7.2 | 2.96 |
| ELGH | 86.2 | 75.9 | 85.2 | 65.1 | 5.1 | 2.31 |
| PMGHA + RFAConv | 87.1 | 76.5 | 86.2 | 65.9 | 6.6 | 2.61 |
| RFAConv + CSFCN | 87.9 | 76.8 | 86.4 | 66.2 | 7.3 | 2.98 |
| PMGHA + RFAConv + CSFCN | 88.8 | 77.9 | 87.1 | 67.0 | 7.6 | 2.98 |
| PMGHA + RFAConv + CSFCN + ELGH | 89.3 | 78.4 | 87.3 | 67.3 | 6.4 | 2.77 |
| Model | Seed | P/% | R/% | mAP@0.5/% | mAP@0.5:0.95/% |
|---|---|---|---|---|---|
| RCAF-Net | 0 | 89.3 | 78.4 | 87.3 | 67.3 |
| RCAF-Net | 42 | 89.1 | 78.2 | 87.1 | 67.1 |
| RCAF-Net | 123 | 89.4 | 78.5 | 87.4 | 67.4 |
| RCAF-Net | 2024 | 89.2 | 78.3 | 87.2 | 67.2 |
| RCAF-Net | 999 | 89.3 | 78.6 | 87.3 | 67.4 |
| RCAF-Net | Mean | 89.26 | 78.40 | 87.26 | 67.28 |
| RCAF-Net | Std | 0.11 | 0.14 | 0.11 | 0.12 |
| Model | P/% | R/% | mAP@0.5/% | mAP@0.5:0.95/% | FLOPs/G | Params/M |
|---|---|---|---|---|---|---|
| YOLOv5n | 80.8 | 76.4 | 82.3 | 62.8 | 7.1 | 2.50 |
| YOLOv8n | 88.2 | 73.4 | 81.8 | 63.3 | 8.1 | 3.00 |
| YOLOv9t | 85.8 | 73.5 | 83.1 | 63.6 | 7.4 | 3.05 |
| YOLOv10n | 88.9 | 73.1 | 82.1 | 63.2 | 8.2 | 2.70 |
| YOLO11n | 85.2 | 75.8 | 83.4 | 63.9 | 6.3 | 2.58 |
| YOLOv12n | 86.3 | 73.6 | 83.3 | 65.9 | 5.8 | 2.50 |
| YOLOv13n | 83.3 | 78.9 | 84.4 | 65 | 6.1 | 2.45 |
| Faster R-CNN | 83.8 | 72.2 | 82.1 | 61.8 | 208.1 | 41.40 |
| RT-DETR | 84.7 | 71.6 | 82.9 | 64.1 | 59.2 | 20.10 |
| RCAF-Net | 89.3 | 78.4 | 87.3 | 67.3 | 6.4 | 2.77 |
| Model | P/% | R/% | mAP@0.5/% | mAP@0.5:0.95/% |
|---|---|---|---|---|
| YOLO11n | 86.25 | 79.84 | 81.75 | 44.68 |
| RCAF-Net | 88.22 | 79.03 | 87.09 | 47.46 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yu, X.; Qu, C.; Xu, Y.; Guo, S.; Fu, L.; Zhou, Y. RCAF-Net: Wildlife Target Detection in Complex Forest Scenarios. Animals 2026, 16, 1484. https://doi.org/10.3390/ani16101484
Yu X, Qu C, Xu Y, Guo S, Fu L, Zhou Y. RCAF-Net: Wildlife Target Detection in Complex Forest Scenarios. Animals. 2026; 16(10):1484. https://doi.org/10.3390/ani16101484
Chicago/Turabian StyleYu, Xiuling, Chenxiao Qu, Yifu Xu, Senyue Guo, Lili Fu, and Yang Zhou. 2026. "RCAF-Net: Wildlife Target Detection in Complex Forest Scenarios" Animals 16, no. 10: 1484. https://doi.org/10.3390/ani16101484
APA StyleYu, X., Qu, C., Xu, Y., Guo, S., Fu, L., & Zhou, Y. (2026). RCAF-Net: Wildlife Target Detection in Complex Forest Scenarios. Animals, 16(10), 1484. https://doi.org/10.3390/ani16101484

