A Novel Object Detection Algorithm Combined YOLOv11 with Dual-Encoder Feature Aggregation
Abstract
1. Introduction
- We propose an improved dual-branch RGB-D fusion object detection framework based on YOLOv11. By designing a symmetric network structure, RGB images and depth maps are processed in parallel. This enables the model to simultaneously leverage the rich texture and semantic information from RGB data and the geometric structural information provided by depth data, significantly enhancing detection performance in complex environments such as low illumination, occlusion, and texture-sparse scenarios.
- We design a Dynamic Branch Enhancement (DBE) module that incorporates a combined spatial-channel attention mechanism. This module achieves adaptive weighting of cross-modal features, dynamically enhancing complementary information while suppressing redundant or noisy feature activations, thereby improving the model’s robustness and feature representation capability.
- We introduce a Dual-Encoder Attention (DEA) module, which consists of two sub-modules: the Dual-Encoder Cross-Attention (DECA) for refining features within each modality, and the Dual-Encoder Feature Aggregation (DEPA) for effectively integrating cross-modal information. This hierarchical design enables multi-scale feature fusion, strengthening the model’s ability to perceive and represent multi-modal information.
- We employ the Depth Anything model to generate pseudo-depth maps from single RGB images, constructing a pseudo-RGBD training set. This approach effectively addresses the limitation of lacking real depth sensor data in conventional visible-light datasets, providing a feasible solution for training and evaluating RGB-D fusion detection models under monocular data constraints.
2. Methods
2.1. Dual-Encoder Attention
2.1.1. Dual-Encoder Cross-Attention
2.1.2. Dual-Encoder Feature Aggregation
2.2. Loss Function
3. Results
3.1. Experimental Configuration and Model Evaluation Metrics
3.2. Evaluation Indicators
3.3. Introduction to the Dataset
3.3.1. M3FD Dataset
3.3.2. VOC2007 Dataset
3.4. Comparison Experiment
3.5. Ablation Experiment
3.6. Analysis of Test Results
4. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Hu, Q.; Wang, K.; Ren, F.; Wang, Z. Research on underwater robot ranging technology based on semantic segmentation and binocular vision. Sci. Rep. 2024, 14, 12309. [Google Scholar] [CrossRef] [PubMed]
- Yaqoob, M.; Ishaq, M.; Ansari, M.Y.; Qaiser, Y.; Hussain, R.; Rabbani, H.S.; Garwood, R.J.; Seers, T.D. Advancing paleontology: A survey on deep learning methodologies in fossil image analysis. Artif. Intell. Rev. 2025, 58, 83. [Google Scholar] [CrossRef]
- Sun, W.; Lin, X.; Shi, Y.; Zhang, C.; Wu, H.; Zheng, S. Sparsedrive: End-to-end autonomous driving via sparse scene representation. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA), Atlanta, GA, USA, 19–23 May 2025; IEEE: New York, NY, USA, 2025; pp. 8795–8801. [Google Scholar]
- Chen, W.; Chi, W.; Ji, S.; Ye, H.; Liu, J.; Jia, Y.; Yu, J.; Cheng, J. A survey of autonomous robots and multi-robot navigation: Perception, planning and collaboration. Biomim. Intell. Robot. 2025, 5, 100203. [Google Scholar] [CrossRef]
- Zhang, A.; Song, R.; Cui, H.; Xu, C.; Zhou, C. Visual and semantic fusion fault identification algorithm for power inspection robot. In Proceedings of the 2025 IEEE 5th International Conference on Electronic Technology, Communication and Information (ICETCI), Changchun, China, 23–25 May 2025; IEEE: New York, NY, USA, 2025; pp. 414–422. [Google Scholar]
- Zhou, T.; Fan, D.-P.; Cheng, M.-M.; Shen, J.; Shao, L. RGB-D salient object detection: A survey. Comput. Vis. Media 2021, 7, 37–69. [Google Scholar] [CrossRef] [PubMed]
- Peng, H.W.; Li, B.; Xiong, W.H.; Hu, W.M.; Ji, R.R. RGBD salient object detection: A benchmark and algorithms. In Computer Vision–ECCV 2014; Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2014; Volume 8691, pp. 92–109. [Google Scholar]
- Han, J.W.; Chen, H.; Liu, N.; Yan, C.G.; Li, X.L. CNNs-based RGB-D saliency detection via cross-view transfer and multiview fusion. IEEE Trans. Cybern. 2018, 48, 3171–3183. [Google Scholar] [CrossRef] [PubMed]
- Li, C.; Cong, R.; Kwong, S.; Hou, J.; Fu, H.; Zhu, G.; Zhang, D.; Huang, Q. ASIF-Net: Attention steered interweave fusion network for RGB-D salient object detection. IEEE Trans. Cybern. 2021, 51, 88–100. [Google Scholar] [CrossRef] [PubMed]
- Chen, H.; Li, Y. Progressively complementarity-aware fusion network for RGB-D salient object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 3051–3060. [Google Scholar]
- Chen, H.; Li, Y.-F.; Su, D. M3Net: Multi-scale multi-path multi-modal fusion network and example application to RGB-D salient object detection. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada, 24–28 September 2017; pp. 4911–4916. [Google Scholar]
- Liu, N.; Zhang, N.; Wan, K.; Shao, L.; Han, J. Visual saliency transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 4722–4732. [Google Scholar]
- Chen, G.; Fu, H.; Zhou, T.; Xiao, G.; Fu, K.; Zhang, Y. EM-Trans: Edge-aware multimodal transformer for RGB-D salient object detection. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 3180–3188. [Google Scholar] [CrossRef] [PubMed]
- Pan, C.; Jiang, Q.; Zheng, H.; Wang, P.; Jin, X.; Yao, S.; Zhou, W. DANet: A Dual-Branch Framework With Diffusion-Integrated Autoencoder for Infrared-Visible Image Fusion. IEEE Trans. Instrum. Meas. 2025, 74, 5031713. [Google Scholar] [CrossRef]
- Li, J.; Liu, L.; Song, H.; Huang, Y.; Jiang, J.; Yang, J. DCTNet: A heterogeneous dual-branch multi-cascade network for infrared and visible image fusion. IEEE Trans. Instrum. Meas. 2023, 72, 5030914. [Google Scholar] [CrossRef]
- Hu, X.; Liu, Y.; Yang, F. PFCFuse: A Poolformer and CNN fusion network for infrared-visible image fusion. IEEE Trans. Instrum. Meas. 2024, 73, 5029714. [Google Scholar] [CrossRef]
- Huang, Z.; Lin, C.; Xu, B.; Xia, M.; Li, Q.; Li, Y.; Sang, N. T2EA: Target-Aware Taylor Expansion Approximation Network for Infrared and Visible Image Fusion. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 4831–4845. [Google Scholar] [CrossRef]
- Chen, J.; Ding, J.; Ma, J. HitFusion: Infrared and visible image fusion for high-level vision tasks using transformer. IEEE Trans. Multimed. 2024, 26, 10145–10159. [Google Scholar] [CrossRef]
- Liu, H.; Mao, Q.; Dong, M.; Zhan, Y. Infrared-visible Image Fusion Using Dual-branch Auto-encoder with Invertible High-frequency Encoding. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 2675–2688. [Google Scholar] [CrossRef]
- Li, J.; Bai, L.; Yang, B.; Li, C.; Ma, L.; Cui, L.; Hancock, E.R. Dual-modal prior semantic guided infrared and visible image fusion for intelligent transportation system. IEEE Trans. Intell. Transp. Syst. 2025, 26, 9767–9780. [Google Scholar] [CrossRef]
- Murat, A.A.; Kiran, M.S. A comprehensive review on YOLO versions for object detection. Eng. Sci. Technol. Int. J. 2025, 70, 102161. [Google Scholar] [CrossRef]
- Du, H.; Ren, L.; Wang, Y.; Cao, X.; Sun, C. Advancements in perception system with multi-sensor fusion for embodied agents. Inf. Fusion 2025, 117, 102859. [Google Scholar] [CrossRef]
- Zeng, Y. Research on Obstacle Avoidance Strategies for Intelligent Connected Vehicles Based on Multi-Sensor Fusion and YOLOv5 Algorithm. In Proceedings of the 2025 IEEE the 5th International Conference on Computer Communication and Artificial Intelligence, Haikou, China, 23–25 May 2025; pp. 490–494. [Google Scholar]
- Song, Q.; Sun, S.; Wang, B.; Song, Q.; Wang, T.; Jiang, H. Noise-robust fault diagnosis network based on multiscale feature enhancement and dynamic cross-modal interaction. Measurement 2025, 257, 118564. [Google Scholar] [CrossRef]
- Wu, Y.; Mu, X.; Shi, H.; Hou, M. An object detection model AAPW-YOLO for UAV remote sensing images based on adaptive convolution and reconstructed feature fusion. Sci. Rep. 2025, 15, 16214. [Google Scholar] [CrossRef] [PubMed]
- Everingham, M.; Van Gool, L.; Williams, C.K.I.; Winn, J.M.; Zisserman, A. The Pascal Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303. [Google Scholar] [CrossRef]
- Liu, J.; Fan, X.; Huang, Z.; Wu, G.; Liu, R.; Zhong, W.; Luo, Z. Target-aware Dual Adversarial Learning and a Multi-scenario Multi-Modality Benchmark to Fuse Infrared and Visible for Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022. [Google Scholar]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar]
















| Bibliographic Citation | Description of Method | Advantages |
|---|---|---|
| Zhou et al. [6] | Conduct a comprehensive investigation on RGB-D salient object detection and systematically classify traditional and deep learning methods. | Emphasize the complementarity between RGB and deep modalities to provide benchmark datasets and classification frameworks for fusion research. |
| Peng et al. [7] | The early fusion strategy involves simply concatenating RGB and depth data for fusion at the input or feature layer. | It has low computational overhead, simple processing, and is suitable for resource-constrained scenarios. |
| Han et al. [8] | The late fusion method independently processes RGB and deep modes and then combines the prediction results. | Retain modal-specific features to enhance robustness against noise. |
| Li et al. [9] | The ASIF-Net network adopts an interactive fusion scheme and an attention mechanism to dynamically weight cross-modal features. | Enhance feature representation through attention weighting to improve accuracy under low-light and occlusion conditions. |
| Chen and Li [10] | The PCF network uses complementary perception fusion modules for cross-level interaction. | Explicitly model modal dependencies, reduce fusion ambiguity, and take into account both low-level details and high-level semantics. |
| Liu et al. [12] | Visual Saliency Transformer (VST) applies the self-attention mechanism to capture long-range dependencies. | Handle the global context of large-scale scenes and enhance performance in complex backgrounds. |
| Model | Precision (%) | Recall (%) | mAP50 (%) | mAP50-95 (%) | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|
| YOLOv5n | 83.77 | 64.57 | 73.12 | 46.37 | 2.50 | 7.1 |
| YOLOv8n | 83.63 | 68.92 | 74.91 | 51.40 | 3.01 | 8.1 |
| YOLOv10n | 76.84 | 70.15 | 76.04 | 47.14 | 2.27 | 6.5 |
| YOLOv11n | 80.82 | 65.57 | 71.89 | 50.37 | 2.58 | 6.3 |
| YOLOv12n | 80.79 | 64.38 | 70.88 | 44.42 | 2.51 | 5.8 |
| YOLOv13n | 73.88 | 55.61 | 61.94 | 38.46 | 2.45 | 6.2 |
| RT-DETR | 87.24 | 80.98 | 81.20 | 55.00 | 19.88 | 57.0 |
| TarDAL (RGB + IR) | - | - | 80.70 | - | - | 14.9 |
| MMDetection-DS (RGB + Depth) | 85.01 | 66.19 | 75.38 | 50.20 | 2.51 | 19.8 |
| Ours (RGB + IR) | 87.03 | 73.40 | 81.14 | 54.69 | 2.47 | 19.3 |
| Ours (RGB + Depth) | 88.91 | 68.42 | 77.37 | 51.86 | 2.47 | 19.3 |
| Model | Precision (%) | Recall (%) | mAP50 (%) | mAP50-95 (%) | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|
| YOLOv5n | 64.89 | 51.72 | 56.40 | 35.66 | 2.19 | 5.8 |
| YOLOv8n | 78.13 | 70.83 | 77.21 | 57.81 | 2.69 | 6.8 |
| YOLOv10n | 69.35 | 56.18 | 62.46 | 42.69 | 2.27 | 6.5 |
| YOLOv11n | 78.85 | 69.49 | 77.53 | 68.35 | 2.59 | 6.3 |
| YOLOv12n | 80.04 | 69.25 | 77.17 | 58.38 | 2.51 | 5.8 |
| YOLOv13n | 66.88 | 53.13 | 58.55 | 40.28 | 2.45 | 6.2 |
| RT-DETR | 81.83 | 70.57 | 75.88 | 58.03 | 19.90 | 57.0 |
| Ours | 81.22 | 75.61 | 82.59 | 65.27 | 2.47 | 19.4 |
| Class | YOLOv5n | YOLOv8n | YOLOv10n | YOLOv11n | YOLOv12n | YOLOv13n | RT-DETR | Ours (RGB + IR) | Ours (RGB + Depth) |
|---|---|---|---|---|---|---|---|---|---|
| People | 66.19 | 67.74 | 68.24 | 66.05 | 64.61 | 58.00 | 75.61 | 81.60 | 69.73 |
| Car | 87.72 | 87.92 | 88.56 | 87.15 | 86.46 | 79.92 | 87.82 | 89.89 | 89.61 |
| Bus | 84.63 | 85.60 | 84.06 | 84.24 | 84.50 | 75.14 | 82.75 | 87.41 | 88.01 |
| Motorcycle | 63.69 | 65.56 | 67.94 | 63.60 | 59.73 | 48.37 | 85.48 | 75.42 | 75.42 |
| Lamp | 63.47 | 70.02 | 74.24 | 58.38 | 60.60 | 50.92 | 79.74 | 74.76 | 69.73 |
| Truck | 73.04 | 72.60 | 73.18 | 71.92 | 69.36 | 59.27 | 75.78 | 77.68 | 71.77 |
| All | 73.12 | 74.91 | 76.04 | 71.89 | 70.88 | 61.94 | 81.20 | 81.14 | 77.37 |
| Class | YOLOv5n | YOLOv8n | YOLOv10n | YOLOv11n | YOLOv12n | YOLOv13n | RT-DETR | Ours |
|---|---|---|---|---|---|---|---|---|
| Aeroplane | 70.13 | 84.95 | 75.27 | 86.75 | 84.46 | 70.13 | 86.65 | 88.73 |
| Bicycle | 67.38 | 84.10 | 75.80 | 86.13 | 87.42 | 70.53 | 81.63 | 89.71 |
| Bird | 45.27 | 79.20 | 56.02 | 79.69 | 80.09 | 52.99 | 80.49 | 83.65 |
| Boat | 35.04 | 53.31 | 38.15 | 55.29 | 53.82 | 34.85 | 52.43 | 61.01 |
| Bottle | 28.34 | 57.06 | 38.71 | 59.80 | 53.89 | 29.03 | 65.14 | 75.42 |
| Bus | 63.77 | 82.03 | 73.30 | 80.53 | 81.62 | 69.05 | 80.59 | 86.53 |
| Car | 72.85 | 87.92 | 76.95 | 87.01 | 86.71 | 73.37 | 88.48 | 88.02 |
| Cat | 67.37 | 86.29 | 72.55 | 86.30 | 88.69 | 65.54 | 80.19 | 93.03 |
| Chair | 37.59 | 61.30 | 42.22 | 59.97 | 60.14 | 34.47 | 61.17 | 70.16 |
| Cow | 50.02 | 87.11 | 53.92 | 83.91 | 82.34 | 49.18 | 80.25 | 88.72 |
| Diningtable | 59.10 | 71.04 | 65.57 | 72.46 | 74.78 | 62.84 | 74.05 | 82.73 |
| Dog | 50.15 | 82.01 | 59.06 | 78.12 | 78.82 | 53.49 | 75.76 | 85.70 |
| Horse | 76.27 | 91.85 | 78.41 | 92.56 | 90.24 | 75.83 | 82.22 | 94.03 |
| Motorbike | 76.65 | 89.20 | 78.06 | 90.34 | 90.32 | 78.49 | 89.00 | 92.92 |
| Person | 75.31 | 88.41 | 78.36 | 87.62 | 89.02 | 76.15 | 87.46 | 89.55 |
| Pottedplant | 27.27 | 44.51 | 32.76 | 51.70 | 49.82 | 26.93 | 44.45 | 49.71 |
| Sheep | 54.08 | 78.61 | 62.51 | 78.55 | 81.03 | 61.70 | 76.05 | 84.81 |
| Sofa | 47.95 | 70.75 | 52.96 | 69.73 | 69.73 | 49.84 | 69.18 | 75.16 |
| Train | 66.41 | 86.83 | 75.05 | 86.85 | 84.22 | 76.99 | 86.28 | 87.62 |
| Tvmonitor | 57.03 | 77.67 | 63.62 | 77.18 | 76.34 | 59.53 | 74.29 | 83.40 |
| All | 56.40 | 77.21 | 62.46 | 77.53 | 77.17 | 58.55 | 75.88 | 82.59 |
| Model | Precision (%) | Recall (%) | mAP50 (%) | mAP50-95 (%) | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|
| YOLOv11n | 78.85 | 69.49 | 77.53 | 68.35 | 2.59 | 6.3 |
| YOLOv11n + A | 79.10 | 69.53 | 78.03 | 68.91 | 2.59 | 6.3 |
| YOLOv11n + B | 80.97 | 75.03 | 82.10 | 64.30 | 2.47 | 19.4 |
| YOLOv11n + A + B | 81.22 | 75.61 | 82.59 | 65.27 | 2.47 | 19.4 |
| Modal | mAP50 (%) | mAP50-95 (%) | Params (M) | GFLOPs | ||
|---|---|---|---|---|---|---|
| Visible | Depth | Infrared | ||||
| √ | 71.89 | 50.37 | 2.58 | 6.3 | ||
| √ | 67.23 | 43.60 | 2.58 | 6.3 | ||
| √ | 68.70 | 44.38 | 2.58 | 6.3 | ||
| √ | √ | 77.37 | 51.86 | 2.47 | 19.3 | |
| √ | √ | 81.14 | 54.69 | 2.47 | 19.3 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Chen, H.; Yuan, P.; Liu, W.; Li, F.; Wang, A. A Novel Object Detection Algorithm Combined YOLOv11 with Dual-Encoder Feature Aggregation. Sensors 2025, 25, 7270. https://doi.org/10.3390/s25237270
Chen H, Yuan P, Liu W, Li F, Wang A. A Novel Object Detection Algorithm Combined YOLOv11 with Dual-Encoder Feature Aggregation. Sensors. 2025; 25(23):7270. https://doi.org/10.3390/s25237270
Chicago/Turabian StyleChen, Haisong, Pengfei Yuan, Wenbai Liu, Fuling Li, and Aili Wang. 2025. "A Novel Object Detection Algorithm Combined YOLOv11 with Dual-Encoder Feature Aggregation" Sensors 25, no. 23: 7270. https://doi.org/10.3390/s25237270
APA StyleChen, H., Yuan, P., Liu, W., Li, F., & Wang, A. (2025). A Novel Object Detection Algorithm Combined YOLOv11 with Dual-Encoder Feature Aggregation. Sensors, 25(23), 7270. https://doi.org/10.3390/s25237270

