DTNet: A Novel Infrared-Visible Fruit Object Detection Method Based on Dual-Modal Feature Interaction Fusion and Transformer Decoding
Abstract
1. Introduction
1.1. Background and Significance
1.2. Related Work
1.2.1. Fruit Object Detection Based on RGB Images
1.2.2. Applications of Multispectral and Multimodal Imaging in Agricultural Perception
1.2.3. Multimodal Detection Models
1.3. Main Contributions
2. Materials and Methods
2.1. Experimental Materials and Dataset Description
2.1.1. Grape RGB-IR Dataset
2.1.2. Tomato RGB-IR Dataset
2.1.3. Annotation, Data Selection, and Split Strategy
2.2. Overall Network Architecture
2.3. RGB-IR Dual-Branch Backbone and Dual-Modal Fusion Block (DFB)
2.3.1. Dual-Branch Feature Extraction Backbone
2.3.2. Dual-Modal Fusion Block (DFB)
- (1)
- Initial Fusion and Context Modeling
- (2)
- Pixel-Wise Adaptive Weight Estimation
- (3)
- Modality-Adaptive Fusion and Linear Remapping
2.4. Local Enhancement Module (LEM)
2.4.1. Overall Structure and Residual Learning
2.4.2. Adaptive Normalization Module (LayNorm) and Recalibration
2.4.3. Window Attention (WATT): Local Context Modeling and Relative Positional Bias
2.4.4. Attention Output Projection and Feature Fusion
2.4.5. Channel-Wise Nonlinear Mapping and Second-Stage Residual Enhancement
2.5. Adaptive Spatial Attention Module (ASA)
2.6. Transformer-Based Adaptive Demodulation Detection Head
2.7. Multimodal Balanced Detection Loss (MBDL)
2.7.1. Balanced Classification Loss
2.7.2. PioU v2 Localization Loss
2.7.3. Classification-Localization Consistency Constraint
3. Results
3.1. Experimental Setup and Evaluation Metrics
3.1.1. Evaluation Metrics
3.1.2. Implementation Details and Training Configuration
3.1.3. Edge-Side Deployment Evaluation on Raspberry Pi 5
3.2. Baseline Selection
3.3. Comparative Experiments and Result Analysis
Quantitative Comparison with Mainstream Methods
3.4. Ablation Study
Ablation Results
3.5. Evaluation on an Additional Tomato RGB-IR Dataset
3.5.1. Results on the Tomato RGB-IR Dataset
3.5.2. Analysis of Results on the Additional Dataset
3.6. Qualitative Visualization Comparison in Complex Scenarios
4. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Koirala, A.; Walsh, K.B.; Wang, Z.; McCarthy, C. Deep learning for real-time fruit detection and orchard fruit load estimation: Benchmarking of “MangoYOLO”. Precis. Agric. 2019, 20, 1107–1135. [Google Scholar] [CrossRef]
- Xiao, B.J.; Nguyen, M.; Yan, W.Q. Fruit ripeness identification using YOLOv8 model. Multimed. Tools Appl. 2024, 83, 28039–28056. [Google Scholar] [CrossRef]
- Sun, H.; Wang, B.; Xue, J. YOLO-P: An efficient method for pear fast detection in complex orchard picking environment. Front. Plant Sci. 2023, 13, 1089454. [Google Scholar] [CrossRef]
- Lin, T.; Sun, F.; Li, X.; Guo, X.; Ying, J.; Wu, H.; Li, H. A review of key technologies and recent advances in intelligent fruit-picking robots. Horticulturae 2026, 12, 158. [Google Scholar] [CrossRef]
- Zhang, Y.; Li, N.; Zhang, L.; Lin, J.; Gao, X.; Chen, G. A review on the recent developments in vision-based apple-harvesting robots for recognizing fruit and picking pose. Comput. Electron. Agric. 2025, 231, 109968. [Google Scholar] [CrossRef]
- Shi, X.; Wang, S.; Zhang, B.; Ding, X.; Qi, P.; Qu, H.; Li, N.; Wu, J.; Yang, H. Advances in object detection and localization techniques for fruit harvesting robots. Agronomy 2025, 15, 145. [Google Scholar] [CrossRef]
- Onishi, Y.; Yoshida, T.; Kurita, H.; Fukao, T.; Arihara, H.; Iwai, A. An automated fruit harvesting robot by using deep learning. Robomech J. 2019, 6, 13. [Google Scholar] [CrossRef]
- Guo, C.; Zheng, S.; Cheng, G.; Zhang, Y.; Ding, J. An improved YOLOv4 used for grape detection in unstructured environment. Front. Plant Sci. 2023, 14, 1209910. [Google Scholar] [CrossRef]
- Liu, G.; Hou, Z.; Liu, H.; Liu, J.; Zhao, W.; Li, K. TomatoDet: Anchor-free detector for tomato detection. Front. Plant Sci. 2022, 13, 942875. [Google Scholar] [CrossRef]
- Pinheiro, I.; Moreira, G.; da Silva, D.Q.; Magalhães, S.; Valente, A.; Oliveira, P.M.; Cunha, M.; Santos, F. Deep learning YOLO-based solution for grape bunch detection and assessment of biophysical lesions. Agronomy 2023, 13, 1120. [Google Scholar] [CrossRef]
- Murat, A.A.; Kiran, M.S. A comprehensive review on YOLO versions for object detection. Eng. Sci. Technol. Int. J. 2025, 70, 102161. [Google Scholar] [CrossRef]
- Pagire, V.; Chavali, M.; Kale, A. A comprehensive review of object detection with traditional and deep learning methods. Signal Process. 2025, 237, 110250. [Google Scholar] [CrossRef]
- Mbouembe, P.L.T.; Liu, G.; Sikati, J.; Kim, S.C.; Kim, J.H. An efficient tomato-detection method based on improved YOLOv4-tiny model in complex environment. Front. Plant Sci. 2023, 14, 1150958. [Google Scholar] [CrossRef] [PubMed]
- De Silva, M.; Brown, D. Multispectral plant disease detection with vision transformer-convolutional neural network hybrid approaches. Sensors 2023, 23, 8531. [Google Scholar] [CrossRef]
- Nguyen, C.; Moghadam, P.; Ward, B.; Miller, J.; O’Halloran, K.; Hernandez, E.; Salisbury, J.; Able, A.J. Early detection of plant viral disease using hyperspectral imaging and deep learning. Sensors 2021, 21, 742. [Google Scholar] [CrossRef] [PubMed]
- Xiang, Y.; Chen, Q.; Su, Z.; Zhang, L.; Chen, Z.; Zhou, G.; Yao, Z.; Xuan, Q.; Cheng, Y. Deep learning and hyperspectral images based tomato soluble solids content and firmness estimation. Front. Plant Sci. 2022, 13, 860656. [Google Scholar] [CrossRef]
- Barros, T.; Conde, P.; Gonçalves, G.; Premebida, C.; Monteiro, M.; Ferreira, C.S.; Henriques, R. Multispectral vineyard segmentation: A deep learning comparison study. Comput. Electron. Agric. 2022, 195, 106782. [Google Scholar] [CrossRef]
- Su, S.; Chen, R.; Fang, X.; Zhu, Y.; Zhang, T.; Xu, Z. A novel lightweight grape detection method. Agriculture 2022, 12, 1364. [Google Scholar] [CrossRef]
- Yang, W.; Qiu, X. A lightweight and efficient model for grape bunch detection and biophysical anomaly assessment in complex environments based on YOLOv8s. Front. Plant Sci. 2024, 15, 1395796. [Google Scholar] [CrossRef]
- Wu, X.; Tang, R.; Mu, J.; Niu, Y.; Xu, Z.; Chen, Z. A lightweight grape detection model in natural environments based on an enhanced YOLOv8 framework. Front. Plant Sci. 2024, 15, 1407839. [Google Scholar] [CrossRef]
- Shuai, L.; Li, Z.; Chen, Z.; Luo, D.; Mu, J. A research review on deep learning combined with hyperspectral imaging in multiscale agricultural sensing. Comput. Electron. Agric. 2024, 217, 108577. [Google Scholar] [CrossRef]
- Zhu, X.; Yu, Z.; Li, C. GrapeUL-YOLO: Bidirectional cross-scale fusion with elliptical anchors for robust grape detection in orchards. Front. Plant Sci. 2025, 16, 1701817. [Google Scholar] [CrossRef] [PubMed]
- Gao, F.; Fang, W.; Sun, X.; Wu, Z.; Zhao, G.; Li, G.; Li, R.; Fu, L.; Zhang, Q. A novel apple fruit detection and counting methodology based on deep learning and trunk tracking in modern orchard. Comput. Electron. Agric. 2022, 197, 107000. [Google Scholar] [CrossRef]
- Ji, W.; Huang, X.; Wang, S.; He, X. A comprehensive review of the research of the “eye-brain-hand” harvesting system in smart agriculture. Agronomy 2023, 13, 2237. [Google Scholar] [CrossRef]
- Tu, S.; Huang, Y.; Huang, Q.; Liu, H.; Cai, Y.; Lei, H. Estimation of passion fruit yield based on YOLOv8n + OC-SORT + CRCM algorithm. Comput. Electron. Agric. 2025, 229, 109727. [Google Scholar] [CrossRef]
- De Silva, M.; Brown, D. Tomato disease detection using multispectral imaging with deep learning models. In Proceedings of the 7th International Conference on Artificial Intelligence, Big Data, Computing and Data Communication Systems (ICABCD 2024), Durban, South Africa, 16–18 May 2024. [Google Scholar]
- Seiche, A.T.; Wittstruck, L.; Jarmer, T. Weed detection from unmanned aerial vehicle imagery using deep learning—A comparison between high-end and low-cost multispectral sensors. Sensors 2024, 24, 1544. [Google Scholar] [CrossRef] [PubMed]
- Baltrusaitis, T.; Ahuja, C.; Morency, L.-P. Multimodal machine learning: A survey and taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 423–443. [Google Scholar] [CrossRef]
- Zhao, F.; Zhang, C.; Geng, B. Deep multimodal data fusion. ACM Comput. Surv. 2024, 56, 216. [Google Scholar] [CrossRef]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Computer Vision—ECCV 2020, Part I; Vedaldi, A., Bischof, H., Brox, T., Frahm, J.-M., Eds.; Springer: Cham, Switzerland, 2020; pp. 213–229. [Google Scholar]
- Zhao, Y.; Wong, W.; Bénière, R.; Ye, J.; de Charette, R.; Chavdarova, T. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 16965–16974. [Google Scholar]
- Chiu, M.T.; Xu, X.; Wang, Y.; Wei, J.; Huang, Z.; Schwing, A.G.; Brunner, R.; Khachatrian, H.; Karapetyan, H. Agriculture-Vision: A large aerial image database for agricultural pattern analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 2828–2838. [Google Scholar]
- Xu, K.; Zhu, Y.; Cao, W.; Jiang, X.; Jiang, Z.; Li, S.; Ni, J. Multi-modal deep learning for weeds detection in wheat field based on RGB-D images. Front. Plant Sci. 2021, 12, 732968. [Google Scholar] [CrossRef]
- Wu, Y.; Zhao, C.; Hou, Y.; Li, Y.; Liu, H.; Yang, Y.; He, Y. An RGB-D object detection model with high-generalization ability applied to tea harvesting robot for outdoor cross-variety tea shoots detection. J. Field Robot. 2024, 41, 1167–1186. [Google Scholar] [CrossRef]
- Loghmani, M.R.; Khosravi, H.; Soryani, M.; Nezamabadi-pour, H. Recurrent convolutional fusion for RGB-D object recognition. IEEE Robot. Autom. Lett. 2019, 4, 2878–2885. [Google Scholar] [CrossRef]
- Villacrés, J.; Viscaíno, M.; Delpiano, J.; Vougioukas, S.; Auat Cheein, F. Apple orchard production estimation using deep learning strategies: A comparison of tracking-by-detection algorithms. Comput. Electron. Agric. 2023, 204, 107513. [Google Scholar] [CrossRef]
- Xu, H.; Li, H.; Zhao, J. A lightweight tri-modal few-shot detection framework for fruit diversity recognition toward digital orchard archiving. Front. Plant Sci. 2025, 16, 1696622. [Google Scholar] [CrossRef] [PubMed]
- Sahin, H.M.; Miftahushudur, T.; Grieve, B.; Yin, H. Segmentation of weeds and crops using multispectral imaging and CRF-enhanced U-Net. Comput. Electron. Agric. 2023, 211, 107956. [Google Scholar] [CrossRef]
- Li, K.; Wei, X.; Wang, Q.; Zhang, W. Research on strawberry visual recognition and 3D localization based on lightweight RAFS-YOLO and RGB-D camera. Agriculture 2025, 15, 2212. [Google Scholar] [CrossRef]
- Chen, W.; Rao, Y.; Wang, F.; Zhang, Y.; Yang, Y.; Luo, Q.; Zhang, T.; Wan, T.; Liu, X.; Zhang, M.; et al. A dataset of grape multimodal object detection and semantic segmentation. China Sci. Data 2023, 10, 89–104. [Google Scholar] [CrossRef]
- Zhang, Y.; Rao, Y.; Chen, W.; Hou, W.; Yan, S.; Li, Y.; Zhou, C.; Wang, F.; Chu, Y.; Shi, Y. A dataset of multimodal images of tomato fruits at different stages of maturity. China Sci. Data 2025, 10, 73–88. [Google Scholar] [CrossRef]
- Lei, M.; Li, S.; Wu, Y.; Hu, H.; Zhou, Y.; Zheng, X.; Ding, G.; Du, S.; Wu, Z.; Gao, Y. YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception. arXiv 2025, arXiv:2506.17733. [Google Scholar]
- Sapkota, R.; Meng, Z.; Churuvija, M.; Du, X.; Ma, Z.; Karkee, M. Comprehensive performance evaluation of YOLOv12, YOLO11, YOLOv10, YOLOv9 and YOLOv8 on detecting and counting fruitlet in complex orchard environments. Agric. Commun. 2026, 4, 100125. [Google Scholar] [CrossRef]
- Li, G.; Fang, J. LCW-YOLO: A lightweight multi-scale object detection method based on YOLOv11 and its performance evaluation in complex natural scenes. Sensors 2025, 25, 6209. [Google Scholar] [CrossRef]
- Tang, L.; Zhang, H.; Xu, H.; Ma, J. Rethinking the necessity of image fusion in high-level vision tasks: A practical infrared and visible image fusion network based on progressive semantic injection and scene fidelity. Inf. Fusion 2023, 99, 101870. [Google Scholar] [CrossRef]
- Shen, J.; Chen, Y.; Liu, Y.; Zuo, X.; Fan, H.; Yang, W. ICAFusion: Iterative cross-attention guided feature fusion for multispectral object detection. Pattern Recognit. 2024, 145, 109913. [Google Scholar] [CrossRef]
- Ghorbani Kolahi, S.; Chaharsooghi, S.K.; Khatibi, T.; Bozorgpour, A.; Azad, R.; Heidari, M.; Hacihaliloglu, I.; Merhof, D. MSA^2Net: Multi-scale Adaptive Attention-guided Network for Medical Image Segmentation. In Proceedings of the 35th British Machine Vision Conference (BMVC 2024), Glasgow, UK, 25–28 November 2024. [Google Scholar]
- Chen, Z.; He, Z.; Lu, Z.-M. DEA-Net: Single image dehazing based on detail-enhanced convolution and content-guided attention. IEEE Trans. Image Process. 2024, 33, 1002–1015. [Google Scholar] [CrossRef]
- Zhang, Y.; Zhou, S.; Li, H. Depth Information Assisted Collaborative Mutual Promotion Network for Single Image Dehazing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 2846–2855. [Google Scholar]
- Zhang, J.; Li, X.; Li, J.; Liu, L.; Xue, Z.; Zhang, B.; Jiang, Z.; Huang, T.; Wang, Y.; Wang, C. Rethinking Mobile Block for Efficient Attention-based Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 1389–1400. [Google Scholar]
- Wang, X.; Wen, X.; Li, Y.; Du, C.; Zhang, D.; Sun, C.; Chen, B. A precise detection method for tomato fruit ripeness and picking points in complex environments. Horticulturae 2025, 11, 585. [Google Scholar] [CrossRef]









| Model | Precision | Recall | mAP@0.5 | mAP@0.5:0.95 | Params | FLOPs |
|---|---|---|---|---|---|---|
| YOLOv8 | 0.8769 | 0.8024 | 0.8954 | 0.6598 | 2,690,598 | 6.9 G |
| YOLOv11 | 0.8963 | 0.8021 | 0.8810 | 0.6485 | 2,590,230 | 6.4 G |
| YOLOv12 | 0.8885 | 0.7837 | 0.8957 | 0.6540 | 2,538,486 | 6.0 G |
| YOLOv13 | 0.8760 | 0.7891 | 0.9031 | 0.7058 | 2,460,301 | 6.4 G |
| Methods | Precision | Recall | mAP@0.5 | mAP@0.5:0.95 | Params | FLOPs |
|---|---|---|---|---|---|---|
| DTNet | 0.9132 | 0.8931 | 0.9552 | 0.8001 | 4,141,053 | 11.9 G |
| SDFMNet [45] | 0.9085 | 0.8425 | 0.9286 | 0.7408 | 4,038,653 | 10.7 G |
| ICANet [46] | 0.9055 | 0.8372 | 0.9291 | 0.7310 | 6,908,449 | 11.4 G |
| MSGANet [47] | 0.8789 | 0.8303 | 0.9180 | 0.7279 | 4,741,138 | 13.7 G |
| CGANet [48] | 0.8953 | 0.8413 | 0.9276 | 0.7266 | 3,722,086 | 10.4 G |
| MFMNet [49] | 0.9107 | 0.8375 | 0.9265 | 0.7223 | 2,486,525 | 8.5 G |
| iRMBNet [50] | 0.8902 | 0.8444 | 0.9237 | 0.7328 | 4,096,509 | 11.8 G |
| No. | RGB Input | IR Input | +YOLOv13-MidFusion (Simple Addition-Based) | +DFB | +LEM | +ASA | +Transformer-Based Head | +MBDL | Precision | Recall | mAP@0.5 | mAP@0.5:0.95 | Params | FLOPs |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ✔ | 0.8760 | 0.7891 | 0.9030 | 0.7058 | 2,460,301 | 6.4 G | |||||||
| 2 | ✔ | 0.8531 | 0.7642 | 0.8812 | 0.6898 | 2,460,301 | 6.4 G | |||||||
| 3 | ✔ | ✔ | ✔ | 0.8833 | 0.7911 | 0.9117 | 0.7158 | 5,168,124 | 12.9 G | |||||
| 4 | ✔ | ✔ | ✔ | 0.9147 | 0.8567 | 0.9318 | 0.7412 | 11,139,217 | 22.9 G | |||||
| 5 | ✔ | ✔ | ✔ | ✔ | 0.9181 | 0.8715 | 0.9439 | 0.7759 | 3,782,645 | 10.7 G | ||||
| 6 | ✔ | ✔ | ✔ | ✔ | ✔ | 0.9169 | 0.8874 | 0.9520 | 0.7944 | 4,141,053 | 11.9 G | |||
| 7 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | 0.9140 | 0.8907 | 0.9537 | 0.7974 | 4,141,053 | 11.9 G | ||
| 8 | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | 0.9132 | 0.8931 | 0.9552 | 0.8001 | 4,141,053 | 11.9 G |
| Dataset | Precision | Recall | mAP@0.5 | mAP@0.5:0.95 | Params | FLOPs |
|---|---|---|---|---|---|---|
| Grape dataset | 0.9132 | 0.8931 | 0.9552 | 0.8001 | 4,141,053 | 11.9 G |
| Tomato dataset | 0.9098 | 0.8843 | 0.9463 | 0.7863 | 4,141,053 | 11.9 G |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Di, Z.; Sun, G. DTNet: A Novel Infrared-Visible Fruit Object Detection Method Based on Dual-Modal Feature Interaction Fusion and Transformer Decoding. Agriculture 2026, 16, 1044. https://doi.org/10.3390/agriculture16101044
Di Z, Sun G. DTNet: A Novel Infrared-Visible Fruit Object Detection Method Based on Dual-Modal Feature Interaction Fusion and Transformer Decoding. Agriculture. 2026; 16(10):1044. https://doi.org/10.3390/agriculture16101044
Chicago/Turabian StyleDi, Ziqian, and Guoxiang Sun. 2026. "DTNet: A Novel Infrared-Visible Fruit Object Detection Method Based on Dual-Modal Feature Interaction Fusion and Transformer Decoding" Agriculture 16, no. 10: 1044. https://doi.org/10.3390/agriculture16101044
APA StyleDi, Z., & Sun, G. (2026). DTNet: A Novel Infrared-Visible Fruit Object Detection Method Based on Dual-Modal Feature Interaction Fusion and Transformer Decoding. Agriculture, 16(10), 1044. https://doi.org/10.3390/agriculture16101044

