DAMFusion: Multi-Spectral Image Segmentation via Competitive Query and Boundary Region Attention
Highlights
- A novel multi-source image fusion branch named DAMFusion is proposed, which integrates a Competitive Query Module, a Multimodal Fusion Module, and a Boundary-Aware Attention Multi-level Fusion Module. This architecture effectively alleviates cross-modal interference and low small-target segmentation accuracy in multispectral farmland imagery. It achieves state-of-the-art performance on the self-built OUC-UAV-MSEG dataset, with an F1-score of 0.927 and a mean Intersection over Union (mIoU) of 0.897.
- Ablation experiments verify the independent effectiveness of each core module in DAMFusion: the modal competitive query selection strategy outperforms learnable queries in feature representation. The combination of cross-modal attention fusion and boundary-region attention enhancement significantly improves the model’s ability to recover fine boundary details and distinguish similar farmland objects.
- The dynamic modal feature selection and cross-layer detail fusion design of DAMFusion provide a new technical paradigm for multispectral remote sensing image segmentation, which can be extended to other complex scene segmentation tasks such as urban land cover classification and natural disaster damage assessment.
- The self-built OUC-UAV-MSEG multispectral farmland dataset fills the gap of dedicated benchmark datasets in the field of agricultural UAV image segmentation. The high-precision segmentation capability of DAMFusion lays a technical foundation for practical smart agriculture applications including precise farmland management, variable-rate operations, and crop growth monitoring.
Abstract
1. Introduction
2. Related Works
2.1. Unsupervised Fusion
2.2. Contrastive Fusion
2.3. Based on Generative Adversarial Networks
2.4. Multi-Task Fusion Optimization
2.5. Comparison with State-of-the-Art Remote Sensing Segmentation Methods
3. Difficulty Analysis
3.1. Difficulties in RGB and Multispectral Image Fusion
3.2. Difficulties Caused by Farmland Application Scenarios
3.3. Difficulties in Model Design
4. Algorithm Design
4.1. Overall Algorithm Framework
4.2. Competitive Query Module
4.3. Multimodal Fusion Module
4.4. Boundary Region Attention Multi-Level Fusion Module
4.5. Loss
5. Experiments and Evaluation
5.1. Dataset Introduction
5.2. Experimental Details
5.2.1. Qualitative Experiments
5.2.2. Quantitative Experiments
5.3. Ablation Experiments
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Reedha, R.; Dericquebourg, E.; Canals, R.; Hafiane, A. Vision Transformers for Weeds and Crops Classification of High Resolution UAV Images. arXiv 2021, arXiv:2109.02716. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Cong, R.; Li, C.; Cheng, M.M.; Fang, Y.; Cao, X.; Zhao, Y.; Kwong, S. Dense attention fluid network for salient object detection in optical remote sensing images. IEEE Trans. Image Process. 2020, 30, 1305–1317. [Google Scholar] [CrossRef] [Scilit]
- Tu, B.; Ren, Q.; Li, J.; Cao, Z.; Chen, Y.; Plaza, A. NCGLF2: Network combining global and local features for fusion of multisource remote sensing data. Inf. Fusion 2024, 104, 102192. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Gao, K.; Wang, H.; Yang, Z.; Wang, P.; Ji, S.; Huang, Y.; Zhu, Z.; Zhao, X. A Transformer-based multi-modal fusion network for semantic segmentation of high-resolution remote sensing imagery. Int. J. Appl. Earth Obs. Geoinf. 2024, 133, 104083. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Yang, J.; Zhang, Y.; Wang, F.; Zheng, F. Depth-Aware Concealed Crop Detection in Dense Agricultural Scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 17201–17211. [Google Scholar]
- Yu, J.; Wang, A.; Dong, W.; Xu, M.; Islam, M.; Wang, J.; Bai, L.; Ren, H. Sam 2 in robotic surgery: An empirical evaluation for robustness and generalization in surgical video segmentation. arXiv 2024, arXiv:2408.04593. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Shi, Y.; Yi, C.; Du, H.; Kang, J.; Niyato, D. Generative AI-Driven Human Digital Twin in IoT-Healthcare: A Comprehensive Survey. arXiv 2024, arXiv:2401.13699. [Google Scholar] [CrossRef] [Scilit]
- Fu, J.; Liu, J.; Tian, H.; Li, Y.; Bao, Y.; Fang, Z.; Lu, H. Dual attention network for scene segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 3146–3154. [Google Scholar]
- Liu, Y.; Li, H.; Cheng, J.; Chen, X. MSCAF-Net: A general framework for camouflaged object detection via learning multi-scale context-aware features. IEEE Trans. Circuits Syst. Video Technol. 2023, 33, 4934–4947. [Google Scholar] [CrossRef] [Scilit]
- Zhou, X.; Liang, F.; Chen, L.; Liu, H.; Song, Q.; Vivone, G.; Chanussot, J. Mesam: Multiscale enhanced segment anything model for optical remote sensing images. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5623515. [Google Scholar] [CrossRef] [Scilit]
- Yu, Z.; Zhang, X.; Zhao, L.; Bin, Y.; Xiao, G. Exploring Deeper! Segment Anything Model with Depth Perception for Camouflaged Object Detection. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, Australia, 28 October–1 November 2024; pp. 4322–4330. [Google Scholar]
- Lan, X.; Gu, X.; Gu, X. MMNet: Multi-modal multi-stage network for RGB-T image semantic segmentation. Appl. Intell. 2022, 52, 5817–5829. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Liu, H.; Yang, K.; Hu, X.; Liu, R.; Stiefelhagen, R. CMX: Cross-modal fusion for RGB-X semantic segmentation with transformers. IEEE Trans. Intell. Transp. Syst. 2023, 24, 14679–14694. [Google Scholar] [CrossRef] [Scilit]
- Hu, S.; Bonardi, F.; Bouchafa, S.; Sidibé, D. Multi-modal unsupervised domain adaptation for semantic image segmentation. Pattern Recognit. 2023, 137, 109299. [Google Scholar] [CrossRef] [Scilit]
- Fan, R.; Wang, Z.; Zhu, Q. EGFNet: Efficient guided feature fusion network for skin cancer lesion segmentation. In Proceedings of the 2022 6th International Conference on Innovation in Artificial Intelligence, Guangzhou, China, 4–6 March 2022; pp. 95–99. [Google Scholar]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 801–818. [Google Scholar]
- Cai, Y.; Shang, Y.; Yin, J. MultiDAN: Unsupervised, Multistage, Multisource and Multitarget Domain Adaptation for Semantic Segmentation of Remote Sensing Images. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, Australia, 28 October–1 November 2024; pp. 1168–1177. [Google Scholar]
- Maiti, A.; Elberink, S.O.; Vosselman, G. TransFusion: Multi-modal fusion network for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 6537–6547. [Google Scholar]
- Yi, M.; Wang, X.; Liu, J.; Zhang, Y.; Hou, R. Meta-Reinforcement Learning for Timely and Energy-efficient Data Collection in Solar-powered UAV-assisted IoT Networks. arXiv 2023, arXiv:2311.06742. [Google Scholar]
- Yi, M.; Wang, X.; Liu, J.; Zhang, Y.; Bai, B. Deep Reinforcement Learning for Fresh Data Collection in UAV-assisted IoT Networks. arXiv 2020, arXiv:2003.00391. [Google Scholar] [CrossRef] [Scilit]
- Geraci, G.; Garcia-Rodriguez, A.; Giordano, L.G.; López-Pérez, D.; Björnson, E. Understanding UAV Cellular Communications: From Existing Networks to Massive MIMO. arXiv 2018, arXiv:1804.08489. [Google Scholar] [CrossRef] [Scilit]
- Fikri, M.R.; Candra, T.; Saptaji, K.; Noviarini, A.N.; Wardani, D.A. A review of Implementation and Challenges of Unmanned Aerial Vehicles for Spraying Applications and Crop Monitoring in Indonesia. arXiv 2023, arXiv:2301.00379. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Zhang, X.; Pun, M.O.; Liu, M. A multilevel multimodal fusion transformer for remote sensing semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5403215. [Google Scholar] [CrossRef] [Scilit]
- Son, N.T.; Hoang, Q.C.; Giang, D.T.H.; Trung, V.M.; Huy, V.Q.; Tuan, M.A. Developing system of wireless sensor network and unmaned aerial vehicle for agriculture inspection. arXiv 2021, arXiv:2107.01008. [Google Scholar] [CrossRef] [Scilit]
- Nomikos, N.; Gkonis, P.K.; Bithas, P.S.; Trakadas, P. A Survey on UAV-Aided Maritime Communications: Deployment Considerations, Applications, and Future Challenges. IEEE Open J. Commun. Soc. 2023, 4, 56–78. [Google Scholar] [CrossRef] [Scilit]
- Pal, O.K.; Shovon, M.S.H.; Mridha, M.F.; Shin, J. A Comprehensive Review of AI-enabled Unmanned Aerial Vehicle: Trends, Vision, and Challenges. arXiv 2023, arXiv:2310.16360. [Google Scholar] [CrossRef] [Scilit]
- Zhao, H.; Li, W.; Huang, D.; Huang, J.; Zhang, L. M-GAN: Multiattribute learning and multimodal feature fusion-based generative adversarial network for text-to-image synthesis. Vis. Comput. 2024, 41, 3017–3035. [Google Scholar] [CrossRef] [Scilit]
- Yi, X.; Tang, L.; Zhang, H.; Xu, H.; Ma, J. Diff-IF: Multi-modality image fusion via diffusion model with fusion knowledge prior. Inf. Fusion 2024, 110, 102450. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Cao, M.; Xie, W.; Lei, J.; Li, D.; Huang, W.; Li, Y.; Yang, X. E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection. Adv. Neural Inf. Process. Syst. 2025, 37, 52296–52322. [Google Scholar]
- LeCun, Y.; Boser, B.; Denker, J.S.; Henderson, D.; Howard, R.E.; Hubbard, W.; Jackel, L.D. Backpropagation applied to handwritten zip code recognition. Neural Comput. 1989, 1, 541–551. [Google Scholar] [CrossRef] [Scilit]
- Yang, Y.; Tong, S.; Huang, S.; Lin, P. Dual-tree complex wavelet transform and image block residual-based multi-focus image fusion in visual sensor networks. Sensors 2014, 14, 22408–22430. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Xu, S.; Liu, J.; Zhao, Z.; Zhang, C.; Zhang, J. MFIF-GAN: A new generative adversarial network for multi-focus image fusion. Signal Process. Image Commun. 2021, 96, 116295. [Google Scholar] [CrossRef] [Scilit]
- Badrinarayanan, V.; Kendall, A.; Cipolla, R. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 2481–2495. [Google Scholar] [CrossRef] [Scilit]
- Lei, Z.; Fang, T.; Huo, H.; Li, D. Bi-temporal texton forest for land cover transition detection on remotely sensed imagery. IEEE Trans. Geosci. Remote Sens. 2013, 52, 1227–1237. [Google Scholar] [CrossRef]
- Chen, L.C. Semantic image segmentation with deep convolutional nets and fully connected CRFs. arXiv 2014, arXiv:1412.7062. [Google Scholar]
- Dong, A.; Wang, L.; Liu, J.; Xu, J.; Zhao, G.; Zhai, Y.; Lv, G.; Cheng, J. Co-Enhancement of Multi-modality Image Fusion and Object Detection via Feature Adaptation. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 12624–12637. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Liu, Z.; Wu, G.; Ma, L.; Liu, R.; Zhong, W.; Luo, Z.; Fan, X. Multi-interactive feature learning and a full-time multi-modality benchmark for image fusion and segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 8115–8124. [Google Scholar]
- Shim, J.h.; Yu, H.; Kong, K.; Kang, S.J. Feedformer: Revisiting transformer decoder for efficient semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 2263–2271. [Google Scholar]
- Yu, C.; Gao, C.; Wang, J.; Yu, G.; Shen, C.; Sang, N. Bisenet v2: Bilateral network with guided aggregation for real-time semantic segmentation. Int. J. Comput. Vis. 2021, 129, 3051–3068. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar]
- Yao, Z.; Ai, J.; Li, B.; Zhang, C. Efficient detr: Improving end-to-end object detector with dense prior. arXiv 2021, arXiv:2104.01318. [Google Scholar]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H.Y. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv 2022, arXiv:2203.03605. [Google Scholar]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable detr: Deformable transformers for end-to-end object detection. arXiv 2020, arXiv:2010.04159. [Google Scholar]
- Zhang, H.; Wang, Y.; Dayoub, F.; Sunderhauf, N. Varifocalnet: An iou-aware dense object detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 8514–8523. [Google Scholar]
- Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P.H.; et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 6881–6890. [Google Scholar]
- Chen, S.; Tan, X.; Wang, B.; Hu, X. Reverse Attention for Salient Object Detection. arXiv 2018, arXiv:1807.09940. [Google Scholar]
- Markus Gerke, I. Use of the Stair Vision Library Within the ISPRS 2D Semantic Labeling Benchmark (Vaihingen); ResearcheGate: Berlin, Germany, 2014. [Google Scholar]









| Pub./Year | OA | MP | Recall | F1 | mIoU | FWIoU | |
|---|---|---|---|---|---|---|---|
| DeepLabV3+ [16] | ECCV/2018 | 79.84 | 73.66 | 75.07 | 73.80 | 69.03 | 70.03 |
| MMNet [12] | Appl. Intell/2022 | 83.97 | 78.63 | 82.21 | 81.28 | 78.33 | 78.39 |
| CMX [13] | T-ITS/2023 | 83.12 | 83.49 | 84.07 | 84.27 | 81.97 | 83.94 |
| NCGLF2 [3] | Inform Fusion/2024 | 85.83 | 86.33 | 84.88 | 86.10 | 83.03 | 84.16 |
| AMMFuseNet [4] | Appl Earth Obs/2024 | 87.49 | 88.33 | 86.69 | 87.51 | 85.73 | 86.34 |
| Ours | 93.25 | 92.70 | 91.87 | 91.71 | 89.70 | 88.83 |
| Method | Evaluation Metrics (%) | |||||
|---|---|---|---|---|---|---|
| OA | MP | Recall | F1 | mIoU | FWIoU | |
| Learnable Query | 91.00 | 90.74 | 88.87 | 89.30 | 89.09 | 88.32 |
| Modal Competitive Selection Query | 93.25 | 92.70 | 91.87 | 91.71 | 89.70 | 90.73 |
| Method | OA | MP | Recall | F1 | mIoU | FWIoU | ||
|---|---|---|---|---|---|---|---|---|
| MMformer | w/o Token | BRM | ||||||
| ✓ | 85.43 | 86.30 | 88.16 | 83.04 | 80.24 | 79.97 | ||
| ✓ | 84.21 | 85.00 | 86.45 | 82.37 | 79.64 | 78.02 | ||
| ✓ | ✓ | 87.24 | 89.06 | 90.11 | 87.45 | 82.90 | 83.27 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yu, M.; Lu, X.; Yang, Z.; Gao, D.; Zhong, G. DAMFusion: Multi-Spectral Image Segmentation via Competitive Query and Boundary Region Attention. Remote Sens. 2026, 18, 1064. https://doi.org/10.3390/rs18071064
Yu M, Lu X, Yang Z, Gao D, Zhong G. DAMFusion: Multi-Spectral Image Segmentation via Competitive Query and Boundary Region Attention. Remote Sensing. 2026; 18(7):1064. https://doi.org/10.3390/rs18071064
Chicago/Turabian StyleYu, Miao, Xing Lu, Ziyao Yang, Daoxing Gao, and Guoqiang Zhong. 2026. "DAMFusion: Multi-Spectral Image Segmentation via Competitive Query and Boundary Region Attention" Remote Sensing 18, no. 7: 1064. https://doi.org/10.3390/rs18071064
APA StyleYu, M., Lu, X., Yang, Z., Gao, D., & Zhong, G. (2026). DAMFusion: Multi-Spectral Image Segmentation via Competitive Query and Boundary Region Attention. Remote Sensing, 18(7), 1064. https://doi.org/10.3390/rs18071064

