MPFT-UNet: A Boundary-Refined and Multi-Scale Dynamic Fusion Network for UAV-Based Port Ship Segmentation
Abstract
1. Introduction
- A Boundary Supervision Refinement (BSR) module is introduced to address the low contrast and irregular characteristics of ship boundaries under complex maritime conditions. This module enhances boundary-aware feature representation and improves the continuity of object contours.
- A Multi-scale Dynamic Sparse Cross-gating (MDSC) module is designed to cope with significant scale variation and the predominance of small-scale targets in UAV imagery. By incorporating a dynamic sparse attention mechanism, the module enables effective selection of informative cross-scale features while reducing the influence of background noise.
- A Transformer-based Feature Fusion (FFT) module is embedded at the bottleneck layer to mitigate the limitations of conventional convolutional networks in capturing long-range dependencies. This design contributes to improved global semantic consistency in complex maritime scenes.
- Extensive experiments conducted on the LSRS-Ship dataset demonstrate that the proposed MPFT-UNet achieves superior performance compared with existing state-of-the-art methods. The results indicate notable improvements in segmentation accuracy and robustness under complex UAV-based maritime conditions.
2. Related Work
2.1. Boundary-Aware Segmentation and Contour Refinement
2.2. Multi-Scale Feature Fusion and Small-Target Modeling
2.3. Transformer-Based Global Context Modeling
3. Proposed Methodology
3.1. Overall Framework
3.2. Boundary Supervision Refinement Module (BSR)
3.3. Multi-Scale Dynamic Sparse Cross-Gating Module (MDSC)
3.3.1. Cross-Scale Multi-Level Fusion
3.3.2. Dynamic Sparse Cross-Scale Fusion
3.4. Transformer-Based Feature Fusion Module (FFT)
3.4.1. Local and Global Path Modeling
3.4.2. Feature Fusion and Enhancement
4. Experiments and Results
4.1. Dataset
4.1.1. LSRS-Ship Dataset
4.1.2. iSAID-Ship Dataset
4.2. Experimental Environment and Settings
4.3. Loss Function
4.4. Comparative Experiments
4.4.1. Qualitative Evaluation
4.4.2. Quantitative Evaluation
4.5. Ablation Experiments
4.6. Parameter Sensitivity Analysis
4.7. Cross-Dataset Generalization Experiments
5. Conclusions
Author Contributions
Funding
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Duan, G.J.; Zhang, P.F. Research on application of UAV for maritime supervision. J. Shipp. Ocean Eng. 2014, 4, 322–326. [Google Scholar]
- Koritarov, T.; Dimitrakiev, D. Unmanned Aerial Vehicles in Port Operations–Use Cases and Benefits. In Proceedings of the International Scientific Conference Innovative Education for Emerging Maritime Issues, Varna, Bulgaria, 25 February 2021; Volume 25, pp. 70–78. [Google Scholar]
- Song, A. Deep Learning-Based Semantic Segmentation of Urban Areas Using Heterogeneous Unmanned Aerial Vehicle Datasets. Aerospace 2023, 10, 880. [Google Scholar] [CrossRef] [Scilit]
- Luo, X.; Wu, Y.; Chen, J. Research progress on deep learning methods for object detection and semantic segmentation in UAV aerial images. Acta Aeronaut. Astronaut. Sin. 2024, 45, 028822. [Google Scholar]
- Han, Y.; Guo, J.; Yang, H.; Guan, R.; Zhang, T. SSMA-YOLO: A lightweight YOLO model with enhanced feature extraction and Fusion capabilities for drone-aerial ship image detection. Drones 2024, 8, 145. [Google Scholar] [CrossRef] [Scilit]
- Zhao, T.; Wang, Y.; Li, Z.; Gao, Y.; Chen, C.; Feng, H.; Zhao, Z. Ship detection with deep learning in optical remote-sensing images: A survey of challenges and advances. Remote Sens. 2024, 16, 1145. [Google Scholar] [CrossRef] [Scilit]
- Huang, L.; Jiang, B.; Lv, S.; Liu, Y.; Fu, Y. Deep-learning-based semantic segmentation of remote sensing images: A survey. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2023, 17, 8370–8396. [Google Scholar]
- Cheng, J.; Deng, C.; Su, Y.; An, Z.; Wang, Q. Methods and datasets on semantic segmentation for Unmanned Aerial Vehicle remote sensing images: A review. ISPRS J. Photogramm. Remote Sens. 2024, 211, 1–34. [Google Scholar]
- Liu, X.; Deng, Z.; Yang, Y. Recent progress in semantic image segmentation. Artif. Intell. Rev. 2019, 52, 1089–1106. [Google Scholar]
- Majidizadeh, A.; Hasani, H.; Jafari, M. Semantic segmentation of UAV images based on U-NET in urban area. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2023, 10, 451–457. [Google Scholar]
- Kumar, S.; Kumar, A.; Lee, D.G. Semantic segmentation of UAV images based on transformer framework with context information. Mathematics 2022, 10, 4735. [Google Scholar] [CrossRef] [Scilit]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. Adv. Neural Inf. Process. Syst. 2015, 28, 91. [Google Scholar] [CrossRef] [Scilit]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. Ssd: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; IEEE: New York, NY, USA, 2016; pp. 779–788. [Google Scholar]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; IEEE: New York, NY, USA, 2015; pp. 3431–3440. [Google Scholar]
- Badrinarayanan, V.; Kendall, A.; Cipolla, R. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 2481–2495. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar]
- Shamsolmoali, P.; Zareapoor, M.; Wang, R.; Zhou, H.; Yang, J. A novel deep structure U-Net for sea-land segmentation in remote sensing images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2019, 12, 3219–3232. [Google Scholar] [CrossRef] [Scilit]
- Hordiiuk, D.; Oliinyk, I.; Hnatushenko, V.; Maksymov, K. Semantic segmentation for ships detection from satellite imagery. In Proceedings of the 2019 IEEE 39th International Conference on Electronics and Nanotechnology (ELNANO), Kyiv, Ukraine, 16–18 April 2019; IEEE: New York, NY, USA, 2019; pp. 454–457. [Google Scholar]
- Wu, K.; Zhao, S.; Li, W.; Jiang, R. Spatial global context information network for semantic segmentation of remote sensing image. J. Zhejiang Univ. Eng. Sci. 2022, 56, 795–802. [Google Scholar]
- Li, G.; Wang, R.; Zhang, Y.; Xu, C.; Fan, X.; Zhou, Z.; Lv, P.; Ruan, Z. LR-Net: Lossless Feature Fusion and Revised SIoU for Small Object Detection. Comput. Mater. Contin. 2025, 85, 3267. [Google Scholar] [CrossRef] [Scilit]
- Spasev, V.; Dimitrovski, I.; Chorbev, I.; Kitanovski, I. Semantic segmentation of unmanned aerial vehicle remote sensing images using SegFormer. In Proceedings of the International Conference on Intelligent Systems and Pattern Recognition; Springer: Cham, Switzerland, 2024; pp. 108–122. [Google Scholar]
- Bokhovkin, A.; Burnaev, E. Boundary loss for remote sensing imagery semantic segmentation. In Proceedings of the International Symposium on Neural Networks; Springer: Cham, Switzerland, 2019; pp. 388–401. [Google Scholar]
- Dharampal, V.M. Methods of image edge detection: A review. J. Electr. Electron. Syst. 2015, 4, 150. [Google Scholar]
- Chen, Y.; Ge, P.; Wang, G.; Weng, G.; Chen, H. An overview of intelligent image segmentation using active contour models. Intell. Robot. 2023, 3, 23–55. [Google Scholar] [CrossRef] [Scilit]
- Takikawa, T.; Acuna, D.; Jampani, V.; Fidler, S. Gated-scnn: Gated shape cnns for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; IEEE: New York, NY, USA, 2019; pp. 5229–5238. [Google Scholar]
- Kervadec, H.; Bouchtiba, J.; Desrosiers, C.; Granger, E.; Dolz, J.; Ayed, I.B. Boundary loss for highly unbalanced segmentation. In Proceedings of the International Conference on Medical Imaging with Deep Learning, London, UK, 8–10 July 2019; PMLR: Cambridge, MA, USA, 2019; pp. 285–296. [Google Scholar]
- Karimi, D.; Salcudean, S.E. Reducing the hausdorff distance in medical image segmentation with convolutional neural networks. IEEE Trans. Med. Imaging 2019, 39, 499–513. [Google Scholar] [CrossRef] [Scilit]
- Fu, J.; Liu, J.; Tian, H.; Li, Y.; Bao, Y.; Fang, Z.; Lu, H. Dual attention network for scene segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; IEEE: New York, NY, USA, 2019; pp. 3146–3154. [Google Scholar]
- McEver, R.A.; Manjunath, B. Pcams: Weakly supervised semantic segmentation using point supervision. arXiv 2020, arXiv:2007.05615. [Google Scholar] [CrossRef] [Scilit]
- Dolgopolov, A.V.; Kazantsev, P.A.; Bezuhliy, N.; Dolgopolov, A.; Kazantsev, P.; Bezuhliy, N. Ship detection in images obtained from the unmanned aerial vehicle (UAV). Indian J. Sci. Technol. 2017, 9, 1–7. [Google Scholar]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; IEEE: New York, NY, USA, 2017; pp. 2117–2125. [Google Scholar]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; IEEE: New York, NY, USA, 2017; pp. 2881–2890. [Google Scholar]
- Chen, L.C.; Papandreou, G.; Schroff, F.; Adam, H. Rethinking atrous convolution for semantic image segmentation. arXiv 2017, arXiv:1706.05587. [Google Scholar] [CrossRef] [Scilit]
- Shang, R.; Zhang, J.; Jiao, L.; Li, Y.; Marturi, N.; Stolkin, R. Multi-scale adaptive feature fusion network for semantic segmentation in remote sensing images. Remote Sens. 2020, 12, 872. [Google Scholar] [CrossRef] [Scilit]
- Zhang, G.; Jiang, W. Remote Sensing Image Semantic Segmentation Method Based on a Deep Convolutional Neural Network and Multiscale Feature Fusion. Int. J. Semant. Web Inf. Syst. 2023, 19, 16. [Google Scholar]
- Meng, T.; Ghiasi, G.; Mahjourian, R.; Le, Q.V.; Tan, M. Revisiting multi-scale feature fusion for semantic segmentation. arXiv 2022, arXiv:2203.12683. [Google Scholar] [CrossRef] [Scilit]
- Chan, Y.T. Maritime filtering for images and videos. Signal Process. Image Commun. 2021, 99, 116477. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5999–6009. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 10012–10022. [Google Scholar]
- Wang, W.; Xie, E.; Li, X.; Fan, D.P.; Song, K.; Liang, D.; Lu, T.; Luo, P.; Shao, L. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 568–578. [Google Scholar]
- Strudel, R.; Garcia, R.; Laptev, I.; Schmid, C. Segmenter: Transformer for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 7262–7272. [Google Scholar]
- Yuan, L.; Chen, Y.; Wang, T.; Yu, W.; Shi, Y.; Jiang, Z.H.; Tay, F.E.; Feng, J.; Yan, S. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 558–567. [Google Scholar]
- Zhou, Z.; Rahman Siddiquee, M.M.; Tajbakhsh, N.; Liang, J. Unet++: A nested u-net architecture for medical image segmentation. In Proceedings of the International Workshop on Deep Learning in Medical Image Analysis; Springer: Cham, Switzerland, 2018; pp. 3–11. [Google Scholar]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; Springer: Berlin/Heidelberg, Germany, 2018; pp. 801–818. [Google Scholar]
- Oktay, O.; Schlemper, J.; Folgoc, L.L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N.Y.; Kainz, B.; et al. Attention u-net: Learning where to look for the pancreas. arXiv 2018, arXiv:1804.03999. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. Transunet: Transformers make strong encoders for medical image segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. Vmamba: Visual state space model. Adv. Neural Inf. Process. Syst. 2024, 37, 103031–103063. [Google Scholar]











| Category | Setting |
|---|---|
| Framework | PyTorch 2.8.0 |
| GPU | NVIDIA RTX 5090 (32 GB) |
| Input Size | |
| Optimizer | Adam |
| Learning Rate | (Cosine Annealing) |
| Batch Size | 8 |
| Epochs | 80 |
| Loss Function | () |
| Model | IoU | Dice | Precision | Recall | AP | Inference Time (ms) |
|---|---|---|---|---|---|---|
| UNet | 0.7854 | 0.8591 | 0.9030 | 0.8384 | 0.9371 | 7.00 |
| UNet++ | 0.6857 | 0.7663 | 0.8525 | 0.7336 | 0.8712 | 22.51 |
| DeepLabV3+ | 0.7891 | 0.8648 | 0.9076 | 0.8488 | 0.9406 | 4.07 |
| Attention UNet | 0.6263 | 0.7149 | 0.8245 | 0.6782 | 0.8531 | 10.40 |
| TransUNet | 0.7347 | 0.8168 | 0.8810 | 0.7922 | 0.9028 | 13.28 |
| Mamba UNet | 0.6912 | 0.7731 | 0.8533 | 0.7408 | 0.8878 | 4.01 |
| MPFT-UNet | 0.8365 | 0.9028 | 0.9346 | 0.8881 | 0.9573 | 8.52 |
| BSR | MDSC | FFT | IoU | Dice | Precision | Recall | AP |
|---|---|---|---|---|---|---|---|
| × | × | × | 0.7854 | 0.8591 | 0.9030 | 0.8384 | 0.9371 |
| ✓ | × | × | 0.8181 | 0.8874 | 0.9184 | 0.8762 | 0.9441 |
| × | ✓ | × | 0.8032 | 0.8727 | 0.9009 | 0.8664 | 0.9403 |
| × | × | ✓ | 0.8247 | 0.8909 | 0.9230 | 0.8757 | 0.9471 |
| ✓ | ✓ | × | 0.8095 | 0.8776 | 0.9160 | 0.8630 | 0.9445 |
| ✓ | × | ✓ | 0.8261 | 0.8926 | 0.9207 | 0.8816 | 0.9448 |
| × | ✓ | ✓ | 0.8296 | 0.8952 | 0.9254 | 0.8818 | 0.9504 |
| ✓ | ✓ | ✓ | 0.8365 | 0.9028 | 0.9346 | 0.8881 | 0.9573 |
| IoU | Dice | |
|---|---|---|
| 0.3 | 0.8317 | 0.8991 |
| 0.5 | 0.8365 | 0.9028 |
| 0.7 | 0.8228 | 0.8943 |
| Model | IoU | Dice | Precision | Recall | AP | Inference Time (ms) |
|---|---|---|---|---|---|---|
| UNet | 0.6112 | 0.7245 | 0.8012 | 0.6724 | 0.8410 | 7.00 |
| UNet++ | 0.5887 | 0.7084 | 0.7796 | 0.6513 | 0.8295 | 22.51 |
| DeepLabV3+ | 0.6320 | 0.7550 | 0.8157 | 0.6832 | 0.8578 | 4.06 |
| Attention UNet | 0.5902 | 0.7226 | 0.7823 | 0.6481 | 0.8263 | 10.40 |
| TransUNet | 0.5843 | 0.7128 | 0.7714 | 0.6372 | 0.8209 | 13.28 |
| Mamba UNet | 0.6095 | 0.7296 | 0.7964 | 0.6480 | 0.8315 | 4.01 |
| MPFT-UNet | 0.6927 | 0.7793 | 0.8414 | 0.7552 | 0.8679 | 8.52 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Shi, M.; Qiu, X.; Li, A.; Yang, Y.; Ke, Y.; Chen, Y. MPFT-UNet: A Boundary-Refined and Multi-Scale Dynamic Fusion Network for UAV-Based Port Ship Segmentation. J. Mar. Sci. Eng. 2026, 14, 945. https://doi.org/10.3390/jmse14100945
Shi M, Qiu X, Li A, Yang Y, Ke Y, Chen Y. MPFT-UNet: A Boundary-Refined and Multi-Scale Dynamic Fusion Network for UAV-Based Port Ship Segmentation. Journal of Marine Science and Engineering. 2026; 14(10):945. https://doi.org/10.3390/jmse14100945
Chicago/Turabian StyleShi, Mengna, Xiulin Qiu, Ang Li, Yuwang Yang, Yaqi Ke, and Yilan Chen. 2026. "MPFT-UNet: A Boundary-Refined and Multi-Scale Dynamic Fusion Network for UAV-Based Port Ship Segmentation" Journal of Marine Science and Engineering 14, no. 10: 945. https://doi.org/10.3390/jmse14100945
APA StyleShi, M., Qiu, X., Li, A., Yang, Y., Ke, Y., & Chen, Y. (2026). MPFT-UNet: A Boundary-Refined and Multi-Scale Dynamic Fusion Network for UAV-Based Port Ship Segmentation. Journal of Marine Science and Engineering, 14(10), 945. https://doi.org/10.3390/jmse14100945

