DFENet: A Novel Dual-Path Feature Extraction Network for Semantic Segmentation of Remote Sensing Images
Abstract
1. Introduction
2. Related Work
2.1. Methods Based on Spatial-Domain
2.2. Methods Based on Frequency-Domain
3. Methodology
3.1. Framework of DFENet
3.2. Structural Design of the Global Path
3.2.1. The Details of PAB
3.2.2. The Details of MAAB
3.2.3. The Details of SCAB
3.2.4. The Details of FFEB
4. Experiment
4.1. Datasets
4.1.1. ISPRS Vaihingen
4.1.2. ISPRS Potsdam
4.2. Experimental Setup
4.3. Performance Comparison
4.4. Analysis and Discussion
4.5. Ablation Study
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Jonnala, N.S.; Bheemana, R.C.; Prakash, K.; Bansal, S.; Jain, A.; Pandey, V.; Faruque, M.R.I.; Al-Mugren, K.S. DSIA U-Net: Deep shallow interaction with attention mechanism UNet for remote sensing satellite images. Sci. Rep. 2025, 15, 549. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jonnala, N.S.; Siraaj, S.; Prastuti, Y.; Chinnababu, P.; Babu, B.P.; Bansal, S.; Upadhyaya, P.; Prakash, K.; Faruque, M.R.I.; Al-Mugren, K.S. AER U-Net: Attention-enhanced multi-scale residual U-Net structure for water body segmentation using Sentinel-2 satellite images. Sci. Rep. 2025, 15, 16099. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 3431–3440. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid Scene Parsing Network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017. [Google Scholar]
- Kirillov, A.; Girshick, R.; He, K.-M.; Dollar, P. Panoptic Feature Pyramid Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
- Ma, A.-L.; Wang, J.-J.; Zhong, Y.-F.; Zheng, Z. Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing Imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020. [Google Scholar]
- Wang, J.-X.; Feng, Z.-X.; Jiang, Y.; Yang, S.-Y.; Meng, H.-X. Orientation Attention Network for semantic segmentation of remote sensing images. Knowl. Based Syst. 2023, 267, 110415. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Hua, W.; Zhang, W.; Liu, F.; Xiao, L. Stair fusion network with context-refined attention for remote sensing image semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4701517. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Zhang, X.; Wang, Z.; Pun, M.-O. Unsupervised domain adaptation augmented by mutually boosted attention for semantic segmentation of VHR remote sensing images. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5400515. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Mei, J.; Li, X.; Lu, Y.; Yu, Q.; Wei, Q.; Luo, X.; Xie, Y.; Adeli, E.; Wang, Y.; et al. TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformers. Med. Image Anal. 2024, 97, 103280. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar]
- You, C.; Jiao, L.; Liu, X.; Li, L.; Liu, F.; Ma, W.; Yang, S. Boundary-aware multiscale learning perception for remote sensing image segmentation. IEEE Trans. Geosci. Remote Sens. 2023, 61, 4407115. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Geng, J.; Jiang, W. Mmt: Mixed-mask transformer for remote sensing image semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5613415. [Google Scholar] [CrossRef] [Scilit]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar] [CrossRef] [Scilit]
- Gu, A.; Goel, K.; Ré, C. Efficiently modeling long sequences with structured state spaces. arXiv 2021, arXiv:2111.00396. [Google Scholar]
- He, X.; Cao, K.; Yan, K.; Li, R.; Xie, C.; Zhang, J.; Zhou, M. Pan-Mamba: Effective pan-sharpening with state space model. arXiv 2024, arXiv:2402.12192. [Google Scholar] [CrossRef] [Scilit]
- Chen, K.; Chen, B.; Liu, C.; Li, W.; Zou, Z.; Shi, Z. RS-Mamba: Remote sensing image classification with state space model. IEEE Geosci. Remote Sens. Lett. 2024, 21, 8002605. [Google Scholar]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 10781–10790. [Google Scholar]
- Qiao, S.; Chen, L.-C.; Yuille, A. DetectoRS: Detecting objects with recursive feature pyramid and switchable atrous convolution. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 10213–10224. [Google Scholar]
- Ma, X.; Zhang, X.; Pun, M.-O. RS-3-Mamba: Visual state space model for remote sensing image semantic segmentation. IEEE Geosci. Remote Sens. Lett. 2024, 21, 6011405. [Google Scholar] [CrossRef] [Scilit]
- Lu, W.; Chen, S.-B.; Ding, C.H.Q.; Tang, J.; Luo, B. LWGANet: A lightweight group attention backbone for remote sensing visual tasks. arXiv 2025, arXiv:2501.10040. [Google Scholar] [CrossRef] [Scilit]
- Hwang, S.; Han, D.; Jung, C.; Jeon, M. WaveDH: Wavelet sub-bands guided ConvNet for efficient image dehazing. arXiv 2024, arXiv:2404.01604. [Google Scholar]
- Xu, Z.; Zhang, W.; Zhang, T.; Yang, Z.; Li, J. Efficient Transformer for remote sensing image segmentation. Remote Sens. 2021, 13, 3585. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 10012–10022. [Google Scholar]
- Yang, C.; Wang, Y.; Zhang, J.; Zhang, H.; Wei, Z.; Lin, Z.; Yuille, A. Lite Vision Transformer with enhanced self-attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 11998–12008. [Google Scholar]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision Mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024, arXiv:2401.09417. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. VMamba: Visual state space model. Adv. Neural Inf. Process. Syst. 2024, 37, 103031–103063. [Google Scholar]
- Fan, J.; Li, J.; Liu, Y.; Zhang, F. Frequency-aware robust multidimensional information fusion framework for remote sensing image segmentation. Eng. Appl. Artif. Intell. 2024, 129, 107638. [Google Scholar] [CrossRef] [Scilit]
- Zou, Z.; Yu, H.; Huang, J.; Zhao, F. FreqMamba: Viewing Mamba from a frequency perspective for image deraining. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia, 28 October–1 November 2024; pp. 1905–1914. [Google Scholar]
- Li, Y.; Liu, Z.; Yang, J.; Zhang, H. Wavelet transform feature enhancement for semantic segmentation of remote sensing images. Remote Sens. 2023, 15, 5644. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Y.; Huang, J.; Wang, C.; Song, L.; Yang, G. XNet: Wavelet-based low and high frequency fusion networks for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; p. 462. [Google Scholar]
- Wei, G.; Xu, J.; Yan, W.; Chong, Q.; Xing, H.; Ni, M. Dual-domain fusion network based on wavelet frequency decomposition and fuzzy spatial constraint for remote sensing image segmentation. Remote Sens. 2024, 16, 3594. [Google Scholar] [CrossRef] [Scilit]
- Huang, Y.-P.; Bhalla, K.; Chu, H.-C.; Lin, Y.-C.; Kuo, H.-C.; Chu, W.-J.; Lee, J.-H. Wavelet K-means clustering and fuzzy-based method for segmenting MRI images depicting Parkinson’s disease. Int. J. Fuzzy Syst. 2021, 23, 1600–1612. [Google Scholar] [CrossRef] [Scilit]
- Zeng, Y.; Li, J.; Zhao, Z.; Liang, W.; Zeng, P.; Shen, S.; Zhang, K.; Shen, C. WET-UNet: Wavelet integrated efficient transformer networks for nasopharyngeal carcinoma tumor segmentation. Sci. Prog. 2024, 107, 00368504241232537. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Li, R.; Zhang, C.; Fang, S.; Duan, C.; Meng, X.; Atkinson, P.M. UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery. ISPRS J. Photogramm. Remote Sens. 2022, 190, 196–214. [Google Scholar] [CrossRef] [Scilit]
- Hendrycks, D.; Gimpel, K. Gaussian error linear units (GELUs). arXiv 2016, arXiv:1606.08415. [Google Scholar]
- Liu, Y.; Meng, F.; Zhang, J.; Zhou, J.; Chen, Y.; Xu, J. CM-Net: A novel collaborative memory network for spoken language understanding. arXiv 2019, arXiv:1909.06937. [Google Scholar]
- Zeng, Y.; Luo, A.; Zhan, K.; Li, J.; Zhang, Y.; Hu, K. Multiscale feature enhancement and adaptive receptive field for tiny object detection in remote sensing images. In Proceedings of the 2025 International Conference on Multimedia Retrieval, Chicago, IL, USA, 30 June–3 July 2025; pp. 1758–1766. [Google Scholar]
- Zhao, C.; Cai, W.; Dong, C.; Hu, C. Wavelet-based Fourier information interaction with frequency diffusion adjustment for underwater image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 8281–8291. [Google Scholar]
- Yang, Y.; Yuan, G.; Li, J. SFFNet: A wavelet-based spatial and frequency domain fusion network for remote sensing segmentation. IEEE Trans. Geosci. Remote Sens. 2024, 62, 3000617. [Google Scholar] [CrossRef] [Scilit]
- Li, R.; Zheng, S.; Duan, C.; Su, J.; Zhang, C. Multistage attention ResU-Net for semantic segmentation of fine-resolution remote sensing images. IEEE Geosci. Remote Sens. Lett. 2021, 19, 1–5. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. TransUNet: Transformers make strong encoders for medical image segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Huang, P.; Zhang, M.; Tang, W.; Yu, X. CMTFNet: CNN and multiscale transformer fusion network for remote-sensing image semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2023, 61, 2004612. [Google Scholar] [CrossRef] [Scilit]
- Zhu, E.; Chen, Z.; Wang, D.; Shi, H.; Liu, X.; Wang, L. UNetMamba: An efficient UNet-like Mamba for semantic segmentation of high-resolution remote sensing images. IEEE Geosci. Remote Sens. Lett. 2024, 22, 6001205. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Zeng, Z.; Sharma, P.K.; Alfarraj, O.; Tolba, A.; Wang, J. A dual encoder crack segmentation network with Haar wavelet-based high–low frequency attention. Expert Syst. Appl. 2024, 256, 14. [Google Scholar] [CrossRef] [Scilit]







| Method | Impervious Surface | Building | Low Vegetation | Tree | Car | OA (%) | mF1 (%) | mIoU (%) |
|---|---|---|---|---|---|---|---|---|
| SFFNet | 89.11/80.36 | 93.82/88.35 | 75.58/60.74 | 88.81/79.86 | 74.93/59.90 | 90.79 | 84.45 | 73.85 |
| UNetMamba | 91.82/84.88 | 96.10/92.49 | 79.86/66.47 | 90.77/83.09 | 85.81/75.13 | 90.87 | 88.87 | 80.42 |
| TransUNet | 92.21/85.54 | 96.10/92.48 | 80.79/67.77 | 90.87/83.27 | 89.60/81.16 | 91.21 | 89.91 | 82.04 |
| UNetFormer | 92.23/85.58 | 96.34/92.93 | 80.54/67.70 | 91.04/83.55 | 90.37/82.43 | 91.29 | 90.14 | 82.44 |
| MAResU-Net | 92.66/86.33 | 96.84/93.87 | 80.57/67.47 | 90.84/83.22 | 89.93/81.71 | 91.50 | 90.17 | 82.51 |
| CMTFNet | 92.68/86.37 | 96.71/93.63 | 80.47/67.33 | 90.78/83.11 | 90.22/82.18 | 91.42 | 90.17 | 82.52 |
| RS3Mamba | 92.69/86.38 | 96.67/93.55 | 80.54/67.42 | 90.59/82.79 | 90.49/82.64 | 91.30 | 90.20 | 82.56 |
| DECSNet | 92.56/86.14 | 96.87/93.92 | 79.85/66.46 | 90.85/83.23 | 88.36/79.15 | 91.34 | 89.70 | 81.79 |
| MIFNet | 92.50/86.04 | 96.78/93.75 | 80.75/67.72 | 91.05/83.57 | 90.38/82.44 | 91.43 | 90.29 | 82.71 |
| DFENet | 92.68/86.38 | 96.75/93.70 | 80.81/67.70 | 91.27/83.94 | 90.95/83.40 | 91.55 | 90.41 | 83.09 |
| Method | Impervious Surface | Building | Low Vegetation | Tree | Car | OA (%) | mF1 (%) | mIoU (%) |
|---|---|---|---|---|---|---|---|---|
| SFFNet | 91.61/82.47 | 96.82/89.75 | 83.58/69.58 | 90.81/64.85 | 92.93/88.78 | 87.15 | 87.96 | 79.09 |
| UNetMamba | 92.54/86.12 | 96.54/94.63 | 85.96/75.39 | 86.37/76.01 | 95.47/91.32 | 90.39 | 91.37 | 84.41 |
| TransUNet | 93.08/87.06 | 96.88/93.94 | 86.74/76.59 | 87.66/78.03 | 96.40/93.05 | 91.03 | 92.15 | 85.73 |
| UNetFormer | 93.02/86.95 | 97.14/94.43 | 86.21/75.76 | 86.93/76.88 | 96.35/92.96 | 90.77 | 91.89 | 85.33 |
| MAResU-Net | 93.15/87.17 | 97.21/94.57 | 86.73/76.57 | 87.14/77.21 | 96.67/93.56 | 91.05 | 92.18 | 85.82 |
| CMTFNet | 93.08/87.06 | 97.30/94.73 | 87.13/77.20 | 86.89/76.97 | 90.22/82.18 | 90.97 | 91.89 | 85.28 |
| RS3Mamba | 92.95/86.12 | 97.32/94.79 | 85.97/75.39 | 86.60/76.37 | 96.44/93.13 | 90.77 | 91.89 | 85.24 |
| DECSNet | 92.35/85.78 | 97.07/94.31 | 85.43/74.57 | 86.33/75.94 | 95.24/90.92 | 90.29 | 91.29 | 84.31 |
| MIFNet | 92.97/86.86 | 97.23/94.60 | 86.74/76.58 | 87.23/77.35 | 96.91/94.01 | 91.01 | 92.07 | 85.81 |
| DFENet | 93.27/87.44 | 97.31/95.83 | 86.41/76.07 | 87.35/77.53 | 96.49/93.22 | 91.08 | 92.21 | 85.89 |
| Methods | FLOPs(G) | Param(M) | mIoU (%) |
|---|---|---|---|
| SFFNet | 12.99 | 34.18 | 73.85 |
| UNetMamba | 4.34 | 13.89 | 80.42 |
| TransUNet | 38.57 | 105.32 | 82.04 |
| UNetFormer | 2.94 | 11.69 | 82.44 |
| MAResU-Net | 7.19 | 26.28 | 82.51 |
| CMTFNet | 8.87 | 30.07 | 82.52 |
| RS3Mamba | 9.87 | 43.32 | 82.56 |
| DECSNet | 10.07 | 47.41 | 81.79 |
| MIFNet | 8.84 | 25.36 | 82.71 |
| DFENet | 16.56 | 95.05 | 83.09 |
| Local Path | Global Path | mIoU (%) |
|---|---|---|
| ✓ | 81.97 | |
| ✓ | 78.79 | |
| ✓ | ✓ | 83.09 |
| PAB | MAAB | SCAB | FFEB | mIoU (%) |
|---|---|---|---|---|
| ✓ | ✓ | ✓ | 78.31 | |
| ✓ | ✓ | ✓ | 77.17 | |
| ✓ | ✓ | ✓ | 76.42 | |
| ✓ | ✓ | ✓ | 76.56 | |
| ✓ | ✓ | ✓ | ✓ | 78.79 |
| FMU | HAU | mIoU (%) |
|---|---|---|
| ✓ | 82.41 | |
| ✓ | 82.36 | |
| ✓ | ✓ | 83.09 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cao, L.; Liu, Z.; Wang, Y.; Gao, R. DFENet: A Novel Dual-Path Feature Extraction Network for Semantic Segmentation of Remote Sensing Images. J. Imaging 2026, 12, 141. https://doi.org/10.3390/jimaging12030141
Cao L, Liu Z, Wang Y, Gao R. DFENet: A Novel Dual-Path Feature Extraction Network for Semantic Segmentation of Remote Sensing Images. Journal of Imaging. 2026; 12(3):141. https://doi.org/10.3390/jimaging12030141
Chicago/Turabian StyleCao, Li, Zishang Liu, Yan Wang, and Run Gao. 2026. "DFENet: A Novel Dual-Path Feature Extraction Network for Semantic Segmentation of Remote Sensing Images" Journal of Imaging 12, no. 3: 141. https://doi.org/10.3390/jimaging12030141
APA StyleCao, L., Liu, Z., Wang, Y., & Gao, R. (2026). DFENet: A Novel Dual-Path Feature Extraction Network for Semantic Segmentation of Remote Sensing Images. Journal of Imaging, 12(3), 141. https://doi.org/10.3390/jimaging12030141

