RFM-UNet: Hybrid Frequency–Mamba UNet for Remote-Sensing Road Extraction
Highlights
- We propose RFM-UNet, a hybrid Frequency–Mamba network featuring an ADAMamba encoder that jointly captures local geometries and global dependencies at linear computational complexity.
- By integrating a Multi-Scale Adaptive Fusion Module (MAFM) and a Dual-Spectrum Aggregation Module (DualSpec), the network effectively preserves narrow-road connectivity and mitigates shadow occlusions, achieving a superior performance–efficiency balance across three public datasets.
- The success of the ADAMamba block demonstrates that rigorous topological continuity in remote-sensing imagery can be achieved efficiently without the quadratic computational bottleneck of traditional Vision Transformers.
- The proposed spatial–frequency synergy offers a highly robust approach for extracting continuous structures under severe occlusions, significantly reducing the need for manual post-processing in applications like GIS updating and autonomous navigation.
Abstract
1. Introduction
- 1.
- We propose an ADAMamba encoder to reconcile local structural details with global context modeling without incurring quadratic computational complexity. By integrating an Anisotropic Directional Attention (ADA) module with 2D cross-directional scanning, it jointly captures local geometric anisotropy and global semantic dependencies.
- 2.
- To mitigate semantic oversmoothing and cross-scale interference prevalent in hierarchical decoding, we design a Multi-Scale Adaptive Fusion Module (MAFM). This module dynamically recalibrates multi-stage features, thereby preserving edge connectivity and topological consistency of narrow roads during feature fusion.
- 3.
- To enhance robustness against shadow-induced occlusions and background noise characterized by similar textures, we introduce a Dual-Spectrum Aggregation Module (DualSpec). Employing FFT-based spatial–frequency decoupling, it independently refines phase and amplitude spectra before adaptively fusing them with spatial features, thereby significantly enhancing structural integrity.
2. Related Work
2.1. CNN- and Transformer-Based Road Extraction Methods
2.2. Application of Mamba Architectures in Deep Learning
2.3. Application of Frequency-Domain Information in Deep Learning
3. Methodology
3.1. Overall Architecture
3.2. ADAMamba Block
| Algorithm 1 ADAMamba Block |
|
3.3. Multi-Scale Adaptive Fusion Module (MAFM)
| Algorithm 2 Multi-Scale Adaptive Fusion Module (MAFM) |
|
3.4. Dual-Spectrum Aggregation Module (DualSpec)
4. Experimental Settings and Results
4.1. Datasets
- Massachusetts Roads: This dataset comprises 1171 high-resolution aerial images from the state of Massachusetts, each with a size of pixels, covering approximately per image and over in total. The road structures exhibit significant variations in width, shape, texture, and color, making the dataset highly challenging. Following the official data split, all images were padded to pixels and cropped into patches, resulting in 9972 training images, 126 validation images, and 441 test images.
- DeepGlobe Road: Released for the CVPR 2018 DeepGlobe Road Extraction Challenge, this dataset contains 8570 high-resolution satellite images from regions such as India, Indonesia, and Thailand. Each image has a size of pixels with a spatial resolution of , covering diverse landscapes from urban to rural areas. Since official test labels were not publicly available, we utilized 6226 labeled images and split them into 4696 training samples and 1530 testing samples for our experiments. All images were directly cropped into patches for training and evaluation.
- CHN6-CUG Roads: Developed by the URSmart team at the China University of Geosciences (Wuhan), this VHR imagery dataset contains pixel-level manual annotations. It covers six representative urban areas: Beijing, Shanghai, Wuhan, Shenzhen, Hong Kong, and Macau, capturing diverse urbanization levels and city characteristics. Road annotations include both covered and uncovered roads, categorized according to coverage extent, and encompass various types such as railways, highways, urban roads, and rural roads. The dataset consists of 4511 images of size pixels, with 3608 used for training and 903 reserved for testing and evaluation.
4.2. Evaluation Metrics
4.3. Experimental Setup
4.4. Comparative Experiments with Different Image Sizes
4.5. Experimental Results
4.6. Experimental Analysis of Road Topology Accuracy
4.7. Model Complexity and Computational Efficiency Analysis
4.8. Ablation Study
4.8.1. Effectiveness of ADAMamba
4.8.2. Effectiveness of MAFM
4.8.3. Effectiveness of DualSpec
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Barzohar, M.; Cooper, D.B. Automatic finding of main roads in aerial images by using geometric-stochastic models and estimation. IEEE Trans. Pattern Anal. Mach. Intell. 2002, 18, 707–721. [Google Scholar]
- Chai, D.; Forstner, W.; Lafarge, F. Recovering line-networks in images by junction-point processes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA, 23–28 June 2013; IEEE: Piscataway, NJ, USA, 2013; pp. 1894–1901. [Google Scholar]
- Wang, D.; Li, J.; Zhu, S. Detecting urban hot regions by using massive geo-tagged image data. Neurocomputing 2021, 428, 325–331. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015: 18th International Conference, Munich, Germany, 5–9 October 2015; Proceedings, Part III 18; Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar]
- Fan, Y.; Hong, C.; Zeng, G.; Liu, L. A deep convolutional encoder–decoder–restorer architecture for image deblurring. Neural Process. Lett. 2024, 56, 27. [Google Scholar] [CrossRef] [Scilit]
- Jiang, X.; Li, Y.; Jiang, T.; Xie, J.; Wu, Y.; Cai, Q.; Jiang, J.; Xu, J.; Zhang, H. RoadFormer: Pyramidal deformable vision transformers for road network extraction with remote sensing images. Int. J. Appl. Earth Obs. Geoinf. 2022, 113, 102987. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Yin, H.; Zhang, D. A self-adaptive classification method for plant disease detection using GMDH-Logistic model. Sustain. Comput. Inform. Syst. 2020, 28, 100415. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. [Google Scholar] [CrossRef] [Scilit]
- Weng, W.; Hou, F.; Gong, S.; Chen, F.; Lin, D. Attribute graph clustering via transformer and graph attention autoencoder. Intell. Data Anal. 2025, 29, 306–319. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 10012–10022. [Google Scholar]
- Liu, L.; Li, P.; Wang, D.; Zhu, S. A wind turbine damage detection algorithm designed based on YOLOv8. Appl. Soft Comput. 2024, 154, 111364. [Google Scholar] [CrossRef] [Scilit]
- Lin, R.; He, Y.; Xu, M. Method of sensitive data mining based on Pan-Bull algebra. Wirel. Netw. 2022, 28, 2733–2741. [Google Scholar]
- Xie, Y.; Hong, C.; Zhuang, W.; Liu, L.; Li, J. HOGFormer: High-order graph convolution transformer for 3D human pose estimation. Int. J. Mach. Learn. Cybern. 2025, 16, 599–610. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. In Proceedings of the First Conference on Language Modeling, Philadelphia, PA, USA, 7–9 October 2024. [Google Scholar]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024, arXiv:2401.09417. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. Vmamba: Visual state space model. Adv. Neural Inf. Process. Syst. 2024, 37, 103031–103063. [Google Scholar] [CrossRef] [Scilit]
- Zhao, S.; Chen, H.; Zhang, X.; Xiao, P.; Bai, L.; Ouyang, W. Rs-mamba for large remote sensing image dense prediction. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5633314. [Google Scholar] [CrossRef] [Scilit]
- Chen, K.; Chen, B.; Liu, C.; Li, W.; Zou, Z.; Shi, Z. Rsmamba: Remote sensing image classification with state space model. IEEE Geosci. Remote Sens. Lett. 2024, 21, 8002605. [Google Scholar] [CrossRef] [Scilit]
- Liu, M.; Dan, J.; Lu, Z.; Yu, Y.; Li, Y.; Li, X. CM-UNet: Hybrid CNN-Mamba UNet for remote sensing image semantic segmentation. arXiv 2024, arXiv:2405.10530. [Google Scholar]
- Zheng, Y.; Zheng, W.; Du, X. Paddy-YOLO: An accurate method for rice pest detection. Comput. Electron. Agric. 2025, 238, 110777. [Google Scholar] [CrossRef] [Scilit]
- Yu, S.; Wu, Y.; Li, W.; Song, Z.; Zeng, W. A model for fine-grained vehicle classification based on deep learning. Neurocomputing 2017, 257, 97–103. [Google Scholar] [CrossRef] [Scilit]
- Shunzhi, Z.; Lizhao, L.; Si, C. Image feature detection algorithm based on the spread of hessian source. Multimed. Syst. 2017, 23, 105–117. [Google Scholar]
- Chen, Y.; Feng, S.; Zhao, C.; Su, N.; Li, W.; Tao, R.; Ren, J. High-resolution remote sensing image change detection based on Fourier feature interaction and multi-scale perception. IEEE Trans. Geosci. Remote Sens. 2024, 21, 8002605. [Google Scholar]
- Xu, C.; Ye, Z.; Mei, L.; Yu, H.; Liu, J.; Yalikun, Y.; Jin, S.; Liu, S.; Yang, W.; Lei, C. Hybrid attention-aware transformer network collaborative multiscale feature alignment for building change detection. IEEE Trans. Instrum. Meas. 2024, 73, 5012914. [Google Scholar] [CrossRef] [Scilit]
- Zhu, X.; Hu, H.; Lin, S.; Dai, J. Deformable convnets v2: More deformable, better results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 9308–9316. [Google Scholar]
- Lee, S.; Bae, J.; Kim, H.Y. Decompose, adjust, compose: Effective normalization by playing with frequency for domain generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 11776–11785. [Google Scholar]
- Cao, B.; Wang, Q.; Zhu, P.; Hu, Q.; Ren, D.; Zuo, W.; Gao, X. Multi-view knowledge ensemble with frequency consistency for cross-domain face translation. IEEE Trans. Neural Netw. Learn. Syst. 2023, 35, 9728–9742. [Google Scholar]
- Zhang, G.; Zhang, W.; Li, W.; Wang, L.; Cui, H. A dynamic attention mechanism for object detection in road or strip environments. Vis. Comput. 2025, 41, 4171–4181. [Google Scholar]
- Xiong, F.; Li, T.; Yang, Y.; Zhou, J.; Lu, J.; Qian, Y. Wavelet siamese network with semi-supervised domain adaptation for remote sensing image change detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5633613. [Google Scholar] [CrossRef] [Scilit]
- Yi, J.; Han, W.; Lai, F. YOLOv8n-DDSW: An efficient fish target detection network for dense underwater scenes. PeerJ Comput. Sci. 2025, 11, e2798. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Hinton, G.E. Learning to detect roads in high-resolution aerial images. In Proceedings of the Computer Vision—ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, 5–11 September 2010; Proceedings, Part VI 11; Springer: Cham, Switzerland, 2010; pp. 210–223. [Google Scholar]
- Long, J.; Shelhamer, E.; Darrell, T. Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 3431–3440. [Google Scholar]
- Zhou, L.; Zhang, C.; Wu, M. D-LinkNet: LinkNet with pretrained encoder and dilated convolution for high resolution satellite imagery road extraction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA, 18–22 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 182–186. [Google Scholar]
- Li, Y.; Peng, B.; He, L.; Fan, K.; Li, Z.; Tong, L. Road extraction from unmanned aerial vehicle remote sensing images based on improved neural networks. Sensors 2019, 19, 4115. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yi, J.; Shen, Z.; Chen, F.; Zhao, Y.; Xiao, S.; Zhou, W. A lightweight multiscale feature fusion network for remote sensing object counting. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5902113. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Bai, L.; Xue, D.; Momi, M.C.; Ye, Z.; Quan, S. FRCFNet: Feature reassembly and context information fusion network for road extraction. IEEE Geosci. Remote Sens. Lett. 2024, 21, 2502805. [Google Scholar] [CrossRef] [Scilit]
- Hua, Z.T.; Chen, S.B.; Lu, W.; Tang, J.; Luo, B. Multi-Scale Adaptive Decoder and Diverse Selection for Road Extraction in Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2025, 62, 4411713. [Google Scholar]
- Li, R.; Zheng, S.; Zhang, C.; Duan, C.; Wang, L.; Atkinson, P.M. ABCNet: Attentive bilateral contextual network for efficient semantic segmentation of Fine-Resolution remotely sensed imagery. ISPRS J. Photogramm. Remote Sens. 2021, 181, 84–98. [Google Scholar] [CrossRef] [Scilit]
- Luo, L.; Wang, J.X.; Chen, S.B.; Tang, J.; Luo, B. BDTNet: Road extraction by bi-direction transformer from remote sensing images. IEEE Geosci. Remote Sens. Lett. 2022, 19, 2505605. [Google Scholar] [CrossRef] [Scilit]
- Yang, R.; Zhong, Y.; Liu, Y.; Lu, X.; Zhang, L. Occlusion-aware road extraction network for high-resolution remote sensing imagery. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5619316. [Google Scholar] [CrossRef] [Scilit]
- Yu, Z.; Chen, Z.; Xiao, K.; Lei, X.; Tang, R.; He, Q.; Sun, Z.; Guo, H. SegRoadv2: A hybrid deformable self-attention and convolutional network for road extraction with connectivity structure. Int. J. Digit. Earth 2025, 18, 2480267. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Yu, P.; Li, M.; Zhong, X.; He, Y.; Su, H.; Zhou, Y. TDFNet: Twice decoding V-Mamba-CNN Fusion features for building extraction. Geo-Spat. Inf. Sci. 2026, 29, 19–38. [Google Scholar]
- Wang, L.; Li, D.; Dong, S.; Meng, X.; Zhang, X.; Hong, D. PyramidMamba: Rethinking pyramid feature fusion with selective space state model for semantic segmentation of remote sensing imagery. Int. J. Appl. Earth Obs. Geoinf. 2025, 144, 104884. [Google Scholar] [CrossRef] [Scilit]
- Zhou, M.; Huang, J.; Yan, K.; Hong, D.; Jia, X.; Chanussot, J.; Li, C. A general spatial-frequency learning framework for multimodal image fusion. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 47, 5281–5298. [Google Scholar] [CrossRef] [Scilit]
- Tang, Y.; Feng, S.; Zhao, C.; Fan, Y.; Shi, Q.; Li, W.; Tao, R. An object fine-grained change detection method based on frequency decoupling interaction for high-resolution remote sensing images. IEEE Trans. Geosci. Remote Sens. 2023, 62, 5600213. [Google Scholar]
- Yu, B.; Yang, J.; Du, Z.; Huang, Y.; Li, C.; Wang, L. Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Daejeon, South Korea, 23–27 September 2025; Springer: Cham, Switzerland, 2025; pp. 278–288. [Google Scholar]
- Song, P.; Yu, P.; Zhong, X.; Wang, L.; He, Y.; Li, G.; Qi, Z.; Su, H. A Cross-Domain Mamba Network with joint spatial-frequency learning for robust SAR oil spill detection. Mar. Pollut. Bull. 2026, 229, 119729. [Google Scholar] [CrossRef] [Scilit]
- Xue, D.; Bai, L.; Zhou, W.; Zhang, Y.; Gozho, A.; Quan, S. SDFFNet: A Spatial-Frequency Domain Feature Fusion Network for Road Extraction From Very High-Resolution Satellite Images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2026, 19, 18207–18223. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V. Machine Learning for Aerial Image Labeling. Ph.D. Thesis, University of Toronto, Toronto, ON, Canada, 2013. [Google Scholar]
- Demir, I.; Koperski, K.; Lindenbaum, D.; Pang, G.; Huang, J.; Basu, S.; Hughes, F.; Tuia, D.; Raskar, R. Deepglobe 2018: A challenge to parse the earth through satellite images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA, 18–22 June 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 172–181. [Google Scholar]
- Zhu, Q.; Zhang, Y.; Wang, L.; Zhong, Y.; Guan, Q.; Lu, X.; Zhang, L.; Li, D. A global context-aware and batch-independent network for road extraction from VHR satellite imagery. ISPRS J. Photogramm. Remote Sens. 2021, 175, 353–365. [Google Scholar] [CrossRef] [Scilit]
- Chaurasia, A.; Culurciello, E. Linknet: Exploiting Encoder Representations for Efficient Semantic Segmentation. In Proceedings of the 2017 IEEE Visual Communications and Image Processing (VCIP), St. Petersburg, FL, USA, 10–13 December 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 1–4. [Google Scholar]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; Springer: Cham, Switzerland, 2018; pp. 801–818. [Google Scholar]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. Transunet: Transformers Make Strong Encoders for Medical Image Segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef] [Scilit]
- Cao, H.; Wang, Y.; Chen, J.; Jiang, D.; Zhang, X.; Tian, Q.; Wang, M. Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; Springer: Cham, Switzerland, 2022; pp. 205–218. [Google Scholar]
- Yang, Z.; Zhou, D.; Yang, Y.; Zhang, J.; Chen, Z. Road extraction from satellite imagery by road context and full-stage feature. IEEE Geosci. Remote Sens. Lett. 2022, 20, 8000405. [Google Scholar] [CrossRef] [Scilit]
- Chen, H.; Li, Z.; Wu, J.; Xiong, W.; Du, C. SemiRoadExNet: A Semi-Supervised Network for Road Extraction from Remote Sensing Imagery via Adversarial Learning. ISPRS J. Photogramm. Remote Sens. 2023, 198, 169–183. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Zhang, X.; Pun, M.O. Rs 3 mamba: Visual state space model for remote sensing image semantic segmentation. IEEE Geosci. Remote Sens. Lett. 2024, 21, 6011405. [Google Scholar] [CrossRef] [Scilit]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 618–626. [Google Scholar]












| Dataset | GSD (m)/Bands | Patch Size | Train | Val/Test | Region | Release |
|---|---|---|---|---|---|---|
| Massachusetts | 1/RGB | 1108 | 14/49 | Massachusetts | 2013 | |
| DeepGlobe | 0.5/RGB | 4696 | -/1530 | South and Southeast Asia | 2018 | |
| CHN6-CUG | 0.5/RGB | 3608 | -/903 | China | 2021 |
| Parameter | Value | Description |
|---|---|---|
| Language | Python 3.12 | Programming language used |
| Framework | PyTorch 2.4.0 | Deep-learning framework |
| GPU | NVIDIA RTX 4090 | Graphics processing unit used for training |
| Random Seed | 42 | Seed for reproducibility |
| Optimizer | AdamW | Optimization algorithm |
| Learning Rate | Initial learning rate (Massachusetts) | |
| Batch Size | 12 | Number of samples per batch (Massachusetts) |
| Epochs | 105 | Number of training epochs |
| Patch Size | P (%) | R (%) | F1 (%) | IoU (%) | cIDice (%) | APLS (%) |
|---|---|---|---|---|---|---|
| 79.20 | 78.82 | 79.01 | 65.30 | 86.98 | 69.58 | |
| 80.56 | 78.75 | 79.65 | 66.18 | 87.84 | 72.62 | |
| 81.51 | 81.24 | 81.37 | 68.59 | 89.13 | 74.04 | |
| 81.20 | 81.13 | 81.15 | 68.27 | 89.09 | 73.71 |
| Method | Year | Massachusetts Roads | DeepGlobe Road | CHN6-CUG Roads | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| P | R | F1 | IoU | P | R | F1 | IoU | P | R | F1 | IoU | ||
| UNet | 2015 | 77.29 | 72.13 | 74.62 | 59.51 | 78.29 | 70.92 | 74.42 | 59.26 | 70.67 | 71.66 | 71.16 | 55.23 |
| LinkNet | 2017 | 79.18 | 73.71 | 75.52 | 61.74 | 75.62 | 73.75 | 74.67 | 59.85 | 74.92 | 71.43 | 73.13 | 57.64 |
| D-LinkNet | 2018 | 77.88 | 74.52 | 76.16 | 61.50 | 77.01 | 76.27 | 76.64 | 62.12 | 76.28 | 71.89 | 74.01 | 58.75 |
| DeepLabV3+ | 2018 | 75.47 | 77.97 | 76.70 | 62.25 | 78.20 | 76.24 | 77.20 | 62.87 | 75.21 | 74.59 | 74.90 | 59.87 |
| TransUNet | 2021 | 78.37 | 77.29 | 77.82 | 63.70 | 77.80 | 77.02 | 77.41 | 63.14 | 74.32 | 77.02 | 75.64 | 60.83 |
| SwinUNet | 2022 | 78.12 | 77.07 | 77.59 | 63.59 | 78.44 | 77.30 | 77.87 | 63.76 | 74.43 | 76.47 | 75.44 | 60.56 |
| UNet* | 2022 | 79.81 | 78.08 | 78.93 | 65.19 | 78.63 | 77.82 | 78.22 | 64.23 | 74.81 | 76.79 | 75.78 | 61.01 |
| RCFSNet | 2022 | 79.23 | 77.62 | 78.43 | 64.49 | 78.85 | 79.35 | 79.10 | 65.43 | 75.65 | 77.29 | 76.46 | 61.89 |
| RoadExNet | 2023 | 80.38 | 75.98 | 78.18 | 64.17 | 76.94 | 79.42 | 78.16 | 64.15 | 75.23 | 77.30 | 76.25 | 61.61 |
| OARENet | 2024 | 77.48 | 79.92 | 78.70 | 64.88 | 81.72 | 81.95 | 81.84 | 69.26 | 77.83 | 72.43 | 75.03 | 60.04 |
| RS-Mamba | 2024 | 86.27 | 76.03 | 80.82 | 67.82 | 83.14 | 80.00 | 81.54 | 68.84 | 74.63 | 79.77 | 77.11 | 62.75 |
| RS3Mamba | 2024 | 80.87 | 77.65 | 79.26 | 65.65 | 80.45 | 79.30 | 79.87 | 66.49 | 77.82 | 72.80 | 75.23 | 60.29 |
| SegRoadv2 | 2025 | 79.84 | 79.02 | 79.43 | 65.87 | 81.06 | 80.98 | 81.02 | 68.09 | 72.10 | 78.76 | 75.28 | 60.36 |
| MADSNet | 2025 | 79.40 | 81.36 | 80.38 | 67.18 | 80.58 | 80.79 | 80.69 | 67.63 | 77.49 | 79.11 | 78.29 | 64.33 |
| RFM-UNet | 81.51 | 81.24 | 81.37 | 68.59 | 81.77 | 83.95 | 82.84 | 70.71 | 80.07 | 76.87 | 78.44 | 64.52 | |
| Method | Massachusetts Roads | DeepGlobe Road | CHN6-CUG Roads | |||
|---|---|---|---|---|---|---|
| cIDice (%) | APLS (%) | cIDice (%) | APLS (%) | cIDice (%) | APLS (%) | |
| UNet | 83.98 | 69.12 | 84.62 | 66.45 | 71.66 | 64.50 |
| LinkNet | 84.61 | 69.98 | 84.90 | 67.05 | 73.92 | 65.28 |
| D-LinkNet | 84.44 | 68.79 | 85.32 | 67.20 | 74.53 | 65.74 |
| DeepLabV3+ | 85.75 | 69.18 | 85.60 | 67.45 | 74.61 | 66.50 |
| TransUNet | 85.95 | 69.35 | 85.74 | 67.74 | 72.84 | 66.88 |
| SwinUNet | 85.90 | 69.26 | 86.12 | 67.82 | 75.06 | 67.10 |
| UNet* | 87.33 | 72.52 | 86.55 | 68.21 | 75.47 | 67.29 |
| RCFSNet | 86.85 | 71.27 | 86.89 | 68.17 | 76.65 | 67.89 |
| RoadExNet | 86.64 | 71.01 | 87.58 | 68.06 | 76.52 | 67.97 |
| OARENet | 87.09 | 71.69 | 89.89 | 69.05 | 75.53 | 68.04 |
| RS-Mamba | 88.82 | 73.20 | 89.15 | 68.92 | 77.63 | 70.15 |
| RS3Mamba | 87.72 | 72.12 | 88.18 | 68.03 | 75.82 | 68.29 |
| SegRoadv2 | 87.97 | 72.33 | 88.80 | 68.62 | 76.10 | 68.36 |
| MADSNet | 88.02 | 72.85 | 88.62 | 68.26 | 78.11 | 72.03 |
| RFM-UNet | 89.13 | 74.04 | 90.91 | 69.39 | 78.26 | 72.17 |
| Method | Backbone | Year | Params (M) | FLOPs (G) | Memory (MB) | FPS |
|---|---|---|---|---|---|---|
| LinkNet | CNN | 2017 | 21.64 | 23.26 | 186.28 | 97.19 |
| D-LinkNet | CNN | 2018 | 31.09 | 26.71 | 230.40 | 258.83 |
| MADSNet | CNN | 2025 | 29.73 | 32.58 | 449.42 | 92.20 |
| RoadExNet | ViT | 2023 | 59.20 | 277.40 | 1035.78 | 26.44 |
| OARENet | ViT | 2024 | 33.48 | 36.27 | 370.51 | 85.35 |
| SegRoadv2 | ViT | 2025 | 54.99 | 109.42 | 717.61 | 44.90 |
| RS-Mamba | SSM | 2024 | 50.37 | 46.33 | 644.54 | 48.87 |
| RS3Mamba | SSM | 2024 | 43.32 | 39.56 | 472.62 | 51.99 |
| RFM-UNet | SSM | 29.86 | 30.13 | 428.69 | 60.70 |
| Baseline | ADAMamba | MAFM | DualSpec | Params | FLOPs | P | R | F1 | IoU | cIDice | APLS |
|---|---|---|---|---|---|---|---|---|---|---|---|
| √ | 23.13 M | 24.09 G | 74.21 | 74.33 | 74.27 | 59.07 | 73.65 | 65.49 | |||
| √ | √ | 25.55 M | 26.78 G | 75.42 | 77.53 | 76.47 | 60.54 | 74.49 | 67.75 | ||
| √ | √ | 26.39 M | 26.92 G | 74.14 | 78.04 | 76.04 | 61.34 | 75.78 | 69.72 | ||
| √ | √ | 24.20 M | 24.60 G | 76.44 | 78.49 | 77.45 | 63.20 | 76.43 | 71.86 | ||
| √ | √ | √ | √ | 29.86 M | 30.13 G | 80.07 | 76.87 | 78.44 | 64.52 | 78.26 | 72.17 |
| Model | P | R | F1 | IoU | cIDice | APLS |
|---|---|---|---|---|---|---|
| Spatial-only | 74.84 | 78.67 | 76.71 | 62.21 | 75.92 | 69.33 |
| Frequency-only | 75.43 | 78.66 | 77.01 | 62.62 | 76.31 | 70.48 |
| DualSpec | 76.44 | 78.49 | 77.45 | 63.20 | 76.43 | 71.86 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Song, P.; Yu, P.; Zhong, X.; Wu, S.; Tong, J.; He, Y.; Zhang, L.; Li, G.; Li, M. RFM-UNet: Hybrid Frequency–Mamba UNet for Remote-Sensing Road Extraction. Remote Sens. 2026, 18, 2546. https://doi.org/10.3390/rs18152546
Song P, Yu P, Zhong X, Wu S, Tong J, He Y, Zhang L, Li G, Li M. RFM-UNet: Hybrid Frequency–Mamba UNet for Remote-Sensing Road Extraction. Remote Sensing. 2026; 18(15):2546. https://doi.org/10.3390/rs18152546
Chicago/Turabian StyleSong, Pu, Peng Yu, Xiaojing Zhong, Shuizhen Wu, Junbin Tong, Yuanrong He, Lujun Zhang, Guangchun Li, and Mengmeng Li. 2026. "RFM-UNet: Hybrid Frequency–Mamba UNet for Remote-Sensing Road Extraction" Remote Sensing 18, no. 15: 2546. https://doi.org/10.3390/rs18152546
APA StyleSong, P., Yu, P., Zhong, X., Wu, S., Tong, J., He, Y., Zhang, L., Li, G., & Li, M. (2026). RFM-UNet: Hybrid Frequency–Mamba UNet for Remote-Sensing Road Extraction. Remote Sensing, 18(15), 2546. https://doi.org/10.3390/rs18152546

