DSD-Mamba: Dual-Stream Semantic Segmentation of Remote Sensing Imagery via Dense-Sparse Fusion
Abstract
1. Introduction
- We propose DSD-Mamba, a hybrid U-shaped framework for high-resolution remote sensing semantic segmentation. The model integrates CNN-based local representation, Mamba-based global sequence modeling, and Top-k selective feature aggregation to address complex urban scenes.
- We design the Dense-Sparse Pyramid Fusion Module (DSPFM) for bottleneck feature fusion. DSPFM performs dense multi-scale alignment and triangular cross-scale interaction, followed by Top-k selective value aggregation to reduce the influence of redundant background responses. We clarify that this operation is used for feature selection rather than theoretical memory reduction.
- We introduce Scale-Aware Strip Attention (SASA) in the skip-fusion pathway. SASA models horizontal and vertical contextual dependencies and is designed to improve the geometric continuity of elongated regions such as roads, rivers, and bridges.
- We develop the Dual-Stream Context Decoder (DSCD), which combines a Mamba-based global semantic branch with a CNN-based local-detail branch. This design decouples global semantic consistency from local boundary refinement during upsampling.
- We conduct experiments on UAVid, ISPRS Vaihingen, and ISPRS Potsdam, together with ablation and complexity analyses, to evaluate the accuracy–cost trade-off of the proposed method under the adopted high-resolution evaluation protocols.
2. Related Work
2.1. Machine-Learning-Based Land-Use Identification
2.2. CNN-Based Semantic Segmentation
2.3. Transformer-Based Dense Prediction
2.4. State Space Models for Vision
3. Methodology
3.1. Overall Architecture
3.2. Dense-Sparse Pyramid Fusion Module (DSPFM)
3.2.1. Scale Alignment and CNN Residual Fusion
3.2.2. MSC-Based Selective Dense-Sparse Cross-Scale Interaction
3.3. Scale-Aware Strip Attention (SASA)
3.4. Dual-Stream Context Decoder (DSCD)
3.4.1. Global Mamba Stream
3.4.2. Local Context Stream
3.4.3. Cosine-Guided Global–Local Fusion
4. Experiments
4.1. Datasets and Data Processing
4.2. Implementation Details and Inference Strategy
- ISPRS Vaihingen Dataset: The model was trained for 250 epochs with a batch size of 4. The loss was the same joint Soft Cross-Entropy and Dice Loss. Cosine annealing with warm restarts was used with and [38].
- ISPRS Potsdam Dataset: The model was trained for 600 epochs with a batch size of 8. Cosine annealing with warm restarts was used with and . Dice Loss with a smoothing factor of 0.05 was used to optimize regional overlap.
4.3. Evaluation Metrics
4.4. Comparative Analysis
4.4.1. Performance on UAVid Dataset
4.4.2. Performance on ISPRS Vaihingen Dataset
4.4.3. Performance on ISPRS Potsdam Dataset
4.4.4. Class-Wise Performance Analysis
4.5. Ablation Study
4.6. Generalizability Across Datasets
4.7. Computational Complexity and Accuracy–Cost Trade-Off
4.8. Practical Applications and Deployment Considerations
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Lyu, Y.; Vosselman, G.; Xia, G.S.; Yilmaz, A.; Yang, M.Y. UAVid: A semantic segmentation dataset for UAV imagery. ISPRS J. Photogramm. Remote Sens. 2020, 165, 108–119. [Google Scholar] [CrossRef]
- Rottensteiner, F.; Sohn, G.; Jung, J.; Gerke, M.; Baillard, C.; Benitez, S.; Breitkopf, U. The ISPRS benchmark on urban object detection and 3D building reconstruction. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2012, I-3, 293–298. [Google Scholar] [CrossRef]
- Li, R.; Zheng, S.; Zhang, C.; Duan, C.; Wang, L.; Atkinson, P.M. ABCNet: Attentive bilateral contextual network for efficient semantic segmentation of fine-resolution remotely sensed imagery. ISPRS J. Photogramm. Remote Sens. 2021, 181, 84–98. [Google Scholar] [CrossRef]
- Wang, L.; Li, R.; Zhang, C.; Fang, S.; Duan, C.; Meng, X.; Atkinson, P.M. UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery. ISPRS J. Photogramm. Remote Sens. 2022, 190, 196–214. [Google Scholar] [CrossRef]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Chen, L.-C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 40, 834–848. [Google Scholar] [CrossRef] [PubMed]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Houlsby, N. An image is worth 16 × 16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Zhou, Y. TransUNet: Transformers make strong encoders for medical image segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar]
- Strudel, R.; Garcia, R.; Laptev, I.; Schmid, C. Segmenter: Transformer for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 7242–7252. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. VMamba: Visual state space model. arXiv 2024, arXiv:2401.10166. [Google Scholar]
- Ruan, J.; Li, J.; Xiang, S. VM-UNet: Vision mamba UNet for medical image segmentation. arXiv 2024, arXiv:2402.02491. [Google Scholar]
- Wang, Z.; Zheng, J.Q.; Zhang, Y.; Cui, G.; Li, L. Mamba-UNet: UNet-like pure visual mamba for medical image segmentation. arXiv 2024, arXiv:2402.05079. [Google Scholar]
- Liu, J.; Yang, H.; Zhou, H.-Y.; Xi, Y.; Yu, L.; Li, C.; Liang, Y.; Shi, G.; Yu, Y.; Zhang, S.; et al. Swin-UMamba: Mamba-based UNet with ImageNet-based pretraining. arXiv 2024, arXiv:2402.03302. [Google Scholar]
- Ma, X.; Zhang, X.; Pun, M.-O. RS3Mamba: Visual state space model for remote sensing image semantic segmentation. IEEE Geosci. Remote Sens. Lett. 2024, 21, 6011405. [Google Scholar] [CrossRef]
- Li, L.; Yi, J.; Fan, H.; Lin, H. A Lightweight Semantic Segmentation Network Based on Self-Attention Mechanism and State Space Model for Efficient Urban Scene Segmentation. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4703215. [Google Scholar] [CrossRef]
- Li, Z.; Zhao, L.; Lu, Y.; Ma, Y.; Li, G. Mamba for Remote Sensing: Architectures, Hybrid Paradigms, and Future Directions. Remote Sens. 2026, 18, 243. [Google Scholar] [CrossRef]
- Bao, M.; Lyu, S.; Xu, Z.; Zhou, H.; Ren, J.; Xiang, S.; Li, X.; Cheng, G. Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook. Remote Sens. 2026, 18, 594. [Google Scholar] [CrossRef]
- Cao, Y.; Liu, C.; Wu, Z.; Zhang, L.; Yang, L. Remote Sensing Image Segmentation Using Vision Mamba and Multi-Scale Multi-Frequency Feature Fusion. Remote Sens. 2025, 17, 1390. [Google Scholar] [CrossRef]
- Li, B.; Yang, X.; Fan, Y. MAFMamba: A Multi-Scale Adaptive Fusion Network for Semantic Segmentation of High-Resolution Remote Sensing Images. Sensors 2026, 26, 531. [Google Scholar] [CrossRef] [PubMed]
- Meedeniya, D.A.; Jayanetti, J.A.A.M.; Dilini, M.D.N.; Wickramapala, M.H.; Madushanka, J.H. Land-Use Classification with Integrated Data. In Machine Vision Inspection Systems: Image Processing, Concepts, Methodologies and Applications; Malarvel, M., Nayak, S.R., Panda, S.N., Pattnaik, P.K., Muangnak, N., Eds.; John Wiley & Sons: New York, NY, USA, 2020; pp. 1–38. [Google Scholar] [CrossRef]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 3431–3440. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention (MICCAI), Munich, Germany, 5–9 October 2015; pp. 234–241. [Google Scholar]
- Oršić, M.; Šegvić, S. Efficient semantic segmentation with pyramidal fusion. Pattern Recognit. 2021, 110, 107611. [Google Scholar] [CrossRef]
- Li, R.; Zheng, S.; Zhang, C.; Duan, C.; Su, J.; Wang, L.; Atkinson, P.M. Multiattention network for semantic segmentation of fine-resolution remote sensing images. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5607713. [Google Scholar] [CrossRef]
- Zheng, X.; Huan, L.; Xia, G.-S.; Gong, J. Parsing very high resolution urban scene images by learning deep ConvNets with edge-aware loss. ISPRS J. Photogramm. Remote Sens. 2020, 170, 15–28. [Google Scholar] [CrossRef]
- Diakogiannis, F.I.; Waldner, F.; Caccetta, P.; Wu, C. ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data. ISPRS J. Photogramm. Remote Sens. 2020, 162, 94–114. [Google Scholar] [CrossRef]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 9992–10002. [Google Scholar]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Online, 6–14 December 2021; pp. 12077–12090. [Google Scholar]
- Wang, L.; Li, R.; Wang, D.; Duan, C.; Wang, T.; Meng, X. Transformer meets convolution: A bilateral awareness network for semantic segmentation of very fine resolution urban scene images. Remote Sens. 2021, 13, 3065. [Google Scholar] [CrossRef]
- Wu, H.; Huang, P.; Zhang, M.; Tang, W.; Yu, X. CMTFNet: CNN and multi-scale transformer fusion network for remote sensing image semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2023, 61, 2004612. [Google Scholar] [CrossRef]
- Hou, Q.; Zhang, L.; Cheng, M.M.; Feng, J. Strip pooling: Rethinking spatial pooling for scene parsing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 4003–4012. [Google Scholar]
- Yalniz, I.Z.; Jégou, H.; Chen, K.; Paluri, M.; Mahajan, D. Billion-scale semi-supervised learning for image classification. arXiv 2019, arXiv:1905.00546. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 1026–1034. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. In Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Zhang, M.R.; Lucas, J.; Hinton, G.; Ba, J. Lookahead optimizer: K steps forward, 1 step back. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
- Milletari, F.; Navab, N.; Ahmadi, S.A. V-Net: Fully convolutional neural networks for volumetric medical image segmentation. In Proceedings of the International Conference on 3D Vision (3DV), Stanford, CA, USA, 25–28 October 2016; pp. 565–571. [Google Scholar]
- Loshchilov, I.; Hutter, F. SGDR: Stochastic gradient descent with warm restarts. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Srinivas, A.; Lin, T.-Y.; Parmar, N.; Shlens, J.; Abbeel, P.; Vaswani, A. Bottleneck transformers for visual recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Montreal, QC, Canada, 10–17 October 2021; pp. 16514–16524. [Google Scholar]
- Xu, W.; Xu, Y.; Chang, T.; Tu, Z. Co-scale conv-attentional image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 9961–9970. [Google Scholar]
- Lu, W.; Chen, S.-B.; Shu, Q.-L.; Tang, J.; Luo, B. DecoupleNet: A lightweight backbone network with efficient feature decoupling for remote sensing visual tasks. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4414613. [Google Scholar] [CrossRef]
- Wang, Z.; Yi, J.; Chen, A.; Chen, L.; Lin, H.; Xu, K. Accurate semantic segmentation of very high-resolution remote sensing images considering feature state sequences. ISPRS J. Photogramm. Remote Sens. 2025, 220, 824–840. [Google Scholar] [CrossRef]
- Hu, P.; Perazzi, F.; Heilbron, F.C.; Wang, O.; Lin, Z.; Saenko, K.; Sclaroff, S. Real-time semantic segmentation with fast attention. IEEE Robot. Autom. Lett. 2021, 6, 263–270. [Google Scholar] [CrossRef]
- Li, R.; Zheng, S.; Duan, C.; Su, J.; Zhang, C. Multistage attention ResU-net for semantic segmentation of fine-resolution remote sensing images. IEEE Geosci. Remote Sens. Lett. 2022, 19, 8009205. [Google Scholar] [CrossRef]
- Li, X.; Xie, L.; Wang, C.; Miao, J.; Shen, H.; Zhang, L. Boundary-enhanced dual-stream network for semantic segmentation of high-resolution remote sensing images. GISci. Remote Sens. 2024, 61, 2356355. [Google Scholar] [CrossRef]
- Li, B.; Zhang, Y.; Zhang, Y.; Li, B.; Li, Z. Dual-path feature fusion network for semantic segmentation of remote sensing images. IEEE Geosci. Remote Sens. Lett. 2024, 21, 2503205. [Google Scholar] [CrossRef]
- Li, X.; Wen, C.; Wang, L.; Fang, Y. Geometry-aware segmentation of remote sensing images via joint height estimation. IEEE Geosci. Remote Sens. Lett. 2022, 19, 8007905. [Google Scholar] [CrossRef]
- Wang, X.; Zhang, Y.; Lei, T.; Wang, Y.; Zhai, Y.; Nandi, A.K. Dynamic convolution self-attention network for land-cover classification in VHR remote-sensing images. Remote Sens. 2022, 14, 4941. [Google Scholar] [CrossRef]
- Zhang, X.; Weng, Z.; Zhu, P.; Han, X.; Zhu, J.; Jiao, L. ESDINet: Efficient shallow-deep interaction network for semantic segmentation of high-resolution aerial images. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5607615. [Google Scholar] [CrossRef]
- UAVid Semantic Segmentation Dataset. Available online: https://uavid.nl/ (accessed on 26 May 2026).
- ISPRS 2D Semantic Labeling Contest—Vaihingen. Available online: https://www.isprs.org/resources/datasets/benchmarks/UrbanSemLab/2d-sem-label-vaihingen.aspx (accessed on 26 May 2026).
- ISPRS 2D Semantic Labeling Contest—Potsdam. Available online: https://www.isprs.org/resources/datasets/benchmarks/UrbanSemLab/2d-sem-label-potsdam.aspx (accessed on 26 May 2026).







| Method | Clutter | Bui. | Road | Tree | Veg. | Mo.Car | St.Car | Human | mIoU |
|---|---|---|---|---|---|---|---|---|---|
| SwiftNet [24] | 64.1 | 85.3 | 61.5 | 78.3 | 76.4 | 51.1 | 62.1 | 15.7 | 61.1 |
| MANet [25] | 64.5 | 85.4 | 77.8 | 77.0 | 60.3 | 67.2 | 53.6 | 14.9 | 62.6 |
| ABCNet [3] | 67.4 | 86.4 | 81.2 | 79.9 | 63.1 | 69.8 | 48.4 | 13.9 | 63.8 |
| Segmenter [9] | 64.2 | 84.4 | 79.8 | 76.1 | 57.6 | 69.2 | 34.5 | 14.2 | 58.7 |
| SegFormer [29] | 66.6 | 86.3 | 80.1 | 79.6 | 62.3 | 72.5 | 52.5 | 28.5 | 66.0 |
| BANet [30] | 66.7 | 85.4 | 80.7 | 78.9 | 62.1 | 69.3 | 52.8 | 21.0 | 64.6 |
| BoTNet [39] | 64.5 | 84.9 | 78.6 | 77.4 | 60.5 | 65.8 | 51.9 | 22.4 | 63.2 |
| CoaT [40] | 69.0 | 88.5 | 80.0 | 79.3 | 62.0 | 70.0 | 59.1 | 18.9 | 65.9 |
| DecoupleNet [41] | 65.1 | 85.4 | 80.6 | 78.8 | 62.1 | 74.1 | 49.7 | 30.8 | 65.8 |
| UNetFormer [4] | 64.4 | 85.2 | 78.8 | 79.2 | 62.6 | 71.1 | 60.7 | 29.4 | 66.4 |
| Mamba-UNet [13] | 56.6 | 78.5 | 76.1 | 71.5 | 53.0 | 66.7 | 40.5 | 15.9 | 57.3 |
| Swin-UMamba [14] | 52.3 | 76.8 | 73.9 | 71.0 | 50.3 | 61.6 | 26.1 | 15.4 | 53.4 |
| VM-UNet [12] | 56.0 | 78.1 | 75.5 | 73.1 | 54.9 | 65.1 | 31.0 | 11.5 | 55.7 |
| UrbanSSF-T [42] | 65.0 | 85.7 | 79.1 | 79.6 | 63.5 | 69.0 | 54.7 | 29.3 | 65.7 |
| UMFormer [16] | 66.4 | 86.7 | 78.9 | 79.2 | 62.2 | 71.4 | 65.3 | 30.3 | 67.6 |
| DSD-Mamba (Ours) | 66.3 | 91.5 | 81.3 | 78.8 | 69.9 | 76.7 | 73.0 | 50.0 | 73.4 |
| Method | Imp.surf. | Bui. | Lowveg. | Tree | Car | MeanF1 | OA | mIoU |
|---|---|---|---|---|---|---|---|---|
| SwiftNet [24] | 92.2 | 94.8 | 84.1 | 89.3 | 81.2 | 88.3 | 90.2 | 79.6 |
| ABCNet [3] | 92.7 | 95.2 | 84.5 | 89.7 | 85.3 | 89.5 | 90.7 | 81.3 |
| FANet [43] | 90.7 | 93.8 | 82.6 | 88.6 | 71.6 | 85.4 | 88.9 | 75.6 |
| EaNet [26] | 91.7 | 94.5 | 83.1 | 89.2 | 80.0 | 87.7 | 89.7 | 78.7 |
| MAResU-Net [44] | 92.0 | 95.0 | 83.7 | 89.3 | 78.3 | 87.7 | 90.1 | 78.6 |
| BEDSN [45] | 92.3 | 94.7 | 83.7 | 89.2 | 86.3 | 89.2 | 90.1 | 80.8 |
| DPFE-AFF [46] | 93.3 | 96.0 | 84.7 | 90.2 | 88.3 | 90.4 | 91.3 | 82.9 |
| GANet [47] | 93.1 | 95.9 | 84.6 | 90.1 | 88.4 | 90.4 | 91.3 | – |
| TransUNet [8] | 90.8 | 94.3 | 79.0 | 90.5 | 82.7 | 87.5 | – | 78.2 |
| Swin-UperNet [28] | 90.3 | 94.1 | 81.1 | 87.4 | 81.6 | 86.8 | 88.2 | 77.1 |
| Segmenter [9] | 89.8 | 93.0 | 81.2 | 88.9 | 67.6 | 84.1 | 88.1 | 73.6 |
| BoTNet [39] | 89.9 | 92.1 | 81.8 | 88.7 | 71.3 | 84.8 | 88.0 | 74.3 |
| CMTFNet [31] | 90.6 | 94.2 | 81.9 | 87.6 | 82.8 | 87.4 | 88.7 | 78.0 |
| DCSA-Net [48] | 92.1 | 96.2 | 83.0 | 90.3 | 82.4 | 88.8 | 90.6 | 78.9 |
| ESDINet [49] | 92.7 | 95.5 | 84.5 | 90.0 | 87.2 | 90.0 | 90.9 | 82.0 |
| UNetFormer [4] | 92.7 | 95.3 | 84.9 | 90.6 | 88.5 | 90.4 | 91.0 | 82.7 |
| Mamba-UNet [13] | 96.4 | 94.5 | 83.7 | 89.4 | 84.3 | 89.7 | 92.6 | 81.6 |
| Swin-UMamba [14] | 96.0 | 94.6 | 81.5 | 91.0 | 83.9 | 89.4 | 92.4 | 81.3 |
| VM-UNet [12] | 96.2 | 93.6 | 83.9 | 89.5 | 78.3 | 88.3 | 92.3 | 79.6 |
| RS3Mamba [15] | 92.8 | 96.8 | 80.8 | 91.1 | 90.9 | 90.5 | – | 82.8 |
| UMFormer [16] | 96.7 | 95.2 | 83.8 | 89.5 | 88.1 | 90.7 | 93.0 | 83.3 |
| DSD-Mamba (Ours) | 97.3 | 96.7 | 82.9 | 91.7 | 90.3 | 91.8 | 94.6 | 85.2 |
| Method | Imp.surf. | Bui. | Lowveg. | Tree | Car | MeanF1 | OA | mIoU |
|---|---|---|---|---|---|---|---|---|
| EaNet [26] | 92.0 | 95.7 | 84.3 | 85.7 | 95.1 | 90.6 | 88.7 | 83.4 |
| MAResU-Net [44] | 91.4 | 95.6 | 85.8 | 86.6 | 93.3 | 90.5 | 89.0 | 83.9 |
| SwiftNet [24] | 91.8 | 95.9 | 85.7 | 86.8 | 94.5 | 91.0 | 89.3 | 83.8 |
| FANet [43] | 92.0 | 96.1 | 86.0 | 87.8 | 94.5 | 91.3 | 89.8 | 84.2 |
| BEDSN [45] | 91.8 | 95.6 | 85.9 | 86.7 | 95.0 | 91.0 | 89.2 | 83.8 |
| DPFE-AFF [46] | 92.3 | 95.4 | 85.8 | 87.5 | 93.7 | 90.5 | 89.7 | 82.8 |
| ResUNet-a [27] | 92.7 | 97.1 | 86.4 | 85.8 | 95.8 | 91.6 | 90.1 | – |
| Swin-UperNet [28] | 91.6 | 96.0 | 86.1 | 87.0 | 91.7 | 90.5 | 89.4 | 82.2 |
| Segmenter [9] | 91.5 | 95.3 | 85.4 | 85.0 | 88.5 | 89.2 | 88.7 | 80.7 |
| CMTFNet [31] | 92.1 | 96.4 | 86.4 | 87.3 | 92.4 | 90.9 | 89.9 | 83.6 |
| ESDINet [49] | 92.7 | 96.3 | 87.3 | 88.1 | 95.4 | 92.0 | 90.5 | 85.3 |
| UNetFormer [4] | 93.0 | 95.6 | 86.7 | 87.9 | 95.0 | 91.6 | 90.5 | 84.8 |
| Mamba-UNet [13] | 92.1 | 94.9 | 82.5 | 85.0 | 93.3 | 90.1 | 88.9 | 82.3 |
| Swin-UMamba [14] | 92.0 | 94.5 | 85.5 | 86.2 | 93.7 | 90.4 | 89.1 | 82.7 |
| VM-UNet [12] | 91.8 | 94.6 | 84.5 | 83.4 | 92.0 | 89.3 | 88.2 | 80.9 |
| UMFormer [16] | 93.7 | 96.4 | 86.7 | 87.8 | 95.5 | 92.0 | 90.9 | 85.5 |
| DSD-Mamba (Ours) | 93.9 | 97.5 | 91.3 | 84.9 | 97.1 | 93.0 | 91.1 | 87.2 |
| DSPFM | SASA | DSCD | mIoU (%) | OA (%) |
|---|---|---|---|---|
| – | – | – | 83.3 | 93.0 |
| ✓ | – | – | 84.0 | 94.2 |
| – | ✓ | – | 84.4 | 94.4 |
| – | – | ✓ | 84.1 | 94.1 |
| ✓ | ✓ | – | 84.5 | 94.3 |
| ✓ | – | ✓ | 84.7 | 94.4 |
| – | ✓ | ✓ | 84.9 | 94.5 |
| ✓ | ✓ | ✓ | 85.2 | 94.6 |
| Scale | Target | Grid | Rel. Affinity | mIoU (%) | OA (%) |
|---|---|---|---|---|---|
| 85.20 | 94.61 | ||||
| 81.87 | 93.81 |
| Method | Params (M) | FLOPs (G) | mIoU (%) on Vaihingen |
|---|---|---|---|
| UNetFormer [4] | 11.84 | 54.70 | 82.7 |
| Swin-UMamba [14] | 39.00 | 96.00 | 81.3 |
| UMFormer [16] | 12.33 | 47.75 | 83.3 |
| DSD-Mamba (Ours) | 26.17 | 117.03 | 85.2 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Feng, X.; Jiang, S.; Wang, L.; Gao, B. DSD-Mamba: Dual-Stream Semantic Segmentation of Remote Sensing Imagery via Dense-Sparse Fusion. Sensors 2026, 26, 3864. https://doi.org/10.3390/s26123864
Feng X, Jiang S, Wang L, Gao B. DSD-Mamba: Dual-Stream Semantic Segmentation of Remote Sensing Imagery via Dense-Sparse Fusion. Sensors. 2026; 26(12):3864. https://doi.org/10.3390/s26123864
Chicago/Turabian StyleFeng, Xinyi, Shaochen Jiang, Liejun Wang, and Beibei Gao. 2026. "DSD-Mamba: Dual-Stream Semantic Segmentation of Remote Sensing Imagery via Dense-Sparse Fusion" Sensors 26, no. 12: 3864. https://doi.org/10.3390/s26123864
APA StyleFeng, X., Jiang, S., Wang, L., & Gao, B. (2026). DSD-Mamba: Dual-Stream Semantic Segmentation of Remote Sensing Imagery via Dense-Sparse Fusion. Sensors, 26(12), 3864. https://doi.org/10.3390/s26123864

