CoMS-UNet: A Cohesive Multi-Scale Network for Semantic Reorganization and Scale-Aware Context Modeling in Remote Sensing Image Segmentation
Abstract
1. Introduction
- We propose the Multi-Scale Semantic Reorganization Module (MSRM), which strengthens cross-scale semantic connections and promotes semantic reorganization of multi-stage encoder features in skip connections. By adopting a fusion–separation–refusion strategy, MSRM aligns, fuses, and reassembles hierarchical encoder features, thereby alleviating the interference caused by noisy shallow spatial details and the spatial detail degradation of deep contextual representations in high-resolution remote sensing scenes.
- We design a Scale-Aware Module (SAM) to improve the model’s sensitivity to target scale variations in remote sensing imagery. SAM constructs receptive-field branches with predefined dilation rates and adaptively weights their contextual responses according to the input feature. By learning input-dependent branch weights at the bottleneck layer, SAM selectively emphasizes scale-relevant contextual cues, thereby providing semantic guidance for decoder reconstruction.
- We propose the Channel-Spatial Shuffle Mamba Block (CSSMBlock) to strengthen channel-aware decoder reconstruction, thereby facilitating the effective recovery of fine structural details in remote sensing imagery. CSSMBlock integrates channel shuffle-based local extraction, SS2D-based long-range modeling, and dual attention to alleviate the channel-wise information isolation caused by depthwise separable convolution. By promoting cross-channel joint modeling during decoder reconstruction, CSSMBlock improves the recovery of object boundaries, narrow structures, and fragmented land-cover regions.
2. Materials and Methods
2.1. Architecture Overview
2.2. Multi-Scale Semantic Reorganization Module (MSRM)
2.3. Scale-Aware Module (SAM)
2.4. Channel-Spatial Shuffle Mamba Block (CSSMBlock)
3. Results
3.1. Experimental Setup
3.1.1. Dataset Settings
3.1.2. Implementation Details
3.1.3. Evaluation Metrics and Comparison Protocol
3.2. Results on ISPRS Vaihingen Dataset
3.3. Results on LoveDA Dataset
3.4. Ablation Study
3.5. Complexity and Efficiency Analysis
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| CoMS-UNet | Cohesive Multi-Scale U-Net |
| MSRM | Multi-Scale Semantic Reorganization Module |
| SAM | Scale-Aware Module |
| CSSMBlock | Channel-Spatial Shuffle Mamba Block |
| ECSBlock | Enhanced Channel-Shuffle Block |
| LCSBlock | Local Channel-Shuffle Block |
| DSC | Depthwise Separable Convolution |
| SS2D | Two-Dimensional Selective Scan |
| MSAA | Multi-Scale Attention Aggregation |
| GAP | Global Average Pooling |
| BN | Batch Normalization |
| LN | Layer Normalization |
| GConv | Grouped Convolution |
| DWConv | Depthwise Convolution |
| C-Shuffle | Channel Shuffle |
| C-Attn | Channel Attention |
| S-Attn | Spatial Attention |
| CS-Attn | Channel-Spatial Attention |
| FLOPs | Floating-Point Operations |
| RGB | Red–Green–Blue |
References
- Zhang, C.; Kerner, H.; Wang, S.; Hao, P.; Li, Z.; Hunt, K.A.; Abernethy, J.; Zhao, H.; Gao, F.; Di, L.; et al. Remote sensing for crop mapping: A perspective on current and future crop-specific land cover data products. Remote. Sens. Environ. 2025, 330, 114995. [Google Scholar] [CrossRef] [Scilit]
- Zheng, J.; Ye, Z.; Wen, Y.; Huang, J.; Zhang, Z.; Li, Q.; Hu, Q.; Xu, B.; Zhao, L.; Fu, H. A comprehensive review of agricultural parcel and boundary delineation from remote sensing images: Recent progress and future perspectives. IEEE Geosci. Remote. Sens. Mag. 2026, 14, 206–237. [Google Scholar] [CrossRef] [Scilit]
- Nuradili, P.; Zhou, J.; Zhou, G.; Melgani, F. Deep learning method for wetland segmentation in unmanned aerial vehicle multispectral imagery. Remote. Sens. 2024, 16, 4777. [Google Scholar] [CrossRef] [Scilit]
- Hertel, V.; Geiß, C.; Wieland, M.; Taubenböck, H. Rapid domain adaptation for disaster impact assessment: Remote sensing of building damage after the 2021 Germany floods. Sci. Remote. Sens. 2025, 12, 100287. [Google Scholar] [CrossRef] [Scilit]
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef] [Scilit]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet classification with deep convolutional neural networks. In Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA, 3–6 December 2012; Volume 25, pp. 1097–1105. [Google Scholar]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 3431–3440. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 5–9 October 2015; Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar] [CrossRef] [Scilit]
- Oktay, O.; Schlemper, J.; Folgoc, L.L.; Lee, M.; Heinrich, M.; Misawa, K.; Mori, K.; McDonagh, S.; Hammerla, N.Y.; Kainz, B.; et al. Attention U-Net: Learning where to look for the pancreas. arXiv 2018, arXiv:1804.03999. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 833–851. [Google Scholar] [CrossRef] [Scilit]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2881–2890. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Zhou, X.; Lin, M.; Sun, J. ShuffleNet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 6848–6856. [Google Scholar] [CrossRef] [Scilit]
- Liu, M.; Dan, J.; Lu, Z.; Yu, Y.; Li, Y.; Li, X. CM-UNet: Hybrid CNN-Mamba UNet for remote sensing image semantic segmentation. arXiv 2024, arXiv:2405.10530. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Chen, T.; Zheng, L.; Tie, J.; Zhang, Y.; Chen, P.; Luo, Z.; Song, Q. A multi-scale remote sensing semantic segmentation model with boundary enhancement based on UNetFormer. Sci. Rep. 2025, 15, 14737. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ma, Q.; Liu, H.; Jin, Y.; Liu, X. Multi-scale context enhancement network with local–global synergy modeling strategy for semantic segmentation on remote sensing images. Electronics 2025, 14, 2526. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Liu, C.; Wu, Z.; Zhang, L.; Yang, L. Remote sensing image segmentation using Vision Mamba and multi-scale multi-frequency feature fusion. Remote. Sens. 2025, 17, 1390. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Li, D.; Dong, S.; Meng, X.; Zhang, X.; Hong, D. PyramidMamba: Rethinking pyramid feature fusion with selective space state model for semantic segmentation of remote sensing imagery. Int. J. Appl. Earth Obs. Geoinf. 2025, 144, 104884. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Q.; Li, H.; He, L.; Fan, L. SwinMamba: A hybrid local–global mamba framework for enhancing semantic segmentation of remotely sensed images. Digit. Signal Process. 2026, 175, 106029. [Google Scholar] [CrossRef] [Scilit]
- Li, R.; Zheng, S.; Zhang, C.; Duan, C.; Wang, L.; Atkinson, P.M. ABCNet: Attentive bilateral contextual network for efficient semantic segmentation of fine-resolution remotely sensed imagery. ISPRS J. Photogramm. Remote. Sens. 2021, 181, 84–98. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Li, R.; Wang, D.; Duan, C.; Wang, T.; Meng, X. Transformer meets convolution: A bilateral awareness network for semantic segmentation of very fine resolution urban scene images. Remote. Sens. 2021, 13, 3065. [Google Scholar] [CrossRef] [Scilit]
- Strudel, R.; Garcia, R.; Laptev, I.; Schmid, C. Segmenter: Transformer for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 7262–7272. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Li, R.; Zhang, C.; Fang, S.; Duan, C.; Meng, X.; Atkinson, P.M. UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery. ISPRS J. Photogramm. Remote. Sens. 2022, 190, 196–214. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Huang, P.; Zhang, M.; Tang, W.; Yu, X. CMTFNet: CNN and multiscale transformer fusion network for remote-sensing image semantic segmentation. IEEE Trans. Geosci. Remote. Sens. 2023, 61, 2004612. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Zhang, X.; Pun, M.O. RS3Mamba: Visual state space model for remote sensing image semantic segmentation. IEEE Geosci. Remote. Sens. Lett. 2024, 21, 6011405. [Google Scholar] [CrossRef] [Scilit]





| Method | Backbone | Imp. | Build. | Low.v | Tree | Car | mF1 | OA | mIoU |
|---|---|---|---|---|---|---|---|---|---|
| ABCNet [19] | R18 | 96.25 | 94.41 | 83.96 | 89.94 | 81.12 | 89.14 | 92.56 | 80.90 |
| BANet [20] | ResT | 96.61 | 94.88 | 84.20 | 89.34 | 83.35 | 89.68 | 92.75 | 81.72 |
| Segmenter [21] | ViT-T | 95.33 | 93.61 | 76.31 | 88.05 | 58.87 | 82.43 | 91.45 | 72.22 |
| UNetFormer [22] | R18 | 96.60 | 95.45 | 83.44 | 89.39 | 86.83 | 90.34 | 92.88 | 82.77 |
| FTUNetFormer [22] | Swin-B | 97.01 | 95.76 | 84.57 | 89.99 | 88.62 | 91.19 | 93.39 | 84.14 |
| CMTFNet [23] | R50 | 96.13 | 94.43 | 83.66 | 89.20 | 85.62 | 89.81 | 92.33 | 81.86 |
| RS3Mamba [24] | R18-Mamba | 96.77 | 96.10 | 81.97 | 91.11 | 89.72 | 91.13 | 94.05 | 84.14 |
| CM-UNet [13] | R18 | 97.14 | 96.51 | 82.67 | 91.73 | 91.00 | 91.81 | 94.49 | 85.28 |
| CoMS-UNet (Ours) | R18 | 97.46 | 96.75 | 84.58 | 92.42 | 91.68 | 92.58 | 95.01 | 86.51 |
| Method | Backbone | Bkg. | Build. | Road | Water | Barr. | Forest | Agri. | mIoU |
|---|---|---|---|---|---|---|---|---|---|
| UNetFormer [22] | R18 | 54.12 | 59.58 | 51.76 | 63.50 | 33.35 | 43.97 | 50.20 | 50.93 |
| RS3Mamba [24] | R18-Mamba | 50.67 | 59.29 | 51.97 | 52.71 | 28.26 | 42.27 | 50.69 | 47.98 |
| Segmenter [21] | ViT-T | 51.83 | 55.58 | 49.95 | 71.06 | 24.12 | 39.62 | 59.73 | 50.27 |
| ABCNet [19] | R18 | 51.97 | 59.26 | 52.54 | 62.97 | 30.39 | 37.22 | 45.82 | 48.59 |
| CM-UNet [13] | R18 | 53.31 | 62.86 | 53.42 | 64.28 | 34.40 | 41.48 | 50.51 | 51.47 |
| CoMS-UNet (Ours) | R18 | 54.21 | 63.79 | 53.16 | 65.85 | 34.77 | 44.20 | 51.35 | 52.48 |
| MSRM | SAM | CSSMBlock | mF1 | OA | mIoU |
|---|---|---|---|---|---|
| × | × | × | |||
| ✓ | × | × | |||
| × | ✓ | × | |||
| × | × | ✓ | |||
| ✓ | ✓ | × | |||
| ✓ | × | ✓ | |||
| × | ✓ | ✓ | |||
| ✓ | ✓ | ✓ |
| Model | FLOPs@512 ↓ | FLOPs@1024 ↓ | Param. ↓ | mIoU (%) ↑ |
|---|---|---|---|---|
| ABCNet [19] | 15.63 | 62.50 | 13.39 | 80.90 |
| UNetFormer [22] | 11.74 | 46.97 | 11.69 | 82.77 |
| FTUNetFormer [22] | 126.30 | 499.19 | 96.14 | 84.14 |
| CMTFNet [23] | 33.07 | 131.87 | 30.07 | 81.86 |
| CM-UNet [13] | 12.02 | 48.08 | 12.89 | 85.28 |
| CoMS-UNet (Ours) | 38.48 | 153.91 | 30.86 | 86.51 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Y.; Jiang, S.; Wang, L.; Gao, B. CoMS-UNet: A Cohesive Multi-Scale Network for Semantic Reorganization and Scale-Aware Context Modeling in Remote Sensing Image Segmentation. Sensors 2026, 26, 4987. https://doi.org/10.3390/s26154987
Wang Y, Jiang S, Wang L, Gao B. CoMS-UNet: A Cohesive Multi-Scale Network for Semantic Reorganization and Scale-Aware Context Modeling in Remote Sensing Image Segmentation. Sensors. 2026; 26(15):4987. https://doi.org/10.3390/s26154987
Chicago/Turabian StyleWang, Yankai, Shaochen Jiang, Liejun Wang, and Beibei Gao. 2026. "CoMS-UNet: A Cohesive Multi-Scale Network for Semantic Reorganization and Scale-Aware Context Modeling in Remote Sensing Image Segmentation" Sensors 26, no. 15: 4987. https://doi.org/10.3390/s26154987
APA StyleWang, Y., Jiang, S., Wang, L., & Gao, B. (2026). CoMS-UNet: A Cohesive Multi-Scale Network for Semantic Reorganization and Scale-Aware Context Modeling in Remote Sensing Image Segmentation. Sensors, 26(15), 4987. https://doi.org/10.3390/s26154987

