Semi-Supervised Remote Sensing Image Semantic Segmentation Based on Multi-Scale Consistency and Cross-Attention
Highlights
- The proposed MSCA-TSN effectively improves semi-supervised remote sensing image segmentation by integrating uncertainty-aware multi-scale consistency learning and cross-network attention interaction.
- Extensive experiments on the LoveDA and ISPRS Potsdam datasets demonstrate that MSCA-TSN achieves superior segmentation accuracy and boundary integrity under limited supervision, reaching up to 52.41% mIoU on LoveDA and 76.34% mIoU on Potsdam with only 10% labeled data.
- The results indicate that uncertainty-guided multi-scale feature consistency can effectively suppress noisy supervision and enhance the robustness of semi-supervised segmentation in complex remote sensing scenes.
- The proposed cross-teacher–student attention interaction provides a promising direction for improving discriminative feature learning in semi-supervised remote sensing tasks, especially when labeled data are scarce.
Abstract
1. Introduction
- Limited multi-scale feature utilization. Existing methods fail to dynamically evaluate the reliability of multi-scale features and effectively fuse valid information, which cannot mitigate the impact of boundary misalignment caused by intra-class scale variation, leading to reduced geometric completeness of segmentation results.
- Insufficient feature discriminability. Lack of effective cross-network interaction between teacher and student models in semi-supervised frameworks results in confounding representations among heterogeneous land-cover classes, making it difficult to distinguish visually similar categories and limiting segmentation accuracy.
- We develop a semi-supervised segmentation network, MSCA-TSN, by combining the teacher–student architecture with multi-scale consistency and cross-attention mechanisms. During the training stage, multi-scale feature fusion and cross-attention guidance enhance the capability of feature representation, enabling more effective extraction of semantic information from remote sensing objects and addressing the core challenges of semi-supervised RSI segmentation under limited annotated data.
- We propose the AMUC module that enables the model to learn consistency constraints of multi-scale features from unlabeled data, supporting adaptive modeling of land-cover objects at different scales. This module dynamically evaluates feature uncertainty across multiple encoder stages using Monte Carlo Dropout, assigns adaptive weights to hierarchical consistency losses, and effectively fuses multi-scale features, alleviating classification errors caused by scale variation and reducing the loss of topological integrity in segmentation results.
- We propose the CCAM module to strengthen the discriminative capability of feature representations. This module exploits the complementary characteristics of the teacher and student networks. The teacher network provides stable and reliable category features while the student network captures fine-grained image details. By constructing cross-network attention interactions (taking student features as queries and teacher features as keys and values), CCAM guides the student network to reconstruct strongly discriminative feature representations, suppressing category confusion among visually similar land-cover types.
- Extensive experiments on the LoveDA and ISPRS datasets evaluate the effectiveness of MSCA-TSN. The results show that the proposed framework consistently improves segmentation accuracy and boundary quality under limited-label settings, and ablation studies further verify the effectiveness of the proposed AMUC and CCAM modules.
2. Related Works
2.1. Semantic Segmentation of RSIs
2.2. Semi-Supervised RSI Segmentation
3. The Proposed Method
3.1. Problem Definition
3.2. Cross-Teacher–Student Cross-Attention Module

3.3. Adaptive Multi-Scale Uncertainty Consistency Module
3.4. Training Strategy and Hyperparameter Setting
3.5. Loss Function
4. Results
4.1. Experimental Settings
4.1.1. Datasets
4.1.2. Experimental Details
4.2. Comparison with Other Methods
4.2.1. Results on LoveDA Dataset
4.2.2. Results on ISPRS Potsdam Dataset
4.3. Ablation Study
4.4. Parameter Sensitivity Analysis
4.4.1. Sensitivity Analysis of the Uncertainty Threshold
4.4.2. Sensitivity Analysis of the Temperature Coefficient
4.4.3. Sensitivity Analysis of the Number of Monte Carlo Dropout Passes T
4.5. Efficiency Analysis
4.5.1. Comparison with Representative Methods
4.5.2. Internal Overhead Analysis of MSCA-TSN
4.6. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- He, Y.; Liu, L.; You, S.; Mu, X.; Zhang, T.; Liu, A.; Han, X. Remote Sensing Monitoring of Water and Wetland on GFDM-1 Satellite Images. In Proceedings of the 44th IEEE International Geoscience and Remote Sensing Symposium, Athens, Greece, 7–12 July 2024; IEEE: New York, NY, USA, 2024; pp. 4868–4871. [Google Scholar]
- Guo, H.A.; Du, B.; Zhang, L.P.; Su, X. A Coarse-to-fine Boundary Refinement Network for Building Footprint Extraction from Remote Sensing Imagery. ISPRS J. Photogramm. Remote Sens. 2022, 183, 240–252. [Google Scholar] [CrossRef] [Scilit]
- Kussul, N.; Lavreniuk, M.; Skakun, S.; Shelestov, A. Deep Learning Classification of Land Cover and Crop Types Using Remote Sensing Data. IEEE Geosci. Remote Sens. Lett. 2017, 14, 778–782. [Google Scholar] [CrossRef] [Scilit]
- Saif, A.; Dimyati, K.; Noordin, K.A.; Mosali, N.A.; Deepak, G.C.; Alsamhi, S.H. Skyward Bound: Empowering Disaster Resilience with Multi-UAV-Assisted B5G Networks for Enhanced Connectivity and Energy Efficiency. Internet Things 2023, 23, 100885. [Google Scholar] [CrossRef] [Scilit]
- Toth, C.; Jóźków, G. Remote Sensing Platforms and Sensors: A Survey. ISPRS J. Photogramm. Remote Sens. 2016, 115, 22–36. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Xu, F.; Zhang, J.; Zhang, H.; Lyu, X.; Liu, F.; Gao, H.; Kaup, A. Frequency-Guided Denoising Network for Semantic Segmentation of Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2025, 64, 5400217. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Xu, F.; Li, J.; Su, Y.; Li, L.; Lyu, X.; Xu, Z.; Kaup, A. Frequency Domain-Enhanced Spectral-Spatial Fusion Transformer for Semantic Segmentation of Remote Sensing Images. Inf. Fusion 2026, 132, 104248. [Google Scholar] [CrossRef] [Scilit]
- Sezgin, M.; Sankur, B. Survey over Image Thresholding Techniques and Quantitative Performance Evaluation. J. Electron. Imaging 2004, 13, 146–165. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Xu, F.; Liu, F.; Lyu, X.; Gao, H.; Zhou, J.; Kaup, A. A Euclidean Affinity-Augmented Hyperbolic Neural Network for Semantic Segmentation of Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5636718. [Google Scholar] [CrossRef] [Scilit]
- Yu, X.; Ouyang, B.; Principe, J.C.; Farrington, S.; Reed, J.; Li, Y. Weakly supervised learning of point-level annotation for coral image segmentation. In Proceedings of the OCEANS 2019 MTS/IEEE SEATTLE, Seattle, WA, USA, 24–31 October 2019; IEEE: New York, NY, USA, 2019; pp. 1–7. [Google Scholar]
- Long, J.; Shelhamer, E.; Darrell, T. Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; IEEE: New York, NY, USA, 2015; pp. 3431–3440. [Google Scholar] [CrossRef] [Scilit]
- Badrinarayanan, V.; Kendall, A.; Cipolla, R. SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 2481–2495. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 5–9 October 2015; Springer: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar]
- Li, X.; Xu, F.; Liu, F.; Tong, Y.; Lyu, X.; Zhou, J. Semantic Segmentation of Remote Sensing Images by Interactive Representation Refinement and Geometric Prior-Guided Inference. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5400318. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Proceedings of the 15th European Conference on Computer Vision, Munich, Germany, 8–14 September 2018; Springer: Cham, Switzerland, 2018; pp. 833–851. [Google Scholar]
- Xu, Z.; Zhang, W.; Zhang, T.; Li, J. HRCNet: High-Resolution Context Extraction Network for Semantic Segmentation of Remote Sensing Images. Remote Sens. 2021, 13, 71. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Zhang, Q.; Li, Y.; Li, X. Allspark: Reborn labeled features from unlabeled in transformer for semi-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; IEEE: New York, NY, USA, 2024; pp. 3627–3636. [Google Scholar]
- Jin, J.; Zhou, W.; Yang, R.; Ye, L.; Yu, L. Edge Detection Guide Network for Semantic Segmentation of Remote-Sensing Images. IEEE Geosci. Remote Sens. Lett. 2023, 20, 5000505. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; IEEE: New York, NY, USA, 2018; pp. 7132–7141. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. Cbam: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; Springer: Cham, Switzerland, 2018; pp. 3–19. [Google Scholar]
- Fu, J.; Liu, J.; Tian, H.; Li, Y.; Bao, Y.; Fang, Z.; Lu, H. Dual Attention Network for Scene Segmentation. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; IEEE: New York, NY, USA, 2019; pp. 3141–3149. [Google Scholar] [CrossRef] [Scilit]
- Bai, B.; Fu, W.; Lu, T.; Li, S. Edge-Guided Recurrent Convolutional Neural Network for Multitemporal Remote Sensing Image Building Change Detection. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5610613. [Google Scholar] [CrossRef] [Scilit]
- Shang, R.; Liu, M.; Jiao, L.; Feng, J.; Li, Y.; Stolkin, R. Region-Level SAR Image Segmentation Based on Edge Feature and Label Assistance. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5237216. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.Z.; Jiang, J. High-Resolution Boundary-Constrained and Context-Enhanced Network for Remote Sensing Image Segmentation. Remote Sens. 2022, 14, 1859. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Zhang, Y.F.; Wang, X.; Li, J. A Boundary Guided Cross Fusion Approach for Remote Sensing Image Segmentation. IEEE Geosci. Remote Sens. Lett. 2024, 21, 6002305. [Google Scholar] [CrossRef] [Scilit]
- Zheng, X.; Huan, L.; Xia, G.; Gong, J. Parsing Very High Resolution Urban Scene Images by Learning Deep ConvNets with Edge-aware Loss. ISPRS J. Photogramm. Remote Sens. 2020, 170, 15–28. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Che, R.; Wang, X.; Ma, M.; Wu, S.; Feng, T.; Zhang, W. DOCNet: Dual-Domain Optimized Class-Aware Network for Remote Sensing Image Segmentation. IEEE Geosci. Remote Sens. Lett. 2024, 21, 2500905. [Google Scholar] [CrossRef] [Scilit]
- Volpi, M.; Tuia, D. Dense Semantic Labeling of Subdecimeter Resolution Images with Convolutional Neural Networks. IEEE Trans. Geosci. Remote Sens. 2017, 55, 881–893. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Xu, F.; Zhang, J.; Yu, A.; Lyu, X.; Gao, H.; Zhou, J. Dual-domain decoupled fusion network for semantic segmentation of remote sensing images. Inf. Fusion 2025, 124, 103359. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Sun, X.; Chen, C.; Hong, D.; Han, J. Semi-Supervised Semantic Segmentation for Remote Sensing Images via Multiscale Uncertainty Consistency and Cross-Teacher-Student Attention. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5517115. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar]
- Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Deeplab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 834–848. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sun, X.; Xia, M.; Dai, T. Controllable Fused Semantic Segmentation with Adaptive Edge Loss for Remote Sensing Parsing. Remote Sens. 2022, 14, 207. [Google Scholar] [CrossRef] [Scilit]
- Neupane, B.; Aryal, J.; Rajabifard, A. Rethinking the U-Net, ResUNet and U-Net3+ Architectures with Dual Skip Connections for Building Footprint Extraction. arXiv 2023, arXiv:2303.09064v4. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Proceedings of the 31st Annual Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 5998–6008. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
- Zhang, Z.; Huang, X.; Li, J. DWin-HRFormer: A High-Resolution Transformer Model with Directional Windows for Semantic Segmentation of Urban Construction Land. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5400714. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Shi, C.; Xu, N.; Su, Y.; Kaup, A.; Liu, D.; Li, X. Position-Aware Differential Denoising Transformer for Semantic Segmentation of Remote Sensing Images. IEEE Geosci. Remote Sens. Lett. 2026, 23, 5000405. [Google Scholar] [CrossRef] [Scilit]
- Xiang, P.; Ali, S.; Zhang, J.; Jung, S.K. Huixin Zhou aPixel-associated autoencoder for hyperspectral anomaly detection. Int. J. Appl. Earth Obs. Geoinf. 2024, 129, 103816. [Google Scholar] [CrossRef] [Scilit]
- Tarvainen, A.; Valpola, H. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Proceedings of the 31st Annual Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 1195–1204. [Google Scholar]
- French, G.; Laine, S.; Aila, T.; Mackiewicz, M.; Finlayson, G. Semi-supervised semantic segmentation needs strong, varied perturbations. In Proceedings of the 31st British Machine Vision Conference, Virtual, 7–10 September 2020. [Google Scholar]
- Devries, T.; Taylor, G.W. Improved Regularization of Convolutional Neural Networks with Cutout. arXiv 2017, arXiv:1708.04552. [Google Scholar] [CrossRef] [Scilit]
- Yun, S.; Han, D.; Chun, S.; Oh, S.J.; Yoo, Y.; Choe, J. CutMix: Regularization Strategy to Train Strong Classifiers with Localizable Features. In Proceedings of the 17th IEEE International Conference on Computer Vision, Seoul, South Korea, 27 October–2 November 2019; IEEE: New York, NY, USA, 2019; pp. 6022–6031. [Google Scholar]
- Olsson, V.; Tranheden, W.; Pinto, J.; Svensson, L. ClassMix: Segmentation-Based Data Augmentation for Semi-Supervised Learning. In Proceedings of the 24th IEEE Winter Conference on Applications of Computer Vision, Virtual, 3–8 January 2021; IEEE: New York, NY, USA, 2021; pp. 1368–1377. [Google Scholar]
- Lu, X.; Jiao, L.; Li, L.; Liu, F.; Liu, X.; Yang, S.; Feng, Z.; Chen, P. Weak-to-strong consistency learning for semisupervised image segmentation. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5510715. [Google Scholar] [CrossRef] [Scilit]
- Zhang, B.; Zhang, Y.; Li, Y.; Wan, Y.; Guo, H.; Zheng, Z.; Yang, K. Semi-supervised Deep Learning via Transformation Consistency Regularization for Remote Sensing Image Semantic Segmentation. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2023, 16, 5782–5796. [Google Scholar] [CrossRef] [Scilit]
- Ouali, Y.; Hudelot, C.; Tami, M. Semi-Supervised Semantic Segmentation With Cross Consistency Training. In Proceedings of the 33th IEEE Conference on Computer Vision and Pattern Recognition, Virtual, 13–19 June 2020; IEEE: New York, NY, USA, 2020; pp. 12671–12681. [Google Scholar]
- Sohn, K.; Berthelot, D.; Li, C.L.; Zhang, Z.; Carlini, N.; Cubuk, E.D.; Kurakin, A.; Zhang, H.; Raffel, C. FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence. In Proceedings of the 34th Annual Conference on Neural Information Processing Systems, Virtual, 6–12 December 2020; Curran Associates Inc.: Red Hook, NY, USA, 2020; pp. 596–608. [Google Scholar]
- Chen, X.; Yuan, Y.; Zeng, G.; Wang, J. Semi-Supervised Semantic Segmentation with Cross Pseudo Supervision. In Proceedings of the 34th IEEE Conference on Computer Vision and Pattern Recognition, Virtual, 20–25 June 2021; IEEE: New York, NY, USA, 2021; pp. 2613–2622. [Google Scholar]
- Yang, L.H.; Qi, L.; Feng, L.T.; Zhang, W.; Shi, Y.H. Revisiting Weak-to-Strong Consistency in Semi-Supervised Semantic Segmentation. In Proceedings of the 36th IEEE Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; IEEE: New York, NY, USA, 2023; pp. 7236–7246. [Google Scholar]
- Wang, J.; Ding, C.H.Q.; Chen, S.; He, C.; Luo, B. Semi-Supervised Remote Sensing Image Semantic Segmentation via Consistency Regularization and Average Update of Pseudo-Label. Remote Sens. 2020, 12, 3603. [Google Scholar] [CrossRef] [Scilit]
- Yang, L.; Zhuo, W.; Qi, L.; Shi, Y.; Gao, Y. ST++: Make Self-training Work Better for Semi supervised Semantic Segmentation. In Proceedings of the 35th IEEE Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; IEEE: New York, NY, USA, 2022; pp. 4258–4267. [Google Scholar]
- Souly, N.; Spampinato, C.; Shah, M. Semi-supervised semantic segmentation using adversarial networks. In Proceedings of the 16th IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; IEEE: New York, NY, USA, 2017; pp. 5677–5686. [Google Scholar]
- Li, D.; Yang, J.; Kreis, K.; Torralba, A.; Fidler, S. Semantic Segmentation with Generative Models: Semi Supervised Learning and Strong Out-of-Domain Generalization. In Proceedings of the 34th IEEE Conference on Computer Vision and Pattern Recognition, Virtual, 19–25 June 2021; IEEE: New York, NY, USA, 2021; pp. 8296–8307. [Google Scholar]
- Hung, W.C.; Tsai, Y.H.; Liou, Y.T.; Lin, Y.Y.; Yang, M.H. Adversarial Learning for Semi-Supervised Semantic Segmentation. In Proceedings of the 29th British Machine Vision Conference, Newcastle, UK, 3–6 September 2018. [Google Scholar]
- Zhai, D.; Hu, B.; Gong, X.; Zou, H.; Luo, J. ASS-GAN: Asymmetric semi-supervised GAN for breast ultrasound image segmentation. Neurocomputing 2022, 493, 204–216. [Google Scholar] [CrossRef] [Scilit]
- Chen, G.C.; Shi, B.J.; Zhang, Y.H.; He, Z.F.; Zhang, P.C. CGSNet: Cross-consistency guiding semi-supervised semantic segmentation network for remote sensing of plateau lake. J. Netw. Comput. Appl. 2024, 230, 103974. [Google Scholar] [CrossRef] [Scilit]
- Luo, Y.; Sun, B.; Li, S.; Hu, Y. Hierarchical Augmentation and Region-Aware Contrastive Learning for Semi-Supervised Semantic Segmentation of Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2024, 63, 4401311. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Y.L.; Yang, M.Y.; Wang, M.; Qian, X.; Yang, R.; Zhang, X.; Dong, W. Semi-Supervised Adversarial Semantic Segmentation Network Using Transformer and Multiscale Convolution for High-Resolution Remote Sensing Imagery. Remote Sens. 2022, 14, 1786. [Google Scholar] [CrossRef] [Scilit]
- Huang, W.; Shi, Y.; Xiong, Z.; Zhu, X.X. Decouple and Weight Semi-Supervised Semantic Segmentation of Remote Sensing Images. ISPRS J. Photogramm. Remote Sens. 2024, 212, 13–26. [Google Scholar] [CrossRef] [Scilit]
- Xue, X.; Zhu, H.; Li, X.; Wang, J.; Qu, L.; Hou, B. EGPO: Enhanced Guidance and Pseudo-Label Optimization for Semi-Supervised Semantic Segmentation of Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5651913. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Zheng, Z.; Ma, A.; Lu, X.; Zhong, Y. LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation. In Proceedings of the 35th Neural Information Processing Systems Track on Datasets and Benchmarks, Montreal, QC, Canada, 6–14 December 2021. [Google Scholar]
- Rottensteiner, F.; Sohn, G.; Jung, J.; Gerke, M.; Baillard, C.; Benitez, S.; Breitkopf, U. The ISPRS Benchmark on Urban Object Classification and 3D Building Reconstruction. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2012, 1, 293–298. [Google Scholar] [CrossRef] [Scilit]
- Lu, X.; Jiao, L.; Liu, F.; Yang, S.; Liu, X.; Feng, Z.; Li, L.; Chen, P. Simple and Efficient: A Semisupervised Learning Framework for Remote Sensing Image Semantic Segmentation. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5543516. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Wang, H.; Shen, Y.; Fei, J.; Li, W.; Jin, G.; Wu, L.; Zhao, R.; Le, X. Semi-Supervised Semantic Segmentation Using Unreliable Pseudo-Labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; IEEE: New York, NY, USA, 2022; pp. 4248–4257. [Google Scholar]
- Wang, Z.; Zhao, Z.; Xing, X.; Xu, D.; Kong, X.; Zhou, L. Conflict-Based Cross-View Consistency for Semi-Supervised Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; IEEE: New York, NY, USA, 2023; pp. 19585–19595. [Google Scholar]





| Dataset | Label Ratio | Labeled Training Set | Unlabeled Training Set | Validation Set | Test Set |
|---|---|---|---|---|---|
| LoveDA Dataset | 5% | 504 | 9584 | 3338 | 3338 |
| 10% | 1008 | 9080 | 3338 | 3338 |
| Hyperparameter | Value |
|---|---|
| Batch Size | 16 |
| Optimizer | SGD |
| Initial Learning Rate | 0.007 |
| Maximum Iterations | 200 |
| Momentum | 0.9 |
| Weight Decay | 0.0001 |
| Image Size | Pixels |
| Label Ratio | Method | IoU Metrics (%) | Metrics | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Background | Building | Road | Water | Barren Land | Forest | Farmland | mIoU (%) | mF1 (%) | Kappa | BF | |||
| 5% | Mean Teacher [41] | 49.73 | 46.22 | 42.34 | 60.93 | 31.51 | 35.79 | 44.22 | 44.39 | 61.81 | 0.5151 | 59.42 | |
| FixMatch [49] | 45.40 | 53.05 | 51.22 | 66.73 | 28.53 | 27.25 | 54.30 | 44.64 | 62.45 | 0.5378 | 60.18 | ||
| CPS [50] | 48.90 | 49.64 | 47.97 | 60.27 | 4.67 | 36.09 | 47.32 | 42.12 | 56.90 | 0.4976 | 55.03 | ||
| LSST [65] | 51.48 | 45.66 | 52.66 | 67.63 | 33.52 | 35.80 | 48.60 | 47.91 | 64.10 | 0.5434 | 61.52 | ||
| U2PL [66] | 52.58 | 53.12 | 50.97 | 65.75 | 16.48 | 38.16 | 47.89 | 46.42 | 61.93 | 0.5391 | 59.86 | ||
| CCVC [67] | 44.17 | 42.82 | 35.09 | 51.20 | 3.17 | 31.40 | 44.62 | 36.07 | 50.93 | 0.4397 | 48.71 | ||
| UniMatch [51] | 50.20 | 54.49 | 50.46 | 67.18 | 26.79 | 30.06 | 54.86 | 47.72 | 63.46 | 0.5543 | 61.89 | ||
| AllSpark [17] | 65.09 | 55.06 | 47.59 | 67.10 | 34.67 | 26.86 | 51.87 | 49.75 | 64.91 | 0.5682 | 62.47 | ||
| DWL [61] | 48.75 | 55.00 | 51.53 | 69.49 | 29.46 | 36.59 | 52.11 | 48.99 | 64.88 | 0.5597 | 63.02 | ||
| MSCA-TSN | 52.00 | 62.22 | 51.99 | 71.13 | 24.44 | 37.43 | 58.13 | 51.05 | 66.28 | 0.6150 | 65.21 | ||
| 10% | Mean Teacher [41] | 50.45 | 55.75 | 43.56 | 66.15 | 35.24 | 36.96 | 45.64 | 47.68 | 64.18 | 0.5387 | 61.33 | |
| FixMatch [49] | 51.02 | 54.59 | 52.20 | 56.91 | 24.86 | 39.83 | 56.50 | 47.99 | 64.97 | 0.5676 | 62.04 | ||
| CPS [50] | 51.30 | 54.93 | 52.57 | 53.37 | 18.39 | 37.59 | 53.24 | 45.91 | 61.78 | 0.5479 | 57.46 | ||
| LSST [65] | 50.69 | 49.50 | 52.63 | 69.85 | 27.25 | 36.24 | 52.06 | 48.32 | 64.17 | 0.5565 | 62.11 | ||
| U2PL [66] | 51.44 | 53.44 | 53.43 | 56.82 | 29.44 | 39.55 | 51.89 | 48.00 | 64.32 | 0.5537 | 61.68 | ||
| CCVC [67] | 46.79 | 40.25 | 46.96 | 45.79 | 19.57 | 26.58 | 38.38 | 37.76 | 54.01 | 0.4389 | 50.72 | ||
| UniMatch [51] | 51.80 | 53.95 | 51.17 | 58.15 | 25.60 | 38.72 | 54.86 | 47.75 | 63.86 | 0.5639 | 62.27 | ||
| AllSpark [17] | 67.13 | 56.16 | 40.67 | 63.58 | 32.54 | 32.03 | 56.91 | 49.86 | 63.97 | 0.5751 | 63.15 | ||
| DWL [61] | 49.94 | 56.66 | 53.89 | 70.35 | 30.62 | 41.49 | 53.13 | 50.87 | 66.64 | 0.5753 | 64.28 | ||
| MSCA-TSN | 50.77 | 61.58 | 55.17 | 71.66 | 29.94 | 37.93 | 59.82 | 52.41 | 67.73 | 0.6257 | 66.41 | ||
| Label Ratio | Method | IoU Metrics (%) | Metrics | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Building | Low Veg. | Tree | Car | Imp. Surfaces | mIoU (%) | mF1 (%) | Kappa | BF | |||
| 5% | Mean teacher [41] | 82.15 | 65.92 | 67.11 | 72.21 | 74.60 | 72.40 | 83.86 | 0.7403 | 82.64 | |
| CutMix [44] | 52.94 | 68.86 | 41.51 | 58.33 | 54.82 | 55.29 | 70.79 | 0.5783 | 68.95 | ||
| CCT [48] | 72.90 | 80.25 | 64.23 | 58.32 | 74.42 | 70.02 | 82.12 | 0.7236 | 80.46 | ||
| CPS [50] | 76.53 | 84.34 | 57.98 | 69.45 | 75.39 | 72.74 | 83.78 | 0.7492 | 82.31 | ||
| LSST [65] | 69.26 | 84.55 | 67.33 | 67.49 | 73.86 | 72.50 | 83.67 | 0.7399 | 82.18 | ||
| FixMatch [49] | 78.12 | 74.87 | 68.89 | 66.58 | 75.30 | 72.75 | 84.15 | 0.7497 | 82.74 | ||
| UniMatch [51] | 78.24 | 73.59 | 67.17 | 66.64 | 75.07 | 72.14 | 83.73 | 0.7432 | 82.26 | ||
| DWL [61] | 74.81 | 85.64 | 63.68 | 62.99 | 75.68 | 73.10 | 84.22 | 0.7507 | 82.87 | ||
| AllSpark [17] | 85.57 | 67.62 | 60.61 | 73.48 | 77.15 | 72.88 | 84.04 | 0.7989 | 83.22 | ||
| MUCA [30] | 88.45 | 69.53 | 61.39 | 74.18 | 79.56 | 74.62 | 85.15 | 0.8166 | 84.36 | ||
| MSCA-TSN | 89.12 | 70.25 | 62.18 | 75.05 | 80.15 | 75.35 | 85.88 | 0.8258 | 85.17 | ||
| 10% | Mean teacher [41] | 84.76 | 69.28 | 68.83 | 71.66 | 76.51 | 74.21 | 85.07 | 0.7578 | 83.78 | |
| CutMix [44] | 64.55 | 80.99 | 64.79 | 65.50 | 68.01 | 68.77 | 81.34 | 0.7109 | 79.63 | ||
| CCT [48] | 73.09 | 83.94 | 61.12 | 60.45 | 73.06 | 70.33 | 82.27 | 0.7265 | 80.72 | ||
| CPS [50] | 77.80 | 87.15 | 61.12 | 68.48 | 75.89 | 74.09 | 84.55 | 0.7533 | 82.97 | ||
| LSST [65] | 70.92 | 86.06 | 68.91 | 70.22 | 74.89 | 74.20 | 84.95 | 0.7549 | 83.46 | ||
| FixMatch [49] | 77.97 | 76.17 | 70.09 | 70.97 | 76.14 | 74.27 | 85.20 | 0.7606 | 83.71 | ||
| UniMatch [51] | 77.34 | 87.75 | 70.79 | 56.65 | 76.46 | 73.80 | 84.52 | 0.7599 | 83.02 | ||
| DWL [61] | 76.37 | 88.42 | 66.54 | 64.37 | 77.14 | 74.57 | 85.16 | 0.7628 | 83.95 | ||
| AllSpark [17] | 86.29 | 69.83 | 64.17 | 75.23 | 78.31 | 74.76 | 85.35 | 0.8144 | 84.47 | ||
| MUCA [30] | 88.02 | 70.58 | 64.53 | 75.20 | 79.92 | 75.65 | 85.90 | 0.8245 | 85.08 | ||
| MSCA-TSN | 89.05 | 71.12 | 65.08 | 75.88 | 80.56 | 76.34 | 86.52 | 0.8325 | 85.89 | ||
| Method | Modules | IoU Metrics (%) | Metrics | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AMUC | CCAM | Background | Building | Road | Water | Barren Land | Forest | Farmland | mIoU (%) | mF1 (%) | Kappa | BF | |||
| MSCA-TSN | 48.00 | 43.34 | 50.56 | 61.42 | 23.14 | 37.63 | 45.00 | ||||||||
| ✓ | 50.31 | 57.03 | 42.86 | 69.11 | 25.03 | 40.32 | 59.02 | ||||||||
| ✓ | 51.00 | 62.08 | 50.99 | 71.13 | 24.68 | 36.43 | 58.83 | ||||||||
| ✓ | ✓ | 50.77 | 61.58 | 55.17 | 71.66 | 29.94 | 37.93 | 59.82 | |||||||
| LoveDA | ISPRS Potsdam | |||
|---|---|---|---|---|
| 5% Label | 10% Label | 5% Label | 10% Label | |
| 0.4 | 50.21 | 51.63 | 74.48 | 75.36 |
| 0.5 | 50.79 | 52.12 | 75.01 | 75.89 |
| 0.6 | 51.05 | 52.41 | 75.35 | 76.34 |
| 0.7 | 50.87 | 52.26 | 75.16 | 76.12 |
| 0.8 | 50.33 | 51.78 | 74.72 | 75.61 |
| LoveDA | ISPRS Potsdam | |||
|---|---|---|---|---|
| 5% Label | 10% Label | 5% Label | 10% Label | |
| 0.5 | 50.31 | 51.84 | 74.83 | 75.77 |
| 1.0 | 50.76 | 52.19 | 75.17 | 76.05 |
| 2.0 | 51.05 | 52.41 | 75.35 | 76.34 |
| 3.0 | 50.91 | 52.28 | 75.12 | 76.18 |
| 4.0 | 50.58 | 51.96 | 74.89 | 75.94 |
| T | LoveDA | ISPRS Potsdam | ||
|---|---|---|---|---|
| 5% Label | 10% Label | 5% Label | 10% Label | |
| 2 | 49.88 | 51.37 | 74.46 | 75.31 |
| 4 | 50.36 | 51.82 | 74.87 | 75.79 |
| 6 | 50.72 | 52.11 | 75.08 | 76.02 |
| 8 | 50.96 | 52.32 | 75.27 | 76.22 |
| 10 | 51.05 | 52.41 | 75.35 | 76.34 |
| Method | Params (M) | FLOPs (G) | Inference Time (ms) |
|---|---|---|---|
| Mean Teacher [41] | 36.2 | 182.5 | |
| UniMatch [51] | 36.2 | 182.5 | |
| AllSpark [17] | 41.7 | 214.8 | |
| MUCA [30] | 38.8 | 200.6 | |
| MSCA-TSN (ours) | 39.2 | 203.4 |
| Framework | AMUC | CCAM | Params (M) | FLOPs (G) | Inference Time (ms) |
|---|---|---|---|---|---|
| MSCA-TSN | 36.2 | 182.5 | |||
| MSCA-TSN | ✓ | 37.1 | 188.7 | ||
| MSCA-TSN | ✓ | 38.4 | 196.9 | ||
| MSCA-TSN | ✓ | ✓ | 39.2 | 203.4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cao, Y.; Chang, L.; Sun, J.; Li, X.; Liu, J.; Li, X.; Liu, D. Semi-Supervised Remote Sensing Image Semantic Segmentation Based on Multi-Scale Consistency and Cross-Attention. Remote Sens. 2026, 18, 1256. https://doi.org/10.3390/rs18081256
Cao Y, Chang L, Sun J, Li X, Liu J, Li X, Liu D. Semi-Supervised Remote Sensing Image Semantic Segmentation Based on Multi-Scale Consistency and Cross-Attention. Remote Sensing. 2026; 18(8):1256. https://doi.org/10.3390/rs18081256
Chicago/Turabian StyleCao, Yuan, Lin Chang, Jiahao Sun, Xinyu Li, Jing Liu, Xin Li, and Daofang Liu. 2026. "Semi-Supervised Remote Sensing Image Semantic Segmentation Based on Multi-Scale Consistency and Cross-Attention" Remote Sensing 18, no. 8: 1256. https://doi.org/10.3390/rs18081256
APA StyleCao, Y., Chang, L., Sun, J., Li, X., Liu, J., Li, X., & Liu, D. (2026). Semi-Supervised Remote Sensing Image Semantic Segmentation Based on Multi-Scale Consistency and Cross-Attention. Remote Sensing, 18(8), 1256. https://doi.org/10.3390/rs18081256

