Enhancing High-Resolution Land Cover Classification Using Multi-Level Cross-Modal Attention Fusion
Abstract
1. Introduction
- Novel Cross-Modal Cross-Attention Fusion Mechanism: The proposed CMCA module enables deep bidirectional information exchange across modalities, effectively capturing cross-modal long-distance dependencies critical for land cover classification, thereby significantly enhancing the discriminative power of feature representation.
- Hierarchical Feature Fusion Architecture: CMCAUNet integrates CMCA and SCAG modules through a dual-branch encoder and a carefully designed decoder structure, creating a multi-level feature fusion framework that progresses from shallow to deep layers, significantly enhancing the model’s ability to focus on semantically important regions.
- Comprehensive Experimental Validation and Performance Comparison: Extensive experiments on the public ISPRS Vaihingen and Potsdam datasets demonstrate that CMCAUNet outperforms existing state-of-the-art models in multimodal land cover classification, validating the effectiveness and superiority of the proposed method.
2. Related Works
2.1. Single-Modal Semantic Segmentation
2.2. Multimodal Semantic Segmentation Based on CNN
3. Proposed Method
3.1. Architecture of CMCAUNet
3.2. Cross-Modal Cross-Attention Fusion Module
3.3. Skip-Connection Attention Gate Module
4. Experiment
4.1. Datasets
4.1.1. Vaihingen
4.1.2. Potsdam
4.2. Evaluation Metrics
4.3. Implementation Details
4.4. Performance Comparison
4.4.1. Performance Comparison on the Vaihingen Dataset
4.4.2. Performance Comparison on the Potsdam Dataset
4.5. Ablation Study
5. Discussion
5.1. Computational Complexity and Efficiency
5.2. Generalization and Semantic Understanding
5.3. Limitations and Future Work
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Touati, R.; Mignotte, M.; Dahmane, M. Multimodal change detection in remote sensing images using an unsupervised pixel pairwise-based Markov random field model. IEEE Trans. Image Process. 2019, 29, 757–767. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Luppino, L.T.; Hansen, M.A.; Kampffmeyer, M.; Bianchi, F.M.; Moser, G.; Jenssen, R.; Anfinsen, S.N. Code-aligned autoencoders for unsupervised change detection in multimodal remote sensing images. IEEE Trans. Neural Netw. Learn. Syst. 2022, 35, 60–72. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hong, D.; Gao, L.; Yokoya, N.; Yao, J.; Chanussot, J.; Du, Q.; Zhang, B. More diverse means better: Multimodal deep learning meets remote-sensing imagery classification. IEEE Trans. Geosci. Remote Sens. 2020, 59, 4340–4354. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Zhou, Y.; Zhang, Y.; Zhong, L.; Wang, J.; Chen, J. DKDFN: Domain knowledge-guided deep collaborative fusion network for multimodal unitemporal remote sensing land cover classification. ISPRS J. Photogramm. Remote Sens. 2022, 186, 170–189. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Shen, Z.; Li, Y.; Xia, L.; Wang, H.; Li, S.; Jiao, S.; Lei, Y. Road extraction in mountainous regions from highresolution images based on DSDNet and terrain optimization. Remote Sens. 2020, 13, 90. [Google Scholar] [CrossRef] [Scilit]
- Meng, Y.; Chen, S.; Liu, Y.; Li, L.; Zhang, Z.; Ke, T.; Hu, X. Unsupervised building extraction from multimodal aerial data based on accurate vegetation removal and image feature consistency constraint. Remote Sens. 2022, 14, 1912. [Google Scholar] [CrossRef] [Scilit]
- Palsson, F.; Sveinsson, J.R.; Ulfarsson, M.O.; Benediktsson, J.A. Model-based fusion of multi- and hyperspectral images using PCA and wavelets. IEEE Trans. Geosci. Remote Sens. 2014, 53, 2652–2663. [Google Scholar] [CrossRef] [Scilit]
- Wei, Q.; Bioucas-Dias, J.; Dobigeon, N.; Tourneret, J.-Y. Hyperspectral and multispectral image fusion based on a sparse representation. IEEE Trans. Geosci. Remote Sens. 2015, 53, 3658–3668. [Google Scholar] [CrossRef] [Scilit]
- Shen, Y.; Chen, J.; Xiao, L.; Pan, D. Optimizing multiscale segmentation with local spectral heterogeneity measure for high resolution remote sensing images. ISPRS J. Photogramm. Remote Sens. 2019, 157, 13–25. [Google Scholar] [CrossRef] [Scilit]
- Gislason, P.O.; Benediktsson, J.A.; Sveinsson, J.R. Random forests for land cover classification. Pattern Recognit. Lett. 2006, 27, 294–300. [Google Scholar] [CrossRef] [Scilit]
- Lu, X.; Zhang, J.; Li, T.; Zhang, G. Synergetic classification of long-wave infrared hyperspectral and visible images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2015, 8, 3546–3557. [Google Scholar] [CrossRef] [Scilit]
- Gao, L.; Li, J.; Khodadadzadeh, M.; Plaza, A.; Zhang, B.; He, Z.; Yan, H. Subspace-based support vector machines for hyperspectral image classification. IEEE Geosci. Remote Sens. Lett. 2014, 12, 349–353. [Google Scholar]
- Krähenbühl, P.; Koltun, V. Efficient inference in fully connectedCRFs with Gaussian edge potentials. In Advances in Neural Information Processing Systems 24, Proceedings of the 25th Annual Conference on Neural Information Processing Systems 2011, Granada, Spain, 12–14 December 2011; Curran Associates Inc.: Red Hook, NY, USA, 2011. [Google Scholar]
- Hazirbas, C.; Ma, L.; Domokos, C.; Cremers, D. FuseNet: Incorporating depth into semantic segmentation via fusion-based CNN architecture. In Computer Vision—ACCV 2016; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2016; pp. 213–228. [Google Scholar]
- Audebert, N.; Le Saux, B.; Lefèvre, S. Beyond RGB: Very high resolution urban remote sensing with multimodal deep networks. ISPRS J. Photogramm. Remote Sens. 2018, 140, 20–32. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Li, R.; Duan, C.; Zhang, C.; Meng, X.; Fang, S. A novel transformer based semantic segmentation scheme for fine-resolution remote sensing images. IEEE Geosci. Remote Sens. Lett. 2022, 19, 6506105. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.; Hong, D.; Chanussot, J. Convolutional neural networks for multimodal remote sensing data classification. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5517010. [Google Scholar] [CrossRef] [Scilit]
- Hong, D.; Hu, J.; Yao, J.; Chanussot, J.; Zhu, X.X. Multimodal remote sensing benchmark datasets for land cover classification with a shared and specific feature learning model. ISPRS J. Photogramm. Remote Sens. 2021, 178, 68–80. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Zhang, X.; Pun, M.-O. A crossmodal multiscale fusion network for semantic segmentation of remote sensing data. IEEE J. Sel. Topics Appl. Earth Obs. Remote Sens. 2022, 15, 3463–3474. [Google Scholar] [CrossRef] [Scilit]
- He, X.; Zhou, Y.; Zhao, J.; Zhang, D.; Yao, R.; Xue, Y. Swin transformer embedding UNet for remote sensing image semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2022, 60, 4408715. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Zhang, B.; Yu, W.; Kang, X. Federated deep learning with prototype matching for object extraction from very-high-resolution remote sensing images. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5603316. [Google Scholar] [CrossRef] [Scilit]
- Gao, L.; Liu, H.; Yang, M.; Chen, L.; Wan, Y.; Xiao, Z.; Qian, Y. STransFuse: Fusing swin transformer and convolutional neural network for remote sensing image semantic segmentation. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2021, 14, 10990–11003. [Google Scholar] [CrossRef] [Scilit]
- Ma, X.; Zhang, X.; Wang, Z.; Pun, M.-O. Unsupervised domain adaptation augmented by mutually boosted attention for semantic segmentation of VHR remote sensing images. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5400515. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Yu, W.; Pun, M.-O.; Shi, W. Cross-domain landslide mapping from large-scale remote sensing images using prototypeguided domain-aware progressive representation learning. ISPRS J. Photogramm. Remote Sens. 2023, 197, 1–17. [Google Scholar] [CrossRef] [Scilit]
- Zhu, J.; Guo, Y.; Sun, G.; Yang, L.; Deng, M.; Chen, J. Unsupervised domain adaptation semantic segmentation of high-resolution remote sensing imagery with invariant domain-level prototype memory. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5603518. [Google Scholar] [CrossRef] [Scilit]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 3431–3440. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer International Publishing: Cham, Switzerland, 2015. [Google Scholar]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. TransUNet: Transformers make strong encoders for medical image segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef] [Scilit]
- Cao, H.; Wang, Y.; Chen, J.; Jiang, D.; Zhang, X.; Tian, Q.; Wang, M. Swin-Unet: Unet-like pure transformer for medical image segmentation. In European Conference on Computer Vision; Springer International Publishing: Cham, Switzerland, 2022; pp. 205–218. [Google Scholar]
- Seichter, D.; Kohler, M.; Lewandowski, B.; Wengefeld, T.; Gross, H.-M. Efficient RGB-D semantic segmentation for indoor scene analysis. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 30 May–5 June 2021; pp. 13525–13531. [Google Scholar]
- Bai, X.; Hu, Z.; Zhu, X.; Huang, Q.; Chen, Y.; Fu, H.; Tai, C.-L. Transfusion: Robust LiDAR-camera fusion for 3D object detection with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 1090–1099. [Google Scholar]
- Li, J.; Hong, D.; Gao, L.; Yao, J.; Zheng, K.; Zhang, B.; Chanussot, J. Deep learning in multimodal remote sensing data fusion: A comprehensive review. Int. J. Appl. Earth Obs. Geoinf. 2022, 112, 102926. [Google Scholar] [CrossRef] [Scilit]
- Roy, S.K.; Deria, A.; Hong, D.; Rasti, B.; Plaza, A.; Chanussot, J. Multimodal fusion transformer for remote sensing image classification. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5515620. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, Virtual Event, 18–24 July 2021; pp. 3314–3329. [Google Scholar]
- Hong, D.; Gao, L.; Hang, R.; Zhang, B.; Chanussot, J. Deep encoder–decoder networks for classification of hyperspectral and LiDAR data. IEEE Geosci. Remote Sens. Lett. 2022, 19, 5500205. [Google Scholar] [CrossRef] [Scilit]
- Feng, D.; Haase-Schutz, C.; Rosenbaum, L.; Hertlein, H.; Glaser, C.; Timm, F.; Wiesbeck, W.; Dietmayer, K. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Trans. Intell. Transp. Syst. 2020, 22, 1341–1360. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Gao, K.; Wang, H.; Yang, Z.; Wang, P.; Ji, S.; Huang, Y.; Zhu, Z.; Zhao, X. A Transformer-based multi-modal fusion network for semantic segmentation of high-resolution remote sensing imagery. Int. J. Appl. Earth Obs. Geoinf. 2024, 133, 104083. [Google Scholar] [CrossRef] [Scilit]
- Guo, H.; Tian, B.; Liu, W. CCFormer: Cross-Modal Cross-Attention Transformer for Classification of Hyperspectral and LiDAR Data. Sensors 2025, 25, 14248220. [Google Scholar] [CrossRef] [Scilit]
- Marmanis, D.; Schindler, K.; Wegner, J.D.; Galliani, S.; Datcu, M.; Stilla, U. Classification with an edge: Improving semantic image segmentation with boundary detection. ISPRS J. Photogramm. Remote Sens. 2018, 135, 158–172. [Google Scholar] [CrossRef] [Scilit]
- Nogueira, K.; Mura, M.D.; Chanussot, J.; Schwartz, W.R.; Dos Santos, J.A. Dynamic multicontext segmentation of remote sensing images based on convolutional networks. IEEE Trans. Geosci. Remote Sens. 2019, 57, 7503–7520. [Google Scholar] [CrossRef] [Scilit]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 2881–2890. [Google Scholar]
- Hosseinpour, H.; Samadzadegan, F.; Javan, F.D. CMGFNet: A deep cross-modal gated fusion network for building extraction from very highresolution remote sensing images. ISPRS J. Photogramm. Remote Sens. 2022, 184, 96–115. [Google Scholar] [CrossRef] [Scilit]
- Diakogiannis, F.I.; Waldner, F.; Caccetta, P.; Wu, C. ResUNet—A: A deep learning framework for semantic segmentation of remotely sensed data. ISPRS J. Photogramm. Remote Sens. 2020, 162, 94–114. [Google Scholar] [CrossRef] [Scilit]
- Ngiam, J.; Khosla, A.; Kim, M.; Nam, J.; Lee, H.; Ng, A.Y. Multimodal deep learning. In Proceedings of the 28th International Conference on International Conference on Machine Learning (ICML 2011), Bellevue, WA, USA, 28 June–2 July 2011; pp. 689–696. [Google Scholar]
- Baltrusaitis, T.; Ahuja, C.; Morency, L.-P. Multimodal machine learning: A survey and taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 41, 423–443. [Google Scholar] [CrossRef] [Scilit]
- He, Q.; Sun, X.; Diao, W.; Yan, Z.; Yao, F.; Fu, K. Multimodal remote sensing image segmentation with intuition-inspired hypergraph modeling. IEEE Trans. Image Process. 2023, 32, 1474–1487. [Google Scholar] [CrossRef] [Scilit]
- Zhou, W.; Jin, J.; Lei, J.; Yu, L. CIMFNet: Cross-layer interaction and multiscale fusion network for semantic segmentation of high-resolution remote sensing images. IEEE J. Sel. Top. Signal Process. 2022, 16, 666–676. [Google Scholar] [CrossRef] [Scilit]
- Ma, J.; Zhou, W.; Lei, J.; Yu, L. Adjacent bi-hierarchical network for scene parsing of remote sensing images. IEEE Geosci. Remote Sens. Lett. 2023, 20, 3000705. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 1–9. [Google Scholar]
- Li, R.; Zheng, S.; Zhang, C.; Duan, C.; Wang, L.; Atkinson, P.M. ABCNet: Attentive bilateral contextual network for efficient semantic segmentation of fine-resolution remotely sensed imagery. ISPRS J. Photogramm. Remote Sens. 2021, 181, 84–98. [Google Scholar] [CrossRef] [Scilit]
- Li, R.; Zheng, S.; Duan, C.; Su, J.; Zhang, C. Multistage attention ResU-Net for semantic segmentation of fine-resolution remote sensing images. IEEE Geosci. Remote Sens. Lett. 2022, 19, 8009205. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Lin, K.-Y.; Wang, J.; Wu, W.; Qian, C.; Li, H.; Zeng, G. Bi-directional cross-modality feature propagation with separation-and-aggregation gate for RGB-D semantic segmentation. In European Conference on Computer Vision; Springer International Publishing: Cham, Switzerland, 2020; pp. 561–577. [Google Scholar]








| Fusion Strategy | Description |
|---|---|
| Early Fusion [17] | Directly merges raw data at the input layer; ensures spatial alignment accuracy, but is prone to introducing redundancy or task-irrelevant noise. |
| Late Fusion [35] | Integrates predictions from independently trained unimodal networks; ignores potential semantic associations between modalities at the feature level. |
| Intermediate Fusion [36] | Provides a compromise through feature-level interactions (e.g., summation or concatenation); however, single-layer operations may fail to capture hierarchical dependencies. |
| Attention-based Fusion [37] | Employs cross-modal attention mechanisms for coarse-grained tasks; often lacks the spatial accuracy and multi-scale refinement needed for fine-grained mapping. |
| Method | OA (%) | mF1 (%) | mIoU (%) | |||||
|---|---|---|---|---|---|---|---|---|
| Bui. | Tre. | Low. | Car | Imp. | Total | |||
| ABCNet [50] | 94.10 | 90.81 | 78.53 | 64.12 | 89.70 | 89.25 | 85.34 | 75.20 |
| PSPNet [41] | 94.52 | 90.17 | 78.84 | 79.22 | 92.03 | 89.94 | 86.55 | 76.96 |
| MAResU-Net [51] | 94.84 | 89.99 | 79.09 | 85.89 | 92.19 | 90.17 | 88.54 | 79.89 |
| vFuseNet [15] | 95.92 | 91.36 | 77.64 | 76.06 | 91.85 | 90.49 | 87.89 | 78.92 |
| FuseNet [14] | 96.28 | 90.28 | 78.98 | 81.37 | 91.66 | 90.51 | 87.71 | 78.71 |
| ESANet [30] | 95.69 | 90.50 | 77.16 | 85.46 | 91.39 | 90.61 | 88.18 | 79.42 |
| SA-GATE [52] | 94.84 | 92.56 | 81.29 | 87.79 | 91.69 | 91.10 | 89.81 | 81.27 |
| CMGFNet [42] | 97.75 | 91.60 | 80.03 | 87.28 | 92.35 | 91.72 | 90.00 | 82.26 |
| CMCAUNet (our) | 97.76 | 91.48 | 77.41 | 90.85 | 90.31 | 90.74 | 89.47 | 81.49 |
| Method | OA (%) | mF1 (%) | mIoU (%) | |||||
|---|---|---|---|---|---|---|---|---|
| Bui. | Tre. | Low. | Car | Imp. | Total | |||
| ABCNet [50] | 96.23 | 78.92 | 86.40 | 92.92 | 88.90 | 87.52 | 88.14 | 79.26 |
| PSPNet [41] | 97.03 | 83.13 | 85.67 | 88.81 | 90.91 | 88.67 | 88.92 | 80.36 |
| MAResU-Net [51] | 96.82 | 83.97 | 87.70 | 95.88 | 92.19 | 89.82 | 90.86 | 83.61 |
| vFuseNet [15] | 97.23 | 84.29 | 89.03 | 95.49 | 91.62 | 90.22 | 91.26 | 84.26 |
| FuseNet [14] | 97.48 | 85.14 | 87.31 | 96.10 | 92.64 | 90.58 | 91.60 | 84.86 |
| ESANet [30] | 97.10 | 85.31 | 87.81 | 94.08 | 92.76 | 89.74 | 91.22 | 84.15 |
| SA-GATE [52] | 96.54 | 81.18 | 85.35 | 96.63 | 90.77 | 87.91 | 90.26 | 82.53 |
| CMGFNet [42] | 97.41 | 86.80 | 86.68 | 95.68 | 92.60 | 90.21 | 91.40 | 84.53 |
| CMCAUNet (our) | 97.51 | 84.83 | 87.98 | 96.98 | 92.95 | 90.28 | 91.52 | 84.76 |
| Structure | OA (%) | mF1 (%) | mIoU (%) | |
|---|---|---|---|---|
| CMCA | SCAG | |||
| √ | 87.07 | 84.34 | 73.70 | |
| √ | 87.34 | 84.19 | 73.55 | |
| √ | √ | 90.74 | 89.47 | 81.49 |
| Method | Multimodal | FLOPs (G) | Parameter (M) | Memory (MB) | Speed (FPS) | MIoU (%) |
|---|---|---|---|---|---|---|
| ABCNet [50] | N | 3.9 | 13.39 | 1598 | 15.87 | 75.20 |
| PSPNet [41] | N | 49.03 | 46.72 | 3124 | 66.01 | 76.96 |
| MAResU-Net [51] | N | 8.79 | 26.27 | 1908 | 10.62 | 79.89 |
| vFuseNet [15] | Y | 60.36 | 44.17 | 2618 | 16.93 | 78.92 |
| FuseNet [14] | Y | 58.37 | 42.08 | 2284 | 18.92 | 78.71 |
| ESANet [30] | Y | 7.73 | 34.03 | 1914 | 10.42 | 79.42 |
| SA-GATE [52] | Y | 41.28 | 110.85 | 3174 | 10.00 | 81.27 |
| CMGFNet [42] | Y | 19.51 | 64.20 | 2463 | 11.61 | 82.26 |
| CMCAUNet (our) | Y | 73.92 | 71.53 | 1299 | 106.87 | 81.49 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jiang, Y.; Liu, T.; Zhou, J.; Guo, Y.; Hu, T. Enhancing High-Resolution Land Cover Classification Using Multi-Level Cross-Modal Attention Fusion. Land 2026, 15, 181. https://doi.org/10.3390/land15010181
Jiang Y, Liu T, Zhou J, Guo Y, Hu T. Enhancing High-Resolution Land Cover Classification Using Multi-Level Cross-Modal Attention Fusion. Land. 2026; 15(1):181. https://doi.org/10.3390/land15010181
Chicago/Turabian StyleJiang, Yangwei, Ting Liu, Junhao Zhou, Yihan Guo, and Tangao Hu. 2026. "Enhancing High-Resolution Land Cover Classification Using Multi-Level Cross-Modal Attention Fusion" Land 15, no. 1: 181. https://doi.org/10.3390/land15010181
APA StyleJiang, Y., Liu, T., Zhou, J., Guo, Y., & Hu, T. (2026). Enhancing High-Resolution Land Cover Classification Using Multi-Level Cross-Modal Attention Fusion. Land, 15(1), 181. https://doi.org/10.3390/land15010181

