Next Article in Journal
Registration of Large Optical and SAR Images with Non-Flat Terrain by Investigating Reliable Sparse Correspondences
Previous Article in Journal
ARE-Net: An Improved Interactive Model for Accurate Building Extraction in High-Resolution Remote Sensing Imagery
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

CNN and Transformer Fusion for Remote Sensing Image Semantic Segmentation

State Key Laboratory of Geohazard Prevention and Geoenvironment Protection, Chengdu University of Technology, Chengdu 610059, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2023, 15(18), 4455; https://doi.org/10.3390/rs15184455
Submission received: 29 July 2023 / Revised: 4 September 2023 / Accepted: 6 September 2023 / Published: 10 September 2023
(This article belongs to the Section Urban Remote Sensing)

Abstract

Semantic segmentation of remote sensing images has been widely used in environmental protection, geological disaster discovery, and natural resource assessment. With the rapid development of deep learning, convolutional neural networks (CNNs) have dominated semantic segmentation, relying on their powerful local information extraction capabilities. Due to the locality of convolution operation, it can be challenging to obtain global context information directly. However, Transformer has excellent potential in global information modeling. This paper proposes a new hybrid convolutional and Transformer semantic segmentation model called CTFuse, which uses a multi-scale convolutional attention module in the convolutional part. CTFuse is a serial structure composed of a CNN and a Transformer. It first uses convolution to extract small-size target information and then uses Transformer to embed large-size ground target information. Subsequently, we propose a spatial and channel attention module in convolution to enhance the representation ability for global information and local features. In addition, we also propose a spatial and channel attention module in Transformer to improve the ability to capture detailed information. Finally, compared to other models used in the experiments, our CTFuse achieves state-of-the-art results on the International Society of Photogrammetry and Remote Sensing (ISPRS) Vaihingen and ISPRS Potsdam datasets.
Keywords: segmentation; remote sensing; CNN; transformer; attention segmentation; remote sensing; CNN; transformer; attention
Graphical Abstract

Share and Cite

MDPI and ACS Style

Chen, X.; Li, D.; Liu, M.; Jia, J. CNN and Transformer Fusion for Remote Sensing Image Semantic Segmentation. Remote Sens. 2023, 15, 4455. https://doi.org/10.3390/rs15184455

AMA Style

Chen X, Li D, Liu M, Jia J. CNN and Transformer Fusion for Remote Sensing Image Semantic Segmentation. Remote Sensing. 2023; 15(18):4455. https://doi.org/10.3390/rs15184455

Chicago/Turabian Style

Chen, Xin, Dongfen Li, Mingzhe Liu, and Jiaru Jia. 2023. "CNN and Transformer Fusion for Remote Sensing Image Semantic Segmentation" Remote Sensing 15, no. 18: 4455. https://doi.org/10.3390/rs15184455

APA Style

Chen, X., Li, D., Liu, M., & Jia, J. (2023). CNN and Transformer Fusion for Remote Sensing Image Semantic Segmentation. Remote Sensing, 15(18), 4455. https://doi.org/10.3390/rs15184455

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop