Next Article in Journal
Monitoring the Coastal Changes of the Po River Delta (Northern Italy) since 1911 Using Archival Cartography, Multi-Temporal Aerial Photogrammetry and LiDAR Data: Implications for Coastline Changes in 2100 A.D.
Next Article in Special Issue
PCAN—Part-Based Context Attention Network for Thermal Power Plant Detection in Remote Sensing Imagery
Previous Article in Journal
Hand Gestures Recognition Using Radar Sensors for Human-Computer-Interaction: A Review
Previous Article in Special Issue
A Novel Unsupervised Classification Method for Sandy Land Using Fully Polarimetric SAR Data
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

C3Net: Cross-Modal Feature Recalibrated, Cross-Scale Semantic Aggregated and Compact Network for Semantic Segmentation of Multi-Modal High-Resolution Aerial Images

1
Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100190, China
2
Key Laboratory of Network Information System Technology (NIST), Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100190, China
3
University of Chinese Academy of Sciences, Beijing 100190, China
4
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100190, China
5
Key Laboratory on Microwave Imaging Technology, Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2021, 13(3), 528; https://doi.org/10.3390/rs13030528
Submission received: 30 December 2020 / Revised: 27 January 2021 / Accepted: 29 January 2021 / Published: 2 February 2021
(This article belongs to the Special Issue Big Remotely Sensed Data)

Abstract

Semantic segmentation of multi-modal remote sensing images is an important branch of remote sensing image interpretation. Multi-modal data has been proven to provide rich complementary information to deal with complex scenes. In recent years, semantic segmentation based on deep learning methods has made remarkable achievements. It is common to simply concatenate multi-modal data or use parallel branches to extract multi-modal features separately. However, most existing works ignore the effects of noise and redundant features from different modalities, which may not lead to satisfactory results. On the one hand, existing networks do not learn the complementary information of different modalities and suppress the mutual interference between different modalities, which may lead to a decrease in segmentation accuracy. On the other hand, the introduction of multi-modal data greatly increases the running time of the pixel-level dense prediction. In this work, we propose an efficient C3Net that strikes a balance between speed and accuracy. More specifically, C3Net contains several backbones for extracting features of different modalities. Then, a plug-and-play module is designed to effectively recalibrate and aggregate multi-modal features. In order to reduce the number of model parameters while remaining the model performance, we redesign the semantic contextual extraction module based on the lightweight convolutional groups. Besides, a multi-level knowledge distillation strategy is proposed to improve the performance of the compact model. Experiments on ISPRS Vaihingen dataset demonstrate the superior performance of C3Net with 15× fewer FLOPs than the state-of-the-art baseline network while providing comparable overall accuracy.
Keywords: semantic segmentation; multi-modal learning; deep neural network design semantic segmentation; multi-modal learning; deep neural network design

Share and Cite

MDPI and ACS Style

Cao, Z.; Diao, W.; Sun, X.; Lyu, X.; Yan, M.; Fu, K. C3Net: Cross-Modal Feature Recalibrated, Cross-Scale Semantic Aggregated and Compact Network for Semantic Segmentation of Multi-Modal High-Resolution Aerial Images. Remote Sens. 2021, 13, 528. https://doi.org/10.3390/rs13030528

AMA Style

Cao Z, Diao W, Sun X, Lyu X, Yan M, Fu K. C3Net: Cross-Modal Feature Recalibrated, Cross-Scale Semantic Aggregated and Compact Network for Semantic Segmentation of Multi-Modal High-Resolution Aerial Images. Remote Sensing. 2021; 13(3):528. https://doi.org/10.3390/rs13030528

Chicago/Turabian Style

Cao, Zhiying, Wenhui Diao, Xian Sun, Xiaode Lyu, Menglong Yan, and Kun Fu. 2021. "C3Net: Cross-Modal Feature Recalibrated, Cross-Scale Semantic Aggregated and Compact Network for Semantic Segmentation of Multi-Modal High-Resolution Aerial Images" Remote Sensing 13, no. 3: 528. https://doi.org/10.3390/rs13030528

APA Style

Cao, Z., Diao, W., Sun, X., Lyu, X., Yan, M., & Fu, K. (2021). C3Net: Cross-Modal Feature Recalibrated, Cross-Scale Semantic Aggregated and Compact Network for Semantic Segmentation of Multi-Modal High-Resolution Aerial Images. Remote Sensing, 13(3), 528. https://doi.org/10.3390/rs13030528

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop