Next Article in Journal
MPES-YOLO: A Multi-Scale Lightweight Framework with Selective Edge Enhancement for Loess Landslide Detection
Next Article in Special Issue
Nonlinear Responses and Spatial Heterogeneity of Net Ecosystem Productivity to Extreme Weather Events in Central Asia
Previous Article in Journal
Mainlobe Coherent Source 3D Imaging via Monopulse Ratio-Based Spatial Steering Vector and Polarization Diversity
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Vegetation Mapping in Heterogeneous Forest–Shrub–Grass Ecosystems Using Fused High-Resolution Optical and SAR Data

1
School of Surveying and Land Information Engineering, Henan Polytechnic University, Jiaozuo 454000, China
2
Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100094, China
3
School of Remote Sensing and Information Engineering, North China Institute of Aerospace Engineering, Langfang 065000, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(9), 1373; https://doi.org/10.3390/rs18091373
Submission received: 10 March 2026 / Revised: 20 April 2026 / Accepted: 24 April 2026 / Published: 29 April 2026

Highlights

What are the main findings?
  • A high-resolution multimodal dataset (GF23FSG) was constructed based on GF-2 optical imagery and GF-3 SAR imagery. The dataset incorporates diverse remote sensing features, including spectral, textural, and SAR scattering characteristics, providing rich multimodal information for improving the fine classification of forest, shrubland, and grassland.
  • A dual-branch network named CASFNet was proposed, incorporating a cross-modal adaptive structure fusion module (CASF-Module) and a multi-auxiliary supervision strategy (MASLoss) to enhance the collaborative learning of optical and SAR features.
What are the implications of the main findings?
  • The GF23FSG dataset provides a valuable data foundation with rich multi-source information for the fine classification of forest, shrubland, and grassland.
  • The proposed CASFNet framework effectively addresses the challenge of cross-modal feature fusion in high-resolution multimodal remote sensing imagery, significantly improving the accuracy of fine classification of forest, shrubland, and grassland and supporting applications such as carbon stock estimation, ecological monitoring, and ecosystem assessment.

Abstract

Forest, shrubland, and grassland exhibit highly overlapping characteristics, and single-modal remote sensing data cannot simultaneously capture both spectral and structural information. Moreover, multimodal fusion learning of optical and SAR data faces challenges such as the lack of high-quality samples and difficulties in effective cross-modal feature fusion. Therefore, a high-resolution multimodal remote sensing feature dataset (GF23FSG) is constructed for the fine classification of forest, shrubland, and grassland, and a Cross-modal Adaptive Structure Fusion Network (CASFNet) is proposed. In response to the feature heterogeneity of optical and SAR, a cross-modal adaptive fusion module based on spatial alignment and a dynamic weight allocation strategy is proposed, which effectively enhances the learning of spectral–spectrum heterogeneous features. In addition, a multi-level auxiliary supervision mechanism is introduced to strengthen feature representation learning. Gradient constraints are further imposed on deep-level features to improve the model’s ability to capture and learn deep cross-modal representations, thereby effectively mitigating representation degradation during the feature fusion process. Experiments on the self-constructed GF23FSG dataset and the publicly available SEN12MS dataset achieve OA of 77.38% and 71.84%, respectively, demonstrating superior classification performance compared with SOTA methods. In addition, comparative analysis with public land cover products and field samples further confirm the reliability and generalization performance of the proposed dataset and model for the fine classification of forest, shrubland, and grassland. This study provides a new solution for the fine classification of forest, shrubland, and grassland from multimodal remote sensing images from the perspectives of dataset construction and methodological design.

1. Introduction

Vegetation refers to the collective plant communities distributed across the Earth’s surface. Forest, shrubland, and grassland exhibit significant differences in structural, spectral, textural, and temporal characteristics and constitute three major components of terrestrial vegetation ecosystems. They contribute significantly to ecosystem stability and biodiversity while also supporting carbon sequestration, hydrological regulation, soil protection, and biodiversity conservation. Forests have a multi-layered canopy structure with greater vertical height and continuous, closed canopies; shrublands have lower canopies, denser vegetation structure, and no obvious trunks, with relatively simple vertical stratification; grasslands exhibit single-layer coverage, low vegetation height, and simple structural characteristics. These distinct vegetation structures cause forests, shrublands, and grasslands to play different roles in carbon storage estimation and soil and water conservation [1,2]. However, in satellite imagery, forests, shrublands, and grasslands often exhibit highly similar spectral characteristics, and shrublands frequently present transitional properties between forests and grasslands, posing significant challenges for the fine classification of these vegetation types [3].
Current research on vegetation classification mainly focuses on distinguishing vegetation from non-vegetation or forest from non-forest. Studies focusing on detailed classification among forests, shrublands, and grasslands remain limited. Forests, shrublands, and grasslands share typical vegetation characteristics, which are particularly evident during the growing season. Due to spectral similarity, medium-resolution global land cover datasets, such as LCMAP_Val, FROM_GLC [4], GLASS-GLC [5], ESA WorldCover, and ESRI_Land_Cover, often merge forests, shrublands, and grasslands or treat forests and shrublands as a single classification category [6]. Land cover products such as GLC_FCS30 [7] and GLC_FCS10 [8], which utilize time-series data from Sentinel-1 and Sentinel-2 to generate more detailed global land cover maps, include forests, shrublands, and grasslands. However, in the accuracy assessment process, shrublands and grasslands are merged into a single category. Reference [7] constructed a local adaptive random forest model using Landsat images, achieving producer’s accuracies of 56.8% for shrubland and 67.3% for grassland. These results still have room for improvement compared to the accuracies of 94% for forest and 88% for cultivated land. Moreover, integrating multi-source information can further enhance the classification accuracy of complex land-cover types [6]. High-spatial-resolution remote sensing images provide abundant texture information and detailed spatial features, offering an effective solution for the fine classification of forests, shrublands, and grasslands [9,10,11]. Ref. [12] extracted geometric features, including the shape and size of shrubland patches, together with spectral information from optical imagery at 1 m and 0.5 m resolution, and analyzed the spatial distribution changes of shrubland patches. The overall classification accuracy reached 98%, demonstrating the significant advantages of high-resolution imagery in characterizing the spatial scale features of vegetation patches. Ref. [9] utilized 0.5 m resolution optical images from the SuperView-1 satellite to extract multispectral band information and vegetation indices, achieving fine identification of forest types such as coniferous and broad-leaved forests. Although high-resolution optical images are effective in capturing texture features of land cover [11], they have limited capability in representing differences in canopy height and structural complexity [13]. Synthetic Aperture Radar (SAR) is an active microwave remote sensing technique whose imaging mechanism is directly related to the scattering characteristics of ground targets. Structural differences in canopy height, branch density, and surface coverage patterns among vegetation types lead to different interaction processes with microwave signals, which are manifested as distinct scattering characteristics in SAR imagery [13,14,15]. Borlaf-Mena et al. extracted structural differences among forest types using backscatter and coherence features derived from Sentinel-1 SAR imagery and achieved mapping of temperate and tropical forest cover [16]. Yuan et al. combined spectral, textural, and structural features derived from Sentinel-1 SAR and Sentinel-2 optical imagery to produce classification maps of coniferous, broad-leaved, and mixed forests [17]. Sun et al. constructed an optical feature set including spectral, vegetation index, textural, and terrain features, a SAR feature set including backscatter and polarimetric decomposition features, and an optical–SAR fusion feature set using Landsat-8 optical imagery and ALOS-2 SAR imagery. Forest type classification experiments conducted on these feature sets showed that the accuracy achieved using optical–SAR fusion features was superior to that obtained from a single data source [18]. Existing studies indicate that SAR imagery is highly sensitive to vegetation structure and dielectric properties, whereas optical imagery is effective in capturing morphological and textural features of the land surface. These two data sources complement each other in vegetation remote sensing observations and effectively improve the accuracy of fine vegetation classification [19,20].
Currently, the main approaches for the fine classification of forests, shrublands, and grasslands based on optical and SAR data include traditional machine learning methods and deep learning methods [21,22]. Zhang et al. constructed optical and scattering features of tree canopies using Sentinel-2 optical imagery and Sentinel-1 SAR imagery and employed the support vector machine (SVM) method to extract canopy information across the Sahel region of Africa [23]. Symeonakis et al. integrated spectral, vegetation index, backscatter, gray-level co-occurrence matrix (GLCM) texture, and temporal features derived from aerial imagery, Landsat data, and ALOS PALSAR data and used the random forest (RF) algorithm to evaluate classification accuracy for major savanna cover types. Their results indicated that multimodal features are key factors for accurately mapping savanna land cover [24]. Maskell et al. constructed a feature set including spectral, temporal, textural, and terrain information using Sentinel-1 and Sentinel-2 imagery and applied the RF algorithm to produce 10 m classification maps of coffee production systems in highly heterogeneous small-scale agricultural landscapes, demonstrating the importance of complementary SAR and optical information in complex environments [25]. However, machine learning methods rely heavily on manual feature design and primarily extract shallow features, limiting their ability to represent complex land-cover variability and effectively fuse multimodal remote sensing information [26]. Deep learning models offer greater potential for network architecture design and multimodal feature fusion than traditional machine learning approaches [27]. Ren et al. proposed a dual-stream Swin Transformer fusion network that achieves deep interaction between SAR and optical features through cross-attention mechanisms, improving the efficiency and accuracy of land cover classification [28]. Wang et al. developed a multi-channel semantic segmentation model, FCN-ResNet, to achieve fine extraction of mountainous vegetation [29].
Although existing deep learning methods have made significant progress in the fine classification of forests, shrublands, and grasslands by integrating optical and SAR features, speckle noise in SAR imagery and spatial resolution differences between SAR and optical images affect the stability and accuracy of classification [30,31,32,33]. Moreover, the high cost of acquiring and annotating high-resolution multimodal remote sensing data, together with the limited availability of high-quality labeled samples, remains a major challenge for the fine classification of forests, shrublands, and grasslands [34].
In response to the aforementioned challenges, this study focuses on improving the fine classification of forests, shrublands, and grasslands. Based on Gaofen-2 (GF-2) optical imagery and Gaofen-3 (GF-3) SAR imagery, this study systematically extracts and analyzes multimodal features, including spectral, textural, and scattering characteristics of different land cover types, and constructs a high-resolution multimodal remote sensing feature dataset. In addition, a refined land cover information extraction network framework for high-resolution multimodal remote sensing images, termed CASFNet, is proposed. By adaptively fusing cross-modal feature information and incorporating joint supervision, the proposed framework enhances the learning capability for discriminative features of forests, shrublands, and grasslands. The key contributions of this study are as follows:
  • A high-resolution multimodal fine classification dataset for forests, shrublands, and grasslands, named GF23FSG, is constructed based on optical spectral and SAR remote sensing imagery.
  • A cross-modal adaptive structure fusion network, CASFNet, is designed. It introduces a multi-scale residual gated fusion mechanism and a multi-level auxiliary supervision loss function to enable collaborative learning of spectral and structural features for forests, shrublands, and grasslands, thereby achieving fine classification of these vegetation types.
  • Experimental results on the GF23FSG and SEN12MS datasets demonstrate that CASFNet achieves superior performance in the fine classification of forests, shrublands, and grasslands, with overall accuracies (OA) of 77.38% and 71.84%, respectively, validating the effectiveness of the proposed method for multimodal fine classification.

2. Materials and Methods

2.1. Study Area

The study area is located at the junction of Chengde City in Hebei Province and Xilingol League in the Inner Mongolia Autonomous Region (42.0°–42.2°N, 116.4°–116.9°E), within the transitional zone between the northern foothills of the Yanshan Mountains and the southern margin of the Inner Mongolia Plateau. This region has a mid-temperate continental monsoon plateau climate. The annual mean temperature is approximately 2.3 °C, with annual precipitation of about 350 mm. Elevation ranges from approximately 700 to 2300 m. The terrain is highly undulating and complex, encompassing diverse geomorphic units both above and below the dam. Due to the combined influence of topography and climate, vegetation types in the study area exhibit strong spatial heterogeneity. Forests, shrublands, and grasslands exhibit distinct transitional and mosaic distribution patterns. Different vegetation types show significant differences in structure, coverage, and spatial scale. Numerous transitional land cover types, such as forest–shrubland, shrubland–grassland, and mixed tree–shrub vegetation, are also present, making this area highly suitable for detailed land cover extraction of forests, shrublands, and grasslands. Figure 1 shows the study area.

2.2. Data Collection and Preprocessing

2.2.1. Data Preprocessing

The optical and SAR imagery used in this study were obtained from the Land Observation Satellite Data Service Platform, specifically from GF-2 and GF-3 satellite data. The GF-2 satellite carries five spectral bands, with spatial resolutions of 3.24 m for the blue, green, red, and near-infrared bands and 0.8 m for the panchromatic band. The GF-3 satellite supports 12 imaging modes with full polarization capabilities, including spotlight, stripmap, and scan modes. This study employed DH single-polarization imagery acquired in ultra-fine strip (UFS) mode with a spatial resolution of 3 m. This imagery effectively represents the backscatter intensity and spatial variation of different vegetation types. The satellite payload information of the optical and SAR data used in this study are summarized in Table 1. To ensure data quality, cloud-free GF-2 imagery and SAR data acquired under stable weather conditions were selected, and the acquisition times of both GF-2 and GF-3 were constrained to the peak growing season of vegetation.
The preprocessing of GF-2 imagery includes radiometric calibration, orthorectification, geometric registration, and image fusion. The preprocessing of GF-3 imagery includes multi-look processing, speckle filtering, radiometric calibration, and geocoding. To achieve precise geospatial registration between SAR and optical imagery, stable and easily identifiable ground features such as road intersections, building corners, and water boundaries were selected as ground control points. The co-registration between GF-3 SAR imagery and GF-2 optical imagery was performed using manual reference point selection [35,36]. Finally, through image cropping and mosaicking, the spatial overlap between GF-2 optical imagery and GF-3 SAR imagery was extracted to construct consistent multimodal remote sensing data for the study area. The overall preprocessing workflow is shown in Figure 2.

2.2.2. Forest–Shrubland–Grassland Taxonomy

This study adopts the classification system of the global land cover product GLC_FCS10 [8] and, considering the vegetation distribution characteristics of the study area, further optimizes the forest–shrubland–grassland taxonomy, which includes seven categories: evergreen forest, deciduous forest, evergreen shrubland, deciduous shrubland, low-cover grassland, high-cover grassland, and other. Referring to the grassland classification criteria of the FAO LCCS and considering the actual vegetation distribution in the study area, grassland is further subdivided into low-coverage and high-coverage classes. Meanwhile, following the GLC_FCS10 classification system, evergreen forest, deciduous forest, evergreen shrubland, and deciduous shrubland are retained. Table 2 presents the Forest–Shrubland–Grassland taxonomy.

2.2.3. Field Data Collection

Field data collection was conducted based on the classification system established in this study. For different vegetation types (evergreen forest, deciduous forest, evergreen shrubland, deciduous shrubland, low-cover grassland, and high-cover grassland), targeted sampling points were arranged and field investigations were conducted within the study area. During data collection, unmanned aerial vehicle aerial photography, positioning camera recording, and on-site verification were used to accurately identify and verify the actual land cover type of each survey sample point, thereby ensuring the authenticity and reliability of the sample information.
A total of 310 field survey sample points were collected in this study, primarily located in areas with pronounced vegetation transitions and ambiguous land-cover boundaries, such as zones where low shrubs and grassland intersect, sparse forest areas, and mixed forest–shrubland regions. The samples included 75 evergreen forest, 44 deciduous forest, 51 evergreen shrubland, 74 deciduous shrubland, 32 low-cover grassland, and 34 high-cover grassland sites. Based on the field survey results and GF-2 optical imagery, visual interpretation was performed to assign land-cover labels to unlabeled areas. Subsequently, both the sample points and the interpreted regions were delineated using vector polygon mapping to construct the label dataset. The spatial distribution of the survey sample points and representative field photographs are shown in Figure 3.

2.3. Feature Set Construction

This study extracted spectral features, vegetation indices, and the backscatter coefficient (BC) for evergreen forest, deciduous forest, evergreen shrubland, deciduous shrubland, low-cover grassland, and high-cover grassland from GF-2 and GF-3 multimodal remote sensing imagery. Texture features were also derived from the NIR band of GF-2 and the BC of GF-3 imagery. Given the substantial differences in numerical range, statistical distribution, and units among multi-source features, as well as the skewness and presence of outliers in certain SAR and texture features, quantile normalization was applied to standardize all features. This approach effectively mitigates the influence of extreme values, aligns features to a consistent distribution, enhances feature comparability, and improves the stability and convergence of the model during training. Feature analysis was conducted from two perspectives: spectral and scattering features, and texture features, to explore the distribution differences among vegetation types and identify more discriminative features for the fine classification of forest, shrubland, and grassland.

2.3.1. Spectral and Scattering Feature Analysis

The spectral features used in this study include four bands of the GF-2 imagery (GF2_B1, GF2_B2, GF2_B3, GF2_B4), as well as two widely used remote sensing indices: the Normalized Difference Vegetation Index (NDVI) and the Normalized Difference Water Index (NDWI). NDVI effectively quantifies vegetation cover and growth conditions, whereas NDWI reflects differences in vegetation structure and water content. Together with the spectral bands, these indices characterize the radiometric responses of evergreen forest, deciduous forest, evergreen shrubland, deciduous shrubland, low-cover grassland, and high-cover grassland. The backscatter feature refers to the BC obtained after SAR image processing, which reflects canopy structural characteristics of different vegetation types through variations in scattering signals [37]. The statistical distributions of spectral and backscatter features within the regions of interest (ROI) for the six vegetation types are shown in Figure 4.
Overall, significant differences exist in the spectral and backscatter characteristics among the vegetation types. Due to their lower vegetation height and sparse canopies, low-cover and high-cover grasslands are more susceptible to ground background effects in their spectral responses. Their reflectance in the GF2_B1 (Blue), GF2_B2 (Green), and GF2_B3 (Red) bands is generally higher than that of forests and shrublands. Evergreen and deciduous forests are influenced by dense canopy structures and non-leaf components, resulting in increased water content and shadow effects, which generally lead to lower NDWI values than those of shrubland and grassland. Meanwhile, evergreen and deciduous forests possess more intact canopies and higher leaf area indices, resulting in more active photosynthesis and significantly higher NDVI values than shrubland and grassland. A noticeable difference also exists in NDVI values between low-cover and high-cover grasslands, with high-cover grasslands generally exhibiting slightly higher NDVI values due to greater vegetation cover.
Regarding SAR backscatter characteristics, evergreen and deciduous forests exhibit relatively complex vertical canopy structures, leading to multiple scattering of microwave signals between canopy layers and branches and generally higher backscatter intensities. Among these, evergreen forests exhibit slightly higher backscatter intensity than deciduous forests. This difference is mainly attributed to the small, dense leaves and stable branch–leaf structures of evergreen vegetation, which facilitate stronger volume scattering of microwave signals within the canopy. In contrast, shrubland and grassland mainly exhibit surface scattering and therefore present generally lower BC values. These results demonstrate that GF-2 multispectral bands, NDVI, and SAR backscatter features effectively capture both spectral responses and structural differences among the six vegetation types, providing a reliable foundation for subsequent classification.

2.3.2. Texture Feature Analysis

Texture features describe the gray-level distribution and spatial arrangement of pixels within local image regions and provide supplementary discriminative information for vegetation types with similar spectral characteristics. Texture features—including mean, variance, entropy, contrast, correlation, homogeneity, heterogeneity, and angular second moment—were calculated using the GLCM method for the NIR band of GF-2 and the BC of GF-3 imagery, respectively. Figure 5 illustrates the ROI distribution curves of optical and SAR texture features for evergreen forest, deciduous forest, evergreen shrubland, deciduous shrubland, low-cover grassland, and high-cover grassland.
Regarding optical texture feature distributions, clear differences are observed among evergreen forest, low-cover grassland, and high-cover grassland. Due to dense canopy structures and pronounced shadow effects in evergreen forest, spectral variations within pixels are relatively concentrated, resulting in lower angular second moment and homogeneity values. Meanwhile, indicators reflecting gray-level variation and complexity, such as variance, contrast, heterogeneity, and entropy, exhibit relatively high values. In contrast, low-cover and high-cover grasslands exhibit relatively uniform surface structures and fine textures, and their overall texture distribution characteristics differ markedly from those of forest areas.
Regarding SAR texture features, evergreen shrubland and deciduous shrubland exhibit relatively disordered spatial structures and higher surface roughness due to their natural growth patterns, resulting in higher values of variance, contrast, heterogeneity, and entropy compared with forest and grassland types, while showing lower homogeneity and angular second moment values. Deciduous shrublands are relatively sparsely distributed, whereas evergreen shrublands contain densely arranged small leaves, causing microwave scattering energy to be more concentrated. Therefore, evergreen shrublands generally exhibit higher values of variance, contrast, heterogeneity, and entropy than deciduous shrublands.

2.3.3. Feature Selection

Based on the analysis of the distinctive characteristics of evergreen forest, deciduous forest, evergreen shrubland, deciduous shrubland, low-cover grassland, and high-cover grassland, spectral, scattering, and texture features all exhibit strong discriminative capability across land-cover types. However, incorporating all features into model training would introduce substantial information redundancy due to feature overlap, potentially degrading classification performance. Therefore, a random forest algorithm was employed for feature importance ranking. Random forest is a supervised classification method based on bootstrap resampling and ensemble learning of multiple decision trees. By repeatedly sampling the training data with replacement, multiple independent decision trees are constructed, each producing a classification result, and the final output is determined by the majority voting strategy [38].
In this study, spectral, scattering, and texture features were input into the random forest model to compute feature importance, with the results shown in Figure 6. NDVI exhibits the highest importance score among all features, significantly outperforming other variables, indicating that vegetation coverage and growth status play a dominant role in the fine classification of forest, shrubland, and grassland. The red (GF2_B3), blue (GF2_B1), and green (GF2_B2) bands of GF-2 imagery follow, all showing high importance values, suggesting that surface reflectance characteristics captured by visible bands retain strong discriminative power across vegetation types, and that optical multi-band information effectively represents spectral variability. In contrast, the overall importance of texture features derived from the near-infrared band and SAR backscatter coefficients is relatively low. Among SAR texture features, the entropy feature (BC_Entropy) shows the highest importance, indicating that entropy, which reflects the complexity of surface scattering structures, provides a certain level of discriminative contribution.
In summary, contrast, heterogeneity, and entropy derived from both optical and SAR imagery demonstrate strong discriminative capability and effectively capture differences in spatial variability and canopy structural complexity among the six vegetation types. However, contrast and heterogeneity mainly reflect differences between neighboring pixels, whereas entropy characterizes texture patterns from the perspective of global information distribution and provides more stable classification performance across vegetation types. Therefore, considering feature discriminative ability, stability, and redundancy, entropy was selected as the representative texture feature in this study.

2.3.4. GF23FSG Feature Set

Based on the analysis of spectral, scattering, and texture features of six vegetation types, this study used GF2_B1, GF2_B2, GF2_B3, GF2_B4, NDVI, BC, and the BC_Entropy texture feature to construct a multimodal fine classification dataset for forest, shrubland, and grassland (GF23FSG). The introduction of selected features is shown in Table 3.
Among them, x i denotes the original value of the input feature, F X ( · ) represents the empirical cumulative distribution function (ECDF), and Φ 1 ( · ) denotes the inverse cumulative distribution function of the standard normal distribution (i.e., the probit function).
The GF23FSG dataset contains seven land cover types (evergreen forest, deciduous forest, evergreen shrubland, deciduous shrubland, low-cover grassland, high-cover grassland, and other), seven features (five optical features and two SAR features), an image size of 256 × 256 pixels, a spatial resolution of 1 m, and a total of 6691 feature–label pairs. Examples of the GF23FSG dataset are presented in Figure 7.
This study divides the GF23FSG dataset into a training set and a validation set at a ratio of 8:2. The training set contains 5352 images, while the validation set includes 1339 images. The class distribution is presented in Table 4, indicating that the samples are generally well balanced across categories.

2.4. Methodology

We propose a fine classification information extraction network for high-resolution multimodal remote sensing imagery, termed CASFNet. Figure 8 illustrates the overall architecture of the proposed model. Given the substantial differences in imaging mechanisms, noise characteristics, and discriminative features between optical and SAR sensors, a shared encoder struggles to effectively capture both the fine texture and boundary details from optical imagery and the structural scattering information from SAR data. Therefore, two independent encoders are employed to extract features from GF23FSG_OptFeature and GF23FSG_SARFeature, respectively. The Neighborhood Interaction High-Resolution Network (NHRNet) serves as the backbone for extracting GF23FSG_OptFeature, enhancing the semantic hierarchy of features and preserving high-resolution spatial information, whereas the SAR encoder adopts ResNet101 as the backbone to strengthen deep semantic representations. Feature information from different modalities is fused through the cross-modal adaptive structure fusion module (CASF-Module) and subsequently fed into the decoder to generate the segmentation results. In addition, a multi-auxiliary supervision loss function (MASLoss) is designed, which integrates multi-level auxiliary supervision and class imbalance awareness to fully exploit multimodal feature information while mitigating negative effects such as noise interference.

2.4.1. Optical Encoder

The optical feature GF23FSG_OptFeature contains rich fine-grained information, including textures, lines, and shadows. However, in complex regions where forest, shrubland, and grassland are intermixed, conventional downsampling encoders tend to lose high-resolution details. The high-resolution network NHRNet is employed to learn optical modality features. Its architecture inherits the structural advantages of HRNet, including multi-resolution parallel processing and cross-scale information exchange, and introduces a cross-scale neighborhood interaction layer in the final stage to further enhance local structural representation. The architecture is illustrated in Figure 9.
First, the input optical feature data X o p t passes through two convolutional layers for initial feature extraction, producing a preliminary high-resolution representation. Subsequently, the network maintains a high-resolution branch as the main backbone while gradually introducing additional lower-resolution branches, forming a multi-branch parallel architecture. Each branch maintains its own spatial resolution and channel dimensions while exchanging information through cross-scale interaction units, enabling parallel modeling of high-resolution spatial details and low-resolution semantic features.
Unlike HRNet, NHRNet introduces a neighborhood-based interaction layer (NI Layer) after feature extraction to perform cross-scale information fusion. This layer performs cross-scale interaction and feature re-integration within local neighborhood regions of the multi-scale outputs, aiming to enhance texture continuity and shape consistency in boundary areas. The resulting interaction features replace the original features and are directly used to produce the final optical multi-scale feature set. The neighborhood interaction structure in NHRNet enhances local structural modeling capability without sacrificing high-resolution representation, making the network more sensitive to subtle land cover differences among forest, shrubland, and grassland.

2.4.2. SAR Encoder

GF23FSG_SARFeature contains rich information about structural characteristics such as tree canopy roughness and branch–leaf structures; however, it is also affected by speckle noise and coarse spatial details. To obtain a more robust structural semantic representation, the SAR modal encoder adopts the ResNet101 model, which has strong semantic representation capability and a deep residual architecture, as the backbone network. Its structure is shown in Figure 10.
The network first applies a 7 × 7 convolution block followed by a 3 × 3 max pooling layer to process the input GF23FSG_SARFeature. Through multiple 3 × 3 convolution and pooling operations, the SAR features are gradually transformed into stable high-dimensional representations, particularly in regions with weak optical textures. Subsequently, the feature map is passed through four ResNet layers composed of different numbers of bottleneck blocks, progressively expanding the receptive field and extracting higher-level semantic features to obtain a multi-scale SAR feature set.
The deep residual architecture of the ResNet101 model enables speckle noise to be mapped into more stable high-dimensional feature representations through multiple convolution and aggregation layers. Especially in regions with weak optical textures or areas affected by shadows, it provides more reliable structural cues for subsequent cross-modal feature fusion.

2.4.3. Cross-Modal Feature Fusion Strategy

The proposed cross-modal adaptive structure fusion module (CASF-Module) consists of three components: multi-scale residual information gating (Region-aware Gate, R-Gate), local window attention-enhanced cross-modal skip (WinAttnSkip), and confidence-driven soft gating (Soft-Gate). These components enable adaptive fusion of multimodal features in a layer-wise, region-wise, and semantic-wise manner. The structure of the CASF module is shown in Figure 11.
(1)
Multi-scale Residual Information Gating (R-Gate)
Due to the significant difference in the number of feature channels between the optical and SAR branches, four 1 × 1 convolution projection layers are first applied to compress the channels to (192, 96, 48, 24). The optical and SAR features at the same scale are then fed into the Gated Channel Weighting (GCW) module to compute adaptive fusion weights W i , and the features are fused as follows:
F i = W i O F i O + W i S F i S
F i O ,   F i S ( i = 1 ,   2 ,   3 ,   4 ) represent the optical and SAR features at each scale, respectively. The fused features dynamically adjust the contributions of the two modalities according to the learned gating weights. In regions where optical features provide advantages in texture, spectral, and shape information, optical features contribute more to the fused representation. Conversely, in regions where shadow robustness, penetration capability, and vertical structural information are more important, SAR features contribute more significantly. This mechanism enables adaptive multimodal feature fusion.
(2)
Local Window Attention-Enhanced Cross-modal Skip (WinAttnSkip)
To alleviate the loss of boundary details caused by overly smooth deep semantic features, local window attention enhancement is applied to the deepest fused feature F 4 . Specifically, F 4 is divided into multiple non-overlapping 7 × 7 windows C 1 ,   C 2 , , and window-based attention (WinAttn) performs attention computation within local windows using a query–key–value mechanism:
F 4 = W i n A t t n S k i p ( F 4 ) + F 4
This local attention mechanism highlights salient structures within each window and refines local spatial representations. The WinAttnSkip module outputs feature maps with locally enhanced semantics, which better capture texture similarities within vegetation regions and further enhance the discriminative differences among forest, shrubland, and grassland classes.
(3)
Confidence-driven Soft-Gate (Soft-Gate)
To further suppress the speckle noise of SAR, and perform reliability correction for uncertain regions, we construct a confidence branch prediction W 4 on the highest-level fused feature F 4 and re-outputs F 4 with secondary weighting:
To further suppress SAR speckle noise and correct unreliable responses in uncertain regions, a confidence prediction branch is introduced to generate a confidence map W 4 based on the highest-level fused feature F 4 . The feature map is then re-weighted as follows:
F 4 = F 4 ( 0.5 + 0.5 W 4 )
The Soft-Gate constrains the scaling range to [0.5, 1], preventing excessive feature suppression while emphasizing features with higher semantic confidence. This mechanism improves the smoothness, stability, and semantic consistency of the fused multimodal features.

2.4.4. Decoder

At the beginning of the decoder, the fused features F 2 ,   F 3 ,   F 4 generated by the fusion module are first bilinearly upsampled to the same spatial resolution as F 1 , and then concatenated to obtain multi-scale fused features. These features are subsequently fed into the classifier module to generate the final classification prediction. The classifier module first applies a 1 × 1 convolution layer with batch normalization (BN) and ReLU activation, followed by two convolutional layers with kernel sizes of 1 × 1 and 3 × 3. Finally, the decoder output is upsampled to the original image resolution using interpolation. The structure of the Decoder is shown in Figure 12.

2.4.5. Loss Function Design

To enable the model to fully exploit multimodal information during training, this study designs a multi-auxiliary supervision loss (MASLoss), which jointly supervises the main output, optical auxiliary output, and SAR auxiliary output. MASLoss consists of two components: (1) the main loss, corresponding to the final prediction after multimodal fusion; (2) the auxiliary loss, corresponding to the independent predictions from the optical and SAR branches. The structure of the MASLoss is shown in Figure 13.
(1)
Main Loss
The main branch utilizes the fused optical and SAR features to produce the final semantic prediction. The loss function is defined as:
L m a i n = D i c e C E L o s s ( P m a i n , Y )
where P m a i n denotes the logits of the main segmentation head, and DiceCELoss integrates Dice loss and cross-entropy (CE) loss into a unified loss function. Compared with commonly used segmentation losses such as CE, Dice, and Focal loss, DiceCELoss achieves a better balance in the optimization objective. Dice loss emphasizes class imbalance, while CE loss stabilizes the predicted probability distribution and improves overall convergence. The combined DiceCELoss therefore ensures pixel-level classification accuracy while alleviating the gradient bias caused by the imbalance among vegetation classes, thereby enhancing the representation capacity of the fusion network. This makes it particularly suitable for the fine classification of forest, shrubland, and grassland under the multimodal fusion framework adopted in this study. Since the main loss plays a dominant supervisory role in the overall training process, its weight is set to W m a i n = 1.0 .
(2)
Auxiliary Loss
Since the shallow multimodal features contain richer local textures, they can enhance the representation of vegetation structures and perform more robustly when handling complex textures such as vegetation boundaries and striped planting patterns. This enables the model to better distinguish subtle structural differences among vegetation types. Therefore, the F 1 features from the two modalities are used to construct auxiliary prediction heads, and independent auxiliary losses are introduced for supervision:
L o p t = D i c e C E L o s s ( P o p t , Y )
L s a r = D i c e C E L o s s ( P s a r , Y )
where P o p t represents the logits of the optical auxiliary segmentation head, and P s a r represents the logits of the SAR auxiliary segmentation head. Since the optical auxiliary supervision provides stronger positive guidance for the deep fused features, its contribution to the classification of forest, shrubland, and grassland is generally higher than that of SAR. Therefore, the initial weight of the optical auxiliary head is set to w o p t = 0.5 , while the initial weight of the SAR auxiliary head is set to w s a r = 0.2 . This strategy allows the model to utilize structural cues from SAR data, such as edge and roughness information, while preventing the training process from being excessively influenced by SAR speckle noise.
To further enhance the model’s adaptability to quality differences between multimodal data sources, a confidence branch is introduced to enable dynamic weight adjustment. Through adaptive weighting, when a particular modality is strongly affected by noise, its corresponding loss weight can be automatically reduced, which helps improve training stability. Based on the confidence output from the Soft-Gate, the weights of the auxiliary losses are adaptively adjusted as follows:
W s a r = w s a r 0.2 c o n f
W o p t = w o p t + ( 1 c o n f )
The final total loss function is defined as:
L = L m a i n W m a i n + L o p t W o p t + L s a r W s a r
MASLoss combines the main loss and the two auxiliary losses through a weighted summation. Through this multi-level and multimodal joint supervision mechanism, the model not only enhances the representation capability of the fused features but also significantly improves its generalization performance in complex vegetation scenes.

3. Results

3.1. Experimental Details

All experiments in this study were implemented in Python 3.9. The experiments were conducted on two workstations equipped with the Windows 10 operating system and NVIDIA 3090 (24 GB VRAM) and NVIDIA 4090 (24 GB VRAM) GPUs, respectively. To ensure experimental consistency, all models were trained using the same training parameters and strategies: a batch size of 4, an initial learning rate of 0.0005, weight decay of 0.01, the Adam optimizer with a momentum of 0.9, and 80 training epochs. The dataset was divided into a training set and a validation set at a ratio of 8:2. During training, MASLoss was adopted as the training loss function, while DiceCELoss, which combines Dice loss and cross-entropy loss, was used as the validation loss.

3.2. Evaluation Metrics

To evaluate the effectiveness of the proposed method, several evaluation metrics were employed, including overall accuracy (OA), F1-score, intersection over union (IoU), mean F1-score (mF1), mean intersection over union (mIoU), and the Kappa coefficient. Among these metrics, mIoU is adopted as the primary evaluation criterion, as it provides a comprehensive assessment of segmentation performance across all classes and is particularly well suited for multi-class vegetation segmentation with imbalanced class distributions. In contrast, F1-score, OA, and the Kappa coefficient are used as auxiliary metrics to reflect overall classification consistency.

3.3. Experimental Results and Analysis

3.3.1. Experimental Results

From the overall classification metrics, the model achieved an OA of 77.38% and a Kappa coefficient of 75.02% across all vegetation categories. These results indicate that the model demonstrates good overall performance and stability in distinguishing the six refined vegetation classes. From the class-wise results, evergreen forest and high-cover grassland achieved the best classification performance, with F1-scores of 82.68% and 80.76%, respectively, and the corresponding IoU values exceeding 67%. This suggests that for vegetation types with relatively stable canopy structures and high vegetation coverage, multimodal features can provide more discriminative information, thereby improving classification accuracy. In contrast, the classification performance of deciduous forest was relatively lower, with an F1-score of only 62.44%. This is mainly because deciduous forest exhibits significant seasonal variability, and its spectral and canopy structural characteristics partially overlap with those of shrubland and grassland, making precise discrimination more challenging for the model. Table 5 presents the detailed classification performance of the proposed method on the test set.
A normalized confusion matrix is used to analyze the separability among forest, shrubland, and grassland classes. The diagonal elements represent the model’s classification accuracy for each class. High-cover grassland and evergreen forest achieve the highest accuracies, both reaching 81%, indicating strong intra-class consistency. Deciduous shrubland, evergreen shrubland, and low-cover grassland attain accuracies of 78%, 75%, and 71%, respectively, reflecting relatively robust performance. In contrast, deciduous forest shows a lower accuracy of 58%, which is likely influenced by its limited sample size. Inter-class confusion exhibits two primary patterns. The first is transitional confusion between forest and shrubland classes. Owing to similar structural and spectral characteristics, 22% of deciduous forest is misclassified as evergreen forest and 11% as deciduous shrubland, while 10% of evergreen shrubland is misclassified as evergreen forest. This indicates that the model is prone to cross-category confusion in transition zones and areas with low-stature trees. The second pattern arises from feature shifts associated with differences in vegetation coverage. Low-cover grassland is strongly affected by the bare soil background, with approximately 10% of samples misclassified into other categories, and some confusion observed with evergreen and deciduous shrubland. This suggests that its spectral and microwave scattering responses tend to resemble those of sparse shrub classes, thereby increasing classification difficulty. Figure 14 presents the normalized confusion matrix of the proposed method on the test set.
The proposed method effectively preserves the spatial integrity of vegetation patches, producing well-defined class boundaries with minimal misclassification or omission. In areas where evergreen forest is interspersed with evergreen shrubland and deciduous shrubland, the classification results maintain strong spatial continuity. These observations suggest that the multimodal feature fusion strategy improves the robustness and reliability of vegetation classification under complex land-cover conditions. The classification results are presented in Figure 15.
Overall, the differences in classification performance among the vegetation classes reflect variations in their spectral responses, structural characteristics, and scattering properties. These findings also provide a basis for further analysis of classification performance from the perspectives of data sources and model architecture in future research.

3.3.2. Comparison Experiments

To comprehensively evaluate the performance of the proposed method for the fine classification of forest, shrubland, and grassland, several mainstream multimodal remote sensing networks were selected for comparison, including MFNet [39], ASMFNet [40], CMFNet [41], DDHRNet [42], MGFNet [43], and FTransUNet [44]. Among these models, DDHRNet, FTransUNet, CMFNet, and MGFNet employ dual-branch CNN-based encoders, whereas MFNet and ASMFNet utilize Transformer-based feature extraction networks to capture multimodal feature representations. All methods were trained and evaluated using the same training and validation samples to ensure a fair comparison. Table 6 presents the quantitative comparison results on the GF23FSG dataset, including the IoU for each class, as well as OA, mIoU, mF1, and the Kappa coefficient.
According to the quantitative results, the proposed method achieves the best overall performance among all compared models, with OA, mIoU, mF1, and Kappa values of 77.38%, 60.50%, 75.02%, and 72.84%, respectively. Compared with the second-best model FTransUNet, our method improves the performance by 3.37% in OA, 3.32% in mIoU, 2.45% in mF1, and 3.99% in Kappa. Compared with other multimodal approaches, the proposed model shows particularly improved performance in the classification of forest and shrubland categories, indicating that the CASF-Module and MASLoss effectively integrate complementary multimodal information, suppress noise interference, and enhance feature representation capability.
The qualitative comparison results further demonstrate that the proposed model produces more accurate segmentation results than the competing methods. It not only distinguishes the fine categories of forest, shrubland, and grassland, but also captures clear boundary details between different land-cover types, resulting in improved segmentation accuracy and detail preservation.In contrast, several competing methods exhibit different types of classification errors. The FTransUNet model misclassified deciduous shrubland as evergreen forest, and also showed missed detections in low-cover grassland areas. MFNet and CMFNet produced large-area regional misclassification errors. For example, in the first row of the results, CMFNet misclassified evergreen shrubland as low-cover grassland, while MFNet incorrectly identified deciduous shrubland as evergreen forest in the second row.The ASMFNet model exhibited insufficient boundary discrimination capability, resulting in noticeable boundary blurring in the classification maps. DDHRNet showed relatively poor recognition performance for grassland categories, and in the third row of the results, it misclassified low-cover grassland as high-cover grassland. The MGFNet model achieved acceptable segmentation results for large homogeneous regions and boundary areas; however, confusion between evergreen shrubland and deciduous shrubland was still observed. Overall, the comparison results demonstrate that the proposed model achieves superior performance, particularly in preserving fine structural details and maintaining clear boundaries between vegetation categories. The visualization results on the GF23FSG dataset are illustrated in Figure 16.

3.4. Comparison Experiments with Different Input Data

To verify the effectiveness of each modality in the multimodal dataset, optical data, SAR data, and their corresponding feature data were progressively introduced as model inputs for comparative experiments. As shown in Table 7, compared with using optical data alone, the introduction of SAR data improves the model performance, with OA, mIoU, mF1, and Kappa increasing by 2.57%, 3.92%, 3.58%, and 3.21%, respectively. Among the vegetation categories, evergreen shrubland and deciduous shrubland exhibit the most significant improvements. This result is consistent with the differences in backscatter characteristics between shrubland and other vegetation types, indicating that SAR data provides complementary canopy structural information that effectively compensates for the limitations of optical data in representing vegetation structure. After further introducing optical feature data, the IoU of high-cover grassland increases from 60.57% to 65.48%, indicating that optical features can further enhance the discriminative capability for vegetation types with high coverage. Finally, when the model input integrates optical data, optical features, SAR data, and SAR features, the model achieves the best classification performance, with OA and Kappa values reaching 77.38% and 72.84%, respectively. These results demonstrate that integrating multimodal remote sensing data with diverse feature representations enables the model to better capture the refined land-cover characteristics of forest, shrubland, and grassland, thereby improving classification accuracy and stability. The results further confirm the effectiveness of the proposed multimodal feature fusion strategy for vegetation fine classification.
When only GF-2 optical data were used as the model input, large areas of low-cover grassland were misclassified as evergreen shrubland, resulting in significant confusion between these two land-cover categories. After introducing SAR data, this confusion was substantially reduced, and the classification accuracy improved noticeably. In the first-row results, the boundary of deciduous forest became clearer compared with the results obtained using only optical data; however, confusion between high-cover grassland and deciduous shrubland still remained. After further introducing the NDVI feature, the classification accuracy of low-cover grassland and evergreen shrubland improved in the second-row results, and the boundary delineation became more precise. Finally, after incorporating the BC_Entropy texture feature, the confusion between high-cover grassland and deciduous shrubland was eliminated in the fourth-row results, and the extraction performance of all vegetation categories was further improved, resulting in the best overall classification outcome. The classification results of the refined land-cover types using different data sources as model inputs are illustrated in Figure 17.

3.5. Ablation Experiments

This section evaluates the effectiveness of each data source in GF23FSG and the contributions of the individual modules in CASFNet. In the ablation experiments, several key modules and structures in both the data inputs and the network architecture were individually removed or replaced. All ablation experiments were conducted on the GF23FSG dataset using the same training settings to ensure a fair comparison.
CASFNet consists of two main components: feature extraction and feature fusion. In the feature extraction stage, the neighborhood-based interactive NHRNet was designed as the backbone network for the optical feature branch, while SAR data were incorporated to form a dual-modality remote sensing input. In the feature fusion stage, the CASF-Module was introduced to adaptively fuse the features from the two modalities. In addition, a multi-auxiliary supervision loss (MASLoss) was designed, which introduces a confidence branch to enable dynamic weight adjustment during training. Through experiments involving different module configurations, the effectiveness of each component was verified. In the ablation tables, the symbol “✓” indicates that the corresponding module was included in the experiment.
Table 8 demonstrates the advantage of the NHRNet branch over the original HRNet in the feature dataset. The NI Layer enables the model to capture richer local structural information from optical data. When using identical backbone networks for both the optical and SAR branches and introducing the proposed adaptive fusion module, the model achieves an OA of 74.61%, indicating a clear improvement in feature extraction performance. Furthermore, after incorporating MASLoss, and using the same dataset, hyperparameters, and training strategy, the model performance improves significantly, achieving an OA of 77.38%, mIoU of 60.50%, mF1 of 75.02%, and a Kappa coefficient of 72.84%. These results indicate that MASLoss improves the overall segmentation accuracy and stability of the model.

3.6. Cross-Dataset Comparison Experiment

Although the proposed model demonstrates strong performance on the GF23FSG dataset, its applicability to data acquired from different sensors and imaging conditions still needs to be further verified. Therefore, to evaluate the generalization capability of CASFNet under different data sources, additional experiments were conducted on the SEN12MS dataset, which contains data collected from different sensors, acquisition conditions, and land-cover distributions.
SEN12MS [45] is a large-scale multimodal public dataset jointly developed by the Technical University of Munich (TUM) and Helmholtz Zentrum München. The dataset contains 180,662 pairs of corresponding Sentinel-1 dual-polarization SAR images, Sentinel-2 multispectral images, and MODIS-derived land-cover maps from various regions around the world. The data cover four seasons and include four label categories, with an image size of 256 × 256, a spatial resolution of 10 m, and land-cover labels derived from MODIS with a spatial resolution of 500 m.
In this study, 5000 image pairs from the summer vegetation growth period were randomly selected from the SEN12MS dataset for experimentation. The land-cover labels were reorganized and merged, and the category mapping is presented in Table 9. After label remapping, images containing all-zero pixel values were removed. In addition, because the number of Mixed Forest samples is limited and the difficulty of merging this category with other vegetation classes, images containing the Mixed Forest category were excluded from the dataset. After this filtering process, a total of 3518 image pairs were retained for the experiments. A visualization example of the SEN12MS dataset is shown in Figure 18.
The quantitative segmentation results on the SEN12MS dataset are presented in Table 10. The proposed CASFNet achieves an OA of 71.84%, mIoU of 55.51%, mF1 of 70.51%, and a Kappa coefficient of 66.42%. Although the spatial resolution of the SEN12MS dataset is lower than that of the GF23FSG dataset, the proposed model still outperforms the other comparison methods in terms of overall accuracy. Compared with FTransUNet, CASFNet improves OA by 6.30%, mIoU by 6.64%, mF1 by 5.39%, and Kappa by 11.54%. Although the original SEN12MS labels were remapped and categories were merged to align with the classification system used in this study—thereby reducing the complexity of the original scheme—the experimental results still demonstrate the capability of CASFNet to extract fine-grained information for forest, shrubland, and grassland across remote sensing imagery with varying spatial resolutions. In particular, the model shows more significant improvements in shrubland and grassland categories, demonstrating the effectiveness of the proposed multimodal fusion strategy under cross-dataset conditions.
The qualitative comparison results on the SEN12MS dataset are illustrated in Figure 19. Compared with the other models, the segmentation results of CASFNet exhibit higher spatial consistency and clearer land-cover boundaries. In the first row, most competing models fail to accurately identify small-scale grassland regions, whereas CASFNet preserves these fine structures more effectively. In the second row, the MFNet, ASMFNet, CMFNet, and MGFNet models show discontinuous extraction results for tropical savannas, while CASFNet produces more spatially coherent classification results. In the fourth row, MFNet, CMFNet, and FTransUNet exhibit large-area misclassification in land-cover extraction. Although MGFNet and CASFNet produce relatively better results, CASFNet still demonstrates clearer boundaries and improved recognition of small-scale land-cover regions.

4. Discussion

The spectral characteristics of forest, shrubland, and grassland exhibit considerable overlap, and the differences in their spatial structure and scattering properties cannot be fully captured by a single data source. This makes the fine-grained classification of these vegetation types particularly challenging. Therefore, this study further analyzes the influence of classification granularity and sample balance on the classification performance of forest, shrubland, and grassland.

4.1. Discussion on the Classification of Forest, Shrubland and Grassland

To investigate this issue, two classification schemes were designed based on the GF23FSG and SEN12MS datasets: (1) coarse classification, including forest, shrubland, and grassland; (2) fine classification, including evergreen forest, deciduous forest, evergreen shrubland, deciduous shrubland, low-cover grassland, high-cover grassland, and other (see Section 3.3.1). The experimental results are presented in Table 11. For the GF23FSG dataset, the coarse classification achieves an OA of 82.70% and an mIoU of 70.69%. Similarly, for the SEN12MS dataset, the coarse classification achieves an OA of 82.74% and an mIoU of 70.69%. Compared with the results of the fine classification (Table 5), the OA increases by 5.33% and 9.72%, respectively. These results indicate that the feature differences among forest, shrubland, and grassland are relatively subtle. As a result, fine-grained classification becomes more difficult, particularly for vegetation types with similar spectral and structural characteristics, such as evergreen forest vs. deciduous forest, evergreen shrubland vs. deciduous shrubland, and grassland types with different coverage levels. Distinguishing between these vegetation classes remains a challenging task for remote sensing-based classification.

4.2. Discussion on Sample Imbalance

The distribution of the sample dataset has a decisive impact on model performance. Due to the differences in imaging mechanisms between optical imagery and SAR data, the availability of high-quality multimodal remote sensing images that simultaneously cover the study area is relatively limited. This constraint makes it difficult to maintain a balanced distribution of sample quantities across all land-cover categories in the dataset constructed in this study. Among the categories, the proportions of deciduous forest and deciduous shrubland samples are relatively low, with deciduous forest accounting for only 2.97% of the total samples. The insufficient number of samples restricts the model’s ability to effectively learn discriminative features for these classes to some extent. As shown in Figure 20, there is an overall correlation between the classification accuracy of different vegetation types and their sample quantities. Categories with fewer samples often struggle to adequately learn their feature distributions, resulting in relatively lower classification accuracy. However, for categories such as evergreen forest and high-cover grassland, the distinguishing features are more pronounced. Therefore, despite variations in sample distribution, their classification accuracy can still remain at a relatively high level.

4.3. Analysis of the Impact of Imaging Conditions

Forest, shrubland, and grassland exhibit substantial overlap in both spectral and structural characteristics, and their separability is highly dependent on imaging conditions. To enhance discrimination among vegetation types, this study imposes constraints on data acquisition. The GF-2 and GF-3 imagery was acquired during the peak growing season, when vegetation biomass is maximized and canopy closure is high, providing more abundant and stable spectral, structural, and texture information that facilitates the differentiation of vegetation types. In addition, cloud-free GF-2 imagery was selected to minimize the effects of cloud cover and atmospheric scattering, while GF-3 SAR data were acquired under clear weather conditions, reducing fluctuations in vegetation canopy structure and soil moisture caused by rainfall or dew. As a result, SAR backscatter signals can more reliably capture differences in the geometric structure of trees, shrubs, and grasses. However, this strategy may introduce certain limitations. During the peak growing season, different vegetation types may exhibit the “same spectrum, different objects” phenomenon. For example, evergreen and deciduous shrubland can show highly similar spectral responses under strong chlorophyll reflectance, increasing classification difficulty. Moreover, separability under other phenological conditions was not considered. Future work will further investigate classification stability across different phenological stages.
To further evaluate the capability of the proposed method for fine-grained classification of forest, shrubland, and grassland, the classification results were compared with several widely used global land-cover products based on 310 validation sample points collected in the study area. Three representative global land-cover datasets were selected for comparison, including ESA WorldCover 2021, Esri Global Land Cover, and GLC_FCS10 [8]. The detailed information of these datasets is presented in Table 12.
The validation results of the different land-cover products are summarized in Table 13. The method proposed in this study achieves the highest classification accuracy on the same validation samples, with an OA of 88.71%, which is significantly higher than that of Esri Global Land Cover, GLC_FCS10, and ESA WorldCover. These results indicate that the proposed multimodal remote sensing feature representation and classification model can more effectively capture the differences in spectral characteristics, spatial structure, and scattering properties among vegetation types, thereby improving the recognition capability for complex vegetation categories. It should be noted that the dataset used in this study was annotated with reference to field survey sample points, resulting in a certain degree of consistency between the model predictions and the ground validation samples.
Among the three public land-cover products, Esri Global Land Cover achieves the highest OA (53.23%), followed by GLC_FCS10 and ESA WorldCover. The relatively higher accuracy of Esri Global Land Cover is mainly attributed to its coarser classification scheme, in which multiple vegetation types (such as forest, shrubland, and grassland) are aggregated into fewer vegetation categories, leading to higher apparent accuracy in the validation results. In contrast, ESA WorldCover shows significantly lower accuracy because it fails to correctly identify the shrubland category, which is a key class in this study. Although the classification system of ESA WorldCover is more detailed than that of Esri Global Land Cover, the absence of accurate shrubland identification leads to the lowest OA among the compared products. GLC_FCS10, while adopting a more detailed land-cover classification system, still shows limited accuracy in distinguishing the refined vegetation categories.
Compared with existing public land-cover products, which are primarily designed for large-scale global land-cover mapping, the proposed method specifically focuses on the fine classification of forest, shrubland, and grassland. By integrating optical spectral information, SAR backscatter features, and texture structural information, the proposed approach can more effectively characterize differences in canopy structure and spatial distribution among vegetation types, thereby improving classification performance for these refined vegetation categories.
In conclusion, through comprehensive comparisons with publicly available datasets from different satellite platforms and multiple global land-cover products, the proposed method demonstrates superior performance. This indicates strong generalizability in both feature representation and classification framework design, with adaptability to different sensor types and spatial resolutions. Compared with medium-resolution data such as the Sentinel series and existing global land-cover products, the high-resolution forest–shrub–grass dataset constructed from GF-2 and GF-3 imagery enables more precise characterization of vegetation features. Under complex land-cover conditions, it exhibits enhanced class separability and improved fine-scale mapping capability, highlighting the significant potential of high-resolution domestic satellite data for detailed remote sensing applications.

5. Conclusions

This paper analyzes the optical and SAR features of forests, shrublands, and grasslands using GF-2 and GF-3 multimodal high-resolution remote sensing images and constructs the GF23FSG feature dataset. Based on this dataset, the CASFNet information extraction framework for high-resolution multimodal remote sensing images is proposed to address cross-modal feature fusion and learning challenges and to exploit complementary advantages among different data modalities. Experimental results on the GF23FSG feature set demonstrate that the proposed method outperforms SOTA models and achieves strong generalization performance on the SEN12MS dataset. Through dataset construction, method optimization, and multi-dimensional validation, this study provides a feasible solution for the fine-grained classification of forests, shrublands, and grasslands. Due to variations in imaging time and acquisition frequency among multi-source satellites, the dataset constructed in this study exhibits class imbalance. Future work will focus on collecting additional samples and multi-temporal remote sensing data and leveraging large-scale models to explore high-precision classification methods for forests, shrublands, and grasslands under modal and temporal alignment mismatches.

Author Contributions

Conceptualization, Q.P. and X.M.; methodology, Q.P.; validation, Q.P. and X.M.; investigation, Q.P., Z.G., J.Z. (Jilong Zhang) and K.D.; data curation, Q.P.; writing—original draft preparation, Q.P.; writing—review and editing, Q.P., X.M. and W.D.; visualization, Q.P. and J.Z. (Jian Zhang); supervision, X.M., J.Y. and Z.Y.; funding acquisition, X.M., J.Y. and W.D. All authors have read and agreed to the published version of the manuscript.

Funding

This work was funded by the Civil Aerospace Technology Pre-research Project of China’s 14th Five-Year Plan (Grant No. D040404), the Shandong Provincial Key R&D Program of China (Grant No. 2024TSGC0428), the “Double First-Class” discipline cultivation in Surveying and Mapping Science and Technology (Grant No. GCCYJ202418), the Natural Science Foundation of Henan Province (Grant No. 252300421847), Foreign Technical Cooperation Research Project (Aerospace) (Grant No. HE02), and the National Natural Science Foundation of China (Grant No. U22A20620/003).

Data Availability Statement

The raw data will be made available on the request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Huang, J.W.; Li, Z.Y.; Chen, E.X.; Zhao, L.; Mo, B.P. Classification of plantation types based on WFV multispectral imagery of the GF-6 satellite. Natl. Remote Sens. Bull. 2021, 25, 539–548. [Google Scholar]
  2. Irfan, A.; Li, Y.; E, X.; Sun, G. Land Use and Land Cover Classification with Deep Learning-Based Fusion of SAR and Optical Data. Remote Sens. 2025, 17, 1298. [Google Scholar] [CrossRef] [Scilit]
  3. Clark, M.L.; Kilham, N.E. Mapping of land cover in northern California with simulated hyperspectral satellite imagery. ISPRS J. Photogramm. Remote Sens. 2016, 119, 228–245. [Google Scholar] [CrossRef] [Scilit]
  4. Gong, P.; Liu, H.; Zhang, M.; Li, C.; Wang, J.; Huang, H.; Clinton, N.; Ji, L.; Li, W.; Bai, Y.; et al. Stable classification with limited sample: Transferring a 30-m resolution sample set collected in 2015 to mapping 10-m resolution global land cover in 2017. Sci. Bull. 2019, 64, 370–373. [Google Scholar] [CrossRef] [Scilit]
  5. Liu, H.; Gong, P.; Wang, J.; Clinton, N.; Bai, Y.; Liang, S. Annual dynamics of global land cover and its long-term changes from 1982 to 2015. Earth Syst. Sci. Data 2020, 12, 1217–1243. [Google Scholar] [CrossRef] [Scilit]
  6. Lesiv, M.; Fritz, S.; Dürauer, M.; Georgieva, I.; Buchhorn, M.; Bertels, L.; Tsendbazar, N.; Van De Kerchove, R.; Zanaga, D.; Schepaschenko, D.; et al. A global reference data set for land cover mapping at 10 m resolution. Earth Syst. Sci. Data Discuss. 2025, 17, 6149–6155. [Google Scholar] [CrossRef] [Scilit]
  7. Zhang, X.; Liu, L.; Chen, X.; Gao, Y.; Xie, S.; Mi, J. GLC_FCS30: Global land-cover product with fine classification system at 30 m using time-series Landsat imagery. Earth Syst. Sci. Data 2021, 13, 2753–2776. [Google Scholar] [CrossRef] [Scilit]
  8. Zhang, X.; Liu, L.; Zhao, T.; Zhang, W.; Guan, L.; Bai, M.; Chen, X. GLC_FCS10: A global 10 m land-cover dataset with a fine classification system from Sentinel-1 and Sentinel-2 time-series data in Google Earth Engine. Earth Syst. Sci. Data 2025, 17, 4039–4062. [Google Scholar] [CrossRef] [Scilit]
  9. Zeng, W.; Lin, H.; Li, X. Study on extracting forest information based on SV-1 image. J. Cent. South Univ. For. Technol. 2020, 40, 32–40. [Google Scholar] [CrossRef]
  10. Zhou, K.; Yang, Y.; Zhang, Y.; Miao, R.; Yang, Y.; Liu, L. Review of land use classification methods based on optical remote sensing images. Sci. Technol. Eng. 2021, 21, 13603–13613. [Google Scholar]
  11. Ye, Y.; Lu, D.; Wu, Z.; Liao, K.; Zhou, M.; Jian, K.; Li, D. Vertical characteristics of vegetation distribution in Wuyishan National Park based on multi-source high-resolution remotely sensed data. Remote Sens. 2023, 15, 5023. [Google Scholar] [CrossRef] [Scilit]
  12. Guirado, E.; Blanco-Sacristán, J.; Rigol-Sánchez, J.P.; Alcaraz-Segura, D.; Cabello, J. A multi-temporal object-based image analysis to detect long-lived shrub cover changes in drylands. Remote Sens. 2019, 11, 2649. [Google Scholar] [CrossRef] [Scilit]
  13. White, L.; Brisco, B.; Dabboor, M.; Schmitt, A.; Pratt, A. A collection of SAR methodologies for monitoring wetlands. Remote Sens. 2015, 7, 7615–7645. [Google Scholar] [CrossRef] [Scilit]
  14. Meng, M.M. Research on Land Cover Classification Based on Dualpolarized Sar Images. Master’s Thesis, Henan University, Kaifeng, China, 2024. [Google Scholar]
  15. Zheng, P.; Fang, P.; Wang, L.; Ou, G.; Xu, W.; Dai, F.; Dai, Q. Synergism of multi-modal data for mapping tree species distribution—A case study from a mountainous forest in southwest china. Remote Sens. 2023, 15, 979. [Google Scholar] [CrossRef] [Scilit]
  16. Borlaf-Mena, I.; Badea, O.; Tanase, M.A. Assessing the utility of sentinel-1 coherence time series for temperate and tropical forest mapping. Remote Sens. 2021, 13, 4814. [Google Scholar] [CrossRef] [Scilit]
  17. Yuan, X.; Liang, Y.; Feng, W.; Li, J.; Ren, H.; Han, S.; Liu, M. Classification of coniferous and broad-leaf forests in China based on high-resolution imagery and local samples in Google Earth Engine. Remote Sens. 2023, 15, 5026. [Google Scholar] [CrossRef] [Scilit]
  18. Sun, M.; Yue, C.; Duan, Y.; Luo, H.; Yu, Q.; Luo, G.; Xu, T. Research on forest type classification with feature level fusion by integrating optical data with SAR Data. For. Eng. 2024, 40, 115–126. [Google Scholar]
  19. Li, H.X.; Shi, Y.; Ding, Z.J.; Huang, L.; Dong, J.; Liang, Z.G.; Zhu, X.W.; Ma, Y.T.; Wang, T. Combining Multi-source Remote Sensing Data and Object-oriented Information Extraction for Arid Eetlands. Environ. Sci. 2025, 46, 3127–3138. [Google Scholar] [CrossRef] [Scilit]
  20. Erinjery, J.J.; Singh, M.; Kent, R. Mapping and assessment of vegetation types in the tropical rainforests of the Western Ghats using multispectral Sentinel-2 and SAR Sentinel-1 satellite imagery. Remote Sens. Environ. 2018, 216, 345–354. [Google Scholar] [CrossRef] [Scilit]
  21. Li, L.; Tian, X.; Weng, Y.L. Land cover classification based on polarization SAR and optical image features. J. Southeast Univ. 2021, 51, 529–534. [Google Scholar]
  22. Chen, F.; Fu, Z.; Huang, L.; Niu, B.; Chen, P.; Wang, L. Review of deep learning in optical and SAR image fusion. Natl. Remote Sens. Bull. 2022, 26, 1744–1756. [Google Scholar] [CrossRef] [Scilit]
  23. Zhang, W.; Brandt, M.; Wang, Q.; Prishchepov, A.V.; Tucker, C.J.; Li, Y.; Lyu, H.; Fensholt, R. From woody cover to woody canopies: How Sentinel-1 and Sentinel-2 data advance the mapping of woody plants in savannas. Remote Sens. Environ. 2019, 234, 111465. [Google Scholar] [CrossRef] [Scilit]
  24. Symeonakis, E.; Higginbottom, T.P.; Petroulaki, K.; Rabe, A. Optimisation of savannah land cover characterisation with optical and SAR data. Remote Sens. 2018, 10, 499. [Google Scholar] [CrossRef] [Scilit]
  25. Maskell, G.; Chemura, A.; Nguyen, H.; Gornott, C.; Mondal, P. Integration of Sentinel optical and radar data for mapping smallholder coffee production systems in Vietnam. Remote Sens. Environ. 2021, 266, 112709. [Google Scholar] [CrossRef] [Scilit]
  26. Wang, Y.; Liu, H.; Sang, L.; Wang, J. Characterizing forest cover and landscape pattern using multi-source remote sensing data with ensemble learning. Remote Sens. 2022, 14, 5470. [Google Scholar] [CrossRef] [Scilit]
  27. Liu, S.; Zhang, X.; Li, X.; Tian, Y. Cooperative land use classification of hyperspectral and multispectral imagery based on dual branch convolutional neural network. Trans. Chin. Soc. Agric. Eng. 2020, 36, 252–262. [Google Scholar] [CrossRef]
  28. Ren, B.; Liu, B.; Hou, B.; Wang, Z.; Yang, C.; Jiao, L. SwinTFNet: Dual-stream transformer with cross attention fusion for land cover classification. IEEE Geosci. Remote Sens. Lett. 2024, 21, 2501505. [Google Scholar] [CrossRef] [Scilit]
  29. Wang, B.; Yao, Y. Mountain vegetation classification method based on multi-channel semantic segmentation model. Remote Sens. 2024, 16, 256. [Google Scholar] [CrossRef] [Scilit]
  30. Liu, C.; Sun, Y.; Xu, Y.; Sun, Z.; Zhang, X.; Lei, L.; Kuang, G. A review of optical and SAR image deep feature fusion in semantic segmentation. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 12910–12930. [Google Scholar] [CrossRef] [Scilit]
  31. Li, W.; Wu, J.; Liu, Q.; Zhang, Y.; Cui, B.; Jia, Y.; Gui, G. An effective multimodel fusion method for SAR and optical remote sensing images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2023, 16, 5881–5892. [Google Scholar] [CrossRef] [Scilit]
  32. Li, Y.; Min, S.; Song, B.; Yang, H.; Wang, B.; Wu, Y. Multisource high-resolution remote sensing image vegetation extraction with comprehensive multifeature perception. Remote Sens. 2024, 16, 712. [Google Scholar] [CrossRef] [Scilit]
  33. Peng, L.; Yang, W.; Huang, J. Research on Land Cover Classification Using Multi-Channel Interferometric Radar in the Western Sichuan Plateau. J. Southwest Univ. 2016, 38, 125–132. [Google Scholar] [CrossRef]
  34. Yuan, Y.; Wen, Q.; Zhao, X.; Liu, S.; Zhu, K.; Hu, B. Identifying grassland distribution in a mountainous region in Southwest China using multi-source remote sensing images. Remote Sens. 2022, 14, 1472. [Google Scholar] [CrossRef] [Scilit]
  35. Sommervold, O.; Gazzea, M.; Arghandeh, R. A survey on SAR and optical satellite image registration. Remote Sens. 2023, 15, 850. [Google Scholar] [CrossRef] [Scilit]
  36. Wu, W.; Shao, Z.; Huang, X.; Teng, J.; Guo, S.; Li, D. Quantifying the sensitivity of SAR and optical images three-level fusions in land cover classification to registration errors. Int. J. Appl. Earth Obs. Geoinf. 2022, 112, 102868. [Google Scholar] [CrossRef] [Scilit]
  37. Zhao, Y.M.; Hu, K.L.; Tu, K.L.; Qing, Y.X.; Yang, C.; Qi, K.L.; Wu, H.Y. Multi-label scene classification method based on fusion of SAR and optical remote sensing images. Geod. Cartogr. Sin. 2025, 54, 911–923. [Google Scholar] [CrossRef]
  38. Lin, L. Collaborative Classification of Land Cover Based on Polarimetric Sar and Optical Imagery. Master’s Thesis, Southeast University, Nanjing, China, 2021. [Google Scholar] [CrossRef]
  39. Ma, X.; Zhang, X.; Pun, M.O.; Huang, B. A unified framework with multimodal fine-tuning for remote sensing semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5405015. [Google Scholar] [CrossRef] [Scilit]
  40. Ma, X.; Xu, X.; Zhang, X.; Pun, M.O. Adjacent-scale multimodal fusion networks for semantic segmentation of remote sensing data. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 20116–20128. [Google Scholar] [CrossRef] [Scilit]
  41. Ma, X.; Zhang, X.; Pun, M.O. A crossmodal multiscale fusion network for semantic segmentation of remote sensing data. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 3463–3474. [Google Scholar] [CrossRef] [Scilit]
  42. Ren, B.; Ma, S.; Hou, B.; Hong, D.; Chanussot, J.; Wang, J.; Jiao, L. A dual-stream high resolution network: Deep fusion of GF-2 and GF-3 data for land cover classification. Int. J. Appl. Earth Obs. Geoinf. 2022, 112, 102896. [Google Scholar] [CrossRef] [Scilit]
  43. Wei, K.; Dai, J.; Hong, D.; Ye, Y. MGFNet: An MLP-dominated gated fusion network for semantic segmentation of high-resolution multi-modal remote sensing images. Int. J. Appl. Earth Obs. Geoinf. 2024, 135, 104241. [Google Scholar] [CrossRef] [Scilit]
  44. Ma, X.; Zhang, X.; Pun, M.O.; Liu, M. A multilevel multimodal fusion transformer for remote sensing semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5403215. [Google Scholar] [CrossRef] [Scilit]
  45. Schmitt, M.; Hughes, L.H.; Qiu, C.; Zhu, X.X. SEN12MS–A curated dataset of georeferenced multi-spectral sentinel-1/2 imagery for deep learning and data fusion. arXiv 2019, arXiv:1906.07789. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Location of the study area. (a) Administrative division. (b) DEM distribution. (c) Optical image map showing the distribution of typical forest, shrubland and grassland.
Figure 1. Location of the study area. (a) Administrative division. (b) DEM distribution. (c) Optical image map showing the distribution of typical forest, shrubland and grassland.
Remotesensing 18 01373 g001
Figure 2. Data Preprocessing Flowchart.
Figure 2. Data Preprocessing Flowchart.
Remotesensing 18 01373 g002
Figure 3. Distribution of field sample points. (a) DEM of the study area. (b,c) Distribution of sample points. (d) Field photographs.
Figure 3. Distribution of field sample points. (a) DEM of the study area. (b,c) Distribution of sample points. (d) Field photographs.
Remotesensing 18 01373 g003
Figure 4. Spectral and scatter characteristic distribution curves.
Figure 4. Spectral and scatter characteristic distribution curves.
Remotesensing 18 01373 g004
Figure 5. Texture feature distribution curve. (a) Optical texture feature. (b) SAR texture feature.
Figure 5. Texture feature distribution curve. (a) Optical texture feature. (b) SAR texture feature.
Remotesensing 18 01373 g005
Figure 6. Random Forest Feature Importance Evaluation Results.
Figure 6. Random Forest Feature Importance Evaluation Results.
Remotesensing 18 01373 g006
Figure 7. Some images of GF23FSG dataset.
Figure 7. Some images of GF23FSG dataset.
Remotesensing 18 01373 g007
Figure 8. Structure of the CASFNet.
Figure 8. Structure of the CASFNet.
Remotesensing 18 01373 g008
Figure 9. Structure of the Optical Encoder.
Figure 9. Structure of the Optical Encoder.
Remotesensing 18 01373 g009
Figure 10. Structure of the SAR Encoder.
Figure 10. Structure of the SAR Encoder.
Remotesensing 18 01373 g010
Figure 11. Structure of the CASF-Module.
Figure 11. Structure of the CASF-Module.
Remotesensing 18 01373 g011
Figure 12. Structure of the Decoder.
Figure 12. Structure of the Decoder.
Remotesensing 18 01373 g012
Figure 13. Structure of the MASLoss.
Figure 13. Structure of the MASLoss.
Remotesensing 18 01373 g013
Figure 14. Normalized confusion matrix.
Figure 14. Normalized confusion matrix.
Remotesensing 18 01373 g014
Figure 15. Visualization results of CASFNet. (a1,a2) Optical images. (b1,b2) SAR images. (c1,c2) Labels. (d1,d2) Inference Results of CASFNet.
Figure 15. Visualization results of CASFNet. (a1,a2) Optical images. (b1,b2) SAR images. (c1,c2) Labels. (d1,d2) Inference Results of CASFNet.
Remotesensing 18 01373 g015
Figure 16. Visualization results of comparison experiment on the GF23FSG dataset.
Figure 16. Visualization results of comparison experiment on the GF23FSG dataset.
Remotesensing 18 01373 g016
Figure 17. Visualization results of data source validity.
Figure 17. Visualization results of data source validity.
Remotesensing 18 01373 g017
Figure 18. Some images of SEN12MS dataset. (a) Sentinel-2 image. (b) Sentinel-1 image. (c) Original label. (d) Re-mapped label.
Figure 18. Some images of SEN12MS dataset. (a) Sentinel-2 image. (b) Sentinel-1 image. (c) Original label. (d) Re-mapped label.
Remotesensing 18 01373 g018
Figure 19. Visualization results of comparison experiment on the SEN12MS dataset.
Figure 19. Visualization results of comparison experiment on the SEN12MS dataset.
Remotesensing 18 01373 g019
Figure 20. The sample quantity distribution of the GF23FSG dataset.
Figure 20. The sample quantity distribution of the GF23FSG dataset.
Remotesensing 18 01373 g020
Table 1. Satellite payload information.
Table 1. Satellite payload information.
SatelliteSensor/Imaging ModeBandSpectral Range/Polarization ModeSpatial Resolution
GF-2PMSPanchromatic0.45∼0.90 μm0.8 m
Blue0.45∼0.52 μm3.24 m
Green0.52∼0.59 μm
Red0.63∼0.69 μm
Near-infrared0.77∼0.89 μm
GF-3UFSC-bandDH3 m
Table 2. Forest–Shrub–Grassland taxonomy.
Table 2. Forest–Shrub–Grassland taxonomy.
ClassDefinitionExample
Evergreen ForestThe main plants distributed in the land are tall trees, usually over 3 m in height. They have an independent main trunk growing from the root, with a clear distinction between the trunk and the canopy. The leaves remain green throughout the year.Remotesensing 18 01373 i001
Deciduous ForestThe main plants in the land shed all their leaves during the autumn and winter seasons or during periods of drought. From the roots, an independent main trunk emerges, and the trunk and the tree crown are clearly distinguishable. The height of the tree is generally over 3 m.Remotesensing 18 01373 i002
Evergreen ShrublandThe land is dominated by woody plants that are no more than 3 m tall, have evergreen leaves throughout the year, and have no distinct main trunk.Remotesensing 18 01373 i003
Deciduous ShrublandLand consisting of woody plants that are no more than 3 m tall, shed leaves in autumn and winter, have no distinct main trunk, and grow in a clustered manner.Remotesensing 18 01373 i004
Low-cover GrasslandThe land where the coverage of herbaceous plants is no more than 60% consists of mostly perennial or annual grasses, sedges, and other low-growing plants of the grass family and sedge family.Remotesensing 18 01373 i005
High-cover GrasslandLand with herbaceous plant coverage exceeding 60% is typically characterized by growth under moderate humidity conditions.Remotesensing 18 01373 i006
OtherLand types other than forest, shrubland and grassland.Remotesensing 18 01373 i007
Table 3. Introduction of Selected Features.
Table 3. Introduction of Selected Features.
Feature NameDescriptionFeature TypeFormula
GF2_B1Quantile-normalized reflectance of the GF-2 blue bandSpectral Φ 1 F X ( x i )
GF2_B2Quantile-normalized reflectance of the GF-2 green bandSpectral Φ 1 F X ( x i )
GF2_B3Quantile-normalized reflectance of the GF-2 red bandSpectral Φ 1 F X ( x i )
GF2_B4Quantile-normalized reflectance of the GF-2 near-infrared bandSpectral Φ 1 F X ( x i )
NDVIQuantile-normalized NDVIVegetation Index Φ 1 F X N i r R e d N i r + R e d
BCQuantile-normalized backscatter coefficient of GF-3 DH polarizationScattering Φ 1 F X ( x i )
BC_EntropyQuantile-normalized GLCM entropy feature derived from the GF-3 backscatter coefficientTexture Φ 1 F X i , j = 1 W P ( i , j ) log P ( i , j )
Table 4. Dataset Class Distribution.
Table 4. Dataset Class Distribution.
ClassTraining Set Proportion (%)Validation Set Proportion (%)Total Proportion (%)
Evergreen Forest19.7620.1719.84
Deciduous Forest2.913.232.97
Evergreen Shrubland12.7112.8512.74
Deciduous Shrubland10.1110.1910.13
Low-cover Grassland17.8819.1318.13
High-cover Grassland13.2312.3313.05
Other23.4022.1023.14
Table 5. Results of the quantitative analysis of CASFNet.
Table 5. Results of the quantitative analysis of CASFNet.
ClassF1 (%)IoU (%)
Other82.8470.18
Evergreen Forest82.6870.47
Deciduous Forest62.4445.40
Evergreen Shrubland72.7457.16
Deciduous Shrubland70.9654.99
Low-cover Grassland73.0457.54
High-cover Grassland80.7667.73
OA77.38
mIoU60.50
mF175.02
Kappa72.84
Table 6. Quantitative results of comparison experiment on the GF23FSG dataset.
Table 6. Quantitative results of comparison experiment on the GF23FSG dataset.
ClassMFNet
(%)
ASMFNet
(%)
CMFNet
(%)
DDHRNet
(%)
MGFNet
(%)
FTransUNet
(%)
Ours
(%)
Other52.3059.0058.7961.8263.0463.6270.18
Evergreen Forest59.3260.5068.1964.2767.3068.2270.47
Deciduous Forest34.6929.5239.3241.3646.6650.0345.40
Evergreen Shrubland43.9844.6751.0251.4453.2253.8657.17
Deciduous Shrubland42.3341.5650.2248.4951.2553.0554.99
Low-cover Grassland39.2848.7251.0145.4448.0353.1757.54
High-cover Grassland49.0752.1751.4850.8354.9758.3167.73
OA65.1267.8971.3870.3372.4374.0177.38
mIoU45.8548.0252.8651.9554.9257.1860.50
mF161.7864.2568.7968.0470.6472.5775.02
Kappa57.9761.4665.4564.1966.7468.8372.84
The best results are highlighted in bold, and the second-best results are underlined.
Table 7. Quantitative results of data source validity.
Table 7. Quantitative results of data source validity.
DatasetIoU of Each Class (%)
Optical InputOptical + SAR
(Without Feature)
Optical + SAR
(with Optical Feature)
Optical + SAR
(with Feature)
Other62.3867.1568.7670.18
Evergreen Forest66.6566.7168.8170.47
Deciduous Forest36.6640.8942.0245.4
Evergreen Shrubland46.1352.0854.4357.17
Deciduous Shrubland45.6450.4152.5554.99
Low-cover Grassland50.6954.8255.7357.54
High-cover Grassland57.0660.5765.4867.73
OA70.8673.4376.3177.38
mIoU52.1756.0958.460.5
mF168.0371.6173.2675.02
Kappa65.0368.2471.472.84
Table 8. Quantitative analysis of the module ablation experiment.
Table 8. Quantitative analysis of the module ablation experiment.
HRNetNHRNetResnet101CASF-ModuleMASLossOA (%)mIoU (%)mF1 (%)Kappa (%)
72.5254.6770.3866.92
72.8855.5571.1567.57
74.6157.9973.1269.36
77.3860.5075.0272.84
Table 9. Mapping of the SEN12MS dataset.
Table 9. Mapping of the SEN12MS dataset.
SEN12MS LabelForest-Shrub-Herb Taxonomy
Evergreen Needleleaf ForestsEvergreen Forest
Evergreen Broadleaf ForestsEvergreen Forest
Deciduous Needleleaf ForestsDeciduous Forest
Deciduous Broadleaf ForestsDeciduous Forest
Mixed ForestsMixed Forest
Closed ShrublandsShrubland
Open ShrublandsShrubland
Woody SavannasSavannas
SavannasSavannas
GrasslandsGrassland
Permanent WetlandsOther
Croplands
Urban&Built-up
Cropland/Natural Vegetation Mosaics
Permanent Snow and Ice
Barren
Water Bodies
Table 10. Quantitative results of comparison experiment on the SEN12MS dataset.
Table 10. Quantitative results of comparison experiment on the SEN12MS dataset.
ClassMFNet
(%)
ASMFNet
(%)
CMFNet
(%)
DDHRNet
(%)
MGFNet
(%)
FTransUNet
(%)
Ours
(%)
Other55.8559.3552.0559.9157.3058.2163.75
Evergreen Forest58.3256.3232.0757.3057.2065.1359.24
Deciduous Forest34.2427.2027.4331.9339.1141.9848.26
Shrubland33.1748.8936.8144.0244.5045.2852.42
Savannas36.0442.3037.7545.4841.8144.3653.19
Grassland31.4532.6231.6333.0833.2038.2551.47
OA60.8764.5257.5865.2163.3965.5471.84
mIoU41.5146.1136.2945.2945.5248.8755.51
mF157.8462.0153.1761.6062.0465.1270.51
Kappa46.6952.0242.6053.0651.1854.8866.42
Table 11. Overview of Classification Accuracy of the “Coarse” Categories.
Table 11. Overview of Classification Accuracy of the “Coarse” Categories.
DatasetIoU (%)OA (%)mIoU (%)mF1 (%)Kappa (%)
OtherForestShrublandGrassland
GF23FSG69.6377.1165.8670.1482.7470.6982.7676.86
SEN12MS62.7672.0786.0266.2881.5671.7883.2773.44
Table 12. This paper utilizes detailed information from publicly available land cover products.
Table 12. This paper utilizes detailed information from publicly available land cover products.
Product NameSpatial ResolutionData SourceRelated Classes of Forest, Shrubland, Grassland
Esri Global Land Cover10 mSentinel-2Forest, Shrubland/grassland
ESA WorldCover 202110 mSentinel-1, Sentinel-2Forest, Shrubland, Grassland
GLC_FCS1010 mSentinel-1, Sentinel-2Evergreen Broadleaved Forest,
Deciduous Broadleaved Forest,
Evergreen Needleleaved Forest,
Deciduous Needleleaved Forest,
Mixed-leaf Forest, Evergreen Shrubland,
Deciduous Shrubland, Grassland.
Table 13. Public verification results of land cover products.
Table 13. Public verification results of land cover products.
ClassField Survey
Sample Points
The Number of Sample Points That Have Been Correctly Classified
Esri Global
Land Cover
ESA WorldCoverGLC_FCS10Ours
ForestEvergreen Forest7564543872
Deciduous Forest442534
ShrublandEvergreen Shrubland51 03347
Deciduous Shrubland741013059
Grassland66 393363
OA (%)53.2330.0052.2688.71
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Pang, Q.; Yuan, Z.; Mi, X.; Yang, J.; Du, W.; Zhang, J.; Zhang, J.; Du, K.; Guo, Z. Vegetation Mapping in Heterogeneous Forest–Shrub–Grass Ecosystems Using Fused High-Resolution Optical and SAR Data. Remote Sens. 2026, 18, 1373. https://doi.org/10.3390/rs18091373

AMA Style

Pang Q, Yuan Z, Mi X, Yang J, Du W, Zhang J, Zhang J, Du K, Guo Z. Vegetation Mapping in Heterogeneous Forest–Shrub–Grass Ecosystems Using Fused High-Resolution Optical and SAR Data. Remote Sensing. 2026; 18(9):1373. https://doi.org/10.3390/rs18091373

Chicago/Turabian Style

Pang, Qingshuang, Zhanliang Yuan, Xiaofei Mi, Jian Yang, Weibing Du, Jian Zhang, Jilong Zhang, Kang Du, and Zheng Guo. 2026. "Vegetation Mapping in Heterogeneous Forest–Shrub–Grass Ecosystems Using Fused High-Resolution Optical and SAR Data" Remote Sensing 18, no. 9: 1373. https://doi.org/10.3390/rs18091373

APA Style

Pang, Q., Yuan, Z., Mi, X., Yang, J., Du, W., Zhang, J., Zhang, J., Du, K., & Guo, Z. (2026). Vegetation Mapping in Heterogeneous Forest–Shrub–Grass Ecosystems Using Fused High-Resolution Optical and SAR Data. Remote Sensing, 18(9), 1373. https://doi.org/10.3390/rs18091373

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop