Next Article in Journal
GeoAI-Enabled Ensemble Modeling to Assess Land Use and Atmospheric Pollutant Impacts on Land Surface Temperature in the US Southwest
Next Article in Special Issue
Deep Learning-Based Extraction of Urban Blue–Green Spaces and Identification of Influencing Factors of Ecosystem Services: A Case Study of Guilin, China
Previous Article in Journal
A Deep Learning-Driven Spatio-Temporal Framework for Timely Corn Yield Estimation Across Multiple Remote Sensing Scenarios
Previous Article in Special Issue
Quantifying Elevation Changes Under Engineering Measures Using Multisource Remote Sensing and Interpretable Machine Learning: A Case Study of the Chinese Loess Plateau
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Semantic Segmentation of Multispectral Remote Sensing Imagery for Coastal Wetlands with SegFormer

School of Information Science and Technology, Beijing Forestry University, Beijing 100083, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(5), 745; https://doi.org/10.3390/rs18050745
Submission received: 13 January 2026 / Revised: 11 February 2026 / Accepted: 19 February 2026 / Published: 28 February 2026

Highlights

What are the main findings?
  • We developed a deep learning model that enhances spectral representation and boundary delineation for coastal wetland segmentation.
What are the implications of the main findings?
  • The proposed method outperforms SegFormer in complex coastal wetland environments characterized by spectral complexity and class imbalance.
  • It offers a robust solution for the semantic segmentation of multispectral remote sensing data in complex coastal wetland environments.

Abstract

Pixel-level semantic segmentation plays an essential role in coastal wetland monitoring using multispectral remote sensing imagery. However, accurate mapping remains challenging due to spectral confusion among heterogeneous land-cover types, fragmented spatial structures, and pronounced class imbalance. Based on the situation, we used the original SegFormer as the basic framework and developed an improved framework to better suit the characteristics of coastal wetland scenes. Prior to the encoder, we introduced a Spectral-Aware Embedding (SAE) module to strengthen inter-band feature representation through spectral projection and adaptive channel weighting. In the decoder, we constructed a Wetland Boundary-Refined Decoder (WBRD), utilizing a dual-path refinement strategy to capture fine-scale textures and a multi-scale boundary attention mechanism to enhance the delineation of irregular boundaries. Additionally, we incorporated a Wetland Imbalance Loss (WIL) during training to moderate the influence of dominant classes. In this article, we evaluated our framework on the Yan14 dataset. The results showcased the framework’s effectiveness, improving segmentation accuracy and boundary fidelity, particularly for rare and narrow wetland categories, while maintaining reasonable computational efficiency.

1. Introduction

Wetlands play a critical role in biodiversity conservation and ecological regulation. Coastal wetlands, located at the land–sea interface, are characterized by strong spatial heterogeneity and dynamic environmental processes, making them highly sensitive to both natural disturbances and human activities [1,2]. In recent decades, land reclamation, aquaculture expansion, and biological invasions have substantially altered coastal wetland landscapes, leading to increased fragmentation and rapid land-cover changes [3,4]. Against this background, remote sensing has become an indispensable tool for coastal wetland monitoring due to its synoptic coverage and long-term observational capability [5,6,7]. With the growing availability of high-resolution multispectral imagery, pixel-level semantic segmentation has attracted growing attention, as it enables fine-grained mapping of wetland types and spatial structures beyond traditional image classification approaches [8,9].
Recent advances in deep learning, particularly Transformer-based semantic segmentation models, have significantly improved the modeling of long-range contextual information in remote sensing imagery [10]. Benefiting from global self-attention mechanisms, Transformers have demonstrated strong potential in complex land-cover mapping tasks and have been increasingly adopted in multispectral remote sensing segmentation. Nevertheless, when applied to coastal wetlands, existing Transformer-based models still encounter persistent challenges related to spectral discrimination, spatial boundary preservation, and severe class imbalance. Although recent studies have extended Transformer-based segmentation frameworks to wetland and coastal applications by incorporating spectral or temporal information, they consistently report that high spectral similarity among wetland types and fragmented boundary structures remain difficult to resolve in complex wetland environments [11,12,13].
One major limitation arises from the insufficient exploitation of multispectral information. In most Transformer-based segmentation frameworks, multispectral bands are treated as simple channel concatenations, without explicitly modeling inter-band dependencies. While several studies have attempted to enhance spectral utilization through band grouping strategies [14], auxiliary spectral indices [15,16], or attention-based weighting mechanisms, these enhancements are often introduced in a shallow or task-specific manner [17]. As a result, spectral dependencies are not systematically encoded in the learned feature representations. This limitation is particularly pronounced in coastal wetlands, where different land-cover types often exhibit high spectral similarity but differ subtly at the band level.
In addition to spectral complexity, Beyond spectral complexity, coastal wetlands are characterized by distinctive spatial morphologies, including elongated channels, fragmented patches, and smooth transitions between adjacent land-cover types. Existing remote sensing segmentation studies have explored boundary enhancement through multi-scale feature fusion [18], edge supervision, or attention-guided refinement strategies [19,20,21]. While these methods have shown improvements in contour preservation for certain object categories, they are primarily designed for generic scenes or specific targets such as buildings and open water bodies, and do not fully account for the spatial characteristics of coastal wetland landscapes [22]. Consequently, conventional decoders tend to produce over-smoothed predictions during multi-scale feature aggregation, leading to blurred boundaries and the loss of fine structural details that are critical for wetland pattern analysis.
Another persistent challenge in coastal wetland segmentation arises from the highly imbalanced distribution of land-cover classes [23,24,25]. Dominant categories such as open water or aquaculture ponds often occupy extensive areas, whereas ecologically important wetland types are sparsely distributed and spatially discontinuous. When trained with conventional loss functions, segmentation models tend to bias toward majority classes, suppressing the learning of minority and small-scale targets [16,26,27]. This imbalance-driven effect reduces segmentation reliability for rare wetland types and remains a non-negligible limitation in practical wetland mapping applications.
Overall, despite notable progress in multispectral wetland segmentation, the challenges discussed above have not yet been effectively resolved in a unified manner. Many existing approaches focus on representation, boundary refinement, or loss reweighting separately, and their improvements are typically introduced as independent or auxiliary components. In complex coastal wetland environments, however, these factors are inherently interrelated. Subtle spectral differences may be weakened during early feature encoding, fragmented and elongated structures are prone to boundary smoothing during decoding, and severe class imbalance further limits the learning of ecologically important but sparsely distributed wetland types. As a result, segmentation errors often accumulate around wetland boundaries, narrow landforms, and minority classes.
From this perspective, adopting a Transformer architecture such as SegFormer [28,29,30], while beneficial for modeling global context, is not sufficient on its own for reliable coastal wetland segmentation. Accurate mapping in such environments requires explicit consideration of spectral discrimination, sensitivity to fine spatial boundaries, and balanced optimization during training. Based on these observations, this study develops a segmentation framework built on SegFormer, with coordinated adaptations introduced at different stages of the segmentation process to better reflect the characteristics of multispectral coastal wetlands. The main methodological contributions are summarized as follows:
(1)
To address the insufficient utilization of multispectral band information, a Spectral-Aware Embedding (SAE) module is introduced at the encoding stage. This module performs unified modeling and adaptive reweighting of multispectral features, enhancing the effectiveness of spectral information in feature representation;
(2)
Considering the fragmented boundaries and complex spatial structures of coastal wetland land-cover types, a Wetland Boundary-Refined Decoder (WBRD) is designed. By incorporating multi-scale feature refinement and boundary-guided enhancement mechanisms, the proposed decoder improves segmentation accuracy for elongated objects and transitional regions;
(3)
To mitigate the impact of class imbalance during model training, a Wetland Imbalance Loss (WIL) is constructed. By constraining the gradient contributions of dominant classes, this loss function enhances the learning performance for minority and easily confused land-cover types.
Together, these components form a unified methodological framework tailored to coastal wetland environments. By jointly addressing spectral representation, boundary preservation, and class imbalance, the proposed approach provides a robust and generalizable solution for multispectral semantic segmentation in complex coastal wetlands, supporting more accurate wetland mapping and analysis in typical coastal regions.

2. Study Area and Data Preprocessing

2.1. Study Area Overview

The Yancheng coastal wetland is located along the western shore of the Yellow Sea in the central coastal region of Jiangsu Province, northern China, north of the Yangtze River Delta. It represents one of the largest and most well-preserved coastal wetland systems in China, covering an area of approximately 450,000 hectares. The study area encompasses the core zones of the Yancheng National Nature Reserve for Rare Birds and the Jiangsu Yellow Sea Wetlands World Natural Heritage Site, covering representative wetland regions in Sheyang, Binhai, Xiangshui, and Dafeng [31]. The location of the study area is shown in Figure 1.
The region contains a wide range of wetland types and extensive wetland resources. Its surface landscape is jointly shaped by natural processes, including tidal action, river–sea interactions, storm surges, and estuarine sediment deposition, as well as by intensive human activities such as land reclamation, salt production, aquaculture pond construction, and road excavation. As a result, the Yancheng coastal wetland exhibits pronounced spatial heterogeneity and severe landscape fragmentation, posing substantial challenges for accurate land-cover mapping and analysis [32,33].

2.2. Data Sources and Remote Sensing Image Preprocessing

Sentinel-2 multispectral imagery was used in this study. Level-2A products were obtained from the European Space Agency (ESA), with cloud coverage below 5% and acquisition dates rang from January 2023 to December 2024. Representative images from four seasons were selected to account for seasonal variability in coastal wetland landscapes. A total of twelve spectral bands were used, covering the visible, near-infrared, and shortwave infrared regions, while the cirrus band (B10) was excluded due to its limited relevance for land surface analysis.
All images were resampled to a uniform spatial resolution of 10 m using the SNAP 12.0.0, followed by band composition, preprocessing, and spatial clipping in ENVI 5.6 based on the boundary of the protected area. The processed Sentinel-2 multispectral data provide consistent spectral information across visible, near-infrared, and shortwave infrared wavelengths, supporting large-scale discrimination of major coastal wetland land-cover types. Details of the selected Sentinel-2 imagery are summarized in Table 1.

2.3. Wetland Classification

Based on the national wetland resource survey and classification system, existing coastal wetland vegetation classification schemes in China, and the ecological characteristics of the Yancheng coastal wetlands, a hierarchical land-cover classification system suitable for remote sensing semantic segmentation was established in this study [31,33,34,35].
The classification system follows a two-level structure. At the first level, land-cover types are categorized into natural wetlands, artificial wetlands, and non-wetlands. At the second level, more detailed land-cover classes are defined, as summarized in Table 2.
The remote sensing images were annotated using a manual interactive interpretation approach. High-resolution Google Earth panchromatic imagery was used as auxiliary reference data, and multi-temporal Sentinel-2 imagery was compared to enhance interpretation reliability. Spectral indices , including NDVI [36], NDWI [37], and MNDWI [38] were also used to support visual enhancement and class discrimination. Image annotation was conducted using eCognition 10.3, and the annotation results are shown in Figure 2.

2.4. Dataset Construction and Annotation

The Yan14 dataset consists of 5853 samples, each with a spatial size of 512 × 512 pixels. To mitigate spatial autocorrelation and potential data leakage, a spatial block-based dataset partitioning strategy was adopted. Specifically, two geographically independent regions were designated as test areas, and all patches within these regions were assigned to the test set (402 samples). The remaining patches, spatially separated from the test regions, were further divided into training and validation sets, yielding 5049 training samples and 402 validation samples. This partitioning strategy ensures geographic separation between training and test data, thereby enabling a more reliable evaluation of the model’s generalization performance. Detailed statistics of sample numbers and pixel distributions for each land-cover class are reported in Table 3 and Table 4.
The Yan14 dataset exhibits substantial diversity in wetland types as well as pronounced class imbalance. Categories such as suspended sediment, grass flat, reed marshes, sea water, and aquaculture ponds occupy relatively large proportions and constitute the dominant land-cover types in the study area. In contrast, salt pans, Suaeda salsa marshes, and small water bodies contain limited numbers of pixels and are considered rare classes. Some narrow or locally distributed categories, such as drainage channels and Spartina alterniflora marshes, exhibit small spatial scales and complex shapes, further increasing segmentation difficulty.
Overall, the Yan14 dataset covers a wide range of typical landscape types in the Yancheng coastal wetland , while also exhibiting significant class imbalance and spatial fragmentation. The large disparity in pixel counts between dominant and rare classes—some accounting for less than 3% of the total pixels—poses high demands on a model’s capability for small-sample learning, boundary representation, and spectral feature modeling. These data characteristics directly motivate the subsequent design of the segmentation model.

3. Methods

3.1. Overall Method Framework

Coastal wetlands present a unique challenge for semantic segmentation due to spectral similarity among vegetation types, fragmented spatial structures, and blurred land–water boundaries. Among recent Transformer-based segmentation models, SegFormer offers a suitable baseline for this task, as its hierarchical Transformer encoder captures multi-scale contextual information with relatively high computational efficiency [28]. While SegFormer provides a strong baseline with its hierarchical Transformer encoder, it treats multispectral bands as simple channel stacks, lacking explicit modeling of spectral relationships. Additionally, its lightweight decoder tends to oversmooth edges, limiting its ability to delineate narrow tidal channels and irregular wetland boundaries. To address these limitations, three tailored enhancements are integrated into the SegFormer pipeline:
Spectral-Aware Embedding (SAE) is inserted before the encoder to explicitly model inter-band dependencies and highlight discriminative spectral regions.
Wetland Boundary-Refined Decoder (WBRD) replaces the standard decoder, enabling improved recovery of fine spatial structures and sharper transitions between adjacent wetland types.
Wetland Imbalance Loss (WIL) is a combined loss function that counters gradient dominance from large-area classes and improves learning on rare categories.
Figure 3 illustrates the overall network architecture. The model takes 12-band Sentinel-2 imagery as input. The SAE module first remaps the spectral features, which are then passed through the SegFormer encoder to extract multi-scale contextual features. The WBRD decoder subsequently merges these features with explicit boundary guidance, and a classification head outputs the final segmentation map. During training, WIL adjusts learning priorities to mitigate class imbalance.

3.2. Spectral-Aware Embedding Module (SAE)

Multispectral remote sensing imagery contains rich spectral information, with different land-cover types exhibiting distinct response characteristics across visible, near-infrared, and shortwave infrared bands. In the SegFormer framework, multispectral inputs are handled through channel stacking, without explicit modeling of inter-band spectral dependencies at the early encoding stage.
Channel attention mechanisms, particularly the Squeeze-and-Excitation (SE) block [39], have been widely adopted in convolutional neural networks to model inter-channel dependencies and enhance feature representations. Recent studies have also explored integrating SE-style attention into Transformer-based architectures, demonstrating its general applicability beyond purely convolutional backbones [40,41]. However, these approaches generally treat channels as abstract feature dimensions and are not specifically designed to reflect the characteristics of multispectral remote sensing data.
In this work, a Spectral-Aware Embedding (SAE) module is introduced for multispectral inputs within the Transformer-based SegFormer framework. Unlike classical SE modules that operate on convolutional feature maps, SAE is positioned at the embedding stage and enhances spectral representations prior to self-attention modeling. By explicitly considering band normalization, spectral projection, and subsequent spectral–spatial feature mixing, SAE improves the modeling of inter-band dependencies commonly encountered in multispectral wetland segmentation. The structure of the SAE module is illustrated in Figure 4.
Let the input multispectral image be denoted as X R B × 12 × H × W , where B is the batch size and the 12 channels correspond to the input spectral bands. To mitigate variations in dynamic ranges across spectral bands, per-band Batch Normalization (BN) is first applied:
X 1 = BN ( X )
A 1 × 1 convolution is then used to project the normalized spectral bands into a higher-dimensional feature space aligned with the embedding dimension C 1 of the SegFormer encoder:
F p = Conv 1 × 1 ( X 1 ) R B × C 1 × H × W
This projection allows the network to learn linear combinations of the original spectral bands, forming a compact and expressive spectral feature representation.
Based on the projected features, the SAE module applies a spectral attention mechanism to adaptively emphasize informative spectral responses. Global Average Pooling (GAP) is employed to aggregate channel-wise statistics, followed by a lightweight bottleneck mapping to generate spectral weighting coefficients:
a = σ W 2 δ W 1 GAP ( F p ) R B × C 1
F a = F p a
where W 1 and W 2 are learnable parameters with a reduction ratio r, δ ( · ) denotes the ReLU activation function, σ ( · ) denotes the Sigmoid function, and ⊙ represents element-wise multiplication.
While global pooling provides a compact spectral summary, it suppresses local spatial context that is important for delineating wetland boundaries. To reintroduce local spatial information and promote spectral–spatial interaction, a lightweight mixing unit is further applied. This unit consists of a pointwise convolution followed by a 3 × 3 depthwise separable convolution. A residual connection with a learnable scaling factor α is adopted to stabilize feature modulation:
X m i x = DWConv 3 × 3 Conv 1 × 1 ( F a ) + α F p
The parameter α is initialized to 0.1 and optimized during training, allowing the network to adaptively balance the original spectral features and the modulated representations.
Finally, Layer Normalization is applied to the mixed features:
X o u t = LayerNorm ( X m i x )
Overall, the SAE module provides a spectrally enhanced and spatially informed embedding for multispectral remote sensing imagery, improving feature discriminability with minimal additional computational cost.

3.3. Wetland Boundary-Refined Decoder (WBRD)

Although Transformer-based encoders are effective at modeling long-range context, their decoders tend to rely on progressively upsampled semantic features, which are often dominated by low-frequency responses. As a result, fine-grained boundary cues—especially for narrow and elongated wetland structures—are easily attenuated during decoding.
 In coastal wetlands, many land-cover boundaries (e.g., tidal channels, aquaculture pond edges, and vegetation–mudflat transitions) are primarily characterized by local intensity gradients rather than strong semantic contrast. Such gradients are spatially stable but may not be consistently recovered from deep semantic features.
The proposed WBRD is built upon two boundary-oriented refinement modules:
(1)
a Dual-path Refine (DPR) module for joint texture–boundary feature enhancement.
(2)
a Multi-scale Boundary Attention (MBA) module for boundary-aware feature modulation.
These modules are embedded within a progressive upsampling decoding scheme that gradually restores spatial resolution while integrating encoder skip connections. Unlike conventional decoders that perform uniform feature fusion, WBRD emphasizes spatial regions with high structural uncertainty by integrating an explicit boundary prior into both refinement and attention mechanisms. The overall architecture of WBRD is illustrated in Figure 5.
Mathematically, let the encoder outputs be F = { F 1 , F 2 , F 3 , F 4 } . where F 1 R B × 64 × H 4 × W 4 , F 2 R B × 128 × H 8 × W 8 , F 3 R B × 320 × H 16 × W 16 , F 4 R B × 512 × H 32 × W 32 , corresponding to progressively coarser spatial resolutions.

3.3.1. Dual-Path Refine Module (DPR)

The DPR module aims to enhance boundary-sensitive representations by integrating complementary information from learnable texture features and explicit gradient cues. To this end, DPR adopts a dual-path design.
 The first path applies depthwise separable convolution to capture local texture patterns in a parameter-efficient manner. This branch focuses on refining spatial details that can be learned from data, such as local variations in vegetation or sediment textures.
 The second path introduces a Sobel-based gradient operator to generate a coarse boundary prior. Rather than serving as a standalone edge detector, the Sobel operator provides a deterministic approximation of local intensity gradients, which are closely related to physical land-cover transitions in coastal wetlands. Importantly, this gradient prior is non-trainable and does not participate in backpropagation, thereby acting as a stable geometric reference rather than a dominant feature extractor.
Formally, the DPR module is defined as:
F d p r = Conv 1 × 1 ( Concat [ DWConv ( F ) , Conv 3 × 3 ( F ) ] )
where D W C o n v ( · ) denotes depthwise convolution and C o n c a t [ · ] represents channel-wise concatenation, and ∇ denotes Sobel gradient operator. The structure of the DPR module is illustrated in Figure 6.

3.3.2. Multi-Scale Boundary Attention Module (MBA)

The MBA computes a boundary attention map by aggregating multi-scale convolutional responses with different receptive fields ( 3 × 3 , 5 × 5 , 7 × 7 ), together with the Sobel-derived gradient prior. This design allows the decoder to capture both local and contextual boundary patterns while maintaining sensitivity to fine-scale structures. The following is the formula for the MBA:
A = σ C o n v 1 × 1 ( C o n c a t [ M 3 , M 5 , M 7 , F ] )
F m b a = F A
where ∇ denotes the Sobel gradient operator, ⊙ denotes element-wise multiplication, and M i = G e L U B N C o n v i × i ( F ) , i = 3 , 5 , 7 .
The output is a feature set where boundary pixels are amplified, making them more prominent during final prediction.
The full upsampling transformation at step i is therefore:
T ( i ) ( x , f ) = Conv 3 × 3 ( i ) Concat Upsample 2 × MBA ( i ) ( DPR ( i ) ( x ) ) , f
where x is the upsampled feature from the previous stage and f is the corresponding encoder feature for skip connection. The structure of the MBA module is shown on Figure 7.
 Overall, WBRD enhances boundary delineation by explicitly integrating a geometry-aware boundary prior with learnable texture features and attention-based refinement. The Sobel operator is not used to directly predict boundaries, but to provide weak, stable guidance during decoding. All segmentation decisions are ultimately learned in an end-to-end manner through convolutional feature fusion and attention modulation. In essence, WBRD does not just merge features; it actively seeks out and reinforces edges and fine structures that are ecologically meaningful in wetland landscapes, leading to segmentation maps with crisper boundaries and better spatial coherence.

3.4. Evaluation Metrics

This study employs commonly used metrics in semantic segmentation tasks: Intersection over Union ( I o U ), Mean Intersection over Union ( m I o U ), Macro-average precision ( m P r e c i s i o n ), Macro-average recall ( m R e c a l l ), Pixel Accuracy ( P A ), and Macro-average F1-score ( m F 1 ). The formulas for these five evaluation metrics are as follows, where T P , F P , and F N represent True Positive, False Positive, and False Negative, respectively; N is the number of classes; T P i denotes the true positives for class F P i denotes the false positives for class i, and F N i denotes the false negatives for class i.
IoU is the ratio of the intersection to the union of the predicted and ground truth regions, measuring the degree of segmentation match.
I o U i = T P i T P i + F P i + F N i
m I o U = 1 N i = 1 N T P i T P i + F P i + F N i
m R e c a l l represents the average of per-class recall rates, indicating the model’s balanced ability to correctly identify pixels belonging to each class.
m R e c a l l = 1 N i = 1 N T P i T P i + F N i
m P r e c i s i o n measures the proportion of correctly predicted pixels for each class.
m P r e c i s i o n = 1 N i = 1 N T P i T P i + F P i
m F 1 is the harmonic mean of Precision and Recall, commonly used to assess segmentation accuracy.
m F 1 = 1 N i = 1 N 2 T P i 2 T P i + F P i + F N i

3.5. Wetland Class Imbalance Loss (WIL)

In the semantic segmentation of the Yancheng coastal wetlands, land-cover categories exhibit a pronounced imbalance in both spatial extent and sample frequency. Dominant classes often occupy large continuous regions, whereas boundary-related or small-scale wetland features are sparsely distributed and easily suppressed during model training. To alleviate this issue, a weighted joint loss function is adopted to supervise the optimization process.
The proposed loss function combines Dice Loss and Focal Loss to jointly account for region-level overlap, pixel-wise classification accuracy, and the contribution of hard-to-classify samples. Dice Loss is employed to measure the overlap between predicted and reference regions, which is particularly effective for handling class imbalance by emphasizing overall shape consistency rather than individual pixel counts [42,43]. It is defined as:
L D i c e = 1 2 T P 2 T P + F P + F N
Focal Loss [44] is adopted to mitigate the dominance of easily classified samples during training and to emphasize hard or minority-class samples. Specifically, it down-weights the contribution of well-classified pixels, thereby encouraging the model to focus on challenging regions under class-imbalanced conditions. The Focal Loss is formulated as:
L F o c a l = c = 1 C ω c i = 1 N 1 p i , c γ log ( p i , c )
where p i , c denotes the predicted probability of the i-th pixel belonging to class c, ω c is the weight assigned to class c, and γ is the focusing parameter that controls the rate at which easy examples are down-weighted. In this study, γ is set to 2, which corresponds to the commonly recommended value in the original Focal Loss formulation [44]. This setting is widely used in semantic segmentation tasks and provides a balance between focusing on hard examples and maintaining stable training, without introducing additional hyperparameter tuning.
The total loss function is:
L W I L = λ D i c e L D i c e + λ F o c a l L F o c a l
In this formulation, Dice Loss emphasizes region-level overlap and boundary consistency, while Focal Loss alleviates the bias toward majority classes by rebalancing the gradient contributions of easy and hard samples [45]. The weighting coefficients λ D i c e and λ F o c a l control the relative influence of the two terms, enabling the loss function to jointly address boundary preservation and class imbalance during training.
In accordance with common practice in semantic segmentation studies that combine region-based and reweighting loss terms, equal weighting between Dice Loss and Focal Loss is frequently adopted to balance boundary accuracy and class imbalance [46,47]. Following this convention, both weighting coefficients were initially set to λ D i c e = 0.5 and λ F o c a l = 0.5 .
To verify the robustness of this choice, a small-scale sensitivity analysis was conducted by varying the loss weights with a step size of 0.1, i.e., { 0.6 , 0.4 } , { 0.5 , 0.5 } , and { 0.4 , 0.6 } . The results indicate that the balanced configuration consistently yields the best overall performance across multiple evaluation metrics. Therefore, λ D i c e = 0.5 and λ F o c a l = 0.5 are adopted in all experiments. 
In summary, the methodological framework proposed in this study for multispectral semantic segmentation of the Yancheng coastal wetlands integrates several key components. We incorporate a Spectral-Aware Embedding (SAE) mechanism in the encoder to explicitly model inter-band dependencies, enhancing spectral feature representation. A Wetland Boundary-Refined Decoder (WBRD) enhances the recovery of fine spatial details during decoding. Model optimization is guided by a weighted joint loss function that addresses class imbalance and emphasizes boundary preservation. Multiple quantitative evaluation metrics are employed to comprehensively assess segmentation performance. These components collectively provide a robust and generalizable framework for multispectral semantic segmentation in complex coastal wetland environments.

4. Experiments and Results

To validate the effectiveness of the multispectral wetland semantic segmentation method constructed in this study, systematic experiments were conducted on a self-built multispectral remote sensing dataset of the Yancheng coastal wetlands. The experiments consistently adopted the same data partitioning method, input resolution, and training strategy to ensure fairness in comparisons between different models. All models were trained and tested under identical hardware and software environments. The detailed experimental settings and parameter configurations are summarized in Table 5.
This study uses the self-constructed Yancheng coastal wetland multispectral remote sensing dataset, Yan14, which comprises 12 spectral bands covering typical natural wetlands, artificial wetlands, coastal zones, and non-wetland construction areas within the study area. Based on the national wetland classification system and spectral characteristics of ground objects, 14 land cover categories were defined, including silt beach, grass flat, reed marsh, suaeda marsh, spartina marsh, aquaculture pond, channel, paddy field, reclaimed wetland, salt pan, water body, suspended sediment, seawater, and other non-wetland class.

4.1. Baseline Model Comparison Experiments

To comprehensively evaluate the performance of different semantic segmentation architectures and to provide a solid basis for subsequent model optimization, this study selects a set of representative models in the field of remote sensing image semantic segmentation as benchmark methods. The selected models cover the major technical paradigms in this domain.
Specifically, classical convolutional neural network-based architectures that established the encoder–decoder paradigm are included, such as FCN [48], U-Net [49], and PSPNet [50]. To further examine the impact of enhanced context modeling, several advanced convolutional encoder–decoder models are adopted, including DeepLabV3+ [51], which employs atrous spatial pyramid pooling, UPerNet [52], based on feature pyramid networks, and GCNet [53], which emphasizes global context modeling.
In addition, recent Transformer-based and hybrid architectures are considered to reflect current research trends in remote sensing segmentation. These include the pure Transformer-based SegFormer [28], UPerNet-Swin [52,54] with a Swin Transformer backbone, as well as hybrid convolution–Transformer models such as ConvFormer [55] and UNetFormer [56]. Finally, a lightweight segmentation model based on MobileNetV4 [57] is included as a reference baseline to assess computational efficiency.
Figure 8 presents the training curves of m I o U for all compared models. Most models converge rapidly during the early training stage, with substantial performance gains observed within the first 20 epochs, followed by a gradual stabilization phase.
Among the compared methods, SegFormer exhibits the most stable training behavior. After approximately 40 epochs, it establishes a clear performance advantage, characterized by the highest final m I o U and the smallest fluctuation amplitude, indicating robust convergence. In contrast, UPerNet, ConvFormer, and UNetFormer also show relatively smooth convergence trends, but their final accuracies remain slightly lower than that of SegFormer.
GCNet demonstrates a noticeably slower convergence rate and inferior overall performance, suggesting limited effectiveness in capturing the complex spatial–spectral characteristics of coastal wetlands. Traditional convolutional models, such as U-Net and PSPNet, exhibit clear performance limitations. Owing to their architectural constraints and limited context modeling capability, these models struggle to adapt to the heterogeneous and fine-grained structures present in multispectral wetland scenes, resulting in substantially lower final m I o U values.
Table 6 reports the overall segmentation performance and computational characteristics of different baseline models. In terms of accuracy, SegFormer achieves the highest m I o U (0.7022) and m F 1 (0.8236), indicating superior overall segmentation quality. Although several models, including UNetFormer, UPerNet, and ConvFormer, achieved comparable performance in terms of mIoU, SegFormer shows a more balanced behavior across evaluation metrics and land-cover categories, which supports its selection as the primary comparative baseline.
Table 7 provides a quantitative comparison of model complexity and computational efficiency among the evaluated segmentation methods. SegFormer achieves a favorable balance between accuracy and computational cost, with a relatively low parameter count (24.73 M) and moderate FLOPs (56.14 G), while delivering the highest inference speed among high-performing models (148.65 FPS).
In contrast, architectures such as UPerNet-Swin and ConvFormer require substantially higher computational resources, with FLOPs exceeding 240 G, which constrains their efficiency despite competitive segmentation accuracy. Conversely, lightweight models, including MobileNetV4 and GCNet, exhibit high inference speed but noticeably lower segmentation performance, suggesting that efficiency-oriented designs alone may be insufficient to capture the complex spatial structures and spectral variability characteristic of coastal wetland environments.
Table 8 presents the class-wise I o U results for different land-cover categories. Overall, segmentation performance exhibits clear category-dependent characteristics. Classes with relatively distinct spectral properties and continuous spatial distributions, such as suspended sediment, seawater, and paddy fields, achieve consistently high I o U values across most models. Conversely, categories affected by strong spectral mixing or fragmented spatial patterns, including grass flats, channels, and Suaeda marshes, remain challenging and generally show lower segmentation accuracy.
From a model comparison perspective, SegFormer demonstrates stable performance across most wetland-related categories, achieving the highest or near-highest I o U values for complex land-cover types such as mudflats, reed marshes, reclaimed wetlands, and salt pans. UNetFormer and UPerNet also perform competitively for several categories, particularly in vegetated and reclaimed areas, but exhibit larger variability across classes, while traditional CNN-based models tend to underperform for narrow or fragmented land-cover types. Overall, the observed performance differences are consistent with the respective architectural characteristics of the models.

4.2. Ablation Study Design and Results Analysis

To further validate the effectiveness of each proposed module in the semantic segmentation task for the Yancheng coastal wetlands, this study designed a systematic ablation study focusing on the Spectral-Aware Embedding module (SAE), the Wetland Boundary-Refining Decoder (WBRD), and the Wetland Class Imbalance Loss function (WIL). Using the original SegFormer model as the baseline, various comparative models were constructed by progressively introducing different functional modules. Their performance was evaluated under the same dataset and training strategy.
Table 9 summarizes the quantitative results of the ablation experiments under different module combinations on the test set. Compared with the baseline SegFormer, introducing the Spectral-Aware Embedding (SAE) module leads to consistent improvements across all evaluation metrics, with m I o U increasing from 0.7022 to 0.7185 and the m F 1 improving from 0.8236 to 0.8347. This performance gain indicates that explicitly modeling spectral information is beneficial for multispectral remote sensing segmentation, as the SAE module enhances feature discrimination among spectrally similar wetland classes by re-weighting multi-band representations during feature encoding.
When the Wetland Boundary-Refining Decoder (WBRD) is further incorporated, the model achieves additional performance gains, particularly in m I o U and pixel accuracy ( P A ), which increase to 0.7327 and 0.9618, respectively. These improvements suggest that refining boundary representations during the decoding stage is effective for handling complex transition zones and fragmented structures commonly observed in coastal wetland environments. Although the introduction of WBRD increases computational cost, the notable accuracy improvement demonstrates its effectiveness in enhancing fine-grained spatial delineation.
By jointly integrating SAE, WBRD, and the Weighted Imbalance Loss (WIL), the Full Model achieves the best overall performance, with m I o U reaching 0.7459 and the m F 1 increasing to 0.8551. In particular, the marked improvement in m F 1 , recall, and precision reflects the contribution of WIL in alleviating class imbalance and improving recognition of minority and hard-to-segment categories. Overall, the ablation results demonstrate that each proposed module contributes to performance improvement at different stages, and their combined use yields complementary benefits for semantic segmentation in complex coastal wetland scenes.
Table 10 presents the category-level I o U results under different module configurations to analyze the contribution of each proposed component. Compared with the baseline SegFormer, the introduction of the Spectral-Aware Embedding (SAE) module results in consistent accuracy improvements across most land-cover categories, with more pronounced gains for spectrally similar vegetation types such as grass flats and reed marshes. This indicates that explicit spectral modeling enhances category discrimination in multispectral wetland imagery.
After incorporating the Wetland Boundary-Refining Decoder (WBRD), further improvements are observed, particularly for categories with complex boundaries or elongated spatial structures, including channels, aquaculture areas, and reclaimed wetlands. This suggests that boundary-aware feature refinement improves the delineation of fine-scale spatial structures in heterogeneous wetland scenes.
With the additional use of the Weighted Integrated Loss (WIL), the full model achieves the highest I o U values across all reported categories and yields the best overall performance ( m I o U = 0.7459 ). Notably, categories with limited samples or higher inter-class confusion, such as Suaeda marshes, channels, and the “Other” class, benefit further from the combined loss design. Overall, the results demonstrate that the proposed modules contribute in a complementary manner to improving wetland semantic segmentation performance at both category and global levels.

4.3. Visual Comparison and Analysis of Segmentation Results

Although quantitative evaluation metrics can reflect the performance differences of various models in wetland semantic segmentation tasks at an overall level, in complex wetland scenes, a model’s ability to depict spatial structures, feature boundaries, and small-scale targets is often difficult to fully represent through numerical indicators alone. To more intuitively analyze the performance of the proposed method in actual segmentation results, this section selects representative area samples from the test set for visual comparative analysis of the segmentation outcomes.
As shown in Figure 9, Region (a) and Region (d) feature key wetland types such as marshes and channels. For these types, the proposed method yields more complete segmentations and crisper boundaries than the baseline model. Region (b) focuses on aquaculture ponds, salt pan, and channels, where the baseline model incorrectly classifies a salt pan area as aquaculture pond, while the proposed model achieves a correct segmentation, effectively reducing inter-class confusion reducing omission of channels. Region (c) is dominated by elongated linear structures such as channels. In this region, the baseline model exhibits boundary blurring and local discontinuities, whereas the proposed method preserves clearer and more continuous boundaries, demonstrating a stronger capability for complex boundary recovery. Region (e) highlights misclassification of the “reclaimed wetland” and "grass flat" categories, where the proposed model shows improved recognition accuracy and reduces both misclassification.
Overall, the visual comparison across different scenarios indicates that the proposed multi-module enhanced SegFormer model achieves superior performance in terms of segmentation consistency, complex boundary delineation, and small-scale wetland type identification. These visual observations are consistent with the trends observed in quantitative evaluation results, further validating the effectiveness of the proposed method for semantic segmentation of Yancheng coastal wetlands.

5. Discussion

Based on the quantitative results and visual analyses presented above, the improved multispectral SegFormer demonstrates stable and competitive performance in semantic segmentation of the Yancheng coastal wetlands, consistently outperforming the baseline model across multiple evaluation metrics. The proposed framework shows advantages in handling wetland scenes characterized by complex boundaries and strong spectral heterogeneity, such as water–land interfaces, mudflats, and aquaculture areas. Compared with the original SegFormer, the improved model produces more spatially coherent predictions at land-cover transitions and preserves boundary continuity more effectively, indicating that the proposed structural refinements alleviate boundary blurring commonly observed in Transformer-based multi-scale fusion.
 The Spectral-Aware Embedding (SAE) module enhances the utilization of multispectral information by explicitly modeling and re-weighting inter-band relationships during the encoding stage. As demonstrated by the I o U results for categories such as silt beach and grass flat, the introduction of SAE consistently improves the performance of the model in these classes. These improvements indicate that spectral-aware feature modeling benefits both vegetation types with high spectral similarity and land-cover types that are spectrally distinct from each other, such as water bodies and vegetation. This further validates that explicit spectral modeling significantly enhances the model’s ability to discriminate between complex wetland features.
 To address the fragmented spatial patterns and complex morphological characteristics typical of coastal wetlands, the Wetland Boundary-Refining Decoder (WBRD) further improves the delineation of fine-scale structures during decoding. Categories with more complex boundaries or elongated shapes, such as channel, reclaimed wetland, and aquaculture, benefit most from WBRD, with notable improvements in boundary integrity and spatial consistency. This suggests that the integration of boundary-aware refinement improves the model’s ability to preserve detailed geometric structures and handle spatially complex wetland features.
 In addition, the Weighted Integrated Loss (WIL) contributes to improving the model’s robustness against class imbalance. It enhances the model’s ability to recognize minority land-cover categories and mitigates misclassification in challenging areas. As shown by the improved segmentation of categories such as Suaeda Marsh and Suspended Sediment, WIL significantly boosts the recognition accuracy for small and difficult-to-classify classes.
It should be noted that all experiments in this study were conducted within a single coastal wetland region. Although Yancheng wetlands encompass diverse land-cover types, complex spatial patterns, and strong spectral variability, which provide a representative testbed for coastal wetland segmentation, the generalization performance of the proposed method across different coastal regions and environmental conditions has not yet been explicitly validated. Differences in sediment composition, vegetation species, hydrological regimes, and image acquisition conditions may influence model behavior when applied to other study areas. Therefore, further validation using multi-region and cross-domain datasets is necessary to comprehensively assess the robustness and transferability of the proposed framework.
 Moreover, the introduction of additional modules inevitably increases computational complexity, which may affect deployment efficiency in large-scale or ultra-high-resolution applications. In addition, this study focuses on single-temporal imagery and does not explicitly model seasonal dynamics, which are pronounced in wetland ecosystems. Future work could incorporate multi-temporal observations and explore lightweight or transfer-learning-based strategies to enhance both efficiency and generalization capability.

6. Conclusions

6.1. Research Conclusions

This study, focusing on the Yancheng coastal wetlands and addressing the need for fine-scale classification of wetland surface cover from multispectral remote sensing imagery, constructed an improved SegFormer semantic segmentation method. By introducing adaptive improvement strategies at the encoding, decoding, and optimization objective levels, the model’s interpretation performance in complex wetland scenes was effectively enhanced. The main research conclusions are as follows:
(1) Enhanced Multispectral Feature Representation: To address the strong inter-band correlation and underutilization of information in multispectral remote sensing imagery, a Spectral-Aware Embedding (SAE) module was proposed. This module models the physical relationships between bands through a channel attention mechanism, enhancing the model’s sensitivity to specific spectral responses and improving the encoding effectiveness of multispectral features.
(2) Recovery of Complex Boundary Structures: To tackle the challenge of fragmented and semantically confusing boundaries in wetland features, a Wetland Boundary-Refining Decoder (WBRD) was designed. By introducing explicit boundary constraints during the multi-scale feature fusion stage, this method strengthens the capture capability for subtle geometric structures, significantly improving the spatial consistency and edge details of segmentation results.
(3) Optimization for Class Imbalance: To address the distribution imbalance present in the wetland dataset, a Weighted Joint Loss function (WIL) was constructed. This strategy ensures stable overall segmentation performance while increasing attention to minority land cover types through a re-weighting mechanism, effectively enhancing the segmentation accuracy for sparse categories.
Both quantitative and qualitative experimental results demonstrate that the proposed method outperforms mainstream baseline models in metrics such as m I o U , m F 1 , and small-target recognition accuracy, proving its effectiveness in coastal wetland monitoring tasks.

6.2. Research Prospects

With the continuous advancement of remote sensing observation technologies and deep learning methods, fine-scale wetland mapping still faces numerous opportunities and challenges. Future research could focus on the following aspects.
 First, extending the proposed framework to multi-region and cross-domain coastal wetland datasets to explicitly evaluate its generalization capability under varying environmental conditions, including differences in sediment composition, vegetation types, and hydrological regimes.
 Second, integrating multi-source (e.g., SAR and hyperspectral) and multi-temporal remote sensing data to better capture seasonal dynamics and long-term ecological processes, thereby enhancing the applicability of the model for dynamic wetland monitoring.
 Third, exploring lightweight architectures and transfer learning strategies to balance segmentation accuracy and computational efficiency, facilitating large-scale deployment in practical coastal wetland monitoring and management applications. 

Author Contributions

Conceptualization, S.P.; methodology, S.P.; validation, S.P., H.X. and N.L.; formal analysis, H.X.; investigation, N.L.; resources, S.P. and H.X.; data curation, S.P.; writing—original draft preparation, S.P.; writing—review and editing, S.P., H.X. and N.L.; visualization, S.P.; supervision, Y.Z.; project administration, Y.Z.; funding acquisition, Y.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China (32470525) and the National Key Research and Development Program of China (2024YFF1307204).

Data Availability Statement

The data are not publicly available due to the confidentiality of the research projects.

Conflicts of Interest

The authors declare that they have no conflicts of interest.

References

  1. Mitsch, W.J.; Bernal, B.; Hernandez, M.E. Ecosystem services of wetlands. Int. J. Biodivers. Sci. Ecosyst. Serv. Manag. 2015, 11, 1–4. [Google Scholar] [CrossRef] [Scilit]
  2. National Research Council (US) Committee on Low-Frequency Sound and Marine Mammals. Low-Frequency Sound and Marine Mammals; National Academies Press: Washington, DC, USA, 1994. [Google Scholar]
  3. Tian, B.; Wu, W.T.; Yang, Z.Q.; Zhou, Y.X. Drivers, Trends, and Potential Impacts of Long-Term Coastal Reclamation in China from 1985 to 2010. Estuar. Coast. Shelf Sci. 2016, 170, 83–90. [Google Scholar] [CrossRef] [Scilit]
  4. Wang, W.; Liu, H.; Li, Y.Q.; Su, J.L. Development and Management of Land Reclamation in China. Ocean Coast. Manag. 2014, 102, 415–425. [Google Scholar] [CrossRef] [Scilit]
  5. Fournier, R.A.; Grenier, M.; Lavoie, A.; Hélie, R. Towards a Strategy to Implement the Canadian Wetland Inventory Using Satellite Remote Sensing. Can. J. Remote Sens. 2007, 33, S1–S16. [Google Scholar] [CrossRef] [Scilit]
  6. LaRocque, A.; Phiri, C.; Leblon, B.; Pirotti, F.; Connor, K.; Hanson, A. Wetland Mapping with Landsat 8 OLI, Sentinel-1, ALOS-1 PALSAR, and LiDAR Data in Southern New Brunswick, Canada. Remote Sens. 2020, 12, 2095. [Google Scholar] [CrossRef] [Scilit]
  7. Yang, J.; Zhao, H.; Luo, Y.H.; Wang, J.D. A Review of Wetland Classification with High-Resolution Remote Sensing Image Based on Deep Learning. Remote Sens. Technol. Appl. 2025, 40, 923–935. [Google Scholar]
  8. Lv, J.N.; Shen, Q.; Lv, M.Z.; Li, Y.R.; Shi, L.; Zhang, P.Y. Deep Learning-Based Semantic Segmentation of Remote Sensing Images: A Review. Front. Ecol. Evol. 2023, 11, 1201125. [Google Scholar] [CrossRef] [Scilit]
  9. Liu, Q.W.; Huang, T.; Dong, Y.N.; Yang, J.Q.; Xiang, W. From Pixels to Images: Deep Learning Advances in Remote Sensing Image Semantic Segmentation. arXiv 2025, arXiv:2505.15147. [Google Scholar] [CrossRef] [Scilit]
  10. Al-Ruzouq, R.; Gibril, M.B.A.; Shanableh, A.; Bolcek, J.; Lamghari, F.; Hammour, N.A.; El-Keblawy, A.; Jena, R. Spectral-Spatial Transformer-Based Semantic Segmentation for Large-Scale Mapping of Individual Date Palm Trees Using Very High-Resolution Satellite Data. Ecol. Indic. 2024, 163, 112110. [Google Scholar] [CrossRef] [Scilit]
  11. Qian, S.Y.; Xue, Z.H.; Jia, M.M.; Chen, Y.P.; Su, H.J. Temporal-Spectral-Semantic-Aware Convolutional Transformer Network for Multi-Class Tidal Wetland Change Detection in Greater Bay Area. ISPRS J. Photogramm. Remote Sens. 2024, 216, 126–141. [Google Scholar] [CrossRef] [Scilit]
  12. Marjani, M.; Mohammadimanesh, F.; Mahdianpari, M.; Gill, E.W. A Novel Spatio-Temporal Vision Transformer Model for Improving Wetland Mapping Using Multi-Seasonal Sentinel Data. Remote Sens. Appl. Soc. Environ. 2025, 37, 101401. [Google Scholar] [CrossRef] [Scilit]
  13. Wang, R.K.; Ma, L.; He, G.J.; Johnson, B.A.; Yan, Z.Y.; Chang, M.; Liang, Y. Transformers for Remote Sensing: A Systematic Review and Analysis. Sensors 2024, 24, 3495. [Google Scholar] [CrossRef] [Scilit]
  14. Hong, D.F.; Han, Z.; Yao, J.; Gao, L.R.; Zhang, B.; Plaza, A.; Chanussot, J. SpectralFormer: Rethinking Hyperspectral Image Classification with Transformers. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5518615. [Google Scholar] [CrossRef] [Scilit]
  15. Zhang, Y.; Yang, R.Q.; Dai, Q.L.; Zhao, Y.L.; Xu, W.H.; Wang, J.; Wang, L.G. Boosting Semantic Segmentation of Remote Sensing Images by Introducing Edge Extraction Network and Spectral Indices. Remote Sens. 2023, 15, 5148. [Google Scholar] [CrossRef] [Scilit]
  16. Lin, X.F.; Cheng, Y.W.; Chen, G.; Chen, W.J.; Chen, R.; Gao, D.M.; Zhang, Y.L.; Wu, Y.B. Semantic Segmentation of China’s Coastal Wetlands Based on Sentinel-2 and Segformer. Remote Sens. 2023, 15, 3714. [Google Scholar] [CrossRef] [Scilit]
  17. Duan, S.N.; Zhao, J.Y.; Huang, X.Y.; Zhao, S.H. Semantic Segmentation of Remote Sensing Data Based on Channel Attention and Feature Information Entropy. Sensors 2024, 24, 1324. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Yan, G.D.; Jing, H.T.; Li, H.; Guo, H.C.; He, S. Enhancing Building Segmentation in Remote Sensing Images: Advanced Multi-Scale Boundary Refinement with MBR-HRNet. Remote Sens. 2023, 15, 3766. [Google Scholar] [CrossRef] [Scilit]
  19. Liu, C.J.; Qiao, Z.; Yan, H.W.; Wu, X.S.; Wang, J.W.; Xin, Y.Q. Semantic Segmentation Network of Remote Sensing Images Based on Dual Path Supervision. J. Beijing Univ. Aeronaut. Astronaut. 2025, 51, 732–741. (In Chinese) [Google Scholar] [CrossRef]
  20. Yu, J.; Cai, Y.; Lyu, X.; Xu, Z.N.; Wang, X.Y.; Fang, Y.W.; Jiang, W.X.; Li, X. Boundary-Guided Semantic Context Network for Water Body Extraction from Remote Sensing Images. Remote Sens. 2023, 15, 4325. [Google Scholar] [CrossRef] [Scilit]
  21. Wang, J.Q.; Chen, T.; Zheng, L.; Tie, J.; Zhang, Y.B.; Chen, P.T.; Luo, Z.Q.; Song, Q.J. A Multi-Scale Remote Sensing Semantic Segmentation Model with Boundary Enhancement Based on UNetFormer. Sci. Rep. 2025, 15, 14737. [Google Scholar] [CrossRef] [Scilit]
  22. Zhang, Y.; Wang, X.; Cai, J.Y.; Yang, Q. MW-SAM: Mangrove Wetland Remote Sensing Image Segmentation Network Based on Segment Anything Model. IET Image Process. 2024, 18, 4503–4513. [Google Scholar] [CrossRef] [Scilit]
  23. Ke, L.N.; Lu, Y.; Tan, Q.; Zhao, Y.; Wang, Q.M. Precise Mapping of Coastal Wetlands Using Time-Series Remote Sensing Images and Deep Learning Model. Front. For. Glob. Chang. 2024, 7, 1409985. [Google Scholar] [CrossRef] [Scilit]
  24. Di Vittorio, C.A.; Wiles, M.; Rabby, Y.W.; Movahedi, S.; Louie, J.; Hezrony, L.; Cifuentes, E.C.; Hinchman, W.; Schluter, A. Mapping Coastal Wetland Changes from 1985 to 2022 in the US Atlantic and Gulf Coasts Using Landsat Time Series and National Wetland Inventories. Remote Sens. Appl. Soc. Environ. 2025, 37, 101392. [Google Scholar] [CrossRef] [Scilit]
  25. Yuan, S.; Liang, X.G.; Lin, T.W.; Chen, S.; Liu, R.; Wang, J.; Zhang, H.S.; Gong, P. A Comprehensive Review of Remote Sensing in Wetland Classification and Mapping. arXiv 2025, arXiv:2504.10842. [Google Scholar] [CrossRef] [Scilit]
  26. Effah, D.; Zia, A.; Awrangjeb, M.; Gao, Y.S.; Sarpong, K. Advances in Machine Learning for Wetland Classification: A Comprehensive Survey of Methods and Applications. Artif. Intell. Rev. 2025, 59, 24. [Google Scholar] [CrossRef] [Scilit]
  27. Istiak, M.A.; Khan, R.H.; Rony, J.H.; Syeed, M.M.M.; Ashrafuzzaman, M.; Karim, M.R.; Hossain, M.S.; Uddin, M.F. AqUavplant Dataset: A High-Resolution Aquatic Plant Classification and Segmentation Image Dataset Using UAV. Sci. Data 2024, 11, 1411. [Google Scholar] [CrossRef] [Scilit]
  28. Xie, E.; Wang, W.H.; Yu, Z.D.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar] [CrossRef] [Scilit]
  29. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. Adv. Neural Inf. Process. Syst. 2017, 30, 6000–6010. [Google Scholar] [CrossRef] [Scilit]
  30. Dosovitskiy, A. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. [Google Scholar] [CrossRef] [Scilit]
  31. Yu, S.B.; Yang, H.; Wang, C.F.; Luo, D.S. Remote Sensing Information Extraction of Yancheng Wetland Rare Bird Reserve Based on Multi-Feature Optimization. Remote Sens. Technol. Appl. 2025, 40, 734–747. [Google Scholar]
  32. Li, J.L.; Tong, C.; Huang, R.P.; Tian, P.; Liu, R.Q.; Wang, L.J.; Zhou, Z.J. Spatio-Temporal Evolution of Coastal Wetlands Affected by Human Activities: A Case Study of Yancheng City, South Coast of Hangzhou Bay and Xiangshan Harbor Wetlands. J. Ningbo Univ. (Nat. Sci. Eng. Ed.) 2020, 33, 1–9. [Google Scholar]
  33. Li, J.X. Habitat Changes of Red-Crowned Cranes Derived from Remote Sensing Data in the Yancheng Coastal Wetland During 1989–2019. Master’s Thesis, University of Chinese Academy of Sciences (Aerospace Information Research Institute, Chinese Academy of Sciences), Beijing, China, 2021. [Google Scholar]
  34. Tan, C.W.; Wang, D.L.; Zhou, J.; Du, Y.; Luo, M.; Guo, W.S. Estimation of Leaf Nitrogen Concentration in Wheat by the Combinations of Two Vegetation Indexes Using HJ-CCD Images. Int. J. Agric. Biol. 2018, 20, 1908–1914. [Google Scholar]
  35. Wang, C.; Liu, H.Y.; Li, Y.F.; Wang, G.; Dong, B.; Chen, H.; Zhang, Y.N.; Zhao, Y.Q. A Study on Habitat Suitability and Ecological Threshold of Waterbird Guilds in Yancheng Coastal Wetlands: Implications for Habitat Structure Restoration. J. Ecol. Rural Environ. 2022, 38, 897–908. [Google Scholar] [CrossRef]
  36. Rouse, J.W., Jr.; Haas, R.H.; Schell, J.A.; Deering, D.W. Monitoring the Vernal Advancement and Retrogradation (Green Wave Effect) of Natural Vegetation; NASA: Washington, DC, USA, 1973. [Google Scholar]
  37. McFeeters, S.K. The Use of the Normalized Difference Water Index (NDWI) in the Delineation of Open Water Features. Int. J. Remote Sens. 1996, 17, 1425–1432. [Google Scholar] [CrossRef] [Scilit]
  38. Xu, H.Q. Modification of Normalised Difference Water Index (NDWI) to Enhance Open Water Features in Remotely Sensed Imagery. Int. J. Remote Sens. 2006, 27, 3025–3033. [Google Scholar] [CrossRef] [Scilit]
  39. Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 7132–7141. [Google Scholar] [CrossRef] [Scilit]
  40. Wassan, S.; Bilal, A.; Alzahrani, A.; Almohammadi, K.; Alrashidi, M.; Mousavirad, S.J. A Modified Vision Transformer Framework for Image-Based Land Cover Segmentation in Rural Architectural Design and Planning. Sci. Rep. 2025, 15, 32658. [Google Scholar] [CrossRef] [Scilit]
  41. Aouayeb, M.; Hamidouche, W.; Soladie, C.; Kpalma, K.; Seguier, R. Learning Vision Transformer with Squeeze and Excitation for Facial Expression Recognition. arXiv 2021, arXiv:2107.03107. [Google Scholar] [CrossRef] [Scilit]
  42. Dice, L.R. Measures of the Amount of Ecologic Association between Species. Ecology 1945, 26, 297–302. [Google Scholar] [CrossRef] [Scilit]
  43. Milletari, F.; Navab, N.; Ahmadi, S.A. V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation. In Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA, 25–28 October 2016; pp. 565–571. [Google Scholar] [CrossRef] [Scilit]
  44. Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.M.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar] [CrossRef] [Scilit]
  45. Isensee, F.; Jaeger, P.F.; Kohl, S.A.A.; Petersen, J.; Maier-Hein, K.H. nnU-Net: A Self-Configuring Method for Deep Learning-Based Biomedical Image Segmentation. Nat. Methods 2021, 18, 203–211. [Google Scholar] [CrossRef] [Scilit]
  46. Gao, K.; Wang, F.; Liu, Z.; Wang, M. Semantic Segmentation of Remote Sensing Images Based on Improved U-Net. J. Jilin Univ. (Earth Sci. Ed.) 2024, 54, 1752–1763. [Google Scholar]
  47. Elgamily, K.M.; Mohamed, M.A.; Abou-Taleb, A.M.; Ata, M.M. A novel W13 deep CNN structure for improved semantic segmentation of multiple objects in remote sensing imagery. Neural Comput. Appl. 2025, 37, 5397–5427. [Google Scholar] [CrossRef] [Scilit]
  48. Long, J.; Shelhamer, E.; Darrell, T. Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 3431–3440. [Google Scholar] [CrossRef] [Scilit]
  49. Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention (MICCAI 2015), Munich, Germany, 5–9 October 2015; pp. 234–241. [Google Scholar] [CrossRef] [Scilit]
  50. Zhao, H.S.; Shi, J.P.; Qi, X.J.; Wang, X.G.; Jia, J.Y. Pyramid Scene Parsing Network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 6230–6239. [Google Scholar] [CrossRef] [Scilit]
  51. Chen, L.C.; Zhu, Y.K.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018. [Google Scholar] [CrossRef] [Scilit]
  52. Xiao, T.; Liu, Y.C.; Zhou, B.L.; Jiang, Y.N.; Sun, J. Unified Perceptual Parsing for Scene Understanding. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 418–434. [Google Scholar] [CrossRef] [Scilit]
  53. Cao, Y.; Xu, J.R.; Lin, S.; Wei, F.Y.; Hu, H. GCNet: Non-Local Networks Meet Squeeze-Excitation Networks and Beyond. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, Seoul, Korea, 27–28 October 2019. [Google Scholar] [CrossRef] [Scilit]
  54. Liu, Z.; Lin, Y.T.; Cao, Y.; Hu, H.; Wei, Y.X.; Zhang, Z.; Lin, S.; Guo, B.N. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 9992–10002. [Google Scholar] [CrossRef] [Scilit]
  55. Gu, P.; Zhang, Y.; Wang, C.; Chen, D.Z. ConvFormer: Combining CNN and Transformer for Medical Image Segmentation. In Proceedings of the 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), Cartagena, Colombia, 18–21 April 2023; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
  56. Wang, L.; Li, R.; Zhang, C.; Fang, S.; Duan, C.; Meng, X.; Atkinson, P.M. UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery. ISPRS J. Photogramm. Remote Sens. 2022, 190, 196–214. [Google Scholar] [CrossRef] [Scilit]
  57. Qin, D.; Leichner, C.; Delakis, M.; Fornoni, M.; Luo, S.; Yang, F.; Wang, W.; Banbury, C.; Ye, C.; Akin, B.; et al. MobileNetV4: Universal Models for the Mobile Ecosystem. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 78–96. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Overview of the study area.
Figure 1. Overview of the study area.
Remotesensing 18 00745 g001
Figure 2. Examples of annotated wetland classes.
Figure 2. Examples of annotated wetland classes.
Remotesensing 18 00745 g002
Figure 3. Overall architecture of the proposed SegFormer-based framework for multispectral wetland segmentation.
Figure 3. Overall architecture of the proposed SegFormer-based framework for multispectral wetland segmentation.
Remotesensing 18 00745 g003
Figure 4. Structure of the Spectral-Aware Embedding (SAE) module.
Figure 4. Structure of the Spectral-Aware Embedding (SAE) module.
Remotesensing 18 00745 g004
Figure 5. Architecture of the proposed Wetland Boundary-Refined Decoder (WBRD).
Figure 5. Architecture of the proposed Wetland Boundary-Refined Decoder (WBRD).
Remotesensing 18 00745 g005
Figure 6. Structure of the DPR module.
Figure 6. Structure of the DPR module.
Remotesensing 18 00745 g006
Figure 7. Structure of the MBA module.
Figure 7. Structure of the MBA module.
Remotesensing 18 00745 g007
Figure 8. m I o U over training epochs.
Figure 8. m I o U over training epochs.
Remotesensing 18 00745 g008
Figure 9. From left to right in the image are the visualization results of Sentinel-2 432 combination, annotation file, SegFormer model prediction, and the improved model prediction presented in this paper. (a) Marshes and channels: More complete segmentation and crisper boundaries than baseline. (b) Aquaculture ponds, salt pan, channels: Correct classification of salt pan, reduced inter-class confusion. (c) Linear structures (channels): Clearer, more continuous boundaries with stronger complex boundary recovery. (d) Marshes and channels: Improved segmentation completeness and boundary sharpness. (e) Reclaimed wetland and grass flat: Enhanced recognition accuracy, reduced misclassification.
Figure 9. From left to right in the image are the visualization results of Sentinel-2 432 combination, annotation file, SegFormer model prediction, and the improved model prediction presented in this paper. (a) Marshes and channels: More complete segmentation and crisper boundaries than baseline. (b) Aquaculture ponds, salt pan, channels: Correct classification of salt pan, reduced inter-class confusion. (c) Linear structures (channels): Clearer, more continuous boundaries with stronger complex boundary recovery. (d) Marshes and channels: Improved segmentation completeness and boundary sharpness. (e) Reclaimed wetland and grass flat: Enhanced recognition accuracy, reduced misclassification.
Remotesensing 18 00745 g009
Table 1. Sentinel-2 image acquisition information.
Table 1. Sentinel-2 image acquisition information.
Image Acquisition DateBaseline/Orbit NumberSatelliteMosaic Domain Number
3 January 2023/02:41:09N0510/R0892AT51STT
14 March 2023/02:35:29N0510/R0892AT51STT
27 March 2023/02:45:29N0510/R1322AT50SQD
23 May 2023/02:35:39N0510/R0892AT51STU
23 May 2023/02:35:39N0510/R0892AT51STT
31 August 2023/02:35:29N0510/R0892AT51STU
15 October 2023/02:36:51N0510/R0892AT51STT
23 October 2023/02:47:49N0510/R1322AT51STU
30 October 2023/02:38:29N0510/R0892AT51STT
19 November 2023/02:40:09N0510/R0892AT51STT
7 December 2023/02:51:01N0510/R1322AT50SQD
12 February 2024/02:38:31N0510/R0892AT51STT
12 February 2024/02:38:31N0510/R0892AT51STS
12 February 2024/02:38:31N0510/R0892AT51SUS
22 April 2024/02:35:31N0510/R0892AT51STT 1
22 April 2024/02:35:31N0510/R0892AT51STU
7 May 2024/02:35:29N0510/R0892AT51STS
31 July 2024/02:35:51N0511/R0892AT51STS
31 July 2024/02:35:51N0511/R0892AT51STT 2
31 July 2024/02:35:51N0511/R0892AT51STU
24 October 2024/02:36:59N0511/R0892AT51STT
28 November 2024/02:40:31N0511/R0892AT51STT
28 November 2024/02:40:31N0511/R0892AT51STS
28 November 2024/02:40:31N0511/R0892AT51SUS
28 November 2024/02:40:31N0511/R0892AT51STU
28 November 2024/02:40:31N0511/R0892AT50SQC
28 December 2024/02:41:21N0511/R0892AT51STT
28 December 2024/02:41:21N0511/R0892AT51STS
28 December 2024/02:41:21N0511/R0892AT50SQC
28 December 2024/02:41:21N0511/R0892AT51STU
Note: A single Sentinel-2 overpass (same acquisition time and orbit) may cover multiple MGRS tiles. Records with identical acquisition dates, orbit numbers, and satellite identifiers but different mosaic domain numbers therefore represent different spatial tiles. 1 Corresponds to products S2A_MSIL2A_20240422T023531_N0510_R089_T51STT_20240422T061251 and S2A_MSIL2A_20240422T023531_N0510_R089_T51STT_20240422T082350. 2 Corresponds to products S2A_MSIL2A_20240731T023551_N0511_R089_T51STT_20240731T074450 and S2A_MSIL2A_20240731T023551_N0511_R089_T51STT_20240731T082846.
Table 2. Wetland land-cover classification system.
Table 2. Wetland land-cover classification system.
Primary CategorySecondary CategoryCategory Description
Natural WetlandSilt BeachHigh water content, low spectral reflectance, exposed intertidal mudflat
Grass FlatAreas with sparse vegetation coverage and no distinct dominant species
Water BodyResidual water surfaces and other natural water bodies
Reed MarshMarsh vegetation type dominated by reeds
Spartina MarshMarsh vegetation type dominated by Spartina alterniflora
Suaeda MarshMarsh vegetation type dominated by Suaeda salsa
Artificial WetlandAquaculture PondRegular rectangular ponds, primarily used for aquaculture
ChannelDrainage/irrigation ditches and small linear water bodies
Paddy FieldAgricultural cultivation area with periodic vegetation changes
Reclaimed WetlandUndergoing restoration, with complex mixed feature distribution
Salt PanArtificial evaporation ponds, high reflectance, grid-like structure
Non-WetlandSea WaterDeep water areas, high blue light reflectance
Suspended SedimentHighly turbid water bodies influenced by tides or runoff
OthersRoads, buildings, bare land, and other non-wetland categories
Table 3. Number of images containing each category in the training, validation, and test sets 1.
Table 3. Number of images containing each category in the training, validation, and test sets 1.
DatasetSilt B.Grass F.Water B.Reed M.Spart. M.Suaeda M.Aqua. P.Chan.Paddy F.Recl. Wet.Salt P.Sea W.Susp. Sed.Others
TRAIN239224201745177013837472331372016781836233133721152780
VAL1771891381781581511593151007759111171114
TEST185128152154159731132108316828145287163
1 Category abbreviations: Silt B. (Silt Beach), Grass F. (Grass Flat), Water B. (Water Body), Reed M. (Reed Marsh), Spart. M. (Spartina Marsh), Suaeda M. (Suaeda Marsh), Aqua. P. (Aquaculture Pond), Chan. (Channel), Paddy F. (Paddy Field), Recl. Wet. (Reclaimed Wetland), Salt P. (Salt Pan), Sea W. (Sea Water), Susp. Sed. (Suspended Sediment).
Table 4. Pixel statistics for each wetland category.
Table 4. Pixel statistics for each wetland category.
Category NamePixel CountPercentage
Suspended Sediment190,729,66220.73%
Aquaculture Pond105,326,24811.45%
Grass Flat92,850,76210.09%
Sea Water86,967,4849.45%
Reed Marsh85,217,2189.26%
Reclaimed Wetland71,556,0777.78%
Paddy Field69,126,6407.51%
Channel61,247,8996.66%
Spartina Marsh59,873,3456.51%
Others32,649,0963.55%
Silt Beach25,793,7252.80%
Water Body22,130,1352.41%
Suaeda Marsh13,124,8111.43%
Salt Pan3,536,2450.38%
Percentage values are rounded to two decimal places.
Table 5. Experimental platform.
Table 5. Experimental platform.
Experimental BackgroundVersion/Environment
PyTorch2.1.0
Python3.10
Ubuntu22.04
CUDA12.1
GPURTX 3090 (24G)
CPU20 vCPU AMD EPYC 7642 48-Core Processor
Training Epochs120
Batch size8
OptimizerAdamW
Table 6. I o U performance of different semantic segmentation models on various categories.
Table 6. I o U performance of different semantic segmentation models on various categories.
ModelmIoUPAmF1mRecallmPrecisionBest Epoch
SegFormer0.70220.95370.82360.83730.815588
UnetFormer0.69220.93590.81100.82240.8193105
UPerNet0.68910.94460.80910.81750.8127110
ConvFormer0.68730.94910.81510.82850.810497
UperNet-Swin0.64520.93180.78330.79310.7823109
DeepLabV3+0.62780.94010.77960.78830.7762103
MobileNetV40.62760.93550.77280.77580.775193
FCN0.61530.93100.76050.76830.761291
PSPNet0.60570.89180.75420.75580.755198
GCNet0.60040.93820.74530.75210.739999
U-Net0.56790.93780.72640.73920.721891
Bold values indicate the highest score in each column (except for Best Epoch).
Table 7. Quantitative comparison of segmentation performance and computational complexity.
Table 7. Quantitative comparison of segmentation performance and computational complexity.
ModelParameters (M)FLOPs (G)FPS
FCN26.1533.29125.30
U-Net32.5545.21118.20
PSPNet25.3137.99124.91
DeepLabV3+26.7138.85128.27
UPerNet37.31157.3288.88
UperNet-Swin84.18242.5645.23
SegFormer24.7356.14148.65
GCNet26.2925.37167.78
ConvFormer45.70267.0742.86
UnetFormer26.8584.6278.94
MobileNetV412.8142.04139.21
Table 8. I o U for each sub-category across different semantic segmentation models.
Table 8. I o U for each sub-category across different semantic segmentation models.
Class NameFCNU-NetPSPNetDeep. 1UPerNetUP.-Swin 2SegF. 3GCNetConvF. 4UnetF. 5MobileV4 6
Silt Beach0.59960.52290.55230.60590.65560.60940.68790.61560.62250.60690.6015
Grass Flat0.54200.51660.48620.53040.55220.51780.61090.52030.57620.60110.5597
Water Body0.62510.59190.52350.51210.68630.55990.65310.63340.61420.68010.6361
Reed Marsh0.52330.48810.51460.60430.62090.51820.63730.53790.58320.65520.5242
Spartina Marsh0.59580.53020.54490.62590.68560.57240.67870.59350.57950.67740.5159
Suaeda Marsh0.50340.48210.45260.43860.52510.43280.56380.46800.42150.52860.4559
Aquaculture Pond0.63670.46620.55170.61180.69740.54220.70640.59910.66040.69320.6143
Channel0.43790.45660.41910.42560.52830.53580.54540.41560.50690.52790.4946
Paddy Field0.69210.61790.62980.76240.80760.66310.81010.61850.76910.81390.6137
Reclaimed Wetland0.55740.56550.52580.61370.73240.62050.73850.56670.62470.73340.6485
Salt Pan0.61590.52230.51140.69440.72940.58370.76840.62800.61510.73160.7501
Seawater0.81530.78590.74430.81010.84250.75670.84330.77280.81110.82920.8297
Suspended Sediment0.94280.89470.88240.95720.95200.90240.95460.92390.95490.94270.9354
Other0.52610.51040.53620.59670.63230.57250.63260.51200.59640.66880.6064
1 Deep.: DeepLabV3+; 2 UP.-Swin: UperNet-Swin; 3 SegF.: SegFormer; 4 ConvF.: ConvFormer; 5 UnetF.: UnetFormer; 6 MobileV4: MobileNetV4. Best I o U values in each row are bolded.
Table 9. Ablation study results of model overall performance under different module combinations.
Table 9. Ablation study results of model overall performance under different module combinations.
SAEWBRDWILmIoUPAmF1mRecallmPrecisionParams (M)FLOPs (G)FPS
×××0.70220.95370.82360.83730.819724.7356.14148.65
××0.71850.95740.83470.84010.835824.7458.57142.45
×0.73270.96180.83780.84230.839126.9799.9683.4
0.74590.96740.85510.86650.851327.2103.675.26
SAE: Spectral-Aware Embedding module; WBRD: Wetland Boundary Refinement Decoder; WIL: Wetland Imbalance Loss function. “×”: disabled, “✔”: enabled. Params: Parameters.
Table 10. I o U performance for each sub-category across different model configurations.
Table 10. I o U performance for each sub-category across different model configurations.
CategorySegFormerSegFormer+SAESegFormer+SAE+WBRDFull Model
Silt Beach0.68790.69170.71470.7357
Grass Flat0.61090.71030.72140.7198
Water Body0.65310.67480.67230.6851
Reed Marsh0.63730.67310.72130.7228
Spartina Marsh0.67870.68130.69140.6998
Suaeda Marsh0.56380.57550.59410.6046
Aquaculture Pond0.70640.71670.74520.7511
Channel0.54540.55070.59270.6193
Paddy Field0.81010.81970.82700.8222
Reclaimed Wetland0.73850.74280.76170.7685
Salt Pan0.76840.77630.77010.7916
Seawater0.84330.86420.86360.8698
Suspended Sediment0.95460.95480.95110.9629
Other0.63260.62730.63170.6888
mIoU0.70220.71850.73270.7459
Bold values indicate the best I o U score in each row.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Peng, S.; Xie, H.; Liu, N.; Zeng, Y. Semantic Segmentation of Multispectral Remote Sensing Imagery for Coastal Wetlands with SegFormer. Remote Sens. 2026, 18, 745. https://doi.org/10.3390/rs18050745

AMA Style

Peng S, Xie H, Liu N, Zeng Y. Semantic Segmentation of Multispectral Remote Sensing Imagery for Coastal Wetlands with SegFormer. Remote Sensing. 2026; 18(5):745. https://doi.org/10.3390/rs18050745

Chicago/Turabian Style

Peng, Simin, Huachen Xie, Nian Liu, and Yi Zeng. 2026. "Semantic Segmentation of Multispectral Remote Sensing Imagery for Coastal Wetlands with SegFormer" Remote Sensing 18, no. 5: 745. https://doi.org/10.3390/rs18050745

APA Style

Peng, S., Xie, H., Liu, N., & Zeng, Y. (2026). Semantic Segmentation of Multispectral Remote Sensing Imagery for Coastal Wetlands with SegFormer. Remote Sensing, 18(5), 745. https://doi.org/10.3390/rs18050745

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop