An Improved DeepLabV3+ Network for Bare Rock Segmentation in Remote Sensing Images
Abstract
1. Introduction
- A task-oriented DeepLabV3+ improvement strategy is developed for GF-1 bare-rock segmentation in complex plateau urban environments. Based on comparative selection of the ResNet50 backbone and Adam optimizer, SE channel attention, SATM-based spatial attention, and RFB receptive-field enhancement are coordinately integrated to strengthen channel, spatial, and multi-scale feature representation, thereby forming a targeted methodological framework for addressing feature confusion, boundary ambiguity, and scale variation of bare-rock targets.
- The proposed architecture combines channel, spatial, and multi-scale feature modeling to improve feature discrimination and boundary representation in urban scenes containing buildings, vegetation, and shadow-affected areas.
- The proposed GF-1-based framework provides a methodological basis for bare-rock segmentation in plateau cities and may support applications such as geological-hazard monitoring, ecological-environment assessment, and bare-rock interpretation in similar plateau river-valley environments.
2. Materials and Methods
2.1. Study Area and Data Sources
2.1.1. Study Area
2.1.2. Data Sources
2.2. Research Methods
2.2.1. Image Preprocessing
- (1)
- To accommodate the different characteristics of GF-1 PMS2 multispectral and panchromatic data, a “separate preprocessing–fusion enhancement” workflow was adopted, as illustrated in Figure 1. Standardized processing, which includes radiometric calibration, atmospheric correction, orthorectification, and image fusion, was performed on the raw L1A data using ENVI 5.6.2 software to mitigate systematic errors and atmospheric interference. This processing generated a 2-m pan-sharpened multispectral raster that served as the source product for subsequent image-patch generation. The complete multispectral raster was not directly used as the network input; only its red, green, and blue visible bands were retained for the RGB-based deep-learning experiments.
- (i)
- Radiometric calibration converts the original digital number (DN) values into top-of-atmosphere radiance. The official GF-1 absolute calibration coefficients were applied to the original Level-1A multispectral bands. The DN values were converted to at-sensor radiance using the official GF-1 calibration coefficients, thereby reducing radiometric inconsistencies associated with sensor response and gain differences. This procedure provided standardized radiometric data for subsequent atmospheric correction.
- (ii)
- Atmospheric correction was performed using the built-in FLAASH module in ENVI 5.6.2 to reduce the effects of atmospheric scattering, gas absorption, aerosols, and water vapor on the remotely sensed signal and to estimate surface reflectance. The study area is located at approximately 36.7° N, and the imagery used for the internal experiments was acquired on 24 September 2024. According to the latitude–season guidance for the standard MODTRAN atmospheric profiles implemented in FLAASH, the Mid-Latitude Summer profile is specified for September acquisitions near the 40° N latitude band. Therefore, the Mid-Latitude Summer atmospheric profile was adopted for the atmospheric correction of the study-area imagery. Here, “Summer” refers to the name of the predefined MODTRAN atmospheric profile rather than to the calendar season of image acquisition. Other core parameters were set as follows: a rural aerosol model was used, with an initial visibility of 40 km and an aerosol elevation of 1.50 km; the MODTRAN spectral resolution was set to 5 cm−1; adjacency-effect correction was enabled; and no aerosol inversion method was applied. Elevation information was derived from the mean elevation of the corresponding GMTED2010 region. The atmospheric correction parameters were matched to the image location and acquisition time to estimate surface reflectance, thereby reducing atmospheric effects and improving the spectral separability of rocks and vegetation.
- (iii)
- Orthorectification was performed to correct geometric distortions associated with terrain relief and satellite viewing geometry. Using the RPC model and the SRTM 30 m DEM in ENVI 5.6.2, the GF-1 multispectral and 2 m panchromatic images were orthorectified separately. The corrected imagery was projected to WGS 84/UTM Zone 47N and resampled using cubic convolution. The orthorectified imagery provided a consistent spatial reference for subsequent image fusion, patch generation, and sample annotation.
- (iv)
- Image fusion was performed to enhance the spatial resolution of the multispectral imagery using the high-resolution panchromatic image. In ENVI 5.6.2, the orthorectified multispectral imagery was first converted from band sequential (BSQ) to band interleaved by line (BIL) format. Both BSQ and BIL describe data-storage organization and do not alter pixel values, spectral information, or spatial resolution. The format conversion was used only as part of the ENVI preprocessing workflow before NNDiffuse pan-sharpening. Subsequently, the NNDiffuse pan sharpening algorithm was applied to fuse the 2 m panchromatic image with the preprocessed multispectral image, producing a fused raster with a spatial resolution of 2 m. The resulting fused raster was used as the source data for subsequent image-patch generation rather than being directly input into the deep-learning network. For the subsequent segmentation experiments, only the red, green, and blue visible bands of the pan-sharpened product were retained to generate RGB image patches, whereas the near-infrared band was excluded from the network input. Therefore, pan sharpening was treated as an image-preprocessing step for spatial enhancement and RGB image construction, and no generalized conclusion regarding full-multispectral spectral fidelity was made in this study.
- (2)
- In this study, the preprocessed fused images were cropped into non-overlapping 256 × 256-pixel patches using the Split Raster tool in ArcMap 10.8. The fused GF-1 imagery had a spatial resolution of 2 m, corresponding to a ground coverage of approximately 512 × 512 m for each 256 × 256-pixel patch. The complete study-area imagery produced a candidate patch pool larger than that required for the model experiments. Final experimental samples were therefore selected from spatially defined candidate regions according to image validity and the representation of bare-rock and background scenes. For geographic partitioning, four spatially adjacent candidate patches arranged in a 2 × 2 configuration were combined into a Spatial Group with an approximate ground extent of 1.024 × 1.024 km. Spatial Groups were used as the basic geographic units for assigning the training, validation, and internal-test regions. Each Spatial Group was associated with a single dataset partition, thereby preserving the geographic integrity of locally adjacent candidate patches. A minimum edge-to-edge separation threshold of 1000 m was established among the training, validation, and internal-test regions. Because the candidate patches were aligned to a regular 512-m geographic grid, the resulting minimum geographic separation was 1024 m. GIS-based distance measurements were performed in the WGS 84/UTM Zone 47N projected coordinate system. The minimum edge-to-edge distances were 1024 m for the training–validation, training–internal-test, and validation–internal-test pairs. Spatial Groups located within the inter-partition separation zones were reserved as geographic buffers and excluded from experimental sample selection. The spatial partitioning strategy and the distribution of candidate regions for internal and external datasets are illustrated in Figure 2. This visualization demonstrates the geographic separation among dataset subsets and the independent sampling region used for external evaluation. After spatial partitioning, the training, validation, and internal-test candidate regions contained 776, 155, and 156 candidate patches, respectively, as shown in Figure 2a. Following establishment of the geographic partitions, representative samples were selected within the corresponding training, validation, and internal-test candidate regions. Sample selection considered image validity, bare-rock occurrence, and representative background conditions. This procedure resulted in 510 original training patches, 128 validation patches, and 128 internal-test patches.
- (3)
- For construction of the deep-learning dataset, the red, green, and blue visible bands were selected from the fused GF-1 imagery to form a true-color RGB representation, whereas the near-infrared band was excluded from the current network input. For the 16-bit input imagery, each retained RGB channel was independently subjected to percentile-based intensity stretching. Specifically, the 1st and 99th percentiles were calculated from non-zero pixels in each channel, and pixel values between these limits were linearly mapped to the 8-bit range of 0–255, with values outside the limits clipped accordingly. After 8-bit conversion, fixed contrast and color-saturation enhancement factors of 1.3 and 1.2, respectively, were applied. The resulting RGB image patches were saved in JPEG format with a quality setting of 98. JPEG was retained as the image representation used consistently in the established RGB dataset construction, annotation, training, and evaluation pipeline. This format choice was made for consistency with the RGB processing workflow rather than for preservation of the complete quantitative radiometric information of the original multispectral raster. Pixel-level manual annotation was performed using LabelMe by one primary annotator and subsequently cross-checked by a second researcher. The pan-sharpened GF-1 imagery was used as the primary annotation reference, while high-resolution reference imagery accessed through the Ovital Map platform was used as auxiliary visual information for uncertain regions. The auxiliary imagery was used only for label interpretation and verification and was not used as network input. A two-pixel-wide contour along each manually delineated bare-rock boundary was assigned the ignore label 255 to represent ambiguous bare-rock–background transition pixels. The same annotation and label-verification protocol was applied to the internal and external datasets. At the patch level, the training subset contained 430 bare-rock-containing and 80 background-only patches, while the validation subset contained 111 bare-rock-containing and 17 background-only patches. The internal test set contained 115 bare-rock-containing and 13 background-only patches. Data augmentation was applied only to the 510 original training samples. Images and their corresponding semantic labels were synchronously transformed at a ratio of 1:1, producing 510 additional image–label pairs and increasing the training set to 1020 image–label pairs. The validation set and internal test set were not augmented, and each retained 128 original samples. Therefore, the final training, validation, and internal-test sets contained 1020, 128, and 128 image–label pairs, respectively. Geographic partitioning and final sample selection were completed before data augmentation. The augmentation operations included random scaling with a scale factor sampled from 0.25 to 2.0, aspect-ratio perturbation with a jitter factor of 0.3, horizontal flipping with a probability of 0.50, Gaussian blurring with a 5 × 5 kernel and a probability of 0.25, and random rotation within −10° to +10° with a probability of 0.25. HSV-based photometric perturbation was also applied to the RGB images, with hue, saturation, and value parameters set to 0.1, 0.7, and 0.3, respectively. Geometric transformations were synchronously applied to the RGB images and their corresponding masks, whereas photometric transformations were applied only to the RGB images. Random scaling, aspect-ratio perturbation, and HSV-based photometric perturbation were applied to each augmented training sample, corresponding to an application probability of 1.0. No random augmentation was applied to the validation, internal-test, or external-test sets.
2.2.2. Model Construction
- (1)
- Backbone Network Selection
- (i)
- Analysis of Backbone Network Characteristics
- (ii)
- Performance Comparison and Backbone Selection
- (2)
- Optimizer Selection: SGD and Adam
- (3)
- SE (Squeeze-and-Excitation) Channel Attention Module
- (4)
- SATM (Spatial Attention Module)
- (5)
- RFB (Receptive Field Block)
2.2.3. Evaluation Metrics
- (1)
- Intersection over Union (IoU)
- (2)
- Precision
- (3)
- F1-score
- (4)
- Matthews Correlation Coefficient (MCC)
2.2.4. Experimental Configuration
3. Results
3.1. Ablation Experiment Results of the Improved DeepLabV3+ Model
- (1)
- As the backbone network, ResNet50 effectively mitigates the vanishing gradient problem in deep networks through its residual connection architecture. Its bottleneck design facilitates deep feature extraction while regulating the number of parameters, thereby enhancing model stability during training. Its deep feature-representation capability provided the baseline for subsequent module integration and supported discrimination between bare rock and background. Under this configuration, the bare-rock IoU was 70.35%, the bare-rock precision was 75.65%, and the bare-rock F1-score was 82.60%.
- (2)
- Based on the ResNet50 backbone network, the Adam-based training configuration was evaluated against the SGD-based configuration used in the preceding experiment. The Adam optimizer integrates the benefits of momentum gradient descent with adaptive parameter updates. By computing first-order and second-order moment estimates of the gradient and applying bias correction, it dynamically adjusts the parameter update magnitude during optimization. Under the reported experimental setting, the ResNet50 + Adam configuration achieved a bare-rock IoU of 74.17%, a bare-rock precision of 84.05%, and a bare-rock F1-score of 85.17%, compared with 70.35%, 75.65%, and 82.60%, respectively, for the tested ResNet50 + SGD configuration.
- (3)
- The SE module uses global average pooling and fully connected layers to learn channel weights and recalibrate deep feature responses. This mechanism emphasizes feature channels associated with bare-rock color, texture, boundary, and semantic information and is intended to improve the discriminative representation of bare-rock and background features. After the SE module was incorporated into the ResNet50 + Adam baseline, the mean bare-rock IoU, bare-rock precision, and bare-rock F1-score reached 78.60 ± 0.11%, 86.63 ± 0.21%, and 88.02 ± 0.07%, respectively. These results indicate higher mean values for all three reported bare-rock class metrics relative to the ResNet50 + Adam baseline under the evaluated experimental setting.
- (4)
- The SATM-based spatial attention module emphasizes the spatial locations of bare-rock targets by generating spatial attention maps that enhance edge detail features. By integrating spatial information from channel average pooling and max pooling, the learned attention map emphasizes spatially informative regions and suppresses irrelevant background responses. This mechanism enhances the representation of bare-rock boundaries in complex scenes, including regions affected by building shadows and topographic variation. In fine-grained segmentation tasks, the SATM module enhances the representation of boundary-related spatial features between bare rock and background. Under the evaluated experimental setting, the SE + SATM configuration achieved a mean bare-rock IoU of 78.91 ± 0.33%, a mean bare-rock precision of 87.85 ± 1.43%, and a mean bare-rock F1-score of 88.21 ± 0.21%.
- (5)
- The RFB receptive field enhancement module constructs a multi-scale receptive field through multi-branch dilated convolutions with varying dilation rates, thereby expanding the model’s spatial perception range. To address the multi-scale distribution of bare-rock targets in urban Xining, ranging from scattered outcrops to extensive exposed bedrock, the RFB captures feature information at different spatial scales. Compared with the SE + SATM configuration without RFB, adding RFB changed the mean bare-rock IoU from 78.91 ± 0.33% to 79.36 ± 0.30%, the mean bare-rock precision from 87.85 ± 1.43% to 88.65 ± 0.55%, and the mean bare-rock F1-score from 88.21 ± 0.21% to 88.49 ± 0.19%. These results indicate a modest positive change after introducing RFB. Considering the relatively small magnitude of the improvement together with the observed run-to-run variability, the contribution of RFB is conservatively interpreted as supplementary optimization for multi-scale feature representation rather than as a definitive or statistically significant improvement.
- (6)
- The final configuration combines ResNet50, Adam, SE, SATM, and RFB to provide backbone feature extraction, adaptive optimization, channel recalibration, spatial attention, and multi-scale feature modeling within a unified DeepLabV3+ framework. The ablation results indicate that configurations involving SE and SATM were associated with larger mean changes in the reported segmentation metrics, whereas the additional inclusion of RFB produced a smaller incremental change. To compare the training behavior of the original and modified configurations, their training and validation loss curves are shown in Figure 9, and the validation-set mIoU curves are shown in Figure 10. Both configurations were independently trained three times using random seeds 11, 22, and 33. The solid curves represent the mean values across the three independent runs, whereas the shaded regions indicate ±1 standard deviation. As shown in Figure 9, the original DeepLabV3+ configuration exhibited greater fluctuations in validation loss, whereas the modified configuration showed a smoother decrease in training loss and a more stable validation-loss trajectory. Figure 10 further shows that the modified configuration maintained higher mean validation mIoU values during the later stages of training. Together, these convergence curves characterize the optimization behavior and run-to-run variability of the two configurations under the evaluated experimental setting.
- (7)
- To further examine the influence of class imbalance on the final model, supplementary loss-function sensitivity experiments were performed using the complete ResNet50 + Adam + SE + SATM + RFB configuration. Under the fixed-seed supplementary setting, weighted cross-entropy produced a bare-rock IoU of 58.39%, a bare-rock precision of 58.91%, and a bare-rock F1-score of 73.73%. Although increasing the contribution of the minority bare-rock class was intended to mitigate class imbalance, the substantial decreases in bare-rock IoU and precision suggest that stronger class weighting was associated with increased false-positive bare-rock predictions under this supplementary setting. Focal loss achieved a bare-rock IoU of 78.91%, a bare-rock precision of 86.07%, and a bare-rock F1-score of 88.22%, whereas the combined CE + Dice loss achieved a bare-rock IoU of 78.75%, a bare-rock precision of 89.42%, and a bare-rock F1-score of 88.11%, respectively. These two alternative losses maintained relatively high bare-rock segmentation performance but did not show a clear overall improvement over the standard cross-entropy results reported for the final model in Table 2. Taken together, the supplementary experiments indicate that stronger imbalance-oriented loss formulations did not necessarily improve overall bare-rock segmentation performance under the present dataset and training configuration. Under the evaluated configuration, standard cross-entropy provided the best overall balance among the loss functions tested in this supplementary experiment and was therefore retained for the final model.
- (8)
- To provide a more complete evaluation under the highly imbalanced rock–background distribution, additional metrics were calculated for the final model from the pixel-level confusion matrices of the three independent runs. The proposed model achieved a mean background IoU of 98.95 ± 0.02%, a mean mIoU of 89.15 ± 0.16%, and a mean MCC of 0.8796 ± 0.0020. These results indicate that the model maintained high segmentation agreement for both background and bare-rock classes under the evaluated setting. The mean normalized confusion-matrix statistics across the three independent runs showed that 99.48% of background pixels were correctly classified as background, whereas 0.52% were misclassified as bare rock. For the bare-rock class, 88.34% of pixels were correctly classified, whereas 11.66% were misclassified as background. These results further characterize the false-positive and false-negative behavior of the final model.
3.2. Comparison of Network Models
- (1)
- Segmentation Performance Analysis
- (2)
- Full-Area Prediction Visualization
3.3. External Evaluation in an Adjacent Region
4. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Malik, O.A.; Puasa, I.; Lai, D.T.C. Segmentation for Multi-Rock Types on Digital Outcrop Photographs Using Deep Learning Techniques. Sensors 2022, 22, 8086. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Du, S.; Xing, J.; Li, J.; Du, S.; Zhang, C.; Sun, Y. Open-Pit Mine Extraction from Very High-Resolution Remote Sensing Images Using OM-DeepLab. Nat. Resour. Res. 2022, 31, 3173–3194. [Google Scholar] [CrossRef] [Scilit]
- Feng, H.; Hu, Q.; Zhao, P.; Wang, S.; Ai, M.; Zheng, D.; Liu, T. FTransDeepLab: Multimodal Fusion Transformer-Based DeepLabv3+ for Remote Sensing Semantic Segmentation. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4406618. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Wu, J.; Xie, H.; Xiao, D.; Ran, M. Semantic Segmentation of Urban Remote Sensing Images Based on Deep Learning. Appl. Sci. 2024, 14, 7499. [Google Scholar] [CrossRef] [Scilit]
- Wang, D.; Sui, X. Extraction of the Bare Rock Rate in Counties of Typical Karst Areas Based on Unmanned Aerial Vehicle Images and Satellite Data, China. Land Degrad. Dev. 2023, 34, 5499–5513. [Google Scholar] [CrossRef] [Scilit]
- Xi, J.; Jiang, Q.; Liu, H.; Gao, X. Lithological Mapping Research Based on Feature Selection Model of ReliefF-RF. Appl. Sci. 2023, 13, 11225. [Google Scholar] [CrossRef] [Scilit]
- Zhao, C.; Xiao, Z.; Zhang, Y.; Yuan, C.; Yang, J. Alteration Mineral Information Extraction Based on Image Super-Resolution Technology. Int. J. Appl. Earth Obs. Geoinf. 2025, 144, 104872. [Google Scholar] [CrossRef] [Scilit]
- Xie, F.; Yang, Y. Hierarchical Multiscale Fusion with Coordinate Attention for Lithologic Mapping from Remote Sensing. Remote Sens. 2026, 18, 413. [Google Scholar] [CrossRef] [Scilit]
- Baibatsha, A.; Vyazovetsky, Y.; Kembayev, M.; Rais, S.; Agaliyeva, B. On the Use of Remote Sensing Data to Study the Geological Structure and Forecast Mineral Resources of the Shu-Ile Suture. Min. Miner. Depos. 2024, 18, 56–70. [Google Scholar] [CrossRef] [Scilit]
- Munzareen, M.; Bibi, F.; Ihsan, S.; Shah, K.S.; Emad, M.Z. Advanced CNN-Based Remote Sensing for Mineral Mapping of Porphyry Systems in the Gilgit Region. Min. Miner. Depos. 2025, 19, 53–62. [Google Scholar] [CrossRef] [Scilit]
- Baibatsha, A.B.; Kembayev, M.K.; Rais, S.E.; Yan, W.; Amantayev, A.K.; Biyakyshev, Y.T. Geodynamics of the Shu-Ile Ore Zone: Integration of Geophysical, Geochemical and Cosmogeological Methods. Eng. J. Satbayev Univ. 2025, 147, 30–36. [Google Scholar] [CrossRef] [Scilit]
- Fu, J.; Wang, C.; Liu, M.; Li, X.; Liu, Y.; Shi, W.; Wang, R. HyperR3SNet: Leveraging Hyperbolic Space and Vision Foundation Models for Remote Sensing Semantic Segmentation. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5620016. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Li, Y.; Yang, X.; Jiang, R.; Zhang, L. RSAM-Seg: A SAM-Based Model with Prior Knowledge Integration for Remote Sensing Image Semantic Segmentation. Remote Sens. 2025, 17, 590. [Google Scholar] [CrossRef] [Scilit]
- Feng, X.; Wei, C.; Xue, X.; Zhang, Q.; Liu, X. RST-DeepLabv3+: Multi-Scale Attention for Tailings Pond Identification with DeepLab. Remote Sens. 2025, 17, 411. [Google Scholar] [CrossRef] [Scilit]
- Qi, Y.; Zhang, Z.; Hu, Y.; Liu, P.; Gao, M.; Zhai, G. CS-DeepLabV3+: A Fine-Grained Semantic Segmentation Method for Mining Land Use in the Kunlun Mountain Region Using High-Resolution Remote Sensing Imagery. Appl. Sci. 2026, 16, 4820. [Google Scholar] [CrossRef] [Scilit]
- Feng, Y.; Fan, Z.; Yan, Y.; Jiang, Z.; Zhang, S. MFAFNet: Multi-Scale Feature Adaptive Fusion Network Based on DeepLab V3+ for Cloud and Cloud Shadow Segmentation. Remote Sens. 2025, 17, 1229. [Google Scholar] [CrossRef] [Scilit]
- Fu, H.; Li, X.; Zhu, L.; Pan, X.; Wu, T.; Li, W.; Feng, Y. DSC-DeepLabv3+: A Lightweight Semantic Segmentation Model for Weed Identification in Maize Fields. Front. Plant Sci. 2025, 16, 1647736. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, S.; Wang, R.; Wang, L.; Liu, S.; Ye, J.; Xu, H.; Niu, R. An Approach for Monitoring Shallow Surface Outcrop Mining Activities Based on Multisource Satellite Remote Sensing Data. Remote Sens. 2023, 15, 4062. [Google Scholar] [CrossRef] [Scilit]
- Yao, X.; Guo, Q.; Li, A. Light-Weight Cloud Detection Network for Optical Remote Sensing Images with Attention-Based DeeplabV3+ Architecture. Remote Sens. 2021, 13, 3617. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Xu, Y.; Chen, J. Method of Convolutional Neural Networks for Lithological Classification Using Multisource Remote Sensing Data. Remote Sens. 2026, 18, 29. [Google Scholar] [CrossRef] [Scilit]
- Sui, Y.; Zhang, L.; Sun, Z.; Yi, W.; Wang, M. Research on Coal and Rock Recognition in Coal Mining Based on Artificial Neural Network Models. Appl. Sci. 2024, 14, 864. [Google Scholar] [CrossRef] [Scilit]
- He, C.; Liu, Y.; Wang, D.; Liu, S.; Yu, L.; Ren, Y. Automatic Extraction of Bare Soil Land from High-Resolution Remote Sensing Images Based on Semantic Segmentation with Deep Learning. Remote Sens. 2023, 15, 1646. [Google Scholar] [CrossRef] [Scilit]
- Ge, Y.; Wang, H.; Liu, G.; Chen, Q.; Tang, H. Automated Identification of Rock Discontinuities from 3D Point Clouds Using a Convolutional Neural Network. Rock Mech. Rock Eng. 2025, 58, 3683–3700. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Gan, S.; Yuan, X.; Luo, W.; Ma, C.; Li, Y. An Improved MSEM-Deeplabv3+ Method for Intelligent Detection of Rock Mass Fractures. Remote Sens. 2026, 18, 1041. [Google Scholar] [CrossRef] [Scilit]
- Lai, P.; Lv, C.; Zhou, L.; Yang, S.; Xu, J.; Dong, Q.; He, M. Improved Lightweight DeepLabV3+ for Bare Rock Extraction from High-Resolution UAV Imagery. Ecol. Inform. 2025, 89, 103204. [Google Scholar] [CrossRef] [Scilit]
- Gao, Z.; Li, Z.; Yao, W.; Zhang, T.; Qiu, S.; Liu, Z. Forest Road Extraction via Optimized DeepLabv3+ and Multi-Temporal Remote Sensing for Wildfire Emergency Response. Appl. Sci. 2026, 16, 3228. [Google Scholar] [CrossRef] [Scilit]
- Lin, L.; Liu, L.; Liu, M.; Zhang, Q.; Feng, M.; Khalil, Y.S.; Yin, F. DEDNet: Dual-Encoder DeeplabV3+ Network for Rock Glacier Recognition Based on Multispectral Remote Sensing Image. Remote Sens. 2024, 16, 2603. [Google Scholar] [CrossRef] [Scilit]
- Hao, Z.; Li, X.; Zhu, Q.; Li, Y.; Mao, Z.; Chen, J.; Pan, D. AMFA-DeepLab: An Improved Lightweight DeepLabV3+ Adaptive Multi-Statistic Fusion Attention Network for Sea Ice Segmentation in GaoFen-1 Images. Remote Sens. 2026, 18, 783. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Wu, G. Multi-Scale Feature Fusion and Global Context Modeling for Fine-Grained Remote Sensing Image Segmentation. Appl. Sci. 2025, 15, 5542. [Google Scholar] [CrossRef] [Scilit]
- Wan, P.; Han, X.; Zhai, R.; Gan, X. Automated Recognition of Rock Mass Discontinuities on Vegetated High Slopes Using UAV Photogrammetry and an Improved Superpoint Transformer. Remote Sens. 2026, 18, 357. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Wang, S.; Dinavahi, V.; Yang, L.; Wu, D.; Shen, M. Landslide Recognition Based on DeepLabv3+ Framework Fusing ResNet101 and ECA Attention Mechanism. Appl. Sci. 2025, 15, 2613. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Li, Q.; Lu, J.; Zheng, K.; Wei, L.; Xiang, Q. A Transfer Learning Remote Sensing Landslide Image Segmentation Method Based on Nonlinear Modeling and Large Kernel Attention. Appl. Sci. 2025, 15, 3855. [Google Scholar] [CrossRef] [Scilit]
- Yuan, T.; Hu, B. REU-Net: A Remote Sensing Image Building Segmentation Network Based on Residual Structure and the Edge Enhancement Attention Module. Appl. Sci. 2025, 15, 3206. [Google Scholar] [CrossRef] [Scilit]
- Bai, H.; Bai, T.; Li, W.; Liu, X. A Building Segmentation Network Based on Improved Spatial Pyramid in Remote Sensing Images. Appl. Sci. 2021, 11, 5069. [Google Scholar] [CrossRef] [Scilit]
- Peng, L.; Ding, M.; Xue, Q.; Dong, Y.; Li, Y.; Zhou, P.; Li, Z. Shape-Constrained ResU-Net for Old Landslides Detection in the Loess Plateau. Appl. Sci. 2026, 16, 546. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar] [CrossRef] [Scilit]
- Chollet, F. Xception: Deep Learning with Depthwise Separable Convolutions. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1800–1807. [Google Scholar] [CrossRef] [Scilit]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.-C. MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 4510–4520. [Google Scholar] [CrossRef] [Scilit]
- Howard, A.; Sandler, M.; Chu, G.; Chen, L.-C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. Searching for MobileNetV3. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1314–1324. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Wang, H.; Liu, J.; Zhao, X.; Lu, Y.; Qu, T.; Tian, H.; Su, J.; Luo, D.; Yang, Y. A Lightweight Winter Wheat Planting Area Extraction Model Based on Improved DeepLabv3+ and CBAM. Remote Sens. 2023, 15, 4156. [Google Scholar] [CrossRef] [Scilit]
- Liu, K.; Xi, Y.; Liu, J.; Zhou, W.; Zhang, Y. MFFNet: A Building Extraction Network for Multi-Source High-Resolution Remote Sensing Data. Appl. Sci. 2023, 13, 13067. [Google Scholar] [CrossRef] [Scilit]
- Ye, N.; Xu, Y.-H.; Zhou, W.; Yu, G.; Zhou, D. MKF-NET: KAN-Enhanced Vision Transformer for Remote Sensing Image Segmentation. Appl. Sci. 2025, 15, 10905. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Shen, L.; Albanie, S.; Sun, G.; Wu, E. Squeeze-and-Excitation Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 2011–2023. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ai, H.; Zhu, X.; Han, Y.; Ma, S.; Wang, Y.; Ma, Y.; Qin, C.; Han, X.; Yang, Y.; Zhang, X. Extraction of Levees from Paddy Fields Based on the SE-CBAM UNet Model and Remote Sensing Images. Remote Sens. 2025, 17, 1871. [Google Scholar] [CrossRef] [Scilit]
- Wubineh, B.Z.; Rusiecki, A.; Halawa, K. SE-DeepLabV3+: Cervical Cell Segmentation and Classification Using a Novel SE-Based DeepLabV3+ and Ensemble Method. IEEE Access 2025, 13, 116430–116441. [Google Scholar] [CrossRef] [Scilit]
- Liu, K.-H.; Lin, B.-Y. MSCSA-Net: Multi-Scale Channel Spatial Attention Network for Semantic Segmentation of Remote Sensing Images. Appl. Sci. 2023, 13, 9491. [Google Scholar] [CrossRef] [Scilit]
- Zhuang, C.; Yuan, X.; Gu, L.; Wei, Z.; Fan, Y.; Guo, X. Frequency Regulated Channel-Spatial Attention Module for Improved Image Classification. Expert Syst. Appl. 2025, 260, 125463. [Google Scholar] [CrossRef] [Scilit]
- Deepak, G.D.; Hiremath, P.; Bhat, S.K. Taguchi-Optimized DeepLabV3+ for Semantic Segmentation in Autonomous Driving Applications. Ain Shams Eng. J. 2026, 17, 103985. [Google Scholar] [CrossRef] [Scilit]
- Wei, H.; Xu, X.; Ou, N.; Zhang, X.; Dai, Y. DEANet: Dual Encoder with Attention Network for Semantic Segmentation of Remote Sensing Imagery. Remote Sens. 2021, 13, 3900. [Google Scholar] [CrossRef] [Scilit]
- Moghimi, A.; Welzel, M.; Celik, T.; Schlurmann, T. A Comparative Performance Analysis of Popular Deep Learning Models and Segment Anything Model (SAM) for River Water Segmentation in Close-Range Remote Sensing Imagery. IEEE Access 2024, 12, 52067–52085. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Li, X.; Zhou, L.; Chen, J.; He, Z.; Guo, L.; Liu, J. Adaptive Receptive Field Enhancement Network Based on Attention Mechanism for Detecting the Small Target in the Aerial Image. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5600118. [Google Scholar] [CrossRef] [Scilit]
- Mahara, A.; Khan, M.R.K.; Deng, L.; Rishe, N.; Wang, W.; Sadjadi, S.M. Automated Road Extraction from Satellite Imagery Integrating Dense Depthwise Dilated Separable Spatial Pyramid Pooling with DeepLabV3+. Appl. Sci. 2025, 15, 1027. [Google Scholar] [CrossRef] [Scilit]
- Ruan, S.; Wan, Q.; Chen, R.; Hu, M.; Guo, X.; Song, K. Context-Aware Feature Enhancement Network for Remote Sensing Image Semantic Segmentation. Remote Sens. 2026, 18, 543. [Google Scholar] [CrossRef] [Scilit]
- Li, F.; Mou, Y.; Zhang, Z.; Liu, Q.; Jeschke, S. A Novel Model for the Pavement Distress Segmentation Based on Multi-Level Attention DeepLabV3+. Eng. Appl. Artif. Intell. 2024, 137, 109175. [Google Scholar] [CrossRef] [Scilit]












| Backbone | IoU (%) | Precision (%) | F1-Score (%) |
|---|---|---|---|
| Xception | 64.85 | 83.85 | 78.67 |
| ResNet50 | 70.35 | 75.65 | 82.60 |
| MobileNetV2 | 70.01 | 78.20 | 82.36 |
| MobileNetV3 | 67.38 | 76.33 | 80.51 |
| ResNet50 | Adam | SE | SATM | RFB | IoU (%) | Precision (%) | F1-Score (%) |
|---|---|---|---|---|---|---|---|
| √ | √ | × | × | × | 74.17 ± 0.82 | 84.05 ± 0.28 | 85.17 ± 0.54 |
| √ | √ | × | × | √ | 78.43 ± 0.64 | 86.79 ± 0.97 | 87.91 ± 0.40 |
| √ | √ | × | √ | × | 78.57 ± 0.52 | 87.00 ± 1.63 | 87.99 ± 0.32 |
| √ | √ | √ | × | × | 78.60 ± 0.11 | 86.63 ± 0.21 | 88.02 ± 0.07 |
| √ | √ | √ | × | √ | 78.62 ± 0.50 | 86.88 ± 1.02 | 88.03 ± 0.31 |
| √ | √ | × | √ | √ | 78.80 ± 0.19 | 87.70 ± 0.28 | 88.14 ± 0.12 |
| √ | √ | √ | √ | × | 78.91 ± 0.33 | 87.85 ± 1.43 | 88.21 ± 0.21 |
| √ | √ | √ | √ | √ | 79.36 ± 0.30 | 88.65 ± 0.55 | 88.49 ± 0.19 |
| Model | IoU (%) | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|---|
| U-Net | 76.06 | 90.92 | 82.31 | 86.40 |
| SegFormer | 69.82 | 75.80 | 89.83 | 82.22 |
| PSPNet | 62.67 | 85.73 | 69.97 | 77.05 |
| HRNet | 73.55 | 85.19 | 84.33 | 84.76 |
| Modified DeepLabV3+ | 79.36 | 88.65 | 88.34 | 88.49 |
| Hardware Component | Specification |
|---|---|
| Workstation | Lenovo workstation (Lenovo Group Limited, Beijing, China) |
| CPU | Intel Core i7-9700 CPU @ 3.00 GHz |
| Clock Frequency | 3.00 GHz |
| GPU | NVIDIA Quadro P2200 (Professional Computing GPU) |
| GPU VRAM Capacity | 5 GB |
| RAM (System Memory) | 32.0 GB |
| Model | Parameters (M) | MACs (G) | FLOPs (G) | Model Size (MiB) | Inference Time (ms) |
|---|---|---|---|---|---|
| U-Net | 62.92 | 27.89 | 55.78 | 240.64 | 32.88 |
| SegFormer | 84.59 | 24.94 | 49.88 | 323.07 | 60.55 |
| PSPNet | 49.07 | 48.64 | 97.28 | 187.52 | 43.94 |
| HRNet | 65.85 | 23.46 | 46.92 | 252.14 | 40.30 |
| Modified DeepLabV3+ | 57.60 | 12.37 | 24.74 | 220.22 | 24.36 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Q.; Lei, H.; Hu, X.; Guo, R. An Improved DeepLabV3+ Network for Bare Rock Segmentation in Remote Sensing Images. Appl. Sci. 2026, 16, 9218. https://doi.org/10.3390/app16189218
Wang Q, Lei H, Hu X, Guo R. An Improved DeepLabV3+ Network for Bare Rock Segmentation in Remote Sensing Images. Applied Sciences. 2026; 16(18):9218. https://doi.org/10.3390/app16189218
Chicago/Turabian StyleWang, Qiang, Haochuan Lei, Xiasong Hu, and Rongrong Guo. 2026. "An Improved DeepLabV3+ Network for Bare Rock Segmentation in Remote Sensing Images" Applied Sciences 16, no. 18: 9218. https://doi.org/10.3390/app16189218
APA StyleWang, Q., Lei, H., Hu, X., & Guo, R. (2026). An Improved DeepLabV3+ Network for Bare Rock Segmentation in Remote Sensing Images. Applied Sciences, 16(18), 9218. https://doi.org/10.3390/app16189218
