MLCRNet: Multi-Level Context Refinement for Semantic Segmentation in Aerial Images
Abstract
1. Introduction
- We propose a MCT module, which dynamically harvests contextual information from the semantic, image, and local perspectives.
- The EMCT module is designed to address feature redundancy and improve the efficiency of our MCT module. Furthermore, a MLCR module is proposed on the basis of EMCT and FPN to enhance feature representation by aggregating multi-level contextual information.
- We propose a novel MLCRNet based on the feature pyramid framework for accurate semantic segmentation.
2. Related Work
2.1. Semantic Segmentation
2.2. Context Aggregation
2.3. Semantic Segmentation of Aerial Imagery
3. Methods
3.1. General Contextual Refinement Framework
3.1.1. Local-Level Context
3.1.2. Image-Level Context
3.1.3. Semantic-Level Context
3.2. EMCT
3.2.1. Multi-Level Context Transform
3.2.2. Reduction of Computational Complexity
3.3. Multi-Level Context Refinement Module
3.4. MLCRNet
4. Experiments and Results
4.1. Experimental Setup
4.1.1. Benchmarks
4.1.2. Implementation Details
4.1.3. Training Settings
4.1.4. Inference Settings
4.1.5. Reproducibility
4.2. Ablation Study
4.2.1. Ablation Studies of the MLCR Module to Different Layers
4.2.2. Ablation Studies of Different Level Contexts
4.2.3. Ablation Studies of Local-Level Context Receptive Fields
4.2.4. Ablation Studies of Computation Cost
4.3. Comparison with State-of-the-Art
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Appendix A. Robustness Evaluation
Appendix A.1. Incorrect Labels and Rectification

Appendix A.2. Robustness Evaluation Results
| Model | Backbone | Stride | mIoU (%) | Acc (%) | F1 |
|---|---|---|---|---|---|
| ISNet [32] | ResNet50 | 16× | 70.2 | 81.3 | 81.8 |
| FCN [11] | ResNet50 | 16× | 71.5 | 81.9 | 82.8 |
| OCRNet [29] | ResNet50 | 16× | 73.6 | 83.6 | 84.2 |
| DepLabV3+ [15] | ResNet50 | 16× | 74.5 | 84.2 | 84.8 |
| SCARF [31] | ResNet50 | 16× | 74.6 | 83.9 | 84.8 |
| SFNet [17] | ResNet50 | 16× | 74.7 | 84.1 | 84.9 |
| Ours | ResNet50 | 16× | 75.3 | 84.6 | 85.4 |
References
- Kang, Y.; Lu, Z.; Zhao, C.; Xu, Y.; Kim, J.-W.; Gallegos, A.J. InSAR monitoring of creeping landslides in mountainous regions: A case study in Eldorado National Forest, California. Remote Sens. Environ. 2021, 258, 112400. [Google Scholar] [CrossRef] [Scilit]
- Bianchi, F.M.; Grahn, J.; Eckerstorfer, M.; Malnes, E.; Vickers, H. Snow avalanche segmentation in SAR images with fully convolutional neural networks. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2020, 14, 75–82. [Google Scholar] [CrossRef] [Scilit]
- Xu, F.; Somers, B. Unmixing-based Sentinel-2 downscaling for urban land cover mapping. ISPRS J. Photogramm. Remote Sens. 2021, 171, 133–154. [Google Scholar] [CrossRef] [Scilit]
- Luo, X.; Tong, X.; Pan, H. Integrating multiresolution and multitemporal Sentinel-2 imagery for land-cover mapping in the Xiongan New Area, China. IEEE Trans. Geosci. Remote Sens. 2021, 59, 1029–1040. [Google Scholar] [CrossRef] [Scilit]
- Azizi, A.; Abbaspour-Gilandeh, Y.; Vannier, E.; Dusséaux, R.; Mseri-Gundoshmian, T.; Moghaddam, H.A. Semantic segmentation: A modern approach for identifying soil clods in precision farming. Biosyst. Eng. 2020, 196, 172–182. [Google Scholar] [CrossRef] [Scilit]
- Anand, T.; Sinha, S.; Mandal, M.; Chamola, V.; Yu, F.R. AgriSegNet: Deep aerial semantic segmentation framework for iot-assisted precision agriculture. IEEE Sens. J. 2021, 21, 17581–17590. [Google Scholar] [CrossRef] [Scilit]
- Azimi, S.M.; Fischer, P.; Korner, M.; Reinartz, P. Aerial LaneNet: Lane-marking semantic segmentation in aerial imagery using wavelet-enhanced cost-sensitive symmetric fully convolutional neural networks. IEEE Trans. Geosci. Remote Sens. 2019, 57, 2920–2938. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.; Chen, H.; Du, C.; Li, J. MSACon: Mining spatial attention-based contextual information for road extraction. IEEE Trans. Geosci. Remote Sens. 2021, 60, 5604317. [Google Scholar] [CrossRef] [Scilit]
- Song, L.; Xia, M.; Jin, J.; Qian, M.; Zhang, Y. SUACDNet: Attentional change detection network based on siamese U-shaped structure. Int. J. Appl. Earth Obs. Geoinf. 2021, 105, 102597. [Google Scholar] [CrossRef] [Scilit]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems; Pereira, F., Burges, C.J.C., Bottou, L., Wein-berger, K.Q., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2012; Volume 25. [Google Scholar]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; IEEE Computer Society: Los Alamitos, CA, USA, 2015; pp. 3431–3440. [Google Scholar]
- Girshick, R. Fast R-CNN. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 1440–1448. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.-C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 40, 834–848. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.-C.; Papandreou, G.; Schroff, F.; Adam, H. Rethinking atrous convolution for semantic image segmentation. arXiv 2017, arXiv:1706.05587. [Google Scholar]
- Chen, L.-C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Li, X.; You, A.; Zhu, Z.; Zhao, H.; Yang, M.; Yang, K.; Tan, S.; Tong, Y. Semantic flow for fast and accurate scene parsing. In Proceedings of the Computer Vision—ECCV 2020, Glasgow, UK, 23–28 August 2020; Vedaldi, A., Bischof, H., Brox, T., Frahm, J.-M., Eds.; Springer International Publishing: Cham, Switzerland, 2020; pp. 775–793. [Google Scholar]
- Kirillov, A.; Girshick, R.; He, K.; Dollár, P. Panoptic feature pyramid networks. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 6399–6408. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image seg-mentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015, Munich, Germany, 5–9 October 2015; Navab, N., Hornegger, J., Wells, W.M., Frangi, A.F., Eds.; Springer International Publishing: Cham, Switzerland, 2015; pp. 234–241. [Google Scholar]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 936–944. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Girshick, R.; Gupta, A.; He, K. Non-local neural networks. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017. [Google Scholar]
- Fu, J.; Liu, J.; Tian, H.; Li, Y.; Bao, Y.; Fang, Z.; Lu, H. Dual attention network for scene segmentation. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 3141–3149. [Google Scholar] [CrossRef] [Scilit]
- Hu, P.; Caba, F.; Wang, O.; Lin, Z.; Sclaroff, S.; Perazzi, F. Temporally distributed networks for fast video semantic segmentation. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 8815–8824. [Google Scholar]
- Zhang, H.; Wang, C.; Xie, J. Co-occurrent features in semantic segmentation. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 548–557. [Google Scholar]
- Zhu, Z.; Xu, M.; Bai, S.; Huang, T.; Bai, X. Asymmetric non-local neural networks for semantic segmentation. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea, 27 October–2 November 2019; pp. 593–602. [Google Scholar]
- Yan, L.; Huang, J.; Xie, H.; Wei, P.; Gao, Z. Efficient depth fusion transformer for aerial image semantic segmentation. Remote Sens. 2022, 14, 1294. [Google Scholar] [CrossRef] [Scilit]
- Yu, C.; Wang, J.; Gao, C.; Yu, G.; Shen, C.; Sang, N. Context prior for scene segmentation. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 12413–12422. [Google Scholar]
- Bello, I. LambdaNetworks: Modeling long-range interactions without attention. In Proceedings of the International Conference on Learning Representations, Vienna, Austria, 4 May 2021. [Google Scholar]
- Yuan, Y.; Chen, X.; Wang, J. Object-contextual representations for semantic segmentation. In Proceedings of the Computer Vision—ECCV 2020, Glasgow, UK, 23–28 August 2020; Vedaldi, A., Bischof, H., Brox, T., Frahm, J.-M., Eds.; Springer International Publishing: Cham, Switzerland, 2020; pp. 173–190. [Google Scholar]
- Zhang, F.; Chen, Y.; Li, Z.; Hong, Z.; Liu, J.; Ma, F.; Han, J.; Ding, E. ACFNet: Attentional class feature network for semantic segmentation. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea, 27 October–2 November 2019; pp. 6797–6806. [Google Scholar]
- Ding, X.; Shen, C.; Che, Z.; Zeng, T.; Peng, Y. SCARF: A semantic constrained attention refinement network for semantic segmentation. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada, 11–17 October 2021; pp. 3002–3011. [Google Scholar]
- Jin, Z.; Liu, B.; Chu, Q.; Yu, N. ISNet: Integrate image-level and semantic-level context for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 11–17 October 2021. [Google Scholar]
- Niu, R.; Sun, X.; Tian, Y.; Diao, W.; Chen, K.; Fu, K. Hybrid multiple attention network for semantic segmentation in aerial images. IEEE Trans. Geosci. Remote Sens. 2021, 60, 5603018. [Google Scholar] [CrossRef] [Scilit]
- Shang, R.; Zhang, J.; Jiao, L.; Li, Y.; Marturi, N.; Stolkin, R. Multi-scale adaptive feature fusion network for semantic segmentation in remote sensing images. Remote Sens. 2020, 12, 872. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Lin, S.; Ding, L.; Bruzzone, L. Multi-scale context aggregation for semantic segmentation of remote sensing images. Remote Sens. 2020, 12, 701. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Zhang, W.; Zhang, T.; Li, J. Hrcnet: High-resolution context extraction net-work for semantic segmentation of remote sensing images. Remote Sens. 2020, 13, 71. [Google Scholar] [CrossRef] [Scilit]
- Fan, M.; Lai, S.; Huang, J.; Wei, X.; Chai, Z.; Luo, J.; Wei, X. Rethinking BiSeNet for real-time semantic segmentation. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 9711–9720. [Google Scholar]
- Wang, J.; Sun, K.; Cheng, T.; Jiang, B.; Deng, C.; Zhao, Y.; Liu, D.; Mu, Y.; Tan, M.; Wang, X.; et al. Deep high-resolution representation learning for visual recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 3349–3364. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Cao, Y.; Xu, J.; Lin, S.; Wei, F.; Hu, H. Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Korea, 27–28 October 2019; pp. 1971–1980. [Google Scholar] [CrossRef] [Scilit]
- Song, Q.; Li, J.; Li, C.; Guo, H.; Huang, R. Fully attentional network for semantic segmentation. arXiv 2021, arXiv:2112.04108. [Google Scholar]
- Li, R.; Zheng, S.; Zhang, C.; Duan, C.; Su, J.; Wang, L.; Atkinson, P.M. Multiattention network for semantic segmentation of fine-resolution remote sensing images. IEEE Trans. Geosci. Remote Sens. 2021, 60, 5607713. [Google Scholar] [CrossRef] [Scilit]
- Huang, L.; Yuan, Y.; Guo, J.; Zhang, C.; Chen, X.; Wang, J. Interlaced sparse self-attention for semantic segmentation. arXiv 2019, arXiv:1907.12273. [Google Scholar]
- Huang, Z.; Wang, X.; Wei, Y.; Huang, L.; Shi, H.; Liu, W.; Huang, T.S. CCNet: Criss-cross attention for semantic segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 1, 9133304. [Google Scholar] [CrossRef] [Scilit]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 6230–6239. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Yang, G.; Zhang, G. Collaborative network for super-resolution and semantic segmentation of remote sensing images. IEEE Trans. Geosci. Remote Sens. 2021, 60, 4404512. [Google Scholar] [CrossRef] [Scilit]
- Saha, S.; Mou, L.; Qiu, C.; Zhu, X.X.; Bovolo, F.; Bruzzone, L. Unsupervised deep joint segmentation of multitemporal high-resolution images. IEEE Trans. Geosci. Remote Sens. 2020, 58, 8780–8792. [Google Scholar] [CrossRef] [Scilit]
- Du, S.; Du, S.; Liu, B.; Zhang, X. Incorporating DeepLabv3+ and object-based image analysis for semantic segmentation of very high resolution remote sensing images. Int. J. Digit. Earth 2020, 14, 357–378. [Google Scholar] [CrossRef] [Scilit]
- Yang, G.; Zhang, Q.; Zhang, G. EANet: Edge-aware network for the extraction of buildings from aerial images. Remote Sens. 2020, 12, 2161. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Identity mappings in deep residual networks. In Proceedings of the Computer Vision—ECCV 2016, Amsterdam, The Netherlands, 11–14 October 2016; Leibe, B., Matas, J., Sebe, N., Welling, M., Eds.; Springer International Publishing: Cham, Switzerland, 2016; pp. 630–645. [Google Scholar]
- Xie, S.; Girshick, R.; Dollár, P.; Tu, Z.; He, K. Aggregated Residual Transformations for Deep Neural Networks. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 5987–5995. [Google Scholar] [CrossRef] [Scilit]
- ISPRS 2D Semantic Labeling Contest. 2016. Available online: https://www2.isprs.org/commissions/comm2/wg4/benchmark/2d-sem-label-potsdam/ (accessed on 15 December 2020).
- Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. ImageNet large scale visual recognition challenge. Int. J. Comput. Vis. 2015, 115, 211–252. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the International Conference on Computer Vision (CVPR), Las Condes, Chile, 11–18 December 2015; pp. 1026–1034. [Google Scholar]
- Ioffe, S.; Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the 32nd International Conference on International Conference on Machine Learning, Lille, France, 7–9 July 2015; Volume 37, pp. 448–456. [Google Scholar]
- Bulo, S.R.; Porzi, L.; Kontschieder, P. In-place activated batchnorm for memory-optimized training of DNNs. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 5639–5647. [Google Scholar]
- Lee, C.-Y.; Xie, S.; Gallagher, P.; Zhang, Z.; Tu, Z. Deeply-supervised nets. In Proceedings of the Eighteenth International Conference on Artificial Intelligence and Statistics, San Diego, CA, USA, 9–12 May 2015; Lebanon, G., Vishwanathan, S.V.N., Eds.; PMLR: San Diego, CA, USA, 2015; Volume 38, pp. 562–570. [Google Scholar]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, Vancouver, BC, Canada, 8–14 December 2019; Wallach, H.M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E.B., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2019; pp. 8024–8035. [Google Scholar]
- Oktay, O.; Schlemper, J.; Folgoc, L.L.; Lee, M.C.H.; Heinrich, M.P.; Misawa, K.; Mori, K.; McDonagh, S.G.; Hammerla, N.Y.; Kainz, B.; et al. Attention U-Net: Learning where to look for the pancreas. arXiv 2018, arXiv:1804.03999. [Google Scholar]










| Method | 3 | 2 | 1 | mIoU (%) | ∆α (%) |
|---|---|---|---|---|---|
| Baseline | 74.1 | — | |||
| MLCR | ✓ | 74.8 | 0.7 ↑ | ||
| MLCR | ✓ | 75.0 | 0.9 ↑ | ||
| MLCR | ✓ | 75.4 | 1.3 ↑ | ||
| MLCR | ✓ | ✓ | 75.7 | 1.6 ↑ | |
| MLCR | ✓ | ✓ | ✓ | 76.0 | 1.9 ↑ |
| Method | S | I | L | mIoU (%) | ∆α (%) |
|---|---|---|---|---|---|
| Baseline | — | — | — | 74.1 | — |
| ✓ | 75.3 | 1.2 ↑ | |||
| ✓ | 75.4 | 1.3 ↑ | |||
| ✓ | 75.0 | 0.9 ↑ | |||
| ✓ | ✓ | 75.6 | 1.5 ↑ | ||
| ✓ | ✓ | 75.5 | 1.4 ↑ | ||
| ✓ | ✓ | 75.6 | 1.5 ↑ | ||
| ✓ | ✓ | ✓ | 76.0 | 1.9 ↑ |
| Method | mIoU (%) | FLOPs (G) |
|---|---|---|
| k = 1 | 75.5 | 42.7 |
| k = 3 | 76.0 | 43.3 |
| k = 5 | 75.8 | 44.3 |
| k = 7 | 75.8 | 45.9 |
| Method | Memory (Mb) | Parameter (M) | FLOPs (G) | FPS | mIoU (%) |
|---|---|---|---|---|---|
| Baseline | 915 | 25.2 | 42.7 | 90.3 | 74.1 |
| MCT | 1170 (+255) | 27.3 (+2.1) | 50.7 (+8.0) | 63.6 (−26.7) | 75.8 (+1.7) |
| Efficient MCT | 917 (+2) | 25.7 (+0.5) | 43.3 (+0.6) | 80.0 (−10.3) | 76.0 (+1.9) |
| Model | Backbone | Stride | mIoU (%) | Acc (%) | F1 | Parameter (M) | FLOPs (G) |
|---|---|---|---|---|---|---|---|
| FCN [11] | ResNet50 | 16× | 72.5 | 83.0 | 83.5 | 32.9 | 33.7 |
| OCRNet [29] | ResNet50 | 16× | 73.9 | 84.0 | 84.4 | 39.0 | 47.6 |
| CCNet [44] | ResNet50 | 16× | 74.1 | 84.1 | 84.6 | 47.4 | 57.4 |
| ISANet [43] | ResNet50 | 16× | 74.5 | 84.5 | 84.8 | 40.0 | 49.5 |
| PSPNet [45] | ResNet50 | 16× | 74.5 | 84.2 | 84.8 | 46.6 | 52.0 |
| ACFNet [30] | ResNet50 | 16× | 74.7 | 84.3 | 84.9 | 30.1 | 39.3 |
| DANet [22] | ResNet50 | 16× | 74.9 | 84.4 | 85.1 | 47.4 | 198.1 |
| DepLabV3+ [15] | ResNet50 | 16× | 75.1 | 84.7 | 85.1 | 40.3 | 69.3 |
| MANet [42] | ResNet50 | 16× | 75.2 | 84.7 | 85.2 | 33.5 | 49.6 |
| AttUNet [59] | ResNet50 | 16× | 75.3 | 84.6 | 85.3 | 96.5 | 207.8 |
| SFNet [17] | ResNet50 | 16× | 75.4 | 84.9 | 85.4 | 30.6 | 100.1 |
| ISNet [32] | ResNet50 | 16× | 75.7 | 85.0 | 85.6 | 44.5 | 58.8 |
| SCARF [31] | ResNet50 | 16× | 75.7 | 85.3 | 85.6 | 25.9 | 45.0 |
| Ours | ResNet50 | 16× | 76.0 | 85.2 | 85.8 | 25.7 | 43.3 |
| Model | Imp.sur | Building | Low.veg | Tree | Car | Clutter | mIoU(%) |
|---|---|---|---|---|---|---|---|
| FCN [11] | 79.8 | 90.3 | 70.6 | 72.6 | 72.2 | 49.8 | 72.5 |
| OCRNet [29] | 80.9 | 90.9 | 71.6 | 73.5 | 74.7 | 51.9 | 73.9 |
| CCNet [44] | 81.1 | 91.5 | 71.9 | 73.3 | 75.6 | 51.3 | 74.1 |
| ISANet [43] | 81.2 | 91.5 | 72.4 | 74.1 | 74.7 | 52.8 | 74.5 |
| PSPNet [45] | 81.4 | 91.3 | 72.1 | 74.1 | 75.4 | 52.5 | 74.5 |
| ACFNet [30] | 81.3 | 91.4 | 71.5 | 73.4 | 79.4 | 51.0 | 74.7 |
| DANet [22] | 81.7 | 91.5 | 72.0 | 74.4 | 76.4 | 53.2 | 74.9 |
| DepLabV3+ [15] | 81.5 | 91.4 | 72.0 | 73.1 | 80.9 | 51.4 | 75.1 |
| MANet [42] | 81.6 | 91.1 | 72.2 | 73.8 | 81.7 | 50.6 | 75.2 |
| AttUNet [59] | 81.6 | 91.3 | 71.9 | 73.1 | 81.4 | 52.3 | 75.3 |
| SFNet [17] | 81.9 | 91.5 | 72.5 | 73.7 | 81.0 | 51.8 | 75.4 |
| ISNet [32] | 82.1 | 91.7 | 72.7 | 74.3 | 81.1 | 52.1 | 75.7 |
| SCARF [31] | 82.1 | 91.5 | 72.8 | 74.1 | 81.4 | 52.1 | 75.7 |
| Ours | 82.3 | 91.4 | 73.1 | 73.7 | 81.6 | 53.7 | 76.0 |
| Model | Backbone | Stride | mIoU (%) | Acc (%) | F1 |
|---|---|---|---|---|---|
| FCN [11] | ResNet50 | 16× | 64.6 | 74.7 | 77.1 |
| CCNet [44] | ResNet50 | 16× | 65.5 | 75.2 | 77.7 |
| OCRNet [29] | ResNet50 | 16× | 66.3 | 76.5 | 78.6 |
| ISNet [32] | ResNet50 | 16× | 66.4 | 76.7 | 78.6 |
| ISANet [43] | ResNet50 | 16× | 66.6 | 76.4 | 78.7 |
| PSPNet [45] | ResNet50 | 16× | 66.6 | 76.0 | 78.6 |
| ACFNet [30] | ResNet50 | 16× | 66.7 | 76.4 | 78.7 |
| DANet [22] | ResNet50 | 16× | 66.8 | 76.4 | 78.8 |
| DepLabV3+ [15] | ResNet50 | 16× | 66.9 | 76.4 | 78.8 |
| MANet [42] | ResNet50 | 16× | 66.9 | 76.2 | 78.8 |
| AttUNet [59] | ResNet50 | 16× | 67.1 | 76.4 | 79.0 |
| Ours | ResNet50 | 16× | 68.1 | 77.5 | 79.8 |
| Model | Imp.sur | Buildings | Low.veg | Tree | Car | Clutter | mIoU (%) |
|---|---|---|---|---|---|---|---|
| FCN [11] | 78.9 | 86.1 | 63.8 | 72.8 | 49.9 | 36.0 | 64.6 |
| CCNet [44] | 80.1 | 86.7 | 65.0 | 73.5 | 52.5 | 35.3 | 65.5 |
| OCRNet [29] | 79.6 | 86.5 | 64.6 | 73.5 | 54.1 | 39.4 | 66.3 |
| ISNet [32] | 79.8 | 86.1 | 63.8 | 72.9 | 58.8 | 36.9 | 66.4 |
| ACFNet [30] | 80.6 | 87.1 | 65.2 | 74.1 | 57.8 | 35.3 | 66.7 |
| DANet [22] | 80.1 | 86.4 | 65.3 | 73.8 | 59.4 | 36.0 | 66.8 |
| DepLabV3+ [15] | 80.4 | 86.5 | 64.3 | 73.7 | 61.3 | 35.2 | 66.9 |
| MANet [42] | 80.3 | 86.5 | 64.1 | 73.5 | 63.4 | 33.7 | 66.9 |
| AttUNet [59] | 80.4 | 86.6 | 64.3 | 73.7 | 63.2 | 34.4 | 67.1 |
| Ours | 81.3 | 87.2 | 65.4 | 74.3 | 64.4 | 36.1 | 68.1 |
Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. |
© 2022 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Huang, Z.; Zhang, Q.; Zhang, G. MLCRNet: Multi-Level Context Refinement for Semantic Segmentation in Aerial Images. Remote Sens. 2022, 14, 1498. https://doi.org/10.3390/rs14061498
Huang Z, Zhang Q, Zhang G. MLCRNet: Multi-Level Context Refinement for Semantic Segmentation in Aerial Images. Remote Sensing. 2022; 14(6):1498. https://doi.org/10.3390/rs14061498
Chicago/Turabian StyleHuang, Zhifeng, Qian Zhang, and Guixu Zhang. 2022. "MLCRNet: Multi-Level Context Refinement for Semantic Segmentation in Aerial Images" Remote Sensing 14, no. 6: 1498. https://doi.org/10.3390/rs14061498
APA StyleHuang, Z., Zhang, Q., & Zhang, G. (2022). MLCRNet: Multi-Level Context Refinement for Semantic Segmentation in Aerial Images. Remote Sensing, 14(6), 1498. https://doi.org/10.3390/rs14061498

