Mitigating Class Imbalance and False-Negative Supervision in Remote Sensing Semantic Segmentation Using Object-Centric Patch Sampling
Highlights
- Object-centric patch sampling anchors each training patch to the geometric centroid of an annotated object, ensuring that every anchored patch contains at least one target instance. This keeps sampling concentrated around verified annotations and reduces the risk of false-negative supervision from incompletely labeled regions.
- The method was tested on satellite, aerial, and UAV imagery using both U-Net and DeepLabV3+. It consistently achieved higher IoU and F1-scores than sliding-window and random sampling across all configurations, with IoU improvements of up to 19.6 percentage points.
- The results show that how training patches are selected can be as important as, or even more important than, the number of patches used. Patch sampling should therefore be considered an important part of the segmentation methodology rather than simply a preprocessing step.
- The method does not require changes to the network architecture, loss function, or training procedure. Because it only changes how the training data are sampled, it can be incorporated into existing segmentation workflows with little to no additional computational cost.
Abstract
1. Introduction
1.1. Related Works
1.1.1. Patch-Based Training in Remote Sensing
1.1.2. Class Imbalance in Remote Sensing Segmentation
1.1.3. Label Noise and Negative Learning
1.1.4. Data-Centric Learning
1.1.5. Distinction from Existing Foreground-Aware Sampling
2. Methodology
2.1. Problem Formulation
2.2. Conventional Patch Sampling Strategies
2.2.1. Sliding-Window Sampling
2.2.2. Random Patch Sampling
2.3. Object-Centric Patch Sampling
2.3.1. Conceptual Overview
2.3.2. Boundary and Edge Handling
2.3.3. Handling Dense or Overlapping Objects
2.3.4. Mitigation of Negative Learning from Incomplete Labels
2.3.5. Controlled Background Integration
2.4. Experimental Setup
- Cross-architecture and multi-class experiments (DeepLabV3+): To test whether the sampling strategy remains effective under a different architecture and in a multi-class setting, further experiments were conducted using a DeepLabV3+ network equipped with an ImageNet-pretrained ResNet-50 backbone (Figure 6) [24]. This suite evaluated both single-class target detection (cotton) and a multi-class joint segmentation task where scene patches contained co-occurring cotton, water, and background classes.
3. Results
3.1. Sentinel-2 Cotton Field Segmentation
3.2. NAIP Building Segmentation
3.3. UAV Water Body Segmentation
3.4. Multi-Model and Multi-Class Segmentation Validation
3.4.1. Cross-Architecture Validation (DeepLabV3+ vs. U-Net)
3.4.2. Multi-Class Joint Segmentation
4. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Krejcar, O.; Namazi, H. AI in Remote Sensing and Satellite Image Processing—A Review. Environ. Earth Sci. 2026, 85, 78. [Google Scholar] [CrossRef] [Scilit]
- Lu, W.; Shi, X.; Lu, Z. A New Two-Step Road Extraction Method in High Resolution Remote Sensing Images. PLoS ONE 2024, 19, e0305933. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, X.; Cai, F.; Tian, J.; Williams, T.K.-A. Spatiotemporal Fusion of Multisource Remote Sensing Data: Literature Survey, Taxonomy, Principles, Applications, and Future Directions. Remote Sens. 2018, 10, 527. [Google Scholar] [CrossRef] [Scilit]
- Liu, Q.; Huang, T.; Dong, Y.; Yang, J.; Xiang, W. From Pixels to Images: Deep Learning Advances in Remote Sensing Image Semantic Segmentation. arXiv 2025, arXiv:2505.15147. [Google Scholar] [CrossRef] [Scilit]
- Ji, S.; Wei, S.; Lu, M. Fully Convolutional Networks for Multisource Building Extraction from an Open Aerial and Satellite Imagery Data Set. IEEE Trans. Geosci. Remote Sens. 2019, 57, 574–586. [Google Scholar] [CrossRef] [Scilit]
- Silwal, A.; Subedi, A.; Tamrakar, R.; Dahal, K.; Dahal, D.; Ekpetere, K.O.; Zhran, M. A Comprehensive Review of Machine Learning and Deep Learning Methods for Flood Inundation Mapping. Earth 2026, 7, 44. [Google Scholar] [CrossRef] [Scilit]
- Cao, H.; Tian, Y.; Liu, Y.; Wang, R. Water Body Extraction from High Spatial Resolution Remote Sensing Images Based on Enhanced U-Net and Multi-Scale Information Fusion. Sci. Rep. 2024, 14, 16132. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, P.; Ke, Y.; Zhang, Z.; Wang, M.; Li, P.; Zhang, S. Urban Land Use and Land Cover Classification Using Novel Deep Learning Models Based on High Spatial Resolution Satellite Imagery. Sensors 2018, 18, 3717. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015; Navab, N., Hornegger, J., Wells, W.M., Frangi, A.F., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2015; Volume 9351, pp. 234–241. [Google Scholar] [CrossRef] [Scilit]
- Lian, R.; Huang, L. DeepWindow: Sliding Window Based on Deep Learning for Road Extraction from Remote Sensing Images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2020, 13, 1905–1916. [Google Scholar] [CrossRef] [Scilit]
- Siva, S.S.; Cross-Zamirski, J.O. Building Damage Detection Using Satellite Images and Patch-Based Transformer Methods. arXiv 2026, arXiv:2602.08117. [Google Scholar] [CrossRef] [Scilit]
- Buda, M.; Maki, A.; Mazurowski, M.A. A Systematic Study of the Class Imbalance Problem in Convolutional Neural Networks. Neural Netw. 2018, 106, 249–259. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xia, W.; Ma, C.; Liu, J.; Liu, S.; Chen, F.; Yang, Z.; Duan, J. High-Resolution Remote Sensing Imagery Classification of Imbalanced Data Using Multistage Sampling Method and Deep Neural Networks. Remote Sens. 2019, 11, 2523. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Z.; Zheng, C.; Liu, X.; Tian, Y.; Chen, X.; Chen, X.; Dong, Z. A Dynamic Effective Class Balanced Approach for Remote Sensing Imagery Semantic Segmentation of Imbalanced Data. Remote Sens. 2023, 15, 1768. [Google Scholar] [CrossRef] [Scilit]
- Garcia-Garcia, A.; Orts-Escolano, S.; Oprea, S.; Villena-Martinez, V.; Martinez-Gonzalez, P.; Garcia-Rodriguez, J. A Survey on Deep Learning Techniques for Image and Video Semantic Segmentation. Appl. Soft Comput. 2018, 70, 41–65. [Google Scholar] [CrossRef] [Scilit]
- Zhu, X.X.; Tuia, D.; Mou, L.; Xia, G.-S.; Zhang, L.; Xu, F.; Fraundorfer, F. Deep Learning in Remote Sensing: A Comprehensive Review and List of Resources. IEEE Geosci. Remote Sens. Mag. 2017, 5, 8–36. [Google Scholar] [CrossRef] [Scilit]
- Rolnick, D.; Veit, A.; Belongie, S.; Shavit, N. Deep Learning Is Robust to Massive Label Noise. arXiv 2018, arXiv:1705.10694. [Google Scholar] [CrossRef] [Scilit]
- Northcutt, C.G.; Jiang, L.; Chuang, I.L. Confident Learning: Estimating Uncertainty in Dataset Labels. J. Artif. Intell. Res. 2021, 70, 1373–1411. [Google Scholar] [CrossRef] [Scilit]
- Kim, Y.; Yim, J.; Yun, J.; Kim, J. NLNL: Negative Learning for Noisy Labels. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 101–110. [Google Scholar] [CrossRef] [Scilit]
- Roth, K.; Milbich, T.; Sinha, S.; Gupta, P.; Ommer, B.; Cohen, J.P. Revisiting Training Strategies and Generalization Performance in Deep Metric Learning. In Proceedings of the 37th International Conference on Machine Learning (ICML 2020), Virtual, 13–18 July 2020; Proceedings of Machine Learning Research; PMLR: Cambridge, MA, USA, 2020; Volume 119, pp. 8242–8252. Available online: https://proceedings.mlr.press/v119/roth20a.html (accessed on 4 August 2026).
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2999–3007. [Google Scholar] [CrossRef] [Scilit]
- He, H.; Garcia, E.A. Learning from Imbalanced Data. IEEE Trans. Knowl. Data Eng. 2009, 21, 1263–1284. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.; Albrecht, C.M.; Wang, Y.; Zhu, X.X. AIO2: Online Correction of Object Labels for Deep Learning with Incomplete Annotation in Remote Sensing Image Segmentation. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5613917. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.-C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 801–818. [Google Scholar]
- Dalal, N.; Triggs, B. Histograms of Oriented Gradients for Human Detection. In Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR), San Diego, CA, USA, 20–25 June 2005; IEEE: Piscataway, NJ, USA, 2005; Volume 1, pp. 886–893. [Google Scholar] [CrossRef] [Scilit]
- Jarrahi, M.H.; Memariani, A.; Guha, S. The Principles of Data-Centric AI. Commun. ACM 2023, 66, 84–92. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bengio, Y. Practical Recommendations for Gradient-Based Training of Deep Architectures. In Neural Networks: Tricks of the Trade; Montavon, G., Orr, G.B., Müller, K.-R., Eds.; Springer: Berlin/Heidelberg, Germany, 2012; Volume 7700, pp. 437–478. [Google Scholar] [CrossRef] [Scilit]
- Ren, P.; Xiao, Y.; Chang, X.; Huang, P.-Y.; Li, Z.; Gupta, B.B.; Chen, X.; Wang, X. A Survey of Deep Active Learning. ACM Comput. Surv. 2021, 54, 1–40. [Google Scholar] [CrossRef] [Scilit]
- MMSegmentation Contributors. MMSegmentation: OpenMMLab Semantic Segmentation Toolbox and Benchmark. Available online: https://github.com/open-mmlab/mmsegmentation (accessed on 15 August 2026).
- Kaiser, P.; Wegner, J.D.; Lucchi, A.; Jaggi, M.; Hofmann, T.; Schindler, K. Learning Aerial Image Segmentation from Online Maps. IEEE Trans. Geosci. Remote Sens. 2017, 55, 6054–6068. [Google Scholar] [CrossRef] [Scilit]
- Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
- Braden, B. The Surveyor’s Area Formula. Coll. Math. J. 1986, 17, 326–337. [Google Scholar] [CrossRef]
- Preparata, F.P.; Shamos, M.I. Computational Geometry: An Introduction; Springer: New York, NY, USA, 1985. [Google Scholar] [CrossRef] [Scilit]
- Rouault, E.; Warmerdam, F.; Schwehr, K.; Kiselev, A.; Butler, H.; Łoskot, M.; Szekeres, T.; Tourigny, E.; Landa, M.; Miara, I.; et al. GDAL, version 3.13.0; Zenodo: Geneva, Switzerland, 2026. [Google Scholar] [CrossRef]
- Drusch, M.; Del Bello, U.; Carlier, S.; Colin, O.; Fernandez, V.; Gascon, F.; Hoersch, B.; Isola, C.; Laberinti, P.; Martimort, P.; et al. Sentinel-2: ESA’s Optical High-Resolution Mission for GMES Operational Services. Remote Sens. Environ. 2012, 120, 25–36. [Google Scholar] [CrossRef] [Scilit]
- Schroeder, T.A.; Obata, S.; Papeş, M.; Branoff, B. Evaluating Statewide NAIP Photogrammetric Point Clouds for Operational Improvement of National Forest Inventory Estimates in Mixed Hardwood Forests of the Southeastern U.S. Remote Sens. 2022, 14, 4386. [Google Scholar] [CrossRef] [Scilit]
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. In Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015), San Diego, CA, USA, 7–9 May 2015. [Google Scholar] [CrossRef] [Scilit]
- Prechelt, L. Early Stopping—But When? In Neural Networks: Tricks of the Trade; Orr, G.B., Müller, K.-R., Eds.; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 1998; Volume 1524, pp. 55–69. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Milletari, F.; Navab, N.; Ahmadi, S.-A. V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation. In Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA, 25–28 October 2016; pp. 565–571. [Google Scholar] [CrossRef] [Scilit]
- Everingham, M.; Van Gool, L.; Williams, C.K.I.; Winn, J.; Zisserman, A. The PASCAL Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef] [Scilit]
- Rolnick, D.; Donti, P.L.; Kaack, L.H.; Kochanski, K.; Lacoste, A.; Sankaran, K.; Ross, A.S.; Milojevic-Dupont, N.; Jaques, N.; Waldman-Brown, A.; et al. Tackling Climate Change with Machine Learning. ACM Comput. Surv. 2022, 55, 42. [Google Scholar] [CrossRef] [Scilit]
- Xiao, Y.; Zhao, Y.; Shu, K. Understanding and Tackling Label Errors in Individual-Level Nature Language Understanding. arXiv 2025, arXiv:2502.13297. [Google Scholar] [CrossRef] [Scilit]





















| Sampling Strategy | IoU | Dice + BCE Loss | F1-Score | Epoch | Dataset Size | Training Time (min) |
|---|---|---|---|---|---|---|
| Sliding-window sampling | 0.8930 | 0.4912 | 0.8595 | 23 | 5028 | 200.697 |
| Random sampling | 0.8687 | 0.6533 | 0.7565 | 35 | 1597 | 93.996 |
| Object-centric sampling | 0.9285 | 0.3207 | 0.9488 | 39 | 1597 | 108.271 |
| Sampling Strategy | IoU | Dice + BCE Loss | F1-Score | Epoch | Dataset Size | Training Time (min) |
|---|---|---|---|---|---|---|
| Sliding-window sampling | 0.6997 | 0.8092 | 0.5666 | 100 | 629 | 109.78 |
| Random sampling | 0.4904 | 1.0612 | 0.0181 | 30 | 600 | 31.31 |
| Object-centric sampling | 0.8960 | 0.1909 | 0.8936 | 82 | 1728 | 242.8 |
| Sampling Strategy | IoU | Dice + BCE Loss | F1-Score | Epoch | Dataset Size | Training Time (min) |
|---|---|---|---|---|---|---|
| Sliding-window sampling | 0.884 | 0.628 | 0.887 | 52 | 1508 | 2038.823 |
| Random sampling | 0.789 | 0.751 | 0.784 | 47 | 782 | 1045.661 |
| Object-centric sampling | 0.9121 | 0.5475 | 0.9181 | 50 | 782 | 1038.429 |
| Task Complexity | Sampling Strategy | mIoU/IoU | Dice + BCE Loss | F1-Score |
|---|---|---|---|---|
| Single-class (cotton only) | U-Net (baseline) | 0.9285 | 0.3207 | 0.9488 |
| Single-class (cotton only) | DeepLabV3+ | 0.9297 | 0.2065 | 0.9633 |
| Multi-class (cotton + water) | DeepLabV3+ | 0.9504 | 0.0725 | 0.9772 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Regmi, Y.; Gautam, S.; Parajuli, G.; Silwal, A.; Bhandari, R.; Acharya, T.D. Mitigating Class Imbalance and False-Negative Supervision in Remote Sensing Semantic Segmentation Using Object-Centric Patch Sampling. Remote Sens. 2026, 18, 2844. https://doi.org/10.3390/rs18162844
Regmi Y, Gautam S, Parajuli G, Silwal A, Bhandari R, Acharya TD. Mitigating Class Imbalance and False-Negative Supervision in Remote Sensing Semantic Segmentation Using Object-Centric Patch Sampling. Remote Sensing. 2026; 18(16):2844. https://doi.org/10.3390/rs18162844
Chicago/Turabian StyleRegmi, Yogesh, Sandeep Gautam, Gaurav Parajuli, Abinash Silwal, Roshan Bhandari, and Tri Dev Acharya. 2026. "Mitigating Class Imbalance and False-Negative Supervision in Remote Sensing Semantic Segmentation Using Object-Centric Patch Sampling" Remote Sensing 18, no. 16: 2844. https://doi.org/10.3390/rs18162844
APA StyleRegmi, Y., Gautam, S., Parajuli, G., Silwal, A., Bhandari, R., & Acharya, T. D. (2026). Mitigating Class Imbalance and False-Negative Supervision in Remote Sensing Semantic Segmentation Using Object-Centric Patch Sampling. Remote Sensing, 18(16), 2844. https://doi.org/10.3390/rs18162844

