A Multi-Scale Global Fusion-Based Method for Surface Fissure Extraction from UAV Imagery
Abstract
1. Introduction
2. Study Area and Dataset
2.1. Study Area
2.2. Dataset
3. Methods
3.1. MGF-UNet Architecture
3.2. MFS (Multi-Scale Feature Sensing)
3.3. EMA (Efficient Multi-Scale Attention)
3.4. TSCT (Token-Selective Context Transformer)
3.5. FiLM (Feature-Wise Linear Modulation)
3.6. AFF (Adaptive Feature Fusion)
4. Loss Functions and Accuracy Metrics
4.1. Loss Functions
4.2. Accuracy Metrics
5. Experimental Results
5.1. Experimental Details
5.2. Mining-Area Surface Fissure Dataset Evaluation and Large-Scale Real-World Fissure Extraction Validation
5.3. Road Crack Dataset Identification
5.3.1. Comparative Experiments on RFD
5.3.2. Comparative Experiments on the Crack500 Dataset
5.4. Model Complexity Comparison
5.5. Ablation Experiments
5.6. Generalization Study on Non-Crack Data
6. Discussion
6.1. General Interpretation
6.2. Limitations and Future Work
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Fu, Y.; Wu, Y.; Yin, X.; Zhang, Y. Mapping mining-induced ground fissures and their evolution using UAV photogrammetry. Front. Earth Sci. 2023, 11, 1260913. [Google Scholar] [CrossRef]
- Liu, Y.; Zhang, D.; Wang, G.; Liu, C.; Zhang, Y. Discrete element method-based prediction of areas prone to buried hill-controlled earth fissures. J. Zhejiang Univ. Sci. A 2019, 20, 794–803. [Google Scholar] [CrossRef]
- Wang, K.; Wei, B.; Zhao, T.; Wu, G.; Zhang, J.; Zhu, L.; Wang, L. An automated approach for mapping mining-induced fissures using CNNs and UAS photogrammetry. Remote Sens. 2024, 16, 2090. [Google Scholar] [CrossRef]
- Xu, D.; Zhao, Y.; Jiang, Y.; Zhang, C.; Sun, B.; He, X. Using improved edge detection method to detect mining-induced ground fissures identified by unmanned aerial vehicle remote sensing. Remote Sens. 2021, 13, 3652. [Google Scholar] [CrossRef]
- Zou, S.; Xiong, F.; Luo, H.; Lu, J.; Qian, Y. AF-Net: All-scale feature fusion network for road extraction from remote sensing images. In Proceedings of the Digital Image Computing: Techniques and Applications (DICTA), Gold Coast, Australia, 29 November–1 December 2021; pp. 1–8. [Google Scholar] [CrossRef]
- Montagnon, T.; Hollingsworth, J.; Pathier, E.; Marchandon, M.; Dalla Mura, M.; Giffard-Roisin, S. Sub-pixel optical satellite image registration for ground deformation using deep learning. In Proceedings of the IEEE International Conference on Image Processing (ICIP), Bordeaux, France, 16–19 October 2022; pp. 2716–2720. [Google Scholar] [CrossRef]
- Zhang, Z.; Zhu, L. A review on unmanned aerial vehicle remote sensing: Platforms, sensors, data processing methods, and applications. Drones 2023, 7, 398. [Google Scholar] [CrossRef]
- Cheng, Z.; Gong, W.; Jaboyedoff, M.; Chen, J.; Derron, M.-H.; Zhao, F. Landslide Identification in UAV Images Through Recognition of Landslide Boundaries and Ground Surface Cracks. Remote Sens. 2025, 17, 1900. [Google Scholar] [CrossRef]
- Cirillo, D.; Tangari, A.C.; Scarciglia, F.; Lavecchia, G.; Brozzetti, F. UAV-PPK Photogrammetry, GIS, and Soil Analysis to Estimate Long-Term Slip Rates on Active Faults in a Seismic Gap of Northern Calabria (Southern Italy). Remote Sens. 2025, 17, 3366. [Google Scholar] [CrossRef]
- Darmawan, H.; Walter, T.R.; Brotopuspito, K.S.; Nandaka, I.G.M.A. Morphological and structural changes at the Merapi lava dome monitored in 2012–15 using unmanned aerial vehicles (UAVs). J. Volcanol. Geotherm. Res. 2018, 349, 256–267. [Google Scholar] [CrossRef]
- Zhang, R.; Li, H.; Duan, K.; You, S.; Liu, K.; Wang, F.; Hu, Y. Automatic Detection of Earthquake-Damaged Buildings by Integrating UAV Oblique Photography and Infrared Thermal Imaging. Remote Sens. 2020, 12, 2621. [Google Scholar] [CrossRef]
- Zhang, F.; Hu, Z.; Fu, Y.; Yang, K.; Wu, Q.; Feng, Z. A new identification method for surface cracks from UAV images based on machine learning in coal mining areas. Remote Sens. 2020, 12, 1571. [Google Scholar] [CrossRef]
- Huangfu, W.; Qiu, H.; Cui, P.; Yang, D.; Liu, Y.; Ullah, M.; Kamp, U. Automated extraction of mining-induced ground fissures using deep learning and object-based image classification. Earth Surf. Process. Landf. 2024, 49, 2189–2204. [Google Scholar] [CrossRef]
- Zhao, J.; Zhang, D.; Shi, B.; Zhou, Y.; Chen, J.; Yao, R.; Xue, Y. Multi-source collaborative enhanced for remote sensing images semantic segmentation. Neurocomputing 2022, 493, 76–90. [Google Scholar] [CrossRef]
- Liu, Y.; Zhou, T.; Xu, J.; Hong, Y.; Pu, Q.; Wen, X. Rotating target detection method of concrete bridge crack based on YOLOv5. Appl. Sci. 2023, 13, 11118. [Google Scholar] [CrossRef]
- An, J.; Dong, S.; Wang, X.; Li, C.; Zhao, W. Research on UAV aerial imagery detection algorithm for mining-induced surface cracks based on improved YOLOv10. Sci. Rep. 2025, 15, 30101. [Google Scholar] [CrossRef] [PubMed]
- Li, Y.; Ouyang, S.; Zhang, Y. Combining deep learning and ontology reasoning for remote sensing image semantic segmentation. Knowl.-Based Syst. 2022, 243, 108469. [Google Scholar] [CrossRef]
- Zhang, Y.; Zhen, J.; Sun, S.; Liu, T.; Huo, L.; Wang, T. SCAFNet: A Semantic Compensated Adaptive Fusion Network for Remote Sensing Images Change Detection. IEEE Geosci. Remote Sens. Lett. 2026, 23, 6003405. [Google Scholar] [CrossRef]
- Zhang, Y.; Wang, T.; Xue, L.; Lian, W.; Tao, R. ORSI Salient Object Detection via Progressive Interaction and Saliency-Guided Enhancement. IEEE Geosci. Remote Sens. Lett. 2025, 23, 1–5. [Google Scholar] [CrossRef]
- Zou, Q.; Zhang, Z.; Li, Q.; Qi, X.; Wang, Q.; Wang, S. DeepCrack: Learning hierarchical convolutional features for crack detection. IEEE Trans. Image Process. 2019, 28, 1498–1512. [Google Scholar] [CrossRef]
- Alipour, M.; Harris, D.K.; Miller, G.R. Robust pixel-level crack detection using deep fully convolutional neural networks. J. Comput. Civ. Eng. 2019, 33, 04019040. [Google Scholar] [CrossRef]
- Zhu, W.; Guo, Y.; Chen, Y.; Gao, Z.; Zhang, H.; Du, S. Mining area ground fissure identification method based on AR-SE-ResUNet. IEEE Geosci. Remote Sens. Lett. 2025, 22, 3002905. [Google Scholar] [CrossRef]
- Hu, H.; Guo, X.; Xiao, J. Ground fissures identification in mining area based on OrientFuse-Net. IEEE Geosci. Remote Sens. Lett. 2025, 22, 2501205. [Google Scholar]
- Jiang, X.; Mao, S.; Li, M.; Liu, H.; Zhang, H.; Fang, S.; Yuan, M.; Zhang, C. MFPA-Net: An efficient deep learning network for automatic ground fissures extraction in UAV images of coal mining areas. Int. J. Appl. Earth Obs. Geoinf. 2022, 114, 103039. [Google Scholar] [CrossRef]
- Chen, P.; Li, P.; Wang, B.; Ding, X.; Zhang, Y.; Zhang, T.; Yu, T. GFSegNet: A multi-scale segmentation model for mining area ground fissures. Int. J. Appl. Earth Obs. Geoinf. 2024, 128, 103788. [Google Scholar] [CrossRef]
- Tao, T.; Han, K.; Yao, X.; Chen, X.; Wu, Z.; Yao, C.; Tian, X.; Zhou, Z.; Ren, K. Identification of ground fissure development induced by coal mining using UAV images and deep learning. Remote Sens. 2024, 16, 1046. [Google Scholar] [CrossRef]
- Meng, J.; Xu, X.; Li, P.; Zhang, Z.; Zhao, W.; Ren, J.; Li, Y. GF-Former: An accurate UAV-based remote sensing image network for high-precision segmentation of ground fissures. Int. J. Mach. Learn. Cybern. 2025, 16, 10097–10118. [Google Scholar] [CrossRef]
- Sun, Z.; Cao, S.; Yang, Y.; Kitani, K. Rethinking transformer-based set prediction for object detection. In Proceedings of the IEEE/CVF ICCV, Montreal, QC, Canada, 11–17 October 2021; pp. 3591–3600. [Google Scholar]
- Ranftl, R.; Bochkovskiy, A.; Koltun, V. Vision transformers for dense prediction. In Proceedings of the IEEE/CVF ICCV, Montreal, QC, Canada, 11–17 October 2021; pp. 12159–12168. [Google Scholar]
- Li, C.; Wang, J.; Liu, G.; Cao, Q.; Wang, S. Rare metal elements in the primary coal seam of Changzhi coalfield in Southeast Shanxi, China. Energy Explor. Exploit. 2020, 38, 1140–1158. [Google Scholar] [CrossRef]
- Shao, L.Y.; Hou, H.H.; Shao, L.; Yang, Z.; Shang, X.; Xiao, Z.; Wang, S.; Zhang, W.; Zheng, M.; Lu, J. Lithofacies palaeogeography of the Carboniferous and Permian in the Qinshui basin, Shanxi Province, China. J. Palaeogeogr. 2015, 4, 384–412. [Google Scholar] [CrossRef]
- Shao, L.Y.; Wang, H.; Lu, J.; He, Z.; Wang, H.; Zhang, P. Permo-Carboniferous coal measures in the Qinshui basin: Lithofacies paleogeography and its control on coal accumulation. Front. Earth Sci. China 2007, 1, 106–115. [Google Scholar] [CrossRef]
- Liu, B.; Chang, S.; Zhang, S.; Li, Y.; Yang, Z.; Liu, Z.; Chen, Q. Seismic-Geological Integrated Study on Sedimentary Evolution and Peat Accumulation Regularity of the Shanxi Formation in Xinjing Mining Area, Qinshui Basin. Energies 2022, 15, 1851. [Google Scholar] [CrossRef]
- Ouyang, D.; He, S.; Zhang, G.; Luo, M.; Guo, H.; Zhan, J.; Huang, Z. Efficient multi-scale attention module with cross-spatial learning. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; pp. 1–5. [Google Scholar]
- Perez, E.; Strub, F.; de Vries, H.; Dumoulin, V.; Courville, A. FiLM: Visual reasoning with a general conditioning layer. Proc. AAAI Conf. Artif. Intell. 2018, 32, 3942–3951. [Google Scholar] [CrossRef]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv 2017, arXiv:1704.04861. [Google Scholar] [CrossRef]
- Shang, Q.; Wang, G.; Wang, X.; Li, Y.; Wang, H. S-Net: A novel shallow network for enhanced detail retention in medical image segmentation. Comput. Methods Programs Biomed. 2025, 265, 108730. [Google Scholar] [CrossRef] [PubMed]
- Huang, C.; Wang, M.; Zhu, Z.; Li, Y. High-Resolution Remote Sensing Imagery Water Body Extraction Using a U-Net with Cross-Layer Multi-Scale Attention Fusion. Sensors 2025, 25, 5655. [Google Scholar] [CrossRef] [PubMed]
- Jadon, S. A survey of loss functions for semantic segmentation. In Proceedings of the IEEE CIBCB, Viña del Mar, Chile, 27–29 October 2020; pp. 1–7. [Google Scholar]
- Roy, A.G.; Navab, N.; Wachinger, C. Recalibrating fully convolutional networks with spatial and channel squeeze & excitation blocks. arXiv 2018, arXiv:1808.08127. [Google Scholar] [CrossRef]
- Fedus, W.; Zoph, B.; Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv 2021, arXiv:2101.03961. [Google Scholar]
- Chen, L.; Fu, Y.; Gu, L.; Yan, C.; Harada, T.; Huang, G. Frequency-aware feature fusion for dense image prediction. arXiv 2024, arXiv:2408.12879. [Google Scholar] [CrossRef]
- Dai, Y.; Gieseke, F.; Oehmcke, S.; Wu, Y.; Barnard, K. Attentional feature fusion. In Proceedings of the IEEE WACV, Waikoloa, HI, USA, 3–8 January 2021; pp. 3559–3568. [Google Scholar]
- Gao, F.; Fu, M.; Cao, J.; Dong, J.; Du, Q. Adaptive frequency enhancement network for remote sensing image semantic segmentation. arXiv 2025, arXiv:2504.02647. [Google Scholar] [CrossRef]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 318–327. [Google Scholar] [CrossRef]
- Abraham, N.; Khan, N.M. Multimodal segmentation with MGF-Net and the focal Tversky loss function. In Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries; Springer: Cham, Switzerland, 2020; pp. 191–198. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In Proceedings of the MICCAI, Munich, Germany, 5–9 October 2015; pp. 234–241. [Google Scholar]
- Chen, L.-C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder–decoder with atrous separable convolution for semantic image segmentation. In Computer Vision–ECCV 2018; Springer: Cham, Switzerland, 2018; pp. 833–851. [Google Scholar]
- Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
- Cao, H.; Wang, Y.; Chen, J.; Jiang, D.; Zhang, X.; Tian, Q.; Wang, M. Swin-Unet: Unet-like pure transformer for medical image segmentation. arXiv 2021, arXiv:2105.05537. [Google Scholar]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the IEEE CVPR, Honolulu, HI, USA, 21–26 July 2017; pp. 6230–6239. [Google Scholar]
- Jonnalagadda, A.V.; Hashim, H.A. SegNet: A segmented deep learning-based convolutional neural network approach for drones wildfire detection. Remote Sens. Appl. Soc. Environ. 2024, 34, 101181. [Google Scholar] [CrossRef]
- Amit, T.; Shaharbany, T.; Nachmani, E.; Wolf, L. SegDiff: Image segmentation with diffusion probabilistic models. arXiv 2022, arXiv:2211.13202. [Google Scholar] [CrossRef]
- Zhang, Y.; Liu, C. Crack segmentation using discrete cosine transform in shadow environments. Autom. Constr. 2024, 166, 105646. [Google Scholar] [CrossRef]
- Munawar, H.S.; Ullah, F.; Shahzad, D.; Heravi, A.; Qayyum, S.; Akram, J. Image-based crack detection methods: A review. Infrastructures 2021, 6, 115. [Google Scholar] [CrossRef]
- Yang, F.; Zhang, L.; Yu, S.; Prokhorov, D.; Mei, X.; Ling, H. Feature pyramid and hierarchical boosting network for pavement crack detection. arXiv 2019, arXiv:1901.06340. [Google Scholar] [CrossRef]
- International Society for Photogrammetry and Remote Sensing (ISPRS). ISPRS 2D Semantic Labeling Dataset (Vaihingen). Available online: https://www.isprs.org/resources/datasets/benchmarks/UrbanSemLab/2d-sem-label-vaihingen.aspx (accessed on 15 January 2026).















| Predicted Results | Fissures | Non-Fissures | |
|---|---|---|---|
| Ground Truth | |||
| Fissures | TP | FN | |
| Non-Fissures | FP | TN |
| ID | Method | Precision | Recall | Dice | IoU | BFI |
|---|---|---|---|---|---|---|
| 1 | U-Net | 0.734 | 0.839 | 0.782 | 0.642 | 0.847 |
| 2 | DeepLabV3+ | 0.759 | 0.849 | 0.801 | 0.669 | 0.827 |
| 3 | Segformer | 0.749 | 0.838 | 0.791 | 0.654 | 0.825 |
| 4 | SwinUNet | 0.726 | 0.826 | 0.773 | 0.630 | 0.766 |
| 5 | HRNet | 0.657 | 0.833 | 0.734 | 0.580 | 0.758 |
| 6 | PSP-Net | 0.713 | 0.847 | 0.775 | 0.633 | 0.788 |
| 7 | SegNet | 0.743 | 0.826 | 0.782 | 0.642 | 0.832 |
| 8 | SegDiff | 0.801 | 0.761 | 0.780 | 0.640 | 0.798 |
| 9 | MGF-UNet | 0.782 | 0.844 | 0.814 | 0.686 | 0.866 |
| ID | Method | Precision | Recall | Dice | IoU | BFI |
|---|---|---|---|---|---|---|
| 1 | U-Net | 0.831 | 0.946 | 0.885 | 0.794 | 0.818 |
| 2 | PSP-Net | 0.791 | 0.958 | 0.866 | 0.763 | 0.786 |
| 3 | DeepLabV3+ | 0.821 | 0.947 | 0.879 | 0.784 | 0.819 |
| 4 | Segformer | 0.825 | 0.951 | 0.875 | 0.778 | 0.820 |
| 5 | SwinUNet | 0.802 | 0.922 | 0.865 | 0.762 | 0.781 |
| 6 | HRNet | 0.762 | 0.941 | 0.838 | 0.722 | 0.659 |
| 7 | SegNet | 0.818 | 0.947 | 0.875 | 0.778 | 0.731 |
| 8 | SegDiff | 0.829 | 0.955 | 0.888 | 0.798 | 0.786 |
| 9 | MGF-UNet | 0.841 | 0.957 | 0.894 | 0.808 | 0.851 |
| ID | Method | Precision | Recall | Dice | IoU | BFI |
|---|---|---|---|---|---|---|
| 1 | U-Net | 0.641 | 0.834 | 0.725 | 0.569 | 0.635 |
| 2 | DeepLabV3+ | 0.653 | 0.861 | 0.743 | 0.591 | 0.639 |
| 3 | Segformer | 0.669 | 0.836 | 0.742 | 0.590 | 0.610 |
| 4 | SwinUNet | 0.645 | 0.843 | 0.730 | 0.575 | 0.635 |
| 5 | HRNet | 0.636 | 0.847 | 0.727 | 0.572 | 0.573 |
| 6 | SegNet | 0.621 | 0.826 | 0.709 | 0.550 | 0.627 |
| 7 | PSP-Net | 0.646 | 0.862 | 0.738 | 0.585 | 0.611 |
| 8 | SegDiff | 0.652 | 0.718 | 0.684 | 0.520 | 0.674 |
| 9 | Ours | 0.697 | 0.842 | 0.763 | 0.617 | 0.732 |
| Method | Input Size | Params (M) | Time (ms) | FPS | FLOPs (G) |
|---|---|---|---|---|---|
| DeepLabV3+ | 5.812 | 21.437 | 47.20 | 6.587 | |
| HRNet | 28.536 | 42.125 | 23.74 | 10.259 | |
| PSP-Net | 46.582 | 8.744 | 114.37 | 44.382 | |
| SegFormer | 13.678 | 7.106 | 140.72 | 3.454 | |
| U-Net | 14.789 | 7.618 | 131.27 | 32.050 | |
| SegNet | 29.440 | 6.362 | 157.18 | 40.072 | |
| SwinUNet | 22.114 | 10.385 | 96.29 | 4.952 | |
| SegDiff | 170.14 | 27400 | 0.04 | 13500 | |
| MGF-UNet | 22.270 | 8.752 | 114.26 | 4.002 |
| ID | Base | MFS | EMA | TSCT | FiLM | AFF | Precision | Recall | Dice | IoU | BF1 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ✓ | 0.734 | 0.839 | 0.782 | 0.642 | 0.847 | |||||
| 2 | ✓ | ✓ | 0.740 | 0.835 | 0.784 | 0.645 | 0.842 | ||||
| 3 | ✓ | ✓ | ✓ | 0.759 | 0.837 | 0.796 | 0.661 | 0.851 | |||
| 4 | ✓ | ✓ | ✓ | ✓ | 0.765 | 0.841 | 0.801 | 0.669 | 0.855 | ||
| 5 | ✓ | ✓ | ✓ | ✓ | ✓ | 0.763 | 0.853 | 0.805 | 0.675 | 0.857 | |
| 6 | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 0.782 | 0.844 | 0.814 | 0.686 | 0.866 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhou, M.; Ji, M.; Jin, F.; Zhang, Z.; Dou, F.; Fan, X. A Multi-Scale Global Fusion-Based Method for Surface Fissure Extraction from UAV Imagery. Sensors 2026, 26, 1440. https://doi.org/10.3390/s26051440
Zhou M, Ji M, Jin F, Zhang Z, Dou F, Fan X. A Multi-Scale Global Fusion-Based Method for Surface Fissure Extraction from UAV Imagery. Sensors. 2026; 26(5):1440. https://doi.org/10.3390/s26051440
Chicago/Turabian StyleZhou, Mingxi, Min Ji, Fengxiang Jin, Zhaomin Zhang, Fengke Dou, and Xiangru Fan. 2026. "A Multi-Scale Global Fusion-Based Method for Surface Fissure Extraction from UAV Imagery" Sensors 26, no. 5: 1440. https://doi.org/10.3390/s26051440
APA StyleZhou, M., Ji, M., Jin, F., Zhang, Z., Dou, F., & Fan, X. (2026). A Multi-Scale Global Fusion-Based Method for Surface Fissure Extraction from UAV Imagery. Sensors, 26(5), 1440. https://doi.org/10.3390/s26051440

