DAH-YOLO: An Accurate and Efficient Model for Crack Detection in Complex Scenarios
Abstract
1. Introduction
1.1. Background and Motivation
- (1)
- Standard YOLOv8 uses static kernels which limit adaptability. We introduce Dynamic Convolution to replace standard convolutions. By leveraging shared dynamic kernels, this module effectively captures complex crack features while optimizing parameter utilization and generalization capability.
- (2)
- An attention-enhanced Dynamic Head (Dyhead) is integrated to replace the traditional detection head. This mechanism refines the focus on critical feature map regions, significantly boosting detection accuracy for multi-scale cracks in complex environments.
- (3)
- Haar wavelet transformation is introduced for feature map down-sampling. Unlike the standard strided convolution in YOLOv8 which causes aliasing and detail loss, this method reduces redundant spatial information while preserving essential high-frequency edge and texture details, thereby mitigating false negatives in low-contrast conditions.
- (4)
- Comprehensive experiments demonstrate that DAH-YOLO outperforms mainstream YOLO variants and classical object detection models, offering a superior balance between robustness, accuracy, and computational efficiency.
1.2. Two-Stage Object Detection Model Based on Candidate Regions
1.3. Single-Stage Object Detection Model Based on Regression Analysis
2. Proposed Model
2.1. YOLOv8 Network Architecture
2.2. Overview of the Proposed Method
2.3. Dynamic Convolution Module
2.4. Dynamic Detection Head
2.5. Haar Wavelet Down-Sampling Module
3. Experimental Analysis
3.1. Experimental Datasets
- (1)
- CFD [37]: This dataset includes 584 images of size , collected using an iPhone 5 on an urban pavement in Beijing, China. All images have the real contours of the cracks traced by hand. This dataset poses significant challenges due to its dynamic background, which includes shadows, oil patches, water stains, and zebra crossings.
- (2)
- Crack-2 [38]: This dataset contains a total of 3102 images recorded in various road and wall scenarios, with an approximate image size of 416 by 416 pixels. The dataset includes a diverse array of images taken in different locations, environments, and densities.
- (3)
- Crack-500 [13]: This dataset is dedicated to the detection and identification of concrete cracks containing 9972 images with corresponding pixel-level annotations. The images have a resolution of and were captured using smartphones around Temple University. The images were taken from different concrete structures under varying angles and lighting situations; thereby encompassing images of cracks in shadows, occlusions, and diverse lighting environments.
3.2. Experimental Environment
3.3. Performance Metrics
3.4. Complexity Analysis
3.5. Results Comparison
- (1)
- Results on the CFD Dataset
- (2)
- Results on the Crack-2 Dataset
- (3)
- Results on the Crack500 Dataset.
3.6. Comparison with Other Models
3.7. Generalization Experimental Verification
3.8. Validation on Cross-Datasets
3.9. Ablation Experiment
4. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Mohan, A.; Poobal, S. Crack detection using image processing: A critical review and analysis. Alex. Eng. J. 2018, 57, 787–798. [Google Scholar] [CrossRef]
- Munawar, H.S.; Hammad, A.W.; Haddad, A.; Soares, C.A.P.; Waller, S.T. Image-based crack detection methods: A review. Infrastructures 2021, 6, 115. [Google Scholar] [CrossRef]
- Flah, M.; Suleiman, A.R.; Nehdi, M.L. Classification and quantification of cracks in concrete structures using deep learning image-based techniques. Cem. Concr. Compos. 2020, 114, 103781. [Google Scholar] [CrossRef]
- Luo, H.; Li, C.; Wu, M.; Cai, L. An enhanced lightweight network for road damage detection based on deep learning. Electronics 2023, 12, 2583. [Google Scholar] [CrossRef]
- Wan, B.; Zhou, X.; Zhu, B.; Xiao, M.; Sun, Y.; Zheng, B.; Zhang, J.; Yan, C. CANet: Context-aware aggregation network for salient object detection of surface defects. J. Vis. Commun. Image Represent. 2023, 93, 103820. [Google Scholar] [CrossRef]
- Pellegrino, F.A.; Vanzella, W.; Torre, V. Edge detection revisited. IEEE Trans. Syst. Man Cybern. Part B (Cybern.) 2004, 34, 1500–1518. [Google Scholar]
- Ehrenfried, K. Processing calibration-grid images using the Hough transformation. Meas. Sci. Technol. 2002, 13, 975. [Google Scholar] [CrossRef]
- Abdel-Qader, I.; Pashaie-Rad, S.; Abudayyeh, O.; Yehia, S. PCA-based algorithm for unsupervised bridge crack detection. Adv. Eng. Softw. 2006, 37, 771–778. [Google Scholar] [CrossRef]
- Salman, M.; Mathavan, S.; Kamal, K.; Rahman, M. Pavement crack detection using the Gabor filter. In Proceedings of the 16th International IEEE Conference on Intelligent Transportation Systems (ITSC 2013), The Hague, The Netherlands, 6–9 October 2013; IEEE: Piscataway, NJ, USA, 2013; pp. 2039–2044. [Google Scholar]
- Gunawan, G.; Nuriyanto, H.; Sriadhi, S.; Fauzi, A.; Usman, A.; Fadlina, F.; Dafitri, H.; Simarmata, J.; Utama Siahaan, A.P.; Rahim, R. Mobile application detection of road damage using canny algorithm. J. Phys. Conf. Ser. 2018, 1019, 012035. [Google Scholar] [CrossRef]
- Xiang, X.; Wang, Z.; Qiao, Y. An improved YOLOv5 crack detection method combined with transformer. IEEE Sens. J. 2022, 22, 14328–14335. [Google Scholar]
- Xu, X.; Zhao, M.; Shi, P.; Ren, R.; He, X.; Wei, X.; Yang, H. Crack detection and comparison study based on faster R-CNN and mask R-CNN. Sensors 2022, 22, 1215. [Google Scholar] [CrossRef] [PubMed]
- Yang, F.; Zhang, L.; Yu, S.; Prokhorov, D.; Mei, X.; Ling, H. Feature pyramid and hierarchical boosting network for pavement crack detection. IEEE Trans. Intell. Transp. Syst. 2019, 21, 1525–1535. [Google Scholar] [CrossRef]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; IEEE: Piscataway, NJ, USA, 2017; pp. 2961–2969. [Google Scholar]
- Dai, J.; Li, Y.; He, K.; Sun, J. R-FCN: Object detection via region-based fully convolutional networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems, Barcelona, Spain, 5–10 December 2016; Curran Associates Inc.: Red Hook, NY, USA, 2016; Volume 29. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 1137–1149. [Google Scholar] [CrossRef] [PubMed]
- Li, S.; Zhou, H.; Wang, G.; Zhu, X.; Kong, L.; Hu, Z. Cracked insulator detection based on R-FCN. J. Phys. Conf. Ser. 2018, 1069, 012147. [Google Scholar]
- Gan, L.; Liu, H.; Yan, Y.; Chen, A. Bridge bottom crack detection and modeling based on faster R-CNN and BIM. IET Image Process. 2024, 18, 664–677. [Google Scholar]
- Hacıefendioğlu, K.; Başağa, H.B. Concrete road crack detection using deep learning-based faster R-CNN method. Iran. J. Sci. Technol. Trans. Civ. Eng. 2022, 46, 1621–1633. [Google Scholar] [CrossRef]
- Pei, Z.; Lin, R.; Zhang, X.; Shen, H.; Tang, J.; Yang, Y. CFM: A consistency filtering mechanism for road damage detection. In Proceedings of the 2020 IEEE International Conference on Big Data (Big Data), Atlanta, GA, USA, 10–13 December 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 5584–5591. [Google Scholar]
- Arya, D.; Maeda, H.; Ghosh, S.K.; Toshniwal, D.; Omata, H.; Kashiyama, T.; Sekimoto, Y. Global road damage detection: State-of-the-art solutions. In Proceedings of the 2020 IEEE International Conference on Big Data (Big Data), Atlanta, GA, USA, 10–13 December 2020; IEEE: Piscataway, NJ, USA, 2020; pp. 5533–5539. [Google Scholar]
- Gonthina, M.; Chamata, R.; Duppalapudi, J.; Lute, V. Deep CNN-based concrete cracks identification and quantification using image processing techniques. Asian J. Civ. Eng. 2023, 24, 727–740. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. SSD: Single shot multibox detector. In Computer Vision—ECCV 2016, Proceedings of the 14th European Conference, Amsterdam, The Netherlands, 11–14 October 2016, Proceedings, Part I; Springer: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar]
- Jiang, P.; Ergu, D.; Liu, F.; Cai, Y.; Ma, B. A Review of Yolo algorithm developments. Procedia Comput. Sci. 2022, 199, 1066–1073. [Google Scholar] [CrossRef]
- Yan, K.; Zhang, Z. Automated asphalt highway pavement crack detection based on deformable single shot multi-box detector under a complex environment. IEEE Access 2021, 9, 150925–150938. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; IEEE: Piscataway, NJ, USA, 2016. [Google Scholar]
- Nie, M.; Wang, C. Pavement Crack Detection based on yolo v3. In Proceedings of the 2019 2nd International Conference on Safety Produce Informatization (IICSPI), Chongqing, China, 28–30 November 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 327–330. [Google Scholar]
- Wan, F.; Sun, C.; He, H.; Lei, G.; Xu, L.; Xiao, T. YOLO-LRDD: A lightweight method for road damage detection based on improved YOLOv5s. EURASIP J. Adv. Signal Process. 2022, 2022, 98. [Google Scholar] [CrossRef]
- Yang, Z.; Li, L.; Luo, W. PDNet: Improved YOLOv5 nondeformable disease detection network for asphalt pavement. Comput. Intell. Neurosci. 2022, 2022, 5133543. [Google Scholar] [CrossRef] [PubMed]
- Xiang, W.; Wang, H.; Xu, Y.; Zhao, Y.; Zhang, L.; Duan, Y. Road disease detection algorithm based on YOLOv5s-DSG. J. Real-Time Image Process. 2023, 20, 56. [Google Scholar] [CrossRef]
- Qiu, Q.; Lau, D. Real-time detection of cracks in tiled sidewalks using YOLO-based method applied to unmanned aerial vehicle (UAV) images. Autom. Constr. 2023, 147, 104745. [Google Scholar] [CrossRef]
- Lan, H.; Zhu, H.; Luo, R.; Ren, Q.; Chen, C. PCB defect detection algorithm of improved YOLOv8. In Proceedings of the 2023 8th International Conference on Image, Vision and Computing (ICIVC), Dalian, China, 27–29 July 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 178–183. [Google Scholar]
- Wang, X.; Gao, H.; Jia, Z.; Li, Z. BL-YOLOv8: An improved road defect detection model based on YOLOv8. Sensors 2023, 23, 8361. [Google Scholar] [CrossRef]
- Chunmei, W.; Huan, L. YOLOv8-VSC: Lightweight algorithm for strip surface defect detection. J. Front. Comput. Sci. Technol. 2024, 18, 151. [Google Scholar]
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 7464–7475. [Google Scholar]
- Chen, Y.; Dai, X.; Liu, M.; Chen, D.; Yuan, L.; Liu, Z. Dynamic ReLU. In Computer Vision—ECCV 2020, Proceedings of the 16th European Conference, Glasgow, UK, 23–28 August 2020, Proceedings, Part XIX; Springer: Cham, Switzerland, 2020; pp. 351–367. [Google Scholar]
- Shi, Y.; Cui, L.; Qi, Z.; Meng, F.; Chen, Z. Automatic road crack detection using random structured forests. IEEE Trans. Intell. Transp. Syst. 2016, 17, 3434–3445. [Google Scholar] [CrossRef]
- Crack Dataset. 2022. Available online: https://universe.roboflow.com/university-bswxt/crack-bphdr (accessed on 23 January 2024).
- Farhadi, A.; Redmon, J. Yolov3: An incremental improvement. In Proceedings of the Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; Springer: Berlin/Heidelberg, Germany, 2018; Volume 1804, pp. 1–6. [Google Scholar]
- Zhang, Y.; Guo, Z.; Wu, J.; Tian, Y.; Tang, H.; Guo, X. Real-time vehicle detection based on improved yolo v5. Sustainability 2022, 14, 12274. [Google Scholar] [CrossRef]
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv 2022, arXiv:2209.02976. [Google Scholar] [CrossRef]
- Terven, J.; Córdova-Esparza, D.M.; Romero-González, J.A. A comprehensive review of yolo architectures in computer vision: From yolov1 to yolov8 and yolo-nas. Mach. Learn. Knowl. Extr. 2023, 5, 1680–1716. [Google Scholar] [CrossRef]
- Wang, C.Y.; Yeh, I.H.; Mark Liao, H.Y. Yolov9: Learning what you want to learn using programmable gradient information. In Computer Vision—ECCV 2024, Proceedings of the 18th European Conference, Milan, Italy, 29 September–4 October 2024, Proceedings, Part XXXI; Springer: Cham, Switzerland, 2025; pp. 1–21. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. Yolov10: Real-time end-to-end object detection. arXiv 2024, arXiv:2405.14458. [Google Scholar]















| Environmental Parameter | Configuration |
|---|---|
| Deep Learning Framework | Pytorch |
| Programming Language | Python3.8 |
| CPU | Intel(R) Core(TM) i7-11700F |
| GPU | NVIDIA GeForce RTX 3090 |
| CUDA | 11.7 |
| Hyperparameters | Configuration |
|---|---|
| Input Size | |
| Training Epochs | 500 |
| Optimizer | Stochastic Gradient Descent (SGD) |
| Learning Rate | |
| Augmentation | Mosaic, Mixup, Random HSV, Random Flip |
| Model | Pr (%) | Re (%) | mAP@0.5 (%) | mAP@0.5: 0.95 (%) | F1 (%) | FLOPs (G) |
|---|---|---|---|---|---|---|
| SSD | 75.4 | 70.2 | 69.8 | 40.5 | 72.7 | 62.4 |
| Faster-RCNN | 80.3 | 75.5 | 78.2 | 45.9 | 77.8 | 370.1 |
| YOLOv3-tiny | 51.2 | 73.3 | 60.2 | 25.1 | 60.3 | 19.0 |
| YOLOv5 | 84.8 | 82.1 | 85.3 | 53.1 | 83.4 | 7.2 |
| YOLOv6 | 74.8 | 79.0 | 79.1 | 48.6 | 76.8 | 11.9 |
| YOLOv8 | 84.1 | 80.0 | 83.5 | 52.2 | 81.9 | 8.1 |
| YOLOv9 | 82.4 | 78.7 | 86.1 | 63.1 | 80.5 | 7.8 |
| YOLOv10 | 85.3 | 73.3 | 82.2 | 55.5 | 78.8 | 8.4 |
| Ours | 84.5 | 84.0 | 86.3 | 52.1 | 84.2 | 7.9 |
| Model | Pr (%) | Re (%) | mAP@0.5 (%) | mAP@0.5: 0.95 (%) | F1 (%) | FLOPs (G) |
|---|---|---|---|---|---|---|
| SSD | 78.3 | 58.7 | 65.4 | 48.7 | 67.1 | 62.4 |
| Faster-RCNN | 80.7 | 60.9 | 69.8 | 52.3 | 69.4 | 370.1 |
| YOLOv3-tiny | 81.9 | 70.5 | 71.2 | 54.0 | 75.8 | 19.0 |
| YOLOv5 | 88.1 | 64.7 | 71.5 | 53.4 | 74.6 | 7.2 |
| YOLOv6 | 79.0 | 66.0 | 70.4 | 53.5 | 71.9 | 11.9 |
| YOLOv8 | 83.8 | 68.9 | 71.5 | 51.8 | 75.6 | 8.1 |
| YOLOv9 | 82.2 | 65.3 | 70.0 | 52.4 | 72.8 | 7.8 |
| YOLOv10 | 86.7 | 64.2 | 72.4 | 55.7 | 73.8 | 8.4 |
| Ours | 83.0 | 71.6 | 75.9 | 56.5 | 76.9 | 7.9 |
| Model | Pr (%) | Re (%) | mAP@0.5 (%) | mAP@0.5: 0.95 (%) | F1 (%) | FLOPs (G) |
|---|---|---|---|---|---|---|
| SSD | 75.4 | 60.5 | 70.6 | 50.4 | 67.1 | 62.4 |
| Faster-RCNN | 78.2 | 70.1 | 74.9 | 54.8 | 73.9 | 370.1 |
| YOLOv3-tiny | 81.8 | 69.7 | 76.5 | 55.5 | 75.3 | 19.0 |
| YOLOv5 | 81.9 | 69.4 | 78.1 | 55.2 | 75.1 | 7.2 |
| YOLOv6 | 83.5 | 69.5 | 76.7 | 54.7 | 75.9 | 11.9 |
| YOLOv8 | 81.2 | 70.4 | 77.3 | 54.8 | 75.4 | 8.1 |
| YOLOv9 | 88.0 | 77.2 | 85.9 | 65.7 | 82.2 | 7.8 |
| YOLOv10 | 87.5 | 71.8 | 82.6 | 62.2 | 78.9 | 8.4 |
| Ours | 90.2 | 84.9 | 91.1 | 73.5 | 87.5 | 7.9 |
| Other Methods | CFD | Crack-2 | Crack500 | FLOPs (G) | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Pr (%) | Re (%) | F1 (%) | Pr (%) | Re (%) | F1 (%) | Pr (%) | Re (%) | F1 (%) | ||
| YOLO-LRDD [28] | 77.8 | 76.0 | 76.9 | 70.8 | 65.5 | 68.0 | 85.5 | 73.8 | 79.2 | 17.4 |
| YOLOv5s-DSG [30] | 89.1 | 78.7 | 83.6 | 86.7 | 64.9 | 74.2 | 88.5 | 79.9 | 83.9 | 19.4 |
| BL-YOLOv8 [33] | 83.5 | 81.9 | 82.7 | 82.6 | 67.6 | 74.4 | 81.0 | 71.6 | 76.0 | 25.5 |
| Ours | 84.5 | 84.0 | 84.2 | 83.0 | 71.6 | 76.9 | 90.2 | 84.9 | 87.5 | 7.9 |
| Testing Dataset | |||||||
|---|---|---|---|---|---|---|---|
| CFD | Crack-2 | Crack500 | |||||
| Pr (%) | Re (%) | Pr (%) | Re (%) | Pr (%) | Re (%) | ||
| Training dataset | CFD | 84.5 | 84.0 | 39.1 | 46.9 | 28.3 | 26.3 |
| Crack-2 | 66.5 | 64.0 | 83.0 | 71.6 | 44.6 | 35.8 | |
| Crack500 | 77.6 | 73.9 | 69.6 | 57.3 | 90.2 | 84.9 | |
| Method | Pr (%) | Re (%) | mAP@0.5 (%) | mAP@0.5: 0.95 (%) | F1 (%) | FLOPs (G) | |||
|---|---|---|---|---|---|---|---|---|---|
| Baseline | Dyconv | Dyhead | HWD | ||||||
| √ | 81.2 | 70.4 | 77.3 | 54.8 | 75.4 | 8.2 | |||
| √ | √ | 79.4 | 72.2 | 78.4 | 56.2 | 75.6 | 6.9 | ||
| √ | √ | √ | 86.3 | 79.9 | 86.8 | 67.8 | 83.0 | 8.5 | |
| √ | √ | √ | √ | 90.2 | 84.9 | 91.1 | 73.5 | 87.5 | 7.9 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Fan, Y.; Li, Q.; Chen, Y.; Yao, Z.; Sun, Y.; Zhang, W. DAH-YOLO: An Accurate and Efficient Model for Crack Detection in Complex Scenarios. Appl. Sci. 2026, 16, 900. https://doi.org/10.3390/app16020900
Fan Y, Li Q, Chen Y, Yao Z, Sun Y, Zhang W. DAH-YOLO: An Accurate and Efficient Model for Crack Detection in Complex Scenarios. Applied Sciences. 2026; 16(2):900. https://doi.org/10.3390/app16020900
Chicago/Turabian StyleFan, Yawen, Qinxin Li, Ye Chen, Zhiqiang Yao, Yang Sun, and Wentao Zhang. 2026. "DAH-YOLO: An Accurate and Efficient Model for Crack Detection in Complex Scenarios" Applied Sciences 16, no. 2: 900. https://doi.org/10.3390/app16020900
APA StyleFan, Y., Li, Q., Chen, Y., Yao, Z., Sun, Y., & Zhang, W. (2026). DAH-YOLO: An Accurate and Efficient Model for Crack Detection in Complex Scenarios. Applied Sciences, 16(2), 900. https://doi.org/10.3390/app16020900

