Real-Time Road Crack Detection on Smartphones Through ConvLSTM-Based Temporal Knowledge Distillation from a CNN-KAN and VMamba Dual-Path Network
Abstract
1. Introduction
2. Methods and Datasets
2.1. Model Feature Extraction
2.1.1. Kolmogorov–Arnold Enhanced Convolutional Neural Network
2.1.2. Vision Mamba
2.1.3. KAN Fusion Module
2.2. KAN Fusion-Based Dual-Path Neural Network for Crack Detection
2.3. Datasets Description
3. Model Implementation and Optimization Evaluation
3.1. Optimization of Generative Adversarial Networks
3.2. Model Distillation
3.3. Segmentation Model Training and Evaluation
3.3.1. Optimization and Model Initialization
3.3.2. Model Evaluation Metrics
3.3.3. Loss Function
4. Experimental Results and Analysis
4.1. Model Evaluation on Public Datasets
4.2. Model Transfer and Evaluation on Road Crack Datasets
4.3. Ablation Experiment
4.4. Evaluation of Model Robustness
4.5. GAN-Driven Optimization of Segmentation Model
4.6. Deployment of Lightweight Distillation Model
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Zhou, Y.; Liang, M.; Yue, X. Deep residual learning for acoustic emission source localization in a steel-concrete composite slab. Constr. Build. Mater. 2023, 411, 134220. [Google Scholar]
- Zhou, Y.; Liu, Y.; Lian, Y.; Pan, T.; Zheng, Y.; Zhou, Y. Ambient vibration measurement-aided multi-1D CNNs ensemble for damage localization framework: Demonstration on a large-scale RC pedestrian bridge. Mech. Syst. Signal Process. 2025, 224, 111937. [Google Scholar] [CrossRef]
- Russel, N.S.; Selvaraj, A. MultiScaleCrackNet: A parallel multiscale deep CNN architecture for concrete crack classification. Expert Syst. Appl. 2024, 249, 123658. [Google Scholar] [CrossRef]
- Deng, J.; Singh, A.; Zhou, Y.; Lu, Y.; Lee, V.C.S. Review on computer vision-based crack detection and quantification methodologies for civil structures. Constr. Build. Mater. 2022, 356, 129238. [Google Scholar] [CrossRef]
- Ali, R.; Chuah, J.H.; Talip, M.S.A.; Mokhtar, N.; Shoaib, M.A. Structural crack detection using deep convolutional neural networks. Autom. Constr. 2022, 133, 103989. [Google Scholar] [CrossRef]
- Matarneh, S.; Elghaish, F.; Rahimian, F.P.; Abdellatef, E.; Abrishami, S. Evaluation and optimisation of pre-trained CNN models for asphalt pavement crack detection and classification. Autom. Constr. 2024, 160, 105297. [Google Scholar] [CrossRef]
- Liu, F.; Wang, L. UNet-based model for crack detection integrating visual explanations. Constr. Build. Mater. 2022, 322, 126265. [Google Scholar] [CrossRef]
- Guo, F.; Liu, J.; Lv, C.; Yu, H. A novel transformer-based network with attention mechanism for automatic pavement crack detection. Constr. Build. Mater. 2023, 391, 131852. [Google Scholar] [CrossRef]
- Wang, Z.; Leng, Z.; Zhang, Z. A weakly-supervised transformer-based hybrid network with multi-attention for pavement crack detection. Constr. Build. Mater. 2024, 411, 134134. [Google Scholar] [CrossRef]
- Wang, C.; Liu, H.; An, X.; Gong, Z.; Deng, F. SwinCrack: Pavement crack detection using convolutional swin-transformer network. Digit. Signal Process. 2024, 145, 104297. [Google Scholar] [CrossRef]
- Zhu, G.; Liu, J.; Fan, Z.; Yuan, D.; Ma, P.; Wang, M.; Sheng, W.; Wang, K.C. A lightweight encoder–decoder network for automatic pavement crack detection. Comput. Aided Civ. Infrastruct. Eng. 2024, 39, 1743–1765. [Google Scholar] [CrossRef]
- Liu, Z.; Wang, Y.; Vaidya, S.; Ruehle, F.; Halverson, J.; Soljačić, M.; Hou, T.Y.; Tegmark, M. Kan: Kolmogorov-arnold networks. arXiv 2024, arXiv:2404.19756. [Google Scholar]
- Li, C.; Liu, X.; Li, W.; Wang, C.; Liu, H.; Liu, Y.; Chen, Z.; Yuan, Y. U-kan makes strong backbone for medical image segmentation and generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; Volume 39, pp. 4652–4660. [Google Scholar]
- Yang, X.; Wang, D. SKPNet: Snake KAN perceive bridge cracks through semantic segmentation. Intell. Robot. 2025, 5, 105–118. [Google Scholar] [CrossRef]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. In Proceedings of the First Conference on Language Modeling, Philadelphia, PA, USA, 7–9 October 2024. [Google Scholar]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. Vmamba: Visual state space model. Adv. Neural Inf. Process. Syst. 2024, 37, 103031–103063. [Google Scholar] [CrossRef]
- Han, C.; Yang, H.; Yang, Y. Enhancing pixel-level crack segmentation with visual mamba and convolutional networks. Autom. Constr. 2024, 168, 105770. [Google Scholar] [CrossRef]
- Ji, T.; Hou, Y.; Zhang, D. A comprehensive survey on Kolmogorov Arnold networks (KAN). arXiv 2024, arXiv:2407.11075v7. [Google Scholar]
- Liang, J.; Gu, X.; Jiang, D.; Zhang, Q. CNN-based network with multi-scale context feature and attention mechanism for automatic pavement crack segmentation. Autom. Constr. 2024, 164, 105482. [Google Scholar] [CrossRef]
- Cheon, M.; Mun, C. Combining KAN with CNN: KonvNeXt’s performance in remote sensing and patent insights. Remote Sens. 2024, 16, 3417. [Google Scholar] [CrossRef]
- Shi, T.; Luo, H. Deep learning for automated detection and classification of crack severity level in concrete structures. Constr. Build. Mater. 2025, 472, 140793. [Google Scholar] [CrossRef]
- Zhang, Z.; Peng, B.; Zhao, T. An ultra-lightweight network combining Mamba and frequency-domain feature extraction for pavement tiny-crack segmentation. Expert Syst. Appl. 2025, 264, 125941. [Google Scholar] [CrossRef]
- Zuo, X.; Sheng, Y.; Shen, J.; Shan, Y. Topology-aware mamba for crack segmentation in structures. Autom. Constr. 2024, 168, 105845. [Google Scholar] [CrossRef]
- Chu, H.; Wang, W.; Deng, L. Tiny-Crack-Net: A multiscale feature fusion network with attention mechanisms for segmentation of tiny cracks. Comput. Aided Civ. Infrastruct. Eng. 2022, 37, 1914–1931. [Google Scholar] [CrossRef]
- Qu, Z.; Chen, W.; Wang, S.Y.; Yi, T.M.; Liu, L. A crack detection algorithm for concrete pavement based on attention mechanism and multi-features fusion. IEEE Trans. Intell. Transp. Syst. 2021, 23, 11710–11719. [Google Scholar] [CrossRef]
- Chen, Y.; Li, J.; Xiao, H.; Jin, X.; Yan, S.; Feng, J. Dual path networks. Adv. Neural Inf. Process. Syst. 2017, 30, 4467–4475. [Google Scholar]
- Zhang, L.; Yang, F.; Zhang, Y.D.; Zhu, Y.J. Road crack detection using deep convolutional neural network. In Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA, 25–28 September 2016; pp. 3708–3712. [Google Scholar]
- Liu, Y.; Yao, J.; Lu, X.; Xie, R.; Li, L. DeepCrack: A deep hierarchical feature learning architecture for crack segmentation. Neurocomputing 2019, 338, 139–153. [Google Scholar] [CrossRef]
- Cui, L.; Qi, Z.; Chen, Z.; Meng, F.; Shi, Y. Pavement distress detection using random decision forests. In Proceedings of the International Conference on Data Science, Sydney, Australia, 8–9 August 2015; Springer: Berlin/Heidelberg, Germany, 2015; pp. 95–102. [Google Scholar]
- Zou, Q.; Cao, Y.; Li, Q.; Mao, Q.; Wang, S. CrackTree: Automatic crack detection from pavement images. Pattern Recognit. Lett. 2012, 33, 227–238. [Google Scholar] [CrossRef]
- Ni, F.; He, Z.; Jiang, S.; Wang, W.; Zhang, J. A generative adversarial learning strategy for enhanced lightweight crack delineation networks. Adv. Eng. Inform. 2022, 52, 101575. [Google Scholar] [CrossRef]
- Chu, C.; Zhmoginov, A.; Sandler, M. Cyclegan, a master of steganography. arXiv 2017, arXiv:1712.02950. [Google Scholar]
- Anwar, S.; Khan, S.; Barnes, N. A deep journey into super-resolution: A survey. ACM Comput. Surv. (CSUR) 2020, 53, 1–34. [Google Scholar] [CrossRef]
- An, Y.K.; Kang, M.S. Crack growth prediction on a concrete structure using deep ConvLSTM. Smart Struct. Syst. 2024, 33, 301–311. [Google Scholar]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. Pytorch: An imperative style, high-performance deep learning library. Adv. Neural Inf. Process. Syst. 2019, 32, 8024–8035. [Google Scholar]
- Sudre, C.H.; Li, W.; Vercauteren, T.; Ourselin, S.; Jorge Cardoso, M. Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations. In Proceedings of the International Workshop on Deep Learning in Medical Image Analysis, Québec City, QC, Canada, 14 September 2017; Springer: Berlin/Heidelberg, Germany, 2017; pp. 240–248. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 5–9 October 2015; Springer: Berlin/Heidelberg, Germany, 2015; pp. 234–241. [Google Scholar]
- Diakogiannis, F.I.; Waldner, F.; Caccetta, P.; Wu, C. ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data. ISPRS J. Photogramm. Remote Sens. 2020, 162, 94–114. [Google Scholar] [CrossRef]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2881–2890. [Google Scholar]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar]
- Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 834–848. [Google Scholar] [CrossRef] [PubMed]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 801–818. [Google Scholar]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. Transunet: Transformers make strong encoders for medical image segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar]
- Liu, H.; Miao, X.; Mertz, C.; Xu, C.; Kong, H. Crackformer: Transformer network for fine-grained crack detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Virtual, 11–17 October 2021; pp. 3783–3792. [Google Scholar]
- Xiong, B.; Hong, R.; Wang, J.; Li, W.; Zhang, J.; Lv, S.; Ge, D. DefNet: A multi-scale dual-encoding fusion network aggregating Transformer and CNN for crack segmentation. Constr. Build. Mater. 2024, 448, 138206. [Google Scholar] [CrossRef]



















| Dataset | Crack500 | DeepCrack | CFD | CrackTree | ||||
|---|---|---|---|---|---|---|---|---|
| Metric | mIoU | mDice | mIoU | mDice | mIoU | mDice | mIoU | mDice |
| Proposed | 0.82 | 0.89 | 0.86 | 0.92 | 0.76 | 0.86 | 0.72 | 0.82 |
| SKPNet | 0.75 | 0.86 | 0.80 | 0.89 | 0.72 | 0.83 | 0.70 | 0.83 |
| MambaCrackNet | 0.73 | 0.85 | 0.77 | 0.87 | 0.68 | 0.81 | 0.67 | 0.79 |
| UNet | 0.66 | 0.77 | 0.68 | 0.83 | 0.63 | 0.79 | 0.57 | 0.70 |
| ResUNet | 0.69 | 0.80 | 0.71 | 0.85 | 0.65 | 0.78 | 0.61 | 0.74 |
| PSPNet | 0.47 | 0.63 | 0.53 | 0.68 | 0.40 | 0.58 | 0.34 | 0.52 |
| FPN | 0.67 | 0.81 | 0.67 | 0.85 | 0.66 | 0.80 | 0.63 | 0.75 |
| DeepLabV3 | 0.62 | 0.72 | 0.61 | 0.75 | 0.56 | 0.68 | 0.50 | 0.64 |
| DeepLabV3+ | 0.65 | 0.83 | 0.68 | 0.86 | 0.67 | 0.81 | 0.62 | 0.76 |
| TransUNet | 0.70 | 0.82 | 0.71 | 0.84 | 0.69 | 0.77 | 0.63 | 0.79 |
| CrackFormer | 0.72 | 0.80 | 0.76 | 0.82 | 0.68 | 0.81 | 0.73 | 0.81 |
| DefNet | 0.76 | 0.84 | 0.79 | 0.88 | 0.71 | 0.82 | 0.70 | 0.83 |
| Model | Params (M) | FLOPs (G) | P (%) | R (%) | (%) | mIoU | mDice |
|---|---|---|---|---|---|---|---|
| Proposed | 57.80 | 60.07 | 90.29 | 88.99 | 90.10 | 0.8087 | 0.9028 |
| SKPNet | 31.14 | 168.02 | 86.82 | 85.99 | 86.40 | 0.7606 | 0.8640 |
| MambaCrackNet | 41.78 | 219.50 | 81.84 | 89.46 | 85.48 | 0.7464 | 0.8548 |
| UNet | 31.04 | 100.27 | 80.19 | 80.70 | 79.77 | 0.6634 | 0.7977 |
| ResUNet | 57.57 | 66.72 | 81.85 | 83.20 | 81.96 | 0.6944 | 0.8196 |
| PSPNet | 19.99 | 34.57 | 64.17 | 70.98 | 64.34 | 0.4742 | 0.6434 |
| FPN | 16.07 | 55.81 | 85.12 | 82.37 | 83.13 | 0.7112 | 0.8313 |
| DeepLabV3 | 13.76 | 22.95 | 71.24 | 83.26 | 76.12 | 0.6145 | 0.7312 |
| DeepLabV3+ | 14.46 | 52.06 | 86.67 | 82.44 | 83.90 | 0.7227 | 0.8390 |
| TransUNet | 87.91 | 122.47 | 84.72 | 83.18 | 81.77 | 0.6989 | 0.8177 |
| CrackFormer | 56.60 | 81.54 | 87.21 | 83.81 | 83.95 | 0.7324 | 0.8395 |
| DefNet | 60.75 | 71.26 | 87.49 | 85.64 | 86.51 | 0.7622 | 0.8651 |
| Model | (%) | (%) | (%) | mIoU ∗ (, ) | mDice ∗ (, ) |
|---|---|---|---|---|---|
| Proposed | 81.71 | 82.90 | 80.41 | 0.7084, 0.2989 | 0.8141, 0.2648 |
| SKPNet | 75.21 | 77.86 | 76.32 | 0.6707, 0.3004 | 0.7555, 0.2719 |
| MambaCrackNet | 72.45 | 74.53 | 73.71 | 0.6507, 0.3016 | 0.7418, 0.2628 |
| UNet | 61.06 | 69.34 | 65.61 | 0.4883, 0.3307 | 0.6561, 0.3039 |
| ResUNet | 57.16 | 69.62 | 60.73 | 0.5334, 0.3391 | 0.6673, 0.3051 |
| PSPNet | 31.55 | 67.58 | 32.47 | 0.1938, 0.3279 | 0.3247, 0.4308 |
| FPN | 57.98 | 67.34 | 61.35 | 0.4425, 0.3367 | 0.6135, 0.3026 |
| DeepLabV3 | 28.67 | 64.63 | 35.04 | 0.2124, 0.2759 | 0.3504, 0.2496 |
| DeepLabV3+ | 51.36 | 65.88 | 56.85 | 0.3972, 0.3440 | 0.5685, 0.3080 |
| TransUNet | 76.58 | 75.01 | 74.52 | 0.6591, 0.2993 | 0.7452, 0.2712 |
| CrackFormer | 76.55 | 77.41 | 75.77 | 0.6614, 0.2790 | 0.7577, 0.2548 |
| DefNet | 78.01 | 79.52 | 78.88 | 0.6899, 0.3140 | 0.7888, 0.2779 |
| Dataset | Public Datasets | Road Crack Datasets | ||||
|---|---|---|---|---|---|---|
| CNN-KAN | VMamba | KAN Fusion | mIoU | mDice | mIoU | mDice |
| ✓ | ✓ | ✓ | 0.8087 | 0.9028 | 0.7084 | 0.8141 |
| ✓ | ✗ | ✗ | 0.7690 | 0.8733 | 0.6383 | 0.7523 |
| ✗ | ✓ | ✗ | 0.6788 | 0.7927 | 0.5827 | 0.7141 |
| ✓ | ✓ | ✗ | 0.7726 | 0.8765 | 0.6642 | 0.7765 |
| Training Strategy | Params (M) | Inference Time (ms) | mIoU | mDice | |
|---|---|---|---|---|---|
| Without distillation | ✗ | 1.57 | 1.46 | 0.6021 | 0.7092 |
| With distillation | ✗ | 1.57 | 1.46 | 0.6785 | 0.7867 |
| With distillation | ✓ | 1.57 | 1.46 | 0.7078 | 0.8137 |
| Model | Params (M) | Size (MB) | Inference Time ∗ (ms) | mIoU | mDice |
|---|---|---|---|---|---|
| Proposed | 57.80 | 220.65 | 99.56 | 0.7084 | 0.8141 |
| EfficientNet | 2.44 | 9.33 | 1.60 | 0.5561 | 0.7059 |
| MobileNet | 3.23 | 12.38 | 6.64 | 0.6954 | 0.8056 |
| ShuffleNet | 4.41 | 16.87 | 7.18 | 0.6975 | 0.8070 |
| Fast-SCNN | 4.68 | 18.98 | 1.02 | 0.6530 | 0.7770 |
| MobileViT | 1.57 | 6.01 | 1.46 | 0.7078 | 0.8137 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Nie, M.; Huang, H.; Guo, M.; Dang, M.; Tang, M. Real-Time Road Crack Detection on Smartphones Through ConvLSTM-Based Temporal Knowledge Distillation from a CNN-KAN and VMamba Dual-Path Network. Sensors 2026, 26, 5071. https://doi.org/10.3390/s26165071
Nie M, Huang H, Guo M, Dang M, Tang M. Real-Time Road Crack Detection on Smartphones Through ConvLSTM-Based Temporal Knowledge Distillation from a CNN-KAN and VMamba Dual-Path Network. Sensors. 2026; 26(16):5071. https://doi.org/10.3390/s26165071
Chicago/Turabian StyleNie, Mengzhao, Hua Huang, Mengxue Guo, Mingxia Dang, and Ming Tang. 2026. "Real-Time Road Crack Detection on Smartphones Through ConvLSTM-Based Temporal Knowledge Distillation from a CNN-KAN and VMamba Dual-Path Network" Sensors 26, no. 16: 5071. https://doi.org/10.3390/s26165071
APA StyleNie, M., Huang, H., Guo, M., Dang, M., & Tang, M. (2026). Real-Time Road Crack Detection on Smartphones Through ConvLSTM-Based Temporal Knowledge Distillation from a CNN-KAN and VMamba Dual-Path Network. Sensors, 26(16), 5071. https://doi.org/10.3390/s26165071

