Deep Learning for Surface-Conditioned Structural Crack Recognition: Integrating CNN-Attention Models and Calibrated Confidence Estimation
Abstract
1. Introduction
- A traceable 49,124-image audit with the original 39,296/4902/4926 train/validation/test split, an explicit correction to the perceptual-hash isolation claim, and a 4819-image hash-isolated sensitivity analysis. Source-folder controls and a separate source-held-out linear probe quantify important limits to generalization.
- A matched comparison of ten CNN and ten attention-based residual-head configurations on one frozen ResNet-18 encoder, including an explicit linear control and identification of four selected zero-correction heads. Conclusions concern this representation and budget, not twenty independently trained end-to-end architectures.
- Validation-fitted confidence calibration evaluated with class-balanced discrimination, NLL, Brier score, ECE, class-conditional diagnostics, and additional balanced-temperature and regularized-vector comparisons. Test outcomes are not used to replace the originally validation-selected head.
- Paired component-bootstrap uncertainty, repeated head-training seeds, complete-test image-quality stress tests, quantitative Grad-CAM deletion analysis, and traceable real-image examples. Revision analyses are identified as post hoc and retain unfavorable outcomes as well as favorable ones.
2. Materials and Methods
2.1. Structural Image Corpus, Taxonomy, and Integrity Control
2.1.1. Corpus Hierarchy
2.1.2. Exact and Perceptual Duplicate Grouping
2.1.3. Grouped 80:10:10 Partition
| Algorithm 1 Corpus audit and duplicate-controlled partitioning |
| Input: Original 49,124-image corpus and supplied category labels Output: Original split plus a corrected revision integrity audit 1. Enumerate images and retain file size, SHA-256, dHash, class and source-folder string. 2. Reproduce the original grouping: exact-hash duplicates take priority; other files use class-specific equal dHash. 3. Shuffle original groups within class with seed 20260903 and allocate approximately 80:10:10. 4. Retain the original 39,296/4902/4926 split; verify zero exact-hash overlap. 5. Revision: form connected components using equal SHA-256 OR within-class equal dHash. 6. Count cross-partition matches and exclude affected test components for sensitivity analysis. 7. Preserve all original assignments and report the post-hoc status of the corrected audit. |
2.2. Common Encoder and Matched Twenty-Method Recognition Benchmark
2.2.1. Common Representation Learning
2.2.2. Residual Recognition Heads
2.2.3. Controlled Head Training and Validation Selection
| Algorithm 2 Common-encoder training and matched residual-head selection |
| Input: Original training and validation partitions Output: Common encoder and twenty validation-selected conditional heads 1. Train the encoder under the recorded eight-epoch schedule; select by validation macro-F1. 2. Freeze the selected encoder and its original linear classifier. 3. For each CNN or attention adapter, initialize its final correction layer to zero and gate parameter to −1.5. 4. Evaluate epoch 0 and five trained epochs with the common head budget. 5. Retain each head’s best validation macro-F1 checkpoint, including epoch 0 when selected. 6. Select TokenFormer-8H by validation macro-F1; do not promote a head using test rank. 7. Record the four zero-correction selections and preserve the linear control. |
2.3. Confidence Calibration, Evaluation Metrics, and Statistical Analysis
2.3.1. Validation-Only Temperature Scaling
2.3.2. Discrimination and Calibration Measures
2.3.3. Paired Statistical Testing and Bootstrap Uncertainty
| Algorithm 3 Calibration, final evaluation, and paired statistical analysis |
| Input: Fixed checkpoints, validation logits and original test images Output: Calibrated predictions and conditional uncertainty estimates 1. Fit one positive temperature to each head by validation NLL; freeze it. 2. Reproduce predictions and compute class-balanced and aggregate test metrics. 3. Revision: construct union-hash components and resample the 3920 test components 5000 times. 4. Use identical resamples for all heads and the linear control. 5. Report percentile intervals for accuracy, macro-F1, balanced accuracy and paired differences. 6. Retain image-level Cochran’s Q and McNemar/Holm results only as secondary dependent-data diagnostics. 7. Fit balanced-temperature and regularized-vector alternatives on validation data; report all test outcomes without selecting a replacement. 8. Interpret repeated seeds as head-training variation with one fixed encoder. |
2.4. Robustness, Visual Explanation, and Implementation Details
| Algorithm 4 Robustness and explanation protocol |
| Input: Validation-selected model with the original temperature fixed Output: Image-quality sensitivity, attribution diagnostics and traceable examples 1. Evaluate the original five fixed corruptions on all 4926 test images. 2. Add six documented acquisition-like transformations to the same full test set. 3. Sample up to 50 component-distinct images per class without correctness selection. 4. Compute layer-3 Grad-CAM; count empty heatmaps. 5. Compare top-attribution and seeded random-pixel deletion at 10%, 20% and 30% area. 6. Bootstrap paired target-probability-drop differences over 323 component-distinct images. 7. Show actual images and traceable failures; do not claim crack-mask, physical-scale or field validation. |
2.5. Reproducibility and Computational Record
2.6. Revision-Stage Diagnostic Protocols
2.6.1. Source Controls and Hash-Isolated Evaluation
2.6.2. Additional Calibration Comparisons
2.6.3. Quantitative Attribution and Retrospective Examples
3. Results
3.1. Experimental Setup
3.1.1. Dataset Preparation
3.1.2. Implementation Details
3.2. Evaluation Metrics
3.3. Comparison of Twenty Matched Recognition Methods
3.4. Controlled Calibration, Uncertainty, and Robustness Analysis
3.4.1. Calibration and Bootstrap Uncertainty
3.4.2. Repeated-Seed Stability
3.4.3. Paired Statistical Comparison
3.4.4. Robustness and Failure Analysis
3.5. Integrity, Source Controls, and Subset Performance
3.6. Class-Dependent Calibration Trade-Offs
3.7. Quantitative Attribution Results
3.8. Acquisition Sensitivity and Retrospective Walkthrough
4. Discussion
4.1. Positioning in Relation to Recent Works
4.2. Overall Performance Evaluation
4.3. Class-Wise Error and Minority-Category Behavior
4.4. CNN and Attention-Family Evidence
4.5. Calibration and Confidence Reliability
4.6. Robustness and Visual Failure Analysis
4.7. Representation Behaviour, Explainability, and Practical Model Choice
4.8. Engineering Scope and Required Ground Truth
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Zou, Z.; Wang, M.; Song, B.; He, J.; Yang, S. Deep learning and digital twin driven structural health monitoring of masonry structures: Progress and challenges. Autom. Constr. 2026, 188, 106993. [Google Scholar] [CrossRef] [Scilit]
- Zhuang, H.; Cheng, Y.; Zhou, M.; Yang, Z. Deep learning for surface crack detection in civil engineering: A comprehensive review. Measurement 2025, 248, 116908. [Google Scholar] [CrossRef] [Scilit]
- Song, Y.; Zhang, Q.; Su, Y.; Zhang, S.; Wang, R.; Zhang, W.; Bi, Z.; Yu, Y. Advances in crack dataset development and deep learning-based detection models. J. Build. Eng. 2025, 116, 114734. [Google Scholar] [CrossRef] [Scilit]
- Zeng, Y.; Lei, D.; He, J.; Zhou, K.; Wang, D. Deep learning-based classification of concrete crack evolution stages with reference to the double-K fracture criterion. J. Build. Eng. 2026, 119, 115176. [Google Scholar] [CrossRef] [Scilit]
- Xu, G.; Zhang, Y.; Yue, Q.; Liu, X. A deep learning framework for real-time multi-task recognition and measurement of concrete cracks. Adv. Eng. Inform. 2025, 65, 103127. [Google Scholar] [CrossRef] [Scilit]
- Ganduri, K.V.; Pathri, B.P.; Hemachandran, K.; Kommu, P.K.; Sharma, V.K. A deep learning and unmanned aerial vehicle framework for autonomous structural monitoring and crack detection. Eng. Appl. Artif. Intell. 2026, 182, 115922. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Su, D.; Zhang, Y.; Lin, D.-W.; Sun, Z. An enhanced YOLO-based deep learning framework for automated structural defect inspection in complex backgrounds. Eng. Struct. 2026, 366, 123483. [Google Scholar] [CrossRef] [Scilit]
- Liu, S.; Isobe, K.; Si, J.; Li, D.; Ren, D. A multi-module deep learning framework with graph-based network and crack attention for tunnel lining crack segmentation from LiDAR point cloud. Constr. Build. Mater. 2025, 494, 143383. [Google Scholar] [CrossRef] [Scilit]
- Majidi, S.; Sharifi, M.A.; Omidalizarandi, M. Deep learning-based crack detection and 3D reconstruction for cost-effective structural health monitoring. Meas. Digit. 2026, 7, 100040. [Google Scholar] [CrossRef] [Scilit]
- Duan, Z.; Cai, Z.; Li, Q.; Liu, Y.; Mao, Y.; Zhou, X. Adaptive deep learning framework for crack contour recognition and dimensional measurement in concrete structures. Constr. Build. Mater. 2026, 512, 145400. [Google Scholar] [CrossRef] [Scilit]
- Jin, T.; Shou, Z.; Liu, H.; Shao, Y. Attention mechanisms and FFM feature fusion module-based modification of the deep neural network for detection of structural cracks. CMES Comput. Model. Eng. Sci. 2026, 146, 11. [Google Scholar] [CrossRef] [Scilit]
- Dias Júnior, L.T.; Finotti, R.P.; Barbosa, F.S.; Cury, A.A. The trajectory of data-driven structural health monitoring: A review from traditional methods to deep learning and future trends for civil infrastructures. CMES Comput. Model. Eng. Sci. 2026, 146, 3. [Google Scholar] [CrossRef] [Scilit]
- Chauhan, S.; Ranjan, R.; Dutta, B.R. A hybrid deep learning method for the crack detection and classification in surface images. Frankl. Open 2026, 16, 100705. [Google Scholar] [CrossRef] [Scilit]
- Djerrad, A.; Zhou, Y.; Meng, S. A transformer-based deep learning framework for predicting crack patterns and structural responses in RC shear walls. Eng. Struct. 2025, 343, 121051. [Google Scholar] [CrossRef] [Scilit]
- Ahani, E.; Yang, J. Deep learning framework for crack type detection in laminated glass based on ultrasonic and modal analysis using finite element simulations. J. Non-Cryst. Solids 2026, 676, 123950. [Google Scholar] [CrossRef] [Scilit]
- Talaghat, M.A.; Golroo, A.; Shahhosseini, V.; Rasti, M. Fully automated pavement crack quantification and PCI estimation via a digital twin-ready vision-based deep learning pipeline. Results Eng. 2026, 32, 111659. [Google Scholar] [CrossRef] [Scilit]
- Shah, S.M.H.; Qureshi, W.S.; O’Dea, G.; Power, D.; Ullah, I. Automated pavement condition rating for cycle routes and greenways using deep learning. Autom. Constr. 2026, 192, 107210. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Wang, C.; Xie, M.; Qi, W.; Feng, G.; Zhang, Z.; Li, Y. Soil crack healing, recurrence and the temporal persistence of preferential flow under wet-dry cycles revealed by deep-learning image analysis and breakthrough curves. Soil Tillage Res. 2026, 259, 107054. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Tian, C.; Yang, J.; Xiao, P.; Wu, Z.; Li, H.; Sun, W. Vision and deep learning-based tunnel surface defect detection: Technological evolution, challenges, and prospects. Measurement 2026, 257, 118853. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Wang, X.; Wu, S.; Guo, S.; Li, B.; Zhu, R.; Xu, L.; Li, D.; Zou, Z.; Zhang, C. Intelligent crack detection in underwater concrete structures: An image enhancement and deep learning fusion-driven approach. Case Stud. Constr. Mater. 2026, 25, e06421. [Google Scholar] [CrossRef] [Scilit]
- Yuan, J.; Ren, Q.; Jia, C.; Zhang, J.; Fu, J.; Li, M. Automated pixel-level crack detection and quantification using deep convolutional neural networks for structural condition assessment. Structures 2024, 59, 105780. [Google Scholar] [CrossRef] [Scilit]
- Zhuang, X.; Tran, T.V.; Nguyen-Xuan, H.; Rabczuk, T. Deep learning-based post-earthquake structural damage level recognition. Comput. Struct. 2025, 315, 107761. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Zhang, H.; Xiao, D. Research on defect recognition method of mine belts based on deep learning. Eng. Appl. Artif. Intell. 2026, 177, 114973. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Bu, B.; Zai, D.; Meng, L.; Li, J. A high-precision deep learning concrete thin crack segmentation network. Structures 2026, 92, 112936. [Google Scholar] [CrossRef] [Scilit]
- Feng, J.; Wang, Y. Intelligent detection algorithm for concrete structure cracks based on deep learning. Procedia Comput. Sci. 2026, 281, 1095–1104. [Google Scholar] [CrossRef] [Scilit]
- Tang, J.; Shang, Y.; Meng, J.; Li, J.; Xu, M.; Hu, Y.; Zhang, J. A new method for quantitative evaluation of micro-cracks on turbine blade surfaces fusing triboelectric sensing and hybrid deep learning. Mech. Syst. Signal Process. 2026, 247, 113947. [Google Scholar] [CrossRef] [Scilit]
- Lu, M.; Qian, Q.; Yang, F.; Li, M. Deep learning enhanced crack identification on rocks. Artif. Intell. Geosci. 2026, 7, 100200. [Google Scholar] [CrossRef] [Scilit]
- Ruggieri, S.; Cardellicchio, A.; Nettis, A.; Renò, V.; Uva, G. Using Attention for Improving Defect Detection in Existing RC Bridges. IEEE Access 2025, 13, 18994–19015. [Google Scholar] [CrossRef] [Scilit]
- Cardellicchio, A.; Renò, V.; Natali, A.; Di Mucci, V.M.; Nettis, A.; Ruggieri, S.; Uva, G. An automated framework to characterize crack patterns in existing RC bridges. Data-Centric Eng. 2026, 7, e26. [Google Scholar] [CrossRef] [Scilit]
- Aung, P.P.W.; Sam, K.M.; Kulinan, A.S.; Cha, G.; Park, M.; Park, S. Enhancing deep learning in structural damage identification with 3D-engine synthetic data. Autom. Constr. 2025, 175, 106203. [Google Scholar] [CrossRef] [Scilit]
- Türer, A.; Bai, Y.; Sezen, H.; Yilmaz, A. Automated post-earthquake structural damage assessment of concrete buildings using a hybrid deep learning and rule-based framework on image datasets. J. Infrastruct. Intell. Resil. 2026, 5, 100208. [Google Scholar] [CrossRef] [Scilit]
- Awan, M.R.; Chan, C.-W.; Murphy, A.; Kumar, D.; Goel, S.; McClory, C. Deep learning and image data-based surface cracks recognition of laser nitrided titanium alloy. Results Eng. 2024, 22, 102003. [Google Scholar] [CrossRef] [Scilit]
- Aung, P.P.W.; Kulinan, A.S.; Park, M.; Ko, D.; Cha, G.; Park, S. Mitigating class imbalance in deep learning-based multi-class structural damage recognition using an informatics-oriented data augmentation framework. Adv. Eng. Inform. 2026, 71, 104430. [Google Scholar] [CrossRef] [Scilit]
- Raza, A.; Hanif, F. COOT-CNN: A metaheuristic-optimized deep learning framework based on lightweight convolutional architectures for multi-class robust crack detection in concrete infrastructures. Ain Shams Eng. J. 2026, 17, 104293. [Google Scholar] [CrossRef] [Scilit]
- He, X.; Liu, J.; Li, J.; Yang, Z.; Kong, X.; Zhang, Y.; Lu, Y.; Yu, Y. Toward intelligent pavement maintenance: A transferable deep learning framework for cross-domain crack segmentation and UAV-based field inspection. Adv. Eng. Inform. 2026, 73, 104582. [Google Scholar] [CrossRef] [Scilit]
- Fan, C.; Ding, Y.; Geng, F.; Lv, Z.; Yang, K. Anomaly warning classification based on deep learning method and temperature-induced effect of bridge crack monitoring data. Eng. Struct. 2026, 352, 122158. [Google Scholar] [CrossRef] [Scilit]
- Cui, J.; Lv, C.; Du, J. Real-time structural health monitoring of steel structures using acoustic emission signals and a KAN-LSTM deep learning framework. Eng. Struct. 2025, 344, 121328. [Google Scholar] [CrossRef] [Scilit]
- Ijaz, M.; Khan, S.U.R.; Rehman, A.U.; Vollmer, S.; Dengel, A.; Asim, M.N. StructDamage: A Large Scale Unified Crack and Surface Defect Dataset for Robust Structural Damage Detection. arXiv 2026, arXiv:2603.10484. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Wang, S.; Cui, C.; Zhang, D. Crack detection in strengthened steel plates based on ultrasonic guided waves and deep learning. Structures 2026, 90, 112492. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Tian, X.; Li, Q.; Ai, S. Advancing structural health monitoring: Deep learning-enhanced quantitative analysis of damage in composite laminates using surface strain field. Compos. Sci. Technol. 2024, 258, 110880. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.; Shi, X.; Ni, F.; Tian, Y. Semi-dense sub-pixel displacement measurement for structural health monitoring: A framework of deep learning-based detector-free feature matching. Measurement 2025, 254, 117899. [Google Scholar] [CrossRef] [Scilit]
- Bi, Q.; Sun, Y.; Yang, Y.; Sun, H. Integrated analysis of crack evolution in anchored jointed rock using Digital Image Correlation and deep learning-based detection. Appl. Comput. Geosci. 2026, 31, 100402. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Lin, Y.; Duan, J.; Yan, H.; Wang, X.; Xiong, X.; Huang, Y.; Guan, S.; Tao, C. Reinforcement response prediction of composite-concrete beams with crack patterns and deep learning. Comput.-Aided Civ. Infrastruct. Eng. 2025, 40, 6638–6655. [Google Scholar] [CrossRef] [Scilit]
- Ma, Y.; Jiang, S.; Chen, X.; Sun, L.; Xue, B.; Zhang, Y. A rapid prediction method for structural fatigue life based on SBFEM and deep learning. Eng. Anal. Bound. Elem. 2026, 191, 106938. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar] [CrossRef] [Scilit]
- Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On Calibration of Modern Neural Networks. Proc. Mach. Learn. Res. 2017, 70, 1321–1330. [Google Scholar]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar] [CrossRef] [Scilit]



















| Category | All | Training | Validation | Test | Masonry Focus |
|---|---|---|---|---|---|
| brick | 450 | 360 | 45 | 45 | Yes |
| cob | 100 | 80 | 10 | 10 | Yes |
| concrete | 660 | 528 | 66 | 66 | No |
| decks | 2025 | 1621 | 202 | 202 | No |
| pavements | 18,304 | 14,657 | 1809 | 1838 | No |
| road | 23,549 | 18,820 | 2367 | 2362 | No |
| stone | 100 | 80 | 10 | 10 | Yes |
| tile | 85 | 69 | 8 | 8 | Yes |
| walls | 3851 | 3081 | 385 | 385 | Yes |
| Audit Item | Measured Value |
|---|---|
| Readable images | 49,124/49,124 |
| Classes | 9 |
| Masonry-focused images | 4586 |
| Files in exact-duplicate sets | 16,896 |
| Files in repeated dHash sets | 19,683 |
| Image width range | 37–4032 px |
| Image height range | 35–4032 px |
| Cross-class exact-hash conflicts | 0 |
| Stored original group IDs crossing partitions | 0 |
| Cross-partition exact SHA-256 groups | 0 |
| Cross-partition within-class equal-dHash groups | 101 |
| Union-component-isolated test images | 4819 (107 excluded) |
| Test union components for bootstrap | 3920 |
| Method | Family | Defining Operation | Trainable Parameters |
|---|---|---|---|
| Class-Attention (epoch 0) | Attention | class-query cross-attention | 101,802 |
| Cross-Covariance | Attention | channel covariance attention | 101,802 |
| Deep-TokenFormer (epoch 0) | Attention | two transformer layers | 139,402 |
| Gated-TokenMixer | Attention | depthwise token mixing + gate | 101,802 |
| Hybrid-ConvFormer | Attention | token convolution + transformer | 101,802 |
| Pyramid-Attention (epoch 0) | Attention | class/mean/max token fusion | 101,802 |
| TokenFormer-2H | Attention | one transformer layer, 2 heads | 101,802 |
| TokenFormer-4H | Attention | one transformer layer, 4 heads | 101,802 |
| TokenFormer-8H | Attention | one transformer layer, 8 heads | 101,802 |
| Window-Attention | Attention | overlapping local token windows | 101,802 |
| CBAM-CNN | CNN | channel + spatial attention | 44,589 |
| Depthwise-CNN | CNN | depthwise-separable spatial block | 47,242 |
| Dilated-CNN | CNN | dilation-2 spatial block | 79,370 |
| GAP-MLP | CNN | global average pooling + MLP | 42,442 |
| GeM-MLP | CNN | learned generalized-mean pooling | 42,443 |
| Inception-CNN | CNN | 1 × 1/3 × 3/dilated branches | 72,722 |
| Pointwise-CNN | CNN | 1 × 1 channel mixing | 54,890 |
| Residual-CNN | CNN | two-layer residual block | 116,554 |
| SE-CNN | CNN | squeeze-excitation gating | 44,570 |
| SPP-Conv (epoch 0) | CNN | 1/2/3-level spatial pyramid | 148,938 |
| Category | Precision | Recall | F1 | Support |
|---|---|---|---|---|
| Brick | 1.000 | 1.000 | 1.000 | 45 |
| Cob | 0.800 | 0.800 | 0.800 | 10 |
| Concrete | 0.957 | 1.000 | 0.978 | 66 |
| Decks | 0.894 | 1.000 | 0.944 | 202 |
| Pavements | 0.998 | 0.987 | 0.993 | 1838 |
| Road | 0.997 | 0.997 | 0.997 | 2362 |
| Stone | 1.000 | 1.000 | 1.000 | 10 |
| Tile | 1.000 | 1.000 | 1.000 | 8 |
| Walls | 0.981 | 0.961 | 0.971 | 385 |
| Rank | Method | Family | Acc. % | Macro-F1 % | Bal. Acc. % | MCC | ECE % |
|---|---|---|---|---|---|---|---|
| 1 | Residual-CNN | CNN | 99.23 | 97.16 | 96.90 | 0.988 | 0.42 |
| 2 | Gated-TokenMixer | Attention | 99.21 | 97.13 | 96.89 | 0.987 | 0.34 |
| 3 | Inception-CNN | CNN | 99.21 | 97.13 | 96.89 | 0.987 | 0.40 |
| 4 | SE-CNN | CNN | 99.19 | 97.12 | 96.89 | 0.987 | 0.36 |
| 5 | TokenFormer-8H | Attention | 99.21 | 97.11 | 96.87 | 0.987 | 0.33 |
| 6 | Pointwise-CNN | CNN | 99.17 | 97.10 | 96.88 | 0.987 | 0.42 |
| 6 | CBAM-CNN | CNN | 99.17 | 97.10 | 96.88 | 0.987 | 0.39 |
| 8 | GAP-MLP | CNN | 99.17 | 97.08 | 96.86 | 0.987 | 0.38 |
| 9 | Dilated-CNN | CNN | 99.15 | 97.07 | 96.88 | 0.986 | 0.41 |
| 10 | GeM-MLP | CNN | 99.17 | 97.05 | 96.90 | 0.987 | 0.46 |
| 11 | Depthwise-CNN | CNN | 99.11 | 96.99 | 96.82 | 0.986 | 0.42 |
| 12 | TokenFormer-4H | Attention | 99.17 | 96.99 | 96.70 | 0.987 | 0.39 |
| 13 | TokenFormer-2H | Attention | 99.15 | 96.99 | 96.69 | 0.986 | 0.36 |
| 14 | Hybrid-ConvFormer | Attention | 99.13 | 96.98 | 96.87 | 0.986 | 0.42 |
| 14 | Cross-Covariance | Attention | 99.13 | 96.98 | 96.87 | 0.986 | 0.44 |
| 16 | Window-Attention | Attention | 99.13 | 96.97 | 96.68 | 0.986 | 0.38 |
| 17 | SPP-Conv | CNN | 99.07 | 96.48 | 97.18 | 0.985 | 0.37 |
| 17 | Deep-TokenFormer | Attention | 99.07 | 96.48 | 97.18 | 0.985 | 0.37 |
| 17 | Class-Attention | Attention | 99.07 | 96.48 | 97.18 | 0.985 | 0.37 |
| 17 | Pyramid-Attention | Attention | 99.07 | 96.48 | 97.18 | 0.985 | 0.37 |
| Measure | Estimate | 95% Component CI/Note |
|---|---|---|
| Accuracy | 99.21% | 98.95–99.45% |
| Macro-F1 | 97.11% | 94.41–98.89% |
| Balanced accuracy | 96.87% | 93.58–99.33% |
| MCC | 0.987 | Not bootstrapped |
| Macro-AUROC | 0.99980 | Not bootstrapped |
| Calibrated ECE | 0.33% | 15 equal-width bins |
| Method | Family | Seeds | Accuracy %, Mean ± SD | Macro-F1 %, Mean ± SD | ECE %, Mean |
|---|---|---|---|---|---|
| TokenFormer-8H | Attention | 3 | 99.154 ± 0.077 | 97.040 ± 0.146 | 0.371 |
| Gated-TokenMixer | Attention | 3 | 99.147 ± 0.054 | 97.039 ± 0.081 | 0.408 |
| Window-Attention | Attention | 3 | 99.141 ± 0.042 | 96.997 ± 0.130 | 0.385 |
| Condition | Accuracy % | Macro-F1 % | Bal. Acc. % | ECE % | Acc. Drop pp |
|---|---|---|---|---|---|
| Clean | 99.21 | 97.11 | 96.87 | 0.33 | 0.00 |
| Low Light | 92.29 | 78.01 | 74.40 | 6.13 | 6.92 |
| Gaussian Noise | 87.41 | 68.35 | 65.48 | 10.77 | 11.79 |
| Blur | 76.63 | 46.81 | 52.32 | 20.18 | 22.57 |
| Pixelation | 93.26 | 78.25 | 75.83 | 4.48 | 5.95 |
| Low Contrast | 95.94 | 88.89 | 84.57 | 2.32 | 3.27 |
| Experiment/Scope | Images | Accuracy % | Macro-F1 % |
|---|---|---|---|
| TokenFormer-8H: original, 9 classes | 4926 | 99.21 | 97.11 |
| TokenFormer-8H: hash-isolated, 9 classes | 4819 | 99.19 | 97.10 |
| TokenFormer-8H: masonry-focused, 5 labels | 458 | 96.07 | 97.35 |
| Source-folder majority: original, 9 classes | 4926 | 97.97 | 81.85 |
| ImageNet-only probe: original, 9 classes | 4926 | 92.94 | 75.20 |
| ImageNet-only probe: original, eligible-source subset | 2700 | 94.37 | 88.31 |
| ImageNet-only probe: held-out folders, 3 labels | 26,953 | 83.29 | 61.02 |
| Calibrator | Acc. % | Macro-F1 % | NLL | Equal-Class NLL | Brier | ECE % |
|---|---|---|---|---|---|---|
| Raw | 99.21 | 97.11 | 0.1451 | 0.1573 | 0.0291 | 11.18 |
| Scalar temperature | 99.21 | 97.11 | 0.0398 | 0.1485 | 0.0138 | 0.33 |
| Balanced temperature | 99.21 | 97.11 | 0.0475 | 0.1259 | 0.0138 | 1.57 |
| Regularized vector | 99.29 | 96.52 | 0.0295 | 0.1524 | 0.0122 | 0.42 |
| Replaced Area | Grad-CAM Drop | Random Drop | Paired Difference | 95% Interval |
|---|---|---|---|---|
| 10.00% | 0.081 | 0.113 | −0.031 | −0.065 to 0.001 |
| 20.00% | 0.117 | 0.221 | −0.104 | −0.146 to −0.061 |
| 30.00% | 0.178 | 0.289 | −0.111 | −0.157 to −0.064 |
| Condition | Accuracy % | Macro-F1 % | Bal. Acc. % | ECE % |
|---|---|---|---|---|
| Gamma 0.7 | 98.84 | 95.27 | 94.15 | 0.49 |
| Gamma 1.5 | 99.09 | 95.60 | 94.88 | 0.29 |
| Spatial shadow | 85.73 | 68.04 | 69.21 | 12.60 |
| White balance | 88.23 | 73.52 | 72.72 | 10.30 |
| Perspective | 85.71 | 73.88 | 75.99 | 10.81 |
| JPEG quality 30 | 93.91 | 82.19 | 85.43 | 3.83 |
| Method | Role in Comparison | Family | Trainable Params | Acc. % | Macro-F1 % | Bal. Acc. % | ECE % |
|---|---|---|---|---|---|---|---|
| Residual-CNN | Highest test score (descriptive) | CNN | 116,554 | 99.23 | 97.16 | 96.90 | 0.42 |
| Gated-TokenMixer | Best attention test macro-F1 | Attention | 101,802 | 99.21 | 97.13 | 96.89 | 0.34 |
| Inception-CNN | Tied second test macro-F1 | CNN | 72,722 | 99.21 | 97.13 | 96.89 | 0.40 |
| TokenFormer-8H | Validation-selected | Attention | 101,802 | 99.21 | 97.11 | 96.87 | 0.33 |
| GAP-MLP | Compact comparator | CNN | 42,442 | 99.17 | 97.08 | 96.86 | 0.38 |
| Outcome | Required Evidence | Status in This Study |
|---|---|---|
| Surface-category recognition | Image-level supplied category labels | Evaluated; subject to source and class imbalance |
| Crack mechanism/origin | Expert labels plus loading, restraint and history | Not evaluated; no verified targets |
| Length or opening width in mm | Crack delineation, scale and perspective calibration | Not evaluated; no physical scale |
| Element location and stress state | Registered coordinates, geometry and structural model | Not evaluated; no spatial or loading data |
| Joint versus crack rejection | Annotated joints, cracks and intact negatives | Not evaluated; no such label classes |
| Residual life/hazard | Structural assessment and longitudinal validation | Not supported by surface predictions |
| Field generalization | Independent site/device cohort and expert reference | Not established; retrospective diagnostics only |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Thango, B.A.; Thango, S.G. Deep Learning for Surface-Conditioned Structural Crack Recognition: Integrating CNN-Attention Models and Calibrated Confidence Estimation. Appl. Sci. 2026, 16, 9337. https://doi.org/10.3390/app16189337
Thango BA, Thango SG. Deep Learning for Surface-Conditioned Structural Crack Recognition: Integrating CNN-Attention Models and Calibrated Confidence Estimation. Applied Sciences. 2026; 16(18):9337. https://doi.org/10.3390/app16189337
Chicago/Turabian StyleThango, Bonginkosi A., and Sipho G. Thango. 2026. "Deep Learning for Surface-Conditioned Structural Crack Recognition: Integrating CNN-Attention Models and Calibrated Confidence Estimation" Applied Sciences 16, no. 18: 9337. https://doi.org/10.3390/app16189337
APA StyleThango, B. A., & Thango, S. G. (2026). Deep Learning for Surface-Conditioned Structural Crack Recognition: Integrating CNN-Attention Models and Calibrated Confidence Estimation. Applied Sciences, 16(18), 9337. https://doi.org/10.3390/app16189337

