Skip to Content
InfrastructuresInfrastructures
  • Review
  • Open Access

20 July 2026

Deep Learning-Based Surface Crack Detection in Bridge Structures: A Review

,
,
and
1
School of Civil Engineering, Universiti Sains Malaysia, Engineering Campus, Nibong Tebal 14300, Pulau Pinang, Malaysia
2
School of Civil Engineering, Sichuan University of Science and Engineering, Zigong 643000, China
*
Author to whom correspondence should be addressed.

Abstract

Surface cracks pose a significant threat to the durability and operational safety of concrete bridges, making accurate crack detection essential for effective structural health monitoring. Conventional manual inspection is limited by low efficiency, subjective judgment, and safety risks during field operations. Although deep learning computer vision techniques have become the dominant approach for automated crack detection, existing review studies provide limited discussion of their adaptability to different bridge structures and offer insufficient guidance for practical engineering applications. This review systematically summarizes recent advances in deep learning methods for surface crack detection in concrete bridges. Existing approaches are classified into crack classification, object detection, and semantic segmentation. The review further examines the relationships among bridge geometry, crack morphology, inspection conditions, and model performance. It also compares the technical challenges associated with beam, arch, cable stayed, and suspension bridges and discusses suitable model adaptation strategies for different structural characteristics. The analysis shows that bridge specific model architectures can significantly improve detection accuracy and robustness under complex inspection conditions. However, several challenges remain, including the limited availability of high quality datasets, data imbalance, inadequate detection of small cracks on curved surfaces under multiple viewing conditions, and the high cost of field deployment. This review provides practical guidance for selecting appropriate deep learning techniques for concrete bridge inspection and highlights future research directions toward integrating advanced sensing, multimodal data fusion, and intelligent inspection technologies to achieve more accurate, efficient, and reliable bridge condition assessment.

1. Introduction

Concrete bridges play a vital role in modern transportation networks by providing essential links across rivers, valleys, roads, and railways. As critical components of transportation infrastructure, they support socioeconomic development while ensuring the safety and efficiency of public travel [1,2]. According to China’s 2024 Statistical Bulletin on the Development of the Transportation Industry, both the number and the total length of highway bridges have continued to increase (Figure 1). By the end of 2024, China had 1.1081 million highway bridges with a combined length of 101.9758 million linear meters, representing annual increases of 28,700 bridges and 6.6876 million linear meters, respectively. Among these, 11,329 were extra-large bridges, with a total length of 20.6047 million linear meters, while 191,400 were classified as large bridges, totaling 53.9705 million linear meters [3,4].
Figure 1. Total number of highway bridges in the country.
Despite their importance, concrete bridges inevitably deteriorate over time due to environmental exposure, repeated traffic loading, and climatic effects. This deterioration is commonly manifested as surface defects, including cracking, spalling, efflorescence, and reinforcement corrosion [5,6]. As these defects develop over time, they progressively compromise structural integrity, shorten service life, and increase the likelihood of costly maintenance, structural failure, and safety-related incidents.
Regular bridge inspection is therefore essential for evaluating structural condition and ensuring operational safety [7]. Current bridge inspection standards identify surface cracks as one of the most important indicators of structural condition [8]. Once initiated, cracks may propagate under service loads, leading to stress concentrations, fatigue damage, and accelerated corrosion of reinforcement. These processes degrade material properties, reduce fracture resistance, weaken durability, and ultimately increase structural risk. It should also be emphasized that visual crack detection represents only one component of a comprehensive bridge inspection system, which also includes non-destructive testing (NDT), structural health monitoring, and structural system identification. This review focuses specifically on vision-based crack detection using deep learning techniques.
To maintain the long-term performance of bridge infrastructure, comprehensive bridge maintenance systems have been established in many countries. These systems combine routine inspections, daily patrols, structural health monitoring, maintenance, and rehabilitation to ensure structural safety and serviceability throughout the bridge life cycle.
For many years, surface cracks in concrete bridges have been identified primarily through manual visual inspection. Although widely adopted, this approach is labor-intensive, time-consuming, and highly dependent on inspectors’ experience, often resulting in inconsistent assessments and potential safety risks [9,10,11]. More recently, UAVs, climbing robots, and other mobile inspection platforms have been introduced to collect large volumes of bridge images while reducing the risks associated with field inspections. Nevertheless, the acquired images and videos still generally require manual interpretation, which limits inspection efficiency and makes large-scale, routine bridge monitoring difficult to achieve [12,13].
The rapid advancement of computer vision has fundamentally transformed crack detection in concrete structures by enabling image-driven analysis and substantially reducing reliance on manual inspection. Image-based crack detection methods can be broadly grouped into two main categories: traditional image processing techniques and data-driven machine learning approaches [14]. Early studies on concrete bridge surface inspection predominantly relied on classical image processing pipelines. These approaches typically involve sequential operations such as grayscale conversion, thresholding (binarization), filtering, edge detection (e.g., Canny, Roberts, and Sobel operators), and intensity normalization to extract crack-related features [15]. Despite their conceptual simplicity, these methods are highly sensitive to environmental disturbances, including uneven illumination, surface noise, and staining effects. More critically, their dependence on handcrafted rules and empirically tuned parameters significantly constrains their generalization capability, particularly in complex in situ bridge environments.
To overcome these limitations, conventional machine learning methods such as support vector machines and random forests were subsequently introduced. These methods improve classification performance by leveraging manually engineered feature representations [16]. However, their performance is inherently bounded by the discriminative quality of the extracted features, which are often difficult to generalize across varying bridge conditions. Moreover, such models exhibit limited hierarchical representation capability, making them less robust when confronted with heterogeneous backgrounds and structural variability. In contrast, deep learning approaches have emerged as a more powerful paradigm by eliminating the need for manual feature engineering [17]. Through end-to-end learning, these models can automatically extract multi-level feature representations and have demonstrated superior performance in structural crack detection tasks, particularly in scenarios with complex texture patterns and variable environmental conditions [18].
With the rapid evolution of deep learning architectures, the integration of deep learning with digital image processing has become a dominant research direction for bridge surface damage detection [19,20]. In typical implementations, high-resolution bridge images collected using UAVs are used to train convolutional neural networks (CNNs) or related deep architectures for automated damage recognition. Once trained, these models can be deployed on edge devices or embedded platforms mounted on UAVs, enabling near real-time inspection and decision support. Compared with traditional inspection strategies, this paradigm significantly enhances detection efficiency and consistency while reducing dependency on subjective human judgment. Furthermore, it enables scalable post-processing workflows, including automated defect localization and prioritization for manual verification, thereby substantially reducing inspection workload and operational costs [21,22].
Despite substantial progress, recent review studies on deep learning-based crack detection remain limited in their specificity to bridge engineering. For instance, Hamishebahar et al. [23] primarily concentrated on algorithmic architectures and benchmark performance comparisons within computer vision, with insufficient consideration of structural engineering constraints that directly influence real-world deployment in bridge inspection scenarios. Similarly, Zhuang et al. [24] addressed a broad range of civil infrastructure systems. Still, they treated bridges as a generic application domain, without distinguishing the heterogeneous inspection requirements associated with different bridge types. Bai et al. [25] adopted an even broader infrastructure perspective, covering roads, dams, and bridges within a unified framework, yet failing to establish explicit links between structural characteristics and algorithmic behavior, thereby limiting their interpretability for bridge-specific applications. In addition, Prakash et al. [26] focused predominantly on vibration-based structural health monitoring techniques within bridge systems, while only marginally addressing vision-based surface crack detection, leaving a critical gap in image-based inspection analysis.
Overall, existing review literature tends to generalize crack detection methodologies across concrete structures while underrepresenting the substantial variability among bridge types in structural configuration, crack morphology, and inspection constraints. This oversimplification restricts the ability of existing studies to inform model selection, adaptation strategies, and deployment design for specific bridge conditions [27,28]. More importantly, current research often emphasizes algorithmic performance in isolation, without adequately integrating bridge-engineering context and operational constraints, which limits its translational value for practical infrastructure monitoring and future research development [29,30].
To ensure methodological rigor, transparency, and reproducibility, a comprehensive literature retrieval strategy was designed and implemented. The review covered studies published from 2016 onwards, a period that coincided with the rapid development of deep learning-based crack detection techniques. Major academic databases, including Web of Science, Scopus, IEEE Xplore, and Google Scholar, were systematically searched to ensure broad coverage of relevant literature. The search strategy targeted advancements in deep learning applications for bridge crack detection and aimed to synthesize representative research outcomes in this domain.
A combination of keywords was employed to construct search queries, including bridge cracks, deep learning, crack classification, object detection, crack segmentation, CNNs, data fusion, and literature review. To further improve coverage, a backward snowballing approach was adopted by screening reference lists of highly relevant review articles, ensuring that influential and potentially overlooked studies were also included. To guarantee the academic quality and relevance of the selected literature, explicit inclusion and exclusion criteria were defined. Studies were included if they focused on deep learning-based surface crack detection in concrete bridges, were published in peer-reviewed English-language journals, and reported complete experimental setups with verifiable performance metrics. Studies were excluded if they addressed non-visual inspection methods, other types of infrastructure without bridge-specific analysis, duplicate publications, or works lacking original experimental validation.
The screening process followed a three-stage protocol, including deduplication, title and abstract screening, and full-text evaluation. After systematic filtering, 168 core original research articles were selected for detailed analysis. In addition, foundational theoretical studies, relevant industrial standards, and dataset-related publications were incorporated to support contextual understanding, resulting in a total of 199 references used in this review.
To address the aforementioned research gaps, this paper systematically reviews the evolution of deep learning methods for detecting cracks on the surfaces of concrete bridges. For the first time, it compares the crack-detection characteristics and applicability of methods for typical concrete bridges across different structural types, proposing customized model-adaptation rules and technical approaches to address the existing gap, in which previous reviews have emphasized algorithmic paradigms while neglecting structural characteristics. Additionally, it provides an in-depth analysis of specific challenges unique to concrete bridges. It identifies future research directions aligned with practical engineering needs, offering a comprehensive reference for both academic research and technological implementation in this field.

2. Application of Deep Learning Method in Bridge Crack Detection

2.1. Deep Learning Methods

As a subfield of machine learning, deep learning employs deep neural networks to learn hierarchical feature representations directly from data, providing a fundamental technological basis for crack detection in civil engineering structures. Its conceptual origin can be traced back to the McCulloch–Pitts (MP) neuron model proposed in 1943. Subsequently, Rumelhart et al. introduced the backpropagation (BP) algorithm in 1986 to enable effective training of multilayer neural networks by learning distributed representations, thereby overcoming key limitations in early neural computation frameworks [31,32]. However, due to computational resource constraints and limited dataset availability at the time, neural network research experienced a prolonged period of stagnation.
The resurgence of deep learning was marked by the introduction of deep architectures in the mid-2000s. In particular, Hinton and Salakhutdinov proposed the Deep Belief Network (DBN) in 2006, which mitigated the vanishing gradient problem through layer-wise pretraining and demonstrated the feasibility of training deep models, thereby laying the groundwork for subsequent developments in representation learning [33].
A series of landmark advances in computer vision further accelerated the adoption of deep learning. The introduction of AlexNet in 2012, CNNs, achieved breakthrough performance in the ImageNet competition, demonstrating the strong capability of deep architectures for large-scale visual recognition tasks [34,35]. This success significantly influenced feature learning strategies for image-based crack detection in civil engineering. In object detection, the R-CNN framework proposed in 2013 shifted the paradigm from handcrafted feature design to data-driven region-based representation learning, establishing a foundation for subsequent crack localization methods in structural inspection tasks [36]. Shortly thereafter, one-stage detection frameworks such as YOLO (e.g., YOLOv1 introduced in 2016) further improved detection efficiency by reformulating object detection as a unified regression problem, enabling real-time performance suitable for practical deployment scenarios such as UAV-based bridge inspection systems [37,38].
More recently, the increasing demand for large-scale annotated datasets has driven the development of self-supervised learning paradigms. These approaches reduce reliance on manual labeling by enabling models to learn meaningful representations from unlabeled data, thereby improving data efficiency and generalization capability in complex real-world environments [39,40,41].
Collectively, these methodological advances constitute the core technical foundation for modern deep learning-based bridge inspection systems. They have significantly improved the accuracy, robustness, and real-time performance of concrete crack detection under complex field conditions. The evolutionary trajectory of these key technologies and their application in bridge crack detection is illustrated in Figure 2.
Figure 2. The Evolution of Deep Learning and Its Application in Bridge Crack Detection [31,32,34,36,37,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59].
Image classification, object detection, and semantic segmentation constitute the three fundamental paradigms for crack detection on concrete bridge surfaces. These tasks correspond to progressively increasing levels of engineering requirement: rapid condition screening, crack localization with statistical analysis, and pixel-level quantitative characterization, respectively.
The primary challenges in crack detection arise from both intrinsic image characteristics and dataset distribution properties. From a visual perspective, cracks typically exhibit low contrast relative to surrounding concrete surfaces and are highly sensitive to environmental disturbances, including stains, illumination variations, and shadow effects. From a geometric standpoint, crack patterns are often highly irregular, with extreme aspect ratios and discontinuous spatial structures. In addition, publicly available datasets are frequently characterized by class imbalance and insufficient coverage of diverse structural conditions, which collectively impose stringent requirements on algorithmic robustness, feature discriminative power, and cross-domain generalization performance [60,61].
CNNs, as illustrated in Figure 3, have become the dominant framework for addressing these challenges due to their hierarchical feature extraction capability and parameter-sharing mechanism. Compared with traditional image processing techniques, CNN-based methods significantly improve representation learning and have therefore become the mainstream paradigm for bridge crack detection. Given that the fundamental architecture and operational principles of CNNs have been extensively documented in the literature, they are not repeated in detail here [62,63].
Figure 3. The architecture of VGG-16 [62].
Despite their strong representational power, CNN-based models often have a large number of parameters, which can lead to overfitting, particularly with limited or imbalanced training data. To alleviate this issue, architectural improvements such as global average pooling (GAP) have been widely adopted to reduce model complexity and improve generalization. In parallel, the availability of open-source deep learning frameworks such as TensorFlow and PyTorch has significantly lowered the barrier for model development, training, and deployment, thereby accelerating the practical adoption of CNN-based crack detection systems.
At present, CNN-based deep learning methods have been widely extended to crack detection tasks across various infrastructure domains, including tunnels [64], roads [65], and bridges [64]. Building upon this foundation, this section systematically reviews three core detection paradigms—classification, object detection, and semantic segmentation. The corresponding architectural components and representative models are summarized in Figure 4.
Figure 4. CNN-based crack detection architecture.

2.2. Crack Classification

Crack classification is a fundamental prerequisite for detecting surface cracks in concrete bridges. Deep learning-based classification approaches typically operate on images or image patches as basic input units. They can be broadly divided into binary classification (i.e., crack versus non-crack identification) and multi-class classification (i.e., discrimination among different crack types). The former is primarily used for rapid large-scale image screening, while the latter provides prior semantic information that supports subsequent localization and segmentation tasks [60,65].
CNNs have become the dominant framework for crack classification due to their strong feature representation capability and robust learning performance [66,67]. The development of CNN-based classification methods can be broadly summarized along three main directions. First, significant progress has been made in foundational network architectures. Early studies typically employed shallow, task-specific CNNs for crack feature extraction and classification. For example, Schmugge et al. [68] directly applied CNNs for crack identification, while Cha et al. [69] extended patch-based CNN classification to jointly achieve crack recognition and approximate spatial localization through image reconstruction strategies. Subsequently, more advanced pre-trained architectures, including VGGNet [69], GoogLeNet [70], ResNet and DenseNet [70], were introduced and widely adopted. Through transfer learning and fine-tuning strategies, these models effectively alleviate the problem of limited labeled data in crack datasets while substantially improving feature representation and generalization capability.
Second, considerable efforts have focused on improving robustness in complex and interference-rich environments. In practical bridge inspection scenarios, crack images are often affected by stains, shadows, uneven illumination, and surface texture noise. To address these challenges, researchers have incorporated mechanisms such as attention modules and multi-scale feature fusion to enhance the discrimination of crack-related regions while suppressing background interference [71]. For instance, Bhattacharya et al. [71] developed a CNN framework integrating multi-scale attention mechanisms to emphasize crack-relevant features selectively; Yang et al. [72] employed a VGG16-based transfer learning strategy to improve robustness against surface contamination and illumination variations; and Xu et al. [73] proposed a feature fusion-based CNN architecture capable of accurately identifying cracks in bridge steel box girders under complex real-world interference conditions, including markings and surface artifacts. Third, increasing attention has been given to lightweight model design for engineering deployment. Lightweight architectures such as MobileNet significantly reduce computational complexity and parameter size while maintaining competitive classification performance, thereby enabling real-time implementation on resource-constrained platforms such as unmanned aerial vehicles and embedded inspection systems [74,75,76].
Overall, CNN-based architectures remain the most widely adopted approach for crack classification due to their strong performance in small-sample and visually complex environments, as well as their mature ecosystem of pre-trained models and frameworks. However, it is important to note that classification methods inherently provide only image-level labels without spatial or geometric information regarding crack morphology. Consequently, they are generally used as preliminary screening tools within inspection pipelines and cannot replace detection- or segmentation-based approaches for detailed structural assessment.

2.3. Crack Object Detection

Object detection enables both spatial localization and category recognition of cracks, making it a core technique for bridge surface inspection. From an architectural perspective, existing methods can be broadly categorized into anchor-based and anchor-free approaches. Anchor-based methods are further divided into two-stage and single-stage paradigms. Two-stage detectors, represented by the R-CNN family [77], first generate region proposals and subsequently perform classification and bounding-box regression. Although these methods generally achieve high detection accuracy, they are computationally expensive and therefore less suitable for real-time inspection scenarios. In contrast, single-stage detectors, such as YOLO [78] and SSD [79], reformulate detection as a unified regression problem, enabling end-to-end inference with significantly improved computational efficiency. Anchor-free methods, including CornerNet [80] and CenterNet [81], eliminate predefined anchor boxes and instead perform detection via keypoint estimation. These approaches show potential for handling irregular and small-scale cracks; however, they still face challenges in precisely localizing boundaries under complex background conditions.
The evolution of two-stage detectors reflects a continuous trade-off between accuracy and efficiency. Early R-CNN models suffered from substantial computational redundancy, which was progressively alleviated by Fast R-CNN [81] and Faster R-CNN [82], which introduced shared convolutional feature extraction and region proposal networks within an end-to-end training paradigm. These improvements established two-stage detectors as a benchmark framework for high-precision crack detection tasks. Nevertheless, despite ongoing engineering optimizations—including improved robustness under interference, UAV-oriented adaptation, enhanced non-maximum suppression strategies, and model compression techniques—their inference efficiency remains a limiting factor for large-scale real-time bridge inspection [82].
Single-stage detectors have therefore become the dominant choice in practical engineering applications due to their superior balance between accuracy and speed. Among them, the YOLO series has been most widely adopted in bridge crack detection tasks [83,84]. Recent research efforts have primarily focused on three aspects. First, lightweight design strategies aim to reduce computational cost for edge deployment by adopting efficient convolutional structures such as depthwise separable convolutions and inverted residual blocks. For example, Yang et al. [85] enhanced the YOLO framework by integrating GhostBottleneck, ECA-Net, and ASFF modules, achieving high detection accuracy while maintaining real-time performance on edge devices. Similarly, Pandey et al. [86] proposed a lightweight YOLOv11 variant that improves both accuracy and inference efficiency across multiple benchmark datasets. Second, feature enhancement strategies improve representation capability in complex environments through multi-scale fusion and attention mechanisms. Zhou et al. [87] optimized YOLOv3 by strengthening feature aggregation within the detection head, while Yanna et al. [88] incorporated shallow feature fusion and coordinate attention into YOLOX, thereby improving performance in multi-class defect detection. Third, detection strategy refinement—particularly improvements in loss functions and label assignment mechanisms—has further enhanced localization accuracy [89]. Through continuous architectural evolution from YOLOv5 to YOLOv11, together with the integration of anchor-free and attention-based mechanisms, the YOLO family has significantly improved both robustness and engineering adaptability in bridge inspection applications [90,91,92,93,94].
Beyond algorithmic development, YOLO-based detection systems have formed a relatively complete engineering pipeline spanning data preparation, model training, and multi-platform deployment. In practice, datasets are typically constructed and annotated using tools such as LabelImg and LabelMe, which support generating standardized YOLO-format annotations [85]. Model training is commonly implemented in the PyTorch ecosystem, enabling flexible adaptation to custom datasets and network architectures. For deployment, a multi-level hardware–software strategy is typically adopted: high-performance server-side inference relies on TensorRT for GPU acceleration, while edge computing scenarios utilize optimized inference engines such as ONNX Runtime and OpenVINO. For mobile or embedded platforms, lightweight deployment frameworks including NCNN and MNN are widely used to support real-time crack detection under resource-constrained conditions [88].
To further evaluate model performance under consistent conditions, recent studies have increasingly adopted unified benchmarking strategies. Using datasets such as MCDS and CODEBRIM, Huang et al. [95] constructed a standardized evaluation framework to compare multiple generations of lightweight detection models under identical training configurations and hardware environments. The comparative results, summarized in Table 1, demonstrate that continuous algorithmic refinement not only improves detection accuracy but also significantly enhances real-time inference capability, confirming the effectiveness of iterative model design for practical bridge inspection applications.
Table 1. Performance comparison of lightweight bridge damage detection algorithms.
Overall, object detection methods are effective for rapidly localizing crack regions; however, their reliance on rectangular bounding boxes limits their ability to accurately represent the highly elongated, curvilinear, and discontinuous geometry of cracks. In addition, bounding boxes often include substantial background pixels, which reduces the precision of subsequent quantitative analyses of crack morphology. Furthermore, in complex inspection environments, visually similar surface features such as construction joints, stains, and texture variations may be mistaken for cracks, leading to false positives and reduced detection reliability. These limitations in fine-grained geometric representation and semantic discrimination have motivated the development of pixel-level semantic segmentation methods, which provide more precise delineation of crack boundaries and enable more accurate assessment of structural condition.

2.4. Crack Segmentation

Semantic segmentation enables pixel-wise extraction of crack regions. It facilitates quantitative characterization of crack geometry, thereby providing more detailed spatial representation and higher localization accuracy than image classification and object detection methods. As such, it has become a key methodological framework for fine-grained assessment of bridge cracks.
Early segmentation approaches primarily relied on traditional image processing techniques and handcrafted feature engineering. However, these methods generally exhibit limited robustness under complex environmental conditions and poor generalization capability across different scenes. The introduction of the fully convolutional network (FCN) by Long et al. [101] marked a fundamental transition toward deep learning-based semantic segmentation. By replacing fully connected layers with convolutional operations and introducing deconvolution-based upsampling, FCN enabled end-to-end dense prediction for inputs of arbitrary size, thereby establishing the foundational framework for modern semantic segmentation.
In bridge inspection applications, early studies demonstrated the feasibility of FCN-based approaches for crack segmentation in complex scenes. For instance, Yang et al. [102] applied FCN to achieve multi-scale crack segmentation in bridge images, while Liang et al. [103] proposed a dual-stage CNN–FCN framework that first suppresses interference regions and then performs fine-grained segmentation, thereby improving robustness in noisy real-world bridge environments.
Subsequent research has progressively advanced toward more powerful encoder–decoder architectures, with U-Net emerging as a representative framework. Building on this, recent studies have further incorporated multi-scale feature fusion, attention mechanisms, and Transformer-based modules to enhance contextual modeling and boundary refinement. The overall evolution of semantic segmentation architectures for bridge crack detection is summarized in Figure 5.
Figure 5. Evolution of Deep Learning-Based Semantic Segmentation Models [98,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118].
To address common challenges in bridge crack segmentation, including highly elongated crack morphology, significant scale variation, and strong background interference, recent research has primarily focused on four complementary directions.
First, significant efforts have been devoted to enhancing multi-scale feature representation and contextual modeling through attention mechanisms. Techniques such as dilated convolutions, spatial pyramid pooling, and various attention modules enable more effective aggregation of global and local contextual information, thereby improving segmentation robustness in complex scenes with fine and discontinuous cracks [119,120,121,122,123,124,125]. Representative architectures, including BC-DUnet [126], GTAU-UNet [127], BridgeNet [57], BridgeHealthNet [128], and Mask R-CNN-based variants [129,130,131,132,133], have demonstrated improved segmentation performance in bridge inspection scenarios.
Second, to mitigate the limitations imposed by scarce labeled data, research on data-efficient learning and augmentation strategies has gained increasing attention. Two-stage segmentation frameworks, such as the one proposed by Tabernik et al. [131], improve learning stability with limited annotations. In parallel, GAN-based generative approaches have been explored for self-supervised or semi-supervised learning, as demonstrated by Zhang et al. [132]. Pixel-level augmentation strategies for small-sample segmentation have also been investigated, including UNet-based methods [134]. High-fidelity synthetic data generation techniques, such as StyleSPADE [127], further enhance model generalization under limited training, increasing emphasis has been placed on lightweight model design for real-time edge deployment. To balance computational efficiency and segmentation accuracy, several hybrid frameworks have been proposed. For example, Gang et al. [135] developed a lightweight detection–segmentation collaborative architecture for real-time inspection tasks. LI et al. [128] proposed a MobileNetV2-GCN-based lightweight segmentation model, while YU et al. [129] integrated YOLOv5 with UNet 3+ to achieve a trade-off between inference speed and segmentation accuracy, thereby meeting the requirements of practical engineering deployment.
Fourth, recent advances have introduced Transformer-based architectures to address the limitations of CNNs in modeling long-range dependencies. By leveraging self-attention mechanisms, models such as CrackFormer [136] and DefNet [137] can capture global contextual relationships in crack structures, thereby improving the segmentation of slender and discontinuous cracks and overcoming the inherent locality constraints of convolution-based feature extraction.
Overall, Table 2 summarizes representative semantic segmentation models in terms of architectural design, key performance characteristics, and applicable scenarios, providing a comparative reference for model selection in bridge engineering applications and highlighting potential directions for future research.
Table 2. Comparison of Representative Deep Learning Models in Existing Literature.
The results summarized in Table 2 indicate that the U-Net baseline model generally achieves competitive segmentation performance compared with representative architectures such as SegFormer, HRNet, and PSPNet, particularly in terms of accuracy with limited training data. Its improved variant, U-Net3+, exhibits performance comparable to more advanced models such as DeepLabv3+, while maintaining relatively lower computational complexity and reduced deployment overhead.
The strong performance of U-Net can be attributed to its encoder–decoder architecture with skip connections, which effectively preserves multi-scale spatial information and enhances the recovery of fine-grained structural details. This design is particularly advantageous for crack segmentation tasks, where thin, discontinuous, and low-contrast features must be accurately captured. In addition, its relatively high data efficiency makes it well-suited for small-sample engineering scenarios commonly encountered in concrete bridge inspection applications.
Despite these advantages, current bridge crack segmentation research still faces several persistent challenges, including data quality limitations, real-time processing constraints, and gaps between laboratory performance and the requirements of practical engineering deployment.

2.5. Dataset

Datasets constitute a critical foundation for the training and evaluation of deep learning models, as their image quality, scene diversity, annotation accuracy, and data distribution characteristics directly influence model generalization capability and detection performance. In particular, insufficient diversity or imbalanced sample distributions may significantly degrade model robustness under real-world bridge inspection conditions.
Table 3 summarizes the key characteristics of commonly used datasets for crack classification, object detection, and semantic segmentation in concrete structures, while Table 4 presents widely adopted evaluation metrics for performance assessment. Together, these standardized benchmarks provide a consistent reference framework for fair model comparison and quantitative validation across different methodological studies.
Table 3. Partial public datasets for concrete bridge crack detection.
Table 4. Evaluation Metrics for Deep Learning-Based Crack Detection Algorithms.
Current specialized datasets for bridge crack analysis still present notable limitations. Most publicly available datasets are primarily focused on road surface cracks, while bridge-specific datasets remain relatively limited in both scale and structural diversity. Although dedicated datasets such as BridgeDamage have partially addressed this gap by covering small- and medium-span concrete bridges, they still do not adequately represent more complex structural forms, including steel–concrete composite beams, arch bridges, and long-span cable-stayed bridges [57].
In addition, the high cost of pixel-level annotation in semantic segmentation tasks results in substantially smaller dataset sizes compared with those used for classification and object detection. This limitation is particularly pronounced in bridge-specific segmentation datasets, where data scarcity and incomplete structural coverage remain critical challenges. As a result, dataset availability remains a major bottleneck, limiting the practical deployment and generalization capabilities of deep learning-based bridge crack segmentation methods.

3. Crack Detection Technology of Different Bridge Structure Forms

Several established technical approaches are available for crack detection in civil engineering practice. In addition to image-based methods, contact-type and embedded sensing technologies are widely adopted, including linear variable differential transformer (LVDT) displacement sensors, vibrating-wire crack meters, and fiber Bragg grating (FBG) sensors. These methods enable direct measurement of crack-induced displacement, thereby providing high measurement accuracy and reduced susceptibility to environmental disturbances. However, their application is generally limited to point-based or localized monitoring, and they are often associated with high installation and maintenance costs. Moreover, achieving full spatial coverage across large-scale bridge structures remains challenging.
In contrast, vision-based crack detection methods offer non-contact operation, wide-area coverage, and flexible deployment, making them particularly suitable for large-scale inspection tasks on mobile platforms such as UAVs. In practical engineering applications, these approaches are often considered complementary to sensor-based monitoring systems rather than direct substitutes. Against this background, this chapter focuses on vision-based approaches, providing a systematic review of deep learning-based crack detection methods and their applications across various bridge structures, and discussing key challenges and research gaps.

3.1. Beam Bridge

Beam and slab bridges are among the most common structural forms in highway and urban bridge systems and are primarily subjected to bending- and shear-dominated stress states. These two bridge types have therefore served as early and widely studied application scenarios for deep learning-based crack detection methods.
In beam bridges, cracking is typically characterized by vertical flexural cracks in the mid-span region and inclined shear cracks near support zones. In slab bridges, due to the combined effects of bidirectional bending, fatigue loading, and shrinkage, cracks tend to be more densely distributed, exhibit more complex orientations, and are generally finer in scale [157].
Compared with more complex bridge types, beam and slab bridges typically exhibit relatively simple surface geometries, stable imaging perspectives, and lower levels of background interference. As a result, early studies commonly adopted fixed cameras or close-range photography for data acquisition. However, in practical inspection scenarios, cracks located on bridge decks and underside slabs are often partially occluded by box-girder interiors or shadowed regions beneath the structure. In addition, the prevalence of fine-scale cracks with low contrast increases the likelihood of confusion with surface textures or stains. These challenges are further exacerbated by the limited computational resources available on inspection equipment for small- and medium-span bridges, necessitating the use of lightweight model architectures [158].
To address the need for accurate crack delineation and quantitative geometric characterization, U-Net-based semantic segmentation methods have been widely adopted. By enabling pixel-level separation between crack regions and background, these approaches provide explicit crack contours and support quantitative evaluation of geometric parameters such as crack length and width [159]. Existing studies have further benchmarked representative segmentation models on beam-bridge datasets, and Table 5 summarizes their performance comparisons on widely used datasets such as Crack500 and SDNET2018.
Table 5. Comparison of Performance of the Segmented Network for Beam Bridges.
Comparative results indicate that although UNet3+ achieves a higher mIoU, it incurs a substantially increased computational burden: approximately 4 times as many parameters as the lightweight UNet variant and a reduction in inference speed to roughly one-third. Such characteristics limit its suitability for deployment on low-power embedded devices commonly used in small- and medium-span bridge inspection scenarios. In contrast, the baseline UNet model maintains a certain level of segmentation accuracy but does not fully meet the simultaneous requirements of detection precision and real-time performance in mobile or edge computing environments.
By comparison, MobileUNet preserves the fundamental skip-connection architecture of UNet while significantly reducing model complexity. It achieves approximately a 67% reduction in parameter size and nearly doubles inference speed relative to the baseline model, thereby providing a more balanced trade-off between accuracy and computational efficiency. Although Transformer-based SegFormer-B0 demonstrates strong parameter efficiency and fast inference characteristics, its limited use of explicit skip connections may weaken its ability to preserve fine-grained spatial details, which are critical for pixel-level crack delineation. As a result, its mIoU performance is lower than that of MobileUNet, and it exhibits a relatively higher false-negative rate (18.2%) in crack detection tasks.
Considering the inherent characteristics of bridge cracks—such as elongated geometry, discontinuous distribution, and low-contrast appearance—accurate preservation of local features is particularly important. Under such conditions, lightweight models such as SegFormer-B0 and MobileSegNet may struggle to capture fine-grained structural details with sufficient fidelity, despite their computational advantages. In contrast, DeepLabv3+ achieves the highest segmentation accuracy among the evaluated models; however, its relatively large parameter size (34.8M) and high computational cost limit its applicability on embedded platforms used in practical bridge inspection systems.
Overall, for beam-and-slab bridge scenarios characterized by relatively simple backgrounds, a high proportion of fine cracks, and strict on-device computational constraints, lightweight UNet-based variants such as MobileUNet and UNet-Lite provide a more favorable balance between accuracy and efficiency. These models are therefore more suitable for edge deployment in practical engineering applications, where real-time performance and resource efficiency are critical considerations.

3.2. Arch Bridge

Arch bridges, characterized by curved load-bearing ribs and a distinctive compression-dominated force-transfer mechanism, are widely used in mountainous regions and along transportation corridors spanning complex terrain. In contrast to other bridge types, the arch ribs primarily sustain compressive stresses, while secondary tensile stresses arise from bending effects induced by external loading.
Cracks in arch bridges commonly occur in ribs, abutments, connection zones, and inner arch surfaces [163]. Typical crack patterns include longitudinal cracks along the arch axis, induced by temperature gradients and shrinkage, transverse or oblique cracks caused by asymmetric loading, and localized cracks near construction joints or material discontinuities. Owing to the curved structural geometry, these cracks often exhibit multi-directional orientations and continuous or semi-continuous distributions along the arch trajectory.
From a data acquisition perspective, the elevated and often inaccessible regions of arch ribs—particularly the intrados of the central arch—make manual inspection challenging. As a result, UAVs equipped with high-resolution imaging systems have become the primary platform for data collection in arch bridge inspection, enabling flexible multi-view acquisition under constrained field conditions [164,165].
However, the continuous curved surface of arch structures introduces significant variations in crack scale, orientation, and contrast across different viewing angles. This, combined with perspective distortion, non-uniform illumination, and shadow interference, significantly increases the difficulty of robust crack feature extraction [166]. Moreover, arch bridges are frequently located in complex natural environments, such as valleys and river crossings, where environmental effects, including weathering, moisture infiltration, and biological attachments, further complicate crack discrimination in the background noise. These characteristics impose stricter requirements on crack detection algorithms, particularly in terms of geometric adaptability to curved surfaces, preservation of spatial continuity, and robustness under multi-view conditions.
To address these challenges, deep learning-based methods for arch bridge crack detection have evolved from early image-level classification approaches toward more geometry-aware detection and segmentation frameworks. Multi-scale feature learning strategies, such as feature pyramid networks (FPNs), dilated (atrous) convolutions, and attention mechanisms, have been widely adopted to improve robustness against scale variation and curved crack morphology [167,168]. These methods enhance the ability of CNN-based models to capture crack structures distributed along non-linear geometric trajectories by integrating local and global contextual information.
In parallel, hybrid architectures combining CNNs and Transformer-based modules have emerged as a promising direction. By incorporating self-attention mechanisms, these models improve long-range dependency modeling and enhance feature consistency across multi-view observations, thereby addressing the fragmentation of crack representation under complex imaging conditions [169,170].
In addition, to improve detection performance under low-contrast and high-noise conditions typical of arch bridge environments, enhanced matched-filtering approaches have been proposed. These methods use orientation-adaptive Gaussian kernels to enhance crack-response signals. Experimental evaluations on public datasets such as CFD and DeepCrack show that the improved matched-filtering method achieves a pixel accuracy of 97.9%, an F1-score of 72.5%, and an IoU of 58.1%, outperforming classical edge detection operators such as Sobel and Canny in pixel-level crack enhancement tasks [171].
However, conventional image-based methods without explicit semantic modeling generally lack sufficient contextual understanding and exhibit limited robustness under complex backgrounds and varying environmental conditions, making them inadequate for large-scale arch bridge inspection tasks when used in isolation. In contrast, the geometric adaptability and spatial continuity required for curved-structure crack representation can only be effectively achieved through deep learning architectures designed for pixel-level dense prediction.
Consequently, recent research on arch bridge crack detection has primarily focused on encoder–decoder-based semantic segmentation frameworks. Among these, U-Net and its variants have become widely adopted due to their skip-connection mechanism, which simultaneously preserves global contextual information and retains fine-grained boundary details. This characteristic is particularly beneficial for accurately delineating slender and continuously distributed cracks on curved arch surfaces [133].
To further enhance detection performance in complex scenarios, multi-scale feature fusion strategies have been introduced to integrate shallow texture cues with high-level semantic representations, thereby improving sensitivity to micro-cracks, especially at joints and local discontinuities [172]. In addition, geometric perception-based approaches that combine UAV-based photogrammetry with deep learning have attracted increasing attention. By reconstructing three-dimensional structural representations of arch bridges, these methods mitigate perspective distortion and enable more consistent spatial alignment of crack features across curved surfaces. Although they offer clear advantages for high-precision inspection and long-term structural monitoring, their practical deployment is often constrained by complex data processing pipelines and high computational demands [173,174,175].
To systematically assess the applicability of different deep learning strategies for arch bridge crack detection, Table 6 summarizes representative models with respect to their architectural characteristics and structural suitability, and provides a comparative evaluation in terms of curvature adaptability, spatial continuity preservation, and engineering applicability.
Table 6. Representative Deep Learning Models for Crack Detection in Arch Bridges and Their Structural Suitability.
A comparative analysis based on Table 5 indicates that the performance of arch bridge crack detection is primarily governed by a model’s ability to accurately represent curved surface geometry and preserve the spatial continuity of crack structures. Conventional object detection methods generally exhibit limited capability for geometric representation, resulting in fragmented crack outputs that are better suited to rapid structural condition screening than to detailed quantitative assessment.
In contrast, attention-enhanced semantic segmentation frameworks based on encoder–decoder architectures demonstrate improved capability in capturing densely distributed and continuously evolving crack patterns. By incorporating global context modeling mechanisms, these methods can partially mitigate the effects of perspective distortion and complex background interference, while skip connections help preserve fine-grained spatial information along crack boundaries. Furthermore, integrating image rectification and geometry-aware preprocessing strategies can improve the spatial consistency of crack representations under multi-view observation conditions.
Overall, segmentation-based frameworks enhanced with attention mechanisms and geometric consistency constraints represent a promising direction for arch bridge crack detection, particularly for balancing structural detail preservation and environmental robustness. Nevertheless, dedicated deep learning models explicitly designed for the geometric characteristics of arch bridge structures remain relatively limited, highlighting an important research gap and a potential direction for future studies on crack detection in curved structural systems.

3.3. Cable-Stayed Bridges and Suspension Bridges

The cable is the primary load-bearing component of cable-supported bridge systems, and its polyethylene protective sheath is continuously exposed to harsh environmental conditions. Under the combined effects of wind loading, fatigue, and corrosion, cable surfaces are susceptible to damage, including cracks, scratches, rust, and voids. These defects typically exhibit continuous or semi-continuous distributions along the cable axis, with characteristic scales ranging from submillimeter to millimeter levels [176,177,178]. Such damage facilitates moisture ingress and accelerates corrosion of internal steel wires, ultimately reducing load-bearing capacity and compromising structural safety and durability [180].
From an inspection perspective, cables in cable-stayed and suspension bridges are slender, spatially distributed, and highly variable in scale, making close-range manual inspection both difficult and inefficient. These characteristics introduce unique challenges for automated crack detection systems. Consequently, UAV-based high-resolution imaging has become the dominant approach for cable inspection, enabling flexible multi-view data acquisition without disrupting traffic operations [181].
However, compared with beam and arch bridge components, image acquisition for cable-stayed bridges presents additional challenges. The cable often occupies only a small portion of the image, while the background typically contains highly heterogeneous scenes such as sky, water surfaces, and urban environments. Moreover, cable surfaces exhibit periodic textures and high reflectivity, which further increases the difficulty of distinguishing cracks from background patterns and imaging artifacts [182,183]. These factors collectively impose stringent requirements on detection sensitivity and model robustness, particularly for identifying small-scale, low-contrast, and elongated defects under complex visual conditions.
From a methodological perspective, early image-level classification approaches offered high computational efficiency but lacked precise spatial localization or continuous crack representation, thereby limiting their applicability in cable inspection tasks. Object detection methods subsequently improved practical feasibility by enabling rapid localization of defects through bounding boxes, making them suitable for preliminary screening and risk assessment applications [56,184]. However, their ability to capture slender, low-contrast crack structures remains limited.
To address these limitations, recent studies have focused on enhancing model sensitivity to small and linear targets through improved detection and segmentation strategies. Techniques such as adaptive anchor design, multi-scale feature enhancement, and lightweight network architectures have been widely adopted to improve the representation of fine-grained defects [185]. In parallel, encoder–decoder-based segmentation models integrated with attention mechanisms or Transformer modules have been proposed to strengthen cable-region feature modeling, suppress background interference, and improve the continuity of elongated crack representations [186,187].
Table 7 summarizes representative model paradigms and their performance characteristics in cable-stayed bridge crack detection tasks, providing a comparative reference for method selection under different engineering requirements.
Table 7. Performance comparison of crack detection networks for cable-stayed bridges.
The results summarized in the table indicate that visual inspection of cable surface cracks in cable-stayed bridges fundamentally constitutes a small-object detection problem under complex and highly heterogeneous background conditions. Under such conditions, a single modeling paradigm is generally insufficient to simultaneously satisfy the competing requirements of real-time performance and high-precision defect characterization.
A more effective strategy is to adopt a complementary framework that integrates lightweight object detection models with fine-grained segmentation networks. Lightweight YOLO-based models have demonstrated strong applicability in UAV-based inspection systems, enabling efficient, large-scale screening while maintaining real-time inference and reliably localizing potential defect regions.
In contrast, for detecting sub-millimeter cracks or early-stage sheath degradation, attention-based or Transformer-enhanced encoder–decoder segmentation models exhibit greater capability to suppress background interference and preserve the continuity of slender crack structures. These characteristics make them more suitable for high-precision defect delineation and long-term structural condition assessment.

3.4. Comparison of Typical Engineering Application Cases and Technical Solutions

This study presents a structured analysis of crack characteristics and deep learning-based detection methods for beam–slab, arch, and cable-supported bridges. In practical engineering applications, crack detection systems based on deep learning exhibit clear differences in data acquisition platforms, model design strategies, and deployment frameworks, resulting in significant variability in applicable scenarios, detection performance, and implementation cost. To clarify the relationship between technical solutions and bridge structural types and to support the selection of engineering-oriented methods, this section summarizes representative detection paradigms based on data acquisition strategies, compares their applicability and limitations, and highlights typical application scenarios in concrete bridge inspection.
From an engineering implementation perspective, existing deep learning-based crack detection approaches can be broadly categorized into three groups according to differences in data acquisition platforms and system architectures. The first category relies on portable devices such as smartphones or handheld terminals for image acquisition, combined with lightweight detection models (e.g., YOLO-based frameworks). These approaches are characterized by low deployment cost, operational flexibility, and real-time inference capability on edge devices. They have been validated in multiple inspection contexts; for example, Karimi et al. [189] demonstrated rapid crack identification in multi-material building surfaces using a mobile YOLO framework, while Lin et al. [190] applied deep learning-based detection models for automated crack recognition under complex material conditions. In bridge inspection applications, this category is mainly used for close-range inspection and routine verification of beam–slab bridges, with research efforts focusing on model compression and edge inference optimization to meet on-site operational requirements.
The second category adopts UAVs as the primary data acquisition platform and integrates scene-adaptive detection or segmentation models for large-scale structural inspection. Key research directions include adaptation to complex bridge geometries and improving robustness under multi-view and variable illumination conditions [88,191]. This approach is applicable to beam–slab, arch, and cable-supported bridges, particularly for components that are difficult to access manually. It represents a widely used solution for automated inspection of large-span bridge structures in practical engineering scenarios.
The third category involves multi-platform integrated geometric perception systems that combine multi-view imaging, 3D photogrammetry, and deep learning techniques. By incorporating geometric priors, these methods enable perspective correction, enhanced spatial consistency, and more accurate quantitative characterization of crack geometry [192,193]. This approach is primarily used for precision inspection and long-term structural health monitoring of critical infrastructure. However, due to the complexity of data acquisition, calibration, and processing workflows, as well as high computational and equipment costs, its application remains largely restricted to high-value or special-purpose inspection scenarios.
The comparative characteristics of these three categories in terms of application scenarios, detection accuracy, deployment cost, and computational efficiency are summarized in Table 8, providing a systematic reference for selecting appropriate technical solutions in bridge engineering applications.
Table 8. Comparison of advantages and disadvantages of three crack detection techniques.
For the concrete bridge scenarios considered in this study, all three aforementioned technical approaches have been successfully applied in practice. To clarify the correspondence between inspection strategies and bridge structural types and to support the selection of engineering-oriented methods, representative case studies are systematically organized by bridge configuration, as summarized in Table 9.
Table 9. Summary of Typical Application Cases for Deep Learning-Based Crack Depth Detection in Concrete Bridges.
The table provides a structured overview of each study, including the corresponding bridge type, key inspection components, data acquisition strategy, adopted deep learning methods, and principal application outcomes. It covers typical structural forms such as concrete beam–slab bridges, arch bridges, cable-stayed bridges, and self-anchored suspension bridges, thereby providing a comparative reference for selecting methods across different bridge inspection scenarios.
Overall, concrete beam–slab bridges currently represent the most mature application scenario, where lightweight detection and segmentation models have been widely integrated into inspection platforms such as UAV systems and mobile devices. This is primarily due to their relatively simple structural configuration and stable imaging conditions, which facilitate reliable model deployment and performance.
In contrast, applications involving arch bridges and cable-supported bridges remain comparatively limited and are typically concentrated on specific structural components, such as arch ribs, bridge piers, and cable sheaths. From a broader perspective, comprehensive, structure-aware inspection solutions for highly curved geometries and large-scale complex components are still in the early stages of development and require further engineering validation and methodological refinement.

4. Challenges and Future Research Directions

4.1. Structural Scarcity and Distribution Imbalance of Datasets

The dataset serves as a fundamental basis for the generalization capability of cross-bridge detection models. In the current state of research, limitations are not restricted to data quantity, but are more prominently reflected in structural imbalance and inconsistent data distribution. Due to non-uniform annotation standards and uneven scenario coverage, data completeness varies significantly across bridge types, thereby hindering the development of structure-adaptive crack detection methods.
At present, datasets for beam–slab bridges are more readily available; however, they exhibit strong scene homogeneity. Representative public datasets such as SDNET2018 primarily focus on bridge deck cracks, with limited variation in viewing angles, illumination conditions, and structural components. As a result, key structural elements such as web plates, flange regions, and piers are insufficiently represented, reducing model robustness in full-scale bridge inspection scenarios with diverse perspectives and component variations.
For arch bridges, dataset construction is further challenged by geometric complexity. The curved-surface structure introduces significant perspective distortion and limits feasible imaging viewpoints, resulting in a limited availability of high-quality annotated samples. In addition, pixel-level crack annotation requires consideration of spatial continuity and geometric consistency, making the labeling process substantially more labor-intensive and costly than that for planar bridge components. These factors collectively contribute to data scarcity and limit the effective training and validation of deep learning models for curved structural surfaces.
The data gap is even more pronounced for long-span cable-supported bridges, where crack data acquisition at critical regions such as towers, anchorages, and cable anchorage zones depends heavily on specialized inspection equipment, including UAVs and wall-climbing robots. Constraints related to traffic safety, operational risk, and accessibility further complicate data collection, resulting in a limited number of high-quality samples. Moreover, systematic datasets that cover the full lifecycle evolution of micro-cracks and stress-induced damage remain largely unavailable, thereby restricting the development of high-precision, long-term structural health monitoring models for cable-supported systems. In addition, inconsistencies in pixel-level annotation standards—particularly in crack boundary definition and labeling granularity—introduce further challenges in cross-dataset benchmarking and multi-source data fusion, thereby reducing the fairness and reproducibility of comparative model evaluation.
To address these limitations, future research should focus on developing hierarchical and standardized dataset systems. For beam–slab bridges, unified annotation protocols and multi-condition benchmark datasets should be established to support robust general-purpose inspection models. For arch bridge applications, physics-based rendering and simulation-driven data augmentation strategies may be explored to generate realistic crack samples under controlled geometric and optical conditions, thereby reducing reliance on costly field data acquisition. For cable-supported bridges, semi-supervised learning and few-shot learning frameworks, combined with domain adaptation techniques, are promising directions to alleviate the scarcity of annotated data in critical structural regions.

4.2. Challenges in Unified Modeling for Fine-Grained Detection Across Multiple Scenarios

Small-scale cracks, surface geometric distortions, and variations in multi-view imaging represent common visual challenges in bridge crack detection. However, these challenges differ substantially in their mechanisms of manifestation and technical priorities across bridge types. Current research has largely relied on generalized algorithmic improvements, with limited consideration of structure-specific characteristics, which constrains further improvements in detection accuracy and robustness.
For beam–slab bridges, the primary challenge lies in extracting microscale crack features within complex, heterogeneous backgrounds. Early-stage cracks often appear at the millimeter or sub-millimeter scale and are easily confounded by background interference such as stains, water marks, and construction artifacts. Although multi-scale feature fusion strategies can improve small-object recall to some extent, they generally lack dedicated mechanisms for suppressing concrete surface noise, which may lead to false positives or missed detections.
For arch bridges, crack detection is further complicated by geometric deformation induced by curved surfaces. Although cracks remain continuous along the structural surface, perspective distortion under multi-view imaging conditions disrupts their apparent spatial consistency. While attention mechanisms and Transformer-based architectures improve global context modeling, the absence of explicit geometric constraints limits their ability to preserve crack continuity across wide viewing angles, thereby affecting the accuracy of geometric parameter estimation, such as crack length and orientation.
For cable-supported bridge components, significant variations exist in both spatial scale and visual characteristics, ranging from large-scale structural elements such as bridge towers to small-scale defects in cable anchorage zones. These scale disparities, combined with varying imaging distances and viewpoints, reduce consistency in multi-view detection results. Most existing approaches rely on single-view inference without enforcing geometric consistency constraints, leading to discontinuous detection outputs and reduced reliability in quantitative crack assessment, particularly for long-term structural health monitoring applications.
It is important to note that these challenges are not independent but are inherently coupled across structural types. Future research should move beyond path-dependent improvements in generic detection models and instead develop unified frameworks that integrate structural geometry priors with deep learning. For beam–slab bridges, this involves designing small-object enhancement modules with improved background suppression capability. For arch bridges, geometry-aware segmentation models that incorporate structural constraints should be developed to preserve spatial continuity under curved-surface deformation. For cable-supported bridges, multi-view geometric consistency learning frameworks based on three-dimensional structural priors should be introduced to improve the reliability and stability of cross-view crack quantification.

4.3. Insufficient Scenario Adaptability for Engineering Deployment

The key challenge in deploying deep learning-based crack detection systems lies not only in model lightweighting but also in the substantial heterogeneity of inspection modes, deployment platforms, and performance requirements across different bridge types. Existing lightweight techniques primarily rely on generic parameter compression strategies that often fail to adapt to bridge-specific engineering contexts, resulting in noticeable gaps between laboratory performance and real-world deployment outcomes.
Although current lightweight models can achieve end-to-end real-time inference, their sensitivity to small cracks at long distances remains limited, particularly in beam–slab bridge inspections where high recall is required for comprehensive defect screening. In addition, the absence of hierarchical task scheduling mechanisms limits the establishment of multi-stage workflows, such as coarse-to-fine detection pipelines, thereby constraining the balance between computational efficiency and detection accuracy.
For arch bridge inspection scenarios, most UAV-based deployment frameworks do not adequately incorporate lightweight geometric correction modules for curved-surface imaging, leading to reduced detection accuracy under perspective distortion and increasing the complexity of subsequent geometric parameter reconstruction. In cable-supported bridge applications, where long-term fixed-point monitoring is required, current methods often lack adaptive updating mechanisms that account for environmental variations, such as seasonal illumination changes and material aging. As a result, model performance may gradually degrade over time, and limited system-level integration further restricts compatibility with existing structural health monitoring (SHM) platforms.
To address these multi-bridge deployment challenges, Mishra et al. [199] proposed an integrated IoT–UAV bridge health monitoring framework. The system introduces hierarchical computational scheduling to separate preliminary detection from fine-grained segmentation tasks, incorporates geometric correction modules tailored to curved bridge surfaces, and employs continuous learning strategies to mitigate performance degradation under long-term environmental variation. In addition, standardized data interfaces are designed to support interoperability with existing SHM systems, providing a reference architecture for system-level integration in bridge inspection applications.
Building on these developments, future research should focus on tighter integration between detection algorithms and engineering deployment requirements. This includes developing lightweight pre-inspection frameworks and high-precision post-analysis models tailored to different bridge inspection stages. Trajectory-aware scheduling strategies may further enable dynamic allocation of computational resources during UAV-based inspections. For arch bridges, geometry-aware lightweight models with built-in distortion correction mechanisms are required to support end-to-end crack detection under curved-surface conditions. For cable-supported bridges, continuous learning and domain adaptation techniques should be explored to improve long-term robustness under varying environmental conditions, along with unified interface standards to enhance interoperability between detection systems and operational platforms.
In addition, in line with the development of intelligent maintenance systems, emerging directions such as digital twin- and VR-based inspection frameworks may be considered to support immersive visualization of bridge conditions and the evolution of defects. These approaches can enhance the mapping between structural states, defect information, and virtual representations, thereby improving remote inspection and decision-making capabilities. Furthermore, optimizing human–machine interaction on mobile and edge devices—such as simplifying model deployment, parameter configuration, and data management workflows—can further reduce technical barriers and improve the practical usability of deep learning-based bridge inspection systems.

5. Conclusions

Existing literature reviews on deep learning-based bridge crack detection predominantly adopt either network-evolution-driven or task-categorization-based frameworks. However, these paradigms generally do not explicitly reveal the intrinsic coupling relationship between bridge structural characteristics and the selection of appropriate detection algorithms.
In contrast, this study takes the interaction between bridge structural features and crack morphology as the central organizing principle. It systematically reviews recent advances in deep learning-based crack detection for concrete bridges and establishes explicit mapping relationships among bridge types, crack characteristics, and suitable detection methodologies. Based on a comparative analysis of three representative technical paradigms—crack classification, object detection, and semantic segmentation—and their applicability across different bridge types, the following key findings are derived.
First, high-quality bridge-specific datasets constitute the foundation for objective model evaluation and methodological advancement in this field. At present, the development of deep learning algorithms has progressed more rapidly than the construction of specialized bridge inspection datasets. Existing public datasets are often characterized by structural bias (e.g., overrepresentation of slab-type bridges), limited diversity in imaging conditions, and inconsistent annotation standards. These issues collectively restrict model generalization across heterogeneous bridge structures and operational environments. Therefore, the development of standardized, multi-source datasets that cover diverse bridge types, environmental conditions, and acquisition platforms is essential to enable fair and reproducible benchmarking and to support the translation of research outcomes into engineering practice.
Second, detection challenges vary significantly across bridge types, necessitating structure-aware model design strategies. For slab bridges, crack morphology is relatively regular and imaging conditions are more stable; in such cases, lightweight encoder–decoder segmentation networks can effectively balance accuracy and computational efficiency. For arch bridges, curved geometric surfaces introduce perspective distortion and discontinuities in crack representation, requiring geometry-aware preprocessing and attention-enhanced models to improve robustness in preserving crack continuity. For cable-supported bridges, where cracks are typically slender, small-scale, and densely distributed, small-object detection frameworks combined with geometric constraints are required to improve localization accuracy and reliability.
Third, current research has primarily focused on performance improvements using benchmark datasets, whereas additional factors, including limited communication bandwidth in UAV-based inspection, computational restrictions on edge devices, and environmental variability in outdoor imaging conditions constrain practical engineering applications. Moreover, different deployment strategies exhibit distinct applicability domains: UAV-based systems are suitable for large-scale surveys, fixed-point monitoring systems are more effective for long-term condition tracking, and mobile terminal solutions are better suited for rapid on-site inspection, each requiring tailored system configurations.
Overall, deep learning-based crack detection for concrete bridges has progressed from methodological validation to engineering-oriented deployment and is now transitioning from laboratory research to practical application.
Future research should move beyond incremental improvements of individual network architectures and instead focus on developing an integrated framework that combines standardized dataset construction, structure-adaptive model design, and lightweight edge deployment strategies. Such a framework would enable a unified pipeline spanning data acquisition, defect detection, and structural condition assessment, thereby providing a systematic technical foundation for intelligent bridge operation and maintenance.

Author Contributions

Conceptualization, J.L., M.M.Y., and B.T., methodology, J.L.; software, J.L. and M.M.Y.; validation, B.T. and Z.Z.; formal analysis, J.L.; investigation, J.L. and M.M.Y.; resources, B.T. and Z.Z.; data curation, J.L.; writing—original draft preparation, J.L.; writing—review and editing, J.L.; visualization, J.L.; supervision, M.M.Y.; funding acquisition, M.M.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Du, H.; Wang, H.F.; Zhang, X.W.; Peng, H.A.; Gao, R.; Zheng, X.Y.; Tong, Y.X.; Shan, Y.H.; Pan, Z.F.; Huang, H.; et al. Automated intelligent measurement of cracks on bridge piers using a ring-climbing vision scanning operation robot. Measurement 2024, 237, 115197. [Google Scholar] [CrossRef] [Scilit]
  2. Liu, Z.L.; Zhou, A.; Ran, X.R.; Wu, Y.P.; Zhao, W.G.; Zhang, H. A crack detection and quantification method using matched filter and photograph reconstruction. Sci. Rep. 2025, 15, 25266. [Google Scholar] [CrossRef] [Scilit]
  3. Alawieh, A.F.; Daoud, H.; Shaffer, M.; Issa, M.A. Implementing Effective Sealing Techniques with Recommended Products to Mitigate Bridge Deck Cracking Effects. J. Bridge Eng. 2026, 31, 04026005. [Google Scholar] [CrossRef] [Scilit]
  4. Zhang, A.A.; Wang, D.; Peng, Y.; Cheng, H.X.; Wei, Y.F.; Zhan, Y. BSCS-Net: A Lightweight Segmentation Network for Automated Bridge Surface Crack Detection. J. Comput. Civ. Eng. 2026, 40, 04026007. [Google Scholar] [CrossRef] [Scilit]
  5. Xu, G.; Yue, Q.; Liu, X.; Chen, H. Investigation on the effect of data quality and quantity of concrete cracks on the performance of deep learning-based image segmentation. Expert Syst. Appl. 2024, 237, 121686. [Google Scholar] [CrossRef] [Scilit]
  6. Ni, Y.; Mao, J.; Fu, Y.; Wang, H.; Zong, H.; Luo, K. Damage Detection and Localization of Bridge Deck Pavement Based on Deep Learning. Sensors 2023, 23, 5138. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Xiao, J.W.; Fan, J.S.; Liu, Y.F.; Li, B.; Nie, J.G. Region of interest (ROI) extraction and crack detection for UAV-based bridge inspection using point cloud segmentation and 3D-to-2D projection. Autom. Constr. 2024, 158, 653–674. [Google Scholar] [CrossRef] [Scilit]
  8. Kim, B.; Yuvaraj, N.; Preethaa, K.R.S.; Pandian, R.A. Surface crack detection using deep learning with shallow CNN architecture for enhanced computation. Neural Comput. Appl. 2021, 33, 9289–9305. [Google Scholar] [CrossRef] [Scilit]
  9. Nguyen, S.D.; Tran, T.S.; Tran, V.P.; Lee, H.J.; Piran, M.J.; Le, V.P. Deep Learning-Based Crack Detection: A Survey. Int. J. Pavement Res. Technol. 2022, 16, 943–967. [Google Scholar] [CrossRef] [Scilit]
  10. Abdel-Qader, I.; Abudayyeh, O.; Kelly, M.E. Analysis of edge-detection techniques for crack identification in bridges. J. Comput. Civ. Eng. 2003, 17, 255–263. [Google Scholar] [CrossRef] [Scilit]
  11. Azimi, M.; Eslamlou, A.D.; Pekcan, G. Data-Driven Structural Health Monitoring and Damage Detection through Deep Learning: State-of-the-Art Review. Sensors 2020, 20, 2778. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Seo, J.; Duque, L.; Wacker, J. Drone-enabled bridge inspection methodology and application. Autom. Constr. 2018, 94, 112–126. [Google Scholar] [CrossRef] [Scilit]
  13. Ayele, Y.Z.; Aliyari, M.; Griffiths, D.; Droguett, E.L. Automatic Crack Segmentation for UAV-Assisted Bridge Inspection. Energies 2020, 13, 6250. [Google Scholar] [CrossRef] [Scilit]
  14. Gupta, P.; Dixit, M. Image-based crack detection approaches: A comprehensive survey. Multimed. Tools Appl. 2022, 81, 40181–40229. [Google Scholar] [CrossRef] [Scilit]
  15. Park, S.E.; Eem, S.-H.; Jeon, H. Concrete Crack Detection and Quantification Using Deep Learning and Structured Light. Constr. Build. Mater. 2020, 252, 119096. [Google Scholar] [CrossRef] [Scilit]
  16. Lee, J.; Kim, J.; Ahn, E.; Shin, M.; Cho, S. Learning to Detect Cracks on Damaged Concrete Surfaces Using Two-Branched Convolutional Neural Network. Sensors 2019, 19, 4796. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Roy, S.; Yogi, B.; Majumdar, R.; Ghosh, P.; Das, S.K. Deep learning-based crack detection and prediction for structural health monitoring. Discov. Appl. Sci. 2025, 7, 674. [Google Scholar] [CrossRef] [Scilit]
  18. Moreh, F.; Bakkali, M.E.; Laalej, H.; Bal, I.E. Deep neural networks for crack detection inside structures. Sci. Rep. 2024, 14, 4439. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Hamishebahar, Y.; Kovačević, M.; Milovanović, B.; Ghasemi, M. A Comprehensive Review of Deep Learning-Based Crack Detection Approaches. Appl. Sci. 2022, 12, 1374. [Google Scholar] [CrossRef] [Scilit]
  20. Alzubaidi, L.; Zhang, J.; Humaidi, A.J.; Al-Dujaili, A.; Duan, Y.; Al-Shamma, O.; Santamaría, J.; Fadhel, M.A.; Al-Amidie, M.; Farhan, L. Review of deep learning: Concepts, CNN architectures, challenges, applications, future directions. J. Big Data 2021, 8, 53. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Zhang, J.; Chen, W.; Zhang, J. UAV-based quantitative crack measurement for bridges integrating four-point laser metric calibration and mamba segmentation. Autom. Constr. 2026, 182, 106774. [Google Scholar] [CrossRef] [Scilit]
  22. Yiğit, A.Y.; Uysal, M. Virtual reality visualisation of automatic crack detection for bridge inspection from 3D digital twin generated by UAV photogrammetry. Measurement 2025, 242, 115931. [Google Scholar] [CrossRef] [Scilit]
  23. Li, S.; Zhao, X.; Zhou, G. Automatic pixel-level multiple damage detection of concrete structure using fully convolutional network. Comput.-Aided Civ. Infrastruct. Eng. 2019, 34, 616–634. [Google Scholar] [CrossRef] [Scilit]
  24. Kim, B.; Natarajan, Y.; Preethaa, K.S.; Song, S.; An, J.; Mohan, S. Real-time assessment of surface cracks in concrete structures using integrated deep neural networks with autonomous unmanned aerial vehicle. Eng. Appl. Artif. Intell. 2024, 129, 107537. [Google Scholar] [CrossRef] [Scilit]
  25. Bai, Y.; Quan, W.; Shi, X.; Yan, Z.; Yuan, G. A Review of UAV-Based Crack Detection in Civil Infrastructure: A Multi-Level Visual Analysis Framework, Scene Adaptability, and Challenges. Remote Sens. 2026, 18, 1806. [Google Scholar] [CrossRef] [Scilit]
  26. Prakash, V.; Debono, C.J.; Musarat, M.A.; Borg, R.P.; Seychell, D.; Ding, W.; Shu, J. Structural Health Monitoring of Concrete Bridges Through Artificial Intelligence: A Narrative Review. Appl. Sci. 2025, 15, 4855. [Google Scholar] [CrossRef] [Scilit]
  27. Pandey, V.; Mishra, S.S. A review of image-based deep learning methods for crack detection. Multimed. Tools Appl. 2025, 84, 35469–35511. [Google Scholar] [CrossRef] [Scilit]
  28. Pan, X.; Zhang, Y.; Li, J.; Wang, H.; Liu, Y. A review of recent advances in data-driven computer vision methods for structural damage evaluation: Algorithms, applications, challenges, and future opportunities. Arch. Comput. Methods Eng. 2025, 32, 4587–4619. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Yuan, Q.; Shi, Y.; Li, M. A Review of Computer Vision-Based Crack Detection Methods in Civil Infrastructure: Progress and Challenges. Remote Sens. 2024, 16, 2910. [Google Scholar] [CrossRef] [Scilit]
  30. Fan, C.; Ding, Y.; Liu, X.; Yang, K. A review of crack research in concrete structures based on data-driven and intelligent algorithms. Structures 2025, 75, 108800. [Google Scholar] [CrossRef] [Scilit]
  31. McCulloch, W.S.; Pitts, W. A logical calculus of the ideas immanent in nervous activity. Bull. Math. Biol. 1943, 5, 115–133. [Google Scholar] [CrossRef] [Scilit]
  32. Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning representations by back-propagating errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef] [Scilit]
  33. Hinton, G.E.; Salakhutdinov, R.R. Reducing the dimensionality of data with neural networks. Science 2006, 313, 504–507. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Hinton, G.; Deng, L.; Yu, D.; Dahl, G.E.; Mohamed, A.-R.; Jaitly, N.; Senior, A.; Vanhoucke, V.; Nguyen, P.; Sainath, T.N.; et al. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Process. Mag. 2012, 29, 82–97. [Google Scholar] [CrossRef] [Scilit]
  35. Qiu, G.Y.; Wang, H.; Li, Z.; Zhang, Y. Low-illumination and noisy bridge crack image restoration by deep CNN denoiser and normalized flow module. Sci. Rep. 2024, 14, 18270. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Koscák, J.; Jaksa, R.; Sincák, P. Influence of Number of Neurons in Time Delay Recurrent Networks with Stochastic Weight Update on Backpropagation Through Time. In Nostradamus: Modern Methods of Prediction, Modeling and Analysis of Nonlinear Systems; Springer: Berlin/Heidelberg, Germany, 2013; pp. 133–141. [Google Scholar]
  37. Bareither, C.A.; Kwak, S. Assessment of municipal solid waste settlement models based on field-scale data analysis. Waste Manag. 2015, 42, 101–117. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Xu, W.Y.; Li, H.; Li, G.D.; Ji, Y.C.; Xu, J.B.; Zang, Z. Improved YOLOv8n-based bridge crack detection algorithm under complex background conditions. Sci. Rep. 2025, 15, 13074. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Kilari, S.; Wu, P. Deep residual learning for image recognition. Int. Res. J. Eng. Technol. 2022, 6, 780–783. [Google Scholar]
  40. Liu, H.; Zhang, Y.F. Bridge condition rating data modeling using deep learning algorithm. Struct. Infrastruct. Eng. 2020, 16, 1447–1460. [Google Scholar] [CrossRef] [Scilit]
  41. Zhang, Y.; Hou, S.; Shen, H.; Li, Y.; Zhao, W. Automatic Multiple Defect Detection Method for Underwater Concrete Bridge Piers based on Deep Learning. Struct. Eng. Int. 2026, 1–13. [Google Scholar] [CrossRef] [Scilit]
  42. Rosenblatt, F. The Perceptron, a Perceiving and Recognizing Automaton Project Para, Report: Cornell Aeronautical Laboratory; Cornell Aeronautical Laboratory: Buffalo, NY, USA, 1957. [Google Scholar]
  43. Elman, J.L. Finding structure in time. Cogn. Sci. 1990, 14, 179–211. [Google Scholar] [CrossRef]
  44. Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef] [Scilit]
  46. Hinton, G.E.; Osindero, S.; Teh, Y.-W. A fast learning algorithm for deep belief nets. Neural Comput. 2006, 18, 1527–1554. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. 2014Creswell, A.; White, T.; Dumoulin, V.; Arulkumaran, K.; Sengupta, B.; Bharath, A.A. Generative adversarial networks: An overview. IEEE Signal Process. Mag. 2018, 35, 53–65. [Google Scholar] [CrossRef] [Scilit]
  48. Rhazaoui, K.; Keddam, M.; Chevalier, S.; Vivet, L. Towards the 3D Modelling of the Effective Conductivity of Solid Oxide Fuel Cell Electrodes—Validation against experimental measurements and prediction of electrochemical performance. Electrochim. Acta 2015, 168, 139–147. [Google Scholar] [CrossRef] [Scilit]
  49. Silver, D.; Huang, A.; Maddison, C.J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. Mastering the game of Go with deep neural networks and tree search. Nature 2016, 529, 484–489. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N. Attention Is All You Need. arXiv 2017, arXiv:1706.03762. [Google Scholar]
  51. Wang, L.; Zhang, Z. Automatic detection of wind turbine blade surface cracks based on UAV-taken images. IEEE Trans. Ind. Electron. 2018, 64, 7293–7303. [Google Scholar]
  52. Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; Volume 1, pp. 4171–4186. [Google Scholar]
  53. Gong, X.Q.; Wang, Y.; Li, S.; Zhang, H. Analysis and Test of Concrete Surface Crack of Railway Bridge Based On Deep Learning. In Proceedings of 2020 IEEE 5th Information Technology and Mechatronics Engineering Conference (ITOEC 2020); IEEE: New York, NY, USA, 2020; pp. 437–442. [Google Scholar]
  54. Feng, C.; Zhang, H.; Wang, H.; Wang, S.; Li, Y. Automatic pixel-level crack detection on dam surface using deep convolutional network. Sensors 2020, 20, 2069. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Liu, T.; Wang, H.; Li, Z.; Zhang, Y.; Chen, X. BC-DUnet-based segmentation of fine cracks in bridges under a complex background. PLoS ONE 2022, 17, e0265258. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Li, R.; Liu, X.; Wang, J.; Zhang, L. Automatic bridge crack detection using Unmanned aerial vehicle and Faster R-CNN. Constr. Build. Mater. 2023, 362, 129659. [Google Scholar] [CrossRef] [Scilit]
  57. Huang, L.J.; Zhang, Y.; Wang, L.; Chen, H. Deep learning for automated multiclass surface damage detection in bridge inspections. Autom. Constr. 2024, 166, 105601. [Google Scholar] [CrossRef] [Scilit]
  58. Murtaza, S.; Imran, M.; Rana, M.R.R. CrackNet-VGG: A Deep Learning Framework for Automated Detection of Surface Cracks in Concrete Structures. Appl. Comput. Syst. 2025, 30, 187–194. [Google Scholar] [CrossRef] [Scilit]
  59. Han, B.L.; Wang, Z.; Li, Y.; Zhao, J. UMDA: Lightweight and Efficient Crack Segmentation Model. J. Comput. Civ. Eng. 2026, 40, 04025143. [Google Scholar] [CrossRef] [Scilit]
  60. Zhuang, H.Y.; Li, Q.; Zhang, X.; Wang, H.; Chen, Y. Deep learning for surface crack detection in civil engineering: A comprehensive review. Measurement 2025, 248, 116908. [Google Scholar] [CrossRef] [Scilit]
  61. Han, X.F. A Novel Search Strategy-Based Deep Learning for City Bridge Cracks Detection in Urban Planning. Autom. Control Comput. Sci. 2022, 56, 428–437. [Google Scholar] [CrossRef] [Scilit]
  62. Shen, Y.G.; Yu, Z.W.; Wen, Z.L. A crack detection method based on deep transfer learning. In Bridge Maintenance, Safety, Management, Life-Cycle Sustainability and Innovations; CRC Press: Boca Raton, FL, USA, 2021; pp. 271–278. [Google Scholar]
  63. Afifah, V.; Erniwati, S. Yolov8 for object detection: A comprehensive review of advances, techniques, and applications. Int. J. Adv. Comput. Sci. Inf. 2026, 2, 53–61. [Google Scholar]
  64. Lavadiya, D.N.; Dorafshan, S. Deep learning models for analysis of non-destructive evaluation data to evaluate reinforced concrete bridge decks: A survey. Eng. Rep. 2025, 7, e12608. [Google Scholar]
  65. Wang, B.; Li, Y.; Zhang, X.; Liu, Z. Graph convolutional networks fusing motif-structure information. Sci. Rep. 2022, 12, 10735. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  66. Amirkhani, D.; Allili, M.S.; Hebbache, L.; Hammouche, N.; Lapointe, J.F. Visual Concrete Bridge Defect Classification and Detection Using Deep Learning: A Systematic Review. IEEE Trans. Intell. Transp. Syst. 2024, 25, 10483–10505. [Google Scholar] [CrossRef] [Scilit]
  67. Schmugge, S.J.; Rice, L.; Nguyen, N.R.; Lindberg, J.; Grizzi, R.; Joffe, C. Detection of Cracks in Nuclear Power Plant using Spatial-temporal Grouping of Local Patches. In Proceedings of the 2016 IEEE Winter Conference on Applications of Computer Vision (WACV 2016), Lake Placid, New York, USA, 7–10 March 2016. [Google Scholar]
  68. Cha, Y.J.; Choi, W.; Büyüköztürk, O. Deep Learning-Based Crack Damage Detection Using Convolutional Neural Networks. Comput.-Aided Civ. Infrastruct. Eng. 2017, 32, 361–378. [Google Scholar] [CrossRef] [Scilit]
  69. Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv 2014, arXiv:1409.1556. [Google Scholar]
  70. Wang, J.F.; Zeng, W.H.; Zhong, H. A mechanism-based hybrid Transformer-GRU network for bridge pier hysteresis curves prediction: An interpretable research. Sci. Rep. 2026, 16, 4961. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Bhattacharya, G.; Mandal, B.; Puhan, N.B. Multi-Deformation Aware Attention Learning for Concrete Structural Defect Classification. IEEE Trans. Circuits Syst. Video Technol. 2021, 31, 3707–3713. [Google Scholar] [CrossRef] [Scilit]
  72. Yang, Q.; Shi, W.; Zhang, J.; Chen, Z. Deep convolution neural network-based transfer learning method for civil infrastructure crack detection. Autom. Constr. 2020, 116, 103199. [Google Scholar] [CrossRef] [Scilit]
  73. Xu, Y.; Bao, Y.; Chen, J.; Zuo, W.; Li, H. Surface fatigue crack identification in steel box girder of bridges by a deep fusion convolutional neural network based on consumer-grade camera images. Struct. Health Monit. 2019, 18, 653–674. [Google Scholar]
  74. Trach, R. A model classifying four classes of defects in reinforced concrete bridge elements using convolutional neural networks. Innov. Infrastruct. Solut. 2023, 8, 123. [Google Scholar] [CrossRef] [Scilit]
  75. Yu, Z.; Cai, R.; Cui, Y.; Liu, X.; Hu, Y.; Kot, A.C. Rethinking vision transformer and masked autoencoder in multimodal face anti-spoofing. Int. J. Comput. Vis. 2024, 132, 5217–5238. [Google Scholar] [CrossRef] [Scilit]
  76. Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2014. [Google Scholar]
  77. Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016. [Google Scholar]
  78. Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single shot multibox detector. In European Conference on Computer Vision; Springer: Cham, Switzerland, 2016. [Google Scholar]
  79. Law, H.; Deng, J. Cornernet: Detecting objects as paired keypoints. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Cham, Switzerland, 2018. [Google Scholar]
  80. Shelhamer, E.; Long, J.; Darrell, T. Fully convolutional networks for semantic segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 640–651. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  81. Girshick, R. Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2015. [Google Scholar]
  82. Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 1137–1149. [Google Scholar] [PubMed]
  83. Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 7464–7475. [Google Scholar]
  84. Sohaib, M.; Arif, M.; Kim, J.-M. Evaluating YOLO models for efficient crack detection in concrete structures using transfer learning. Buildings 2024, 14, 3928. [Google Scholar] [CrossRef] [Scilit]
  85. Yang, H.; Yang, L.; Wu, T.; Meng, Z.Q.; Huang, Y.J.; Wang, P.S.P.; Li, P.; Li, X. Automatic detection of bridge surface crack using improved YOLOv5s. Math. Biosci. Eng. 2022, 36, 2250047. [Google Scholar] [CrossRef] [Scilit]
  86. Pandey, V.; Mishra, S.S. Lightweight frameworks for real-time crack monitoring in civil infrastructure. Ultrasonics 2026, 162, 107970. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  87. Zhou, Q.; Dong, S.; Qin, Y.; Xia, Z.; Yang, J. Bridge apparent disease detection based on improved YOLOv3. J. Chongqing Univ. 2022, 45, 121–130. [Google Scholar]
  88. Yanna, L.; Liang, Y. Bridge disease detection and recognition based on improved YOLOX algorithm. J. Adv. Opt. 2023, 44, 792–800. [Google Scholar] [CrossRef] [Scilit]
  89. Wang, Y.; He, J. A rapid concrete crack detection method based on improved YOLOv8. IEEE Access 2025, 13, 59227–59243. [Google Scholar] [CrossRef] [Scilit]
  90. Inam, H.; Ali, M.; Ahmed, S.; Khan, M.A.; Ur Rehman, A. Smart and Automated Infrastructure Management: A Deep Learning Approach for Crack Detection in Bridge Images. Sustainability 2023, 15, 1866. [Google Scholar] [CrossRef] [Scilit]
  91. Hsieh, H.Y.; Liu, K.Y.; Kang, S.M. Development of an automated surface crack detection and BIM-integrated management system for concrete bridges. J. Civ. Eng. Manag. 2025, 31, 710–728. [Google Scholar] [CrossRef] [Scilit]
  92. Xia, H.; Li, Q.; Qin, X.; Zhuang, W.; Ming, H.; Yang, X. LiuBridge crack detection algorithm designed based on YOLOv8. Appl. Soft Comput. 2025, 171, 112831. [Google Scholar] [CrossRef] [Scilit]
  93. Hakimi, O.; Liu, H.X.; Abudayyeh, O. Deep learning-driven multi-level data fusion framework for predictive maintenance and structural health monitoring of concrete bridge decks. Autom. Constr. 2025, 175, 106180. [Google Scholar] [CrossRef] [Scilit]
  94. Taffese, W.Z.; Sharma, R.; Afsharmovahed, M.H.; Manogaran, G.; Chen, G. Benchmarking YOLOv8 for optimal crack detection in civil infrastructure. arXiv 2025, arXiv:2501.06922, 2025. [Google Scholar]
  95. Huang, J.; Wang, Y.; Li, Z.; Zhang, Q.; Zhao, M. Efficient bridge damage detection using a lightweight attention-based modeling framework. Comput.-Aided Civ. Infrastruct. Eng. 2025, 40, 4758–4773. [Google Scholar] [CrossRef] [Scilit]
  96. Dong, X.; Yan, S.; Duan, C. A lightweight vehicles detection network model based on YOLOv5. Eng. Appl. Artif. Intell. 2022, 113, 104914. [Google Scholar] [CrossRef] [Scilit]
  97. Ha, M.H.; Chen, O.T.C. Deep neural networks using capsule networks and skeleton-based attentions for action recognition. IEEE Access 2021, 9, 6164–6178. [Google Scholar] [CrossRef] [Scilit]
  98. Tran, T.S.; Nguyen, S.D.; Lee, H.J.; Tran, V.P. Advanced crack detection and segmentation on bridge decks using deep learning. Constr. Build. Materials 2023, 400, 132839. [Google Scholar] [CrossRef] [Scilit]
  99. Zhao, W.F.; Rong, S.H.; Yu, Z.; He, B. A lightweight and robust vision framework for underwater pipeline defect detection. Knowl.-Based Syst. 2026, 343, 115984. [Google Scholar] [CrossRef] [Scilit]
  100. Zhang, H.Y.; Tian, C.; Zhang, A.; Liu, Y.L.; Gao, G.x.; Zhuang, Z.W.; Yin, T.T.; Zhang, N. A Bridge Defect Detection Algorithm Based on UGMB Multi-Scale Feature Extraction and Fusion. Symmetry 2025, 17, 1025. [Google Scholar] [CrossRef] [Scilit]
  101. Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2015. [Google Scholar]
  102. Yang, X.; Li, H.; Yu, Y.; Huang, T. Automatic pixel-level crack detection and measurement using fully convolutional network. Comput.-Aided Civ. Infrastruct. Eng. 2018, 33, 1090–1109. [Google Scholar] [CrossRef] [Scilit]
  103. Liang, D.; Zhou, X.F.; Wang, S.; Liu, C.J. Research on concrete cracks recognition based on dual convolutional neural network. IEEE Access 2019, 23, 3066–3074. [Google Scholar] [CrossRef] [Scilit]
  104. Huang, H.; Lin, L.; Tong, R.; Hu, H.; Zhang, Q.; Iwamoto, Y.; Han, X.; Chen, Y.; Wu, J. Unet 3+: A full-scale connected UNet for medical image segmentation. arXiv 2020, arXiv:2004.08790. [Google Scholar]
  105. Tsai, T.-H.; Tseng, Y.-W. BiSeNet V3: Bilateral segmentation network with coordinate attention for real-time semantic segmentation. Neurocomputing 2023, 532, 33–42. [Google Scholar] [CrossRef] [Scilit]
  106. Xu, J.; Xiong, Z.; Bhattacharyya, S.P. PIDNet: A real-time semantic segmentation network inspired by PID controllers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023. [Google Scholar]
  107. Yu, F.; Koltun, V. Multi-scale context aggregation by dilated convolutions. arXiv 2016, arXiv:1511.07122. [Google Scholar]
  108. Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 834–848. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  109. Cao, Y.; Xu, J.; Lin, S.; Wei, F.; Hu, H. Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops; IEEE: New York, NY, USA, 2019. [Google Scholar]
  110. Zhang, H.; Dana, K.; Shi, J.; Zhang, Z.; Wang, X.; Tyagi, A.; Agrawal, A. Context encoding for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018. [Google Scholar]
  111. Badrinarayanan, V.; Kendall, A.; Cipolla, R. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 2481–2495. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  112. Noh, H.; Hong, S.; Han, B. Learning deconvolution network for semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2015. [Google Scholar]
  113. Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
  114. Zhang, W.Q.; Huang, Z.L.; Luo, G.Z.; Chen, T.; Wang, X.G.; Liu, W.Y.; Yu, G.; Shen, C.H. Topformer: Token pyramid transformer for mobile semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022. [Google Scholar]
  115. Zhou, Q.; Ding, S.; Qing, G.; Hu, J. UAV vision detection method for crane surface cracks based on Faster R-CNN and image segmentation. J. Civ. Struct. Health Monit. 2022, 12, 845–855. [Google Scholar]
  116. Liu, H.J.; Miao, X.Y.; Mertz, C.; Xu, C.Z.; Kong, H. Crackformer: Transformer network for fine-grained crack detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021. [Google Scholar]
  117. Zhang, H.Y.; Chen, N.; Li, M.; Mao, S.J. The crack diffusion model: An innovative diffusion-based method for pavement crack detection. Appl. Sci. 2024, 16, 986. [Google Scholar] [CrossRef] [Scilit]
  118. Wang, B.; Li, J.; Zhang, W.; Zhou, M. MPSU-Net: Quantitative interpretation algorithm for road cracks based on multiscale feature fusion and superimposed U-Net. Autom. Constr. 2024, 153, 104598. [Google Scholar] [CrossRef] [Scilit]
  119. Shan, J.; Huang, Y.; Jiang, W. DCUFormer: Enhancing pavement crack segmentation in complex environments with dual-cross/upsampling attention. Expert Syst. Appl. 2025, 264, 125891. [Google Scholar] [CrossRef] [Scilit]
  120. Wen, L.; Ye, Y.; Zuo, L. GAF-Net: A new automated segmentation method based on multiscale feature fusion and feedback module. Pattern Recognit. Lett. 2025, 187, 86–92. [Google Scholar] [CrossRef] [Scilit]
  121. Li, S.; Gou, W.; Tan, Z.; Hedayat, M.; Sun, W.; Guo, P. Ductility and cracking behavior of UHPC-RC composite beam under bending test based on acoustic emission parameters. Structures 2024, 70, 107624. [Google Scholar] [CrossRef] [Scilit]
  122. Li, Z.; Zhu, H.; Huang, M. A deep learning-based fine crack segmentation network on full-scale steel bridge images with complicated backgrounds. IEEE Access 2021, 9, 114989–114997. [Google Scholar] [CrossRef] [Scilit]
  123. Li, J.F.; Wang, B.; An, P.Y.; Zhao, J.X.; Luo, Y. Research on automatic detection and analysis of concrete cracks based on deep convolutional neural network architecture. Proc. Inst. Mech. Eng. L 2025, 29, 1188–1201. [Google Scholar] [CrossRef] [Scilit]
  124. Li, G.; Fang, Z.Y.; Mohammed, A.M.; Liu, T.; Deng, Z.H. Automated bridge crack detection based on improving encoder–decoder network and strip pooling. J. Struct. Eng. 2023, 29, 04023004. [Google Scholar] [CrossRef] [Scilit]
  125. Fan, Z.; Demirci, C.; Lin, Y.X.; Wang, H.R.; Avci, E. Automatic crack detection on road pavements using encoder-decoder architecture. Appl. Sci. 2020, 13, 2960. [Google Scholar] [CrossRef] [Scilit]
  126. Zhu, S.Y.; Du, J.C.; Li, Y.S. Method for bridge crack detection based on the U-Net convolutional networks. J. Bridge Eng. 2019, 46, 35–42. [Google Scholar]
  127. Sim, J.; Kafatos, M.; Kim, S.H.; Lee, Y.W. StyleSPADE: Realistic Image Augmentation for Robust Infrastructure Crack Segmentation via Ensemble Learning. Appl. Sci. 2026, 16, 837. [Google Scholar] [CrossRef] [Scilit]
  128. Li, G.; Gao, Z.Y.; Zhang, X.C.; Zhao, H.X.; Liu, Z. Improved global convolutional network for pavement crack detection. IEEE Access 2020, 57, 081011. [Google Scholar] [CrossRef] [Scilit]
  129. Yu, J.Y.; Li, F.; Xue, X.K.; Yin, D.; Liu, B.L. Intelligent Identification and Measurement of Bridge Cracks Based On YOLOv5 and U-Net3+. J. Hunan Univ. 2023, 50, 65–73. [Google Scholar]
  130. Gao, T.; Yuanzhou, Z.; Ji, B.; Xia, J. Multitask fatigue crack recognition network based on task similarity analysis. Int. J. Fatigue 2023, 176, 107864. [Google Scholar] [CrossRef] [Scilit]
  131. Xiong, B.; Hong, R.; Wang, J.X.; Li, W.; Zhang, J.; Ge, D.D. DefNet: A multi-scale dual-encoding fusion network aggregating Transformer and CNN for crack segmentation. Anal. Chim. Acta 2024, 448, 138206. [Google Scholar] [CrossRef] [Scilit]
  132. Zhang, C.; Yu, J.; Wu, G.Y. Cracks segmentation of engineering structures in complex backgrounds using a concatenation of Transformer and CNN models driven by scene understanding information. Structures 2024, 65, 106685. [Google Scholar] [CrossRef] [Scilit]
  133. Li, Z.W.; Yang, Y.Q.; Xie, M.Z.; Li, Y.H. Automated Concrete Crack Detection System with Superresolution Enhancement and Hybrid Measurement Techniques. J. Comput. Civ. Eng. 2026, 40, 04025159. [Google Scholar] [CrossRef] [Scilit]
  134. He, A.Z.; Dong, Z.S.; Zhang, H.; Zhang, A.A.; Qiu, S.; Liu, Y.; Wang, K.C.P.; Lin, Z.H. Automated Pixel-Level Detection of Expansion Joints on Asphalt Pavement Using a Deep-Learning-Based Approach. Struct. Control Health Monit. 2023, 30, 7552337. [Google Scholar]
  135. ang, L.; Huang, H.Y.; Kong, S.Y.; Liu, Y.H.; Yu, H.N. PAF-Net: A progressive and adaptive fusion network for pavement crack segmentation. IEEE Trans. Intell. Transp. Syst. 2023, 24, 12686–12700. [Google Scholar] [CrossRef] [Scilit]
  136. Zhou, Q.; Qu, Z.; Li, Y.X.; Ju, F.R. Tunnel crack detection with linear seam based on mixed attention and multiscale feature fusion. Meas. Sci. Technol. 2022, 71, 1–11. [Google Scholar] [CrossRef] [Scilit]
  137. Dorafshan, S.; Thomas, R.J.; Maguire, M.J. SDNET2018: An annotated image dataset for non-contact concrete crack detection using deep convolutional neural networks. Data Brief. 2018, 21, 1664–1668. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  138. Zhang, L.; Yang, F.; Zhang, Y.D.; Zhu, Y. Road crack detection using deep convolutional neural network. In Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP); IEEE: New York, NY, USA, 2016. [Google Scholar]
  139. Żarski, M.; Wójcik, B.; Miszczak, J.A. KrakN: Transfer learning framework and dataset for infrastructure thin crack detection. SoftwareX 2021, 16, 100893. [Google Scholar] [CrossRef] [Scilit]
  140. Ye, X.W.; Jin, T.; Li, Z.X. Structural Crack Detection from Benchmark Data Sets Using Pruned Fully Convolutional Networks. J. Struct. Eng. 2021, 147, 04721008. [Google Scholar] [CrossRef] [Scilit]
  141. Zoubir, H.; Rguig, M.; Elaroussi, M. Crack recognition automation in concrete bridges using Deep Convolutional Neural Networks. In MATEC Web of Conferences; EDP Sciences: Les Ulis, France, 2021. [Google Scholar]
  142. Xu, H.; Wang, Y.; Li, Z.; Zhang, Q. Autonomous bridge crack detection using deep convolutional neural networks. In Proceedings of the 3rd International Conference on Computer Engineering, Information Science & Application Technology (ICCIA 2019); Atlantis Press: Dordrecht, The Netherlands, 2019. [Google Scholar]
  143. Hüthwohl, P.; Lu, R.; Brilakis, I. Multi-classifier for reinforced concrete bridge defects. Autom. Constr. 2019, 105, 102824. [Google Scholar] [CrossRef] [Scilit]
  144. Eisenbach, M.; Stricker, R.; Seichter, D.; Amende, K.; Debes, K.; Sesselmann, M.; Ebersbach, D.; Stöckert, U.; Gross, H.M. How to get pavement distress detection ready for deep learning? A systematic approach. In Proceedings of the 2017 International Joint Conference on Neural Networks (IJCNN); IEEE: New York, NY, USA, 2017. [Google Scholar]
  145. Mundt, M.; Majumdar, S.; Murali, S.; Panetsos, P.; Ramesh, V. Meta-learning convolutional neural architectures for multi-target concrete defect classification with the concrete defect bridge image dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019. [Google Scholar]
  146. Xu, H.; Su, X.; Wang, Y.; Cai, H.; Cui, K.; Chen, X. Automatic bridge crack detection using a convolutional neural network. Appl. Sci. 2019, 9, 2867. [Google Scholar] [CrossRef] [Scilit]
  147. Zhang, C.; Chang, C.; Jamshidi, M. Concrete bridge surface damage detection using a single-stage detector. Comput.-Aided Civ. Infrastruct. Eng. 2020, 35, 389–409. [Google Scholar] [CrossRef] [Scilit]
  148. McEnroe, P.; Wang, S.; Liyanage, M. A survey on the convergence of edge computing and AI for UAVs: Opportunities and challenges. IEEE Internet Things J. 2022, 9, 15435–15459. [Google Scholar] [CrossRef] [Scilit]
  149. Shi, Y.; Cui, L.; Qi, Z.; Meng, F.; Chen, Z. Automatic road crack detection using random structured forests. IEEE Trans. Intell. Transp. Syst. 2016, 17, 3434–3445. [Google Scholar] [CrossRef] [Scilit]
  150. Rao, Y.; Han, X.J.; Xiao, F.; Sun, S. Multiple disease detection of concrete structures based on deep learning. J. Build. Eng. 2021, 51, 1439–1445. [Google Scholar]
  151. Zou, Q.; Zhang, Z.; Li, Q.Q.; Qi, X.B.; Wang, Q.; Wang, S.Y. Deepcrack: Learning hierarchical convolutional features for crack detection. IEEE Trans. Image Process. 2018, 28, 1498–1512. [Google Scholar] [CrossRef] [Scilit]
  152. Yang, F.; Zhang, L.; Yu, S.; Prokhorov, D.; Mei, X.; Ling, H.B. Feature pyramid and hierarchical boosting network for pavement crack detection. IEEE Trans. Intell. Transp. Syst. 2019, 21, 1525–1535. [Google Scholar] [CrossRef] [Scilit]
  153. Mustafa, S.; Singh, P.; Srivastava, A.K. Evaluation of flexural cracks in a concrete girder railway bridge using displacement influence lines. Structures 2025, 80, 109758. [Google Scholar] [CrossRef] [Scilit]
  154. Song, Y.; Li, J.; Wang, H.; Zhang, L. Advances in crack dataset development and deep learning-based detection models. J. Build. Eng. 2025, 116, 114734. [Google Scholar] [CrossRef] [Scilit]
  155. Yu, C.X.; Wang, Z.H.; Li, Y.D.; Chen, H.Y. An improved U-Net model for concrete crack detection. Mach. Learn. Appl. 2022, 10, 100436. [Google Scholar] [CrossRef] [Scilit]
  156. Yuan, J.; Liu, Y.; Zhang, Q.; Wang, H. Bridge crack segmentation method based on parallel attention mechanism and multi-scale features fusion. IEEE Access 2023, 74, 6485. [Google Scholar] [CrossRef] [Scilit]
  157. Maryoosh, A.A.; Pashazadeh, S.; Salehpour, P. A Hybrid Learning Framework for Enhancing Bridge Damage Prediction. Appl. Syst. Innov. 2025, 8, 61. [Google Scholar] [CrossRef] [Scilit]
  158. Chen, B.; Fan, M.; Li, K.; Gao, Y.S.; Wang, Y.F.; Chen, Y.Q.; Yin, S.H.; Sun, J.X. The PFILSTM model: A crack recognition method based on pyramid features and memory mechanisms. Front. Mater. 2024, 10, 1347176. [Google Scholar] [CrossRef] [Scilit]
  159. Yang, J.; Huang, L.; Tong, K.; Tang, Q.; Li, H.; Cai, H.; Xin, J. A review on damage monitoring and identification methods for arch bridges. Buildings 2023, 13, 1975. [Google Scholar] [CrossRef] [Scilit]
  160. Jin, T.; Zhang, W.; Chen, C.; Chen, B.; ZhanG, Y.; Zhang, H. Deep-learning-and unmanned aerial vehicle-based structural crack detection in concrete. Sustainability 2023, 13, 3114. [Google Scholar] [CrossRef] [Scilit]
  161. Zhang, L.; Gong, L.L.; Wang, L.; Wang, Z.; Yan, S. A building crack detection UAV system based on deep learning and linear active disturbance rejection control algorithm. Drones 2025, 14, 2975. [Google Scholar] [CrossRef] [Scilit]
  162. Xiao, Y.M.; Lu, X.L.; Zhang, H.M. Perspective Correction and Deep Learning-Based Crack Detection for Concrete Structures. Fatigue Fract. Eng. M. 2025, 34, e70029. [Google Scholar] [CrossRef] [Scilit]
  163. Rashid, K.I.; Hussain, A.; Yang, C.H.; Huang, C.X. Hybrid semantic segmentation with broad context and attention encoded network for urban street scenario. Eng. Appl. Artif. Intell. 2026, 165, 113528. [Google Scholar]
  164. Zhou, Y.C.; Li, G.J.; Wei, W.; Wang, Y.M.; Jing, Q. A Rapid Crack Detection Technique Based on Attention for Intelligent M&O of Cross-Sea Bridge. China J. Highw. Transp. 2024, 38, 866–876. [Google Scholar] [CrossRef] [Scilit]
  165. Yadav, D.P.; Sharma, B.; Chauhan, S.; Ben Dhaou, I. Bridging convolutional neural networks and transformers for efficient crack detection in concrete building structures. Buildings 2024, 24, 4257. [Google Scholar] [CrossRef] [Scilit]
  166. Feng, J.H.; Li, Y.W.; Zhang, X.Q.; Wang, Q.L. Enhanced Crack Segmentation via Dual-Branch CNN-Transformer Architecture with Linear Perception and Multi-Scale Refinement. arXiv 2025, arXiv:2501.12345. [Google Scholar]
  167. Wang, D.; Shi, P.; An, D.; Gou, X. Evaluating multi-defect parameters in concrete via electrical signals: A three-stage strategy of ERT imaging, image segmentation, parameter quantification. Eng. Fract. Mech. 2026, 111999. [Google Scholar] [CrossRef] [Scilit]
  168. Fang, Y. Pixel-level bridge crack detection and 3D localization: An improved hybrid approach combining segmentation and 3D reconstruction. Structures 2025, 82, 110611. [Google Scholar] [CrossRef] [Scilit]
  169. Qi, Y.L.; Wang, Z.B.; Li, J.W.; Zhang, H.T. Crack detection and 3D visualization of crack distribution for UAV-based bridge inspection using efficient approaches. Structures 2025, 78, 109075. [Google Scholar] [CrossRef] [Scilit]
  170. Fan, Q.; Tan, Y.H.; Liu, Y.; Zhu, S.F. Rapid construction of large-scale bridge crack dataset using residual neural network. Meas. Sci. Technol. 2025, 36, 075011. [Google Scholar] [CrossRef] [Scilit]
  171. Wang, Y.M.; Guo, Y.F. Different Scale Crack Contour Extraction Techniques for Bridges Incorporating Improved YOLOv5 and GCNet. IEEE Access 2025, 13, 99547–99562. [Google Scholar] [CrossRef] [Scilit]
  172. Deng, J.H.; Lu, Y.; Lee, V.C.S. Concrete crack detection with handwriting script interferences using faster region-based convolutional neural network. Comput.-Aided Civ. Infrastruct. Eng. 2020, 35, 373–388. [Google Scholar] [CrossRef] [Scilit]
  173. Fu, H.X.; Li, W.H.; Meng, D.; Wang, Y.C. Bridge Crack Semantic Segmentation Based On Improved Deeplabv3+. J. Mar. Sci. Eng. 2021, 9, 671. [Google Scholar] [CrossRef] [Scilit]
  174. Luo, D.M.; Han, J.; Qiao, X.D.; Niu, D.T. Research on intelligent diagnosis of freeze-thaw damage of concrete based on improved YOLO v8s-Segformer integrated model. Eng. Struct. 2025, 343, 121177. [Google Scholar] [CrossRef] [Scilit]
  175. Xu, Z.; Wang, Y.; Hao, X.; Fan, J. Crack detection of bridge concrete components based on large-scene images using an unmanned aerial vehicle. Sensors 2023, 23, 6271. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  176. Hou, S.T.; Dong, B.; Wang, H.C.; Wu, G. Inspection of surface defects on stay cables using a robot and transfer learning. Autom. Constr. 2020, 119, 103382. [Google Scholar] [CrossRef] [Scilit]
  177. Wang, B.Q.; Feng, D.M.; Fan, Z.C.; Shi, H.Y. Automated surface defect detection of stay cables via UAV route planning and deep learning model. Autom. Constr. 2025, 180, 106579. [Google Scholar] [CrossRef] [Scilit]
  178. Kim, J.; Kim, J.; Kim, T. Deep learning-based framework for time-dependent reliability analysis of a cable-stayed bridge with corroded PSC box girders under gravity loads. Sci. Rep. 2025, 15, 42448. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  179. Yue, Z.X.; Ding, Y.L.; Zhao, H.W.; Wang, Z.W. Case Study of Deep Learning Model of Temperature-Induced Deflection of a Cable-Stayed Bridge Driven by Data Knowledge. Symmetry 2021, 13, 2293. [Google Scholar] [CrossRef] [Scilit]
  180. Panigati, T.; Zini, M.; Striccoli, D.; Giordano, P.F.; Tonelli, D.; Limongelli, M.P.; Zonta, D. Drone-based bridge inspections: Current practices and future directions. Autom. Constr. 2025, 173, 106101. [Google Scholar] [CrossRef] [Scilit]
  181. Yazdanpanah, O.; Park, M.; Chang, M.; Chae, Y. Mastering seismic time series response predictions using an attention-Mamba transformer model for bridge bearings and piers across varied testing conditions. Sci. Rep. 2024, 14, 106101. [Google Scholar] [CrossRef] [Scilit]
  182. Xu, F.Y.; Kalantari, M.; Li, B.J.; Wang, X.S. Nondestructive Testing of Bridge Stay Cable Surface Defects Based on Computer Vision. Comput. Mater. Contin. 2023, 75, 2209–2226. [Google Scholar] [CrossRef] [Scilit]
  183. Wu, Y.J.; Li, S.Q.; Li, J.Q.; Yu, Y.P.; Li, J.C.; Li, Y.C. Deep Learning in Crack Detection: A Comprehensive Scientometric Review. Comput. Sci. Rev. 2025, 4, 100144. [Google Scholar] [CrossRef] [Scilit]
  184. Wang, L.; Tang, C. Effective small crack detection based on tunnel crack characteristics and an anchor-free convolutional neural network. Sci. Rep. 2024, 14, 10355. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  185. Kim, B.; Devi, M.S.; Natarajan, Y.; Sri Preethaa, K.R.; Yazhini, C.S.; Park, J.Y.; Park, C.J.; Yi, C.Y. Gradient transformer Self-Attention U-Net for enhanced crack detection in concrete bridges. Sci. Rep. 2025, 15, 37087. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  186. Zhang, T.; Qin, L.; Zou, Q.; Zhang, L.; Wang, R.; Zhang, H. Crackscopenet: A lightweight neural network for rapid crack detection on resource-constrained drone platforms. Drones 2024, 8, 417. [Google Scholar] [CrossRef] [Scilit]
  187. Meng, K.; Che, H.Y.; Liu, Z.Y.; Lin, H.; Yu, J.T.; Zhang, M.Y. Efficient CrackUNet: Hierarchical spatial-channel attention with multi-scale fusion for pavement and bridge crack segmentation. Phys. Scr. 2025, 100, 105011. [Google Scholar] [CrossRef] [Scilit]
  188. Cao, X.P.; Huang, Q.D.; Cao, L.L.; Zhang, C.; Qiu, J.F.; Niu, F.Z. Attention-enhanced automated detection system for small cracks in concrete bridges. Eng. Struct. 2026, 351, 122010. [Google Scholar] [CrossRef] [Scilit]
  189. Karimi, N.; Mishra, M.; Lourenço, P.B. Automated Surface Crack Detection in Historical Constructions with Various Materials Using Deep Learning-Based YOLO Network. Int. J. Archit. Herit. 2025, 19, 581–597. [Google Scholar] [CrossRef] [Scilit]
  190. Lin, C.; Tian, D.; Duan, X.; Zhou, J. TransCrack: Revisiting fine-grained road crack detection with a transformer design. Philos. Trans. R. Soc. A Math. Phys. Eng. Sci. 2023; 381, pp. 179–187 20220172. [Google Scholar]
  191. Song, F.; Sun, Y.; Yuan, G.X. Autonomous identification of bridge concrete cracks using unmanned aircraft images and improved lightweight deep convolutional networks. Math. Probl. Eng. 2024, 2024, 7857012. [Google Scholar] [CrossRef] [Scilit]
  192. Ni, Y.H.; Mao, J.X.; Wang, H.; Xi, Z.; Chen, Z.Y. Surface Damage Detection and Localization for Bridge Visual Inspection Based on Deep Learning and 3D Reconstruction. Math. Probl. Eng. 2024, 2024, 9988793. [Google Scholar] [CrossRef] [Scilit]
  193. Ozturk, H.Y.; Zappa, E. Automated Crack Width Measurement in 3D Models: A Photogrammetric Approach with Image Selection. Information 2025, 16, 448. [Google Scholar] [CrossRef] [Scilit]
  194. Golding, V.P.; Gharineiat, Z.; Munawar, H.S.; Ullah, F. Crack detection in concrete structures using deep learning. Sustainability 2022, 14, 8117. [Google Scholar] [CrossRef] [Scilit]
  195. Lu, X.C.; Li, Q.Q.; Zhang, L. Deep learning-based method for detection and feature quantification of microscopic cracks on the surface of concrete dams. Measurement 2025, 240, 115587. [Google Scholar] [CrossRef] [Scilit]
  196. Chen, D.H.; Cui, H.; Li, Z.; Xu, S.Z.; Zhang, Y. Indirect Identification and Analysis of Bridge Damage Using Vehicle-Bridge Coupled Vibration and Deep Learning. J. Perform. Constr. Facil. 2024, 38, 04024016. [Google Scholar] [CrossRef] [Scilit]
  197. Alfaro, M.C.; Vidal, R.S.; Delgadillo, R.M.; Moya, L.; Casas, J.R. Structural damage detection using an unmanned aerial vehicle-based 3D model and deep learning on a reinforced concrete arch bridge. Infrastructure 2025, 10, 33. [Google Scholar] [CrossRef] [Scilit]
  198. Wang, Q.P. Bridge Health Status Detection Based On Deep Learning: Integrating Attention Mechanism and Hybrid Feature Extraction Technology. J. Cases Inf. Technol. 2025, 27, 1–24. [Google Scholar] [CrossRef] [Scilit]
  199. Mishra, M.; Lourenço, P.B.; Ramana, G.V. Structural health monitoring of civil engineering structures by using the internet of things: A review. J. Build. Eng. 2022, 48, 103954. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.